Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

mail.encoding

import mail

mail exposes this as mail.encoding, so import mail is enough and the names are called as mail.encoding.*. import mail.encoding reaches the same definitions directly.

The encodings a message uses to get eight bit data through a seven bit pipe.

Mail headers are ASCII and mail bodies were historically ASCII too, so everything else has to be spelled out in it. Bodies use base64 or quoted-printable; headers use the encoded words of RFC 2047 and, for parameters, the continuations of RFC 2231. This module implements all four in both directions, plus the character sets a message is likely to name.

import mail.encoding

echo encoding.encode_word('Grüße', 'utf-8', 'Q')
echo encoding.decode_words('=?utf-8?Q?Gr=C3=BC=C3=9Fe?=')
=?utf-8?Q?Gr=C3=BC=C3=9Fe?=
Grüße

Constants

TRANSFER_ENCODINGS

mail.encoding.TRANSFER_ENCODINGS = [...]

Functions

is_ascii()

mail.encoding.is_ascii(data) -> bool

Whether every byte is plain ASCII, which decides whether a header or a body needs encoding at all.

Parameters

  • data (string|bytes)

Returns bool

encode_base64()

mail.encoding.encode_base64(data, width: ?number) -> string

Encodes data as base64, wrapped into lines.

Parameters

  • data (string|bytes)
  • width (?number) — Characters per line, 76 by default. Pass 0 for one unbroken line, which is what a header needs.

Returns string

decode_base64()

mail.encoding.decode_base64(text) -> bytes

Decodes base64, ignoring the line breaks and stray whitespace a message carries it with.

Parameters

  • text (string|bytes)

Returns bytes

Raises MessageError if what is left is not base64.

encode_quoted_printable()

mail.encoding.encode_quoted_printable(data) -> string

Encodes data as quoted-printable.

Bytes that can stand for themselves do; the rest become =XX. Lines are kept within 76 characters with the soft breaks the encoding defines, and whitespace at the end of a line is encoded so that a transport which trims it cannot change what was sent.

import mail.encoding

echo encoding.encode_quoted_printable('café = good')
caf=C3=A9 =3D good

Parameters

  • data (string|bytes)

Returns string

decode_quoted_printable()

mail.encoding.decode_quoted_printable(text) -> bytes

Decodes quoted-printable.

Tolerant of what real messages contain: a soft break with a bare newline rather than a carriage return and newline, and a lone = that is not followed by two hex digits, which is passed through rather than treated as a failure.

Parameters

  • text (string|bytes)

Returns bytes

decode_text()

mail.encoding.decode_text(raw, charset: ?string) -> string

Turns bytes into text, reading them as the character set names.

The sets a message is likely to name are understood outright: UTF-8, US-ASCII, the ISO 8859 Latin sets 1 and 15, windows-1252, and UTF-16 in either byte order. A set this does not know is read as UTF-8, which leaves the ASCII in it intact rather than discarding the part that would have been readable.

Parameters

  • raw (string|bytes)
  • charset (?string) — utf-8 when not given.

Returns string

encode_word()

mail.encoding.encode_word(text: string, charset: ?string, method: ?string) -> string

Encodes text as one or more RFC 2047 encoded words, so that a header can carry characters a header is not allowed to contain.

Text that is already ASCII is returned untouched, since a word that encodes nothing only makes the header harder to read.

Long text becomes several words separated by a space, because one word may not exceed 75 characters. Words are split on character boundaries, never in the middle of one.

Parameters

  • text (string)
  • charset (?string) — utf-8 when not given.
  • method (?string) — B or Q. Chosen by what the text looks like when not given: Q when most of it is ASCII, B otherwise.

Returns string

decode_words()

mail.encoding.decode_words(text: string) -> string

Reads the encoded words out of a header value and puts back the text they stand for.

Whitespace between two adjacent encoded words is dropped, which is what lets a long subject be split across several of them without a space appearing where none was written.

Anything that is not a well-formed encoded word is left exactly as it is, including text that merely begins with =?.

import mail.encoding

echo encoding.decode_words('=?utf-8?B?SGVsbG8=?= =?utf-8?B?IHdvcmxk?=')
Hello world

Parameters

  • text (string)

Returns string

encode_parameter()

mail.encoding.encode_parameter(name: string, value: string) -> list

Writes one parameter of a header, choosing the form its value needs.

A short ASCII value is quoted and written as it is. A value with characters a header cannot carry is written in the extended form of RFC 2231, which names the character set. A long value is split across numbered continuations, because a parameter may not run past the end of a line.

import mail.encoding

echo '; '.join(encoding.encode_parameter('filename', 'report.pdf'))
echo '; '.join(encoding.encode_parameter('filename', 'résumé.pdf'))
filename="report.pdf"
filename*=utf-8''r%C3%A9sum%C3%A9.pdf

Parameters

  • name (string)
  • value (string)

Returns list — of string, each key=value ready to be joined with '; '

decode_parameters()

mail.encoding.decode_parameters(raw: dict) -> dict

Puts the parameters of a header back together.

Takes them as they were written, continuations and extended forms and all, and hands back one value per parameter with the pieces joined and the character set applied.

Parameters

  • raw (dict) — Parameters by the name they were written under, including the *0* and * suffixes.

Returns dict

encode_body()

mail.encoding.encode_body(data, encoding: string) -> string

Encodes a part’s body the way its Content-Transfer-Encoding says.

7bit, 8bit and binary describe the data rather than change it, so they hand it back as it is.

Parameters

  • data (string|bytes)
  • encoding (string)

Returns string

Raises MessageError if encoding is not one this understands.

decode_body()

mail.encoding.decode_body(data, encoding: ?string) -> bytes

Decodes a part’s body the way its Content-Transfer-Encoding says.

An encoding this does not understand is treated as 8bit and handed back untouched, because a part whose encoding cannot be read is still better delivered than discarded.

Parameters

  • data (string|bytes)
  • encoding (?string) — 7bit when not given.

Returns bytes

guess_encoding()

mail.encoding.guess_encoding(data) -> string

Picks the transfer encoding a body should use.

ASCII text that fits inside the line limit needs none. Text that is mostly ASCII is cheaper as quoted-printable, which leaves it readable; anything else is smaller as base64.

Parameters

  • data (string|bytes)

Returns string — one of 7bit, quoted-printable or base64


2026, Richard Ore and Zuri contributors