Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html.entities

import html.entities

html lifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelled html.entities.* needs import html.entities.

Character reference handling for HTML: turning text into something safe to drop into a document (encode()), and turning a document’s &-style references back into the characters they stand for (decode()).

Both directions follow the WHATWG HTML Living Standard. Decoding understands all 2231 named references, decimal references (©), hexadecimal references (©), the Windows-1252 substitutions the standard mandates for code points in the C1 range, and the awkward legacy rule that lets a handful of names work without their trailing semicolon.

%> import html
%> html.decode('café — ☃')
'café: ☃'
%> html.encode('<a href="x">Tom & Jerry</a>')
'&lt;a href=&quot;x&quot;&gt;Tom &amp; Jerry&lt;/a&gt;'

Constants

REPLACEMENT_CHARACTER

html.entities.REPLACEMENT_CHARACTER = '�'

NO_BREAK_SPACE

html.entities.NO_BREAK_SPACE = ' '

Functions

is_ascii_alphanumeric()

html.entities.is_ascii_alphanumeric(c: string) -> bool

Returns true when c is one of 0-9, A-Z or a-z.

Named references are made up entirely of these, so this is the test that bounds how far reference matching will scan.

Parameters

  • c (string)

Returns bool

is_ascii_digit()

html.entities.is_ascii_digit(c: string) -> bool

Returns true when c is an ASCII decimal digit.

Parameters

  • c (string)

Returns bool

is_ascii_hex_digit()

html.entities.is_ascii_hex_digit(c: string) -> bool

Returns true when c is an ASCII hexadecimal digit, in either case.

Parameters

  • c (string)

Returns bool

code_point_to_string()

html.entities.code_point_to_string(code: number) -> string

Applies the standard’s “numeric character reference end state” rules to a raw code point and returns the string it should decode to.

Null, out-of-range and surrogate code points become U+FFFD. Code points in the C1 range are remapped through Windows-1252. Noncharacters and control characters are parse errors but decode to themselves, which is what browsers do and what round-tripping malformed documents depends on.

Parameters

  • code (number)

Returns string

consume_reference()

html.entities.consume_reference(chars: list, start: number, in_attribute: bool) -> dict

Consumes a character reference from chars beginning at the ampersand at index start.

This is the standard’s “character reference state” as a single call, shared by decode() and by the tokenizer so both agree on every edge case. It always succeeds: when the text at start is not a valid reference the ampersand is returned as an ordinary character and next points just past it.

in_attribute selects the attribute-value flavour of the algorithm. Inside an attribute a semicolon-less legacy reference followed by = or by an alphanumeric is left alone, so that a query string like ?x=1&notit=2 keeps its literal &not rather than turning into ¬it.

The returned dictionary has:

  • text: the decoded text (never nil, never empty)
  • next: index of the first character not consumed
  • error: nil, or a short description of the parse error the standard raises for this input; the text is still produced

Parameters

  • chars (list)
  • start (number)
  • in_attribute (bool)

Returns dict

encode()

html.encode(text: string, options: ?dict) -> string

Escapes text so that it can be embedded in an HTML document without being reinterpreted as markup.

By default this escapes &, <, >, " and ', which is safe for both element content and quoted attribute values. Everything else is passed through unchanged, so the result stays readable and stays valid UTF-8.

options is an optional dictionary:

  • quotes (bool, default true): escape " and '. Turn this off only when the result is going into element content, never into an attribute.
  • non_ascii (bool, default false): also escape every character above U+007F. Useful when the output has to survive a transport that is not UTF-8 clean.
  • named (bool, default true): when a character being escaped has a named reference, use it. With false, non-syntactic characters use hexadecimal numeric references instead. The five syntax characters above always use their fixed forms.
%> import html
%> html.encode('5 > 3 & "quoted"')
'5 &gt; 3 &amp; &quot;quoted&quot;'
%> html.encode("it's fine", { quotes: false })
"it's fine"
%> html.encode('café', { non_ascii: true })
'caf&eacute;'
%> html.encode('café', { non_ascii: true, named: false })
'caf&#xE9;'

Surrogate code points cannot appear in a Zuri string, so no unpaired-surrogate handling is needed here.

Parameters

  • text (string)
  • options (?dict)

Returns string

decode()

html.decode(text: string, in_attribute: ?bool) -> string

Decodes every character reference in text and returns the result.

Named, decimal and hexadecimal references are all understood. Text that merely looks like a reference is left exactly as it was, so decode() is safe to run over arbitrary content: 'a & b' stays 'a & b' and '&nosuchthing;' stays '&nosuchthing;'.

Pass true for in_attribute when the text came from an attribute value. That switches on the standard’s attribute rule, under which a legacy reference written without its semicolon is not decoded if the next character is = or alphanumeric. It is what keeps ?a=1&copy=2 from becoming ?a=1©=2. The default is false, the element-content behaviour.

%> import html
%> html.decode('&lt;b&gt; &amp; &#169; &#xA9; &nbsp;')
'<b> & © ©  '
%> html.decode('?a=1&copy=2')
'?a=1©=2'
%> html.decode('?a=1&copy=2', true)
'?a=1&copy=2'

Parameters

  • text (string)
  • in_attribute (?bool)

Returns string

escape_text()

html.escape_text(text: string) -> string

Escapes text the way the HTML fragment serialization algorithm requires for element content.

Exactly four characters change: & becomes &amp;, U+00A0 becomes &nbsp;, < becomes &lt; and > becomes &gt;. Quotes are left alone because they carry no meaning in element content, and the no-break space is escaped because it is otherwise indistinguishable from a plain space in a source listing.

This is what outer_html() and the serializers use. For escaping text you are about to hand to something other than this module’s own serializer, prefer encode(), which is stricter.

Parameters

  • text (string)

Returns string

escape_attribute()

html.escape_attribute(value: string) -> string

Escapes value the way the HTML fragment serialization algorithm requires for a double-quoted attribute value.

Three characters change: & becomes &amp;, U+00A0 becomes &nbsp; and " becomes &quot;. < and > are deliberately left alone; they are ordinary characters inside a quoted attribute and escaping them would needlessly change the value’s serialized form.

Parameters

  • value (string)

Returns string


2026, Richard Ore and The Zuri Contributors