html.entities
import html.entities
htmllifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelledhtml.entities.*needsimport html.entities.
Character reference handling for HTML: turning text into something safe
to drop into a document (encode()), and turning a document’s
&-style references back into the characters they stand for
(decode()).
Both directions follow the WHATWG HTML Living Standard. Decoding
understands all 2231 named references, decimal references (©),
hexadecimal references (©), the Windows-1252 substitutions the
standard mandates for code points in the C1 range, and the awkward
legacy rule that lets a handful of names work without their trailing
semicolon.
%> import html
%> html.decode('café — ☃')
'café: ☃'
%> html.encode('<a href="x">Tom & Jerry</a>')
'<a href="x">Tom & Jerry</a>'
Constants
REPLACEMENT_CHARACTER
html.entities.REPLACEMENT_CHARACTER = '�'
NO_BREAK_SPACE
html.entities.NO_BREAK_SPACE = ' '
Functions
is_ascii_alphanumeric()
html.entities.is_ascii_alphanumeric(c: string) -> bool
Returns true when c is one of 0-9, A-Z or a-z.
Named references are made up entirely of these, so this is the test that bounds how far reference matching will scan.
Parameters
c(string)
Returns bool
is_ascii_digit()
html.entities.is_ascii_digit(c: string) -> bool
Returns true when c is an ASCII decimal digit.
Parameters
c(string)
Returns bool
is_ascii_hex_digit()
html.entities.is_ascii_hex_digit(c: string) -> bool
Returns true when c is an ASCII hexadecimal digit, in either case.
Parameters
c(string)
Returns bool
code_point_to_string()
html.entities.code_point_to_string(code: number) -> string
Applies the standard’s “numeric character reference end state” rules to a raw code point and returns the string it should decode to.
Null, out-of-range and surrogate code points become U+FFFD. Code points in the C1 range are remapped through Windows-1252. Noncharacters and control characters are parse errors but decode to themselves, which is what browsers do and what round-tripping malformed documents depends on.
Parameters
code(number)
Returns string
consume_reference()
html.entities.consume_reference(chars: list, start: number, in_attribute: bool) -> dict
Consumes a character reference from chars beginning at the ampersand at index start.
This is the standard’s “character reference state” as a single call,
shared by decode() and by the tokenizer so both agree on every edge
case. It always succeeds: when the text at start is not a valid
reference the ampersand is returned as an ordinary character and next
points just past it.
in_attribute selects the attribute-value flavour of the algorithm.
Inside an attribute a semicolon-less legacy reference followed by = or
by an alphanumeric is left alone, so that a query string like
?x=1¬it=2 keeps its literal ¬ rather than turning into ¬it.
The returned dictionary has:
text: the decoded text (nevernil, never empty)next: index of the first character not consumederror:nil, or a short description of the parse error the standard raises for this input; the text is still produced
Parameters
chars(list)start(number)in_attribute(bool)
Returns dict
encode()
html.encode(text: string, options: ?dict) -> string
Escapes text so that it can be embedded in an HTML document without being reinterpreted as markup.
By default this escapes &, <, >, " and ', which is safe for
both element content and quoted attribute values. Everything else is
passed through unchanged, so the result stays readable and stays valid
UTF-8.
options is an optional dictionary:
quotes(bool, defaulttrue): escape"and'. Turn this off only when the result is going into element content, never into an attribute.non_ascii(bool, defaultfalse): also escape every character above U+007F. Useful when the output has to survive a transport that is not UTF-8 clean.named(bool, defaulttrue): when a character being escaped has a named reference, use it. Withfalse, non-syntactic characters use hexadecimal numeric references instead. The five syntax characters above always use their fixed forms.
%> import html
%> html.encode('5 > 3 & "quoted"')
'5 > 3 & "quoted"'
%> html.encode("it's fine", { quotes: false })
"it's fine"
%> html.encode('café', { non_ascii: true })
'café'
%> html.encode('café', { non_ascii: true, named: false })
'café'
Surrogate code points cannot appear in a Zuri string, so no unpaired-surrogate handling is needed here.
Parameters
text(string)options(?dict)
Returns string
decode()
html.decode(text: string, in_attribute: ?bool) -> string
Decodes every character reference in text and returns the result.
Named, decimal and hexadecimal references are all understood. Text that
merely looks like a reference is left exactly as it was, so decode()
is safe to run over arbitrary content: 'a & b' stays 'a & b' and
'&nosuchthing;' stays '&nosuchthing;'.
Pass true for in_attribute when the text came from an attribute
value. That switches on the standard’s attribute rule, under which a
legacy reference written without its semicolon is not decoded if the
next character is = or alphanumeric. It is what keeps ?a=1©=2
from becoming ?a=1©=2. The default is false, the element-content
behaviour.
%> import html
%> html.decode('<b> & © © ')
'<b> & © © '
%> html.decode('?a=1©=2')
'?a=1©=2'
%> html.decode('?a=1©=2', true)
'?a=1©=2'
Parameters
text(string)in_attribute(?bool)
Returns string
escape_text()
html.escape_text(text: string) -> string
Escapes text the way the HTML fragment serialization algorithm requires for element content.
Exactly four characters change: & becomes &, U+00A0 becomes
, < becomes < and > becomes >. Quotes are left
alone because they carry no meaning in element content, and the no-break
space is escaped because it is otherwise indistinguishable from a plain
space in a source listing.
This is what outer_html() and the serializers use. For escaping text
you are about to hand to something other than this module’s own
serializer, prefer encode(), which is stricter.
Parameters
text(string)
Returns string
escape_attribute()
html.escape_attribute(value: string) -> string
Escapes value the way the HTML fragment serialization algorithm requires for a double-quoted attribute value.
Three characters change: & becomes &, U+00A0 becomes
and " becomes ". < and > are deliberately left alone; they
are ordinary characters inside a quoted attribute and escaping them
would needlessly change the value’s serialized form.
Parameters
value(string)
Returns string
2026, Richard Ore and The Zuri Contributors