html
import html
A complete HTML5 parser, DOM and serializer, following the WHATWG HTML Living Standard.
Parse a page, walk it, query it with CSS selectors, change it, and write
it back out. Nothing here fetches anything over the network and nothing
here executes a script; <script> content is inert text and always will
be.
Quick start
import html
var doc = html.parse('
<article>
<h1 id="title">Zuri</h1>
<ul class="links">
<li><a href="/docs">Docs</a>
<li><a href="/spec">Spec</a>
</ul>
</article>
')
echo doc.get_element_by_id('title').text_content()
for link in doc.query_selector_all('ul.links a') {
echo '${link.text_content()} -> ${link.get_attribute("href")}'
}
Notice that neither <li> is closed and there is no <html>, <head>
or <body> anywhere. That is fine: the standard defines a tree for
every possible input, and this module builds it. Parsing never raises on
malformed markup.
Finding things
| Method | Finds |
|---|---|
get_element_by_id(id) | the first element with that id |
get_elements_by_class_name(c) | every element with those classes |
get_elements_by_tag_name(t) | every element with that tag |
query_selector(sel) | the first CSS match |
query_selector_all(sel) | every CSS match |
element.matches(sel) | whether this element matches |
element.closest(sel) | this element or its nearest matching ancestor |
The selector syntax is a deliberately bounded subset: type, *, #id,
.class, [attr], [attr="value"], the combinators >, + and ~,
descendant combination by whitespace, :first-child, :last-child,
:nth-child(), and comma-separated lists of all of those. Anything
outside it raises a SelectorError rather than silently matching
nothing.
Changing things
import html
var doc = html.parse('<div class="box"><p>old</p></div>')
var box = doc.query_selector('.box')
box.set_attribute('data-seen', 'yes')
box.class_list().add('active')
box.set_inner_html('<p>new</p>')
echo box.outer_html()
# <div class="box active" data-seen="yes"><p>new</p></div>
Zuri has no property setters, so every writable DOM property is a pair
of methods: text_content() reads and set_text_content() writes, and
likewise for inner_html and outer_html.
Writing it back out
outer_html() round-trips a node exactly. When you want the output
shaped for a particular audience, minify() strips what a browser does
not need and format() indents for a human:
A string is parsed as a whole document, so <html>, <head> and
<body> show up in the output the way a browser would produce them.
Pass { fragment: true } when the input is a fragment and you want just
that fragment back:
%> import html
%> html.minify('<p> a b </p>\\n<p>c</p>', { fragment: true })
'<p>a b</p><p>c</p>'
%> echo html.format('<ul><li>a<li>b</ul>', { fragment: true })
<ul>
<li>a</li>
<li>b</li>
</ul>
Escaping
encode() makes text safe to embed and decode() resolves character
references. Decoding knows all 2231 named references, not a shortlist:
%> import html
%> html.encode('<b>Tom & Jerry</b>')
'<b>Tom & Jerry</b>'
%> html.decode('café — ☃')
'café: ☃'
Encoding
Input is UTF-8 text that has already been decoded, which is what a Zuri
string always is. There is no charset sniffing, no byte order mark
handling and no <meta charset> processing: if you are reading bytes
off a socket, decode them before they get here.
Parse errors
HTML has no fatal errors, so a parse always succeeds. Every conformance problem the standard defines is collected instead:
%> import html
%> var doc = html.parse('<p><b>x</p>')
%> for e in doc.errors { echo e.to_string() }
missing-doctype at 1:1
unexpected-end-tag at 1:8
Each entry carries the standard’s own error name, so you can look it up, plus the line and column where it was noticed.
The html API
Every public name in html, wherever it is declared. Each links to the
page that documents it.
| Name | Kind | Summary |
|---|---|---|
html.ClassList | class | A live view of one element’s class attribute. |
html.Comment | class | An HTML comment. |
html.Document | class | A whole parsed document. |
html.DocumentFragment | class | A parentless container for a run of nodes. |
html.DocumentType | class | The <!DOCTYPE ...> node at the top of a document. |
html.Element | class | An element in the tree. |
html.HTML_NAMESPACE | constant | Namespace URI of ordinary HTML elements. |
html.HierarchyError | class | Raised when a tree mutation would produce a structure that cannot exist, such as inserting a node before… |
html.MATHML_NAMESPACE | constant | Namespace URI of elements inside a <math> subtree. |
html.NODE_COMMENT | constant | Node type of a Comment. |
html.NODE_DOCUMENT | constant | Node type of a Document. |
html.NODE_DOCUMENT_FRAGMENT | constant | Node type of a DocumentFragment, including the fragment that holds a <template> element’s content. |
html.NODE_DOCUMENT_TYPE | constant | Node type of a DocumentType (the <!DOCTYPE ...> node). |
html.NODE_ELEMENT | constant | Node type of an Element. |
html.NODE_TEXT | constant | Node type of a Text node. |
html.Node | class | The base class every node in a parsed document inherits from. |
html.ParseError | class | A parse error raised while tokenizing or building the tree. |
html.RAW_TEXT_ELEMENTS | constant | Elements whose text children are markup, not content: their text is written out byte for byte and never… |
html.SVG_NAMESPACE | constant | Namespace URI of elements inside an <svg> subtree. |
html.Selector | class | A compiled selector list: h1, h2 is one of these holding two ComplexSelector instances. |
html.SelectorError | class | Raised when a selector cannot be parsed, or uses syntax this module deliberately does not support. |
html.Text | class | A run of character data in the tree. |
html.Token | class | One token from the tokenizer. |
html.Tokenizer | class | Turns markup into tokens, one call to next_token() at a time. |
html.TreeBuilder | class | Builds a document tree from a token stream. |
html.VOID_ELEMENTS | constant | The HTML elements that are written without a closing tag and can hold no content. |
html.XLINK_NAMESPACE | constant | Namespace URI used by the xlink: attribute prefix in SVG. |
html.XMLNS_NAMESPACE | constant | Namespace URI used by the xmlns and xmlns:xlink attributes. |
html.XML_NAMESPACE | constant | Namespace URI used by the xml: attribute prefix. |
html.compile | function | Compiles source into a reusable Selector. |
html.decode | function | Decodes every character reference in text and returns the result. |
html.elements.FOREIGN_ATTRIBUTES | constant | Attributes in foreign content that belong to a namespace, mapping the lowercase name the tokenizer produced… |
html.elements.FOREIGN_BREAKOUT_TAGS | constant | Start tags that are always a mistake inside foreign content and that break out of it, closing SVG or MathML… |
html.elements.FORMATTING_ELEMENTS | constant | The formatting elements: the ones the adoption agency algorithm reopens across a badly nested boundary, so… |
html.elements.HTML4_TRANSITIONAL_PREFIXES | constant | Public identifier prefixes whose mode depends on whether the doctype also carries a system identifier: quirks… |
html.elements.IMPLIED_END_TAGS | constant | Elements whose end tag is implied by the start of a sibling, so that <li>a<li>b produces two list items… |
html.elements.LIMITED_QUIRKS_PUBLIC_PREFIXES | constant | Public identifier prefixes that always mean limited-quirks mode. |
html.elements.MATHML_ATTRIBUTES | constant | The one MathML attribute whose case the tokenizer destroys. |
html.elements.MATHML_TEXT_INTEGRATION_POINTS | constant | The MathML elements whose children are parsed as HTML rather than as MathML. |
html.elements.QUIRKS_PUBLIC_EXACT | constant | The two public identifiers that put a document in quirks mode on an exact match rather than a prefix match,… |
html.elements.QUIRKS_PUBLIC_PREFIXES | constant | Public identifier prefixes that put a document in quirks mode, in lowercase for case-insensitive comparison. |
html.elements.QUIRKS_SYSTEM_ID | constant | The system identifier that alone puts a document in quirks mode, in lowercase. |
html.elements.SCOPE_HTML | constant | The HTML elements that make up the “in scope” barrier every scope check shares. |
html.elements.SPECIAL_HTML | constant | The HTML elements the standard calls “special”: the ones a list item or definition term stops searching past,… |
html.elements.SPECIAL_MATHML | constant | The MathML elements that count as special. |
html.elements.SPECIAL_SVG | constant | The SVG elements that count as special. |
html.elements.SVG_ATTRIBUTES | constant | SVG attribute names the tokenizer lowercased and that have to be put back, keyed by the lowercase form. |
html.elements.SVG_HTML_INTEGRATION_POINTS | constant | The SVG elements whose children are parsed as HTML rather than as SVG. |
html.elements.SVG_TAG_NAMES | constant | SVG tag names the tokenizer lowercased and that have to be put back, keyed by the lowercase form. |
html.elements.THOROUGH_IMPLIED_END_TAGS | constant | Everything in IMPLIED_END_TAGS plus the elements only closed when the standard says to generate implied end… |
html.elements.adjust_mathml_attributes | function | Repairs the case of MathML attribute names in attributes and returns a new dictionary. |
html.elements.adjust_svg_attributes | function | Repairs the case of SVG attribute names in attributes and returns a new dictionary. |
html.elements.adjust_svg_tag_name | function | The correctly cased SVG tag name for the lowercase name the tokenizer produced, or name itself when it… |
html.elements.doctype_mode | function | Which quirks mode a doctype puts a document in. |
html.elements.is_html_integration_point | function | True when element is an HTML integration point. |
html.elements.is_mathml_text_integration_point | function | True when element is a MathML text integration point, meaning its children are parsed as HTML. |
html.elements.is_special | function | True when the element name in namespace namespace is one of the standard’s special elements. |
html.encode | function | Escapes text so that it can be embedded in an HTML document without being reinterpreted as markup. |
html.entities.NO_BREAK_SPACE | constant | |
html.entities.REPLACEMENT_CHARACTER | constant | |
html.entities.code_point_to_string | function | Applies the standard’s “numeric character reference end state” rules to a raw code point and returns the… |
html.entities.consume_reference | function | Consumes a character reference from chars beginning at the ampersand at index start. |
html.entities.is_ascii_alphanumeric | function | Returns true when c is one of 0-9, A-Z or a-z. |
html.entities.is_ascii_digit | function | Returns true when c is an ASCII decimal digit. |
html.entities.is_ascii_hex_digit | function | Returns true when c is an ASCII hexadecimal digit, in either case. |
html.escape_attribute | function | Escapes value the way the HTML fragment serialization algorithm requires for a double-quoted attribute… |
html.escape_text | function | Escapes text the way the HTML fragment serialization algorithm requires for element content. |
html.format | function | Renders source as indented, readable HTML. |
html.minify | function | Removes the markup a browser does not need and returns the result. |
html.parse | function | Parses source as a complete HTML document. |
html.parse_file | function | Reads the file at path and parses it as HTML. |
html.parse_fragment | function | Parses source as a fragment, as though it had been written inside context. |
html.parser.FormattingEntry | class | One entry in the list of active formatting elements. |
html.selector.AttributeTest | class | One name/=value test from a [...] block. |
html.selector.ComplexSelector | class | One selector from a comma-separated list: a chain of compounds joined by combinators, such as ul > li + li. |
html.selector.CompoundSelector | class | A run of simple selectors with nothing between them, such as div#main.active[data-x]:first-child. |
html.selector.PseudoTest | class | One :pseudo or :pseudo(...) test. |
html.serialize.INLINE_ELEMENTS | constant | Elements that flow inside a line of text rather than starting a new block. |
html.serialize.PRESERVE_WHITESPACE | constant | Elements whose text content is significant to the last character. |
html.tokenize | function | Tokenizes source and returns every token, ending with the eof token. |
html.tokenizer.RAWTEXT_ELEMENTS | constant | Tag names whose content is text with neither markup nor character references. |
html.tokenizer.RCDATA_ELEMENTS | constant | Tag names whose content is text with character references but no markup. |
Submodules
| Module | Reached as | Summary |
|---|---|---|
html.elements | html.elements.* | The element tables the tree construction algorithm consults on almost every token: which elements are… |
html.entities | html.entities.* | Character reference handling for HTML: turning text into something safe to drop into a document (encode()),… |
html.namespaces | html.* | The five namespace URIs the HTML parser deals in. |
html.node | html.* | The document tree that html.parse() produces, and everything you can do with it: walking it, querying it,… |
html.parser | html.parser.* | The HTML tree construction stage: the half of the parser that takes the tokenizer’s stream and decides what… |
html.selector | html.selector.* | A small, deliberately bounded CSS selector engine: enough to find things in a parsed document, and no more. |
html.serialize | html.serialize.* | Writing a document back out, shaped for whoever has to read it next: minify() for a browser, format() for… |
html.tokenizer | html.tokenizer.* | The HTML tokenizer from the WHATWG HTML Living Standard: the stage that turns a string of markup into a… |
2026, Richard Ore and The Zuri Contributors