Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html

import html

A complete HTML5 parser, DOM and serializer, following the WHATWG HTML Living Standard.

Parse a page, walk it, query it with CSS selectors, change it, and write it back out. Nothing here fetches anything over the network and nothing here executes a script; <script> content is inert text and always will be.


Quick start

import html

var doc = html.parse('
  <article>
    <h1 id="title">Zuri</h1>
    <ul class="links">
      <li><a href="/docs">Docs</a>
      <li><a href="/spec">Spec</a>
    </ul>
  </article>
')

echo doc.get_element_by_id('title').text_content()

for link in doc.query_selector_all('ul.links a') {
  echo '${link.text_content()} -> ${link.get_attribute("href")}'
}

Notice that neither <li> is closed and there is no <html>, <head> or <body> anywhere. That is fine: the standard defines a tree for every possible input, and this module builds it. Parsing never raises on malformed markup.


Finding things

MethodFinds
get_element_by_id(id)the first element with that id
get_elements_by_class_name(c)every element with those classes
get_elements_by_tag_name(t)every element with that tag
query_selector(sel)the first CSS match
query_selector_all(sel)every CSS match
element.matches(sel)whether this element matches
element.closest(sel)this element or its nearest matching ancestor

The selector syntax is a deliberately bounded subset: type, *, #id, .class, [attr], [attr="value"], the combinators >, + and ~, descendant combination by whitespace, :first-child, :last-child, :nth-child(), and comma-separated lists of all of those. Anything outside it raises a SelectorError rather than silently matching nothing.


Changing things

import html

var doc = html.parse('<div class="box"><p>old</p></div>')
var box = doc.query_selector('.box')

box.set_attribute('data-seen', 'yes')
box.class_list().add('active')
box.set_inner_html('<p>new</p>')

echo box.outer_html()
# <div class="box active" data-seen="yes"><p>new</p></div>

Zuri has no property setters, so every writable DOM property is a pair of methods: text_content() reads and set_text_content() writes, and likewise for inner_html and outer_html.


Writing it back out

outer_html() round-trips a node exactly. When you want the output shaped for a particular audience, minify() strips what a browser does not need and format() indents for a human:

A string is parsed as a whole document, so <html>, <head> and <body> show up in the output the way a browser would produce them. Pass { fragment: true } when the input is a fragment and you want just that fragment back:

%> import html
%> html.minify('<p>  a   b  </p>\\n<p>c</p>', { fragment: true })
'<p>a b</p><p>c</p>'
%> echo html.format('<ul><li>a<li>b</ul>', { fragment: true })
<ul>
  <li>a</li>
  <li>b</li>
</ul>

Escaping

encode() makes text safe to embed and decode() resolves character references. Decoding knows all 2231 named references, not a shortlist:

%> import html
%> html.encode('<b>Tom & Jerry</b>')
'&lt;b&gt;Tom &amp; Jerry&lt;/b&gt;'
%> html.decode('caf&eacute; &#8212; &#x2603;')
'café: ☃'

Encoding

Input is UTF-8 text that has already been decoded, which is what a Zuri string always is. There is no charset sniffing, no byte order mark handling and no <meta charset> processing: if you are reading bytes off a socket, decode them before they get here.


Parse errors

HTML has no fatal errors, so a parse always succeeds. Every conformance problem the standard defines is collected instead:

%> import html
%> var doc = html.parse('<p><b>x</p>')
%> for e in doc.errors { echo e.to_string() }
missing-doctype at 1:1
unexpected-end-tag at 1:8

Each entry carries the standard’s own error name, so you can look it up, plus the line and column where it was noticed.

The html API

Every public name in html, wherever it is declared. Each links to the page that documents it.

NameKindSummary
html.ClassListclassA live view of one element’s class attribute.
html.CommentclassAn HTML comment.
html.DocumentclassA whole parsed document.
html.DocumentFragmentclassA parentless container for a run of nodes.
html.DocumentTypeclassThe <!DOCTYPE ...> node at the top of a document.
html.ElementclassAn element in the tree.
html.HTML_NAMESPACEconstantNamespace URI of ordinary HTML elements.
html.HierarchyErrorclassRaised when a tree mutation would produce a structure that cannot exist, such as inserting a node before…
html.MATHML_NAMESPACEconstantNamespace URI of elements inside a <math> subtree.
html.NODE_COMMENTconstantNode type of a Comment.
html.NODE_DOCUMENTconstantNode type of a Document.
html.NODE_DOCUMENT_FRAGMENTconstantNode type of a DocumentFragment, including the fragment that holds a <template> element’s content.
html.NODE_DOCUMENT_TYPEconstantNode type of a DocumentType (the <!DOCTYPE ...> node).
html.NODE_ELEMENTconstantNode type of an Element.
html.NODE_TEXTconstantNode type of a Text node.
html.NodeclassThe base class every node in a parsed document inherits from.
html.ParseErrorclassA parse error raised while tokenizing or building the tree.
html.RAW_TEXT_ELEMENTSconstantElements whose text children are markup, not content: their text is written out byte for byte and never…
html.SVG_NAMESPACEconstantNamespace URI of elements inside an <svg> subtree.
html.SelectorclassA compiled selector list: h1, h2 is one of these holding two ComplexSelector instances.
html.SelectorErrorclassRaised when a selector cannot be parsed, or uses syntax this module deliberately does not support.
html.TextclassA run of character data in the tree.
html.TokenclassOne token from the tokenizer.
html.TokenizerclassTurns markup into tokens, one call to next_token() at a time.
html.TreeBuilderclassBuilds a document tree from a token stream.
html.VOID_ELEMENTSconstantThe HTML elements that are written without a closing tag and can hold no content.
html.XLINK_NAMESPACEconstantNamespace URI used by the xlink: attribute prefix in SVG.
html.XMLNS_NAMESPACEconstantNamespace URI used by the xmlns and xmlns:xlink attributes.
html.XML_NAMESPACEconstantNamespace URI used by the xml: attribute prefix.
html.compilefunctionCompiles source into a reusable Selector.
html.decodefunctionDecodes every character reference in text and returns the result.
html.elements.FOREIGN_ATTRIBUTESconstantAttributes in foreign content that belong to a namespace, mapping the lowercase name the tokenizer produced…
html.elements.FOREIGN_BREAKOUT_TAGSconstantStart tags that are always a mistake inside foreign content and that break out of it, closing SVG or MathML…
html.elements.FORMATTING_ELEMENTSconstantThe formatting elements: the ones the adoption agency algorithm reopens across a badly nested boundary, so…
html.elements.HTML4_TRANSITIONAL_PREFIXESconstantPublic identifier prefixes whose mode depends on whether the doctype also carries a system identifier: quirks…
html.elements.IMPLIED_END_TAGSconstantElements whose end tag is implied by the start of a sibling, so that <li>a<li>b produces two list items…
html.elements.LIMITED_QUIRKS_PUBLIC_PREFIXESconstantPublic identifier prefixes that always mean limited-quirks mode.
html.elements.MATHML_ATTRIBUTESconstantThe one MathML attribute whose case the tokenizer destroys.
html.elements.MATHML_TEXT_INTEGRATION_POINTSconstantThe MathML elements whose children are parsed as HTML rather than as MathML.
html.elements.QUIRKS_PUBLIC_EXACTconstantThe two public identifiers that put a document in quirks mode on an exact match rather than a prefix match,…
html.elements.QUIRKS_PUBLIC_PREFIXESconstantPublic identifier prefixes that put a document in quirks mode, in lowercase for case-insensitive comparison.
html.elements.QUIRKS_SYSTEM_IDconstantThe system identifier that alone puts a document in quirks mode, in lowercase.
html.elements.SCOPE_HTMLconstantThe HTML elements that make up the “in scope” barrier every scope check shares.
html.elements.SPECIAL_HTMLconstantThe HTML elements the standard calls “special”: the ones a list item or definition term stops searching past,…
html.elements.SPECIAL_MATHMLconstantThe MathML elements that count as special.
html.elements.SPECIAL_SVGconstantThe SVG elements that count as special.
html.elements.SVG_ATTRIBUTESconstantSVG attribute names the tokenizer lowercased and that have to be put back, keyed by the lowercase form.
html.elements.SVG_HTML_INTEGRATION_POINTSconstantThe SVG elements whose children are parsed as HTML rather than as SVG.
html.elements.SVG_TAG_NAMESconstantSVG tag names the tokenizer lowercased and that have to be put back, keyed by the lowercase form.
html.elements.THOROUGH_IMPLIED_END_TAGSconstantEverything in IMPLIED_END_TAGS plus the elements only closed when the standard says to generate implied end…
html.elements.adjust_mathml_attributesfunctionRepairs the case of MathML attribute names in attributes and returns a new dictionary.
html.elements.adjust_svg_attributesfunctionRepairs the case of SVG attribute names in attributes and returns a new dictionary.
html.elements.adjust_svg_tag_namefunctionThe correctly cased SVG tag name for the lowercase name the tokenizer produced, or name itself when it…
html.elements.doctype_modefunctionWhich quirks mode a doctype puts a document in.
html.elements.is_html_integration_pointfunctionTrue when element is an HTML integration point.
html.elements.is_mathml_text_integration_pointfunctionTrue when element is a MathML text integration point, meaning its children are parsed as HTML.
html.elements.is_specialfunctionTrue when the element name in namespace namespace is one of the standard’s special elements.
html.encodefunctionEscapes text so that it can be embedded in an HTML document without being reinterpreted as markup.
html.entities.NO_BREAK_SPACEconstant
html.entities.REPLACEMENT_CHARACTERconstant
html.entities.code_point_to_stringfunctionApplies the standard’s “numeric character reference end state” rules to a raw code point and returns the…
html.entities.consume_referencefunctionConsumes a character reference from chars beginning at the ampersand at index start.
html.entities.is_ascii_alphanumericfunctionReturns true when c is one of 0-9, A-Z or a-z.
html.entities.is_ascii_digitfunctionReturns true when c is an ASCII decimal digit.
html.entities.is_ascii_hex_digitfunctionReturns true when c is an ASCII hexadecimal digit, in either case.
html.escape_attributefunctionEscapes value the way the HTML fragment serialization algorithm requires for a double-quoted attribute…
html.escape_textfunctionEscapes text the way the HTML fragment serialization algorithm requires for element content.
html.formatfunctionRenders source as indented, readable HTML.
html.minifyfunctionRemoves the markup a browser does not need and returns the result.
html.parsefunctionParses source as a complete HTML document.
html.parse_filefunctionReads the file at path and parses it as HTML.
html.parse_fragmentfunctionParses source as a fragment, as though it had been written inside context.
html.parser.FormattingEntryclassOne entry in the list of active formatting elements.
html.selector.AttributeTestclassOne name/=value test from a [...] block.
html.selector.ComplexSelectorclassOne selector from a comma-separated list: a chain of compounds joined by combinators, such as ul > li + li.
html.selector.CompoundSelectorclassA run of simple selectors with nothing between them, such as div#main.active[data-x]:first-child.
html.selector.PseudoTestclassOne :pseudo or :pseudo(...) test.
html.serialize.INLINE_ELEMENTSconstantElements that flow inside a line of text rather than starting a new block.
html.serialize.PRESERVE_WHITESPACEconstantElements whose text content is significant to the last character.
html.tokenizefunctionTokenizes source and returns every token, ending with the eof token.
html.tokenizer.RAWTEXT_ELEMENTSconstantTag names whose content is text with neither markup nor character references.
html.tokenizer.RCDATA_ELEMENTSconstantTag names whose content is text with character references but no markup.

Submodules

ModuleReached asSummary
html.elementshtml.elements.*The element tables the tree construction algorithm consults on almost every token: which elements are…
html.entitieshtml.entities.*Character reference handling for HTML: turning text into something safe to drop into a document (encode()),…
html.namespaceshtml.*The five namespace URIs the HTML parser deals in.
html.nodehtml.*The document tree that html.parse() produces, and everything you can do with it: walking it, querying it,…
html.parserhtml.parser.*The HTML tree construction stage: the half of the parser that takes the tokenizer’s stream and decides what…
html.selectorhtml.selector.*A small, deliberately bounded CSS selector engine: enough to find things in a parsed document, and no more.
html.serializehtml.serialize.*Writing a document back out, shaped for whoever has to read it next: minify() for a browser, format() for…
html.tokenizerhtml.tokenizer.*The HTML tokenizer from the WHATWG HTML Living Standard: the stage that turns a string of markup into a…

2026, Richard Ore and The Zuri Contributors