html.parser
import html.parser
htmllifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelledhtml.parser.*needsimport html.parser.
The HTML tree construction stage: the half of the parser that takes the tokenizer’s stream and decides what the document actually means.
This is where HTML’s famous forgiveness lives. <b>a<p>b</b>c,
<table><td>x, an unclosed <li>, a </div> with nothing to close:
none of them are errors you have to handle, because the standard defines
exactly what tree each one produces, and this module implements those
rules rather than inventing its own.
Most programs want html.parse() and never come here. TreeBuilder is
public for the cases that need the machinery itself: reading errors to
lint a document, or driving a parse whose tokenizer you control.
Functions
parse()
html.parse(source: string, options: ?dict) -> Document
Parses source as a complete HTML document.
The result is always a usable Document, however broken the input was:
HTML defines a tree for every possible string, and this follows those
rules rather than raising. Anything the standard calls a parse error is
collected in document.errors, so a non-empty list there is how you
find out the source was not conformant.
The document always has an <html> root with a <head> and a <body>,
even when the source contains none of those tags, because the tree
construction algorithm creates them.
options takes:
scripting(bool, defaultfalse): parse as though scripting were enabled. With the default,<noscript>content is parsed as real markup, which is what you want when reading a page rather than rendering it. Nothing is ever executed either way.
The source is treated as UTF-8 text that has already been decoded. No
charset sniffing, byte order mark handling or <meta charset>
processing happens; a Zuri string is already Unicode by the time it
reaches here.
%> import html
%> var doc = html.parse('<p>Hello<p>World')
%> doc.query_selector_all('p').length()
2
%> doc.body().outer_html()
'<body><p>Hello</p><p>World</p></body>'
Parameters
source(string)options(?dict)
Returns Document
parse_file()
html.parse_file(path: string, options: ?dict) -> Document
Reads the file at path and parses it as HTML.
The file is read as UTF-8. Bytes that are not valid UTF-8 are replaced with U+FFFD rather than raising, so a mislabelled file still parses into something you can inspect.
import html
var doc = html.parse_file('page.html')
echo doc.title()
Parameters
path(string)options(?dict) — The same optionsparse()takes.
Returns Document
Raises Error when the file cannot be read.
parse_fragment()
html.parse_fragment(source: string, context, options: ?dict) -> DocumentFragment
Parses source as a fragment, as though it had been written inside context.
This is the algorithm behind set_inner_html(), and the context
matters: '<td>x' keeps its cell when the context is a <tr> and loses
it anywhere else, exactly as it would in a browser.
context may be an Element or a tag name. A tag name is treated as an
HTML element with no attributes, which is the common case. Passing nil
parses in a <body> context.
The returned DocumentFragment holds the parsed nodes. Its children are
detached from any document, ready to be inserted wherever you want them.
%> import html
%> html.parse_fragment('<td>x', 'tr').inner_html()
'<td>x</td>'
%> html.parse_fragment('<td>x', 'div').inner_html()
'x'
Parameters
source(string)context(?Element|string)options(?dict) — The same optionsparse()takes.
Returns DocumentFragment
Classes
FormattingEntry
class html.parser.FormattingEntry
One entry in the list of active formatting elements.
The list holds the element and the token it was created from, because the adoption agency algorithm has to build fresh copies of an element from its original attributes long after the tag that opened it has gone.
Fields
| Field | Type | Description |
|---|---|---|
element | The element currently standing for this entry, or nil when the entry is a marker. | |
token | The start tag token the element was created from. |
Constructor
html.parser.FormattingEntry(element, token)
Parameters
element(?Element) —nilmakes this a marker.token(?Token)
FormattingEntry.is_marker()
html.parser.FormattingEntry.is_marker() -> bool
True when this entry is a marker rather than an element.
Markers are pushed when a <table>, <template>, <caption>, <td>
or an applet-like element opens, and they stop formatting elements from
leaking across that boundary.
Returns bool
TreeBuilder
class html.TreeBuilder
Builds a document tree from a token stream.
The usual way in is parse(), which wires a tokenizer to a builder and
runs it. Construct one directly when you want to watch the process:
errors accumulates every conformance problem, and the stack of open
elements is readable at any point.
Fields
| Field | Type | Description |
|---|---|---|
document | The document being built. | |
tokenizer | The tokenizer feeding this builder. | |
errors | Parse errors from both stages, in the order they happened. | |
open_elements | The stack of open elements. | |
formatting | The list of active formatting elements, holding FormattingEntry instances. | |
mode | The current insertion mode, named as the standard names it: 'initial', 'in body', 'in table text' and… | |
original_mode | Where to return to after a RAWTEXT or RCDATA element finishes. | |
template_modes | The stack of template insertion modes, one per open <template>. | |
head_element | The <head> element, once it exists. | |
form_element | The innermost open <form>, which is what makes a stray </form> close the right thing. | |
frameset_ok | Whether a <frameset> could still legally replace the body. | |
scripting | Whether to parse as though scripting were enabled. | |
fragment_context | The context element when this builder is parsing a fragment, or nil for a whole document. | |
foster_parenting | Whether insertions are currently being foster parented out of a table. | |
pending_characters | Character tokens collected by the “in table text” insertion mode before it decides whether they are legal… | |
done | Set once parsing has stopped, either at end of file or because a rule said to stop. |
Constructor
html.TreeBuilder(source, options: ?dict)
Parameters
source(Tokenizer)options(?dict) —scripting(defaultfalse).
TreeBuilder.run()
html.TreeBuilder.run() -> Document
Runs the parse to completion and returns the document.
Returns Document
TreeBuilder.current_node()
html.TreeBuilder.current_node() -> ?Element
The element the parser is currently inside, or nil when the stack is
empty.
Returns ?Element
TreeBuilder.prepare_fragment()
html.TreeBuilder.prepare_fragment(context) -> Element
Puts this builder into the state the fragment parsing algorithm starts
from, with context as the element the markup is being parsed inside,
and returns the synthetic <html> root the parsed nodes will hang off.
parse_fragment() is the friendly form of this and is what you normally
want; this is public so a caller driving the tokenizer by hand can set
the same state up.
Parameters
context(Element)
Returns Element
2026, Richard Ore and The Zuri Contributors