Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html.parser

import html.parser

html lifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelled html.parser.* needs import html.parser.

The HTML tree construction stage: the half of the parser that takes the tokenizer’s stream and decides what the document actually means.

This is where HTML’s famous forgiveness lives. <b>a<p>b</b>c, <table><td>x, an unclosed <li>, a </div> with nothing to close: none of them are errors you have to handle, because the standard defines exactly what tree each one produces, and this module implements those rules rather than inventing its own.

Most programs want html.parse() and never come here. TreeBuilder is public for the cases that need the machinery itself: reading errors to lint a document, or driving a parse whose tokenizer you control.

Functions

parse()

html.parse(source: string, options: ?dict) -> Document

Parses source as a complete HTML document.

The result is always a usable Document, however broken the input was: HTML defines a tree for every possible string, and this follows those rules rather than raising. Anything the standard calls a parse error is collected in document.errors, so a non-empty list there is how you find out the source was not conformant.

The document always has an <html> root with a <head> and a <body>, even when the source contains none of those tags, because the tree construction algorithm creates them.

options takes:

  • scripting (bool, default false): parse as though scripting were enabled. With the default, <noscript> content is parsed as real markup, which is what you want when reading a page rather than rendering it. Nothing is ever executed either way.

The source is treated as UTF-8 text that has already been decoded. No charset sniffing, byte order mark handling or <meta charset> processing happens; a Zuri string is already Unicode by the time it reaches here.

%> import html
%> var doc = html.parse('<p>Hello<p>World')
%> doc.query_selector_all('p').length()
2
%> doc.body().outer_html()
'<body><p>Hello</p><p>World</p></body>'

Parameters

  • source (string)
  • options (?dict)

Returns Document

parse_file()

html.parse_file(path: string, options: ?dict) -> Document

Reads the file at path and parses it as HTML.

The file is read as UTF-8. Bytes that are not valid UTF-8 are replaced with U+FFFD rather than raising, so a mislabelled file still parses into something you can inspect.

import html

var doc = html.parse_file('page.html')
echo doc.title()

Parameters

  • path (string)
  • options (?dict) — The same options parse() takes.

Returns Document

Raises Error when the file cannot be read.

parse_fragment()

html.parse_fragment(source: string, context, options: ?dict) -> DocumentFragment

Parses source as a fragment, as though it had been written inside context.

This is the algorithm behind set_inner_html(), and the context matters: '<td>x' keeps its cell when the context is a <tr> and loses it anywhere else, exactly as it would in a browser.

context may be an Element or a tag name. A tag name is treated as an HTML element with no attributes, which is the common case. Passing nil parses in a <body> context.

The returned DocumentFragment holds the parsed nodes. Its children are detached from any document, ready to be inserted wherever you want them.

%> import html
%> html.parse_fragment('<td>x', 'tr').inner_html()
'<td>x</td>'
%> html.parse_fragment('<td>x', 'div').inner_html()
'x'

Parameters

  • source (string)
  • context (?Element|string)
  • options (?dict) — The same options parse() takes.

Returns DocumentFragment

Classes

FormattingEntry

class html.parser.FormattingEntry

One entry in the list of active formatting elements.

The list holds the element and the token it was created from, because the adoption agency algorithm has to build fresh copies of an element from its original attributes long after the tag that opened it has gone.

Fields

FieldTypeDescription
elementThe element currently standing for this entry, or nil when the entry is a marker.
tokenThe start tag token the element was created from.

Constructor

html.parser.FormattingEntry(element, token)

Parameters

  • element (?Element) — nil makes this a marker.
  • token (?Token)

FormattingEntry.is_marker()

html.parser.FormattingEntry.is_marker() -> bool

True when this entry is a marker rather than an element.

Markers are pushed when a <table>, <template>, <caption>, <td> or an applet-like element opens, and they stop formatting elements from leaking across that boundary.

Returns bool

TreeBuilder

class html.TreeBuilder

Builds a document tree from a token stream.

The usual way in is parse(), which wires a tokenizer to a builder and runs it. Construct one directly when you want to watch the process: errors accumulates every conformance problem, and the stack of open elements is readable at any point.

Fields

FieldTypeDescription
documentThe document being built.
tokenizerThe tokenizer feeding this builder.
errorsParse errors from both stages, in the order they happened.
open_elementsThe stack of open elements.
formattingThe list of active formatting elements, holding FormattingEntry instances.
modeThe current insertion mode, named as the standard names it: 'initial', 'in body', 'in table text' and…
original_modeWhere to return to after a RAWTEXT or RCDATA element finishes.
template_modesThe stack of template insertion modes, one per open <template>.
head_elementThe <head> element, once it exists.
form_elementThe innermost open <form>, which is what makes a stray </form> close the right thing.
frameset_okWhether a <frameset> could still legally replace the body.
scriptingWhether to parse as though scripting were enabled.
fragment_contextThe context element when this builder is parsing a fragment, or nil for a whole document.
foster_parentingWhether insertions are currently being foster parented out of a table.
pending_charactersCharacter tokens collected by the “in table text” insertion mode before it decides whether they are legal…
doneSet once parsing has stopped, either at end of file or because a rule said to stop.

Constructor

html.TreeBuilder(source, options: ?dict)

Parameters

  • source (Tokenizer)
  • options (?dict) — scripting (default false).

TreeBuilder.run()

html.TreeBuilder.run() -> Document

Runs the parse to completion and returns the document.

Returns Document

TreeBuilder.current_node()

html.TreeBuilder.current_node() -> ?Element

The element the parser is currently inside, or nil when the stack is empty.

Returns ?Element

TreeBuilder.prepare_fragment()

html.TreeBuilder.prepare_fragment(context) -> Element

Puts this builder into the state the fragment parsing algorithm starts from, with context as the element the markup is being parsed inside, and returns the synthetic <html> root the parsed nodes will hang off.

parse_fragment() is the friendly form of this and is what you normally want; this is public so a caller driving the tokenizer by hand can set the same state up.

Parameters

  • context (Element)

Returns Element


2026, Richard Ore and The Zuri Contributors