html.tokenizer
import html.tokenizer
htmllifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelledhtml.tokenizer.*needsimport html.tokenizer.
The HTML tokenizer from the WHATWG HTML Living Standard: the stage that turns a string of markup into a stream of doctype, tag, comment and character tokens.
You rarely need this directly. html.parse() drives it to build a
document tree, and that is what almost every program wants. Reach for
the tokenizer when you care about the markup as written: linting a
template, rewriting attributes in place, or checking which parse errors
a document raises.
import html
for token in html.tokenize('<p class="x">hi</p>') {
echo '${token.type} ${token.name}${token.data}'
}
# start-tag p
# character hi
# end-tag p
# eof
Two deliberate departures from a literal reading of the spec
The standard emits one character token per character. This tokenizer emits one per contiguous run, so a paragraph of prose is a single token rather than a few hundred. Nothing observable changes: the tree builder splits runs back apart wherever the algorithm is defined per character.
The standard also spells the “character reference” sub-machine as nine
separate tokenizer states. Here that algorithm lives in
html.entities.consume_reference() and is shared with html.decode(),
so both agree on every edge case by construction rather than by careful
duplication.
Constants
RCDATA_ELEMENTS
html.tokenizer.RCDATA_ELEMENTS = [...]
Tag names whose content is text with character references but no markup.
A <title> may contain & but never a nested element.
RAWTEXT_ELEMENTS
html.tokenizer.RAWTEXT_ELEMENTS = [...]
Tag names whose content is text with neither markup nor character
references. <noscript> joins this list only when the tokenizer is
created with scripting enabled.
Functions
tokenize()
html.tokenize(source: string, options: ?dict) -> list
Tokenizes source and returns every token, ending with the eof token.
This is the whole-document convenience over Tokenizer: it runs with
auto_content_state on, so <script>, <style>, <title>,
<textarea> and <plaintext> have their content tokenized as text
exactly as a browser would, without a tree builder being involved.
%> import html
%> html.tokenize('<b>hi</b>')
[<start-tag b>, <character 'hi'>, <end-tag b>, <eof>]
options takes scripting (default false), which decides whether
<noscript> content is text or markup.
Parameters
source(string)options(?dict)
Returns list
Classes
Token
class html.Token
One token from the tokenizer.
A single class covers all five token types rather than five classes,
because the tree builder dispatches on type on every token and an
inheritance check would be the hottest thing in the parser. Which fields
carry meaning depends on type:
type | fields that matter |
|---|---|
doctype | name, public_id, system_id, force_quirks |
start-tag | name, attributes, self_closing |
end-tag | name, attributes, self_closing |
comment | data |
character | data |
eof | none |
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
type | One of 'doctype', 'start-tag', 'end-tag', 'comment', 'character' or 'eof'. | |
name | The tag or doctype name, lowercased. | |
data | Character data for a character token, or the text between <!-- and --> for a comment. | |
attributes | A tag’s attributes in source order, as a dictionary of lowercased name to value. | |
self_closing | True when the tag was written with a trailing slash, as in . | |
acknowledged_self_closing | Set by the tree builder when it has taken self_closing into account, which only a void or foreign element… | |
public_id | A doctype’s public identifier, or nil when it has none. | |
system_id | A doctype’s system identifier, or nil when it has none. | |
force_quirks | Set on a doctype the tokenizer could not read properly. | |
line | 1-based line the token started on. | |
column | 1-based column the token started at. |
Constructor
html.Token(type: string)
Parameters
type(string)
Token.is_start()
html.Token.is_start(name: string) -> bool
True when this token is a start tag named name.
Parameters
name(string)
Returns bool
Token.is_end()
html.Token.is_end(name: string) -> bool
True when this token is an end tag named name.
Parameters
name(string)
Returns bool
Token.to_string()
html.Token.to_string() -> string
A short, readable rendering of the token, meant for debugging and for test output rather than for round-tripping to markup.
Returns string
ParseError
class html.ParseError
A parse error raised while tokenizing or building the tree.
HTML has no fatal errors: every one of these is recoverable and the parse always produces a usable result. They are collected so that a linter, or anyone auditing markup, can see what a browser silently forgave.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
code | The standard’s name for this error, such as 'unexpected-null-character' or 'eof-in-tag'. | |
line | 1-based line the error was noticed on. | |
column | 1-based column the error was noticed at. |
Constructor
html.ParseError(code: string, line: number, column: number)
Parameters
code(string)line(number)column(number)
ParseError.to_string()
html.ParseError.to_string() -> string
Returns string
Tokenizer
class html.Tokenizer
Turns markup into tokens, one call to next_token() at a time.
The tokenizer’s state is not entirely its own: which state it should be in after a start tag depends on what the tree builder decided to do with that tag. Two ways to bridge that:
- Driven by
html.parse(), the tree builder callsset_state()itself, exactly as the standard describes. - Used on its own (which is what
tokenize()does), the tokenizer switches itself on<script>,<style>,<title>,<textarea>,<plaintext>and friends, which is what the tree builder would have told it to do anyway.
The second behaviour is controlled by the auto_content_state option
and is on by default, so a tokenizer you create by hand does the useful
thing without extra ceremony.
Fields
| Field | Type | Description |
|---|---|---|
input | The source, split into characters, with newlines already normalized: CRLF and lone CR both become LF, as the… | |
pos | How far into input the tokenizer has read. | |
state | The state machine’s current state, named exactly as the standard names it but in snake case: 'data',… | |
errors | Parse errors seen so far, as ParseError instances, in the order they happened. | |
scripting | Whether to enter RAWTEXT for <noscript>, and generally to behave as a browser with scripting turned on… | |
auto_content_state | Whether the tokenizer switches itself into RCDATA, RAWTEXT, script data or PLAINTEXT after the corresponding… | |
cdata_ok | Set by the tree builder when the insertion point is inside foreign content, where <![CDATA[ is meaningful. |
Constructor
html.Tokenizer(source: string, options: ?dict)
Parameters
source(string)options(?dict) —scripting(defaultfalse) andauto_content_state(defaulttrue).
Tokenizer.next_token()
html.Tokenizer.next_token() -> ?Token
The next token, or nil once the eof token has been returned.
The eof token is returned exactly once and is always the last one, so
a loop can stop on nil or on token.type == 'eof' interchangeably.
Returns ?Token
Tokenizer.set_state()
html.Tokenizer.set_state(name: string) -> Tokenizer
Puts the tokenizer into name, one of the state names listed on the
state field.
The tree builder calls this to enter RCDATA or RAWTEXT after a start tag, which is how the standard splits the work between the two stages.
Parameters
name(string)
Returns Tokenizer
Tokenizer.set_last_start_tag()
html.Tokenizer.set_last_start_tag(name: ?string) -> Tokenizer
Tells the tokenizer which start tag the tree builder most recently opened, so that RCDATA and RAWTEXT know which end tag closes them.
Parameters
name(?string)
Returns Tokenizer
2026, Richard Ore and The Zuri Contributors