Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html.tokenizer

import html.tokenizer

html lifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelled html.tokenizer.* needs import html.tokenizer.

The HTML tokenizer from the WHATWG HTML Living Standard: the stage that turns a string of markup into a stream of doctype, tag, comment and character tokens.

You rarely need this directly. html.parse() drives it to build a document tree, and that is what almost every program wants. Reach for the tokenizer when you care about the markup as written: linting a template, rewriting attributes in place, or checking which parse errors a document raises.

import html

for token in html.tokenize('<p class="x">hi</p>') {
  echo '${token.type} ${token.name}${token.data}'
}

# start-tag p
# character hi
# end-tag p
# eof

Two deliberate departures from a literal reading of the spec

The standard emits one character token per character. This tokenizer emits one per contiguous run, so a paragraph of prose is a single token rather than a few hundred. Nothing observable changes: the tree builder splits runs back apart wherever the algorithm is defined per character.

The standard also spells the “character reference” sub-machine as nine separate tokenizer states. Here that algorithm lives in html.entities.consume_reference() and is shared with html.decode(), so both agree on every edge case by construction rather than by careful duplication.

Constants

RCDATA_ELEMENTS

html.tokenizer.RCDATA_ELEMENTS = [...]

Tag names whose content is text with character references but no markup. A <title> may contain &amp; but never a nested element.

RAWTEXT_ELEMENTS

html.tokenizer.RAWTEXT_ELEMENTS = [...]

Tag names whose content is text with neither markup nor character references. <noscript> joins this list only when the tokenizer is created with scripting enabled.

Functions

tokenize()

html.tokenize(source: string, options: ?dict) -> list

Tokenizes source and returns every token, ending with the eof token.

This is the whole-document convenience over Tokenizer: it runs with auto_content_state on, so <script>, <style>, <title>, <textarea> and <plaintext> have their content tokenized as text exactly as a browser would, without a tree builder being involved.

%> import html
%> html.tokenize('<b>hi</b>')
[<start-tag b>, <character 'hi'>, <end-tag b>, <eof>]

options takes scripting (default false), which decides whether <noscript> content is text or markup.

Parameters

  • source (string)
  • options (?dict)

Returns list

Classes

Token

class html.Token

One token from the tokenizer.

A single class covers all five token types rather than five classes, because the tree builder dispatches on type on every token and an inheritance check would be the hottest thing in the parser. Which fields carry meaning depends on type:

typefields that matter
doctypename, public_id, system_id, force_quirks
start-tagname, attributes, self_closing
end-tagname, attributes, self_closing
commentdata
characterdata
eofnone
  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
typeOne of 'doctype', 'start-tag', 'end-tag', 'comment', 'character' or 'eof'.
nameThe tag or doctype name, lowercased.
dataCharacter data for a character token, or the text between <!-- and --> for a comment.
attributesA tag’s attributes in source order, as a dictionary of lowercased name to value.
self_closingTrue when the tag was written with a trailing slash, as in .
acknowledged_self_closingSet by the tree builder when it has taken self_closing into account, which only a void or foreign element…
public_idA doctype’s public identifier, or nil when it has none.
system_idA doctype’s system identifier, or nil when it has none.
force_quirksSet on a doctype the tokenizer could not read properly.
line1-based line the token started on.
column1-based column the token started at.

Constructor

html.Token(type: string)

Parameters

  • type (string)

Token.is_start()

html.Token.is_start(name: string) -> bool

True when this token is a start tag named name.

Parameters

  • name (string)

Returns bool

Token.is_end()

html.Token.is_end(name: string) -> bool

True when this token is an end tag named name.

Parameters

  • name (string)

Returns bool

Token.to_string()

html.Token.to_string() -> string

A short, readable rendering of the token, meant for debugging and for test output rather than for round-tripping to markup.

Returns string

ParseError

class html.ParseError

A parse error raised while tokenizing or building the tree.

HTML has no fatal errors: every one of these is recoverable and the parse always produces a usable result. They are collected so that a linter, or anyone auditing markup, can see what a browser silently forgave.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
codeThe standard’s name for this error, such as 'unexpected-null-character' or 'eof-in-tag'.
line1-based line the error was noticed on.
column1-based column the error was noticed at.

Constructor

html.ParseError(code: string, line: number, column: number)

Parameters

  • code (string)
  • line (number)
  • column (number)

ParseError.to_string()

html.ParseError.to_string() -> string

Returns string

Tokenizer

class html.Tokenizer

Turns markup into tokens, one call to next_token() at a time.

The tokenizer’s state is not entirely its own: which state it should be in after a start tag depends on what the tree builder decided to do with that tag. Two ways to bridge that:

  • Driven by html.parse(), the tree builder calls set_state() itself, exactly as the standard describes.
  • Used on its own (which is what tokenize() does), the tokenizer switches itself on <script>, <style>, <title>, <textarea>, <plaintext> and friends, which is what the tree builder would have told it to do anyway.

The second behaviour is controlled by the auto_content_state option and is on by default, so a tokenizer you create by hand does the useful thing without extra ceremony.

Fields

FieldTypeDescription
inputThe source, split into characters, with newlines already normalized: CRLF and lone CR both become LF, as the…
posHow far into input the tokenizer has read.
stateThe state machine’s current state, named exactly as the standard names it but in snake case: 'data',…
errorsParse errors seen so far, as ParseError instances, in the order they happened.
scriptingWhether to enter RAWTEXT for <noscript>, and generally to behave as a browser with scripting turned on…
auto_content_stateWhether the tokenizer switches itself into RCDATA, RAWTEXT, script data or PLAINTEXT after the corresponding…
cdata_okSet by the tree builder when the insertion point is inside foreign content, where <![CDATA[ is meaningful.

Constructor

html.Tokenizer(source: string, options: ?dict)

Parameters

  • source (string)
  • options (?dict) — scripting (default false) and auto_content_state (default true).

Tokenizer.next_token()

html.Tokenizer.next_token() -> ?Token

The next token, or nil once the eof token has been returned.

The eof token is returned exactly once and is always the last one, so a loop can stop on nil or on token.type == 'eof' interchangeably.

Returns ?Token

Tokenizer.set_state()

html.Tokenizer.set_state(name: string) -> Tokenizer

Puts the tokenizer into name, one of the state names listed on the state field.

The tree builder calls this to enter RCDATA or RAWTEXT after a start tag, which is how the standard splits the work between the two stages.

Parameters

  • name (string)

Returns Tokenizer

Tokenizer.set_last_start_tag()

html.Tokenizer.set_last_start_tag(name: ?string) -> Tokenizer

Tells the tokenizer which start tag the tree builder most recently opened, so that RCDATA and RAWTEXT know which end tag closes them.

Parameters

  • name (?string)

Returns Tokenizer


2026, Richard Ore and The Zuri Contributors