Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html.node

import html

Everything here is re-exported by html, so import html is enough and the names are called as html.*. Importing html.node on its own works too and reaches the same definitions.

The document tree that html.parse() produces, and everything you can do with it: walking it, querying it, reading and changing attributes and content, and serializing any part of it back to markup.

The shape follows the DOM closely enough that the names are familiar, without pretending to be a browser. Where this module departs from the DOM it does so deliberately and says why in the relevant doc block; the two differences worth knowing up front are:

  • children holds every child node (elements, text and comments alike), which is what the DOM calls childNodes. Use child_elements() when you only want elements.
  • There are no property setters in Zuri, so every DOM property that can be written is a pair of methods: text_content() reads and set_text_content() writes.
import html

var doc = html.parse('<ul><li class="a">One<li class="a">Two</ul>')

for item in doc.query_selector_all('li.a') {
  echo item.text_content()
}

Constants

NODE_ELEMENT

html.NODE_ELEMENT = 1

Node type of an Element.

NODE_TEXT

html.NODE_TEXT = 3

Node type of a Text node.

NODE_COMMENT

html.NODE_COMMENT = 8

Node type of a Comment.

NODE_DOCUMENT

html.NODE_DOCUMENT = 9

Node type of a Document.

NODE_DOCUMENT_TYPE

html.NODE_DOCUMENT_TYPE = 10

Node type of a DocumentType (the <!DOCTYPE ...> node).

NODE_DOCUMENT_FRAGMENT

html.NODE_DOCUMENT_FRAGMENT = 11

Node type of a DocumentFragment, including the fragment that holds a <template> element’s content.

VOID_ELEMENTS

html.VOID_ELEMENTS = [...]

The HTML elements that are written without a closing tag and can hold no content. Serializing one of these emits <br> and nothing else, and any child it somehow acquired is not written out, because markup that cannot be read back is worse than markup that loses it.

This is the list the fragment serialization algorithm uses, which is wider than the standard’s modern “void elements” list: it also covers basefont, bgsound, frame, keygen and param, five legacy elements that were dropped from the authoring list but still take no end tag and still turn up in real documents.

RAW_TEXT_ELEMENTS

html.RAW_TEXT_ELEMENTS = [...]

Elements whose text children are markup, not content: their text is written out byte for byte and never escaped, because escaping it would change what a stylesheet or a script means.

The serialization algorithm also lists noscript here, but only when the scripting flag is enabled. This module parses with scripting disabled by default, and in that mode a <noscript> holds real elements rather than text, so leaving it out is right for the default and safe for the other case: the worst that happens to a document parsed with scripting: true is that its <noscript> text comes back escaped, whereas including it would let hand-built text inside a <noscript> serialize into live markup.

Classes

HierarchyError

class html.HierarchyError < Error

Raised when a tree mutation would produce a structure that cannot exist, such as inserting a node before something that is not a child of the target, or making a node its own ancestor.

Constructor

html.HierarchyError(message)

Parameters

  • message (string)

Node

class html.Node

The base class every node in a parsed document inherits from.

You never construct a Node directly; parsing produces Document, Element, Text, Comment and DocumentType instances, and this class is where the behaviour they share lives.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
node_typeOne of the NODE_* constants in this module, identifying which subclass this node actually is.
parent_nodeThe node this one hangs off, or nil for a Document and for any node that has been detached.
childrenEvery child of this node in document order: elements, text nodes and comments together.

Constructor

html.Node(node_type: number)

Parameters

  • node_type (number)

Node.is_element()

html.Node.is_element() -> bool

True when this node is an Element.

Returns bool

Node.is_text()

html.Node.is_text() -> bool

True when this node is a Text node.

Returns bool

Node.is_comment()

html.Node.is_comment() -> bool

True when this node is a Comment.

Returns bool

Node.is_document()

html.Node.is_document() -> bool

True when this node is a Document.

Returns bool

Node.is_doctype()

html.Node.is_doctype() -> bool

True when this node is a DocumentType.

Returns bool

Node.is_fragment()

html.Node.is_fragment() -> bool

True when this node is a DocumentFragment.

Returns bool

Node.first_child()

html.Node.first_child() -> ?Node

The first child of this node, or nil when it has none.

Returns ?Node

Node.last_child()

html.Node.last_child() -> ?Node

The last child of this node, or nil when it has none.

Returns ?Node

Node.index_in_parent()

html.Node.index_in_parent() -> number

Position of this node among its parent’s children, or -1 when it has no parent.

Returns number

Node.next_sibling()

html.Node.next_sibling() -> ?Node

The node immediately after this one under the same parent, or nil when this is the last child or has no parent.

Returns ?Node

Node.previous_sibling()

html.Node.previous_sibling() -> ?Node

The node immediately before this one under the same parent, or nil when this is the first child or has no parent.

Returns ?Node

Node.next_element_sibling()

html.Node.next_element_sibling() -> ?Element

The next sibling that is an element, skipping over text and comments, or nil when there is none.

Returns ?Element

Node.previous_element_sibling()

html.Node.previous_element_sibling() -> ?Element

The previous sibling that is an element, skipping over text and comments, or nil when there is none.

Returns ?Element

Node.child_elements()

html.Node.child_elements() -> list

This node’s children that are elements, in document order.

A fresh list is returned on every call, so changing it does not change the tree.

Returns list

Node.first_element_child()

html.Node.first_element_child() -> ?Element

The first child that is an element, or nil.

Returns ?Element

Node.last_element_child()

html.Node.last_element_child() -> ?Element

The last child that is an element, or nil.

Returns ?Element

Node.root()

html.Node.root() -> Node

The topmost node reachable by following parent_node, which for a parsed document is the Document itself and for a detached subtree is that subtree’s own root.

Returns Node

Node.has_ancestor()

html.Node.has_ancestor(other: instance) -> bool

True when other is this node or one of its ancestors.

Parameters

  • other (Node)

Returns bool

Node.descendants()

html.Node.descendants() -> list

Every node beneath this one, in document order, not including this node itself.

Returns list

Node.walk()

html.Node.walk(callback: function)

Calls callback once for every node beneath this one, in document order. The node is passed as the only argument.

Returning false from the callback prunes that node’s subtree; any other return value (including nil) keeps walking. The walk reads children as it goes, so do not restructure the tree from inside the callback.

Parameters

  • callback (function)

Node.append_child()

html.Node.append_child(node: instance) -> Node

Adds node as this node’s last child, detaching it from wherever it currently lives first. Returns node.

Parameters

  • node (Node)

Returns Node

Raises HierarchyError when node is this node or one of its ancestors, which would make the tree cyclic.

Node.insert_before()

html.Node.insert_before(node: instance, reference: ?instance) -> Node

Inserts node immediately before reference, which must be a child of this node. Passing nil for reference appends. Returns node.

Parameters

  • node (Node)
  • reference (?Node)

Returns Node

Raises HierarchyError when reference is not a child of this node, or when the insertion would make the tree cyclic.

Node.remove_child()

html.Node.remove_child(node: instance) -> Node

Removes node from this node’s children and returns it. The removed node keeps its own children; only its link to this parent is broken.

Parameters

  • node (Node)

Returns Node

Raises HierarchyError when node is not a child of this node.

Node.replace_child()

html.Node.replace_child(replacement: instance, existing: instance) -> Node

Puts replacement where existing currently sits and returns existing, now detached.

Parameters

  • replacement (Node)
  • existing (Node)

Returns Node

Raises HierarchyError when existing is not a child of this node, or when the replacement would make the tree cyclic.

Node.detach()

html.Node.detach() -> Node

Removes this node from its parent, if it has one. Returns this node so calls can be chained.

Returns Node

Node.clear_children()

html.Node.clear_children() -> Node

Removes every child of this node. The children are detached but otherwise untouched.

Returns Node

Node.text_content()

html.Node.text_content() -> string

The concatenated text of every Text node beneath this one, in document order. Comments and doctypes contribute nothing.

Text and Comment override this to return their own data.

Unlike the browser DOM, a Document answers this the same way an element does rather than returning nothing; being told the text of a document you just parsed is far more useful than being told nil.

Returns string

Node.set_text_content()

html.Node.set_text_content(value: string) -> Node

Replaces every child of this node with a single Text node holding value. Passing an empty string just empties the node, matching the DOM.

Parameters

  • value (string)

Returns Node

Node.inner_html()

html.Node.inner_html() -> string

The markup of this node’s children, serialized the way the HTML fragment serialization algorithm specifies.

Returns string

Node.set_inner_html()

html.Node.set_inner_html(source: string) -> Node

Parses source as HTML in the context of this node and replaces all of its children with the result.

The parse runs the fragment parsing algorithm with this node as the context element, so source is interpreted exactly as it would be had it appeared inside this element in the original document. That matters: '<td>x' keeps its cell inside a <tr> and loses it anywhere else, which is what a browser does too.

Nothing is executed and nothing is fetched; <script> content becomes an inert text node.

Parameters

  • source (string)

Returns Node

Node.outer_html()

html.Node.outer_html() -> string

The markup of this node including its own tags.

For a Document this is the whole document; for a Text node it is the escaped text; for a Comment it is <!--...-->.

Returns string

Node.set_outer_html()

html.Node.set_outer_html(source: string) -> list

Parses source as HTML in the context of this node’s parent and puts the result where this node currently sits.

Returns the list of nodes that replaced this one, which may be empty when source produces nothing. This node is detached either way.

Parameters

  • source (string)

Returns list

Raises HierarchyError when this node has no parent, since there would be nowhere to put the result.

Node.serialize_into()

html.Node.serialize_into(out: list)

Appends this node’s markup to out, a list of string pieces the caller is expected to join().

outer_html() is the friendly form of this and is what you normally want. Reach for this one when you are writing a large document out in pieces and would rather not build the whole string in memory first.

Every node class overrides it; the base implementation writes nothing.

Parameters

  • out (list)

Node.clone_node()

html.Node.clone_node(deep: ?bool) -> Node

A deep or shallow copy of this node, detached from any parent.

The copy is shallow by default, matching the DOM: you get the node itself with its attributes but with no children. Pass true for deep to copy the whole subtree.

Parameters

  • deep (?bool)

Returns Node

Node.get_element_by_id()

html.Node.get_element_by_id(value: string) -> ?Element

The first element beneath this node whose id attribute is value, or nil when there is none.

Ids are compared exactly, including case: HTML does not fold the case of id values even though it folds tag names.

A document with duplicate ids is malformed but perfectly parseable, and this returns the first one in document order, which is what browsers settled on.

Parameters

  • value (string)

Returns ?Element

Node.get_elements_by_class_name()

html.Node.get_elements_by_class_name(names) -> list

Every element beneath this node carrying all of the classes in names, in document order.

names may be a single class name, several separated by whitespace, or a list of names; in every form an element must carry all of them to match. Class names are compared exactly, including case. An empty names matches nothing.

%> doc.get_elements_by_class_name('warning')
%> doc.get_elements_by_class_name('warning sticky')
%> doc.get_elements_by_class_name(['warning', 'sticky'])

Parameters

  • names (string|list)

Returns list

Node.get_elements_by_tag_name()

html.Node.get_elements_by_tag_name(name: string) -> list

Every element beneath this node whose tag name is name, in document order.

The comparison folds ASCII case, so 'DIV' and 'div' find the same elements. Passing '*' returns every element.

Parameters

  • name (string)

Returns list

Node.query_selector()

html.Node.query_selector(query: string) -> ?Element

The first element beneath this node matching the CSS selector query, in document order, or nil when nothing matches.

The supported selector syntax is described in html.selector.

Parameters

  • query (string)

Returns ?Element

Raises SelectorError when query is not valid.

Node.query_selector_all()

html.Node.query_selector_all(query: string) -> list

Every element beneath this node matching the CSS selector query, in document order.

Parameters

  • query (string)

Returns list

Raises SelectorError when query is not valid.

Node.to_string()

html.Node.to_string() -> string

Renders this node as markup, the same as outer_html().

Returns string

Text

class html.Text < Node

A run of character data in the tree.

The parser merges adjacent characters into as few Text nodes as the tree construction algorithm allows, so a paragraph of prose is normally one node rather than one per character.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
dataThe characters this node holds, already decoded: character references were resolved during parsing, so data…

Constructor

html.Text(data: string)

Parameters

  • data (string)

Text.text_content()

html.Text.text_content() -> string

The characters this node holds.

Returns string

Text.set_text_content()

html.Text.set_text_content(value: string) -> Text

Replaces this node’s characters with value.

Parameters

  • value (string)

Returns Text

Text.append_data()

html.Text.append_data(value: string) -> Text

Appends value to this node’s characters.

Parameters

  • value (string)

Returns Text

Text.serialize_into()

html.Text.serialize_into(out: list)

Parameters

  • out (list)

Comment

class html.Comment < Node

An HTML comment.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
dataThe text between <!-- and -->, exactly as it appeared.

Constructor

html.Comment(data: string)

Parameters

  • data (string)

Comment.text_content()

html.Comment.text_content() -> string

The comment’s text.

Returns string

Comment.set_text_content()

html.Comment.set_text_content(value: string) -> Comment

Replaces the comment’s text with value.

Parameters

  • value (string)

Returns Comment

Comment.serialize_into()

html.Comment.serialize_into(out: list)

Parameters

  • out (list)

DocumentType

class html.DocumentType < Node

The <!DOCTYPE ...> node at the top of a document.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
nameThe doctype’s name, lowercased by the tokenizer.
public_idThe public identifier, or an empty string when the doctype has none.
system_idThe system identifier, or an empty string when the doctype has none.

Constructor

html.DocumentType(name: string, public_id: ?string, system_id: ?string)

Parameters

  • name (string)
  • public_id (?string)
  • system_id (?string)

DocumentType.text_content()

html.DocumentType.text_content() -> string

A doctype contains no text, so this is always an empty string.

Returns string

DocumentType.serialize_into()

html.DocumentType.serialize_into(out: list)

Parameters

  • out (list)

DocumentFragment

class html.DocumentFragment < Node

A parentless container for a run of nodes.

Fragment parsing produces one of these, and every <template> element owns one holding its content.

  • printable — has a @to_string(), so echo and print() show something useful

Constructor

html.DocumentFragment()

DocumentFragment.serialize_into()

html.DocumentFragment.serialize_into(out: list)

Parameters

  • out (list)

Document

class html.Document < Node

A whole parsed document.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
modeWhich quirks mode the doctype (or the lack of one) put this document in: 'no-quirks', 'limited-quirks' or…
errorsEvery parse error the tokenizer and tree builder raised, in the order they happened.

Constructor

html.Document()

Document.document_element()

html.Document.document_element() -> ?Element

The document’s root element, normally <html>, or nil for an empty document.

Returns ?Element

Document.doctype()

html.Document.doctype() -> ?DocumentType

The document’s <!DOCTYPE> node, or nil when it has none.

Returns ?DocumentType

Document.head()

html.Document.head() -> ?Element

The document’s <head> element, or nil.

The tree construction algorithm always creates one when parsing a whole document, even for input with no <head> tag in it, so this only returns nil for a document that was built by hand.

Returns ?Element

Document.body()

html.Document.body() -> ?Element

The document’s <body> element, or the <frameset> that stands in for it in a frameset document, or nil when there is neither.

Returns ?Element

Document.title()

html.Document.title() -> string

The text of the document’s <title> element, with leading and trailing whitespace removed, or an empty string when there is no title.

Returns string

Document.serialize_into()

html.Document.serialize_into(out: list)

Parameters

  • out (list)

ClassList

class html.ClassList

A live view of one element’s class attribute.

Every method reads and writes the attribute itself, so a class_list never goes stale and two of them taken from the same element always agree.

%> var el = doc.query_selector('div')
%> el.class_list().add('active')
%> el.class_list().contains('active')
true
%> el.get_attribute('class')
'active'
  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
elementThe element this list reads from and writes to.

Constructor

html.ClassList(element)

Parameters

  • element (Element)

ClassList.to_list()

html.ClassList.to_list() -> list

The class names currently on the element, in source order, with duplicates removed.

Returns list

ClassList.length()

html.ClassList.length() -> number

How many distinct classes the element carries.

Returns number

ClassList.contains()

html.ClassList.contains(name: string) -> bool

True when the element carries name. The comparison is exact: HTML class names are case-sensitive in standards mode.

Parameters

  • name (string)

Returns bool

ClassList.add()

html.ClassList.add(...names: list) -> ClassList

Adds every name given to the element’s class attribute, ignoring any it already has. Returns the list so calls can be chained.

An empty name, or one containing whitespace, is rejected: those cannot be written to a class attribute without changing what it means.

Parameters

  • names (...string)

Returns ClassList

Raises ArgumentError

ClassList.remove()

html.ClassList.remove(...names: list) -> ClassList

Removes every name given from the element’s class attribute, ignoring any it does not have. Returns the list.

Parameters

  • names (...string)

Returns ClassList

Raises ArgumentError

ClassList.toggle()

html.ClassList.toggle(name: string, force: ?bool) -> bool

Adds name when the element does not have it and removes it when it does, then returns true when the class is present afterwards.

Passing force turns this into an unconditional add (true) or remove (false), which is handy when the desired state is already in a variable.

Parameters

  • name (string)
  • force (?bool)

Returns bool

Raises ArgumentError

ClassList.to_string()

html.ClassList.to_string() -> string

The class attribute’s value as it would be written out.

Returns string

Element

class html.Element < Node

An element in the tree.

  • printable — has a @to_string(), so echo and print() show something useful

Fields

FieldTypeDescription
tag_nameThe element’s tag name.
namespaceThe namespace this element lives in: one of HTML_NAMESPACE, SVG_NAMESPACE or MATHML_NAMESPACE.
attributesThe element’s attributes, in source order, as a dictionary of name to value.
attribute_namespacesNamespace URIs for the few attributes that have one, keyed by the same qualified name used in attributes.
source_lineThe line in the source where this element’s start tag began, counting from 1, or 0 for an element that was…
source_columnThe column in the source where this element’s start tag began, counting from 1, or 0 for an element that…
contentFor a <template> element, the DocumentFragment holding its content.

Constructor

html.Element(tag_name: string, namespace: ?string, attributes: ?dict)

Parameters

  • tag_name (string)
  • namespace (?string)
  • attributes (?dict)

Element.is_html()

html.Element.is_html() -> bool

True when this is an HTML element, as opposed to one inside an <svg> or <math> subtree.

Returns bool

Element.is_void()

html.Element.is_void() -> bool

True when this element is one of the HTML void elements, which have no closing tag and can hold no content.

Returns bool

Element.get_attribute()

html.Element.get_attribute(name: string) -> ?string

The value of the attribute name, or nil when the element does not carry it.

The lookup folds ASCII case for HTML elements, so get_attribute('HREF') and get_attribute('href') are the same question. For foreign elements the name is matched exactly, because SVG really does distinguish viewBox from viewbox.

Parameters

  • name (string)

Returns ?string

Element.has_attribute()

html.Element.has_attribute(name: string) -> bool

True when the element carries the attribute name, whatever its value.

Parameters

  • name (string)

Returns bool

Element.set_attribute()

html.Element.set_attribute(name: string, value: string) -> Element

Sets the attribute name to value, replacing any existing value and leaving the attribute’s position among the others unchanged. A new attribute is appended after the existing ones.

Pass an empty string for a valueless attribute such as disabled; there is no separate “no value” state.

Parameters

  • name (string)
  • value (string)

Returns Element

Raises ArgumentError when name is empty or contains a character that cannot appear in an attribute name.

Element.remove_attribute()

html.Element.remove_attribute(name: string) -> bool

Removes the attribute name and returns true when it was actually there.

Parameters

  • name (string)

Returns bool

Element.attribute_names()

html.Element.attribute_names() -> list

The attribute names this element carries, in source order.

Returns list

Element.attribute_namespace()

html.Element.attribute_namespace(name: string) -> ?string

The namespace URI of the attribute name, or nil when it has none. Only foreign content produces namespaced attributes.

Parameters

  • name (string)

Returns ?string

Element.set_attribute_namespace()

html.Element.set_attribute_namespace(name: string, uri: string) -> Element

Records that the attribute name belongs to uri. Called by the parser while adjusting foreign attributes; there is no reason to call it by hand.

Parameters

  • name (string)
  • uri (string)

Returns Element

Element.id()

html.Element.id() -> string

The element’s id, or an empty string when it has none.

Returns string

Element.class_name()

html.Element.class_name() -> string

The element’s class attribute verbatim, or an empty string.

Returns string

Element.class_names()

html.Element.class_names() -> list

The element’s classes as a list, in source order and without duplicates. An element with no class attribute gives an empty list.

Returns list

Element.class_list()

html.Element.class_list() -> ClassList

A ClassList view of this element’s class attribute, for adding, removing and toggling classes.

A fresh view is returned on every call; because it reads and writes the attribute directly there is no state to keep, and two views of the same element behave identically.

Returns ClassList

Element.matches()

html.Element.matches(query: string) -> bool

True when this element matches the CSS selector query.

Parameters

  • query (string)

Returns bool

Raises SelectorError when query is not valid.

Element.closest()

html.Element.closest(query: string) -> ?Element

This element, or its nearest ancestor, matching the CSS selector query: or nil when neither this element nor any ancestor does.

The search starts at the element itself, which is what makes el.closest('div') return el when el is a div.

Parameters

  • query (string)

Returns ?Element

Raises SelectorError when query is not valid.

Element.serialize_into()

html.Element.serialize_into(out: list)

Parameters

  • out (list)

Element.inner_html()

html.Element.inner_html() -> string

The markup of this element’s children.

A <template> answers with its content fragment rather than with its own children, which are always empty. That matches the DOM, and it is the only way the question has a useful answer for a template.

Returns string

Element.set_inner_html()

html.Element.set_inner_html(source: string) -> Element

Parses source and replaces this element’s children with the result.

For a <template> the result goes into the content fragment, matching where inner_html() reads from.

Parameters

  • source (string)

Returns Element

Element.clone_node()

html.Element.clone_node(deep: ?bool) -> Element

A deep or shallow copy of this element. Deep-cloning a <template> also copies its content fragment, which is otherwise not part of the element’s children.

Parameters

  • deep (?bool)

Returns Element


2026, Richard Ore and The Zuri Contributors