html.node
import html
Everything here is re-exported by
html, soimport htmlis enough and the names are called ashtml.*. Importinghtml.nodeon its own works too and reaches the same definitions.
The document tree that html.parse() produces, and everything you can
do with it: walking it, querying it, reading and changing attributes and
content, and serializing any part of it back to markup.
The shape follows the DOM closely enough that the names are familiar, without pretending to be a browser. Where this module departs from the DOM it does so deliberately and says why in the relevant doc block; the two differences worth knowing up front are:
childrenholds every child node (elements, text and comments alike), which is what the DOM callschildNodes. Usechild_elements()when you only want elements.- There are no property setters in Zuri, so every DOM property that
can be written is a pair of methods:
text_content()reads andset_text_content()writes.
import html
var doc = html.parse('<ul><li class="a">One<li class="a">Two</ul>')
for item in doc.query_selector_all('li.a') {
echo item.text_content()
}
Constants
NODE_ELEMENT
html.NODE_ELEMENT = 1
Node type of an Element.
NODE_TEXT
html.NODE_TEXT = 3
Node type of a Text node.
NODE_COMMENT
html.NODE_COMMENT = 8
Node type of a Comment.
NODE_DOCUMENT
html.NODE_DOCUMENT = 9
Node type of a Document.
NODE_DOCUMENT_TYPE
html.NODE_DOCUMENT_TYPE = 10
Node type of a DocumentType (the <!DOCTYPE ...> node).
NODE_DOCUMENT_FRAGMENT
html.NODE_DOCUMENT_FRAGMENT = 11
Node type of a DocumentFragment, including the fragment that holds a
<template> element’s content.
VOID_ELEMENTS
html.VOID_ELEMENTS = [...]
The HTML elements that are written without a closing tag and can hold no
content. Serializing one of these emits <br> and nothing else, and any
child it somehow acquired is not written out, because markup that cannot
be read back is worse than markup that loses it.
This is the list the fragment serialization algorithm uses, which is
wider than the standard’s modern “void elements” list: it also covers
basefont, bgsound, frame, keygen and param, five legacy
elements that were dropped from the authoring list but still take no end
tag and still turn up in real documents.
RAW_TEXT_ELEMENTS
html.RAW_TEXT_ELEMENTS = [...]
Elements whose text children are markup, not content: their text is written out byte for byte and never escaped, because escaping it would change what a stylesheet or a script means.
The serialization algorithm also lists noscript here, but only when
the scripting flag is enabled. This module parses with scripting
disabled by default, and in that mode a <noscript> holds real elements
rather than text, so leaving it out is right for the default and safe
for the other case: the worst that happens to a document parsed with
scripting: true is that its <noscript> text comes back escaped,
whereas including it would let hand-built text inside a <noscript>
serialize into live markup.
Classes
HierarchyError
class html.HierarchyError < Error
Raised when a tree mutation would produce a structure that cannot exist, such as inserting a node before something that is not a child of the target, or making a node its own ancestor.
Constructor
html.HierarchyError(message)
Parameters
message(string)
Node
class html.Node
The base class every node in a parsed document inherits from.
You never construct a Node directly; parsing produces Document,
Element, Text, Comment and DocumentType instances, and this
class is where the behaviour they share lives.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
node_type | One of the NODE_* constants in this module, identifying which subclass this node actually is. | |
parent_node | The node this one hangs off, or nil for a Document and for any node that has been detached. | |
children | Every child of this node in document order: elements, text nodes and comments together. |
Constructor
html.Node(node_type: number)
Parameters
node_type(number)
Node.is_element()
html.Node.is_element() -> bool
True when this node is an Element.
Returns bool
Node.is_text()
html.Node.is_text() -> bool
True when this node is a Text node.
Returns bool
Node.is_comment()
html.Node.is_comment() -> bool
True when this node is a Comment.
Returns bool
Node.is_document()
html.Node.is_document() -> bool
True when this node is a Document.
Returns bool
Node.is_doctype()
html.Node.is_doctype() -> bool
True when this node is a DocumentType.
Returns bool
Node.is_fragment()
html.Node.is_fragment() -> bool
True when this node is a DocumentFragment.
Returns bool
Node.first_child()
html.Node.first_child() -> ?Node
The first child of this node, or nil when it has none.
Returns ?Node
Node.last_child()
html.Node.last_child() -> ?Node
The last child of this node, or nil when it has none.
Returns ?Node
Node.index_in_parent()
html.Node.index_in_parent() -> number
Position of this node among its parent’s children, or -1 when it has
no parent.
Returns number
Node.next_sibling()
html.Node.next_sibling() -> ?Node
The node immediately after this one under the same parent, or nil when
this is the last child or has no parent.
Returns ?Node
Node.previous_sibling()
html.Node.previous_sibling() -> ?Node
The node immediately before this one under the same parent, or nil
when this is the first child or has no parent.
Returns ?Node
Node.next_element_sibling()
html.Node.next_element_sibling() -> ?Element
The next sibling that is an element, skipping over text and comments, or
nil when there is none.
Returns ?Element
Node.previous_element_sibling()
html.Node.previous_element_sibling() -> ?Element
The previous sibling that is an element, skipping over text and
comments, or nil when there is none.
Returns ?Element
Node.child_elements()
html.Node.child_elements() -> list
This node’s children that are elements, in document order.
A fresh list is returned on every call, so changing it does not change the tree.
Returns list
Node.first_element_child()
html.Node.first_element_child() -> ?Element
The first child that is an element, or nil.
Returns ?Element
Node.last_element_child()
html.Node.last_element_child() -> ?Element
The last child that is an element, or nil.
Returns ?Element
Node.root()
html.Node.root() -> Node
The topmost node reachable by following parent_node, which for a
parsed document is the Document itself and for a detached subtree is
that subtree’s own root.
Returns Node
Node.has_ancestor()
html.Node.has_ancestor(other: instance) -> bool
True when other is this node or one of its ancestors.
Parameters
other(Node)
Returns bool
Node.descendants()
html.Node.descendants() -> list
Every node beneath this one, in document order, not including this node itself.
Returns list
Node.walk()
html.Node.walk(callback: function)
Calls callback once for every node beneath this one, in document order. The node is passed as the only argument.
Returning false from the callback prunes that node’s subtree; any
other return value (including nil) keeps walking. The walk reads
children as it goes, so do not restructure the tree from inside the
callback.
Parameters
callback(function)
Node.append_child()
html.Node.append_child(node: instance) -> Node
Adds node as this node’s last child, detaching it from wherever it currently lives first. Returns node.
Parameters
node(Node)
Returns Node
Raises HierarchyError when node is this node or one of its
ancestors, which would make the tree cyclic.
Node.insert_before()
html.Node.insert_before(node: instance, reference: ?instance) -> Node
Inserts node immediately before reference, which must be a child of
this node. Passing nil for reference appends. Returns node.
Parameters
node(Node)reference(?Node)
Returns Node
Raises HierarchyError when reference is not a child of this
node, or when the insertion would make the tree cyclic.
Node.remove_child()
html.Node.remove_child(node: instance) -> Node
Removes node from this node’s children and returns it. The removed node keeps its own children; only its link to this parent is broken.
Parameters
node(Node)
Returns Node
Raises HierarchyError when node is not a child of this node.
Node.replace_child()
html.Node.replace_child(replacement: instance, existing: instance) -> Node
Puts replacement where existing currently sits and returns existing, now detached.
Parameters
replacement(Node)existing(Node)
Returns Node
Raises HierarchyError when existing is not a child of this node,
or when the replacement would make the tree cyclic.
Node.detach()
html.Node.detach() -> Node
Removes this node from its parent, if it has one. Returns this node so calls can be chained.
Returns Node
Node.clear_children()
html.Node.clear_children() -> Node
Removes every child of this node. The children are detached but otherwise untouched.
Returns Node
Node.text_content()
html.Node.text_content() -> string
The concatenated text of every Text node beneath this one, in document
order. Comments and doctypes contribute nothing.
Text and Comment override this to return their own data.
Unlike the browser DOM, a Document answers this the same way an
element does rather than returning nothing; being told the text of a
document you just parsed is far more useful than being told nil.
Returns string
Node.set_text_content()
html.Node.set_text_content(value: string) -> Node
Replaces every child of this node with a single Text node holding
value. Passing an empty string just empties the node, matching the
DOM.
Parameters
value(string)
Returns Node
Node.inner_html()
html.Node.inner_html() -> string
The markup of this node’s children, serialized the way the HTML fragment serialization algorithm specifies.
Returns string
Node.set_inner_html()
html.Node.set_inner_html(source: string) -> Node
Parses source as HTML in the context of this node and replaces all of its children with the result.
The parse runs the fragment parsing algorithm with this node as the
context element, so source is interpreted exactly as it would be had
it appeared inside this element in the original document. That matters:
'<td>x' keeps its cell inside a <tr> and loses it anywhere else,
which is what a browser does too.
Nothing is executed and nothing is fetched; <script> content becomes
an inert text node.
Parameters
source(string)
Returns Node
Node.outer_html()
html.Node.outer_html() -> string
The markup of this node including its own tags.
For a Document this is the whole document; for a Text node it is the
escaped text; for a Comment it is <!--...-->.
Returns string
Node.set_outer_html()
html.Node.set_outer_html(source: string) -> list
Parses source as HTML in the context of this node’s parent and puts the result where this node currently sits.
Returns the list of nodes that replaced this one, which may be empty when source produces nothing. This node is detached either way.
Parameters
source(string)
Returns list
Raises HierarchyError when this node has no parent, since there
would be nowhere to put the result.
Node.serialize_into()
html.Node.serialize_into(out: list)
Appends this node’s markup to out, a list of string pieces the caller
is expected to join().
outer_html() is the friendly form of this and is what you normally
want. Reach for this one when you are writing a large document out in
pieces and would rather not build the whole string in memory first.
Every node class overrides it; the base implementation writes nothing.
Parameters
out(list)
Node.clone_node()
html.Node.clone_node(deep: ?bool) -> Node
A deep or shallow copy of this node, detached from any parent.
The copy is shallow by default, matching the DOM: you get the node
itself with its attributes but with no children. Pass true for deep
to copy the whole subtree.
Parameters
deep(?bool)
Returns Node
Node.get_element_by_id()
html.Node.get_element_by_id(value: string) -> ?Element
The first element beneath this node whose id attribute is value, or
nil when there is none.
Ids are compared exactly, including case: HTML does not fold the case of id values even though it folds tag names.
A document with duplicate ids is malformed but perfectly parseable, and this returns the first one in document order, which is what browsers settled on.
Parameters
value(string)
Returns ?Element
Node.get_elements_by_class_name()
html.Node.get_elements_by_class_name(names) -> list
Every element beneath this node carrying all of the classes in names, in document order.
names may be a single class name, several separated by whitespace, or a list of names; in every form an element must carry all of them to match. Class names are compared exactly, including case. An empty names matches nothing.
%> doc.get_elements_by_class_name('warning')
%> doc.get_elements_by_class_name('warning sticky')
%> doc.get_elements_by_class_name(['warning', 'sticky'])
Parameters
names(string|list)
Returns list
Node.get_elements_by_tag_name()
html.Node.get_elements_by_tag_name(name: string) -> list
Every element beneath this node whose tag name is name, in document order.
The comparison folds ASCII case, so 'DIV' and 'div' find the same
elements. Passing '*' returns every element.
Parameters
name(string)
Returns list
Node.query_selector()
html.Node.query_selector(query: string) -> ?Element
The first element beneath this node matching the CSS selector query,
in document order, or nil when nothing matches.
The supported selector syntax is described in html.selector.
Parameters
query(string)
Returns ?Element
Raises SelectorError when query is not valid.
Node.query_selector_all()
html.Node.query_selector_all(query: string) -> list
Every element beneath this node matching the CSS selector query, in document order.
Parameters
query(string)
Returns list
Raises SelectorError when query is not valid.
Node.to_string()
html.Node.to_string() -> string
Renders this node as markup, the same as outer_html().
Returns string
Text
class html.Text < Node
A run of character data in the tree.
The parser merges adjacent characters into as few Text nodes as the
tree construction algorithm allows, so a paragraph of prose is normally
one node rather than one per character.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
data | The characters this node holds, already decoded: character references were resolved during parsing, so data… |
Constructor
html.Text(data: string)
Parameters
data(string)
Text.text_content()
html.Text.text_content() -> string
The characters this node holds.
Returns string
Text.set_text_content()
html.Text.set_text_content(value: string) -> Text
Replaces this node’s characters with value.
Parameters
value(string)
Returns Text
Text.append_data()
html.Text.append_data(value: string) -> Text
Appends value to this node’s characters.
Parameters
value(string)
Returns Text
Text.serialize_into()
html.Text.serialize_into(out: list)
Parameters
out(list)
Comment
class html.Comment < Node
An HTML comment.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
data | The text between <!-- and -->, exactly as it appeared. |
Constructor
html.Comment(data: string)
Parameters
data(string)
Comment.text_content()
html.Comment.text_content() -> string
The comment’s text.
Returns string
Comment.set_text_content()
html.Comment.set_text_content(value: string) -> Comment
Replaces the comment’s text with value.
Parameters
value(string)
Returns Comment
Comment.serialize_into()
html.Comment.serialize_into(out: list)
Parameters
out(list)
DocumentType
class html.DocumentType < Node
The <!DOCTYPE ...> node at the top of a document.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
name | The doctype’s name, lowercased by the tokenizer. | |
public_id | The public identifier, or an empty string when the doctype has none. | |
system_id | The system identifier, or an empty string when the doctype has none. |
Constructor
html.DocumentType(name: string, public_id: ?string, system_id: ?string)
Parameters
name(string)public_id(?string)system_id(?string)
DocumentType.text_content()
html.DocumentType.text_content() -> string
A doctype contains no text, so this is always an empty string.
Returns string
DocumentType.serialize_into()
html.DocumentType.serialize_into(out: list)
Parameters
out(list)
DocumentFragment
class html.DocumentFragment < Node
A parentless container for a run of nodes.
Fragment parsing produces one of these, and every <template> element
owns one holding its content.
- printable — has a
@to_string(), soechoandprint()show something useful
Constructor
html.DocumentFragment()
DocumentFragment.serialize_into()
html.DocumentFragment.serialize_into(out: list)
Parameters
out(list)
Document
class html.Document < Node
A whole parsed document.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
mode | Which quirks mode the doctype (or the lack of one) put this document in: 'no-quirks', 'limited-quirks' or… | |
errors | Every parse error the tokenizer and tree builder raised, in the order they happened. |
Constructor
html.Document()
Document.document_element()
html.Document.document_element() -> ?Element
The document’s root element, normally <html>, or nil for an empty
document.
Returns ?Element
Document.doctype()
html.Document.doctype() -> ?DocumentType
The document’s <!DOCTYPE> node, or nil when it has none.
Returns ?DocumentType
Document.head()
html.Document.head() -> ?Element
The document’s <head> element, or nil.
The tree construction algorithm always creates one when parsing a whole
document, even for input with no <head> tag in it, so this only
returns nil for a document that was built by hand.
Returns ?Element
Document.body()
html.Document.body() -> ?Element
The document’s <body> element, or the <frameset> that stands in for
it in a frameset document, or nil when there is neither.
Returns ?Element
Document.title()
html.Document.title() -> string
The text of the document’s <title> element, with leading and trailing
whitespace removed, or an empty string when there is no title.
Returns string
Document.serialize_into()
html.Document.serialize_into(out: list)
Parameters
out(list)
ClassList
class html.ClassList
A live view of one element’s class attribute.
Every method reads and writes the attribute itself, so a class_list
never goes stale and two of them taken from the same element always
agree.
%> var el = doc.query_selector('div')
%> el.class_list().add('active')
%> el.class_list().contains('active')
true
%> el.get_attribute('class')
'active'
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
element | The element this list reads from and writes to. |
Constructor
html.ClassList(element)
Parameters
element(Element)
ClassList.to_list()
html.ClassList.to_list() -> list
The class names currently on the element, in source order, with duplicates removed.
Returns list
ClassList.length()
html.ClassList.length() -> number
How many distinct classes the element carries.
Returns number
ClassList.contains()
html.ClassList.contains(name: string) -> bool
True when the element carries name. The comparison is exact: HTML class names are case-sensitive in standards mode.
Parameters
name(string)
Returns bool
ClassList.add()
html.ClassList.add(...names: list) -> ClassList
Adds every name given to the element’s class attribute, ignoring any it already has. Returns the list so calls can be chained.
An empty name, or one containing whitespace, is rejected: those cannot be written to a class attribute without changing what it means.
Parameters
names(...string)
Returns ClassList
Raises ArgumentError
ClassList.remove()
html.ClassList.remove(...names: list) -> ClassList
Removes every name given from the element’s class attribute, ignoring any it does not have. Returns the list.
Parameters
names(...string)
Returns ClassList
Raises ArgumentError
ClassList.toggle()
html.ClassList.toggle(name: string, force: ?bool) -> bool
Adds name when the element does not have it and removes it when it
does, then returns true when the class is present afterwards.
Passing force turns this into an unconditional add (true) or remove
(false), which is handy when the desired state is already in a
variable.
Parameters
name(string)force(?bool)
Returns bool
Raises ArgumentError
ClassList.to_string()
html.ClassList.to_string() -> string
The class attribute’s value as it would be written out.
Returns string
Element
class html.Element < Node
An element in the tree.
- printable — has a
@to_string(), soechoandprint()show something useful
Fields
| Field | Type | Description |
|---|---|---|
tag_name | The element’s tag name. | |
namespace | The namespace this element lives in: one of HTML_NAMESPACE, SVG_NAMESPACE or MATHML_NAMESPACE. | |
attributes | The element’s attributes, in source order, as a dictionary of name to value. | |
attribute_namespaces | Namespace URIs for the few attributes that have one, keyed by the same qualified name used in attributes. | |
source_line | The line in the source where this element’s start tag began, counting from 1, or 0 for an element that was… | |
source_column | The column in the source where this element’s start tag began, counting from 1, or 0 for an element that… | |
content | For a <template> element, the DocumentFragment holding its content. |
Constructor
html.Element(tag_name: string, namespace: ?string, attributes: ?dict)
Parameters
tag_name(string)namespace(?string)attributes(?dict)
Element.is_html()
html.Element.is_html() -> bool
True when this is an HTML element, as opposed to one inside an <svg>
or <math> subtree.
Returns bool
Element.is_void()
html.Element.is_void() -> bool
True when this element is one of the HTML void elements, which have no closing tag and can hold no content.
Returns bool
Element.get_attribute()
html.Element.get_attribute(name: string) -> ?string
The value of the attribute name, or nil when the element does not
carry it.
The lookup folds ASCII case for HTML elements, so
get_attribute('HREF') and get_attribute('href') are the same
question. For foreign elements the name is matched exactly, because SVG
really does distinguish viewBox from viewbox.
Parameters
name(string)
Returns ?string
Element.has_attribute()
html.Element.has_attribute(name: string) -> bool
True when the element carries the attribute name, whatever its value.
Parameters
name(string)
Returns bool
Element.set_attribute()
html.Element.set_attribute(name: string, value: string) -> Element
Sets the attribute name to value, replacing any existing value and leaving the attribute’s position among the others unchanged. A new attribute is appended after the existing ones.
Pass an empty string for a valueless attribute such as disabled; there
is no separate “no value” state.
Parameters
name(string)value(string)
Returns Element
Raises ArgumentError when name is empty or contains a character
that cannot appear in an attribute name.
Element.remove_attribute()
html.Element.remove_attribute(name: string) -> bool
Removes the attribute name and returns true when it was actually
there.
Parameters
name(string)
Returns bool
Element.attribute_names()
html.Element.attribute_names() -> list
The attribute names this element carries, in source order.
Returns list
Element.attribute_namespace()
html.Element.attribute_namespace(name: string) -> ?string
The namespace URI of the attribute name, or nil when it has none.
Only foreign content produces namespaced attributes.
Parameters
name(string)
Returns ?string
Element.set_attribute_namespace()
html.Element.set_attribute_namespace(name: string, uri: string) -> Element
Records that the attribute name belongs to uri. Called by the parser while adjusting foreign attributes; there is no reason to call it by hand.
Parameters
name(string)uri(string)
Returns Element
Element.id()
html.Element.id() -> string
The element’s id, or an empty string when it has none.
Returns string
Element.class_name()
html.Element.class_name() -> string
The element’s class attribute verbatim, or an empty string.
Returns string
Element.class_names()
html.Element.class_names() -> list
The element’s classes as a list, in source order and without duplicates.
An element with no class attribute gives an empty list.
Returns list
Element.class_list()
html.Element.class_list() -> ClassList
A ClassList view of this element’s class attribute, for adding,
removing and toggling classes.
A fresh view is returned on every call; because it reads and writes the attribute directly there is no state to keep, and two views of the same element behave identically.
Returns ClassList
Element.matches()
html.Element.matches(query: string) -> bool
True when this element matches the CSS selector query.
Parameters
query(string)
Returns bool
Raises SelectorError when query is not valid.
Element.closest()
html.Element.closest(query: string) -> ?Element
This element, or its nearest ancestor, matching the CSS selector
query: or nil when neither this element nor any ancestor does.
The search starts at the element itself, which is what makes
el.closest('div') return el when el is a div.
Parameters
query(string)
Returns ?Element
Raises SelectorError when query is not valid.
Element.serialize_into()
html.Element.serialize_into(out: list)
Parameters
out(list)
Element.inner_html()
html.Element.inner_html() -> string
The markup of this element’s children.
A <template> answers with its content fragment rather than with its
own children, which are always empty. That matches the DOM, and it is
the only way the question has a useful answer for a template.
Returns string
Element.set_inner_html()
html.Element.set_inner_html(source: string) -> Element
Parses source and replaces this element’s children with the result.
For a <template> the result goes into the content fragment, matching
where inner_html() reads from.
Parameters
source(string)
Returns Element
Element.clone_node()
html.Element.clone_node(deep: ?bool) -> Element
A deep or shallow copy of this element. Deep-cloning a <template> also
copies its content fragment, which is otherwise not part of the
element’s children.
Parameters
deep(?bool)
Returns Element
2026, Richard Ore and The Zuri Contributors