html.serialize
import html.serialize
htmllifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelledhtml.serialize.*needsimport html.serialize.
Writing a document back out, shaped for whoever has to read it next:
minify() for a browser, format() for a person.
Neither of these is the same as outer_html(). That one is exact and
round-trips a tree byte for byte; these two deliberately change
whitespace, which is the whole point of both.
%> import html
%> html.minify('<p> a b </p>\n<p>c</p>', { fragment: true })
'<p>a b</p><p>c</p>'
%> echo html.format('<ul><li>a<li>b</ul>', { fragment: true })
<ul>
<li>a</li>
<li>b</li>
</ul>
Constants
INLINE_ELEMENTS
html.serialize.INLINE_ELEMENTS = [...]
Elements that flow inside a line of text rather than starting a new block. Whitespace around these carries meaning, so neither serializer moves it.
This is the default rendering of each element. CSS can turn any
element into any box type, and neither serializer reads CSS, so a
stylesheet that makes a <div> inline is a case where format() can
change how the page looks. minify() is safe either way: it only ever
collapses runs of whitespace to a single space, which both box types
render identically.
PRESERVE_WHITESPACE
html.serialize.PRESERVE_WHITESPACE = [...]
Elements whose text content is significant to the last character. Nothing inside one of these is ever reflowed, re-indented or collapsed.
Functions
minify()
html.minify(source, options: ?dict) -> string
Removes the markup a browser does not need and returns the result.
source is either a node you already have or a string of HTML. A string
is parsed as a complete document, so <html>, <head> and <body>
appear in the output even when the input had none of them, exactly as a
browser would produce them. Pass { fragment: true } to parse and
return just the markup you gave it instead.
What this changes:
- Runs of whitespace in text become a single space.
- Whitespace-only text between two block level elements is removed entirely, along with leading and trailing whitespace inside a block level element.
- Comments are removed.
What this never touches:
- The content of
<pre>,<textarea>,<script>,<style>,<xmp>,<listing>and<plaintext>. - Whitespace next to an inline element, which is a word separator and cannot be removed without changing the text.
- Attribute values, attribute order, tag names, or the doctype.
options:
fragment(bool, defaultfalse): parse a string source as a fragment rather than a whole document. Ignored when source is already a node.comments(bool, defaultfalse): keep comments. Turn this on when the markup carries conditional comments or a licence header that has to survive.collapse_whitespace(bool, defaulttrue): do the whitespace work at all. Withfalsethe only thing minifying does is drop comments.
The tree passed in is never modified; the work happens on a copy.
%> import html
%> html.minify('<div> <p>a b</p> <!-- x --> </div>', { fragment: true })
'<div><p>a b</p></div>'
%> html.minify('<p>a <b> b </b> c</p>', { fragment: true })
'<p>a <b> b </b> c</p>'
Parameters
source(Node|string)options(?dict)
Returns string
format()
html.format(source, options: ?dict) -> string
Renders source as indented, readable HTML.
This is a pretty printer, not a round-trip. It moves whitespace between
elements to make the structure visible, so the output is equivalent for
reading and reviewing but is not byte-identical to the input. Use
outer_html() when the exact markup matters.
Whitespace inside inline content is left alone, because collapsing it
would change the rendered text; only the gaps between block level
elements are re-indented. The content of <pre>, <textarea>,
<script>, <style>, <xmp>, <listing> and <plaintext> is copied
through untouched.
options:
indent(number or string, default2): a number is that many spaces per level; a string is used as the unit verbatim, so'\t'gives tab indentation.fragment(bool, defaultfalse): parse a string source as a fragment rather than a whole document.comments(bool, defaulttrue): keep comments.
%> import html
%> echo html.format('<div><p>Hi <b>there</b></p></div>', { fragment: true })
<div>
<p>Hi <b>there</b></p>
</div>
%> echo html.format('<div><p>x</p></div>', { fragment: true, indent: '\t' })
<div>
<p>x</p>
</div>
Parameters
source(Node|string)options(?dict)
Returns string
2026, Richard Ore and The Zuri Contributors