Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

html.serialize

import html.serialize

html lifts part of this module out to its own top level; each name below is shown with the path that reaches it. Anything still spelled html.serialize.* needs import html.serialize.

Writing a document back out, shaped for whoever has to read it next: minify() for a browser, format() for a person.

Neither of these is the same as outer_html(). That one is exact and round-trips a tree byte for byte; these two deliberately change whitespace, which is the whole point of both.

%> import html
%> html.minify('<p>  a   b  </p>\n<p>c</p>', { fragment: true })
'<p>a b</p><p>c</p>'
%> echo html.format('<ul><li>a<li>b</ul>', { fragment: true })
<ul>
  <li>a</li>
  <li>b</li>
</ul>

Constants

INLINE_ELEMENTS

html.serialize.INLINE_ELEMENTS = [...]

Elements that flow inside a line of text rather than starting a new block. Whitespace around these carries meaning, so neither serializer moves it.

This is the default rendering of each element. CSS can turn any element into any box type, and neither serializer reads CSS, so a stylesheet that makes a <div> inline is a case where format() can change how the page looks. minify() is safe either way: it only ever collapses runs of whitespace to a single space, which both box types render identically.

PRESERVE_WHITESPACE

html.serialize.PRESERVE_WHITESPACE = [...]

Elements whose text content is significant to the last character. Nothing inside one of these is ever reflowed, re-indented or collapsed.

Functions

minify()

html.minify(source, options: ?dict) -> string

Removes the markup a browser does not need and returns the result.

source is either a node you already have or a string of HTML. A string is parsed as a complete document, so <html>, <head> and <body> appear in the output even when the input had none of them, exactly as a browser would produce them. Pass { fragment: true } to parse and return just the markup you gave it instead.

What this changes:

  • Runs of whitespace in text become a single space.
  • Whitespace-only text between two block level elements is removed entirely, along with leading and trailing whitespace inside a block level element.
  • Comments are removed.

What this never touches:

  • The content of <pre>, <textarea>, <script>, <style>, <xmp>, <listing> and <plaintext>.
  • Whitespace next to an inline element, which is a word separator and cannot be removed without changing the text.
  • Attribute values, attribute order, tag names, or the doctype.

options:

  • fragment (bool, default false): parse a string source as a fragment rather than a whole document. Ignored when source is already a node.
  • comments (bool, default false): keep comments. Turn this on when the markup carries conditional comments or a licence header that has to survive.
  • collapse_whitespace (bool, default true): do the whitespace work at all. With false the only thing minifying does is drop comments.

The tree passed in is never modified; the work happens on a copy.

%> import html
%> html.minify('<div>  <p>a   b</p>  <!-- x --> </div>', { fragment: true })
'<div><p>a b</p></div>'
%> html.minify('<p>a <b> b </b> c</p>', { fragment: true })
'<p>a <b> b </b> c</p>'

Parameters

  • source (Node|string)
  • options (?dict)

Returns string

format()

html.format(source, options: ?dict) -> string

Renders source as indented, readable HTML.

This is a pretty printer, not a round-trip. It moves whitespace between elements to make the structure visible, so the output is equivalent for reading and reviewing but is not byte-identical to the input. Use outer_html() when the exact markup matters.

Whitespace inside inline content is left alone, because collapsing it would change the rendered text; only the gaps between block level elements are re-indented. The content of <pre>, <textarea>, <script>, <style>, <xmp>, <listing> and <plaintext> is copied through untouched.

options:

  • indent (number or string, default 2): a number is that many spaces per level; a string is used as the unit verbatim, so '\t' gives tab indentation.
  • fragment (bool, default false): parse a string source as a fragment rather than a whole document.
  • comments (bool, default true): keep comments.
%> import html
%> echo html.format('<div><p>Hi <b>there</b></p></div>', { fragment: true })
<div>
  <p>Hi <b>there</b></p>
</div>
%> echo html.format('<div><p>x</p></div>', { fragment: true, indent: '\t' })
<div>
	<p>x</p>
</div>

Parameters

  • source (Node|string)
  • options (?dict)

Returns string


2026, Richard Ore and The Zuri Contributors