Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The Zuri Programming Language

by the Zuri Project

Zuri is a language for building whole applications, and it ships with everything that takes. The runtime, the web server, the templating engine, the data formats, the cryptography and the toolchain were designed together and are delivered as one binary, so starting a real program does not begin with assembling an ecosystem out of third-party parts. Zuri has first-class language-level support for package management and vendoring, so the code you bring in from outside needs no third-party tooling either.

The standard library ships with the language, and it is large. Templating, HTTP/1.1 and HTTP/2, WebSockets, TLS, JSON, YAML, CSV, compression, cryptography, image decoding and drawing, HTML parsing, date arithmetic with the IANA time zone database, and OS-thread concurrency are all part of the installation. Every one of them is reachable with a bare import.

The tools ship in the same binary. zuri fmt lays out your code, zuri test runs your tests, and zuri lsp gives every editor completion, navigation and diagnostics. zuri install and zuri publish manage packages against Nyssa, a package registry any installation can serve, and zuri bundle ships a finished program to machines where Zuri is not installed, as a single executable if you like.

Zuri is dynamically typed, and it declines to be vague. A declared parameter type is enforced at the call, classes are sealed, a function cannot be quietly redefined, and a name beginning with an underscore stays inside the module that declared it. Your code runs as bytecode until a piece of it gets hot, and a JIT compiles that piece to machine code while the program is still running.

This book teaches the language from the first line of code to a complete web application. It is written against the Rust implementation of Zuri. Every program printed in these pages was run against that implementation, and the output shown underneath each one is the output it produced.

Introduction

Welcome. Open a terminal, keep it next to this page, and let’s begin.

This book teaches Zuri. It starts with installing the language and printing a line of text, and it ends with a task-board web application that stores data on disk, renders HTML from templates, serves a JSON API, and runs behind a stack of middleware. Everything in between is the road from one to the other.

What Zuri Looks Like

Here is a complete program. You do not need to understand all of it yet; read it the way you would read a paragraph in a language you are learning, and see how much comes through.

class Account {

  @new(owner: string, balance: number) {
    self.owner = owner
    self.balance = balance
  }

  deposit(amount: number) {
    if amount <= 0 {
      raise ValueError('deposit must be positive')
    }

    self.balance += amount
    return self.balance
  }

  to_string() {
    return '${self.owner}: ${self.balance}'
  }
}

var accounts = [
  Account('ada', 100),
  Account('grace', 250),
]

for account in accounts {
  account.deposit(50)
  echo account.to_string()
}
ada: 150
grace: 300

Most of that will be familiar if you have written code before. The pieces worth pointing at now, because they come up on the very first page of Chapter 3:

  • var declares a variable, and class declares a class.
  • A constructor is called @new. Methods whose names begin with @ hook into the language’s own syntax, and Chapter 6 covers the full set.
  • self refers to the instance, and reading a field always goes through it: self.balance, never a bare balance.
  • '${...}' interpolates an expression into a string.
  • echo prints a value and a newline.
  • raise signals an error; Chapter 7 shows how to catch one.
  • Every block is written with braces, and every control-flow statement requires one.

What Comes With Zuri

Installing Zuri installs its whole toolchain with it. Every tool is a subcommand of zuri, and every one has its place in this book:

CommandWhat it doesWhere
zurithe interactive REPLChapter 1
zuri runruns a script, a package or a directoryChapter 1
zuri initstarts a projectChapter 1
zuri fmtlays out Zuri source in the house styleChapter 1
zuri testruns a project’s testsChapter 23
zuri lspthe language server behind your editorChapter 29
zuri installadds packages, resolved into a lockfileChapter 27
zuri publishshares a package on a registryChapter 27
zuri serveruns Nyssa, a package registry of your ownChapter 27
zuri bundleships a program to machines without ZuriChapter 27
zuri upgrademoves the installation to a newer releaseChapter 27

A project adds commands of its own the same way, and so can a package; Appendix I shows how.

Who This Book Is For

Chapters 1 through 6 assume you can open a terminal and nothing else. If this is your first programming language, start at the beginning and go slowly. Every idea is introduced before it is used, and every example is short enough to type out.

If you already write Python, JavaScript, Ruby, Go or Java, you can move through Chapters 3 to 5 quickly. Read Appendix H first: it lists the places where Zuri does something different from what the same syntax does in the language you already know, which is where the hours get lost.

How This Book Is Organised

Getting started, Chapters 1 and 2. Install the language, run something, start a project and format it, then build a small command-line application end to end so you have seen the shape of a real program before we take one apart.

The language, Chapters 3 to 8. Variables, types, operators, control flow, strings, numbers, collections, functions, closures, type annotations, classes, inheritance, decorated methods, errors and modules. Read these in order.

The world outside your program, Chapters 9 to 13. Files, binary data and byte streams, isolates and concurrency, sockets and networking, and a tour of the standard library.

The large modules, Chapters 14 to 20. Wire for templating, HTTP for clients and servers, Imagine for images, SQL for databases, Mail for the three protocols that move messages, env for the configuration all of them read, and args for the command line they are started from. Each is big enough to need a chapter of its own, and each is a reference you will come back to.

Depth, Chapters 21 to 24. Reflection and the compiler API, how Zuri executes your code and how to make it faster, how to test it, and how to debug a program that is doing something you did not expect.

The capstone, Chapter 25. One application, built in seven steps, using almost everything the book has covered.

Interoperability, Chapter 26. Calling C and Rust libraries, and being called back by them.

Sharing and shipping, Chapter 27. Installing and publishing packages, the lockfile, packaging a program to run where Zuri is not installed, and running Nyssa, the package repository, for a team or the public.

Talking to other programs, Chapter 28. JSON-RPC, for calling methods on another program and answering its calls over any connection.

Tooling, Chapter 29. The language server, and setting up your editor to use it.

Appendices. Keywords, operators and precedence, decorated methods, built-in functions, every method on every built-in type, the standard library index, the error hierarchy, the notes for readers arriving from another language, and writing commands of your own. These are reference material; the rest of the book is prose.

How to Read It

Read it with the interpreter running. Zuri has a REPL, every short example in this book can be pasted straight into it, and the fastest way to understand a rule is to break it on purpose and read the error.

Chapters build on each other. When a chapter needs something from later in the book, it says so and links to it, and you can carry on without following the link. When a chapter introduces something you will need again, it says that too.

Conventions

Commands you type in a shell appear with a $ prompt, and the output follows underneath:

$ zuri run main.zu
Hello, world!

Code that belongs in a file appears with the filename above it when the filename matters:

Filename: main.zu

echo 'Hello, world!'

Sessions in the interactive prompt use %> for the first line of an input and .. for its continuations, which is exactly what the REPL itself prints:

%> var name = 'zuri'
%> name.upper()
'ZURI'

Where a rule has an exception, the exception is stated in the same paragraph as the rule. Where a limit exists, it is stated as a limit. Nothing in this book is a guess about how the language behaves; every claim was checked by running it.

Let’s get started.

Getting Started

Three things before the language itself: putting Zuri on your machine, running a file, and using the interactive prompt.

The prompt matters more than it sounds. Zuri prints the value of any expression you type at it, which makes it the fastest way to answer “what does this actually return?” — and that question comes up constantly in the next five chapters.

Installation

Zuri installs with one command, which downloads the release for your machine, checks it, and puts zuri on your PATH. Releases cover:

  • Linux on x86-64 and ARM (x86_64-unknown-linux-gnu, aarch64-unknown-linux-gnu)
  • macOS on Intel and Apple silicon (x86_64-apple-darwin, aarch64-apple-darwin)
  • Windows on x86-64 (x86_64-pc-windows-msvc)

Zuri also builds from source with Cargo, which the end of this page covers.

Linux and macOS

$ curl -fsSL https://zuri-lang.github.io/zuri/install.sh | sh
Downloading zuri 0.1.0 for x86_64-unknown-linux-gnu
Installed zuri 0.1.0 in /home/ada/.zuri/runtime
Added /home/ada/.zuri/bin to PATH in /home/ada/.bashrc

Open a new terminal, or run this one to use zuri straight away:

  export PATH="$HOME/.zuri/bin:$PATH"

Get started with:

  zuri --help

The installer downloads the archive for your machine and checks it against the checksum published beside it. It then runs the new zuri once to confirm the version it reports, and only after that moves it into ~/.zuri/runtime, so a failed or tampered download never replaces a working installation. zuri is linked into ~/.zuri/bin, the same directory that globally installed tools put their launchers in, and that one directory goes on your PATH through your shell’s startup file: .zshrc for zsh, .bashrc for bash (.bash_profile on macOS), config.fish for fish, and .profile for anything else.

Options go after sh -s --:

$ curl -fsSL https://zuri-lang.github.io/zuri/install.sh | sh -s -- --version 0.2.0
OptionWhat it does
--version <version>installs that release rather than the newest
--prereleaseconsiders pre-releases when picking the newest
--no-modify-pathleaves your shell’s startup files alone

Windows

In PowerShell:

> powershell -c "irm https://zuri-lang.github.io/zuri/install.ps1 | iex"

The Windows installer makes the same checks and installs into %USERPROFILE%\.zuri\runtime. That directory goes on your user PATH, and so does %USERPROFILE%\.zuri\bin, where globally installed tools put their launchers. A terminal opened afterwards has both; the one the installer ran in has them already.

The same options are parameters, given by running the installer as a script block:

> & ([scriptblock]::Create((irm https://zuri-lang.github.io/zuri/install.ps1))) -Version 0.2.0

-Version, -Prerelease and -NoModifyPath match the options above.

Checking the Install

$ zuri
Zuri 0.1.0 (running on ZuriVM 0.1.0), REPL/Interactive mode = ON
Build No. => 2026-09-10 23:19:10 UTC
Type ".exit" to quit, ".help" for help or ".credits" for more information
%>

That %> is the Zuri prompt. Type .exit to leave.

zuri --version reports the same build without opening a session, which is the one to reach for from a script or a CI job, and zuri --help adds everything the runtime can run from here:

$ zuri --version
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC

Where It Goes

Both installers keep everything under ZURI_HOME, which is .zuri in your home directory unless the variable says otherwise:

~/.zuri/
  runtime/        zuri, with libs and cmds beside it
  bin/            the zuri link, and globally installed tools

runtime holds the executable with copies of the libs and cmds directories beside it. That pairing matters: the standard library is written mostly in Zuri itself, and the runtime finds it by looking for a libs directory beside the executable. cmds is the same arrangement for the commands the runtime ships.

Running an installer again installs over the top. zuri upgrade does the same from inside an installation, and Bundles and Upgrades covers it. Removing Zuri is deleting ~/.zuri/runtime and the zuri link in ~/.zuri/bin, along with the line the installer added to your shell’s startup file.

ZURI_RELEASES_URL points both installers, and zuri upgrade, at another listing of releases in the shape GitHub publishes them, such as a mirror. GITHUB_TOKEN, when set, is sent to GitHub’s API, where it raises the rate limit.

Downloading by Hand

Every release on the releases page carries an archive per platform, named zuri-<version>-<target>.tar.gz, or .zip for Windows, with a .sha256 beside it. Check the archive against it, unpack it, and put the zuri it holds on your PATH, keeping the libs and cmds directories beside it.

Building From Source

Building needs Rust’s package manager, Cargo. If you do not have Rust installed, get it from rustup.rs; it takes a minute and installs cargo for you.

$ git clone https://github.com/zuri-lang/zuri
$ cd zuri
$ cargo patch-crates
$ cargo build --release

cargo patch-crates prepares the few dependencies Zuri builds with changes of its own. It runs once per clone, and again only when one of those changes is updated; the build says when that is.

The first build compiles a large set of native dependencies, so it takes a few minutes. Builds after that are incremental and quick.

When it finishes you have an executable at target/release/zuri, and, sitting right next to it, copies of the libs/ and cmds/ directories, in the same arrangement a release has.

Putting a Build on Your PATH

The simplest arrangement is to keep the binary and its two directories together and symlink to the binary:

$ sudo ln -s "$PWD/target/release/zuri" /usr/local/bin/zuri

A symlink is resolved before the runtime looks for libs, so this works. Copying just the binary does not.

If you want them somewhere else entirely, set ZURI_ROOT to the directory that contains them:

$ export ZURI_ROOT=/opt/zuri
$ ls /opt/zuri
cmds  libs

ZURI_ROOT wins over the executable-adjacent lookup, which makes it handy when you are hacking on the standard library itself and want a build to pick up your edits immediately:

$ ZURI_ROOT=$PWD ./target/debug/zuri run myscript.zu

A Debug Build, and Why You Might Want One

Plain cargo build produces target/debug/zuri. It keeps the runtime assertions that the optimised build strips out, which makes it the build to reach for when a program is doing something you did not expect. It runs more slowly in exchange. Either build runs everything in this book.

Hello, World!

Make a directory, put one file in it, and run it.

$ mkdir hello
$ cd hello

Create a file called main.zu. The .zu extension is what the runtime looks for when resolving imports, so get in the habit early.

Filename: main.zu

echo 'Hello, world!'

Run it:

$ zuri run main.zu
Hello, world!

That is the whole program. No main function, no imports, no boilerplate. A Zuri file is a script, and its top level is code that runs.

Anatomy of One Line

echo 'Hello, world!'

echo is a keyword, not a function. You do not write echo(...), you write echo followed by an expression. It prints the value and adds a newline.

Strings are written between single or double quotes, and the two are identical in meaning. Most Zuri code uses single quotes and saves double quotes for strings that contain an apostrophe.

There is no semicolon. Zuri ends a statement at the end of the line. You can write a semicolon if you want two statements on one line, but almost nobody does.

echo and print

There is also a print() function, and the difference is worth learning now because it bites people later.

echo 'first'
echo 'second'

print('third')
print('fourth\n')
first
second
thirdfourth

echo appends a newline; print() does not, and it takes any number of arguments. Use echo for output meant for a human reading a terminal, and print() when you are assembling output character by character.

Running a Directory

run takes a directory as readily as a file. Point it at one and it looks for index.zu inside and runs that:

$ zuri run hello
Hello, world!

If there is no index.zu, you get told so plainly:

$ zuri run hello
(Zuri):
  Launch aborted for hello
  Reason: No entrypoint found in the directory

This is the same rule the module system uses for packages, which we will get to in Chapter 8. A directory with an index.zu is a unit you can run or import.

Leave the path off and run launches the directory you are standing in, which is how a project is usually started:

$ cd hello
$ zuri run
Hello, world!

Everything after the path belongs to the program rather than to zuri, so a script reads its own flags exactly as it would anywhere else:

$ zuri run main.zu --name Ada --verbose

Chapter 20 covers reading them.

Starting a Project

One file is where everybody starts, and it stops being enough about as soon as there are two of them. zuri init writes the layout a project grows into:

$ zuri init myproject

Run in a terminal, it asks what the project is called, what version it starts at and who wrote it, offering an answer to each. Press enter to take what it offers. Run anywhere there is nobody to ask — a script, a CI job — it takes those answers itself, which --yes also says outright.

What it leaves behind:

myproject/
  project.toml          what the project is
  index.zu              the entry point
  app/
    index.zu            the application
  tests/
    app.test.zu         its tests
  README.md
  .gitignore
  .gitattributes

It runs, and its tests pass, before you have written anything:

$ cd myproject
$ zuri run
Hello, world!
$ zuri test

  zuri test  1 file in tests

   PASS   app.test.zu  36ms  2 tests

  1 file
  2 passed  •  2 total

The root index.zu is two things at once:

Filename: index.zu

import @.app { * }

if __root__ == __file__ {
  main()
}

The import re-exports everything app declares, so another project can import myproject and reach all of it. The if starts the application, but only when this file is the one that was run: __root__ is the file zuri was pointed at, and __file__ is the file the line is written in. Importing the project leaves main() alone. That is the Zuri spelling of a main guard, and it is worth knowing early because every runnable package uses it.

app/index.zu is the application itself. Whatever it exports, the project exports:

Filename: app/index.zu

def greet(name: ?string) {
  return 'Hello, ${name or "world"}!'
}

def main() {
  echo greet()
}

project.toml is what the project is, for anything that reads it:

[project]
name = "myproject"
version = "0.1.0"
description = ""
authors = ["Ada Lovelace <ada@example.com>"]
license = "MIT"

[dependencies]

The name comes from the directory, the author from git config, and a git repository is created unless there is already one above. zuri init --help has the rest, and every question it asks has a flag that answers it:

$ zuri init myproject --name web-server --license Apache-2.0 --yes

Run inside a directory that already has files, init adds what is missing and writes over nothing. Run where a project already is, it refuses rather than overwrite one.

Chapter 8 is where packages and re-exports are covered properly, and Chapter 25 grows this layout into a full application.

Formatting

zuri fmt lays out Zuri source in one house style: two spaces to a level, one statement to a line, an explicit block wherever one belongs, spacing the way a reader expects it, and lines past eighty columns broken where the expression holds together least. Line breaks you made yourself are kept, so a list you laid out one entry to a line stays that way.

Here is a file written in a hurry:

Filename: rooms.zu

def area(width,height){return width*height}
var rooms=[ [3,4],[5,6] ]
for room in rooms { echo area(room[0],room[1]) }
$ zuri fmt rooms.zu
formatted rooms.zu
Formatted 1 file.

Filename: rooms.zu

def area(width, height) {
  return width * height
}

var rooms = [[3, 4], [5, 6]]
for room in rooms {
  echo area(room[0], room[1])
}

Given a directory, zuri fmt formats every .zu file beneath it apart from the packages a project has installed, and with no path at all it formats the current directory. It never changes what a program does: before it writes a file, it parses its own result and checks that the program and every comment in it are exactly what it started with, and a file it cannot format that way is left as it was.

--dry-run writes nothing. It names each file that would change, and exits with status 1 when any would, which is the check a commit hook or a CI job wants:

$ zuri fmt --dry-run
would format /home/ada/myproject/app/extra.zu
1 file would change. 3 files were already formatted.

Your editor runs the same formatter through the language server, on a whole file or a selection; Chapter 29 sets that up.

Commands

A first word that is not run names a command instead of a path:

$ zuri greet Ada
Hello, Ada!

A command is a .zu file, or a directory with an index.zu, sitting in a cmds directory. A project keeps its own in .zuri/cmds, so either of these answers to zuri greet:

.zuri/cmds/greet.zu           # a command in one file
.zuri/cmds/greet/index.zu     # a command with room to grow

The directory wins if both are there, so a command that has outgrown one file takes over the name as soon as its index.zu lands. Until then the directory is not a command, and the file goes on answering.

The commands the runtime itself ships live in the cmds directory beside the executable, and those win over a project’s. Everything after the name is forwarded to the command untouched, flags included.

A name that matches nothing is refused rather than guessed at:

$ zuri gret Ada
(Zuri):
  Launch aborted for gret
  Reason: Unknown command

A command says what it is in its own doc block, with two tags. zuri --help lists every command it can reach by exactly those:

Filename: .zuri/cmds/greet.zu

/**
 * @command greet
 * @description Say hello to somebody by name.
 */

import os

echo 'Hello, ${os.args[2]}!'
$ zuri --help
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC

Usage: zuri                    start the interactive REPL
       zuri run [PATH]         run a script, a package, or this directory
       zuri <command> [ARGS]   run a command

OPTIONS:
  -h, --help     Show this help message and exit
  -v, --version  Show version information and exit

COMMANDS:
  format  Lay out Zuri source in the project style.
  init    Scaffold a new Zuri project.
  test    Run a project's test files, each in its own process.

PROJECT COMMANDS:
  greet   Say hello to somebody by name.

Run "zuri <command> --help" for help on a specific command.

A description too long for one line carries on below it, indented:

/**
 * @command deploy
 * @description Ship the current build to staging, then wait for
 *    the health check to come back green.
 */

Two flags sit outside all of this and run nothing. zuri --help is the listing above, and zuri --version reports just the build:

$ zuri --version
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC

Both take the short spelling too, -h and -v.

The REPL

Run zuri with no arguments and you get an interactive prompt:

$ zuri
Zuri 0.1.0 (running on ZuriVM 0.1.0), REPL/Interactive mode = ON
Build No. => 2026-09-10 23:19:10 UTC
Type ".exit" to quit, ".help" for help or ".credits" for more information
%>

REPL stands for read-eval-print loop, and the print is the part that makes it useful. At the top level of the REPL, any expression whose value is not nil is printed for you, so you never need echo just to look at something:

%> 1 + 2
3
%> var name = 'Zuri'
%> name.upper()
ZURI
%> [1, 2, 3].map(@(n) { return n * n })
[1, 4, 9]

Notice that var name = 'Zuri' printed nothing. A declaration is not an expression, and anything that evaluates to nil stays quiet. That single rule is what keeps list.append(4) and echo x from doubling up on your screen.

Auto-printing applies only to the outermost level. Type a loop and the body does not print once per iteration:

%> iter var i = 0; i < 3; i++ {
..   i * 10
.. }
%>

Use echo inside a block when you want to see something.

Multi-line Input

The prompt changes from %> to .. when what you have typed so far is not yet a complete statement. Press Enter at the end of a line with an unclosed brace and keep going:

%> def double(n) {
..   return n * 2
.. }
%> double(21)
42

The same applies to an unclosed bracket, an unclosed parenthesis, or a string that has not been closed. The REPL decides when you are finished; you never have to signal it.

Everything Stays Defined

Variables, functions and classes you declare at the top level of the REPL remain defined for the rest of the session, because the whole session shares one namespace:

%> class Point {
..   @new(x, y) {
..     self.x = x
..     self.y = y
..   }
.. }
%> var p = Point(3, 4)
%> (p.x ** 2 + p.y ** 2).sqrt()
5

Redeclaring a name at the REPL’s top level replaces it rather than raising an error, so you can retype a function until it is right.

Getting Out

Type .exit. That is the only thing that ends the session.

Ctrl+C and Ctrl+D do not exit. Both print an interrupt line and hand the prompt back:

%> <KeyboardInterrupt [CtrlC]>
Type '.exit' to exit the REPL session
%>

Ctrl+C is how you abandon a half-typed line: the line is discarded and you start again on a fresh prompt.

Syntax Highlighting

On a terminal, the REPL colours what you type as you type it. Keywords are highlighted and everything else is left plain, so a misspelled retrun or whiel stands out before you press Enter — it simply does not change colour.

The highlighting understands string literals. A keyword inside quotes is text, not a keyword, and is left uncoloured:

%> var note = 'return it later'

return there stays plain, because it is part of the string.

Alongside it, a greyed-out suggestion appears to the right of the cursor when what you have typed so far matches something earlier in your history. Press the right arrow to accept it.

Both are terminal features. Piping input to zuri or redirecting its output produces plain text, so a session captured to a file has no escape codes in it.

The Other Dot Commands

.help reminds you that Tab offers completions.

.credits opens the project’s licence in a pager. Press q to come back.

Both are completed by Tab, along with every keyword in the language.

History

Up and Down walk through previous lines, and Ctrl+R searches backwards through them. History is written to a history.txt file sitting next to the zuri executable, so it survives between sessions.

What the REPL Is Good At

Each line is compiled and run on its own, which makes the REPL the right tool for a specific kind of question:

  • What does this method return? 'a,b,,c'.split(',')
  • Is this value truthy? !!(-1)
  • What type am I actually holding? typeof(x)
  • Does this syntax parse the way I think? Type it and find out.

It is a poor place to write a program, because you cannot go back and edit line four. Once you are past a few lines, put them in a .zu file and run it — which is what the next chapter does.

Programming a Bookmark Keeper

We are going to build something small and real: a command-line bookmark keeper. It asks you for a title and a URL, stores what you give it in a JSON file, and lets you list, search and delete entries later.

By the end of this chapter you will have used variables, functions, lists, dictionaries, loops, branching, string methods, file handling, JSON and error handling. We will not explain any of them thoroughly. That is what Chapters 3 to 8 are for. The goal here is to see a whole program, working, before we take it apart.

Follow along by typing the code rather than copying it. Mistakes are the fastest way to learn what the error messages mean.

Setting Up

$ mkdir bookmarks
$ cd bookmarks

Create main.zu and start with the modules we will need:

Filename: main.zu

import io
import json

var STORE = 'bookmarks.json'

echo 'Bookmark keeper.'
$ zuri run main.zu
Bookmark keeper.

Three things to notice. import brings in a module and binds it to a name you then reach through with a dot. var declares a variable. A .zu file’s top level is just code, running from top to bottom.

Asking a Question

The io module has readline(), which prints a prompt and waits for the user to type a line:

var title = io.readline('Title: ')
echo 'You typed: ' + title
$ zuri run main.zu
Title: Zuri docs
You typed: Zuri docs

The value that comes back includes whatever the user typed, whitespace included, so we almost always follow it with .trim():

var title = io.readline('Title: ').trim()

trim() is a method on strings. Every value in Zuri has methods, numbers, booleans and nil included, and you call them with a dot.

Storing One Bookmark

A bookmark has a title, a URL and some tags. The natural shape for that is a dictionary: a set of keys with values attached.

var bookmark = {
  title: 'Zuri docs',
  url: 'https://zuri.dev',
  tags: ['lang', 'docs'],
}

Square brackets make a list, an ordered sequence. Curly braces make a dictionary. You read a dictionary’s values with a dot, the same way you reach into a module:

echo bookmark.title
echo bookmark.tags
Zuri docs
[lang, docs]

Zuri has a shorthand you will use constantly. When the key you want and the variable holding its value have the same name, write the name once:

var title = 'Zuri docs'
var url = 'https://zuri.dev'

var bookmark = { title, url }

That is exactly the same dictionary as { title: title, url: url }, with half the noise.

Many Bookmarks

One bookmark is a dictionary. Many bookmarks is a list of them:

var bookmarks = []

bookmarks.append({ title: 'Zuri docs', url: 'https://zuri.dev' })
bookmarks.append({ title: 'Cranelift', url: 'https://cranelift.dev' })

echo bookmarks.length()
2

Saving to Disk

The json module turns Zuri values into text and back. file() opens a file; 'w' means open it for writing.

def save(bookmarks) {
  var handle = file(STORE, 'w')
  handle.write(json.encode(bookmarks, false))
  handle.close()
}

def declares a function. The parameter list needs no types (though it can have them; see Chapter 5). The body is a block, and blocks in Zuri always use braces.

That second argument to json.encode() is compact. It defaults to true, which gives you one dense line. Passing false asks for the indented form, which is what you want in a file a human might open:

[
  {
    "title": "Cranelift",
    "url": "https://cranelift.dev",
    "tags": [
      "compilers"
    ]
  }
]

Loading From Disk, Carefully

Reading is where a real program has to start thinking about what can go wrong. The file might not exist yet. It might exist and contain garbage because someone edited it by hand. Neither should stop the program.

def load() {
  var handle = file(STORE)

  if !handle.exists() {
    return []
  }

  catch {
    return json.decode(handle.read())
  } as error {
    echo 'Could not read ${STORE}: ${error.message}'
    return []
  }
}

Two new pieces here.

The first is ${...} inside a string. That is interpolation: the expression between the braces is evaluated and its value spliced into the string. It works in both single- and double-quoted strings.

The second is catch. Zuri does not have try. You write catch { ... } around the code that might fail, and as error { ... } to handle it. If nothing goes wrong, the handler never runs. If something does, error is an object with a message, a type and a stacktrace.

Notice there is no finally. Zuri does not have one. Code after the catch statement runs either way, which covers most of what finally is used for.

Adding a Bookmark

Now we can put the pieces together:

def add(bookmarks) {
  var title = io.readline('Title: ').trim()
  var url = io.readline('URL:   ').trim()

  if title.is_empty() or url.is_empty() {
    echo 'Both a title and a URL are required.'
    return
  }

  var tags = io.readline('Tags (comma separated): ').trim()

  bookmarks.append({
    title,
    url,
    tags: tags.is_empty() ? [] : tags.split(',').map(@(t) { return t.trim() }),
  })

  save(bookmarks)
  echo 'Saved "${title}".'
}

The tags line packs in three ideas.

cond ? a : b is the conditional operator: if cond is truthy the whole expression is a, otherwise b.

split(',') cuts a string into a list at every comma.

map() takes a function and applies it to every element, giving back a new list. @(t) { return t.trim() } is an anonymous function: @ followed by a parameter list and a body. You can also spell it def(t) { ... }; @ is the shorthand and it is what most Zuri code uses for a one-liner.

There is a shorter form still. When the body is a single expression that you want returned, => replaces the braces and the return:

tags.split(',').map(@(t) => t.trim())

That is the same function written three ways. Chapter 5 covers every spelling, and when each one reads best.

Also notice or rather than ||. Zuri spells its logical operators and, or and !.

Printing a Bookmark

def show(bookmark, index) {
  echo '${index + 1}. ${bookmark.title}'
  echo '   ${bookmark.url}'

  if !bookmark.tags.is_empty() {
    echo '   [' + ', '.join(bookmark.tags) + ']'
  }
}

join lives on the string, not on the list, and it reads exactly as it works: take this separator, and stitch that list together with it.

Listing and Searching

def list_all(bookmarks) {
  if bookmarks.is_empty() {
    echo 'Nothing saved yet.'
    return
  }

  bookmarks.each(@(bookmark, index) {
    show(bookmark, index)
  })
}

each() calls your function once per element, handing it the value first and the index second. That order catches people out; it is the same for lists, dictionaries and strings.

Search is a filter plus a test:

def find(bookmarks) {
  var needle = io.readline('Search: ').trim().lower()

  var hits = bookmarks.filter(@(bookmark) {
    if bookmark.title.lower().contains(needle) {
      return true
    }
    return bookmark.tags.some(@(tag) { return tag.lower() == needle })
  })

  if hits.is_empty() {
    echo 'No match for "${needle}".'
    return
  }

  hits.each(@(bookmark, index) {
    show(bookmark, index)
  })
}

filter() keeps the elements your function returns true for. some() answers “is this true of at least one element?” and stops at the first one that matches.

Removing a Bookmark

def remove(bookmarks) {
  list_all(bookmarks)

  if bookmarks.is_empty() {
    return
  }

  var answer = io.readline('Remove which number? ').trim()
  var position = answer.to_number() - 1

  if position < 0 or position >= bookmarks.length() {
    echo 'There is no bookmark ${answer}.'
    return
  }

  var gone = bookmarks[position]
  bookmarks.remove_at(position)
  save(bookmarks)
  echo 'Removed "${gone.title}".'
}

to_number() converts a string; bookmarks[position] indexes a list; remove_at() deletes by position and shifts the rest down.

The bounds check is not optional politeness. Indexing a list past its end raises an error, and an unhandled error ends the program.

The Main Loop

The last piece is a loop that reads a command and dispatches on it:

def main() {
  var bookmarks = load()

  echo 'Bookmark keeper. ${bookmarks.length()} saved.'

  while true {
    var command = io.readline('\n(a)dd (l)ist (f)ind (r)emove (q)uit > ').trim().lower()

    using command {
      when 'a' add(bookmarks)
      when 'l' list_all(bookmarks)
      when 'f' find(bookmarks)
      when 'r' remove(bookmarks)
      when 'q' {
        echo 'Bye.'
        return
      }
      default {
        echo 'Unknown command "${command}".'
      }
    }
  }
}

main()

using is Zuri’s multi-way branch, and when is its branch keyword. It compares the subject against each when value and runs the first match; default catches everything else. A when whose body is a single short statement can stay on one line, which is what makes the dispatch table above readable. Anything longer gets a block.

Unlike a C switch, there is no fall-through and no break to remember. One branch runs, then the statement is over.

The Whole Program

Filename: main.zu

import io
import json

var STORE = 'bookmarks.json'

def load() {
  var handle = file(STORE)

  if !handle.exists() {
    return []
  }

  catch {
    return json.decode(handle.read())
  } as error {
    echo 'Could not read ${STORE}: ${error.message}'
    return []
  }
}

def save(bookmarks) {
  var handle = file(STORE, 'w')
  handle.write(json.encode(bookmarks, false))
  handle.close()
}

def add(bookmarks) {
  var title = io.readline('Title: ').trim()
  var url = io.readline('URL:   ').trim()

  if title.is_empty() or url.is_empty() {
    echo 'Both a title and a URL are required.'
    return
  }

  var tags = io.readline('Tags (comma separated): ').trim()

  bookmarks.append({
    title,
    url,
    tags: tags.is_empty() ? [] : tags.split(',').map(@(t) { return t.trim() }),
  })

  save(bookmarks)
  echo 'Saved "${title}".'
}

def show(bookmark, index) {
  echo '${index + 1}. ${bookmark.title}'
  echo '   ${bookmark.url}'

  if !bookmark.tags.is_empty() {
    echo '   [' + ', '.join(bookmark.tags) + ']'
  }
}

def list_all(bookmarks) {
  if bookmarks.is_empty() {
    echo 'Nothing saved yet.'
    return
  }

  bookmarks.each(@(bookmark, index) {
    show(bookmark, index)
  })
}

def find(bookmarks) {
  var needle = io.readline('Search: ').trim().lower()

  var hits = bookmarks.filter(@(bookmark) {
    if bookmark.title.lower().contains(needle) {
      return true
    }
    return bookmark.tags.some(@(tag) { return tag.lower() == needle })
  })

  if hits.is_empty() {
    echo 'No match for "${needle}".'
    return
  }

  hits.each(@(bookmark, index) {
    show(bookmark, index)
  })
}

def remove(bookmarks) {
  list_all(bookmarks)

  if bookmarks.is_empty() {
    return
  }

  var answer = io.readline('Remove which number? ').trim()
  var position = answer.to_number() - 1

  if position < 0 or position >= bookmarks.length() {
    echo 'There is no bookmark ${answer}.'
    return
  }

  var gone = bookmarks[position]
  bookmarks.remove_at(position)
  save(bookmarks)
  echo 'Removed "${gone.title}".'
}

def main() {
  var bookmarks = load()

  echo 'Bookmark keeper. ${bookmarks.length()} saved.'

  while true {
    var command = io.readline('\n(a)dd (l)ist (f)ind (r)emove (q)uit > ').trim().lower()

    using command {
      when 'a' add(bookmarks)
      when 'l' list_all(bookmarks)
      when 'f' find(bookmarks)
      when 'r' remove(bookmarks)
      when 'q' {
        echo 'Bye.'
        return
      }
      default {
        echo 'Unknown command "${command}".'
      }
    }
  }
}

main()

A session looks like this:

$ zuri run main.zu
Bookmark keeper. 0 saved.

(a)dd (l)ist (f)ind (r)emove (q)uit > a
Title: Zuri docs
URL:   https://zuri.dev
Tags (comma separated): lang, docs
Saved "Zuri docs".

(a)dd (l)ist (f)ind (r)emove (q)uit > l
1. Zuri docs
   https://zuri.dev
   [lang, docs]

(a)dd (l)ist (f)ind (r)emove (q)uit > q
Bye.

What You Just Learned

You wrote a program with persistent state, user input, error recovery and a command loop, in about a hundred lines, using nothing that is not in the box.

Along the way you met: var, def, if/else, while, using/when, catch/as, return, lists, dictionaries, dictionary shorthand, string interpolation, the conditional operator, anonymous functions, and a handful of built-in methods.

The next six chapters take every one of those and explain it properly.

Common Programming Concepts

Every language gives you a way to name a value, a set of types those values come in, operators to combine them, and statements to decide what runs next. This chapter is those four things in the form Zuri gives them.

If this is your first language, read the sections in order; each one uses only what came before it. If you already program, read Operators and Control Flow attentively — those are where the same syntax you already know does something different here.

The Reserved Words

Before anything else, the words you cannot use as names. Zuri reserves 31 keywords:

and       as        assert    break     catch     class     const
continue  def       default   do        echo      else      false
for       if        import    in        iter      nil       or
parent    raise     return    self      static    true      using
var       when      while

Every one is lowercase, and every one is explained somewhere in this book. Appendix A is the index, with a one-line summary and a pointer for each.

Three things are not in that list, and are worth noticing now:

  • print is not a keyword. It is an ordinary built-in function, and you could shadow it with a variable of your own if you wanted to. echo is a keyword.
  • There is no function, int, string or bool keyword. Type names are ordinary identifiers, which is why they can appear as annotations without being reserved.
  • There is no try, finally, switch, case, new, this, public or private. The equivalents are catch, using, when, calling the class directly, self, and a leading underscore.

Variables and Constants

A variable is a name for a value. In Zuri you introduce one with var, and from then on the name stands for whatever you last put in it.

var greeting = 'hello'
var count = 0

echo greeting
echo count
hello
0

That is the whole idea. The rest of this section is the detail: what a name may be called, what happens when you leave the value out, where the name can be seen, and what const adds.

Declaring

A var with no initializer holds nil, the value that means “nothing here yet”:

var pending

echo pending
nil

You can declare several names in one statement by separating them with commas, and each one may or may not have a value:

var a = 1, b = 2, c

echo [a, b, c]
[1, 2, nil]

This is worth using when the names genuinely belong together — the three components of a colour, the lower and upper bound of a window — and worth avoiding when they do not, because one long line of declarations is harder to read than three short ones.

var is a declaration, not an expression. It has no value, which is why you cannot write if var x = f() or echo var y = 1. Declare first, then use the name.

Naming

A name is made of letters, digits and underscores, and may not start with a digit. firstName, first_name and _first_name are all legal, and all three are different names.

Zuri code follows one set of conventions everywhere, and the standard library is written in it:

Kind of nameConventionExample
variable, function, methodsnake_caseread_line, total_count
classPascalCaseAccount, HttpClient
constantSCREAMING_SNAKE_CASEMAX_RETRIES
private anythingleading underscore_cache, _retry()

The leading underscore is the one convention that is not only a convention. A name beginning with _ is private, and the compiler enforces it: another file cannot reach module._helper, and code outside a class cannot reach instance._field. Chapter 6 and Chapter 8 cover what that means in each case.

The 31 keywords listed at the end of the previous section cannot be used as names at all. Trying produces a syntax error at the point of declaration.

Reassignment

Zuri is dynamically typed. A variable holds a value, not a type, and assigning a different kind of value to the same name is legal:

var value = 42
value = 'now a string'
value = [1, 2, 3]

echo value
[1, 2, 3]

Legal is not the same as advisable. A name that holds a number on one line and a list twenty lines later is a name no reader can predict. Reach for a second variable instead; they are free.

Assignment is an expression, and it evaluates to the value assigned. That lets you chain:

var x, y

x = y = 5

echo '${x} ${y}'
5 5

Compound Assignment

Every arithmetic and bitwise operator has an assignment form, which applies the operator to the variable’s current value and stores the result:

var n = 5

n += 10     # 15
n **= 2     # 225
n //= 7     # 32
n <<= 1     # 64

echo n
64

The full set is +=, -=, *=, /=, //=, **=, %=, &=, |=, ^=, ~=, <<=, >>= and >>>=. Each one means exactly what the matching binary operator means, applied in place.

Compound assignment works on anything you can assign to, not just plain variables:

var totals = { food: 0 }
var scores = [1, 2, 3]

totals.food += 12
scores[0] *= 10

echo totals
echo scores
{food: 12}
[10, 2, 3]

Increment and Decrement

++ and -- add or subtract one in place:

var n = 5

n++
echo n

n--
echo n
6
5

Two rules to remember. First, both are postfix only. ++n is a syntax error; the operator goes after the name.

Second, when n++ appears inside a larger expression it updates n and then evaluates to the new value:

var j = 5

echo j++
echo j
6
6

If you have written C, Java or JavaScript, this is the opposite of what you expect: there, j++ gives you the old value. In Zuri it gives you the new one. The habit that avoids the question entirely is to let ++ be a statement of its own, or to use it in the update clause of an iter loop where the value is discarded:

iter var i = 0; i < 3; i++ {
  echo i
}
0
1
2

const

const declares a name that cannot be reassigned:

def area(radius) {
  const PI = 3.141592653589793

  return PI * radius * radius
}

echo area(2)
12.566370614359172

Writing to it is caught when the file is compiled, before anything runs:

def f() {
  const K = 1
  K = 2
}
SyntaxError: cannot assign to constant 'K'
  --> /path/to/main.zu:3:3
  |
3 |   K = 2
  |   ^

A const must be given a value where it is declared. const K on its own does not compile, which is the point: a constant with no value would be a constant nil forever.

The const keyword is enforced in local scopes — inside functions, methods, and blocks. At the top level of a module, const declares an ordinary module global and the reassignment check does not apply. Use it where it does its work, which is inside the code that would otherwise be tempted to reassign.

Constant Does Not Mean Frozen

const prevents rebinding the name. It says nothing about the value. A constant list is still a list, and a list is still mutable:

def demo() {
  const items = [1, 2]

  items.append(3)
  echo items
}

demo()
[1, 2, 3]

items = [4] would be an error; items.append(3) is not. If you need a collection nothing can change, copy it at the boundary where you hand it out rather than relying on the declaration.

Scope

A block is a scope. Braces open one, and every name declared inside it is gone when the block closes:

var x = 'outer'

{
  var x = 'inner'
  echo x
}

echo x
inner
outer

The inner x hides the outer one for the length of the block. This is called shadowing, and it is allowed deliberately: a short block can use a short name without worrying about what that name means outside it.

Function bodies, loop bodies, if branches and bare { ... } blocks all work this way. So does the initialiser of an iter loop, whose variable belongs to the loop:

iter var i = 0; i < 2; i++ {
  echo i
}

echo i
0
1
Unhandled UndefinedError: undefined global 'i'
  --> /path/to/main.zu:4

After the loop, i is not a variable at all, and reading it is an error. That error is the good outcome: a name that does not exist is a mistake worth hearing about, and Zuri never invents a silent nil to paper over one.

One Declaration Per Scope

Shadowing across scopes is allowed. Declaring the same name twice in one scope is not:

def f() {
  var a = 1
  var a = 2
}
SyntaxError: 'a' is already declared in this scope
  --> /path/to/main.zu:3:7
  |
3 |   var a = 2
  |       ^

The exception is the top level of a script or module, where a second var a rebinds the existing global instead of erroring. That is what lets you retype a declaration in the REPL while you are experimenting.

Functions Follow the Same Rule

A def is scoped exactly like a var. One written at the top level of a file binds a module-level name; one written inside a function is a local of that function, and nothing outside can reach it:

def outer() {
  def helper() {
    return 'from helper'
  }

  return helper()
}

echo outer()
echo helper()
from helper
Unhandled UndefinedError: undefined global 'helper'
  --> /path/to/main.zu:9

The first call works because helper is a local of outer. The second fails because, outside outer, no such name was ever created.

Chapter 5 covers this in full.

Data Types

Zuri has ten built-in types. typeof() names the one you are holding:

var values = [1, 1.5, 'text', true, nil, [1], {a: 1}, 1..3, 10n, bytes(2)]

for v in values {
  echo typeof(v)
}
number
number
string
bool
nil
list
dict
range
bigint
bytes

On top of those there are function, class, instance, module, file and ptr, which you get from declaring things rather than from a literal.

number

A number is an IEEE-754 double. There is no separate integer type, which means 1 and 1.0 are the same value:

echo 1 == 1.0
true

is_int() asks whether a number’s value is integral, not whether it was written without a decimal point:

echo is_int(1)
echo is_int(1.0)
echo is_int(1.5)
true
true
false

Doubles hold integers exactly up to 2^53. Past that, precision goes:

echo 9007199254740993
9007199254740992

Numeric literals come in five forms:

var decimal = 1_000_000
var binary = 0b1010          # 10
var octal = 0c17             # 15
var hexadecimal = 0xff       # 255
var scientific = 6.02e23

echo [decimal, binary, octal, hexadecimal, scientific]
[1000000, 10, 15, 255, 602000000000000000000000]

Underscores are digit separators, and they are accepted in decimal literals only. 1_000_000 and 1_0.5_5 are fine; 0xdead_beef and 0b1010_1010 are not. Numbers has the full rules.

The special values behave the way the standard says:

echo 1 / 0
echo -1 / 0
echo 0 / 0
inf
-inf
NaN

Every number carries methods, so mathematics reads left to right:

echo 2.sqrt()
echo 16.log2()
echo (-3).abs()
echo 3.7.round()
echo 3.14159.fixed(2)
echo 255.hex()
1.4142135623730951
4
3
4
3.14
ff

A literal takes a method directly; the parentheses around (-3) are there because a method call binds tighter than the minus sign, so -3.abs() would negate the result instead of the operand. Numbers covers the rule in full.

bigint

When 2^53 is not enough, suffix the literal with n and you get an arbitrary-precision integer:

echo 9007199254740993n + 1n
9007199254740994n

Bigints never lose precision and never overflow. They are slower than numbers and they do not silently mix with them, so convert explicitly with to_bigint() and to_number(). Chapter 4 covers them properly.

bool

true and false. Every value in Zuri is either truthy or falsy, and that determines what if, while, and, or, ! and ? : do with it. The falsy values are:

ValueWhy
falseitself
nilabsence of a value
0, 0.0, -0.0zero
NaNnot a number at all
0nthe bigint zero
''the empty string
bytes(0)an empty byte buffer

Everything else is truthy, including every negative number, [], {} and '0'.

Two consequences deserve a second look.

A position of 0 is falsy. index_of() returns -1 when it finds nothing, which is truthy, and 0 when it finds the item first, which is falsy. So if list.index_of(x) { ... } gets both cases backwards. Compare the result explicitly:

var haystack = ['a', 'b', 'c']
var position = haystack.index_of('z')

if position == -1 {
  echo 'not found'
}
not found

An empty list is truthy. [] and {} are objects, and objects are truthy. Use is_empty():

var items = []

if items.is_empty() {
  echo 'nothing here'
}
nothing here

nil

nil is the absence of a value. An uninitialised var is nil, a function that falls off the end returns nil, and a missing dictionary key read through get() gives nil.

nil is falsy, and it still has a to_string():

echo nil.to_string()
nil

Calling any other method on nil raises a TypeError, which is usually exactly the error you wanted.

string

This is the short version. Strings covers quoting, escapes, interpolation, concatenation, repetition and regular expressions in full.

Strings are written in single or double quotes, with no difference in meaning:

echo 'single'
echo "double"

Both kinds may span multiple lines:

echo 'multi
line'
multi
line

The escape sequences are \0, \a, \b, \f, \n, \r, \t, \v, \\, \', \", plus \xNN for a byte, \uNNNN for a code point and \UNNNNNNNN for one outside the basic plane:

echo "tab:\there"
echo "hex: \x41"
echo "emoji: \U0001F600"
tab:	here
hex: A
emoji: 😀

A backslash followed by anything else is left alone, backslash and all.

Interpolation

${...} inside a string evaluates the expression and splices the result in. It works in both quote styles and the expression can be anything:

var name = 'Zuri'
echo 'Hi ${name}, ${1 + 2} and ${name.upper()}'
Hi Zuri, 3 and ZURI

To produce a literal ${, build it by concatenation:

echo 'B: $' + '{x}'
B: ${x}

Strings Are Sequences

Indexing gives you a one-character string, and negative indices count from the end:

var s = 'hello world'
echo s[0]
echo s[0,5]
echo s[-5,]
h
hello
world

s[a,b] is a slice: from a up to but not including b. Either side may be left out, so s[,3] is the first three characters and s[3,] is everything from index three onward.

+ concatenates and * repeats:

echo 'ab' + 'cd'
echo 'ab' * 3
abcd
ababab

list

An ordered, growable sequence, written in square brackets. Elements may be of any type, including other lists:

var mixed = [1, 'two', [3], { four: 4 }]
echo mixed.length()
4

Lists index and slice exactly like strings, negative indices included. Chapter 4 covers the forty methods they carry.

dict

An insertion-ordered mapping from keys to values:

var config = { host: 'localhost', port: 8080, debug: true }

A bare word key is taken as a string, so { host: ... } and { 'host': ... } are the same dictionary. Keys can also be numbers, and they can be computed:

var d = { name: 'Ada', 'age': 36, 3: 'three' }
echo d['name']
echo d.name
echo d[3]
Ada
Ada
three

Dot access and bracket access are the same operation. Use the dot when the key is a fixed name, brackets when it is computed or not a valid identifier.

When a key and the variable holding its value share a name, write it once:

var host = 'localhost'
var port = 8080

var config = { host, port }

range

a..b describes the integers from a up to but not including b:

echo 1..5
echo 1..5.to_list()
1..5
[1, 2, 3, 4]

A range is a real value, not loop syntax. You can store one, pass it around, and ask it questions:

var r = 0..10
echo r.lower()
echo r.upper()
echo r.within(7)
0
10
true

bytes

A fixed-size buffer of 8-bit values, for binary data. bytes(n) allocates n zero bytes; bytes(list) builds one from numbers in 0..256:

var b = bytes([72, 101, 108, 108, 111])
echo b
echo b.to_string()
echo b[0]
(48 65 6c 6c 6f)
Hello
72

Bytes print as hexadecimal in parentheses, which is how you can always tell one from a list at a glance. Chapter 10 is the full treatment.

Checking Types

typeof() gives you a name. The is_* family gives you a boolean, and there are fifteen of them:

is_bigint   is_bool     is_bytes    is_callable  is_class
is_dict     is_file     is_function is_instance  is_int
is_iterable is_list     is_number   is_object    is_string
echo is_string('x')
echo is_callable(print)
echo is_iterable([1, 2])
true
true
true

Use instance_of(value, SomeClass) for classes, which walks the inheritance chain. Chapter 6 covers that.

Operators

Arithmetic

OperatorMeaningExample
+addition2 + 3 is 5
-subtraction5 - 2 is 3
*multiplication4 * 3 is 12
/division, always floating point7 / 2 is 3.5
//floor division7 // 2 is 3
%remainder7 % 3 is 1
**exponentiation2 ** 10 is 1024

/ never truncates. 7 / 2 is 3.5 even though both operands look like integers, because there is only one numeric type. When you want the integer, ask for it with //.

// rounds towards negative infinity, while % keeps the sign of the left operand:

echo -7 // 2
echo 7 // -2
echo -7 % 3
echo 7 % -3
-4
-4
-1
1

% works on non-integers too: 7.5 % 2 is 1.5.

Comparison

OperatorMeaning
==equal
!=not equal
< <= > >=ordering, numbers only

== compares by value for numbers, strings, lists and dictionaries, and by identity for everything else. A class changes that for its instances with @eq.

echo 1 == 1.0
echo [1, 2] == [1, 2]
echo {a: 1} == {a: 1}
echo '1' == 1
true
true
true
false

There is no coercion in ==. A string is never equal to a number.

The ordering operators are for numbers. Comparing two strings with < raises a TypeError; use compare(), which returns -1, 0 or 1:

echo 'abc'.compare('abd')
-1

Logic

Zuri spells these as words:

OperatorMeaning
andtrue when both sides are truthy
ortrue when either side is truthy
!negation

and and or short-circuit, and they return the operand, not a boolean:

echo 1 and 2
echo nil or 'fallback'
2
fallback

That is what makes var name = given or 'anonymous' work. Remember that 0, NaN, '' and false are falsy, so this idiom is only safe when those are not legitimate values.

The Conditional Operator

echo true ? 'yes' : 'no'
yes

cond ? a : b evaluates cond, then exactly one of the branches.

Bitwise

OperatorMeaning
&and
|or
^xor
~not
<<left shift
>>arithmetic right shift, sign preserving
>>>logical right shift, zero filling
echo ~5
echo 5 >>> 1
echo 1 << 10
-6
2
1024

Bitwise operators work on the integer value of a number, and on bigints.

Concatenation and Repetition

+ on a string concatenates, and a number on either side is converted:

echo 'n=' + 5
n=5

+ on a list concatenates; * repeats:

echo [1, 2] + [3]
echo [1, 2] * 2
[1, 2, 3]
[1, 2, 1, 2]

Dictionaries do not support +. Use extend().

Anything the language does not define raises a TypeError that names the exact signature:

catch {
  echo nil + 1
} as e {
  echo e.message
}
operator '+' not defined for call signature (nil, number)

Classes can define what + and every other operator mean for their own instances. That is Chapter 6.

Member, Index and Slice

SyntaxMeaning
x.namemember of an object, dictionary or module
x[key]index with a computed key
x[a, b]slice from a up to but not including b
x[, b]slice from the start
x[a, ]slice to the end
a..ba range value

Negative indices count back from the end, for both strings and lists.

Precedence

From tightest to loosest. Everything on one row binds equally and associates left to right, except **, which associates right to left.

LevelOperators
1literals, (...), [...], {...}, self, parent
2..
3. () []
4++ --
5**
6! - ~ (unary)
7* / // %
8+ -
9<< >> >>>
10&
11^
12|
13< <= > >= == !=
14and
15or
16? :
17= and every compound assignment

** follows the mathematical convention. It binds tighter than multiplication and tighter than a unary operator written before it, and it groups from the right:

echo 2 * 3 ** 2
echo 2 ** 3 ** 2
echo -2 ** 2
echo 2 ** -1
18
512
-4
0.5

2 * 3 ** 2 is 2 * (3 ** 2), 2 ** 3 ** 2 is 2 ** (3 ** 2), and -2 ** 2 is -(2 ** 2). The exponent itself may carry a sign, so 2 ** -1 needs no parentheses. Write (-2) ** 2 to raise a negative number.

.. binds very tightly, to primaries only. 1 + 2..5 parses as 1 + (2..5). Parenthesise any range whose endpoints are expressions.

Operators Zuri Does Not Have

There is no in operator for membership; use contains(). There is no ?. optional chaining and no ?? null coalescing; or covers the common case. There is no comma operator, and no += on a member that does not already exist.

Comments and Doc Blocks

Line Comments

# runs to the end of the line:

# Rates are quoted per thousand, not per unit.
var rate = 0.0125  # 1.25%

Block Comments

/* ... */ spans as many lines as you like, and nests:

/* This whole section is off.
   /* Including this inner comment. */
   Still off. */

Nesting is worth knowing about, because it means commenting out a region that already contains a block comment works the way you expect, and it also means an unbalanced /* inside a comment swallows the rest of your file.

Doc Blocks

A block comment that opens with /** is a doc block. The parser keeps doc blocks in the syntax tree rather than discarding them, which is what lets the zuri module read a file’s documentation without a separate parser. Put one directly above the thing it documents:

/**
 * Converts a duration in seconds to a human-readable string.
 *
 * Rounds to the nearest whole second. Durations below one second
 * render as `'0s'`.
 *
 * @param number seconds
 * @returns string
 */
def humanize(seconds) {
  # ...
}

The tag vocabulary used across the standard library is:

TagMeaning
@param {type} name: descriptionone parameter
@returns typewhat the function gives back
@throws ErrorClassan error it can raise
@notesomething the caller must know
@defaultthe default a parameter falls back to

Two conventions from the standard library are worth copying. State the default of every optional parameter, and state what happens at the edges: an empty input, a zero length, a value out of range. Anything a caller would otherwise have to discover by experiment belongs in the doc block.

Doc blocks are not only for readers. Because the parser keeps them, a program can read them: zuri.parse() returns each one as a DocBlock node sitting immediately before the declaration it documents, which is enough to build a documentation generator in a few dozen lines. Chapter 21 shows how.

Commenting Style

A comment earns its place by explaining why, not what. The code already says what it does:

# Bad: restates the code.
# Add one to the counter.
counter++

# Good: explains a decision the code cannot.
# Servers count from one, and the wire protocol has no zero frame.
counter++

Control Flow

Control flow is how a program decides what to run next: run this only when that is true, run this until that stops being true, run this once for every item in a collection. Zuri has seven statements for it, and this section covers all of them.

Blocks and Single Statements

Every control-flow statement takes a body. A body is either a block in braces, or a single statement:

var ready = true

if ready {
  echo 'launching'
}

if ready echo 'launching again'
launching
launching again

Both forms are the language. if, else, while, do, for, and a when branch inside using all accept either one.

iter is the exception. Its body must be a block, because the semicolons in its header would otherwise be ambiguous with the statement that follows:

iter var i = 0; i < 3; i++ echo i
SyntaxError: Expected '{' at the start of iter body.

The standard library and the examples in this book brace every body except a short when branch. Braces make a one-line body easy to extend into a three-line one later, and they keep the shape of a nested statement obvious at a glance. It is a convention, not a rule — write whichever form reads better in the code you are writing.

if / else

var score = 73

if score >= 90 {
  echo 'A'
} else if score >= 70 {
  echo 'B'
} else {
  echo 'C'
}
B

There is no elif; else if is two keywords, and it works because the body of an else can itself be an if statement.

The condition is any expression at all, and it is judged by truthiness rather than by being a bool:

var name = ''

if name {
  echo 'have a name'
} else {
  echo 'no name'
}
no name

That is convenient and it is also where most bugs in new Zuri code come from, because a legitimate 0 is falsy. Re-read the table in Data Types before you write if count or if position.

The Conditional Expression

? : is the expression form of if. It chooses between two values rather than between two statements:

var age = 20
var status = age >= 18 ? 'adult' : 'minor'

echo status
adult

It nests, and it can span several lines. When breaking one across lines, ? and : may either end a line or begin the next, whichever reads better; leading each branch with its operator keeps the shape of the choice visible:

var n = 7

var size = n > 100
  ? 'large'
  : n > 5
    ? 'medium'
    : 'small'

echo size
medium

Use ? : when you are producing a value and if when you are performing an action. A conditional expression whose branches are both side effects is harder to read than the if it replaced.

while

while tests before each pass, so a body may run zero times:

var n = 0

while n < 3 {
  echo 'while ${n}'
  n++
}
while 0
while 1
while 2

do / while

do runs the body once and then tests, so the body always runs at least once. Reach for it when the test depends on something the body produces:

var m = 10

do {
  echo 'do ${m}'
  m++
} while m < 3
do 10

The condition was false from the very start, and the body still ran.

iter

iter is the counting loop. Its header has three clauses separated by semicolons: an initialiser, a condition, and an update.

iter var i = 0; i < 3; i++ {
  echo 'iter ${i}'
}
iter 0
iter 1
iter 2

The initialiser runs once. The condition is tested before every pass. The update runs after every pass. A variable declared in the initialiser belongs to the loop and does not exist after it.

The initialiser may declare several variables and the update may hold several expressions, each list separated by commas. The updates run in the order they are written:

iter var i = 0, j = 10; i < 3; i++, j-- {
  echo '${i} ${j}'
}
0 10
1 9
2 8

Every clause is optional:

var k = 0

iter ; k < 2; {
  echo 'bare ${k}'
  k++
}
bare 0
bare 1

iter ; ; { ... } loops forever, though while true says it more plainly.

The header may be spread over several lines, which is worth doing when the condition is long:

iter var i = 1;
  i <= 3;
  i++
{
  echo i
}
1
2
3

for / in

for walks an iterable: a list, a dictionary, a string, a range, a byte stream, or any class that implements the iterator protocol.

With one variable you get the value:

for ch in 'hey' {
  echo ch
}
h
e
y

With two, you get the key first and the value second:

for index, value in ['a', 'b'] {
  echo '${index} ${value}'
}

for key, value in { x: 1, y: 2 } {
  echo '${key}=${value}'
}
0 a
1 b
x=1
y=2

For a list the key is the index, for a dictionary it is the key, and for a string it is the character position. Dictionaries iterate in insertion order, and so does everything else with an order to preserve.

Ranges are exclusive at the top:

for i in 0..3 {
  echo i
}
0
1
2

The iterable expression is evaluated exactly once, before the first pass. Calling a function in that position calls it once, not once per element:

def source() {
  echo 'source() called'
  return [1, 2, 3]
}

for n in source() {
  echo n
}
source() called
1
2
3

Chapter 6 shows how to make your own class work with for by defining @key() and @value().

break and continue

continue skips to the next pass, and break leaves the loop entirely:

iter var i = 0; i < 5; i++ {
  if i == 1 {
    continue
  }

  if i == 3 {
    break
  }

  echo i
}
0
2

Both apply to the innermost enclosing loop, and there are no loop labels. To leave a nested loop you either carry a flag:

var found = false

iter var i = 1; i < 4; i++ {
  iter var j = 1; j < 4; j++ {
    if i * j == 4 {
      echo 'found ${i} x ${j}'
      found = true
      break
    }
  }

  if found {
    break
  }
}
found 2 x 2

or, more often, put the nested loop in a function and return out of it:

def first_product(target) {
  iter var i = 1; i < 4; i++ {
    iter var j = 1; j < 4; j++ {
      if i * j == target {
        return [i, j]
      }
    }
  }

  return nil
}

echo first_product(4)
echo first_product(99)
[2, 2]
nil

The second version says what it is looking for in its name, and it has no flag to keep in sync. Prefer it.

using / when

using compares one subject against several candidate values:

var grade = 'B'

using grade {
  when 'A' echo 'excellent'
  when 'B', 'C' echo 'good'
  default echo 'try again'
}
good

Several values on one when are alternatives, not a sequence: this when matches 'B' or 'C'. The first branch that matches runs, and then the statement is over — there is no fall-through, and no break to remember.

Matching uses the same equality as ==, which means using works on numbers, strings and booleans, and compares instances by identity, or through @eq when their class defines one.

default is optional. When nothing matches and there is no default, nothing happens:

using 99 {
  when 1 echo 'one'
}

echo 'nothing ran, and that is not an error'
nothing ran, and that is not an error

A when branch takes a block when it needs one:

var command = 'q'

using command {
  when 'a' echo 'adding'
  when 'q' {
    echo 'saving'
    echo 'goodbye'
  }
  default echo 'unknown command'
}
saving
goodbye

using is the right shape whenever you are dispatching on one value to many outcomes. A chain of else if comparing the same variable over and over is the same thing written longer.

assert

assert checks a condition and raises an AssertError when it is false:

var items = [1]

assert 1 + 1 == 2
assert items.length() == 1, 'list should have one item'

echo 'both assertions held'
both assertions held

The optional second argument is the message the error carries:

assert 1 == 2, 'arithmetic is broken'
Unhandled AssertError: arithmetic is broken

Assertions state what you believe is already true. They are for conditions that should be impossible, not for checking input that might legitimately be wrong — raise a ValueError for that. Chapter 7 draws the line in detail.

return

return leaves the current function, with a value if you give it one and nil if you do not:

def classify(n) {
  if n < 0 {
    return 'negative'
  }

  if n == 0 {
    return 'zero'
  }

  return 'positive'
}

echo classify(-5)
echo classify(0)
echo classify(3)
negative
zero
positive

A function that reaches its closing brace without a return yields nil. There are no implicit returns; the last expression in a body is not its value.

return is only legal inside a function. At the top level of a script it is a syntax error, because there is nothing to return from.

Text, Numbers and Collections

Chapter 3 introduced the built-in types and showed what the operators do to them. This chapter is the working manual for the five you will reach for every day: strings, numbers, lists, dictionaries and ranges.

One idea shapes all five. Zuri puts behaviour on the value itself. There is no len(x) and no str.upper(x); there is x.length() and x.upper(). That holds for numbers too, which carry the whole of the mathematics library as methods:

echo 'zuri'.upper()
echo [3, 1, 2].sort()
echo 16.sqrt()
echo 255.hex()
ZURI
[1, 2, 3]
4
ff

Because the methods live on the value, they chain, and a chain reads left to right in the order the work happens:

var line = '  Ada, Grace , Alan  '

echo line.trim().split(',').map(@(name) => name.trim()).length()
3

Each section here covers the methods worth knowing by name, the ones with a behaviour you would not guess, and the mistakes that come up most often. Appendix E is the complete list, generated from the same documentation the runtime ships.

Strings

A Zuri string is UTF-8 text. Indexing, slicing and length() work in characters, not bytes, so a string of emoji behaves the way a reader expects.

Text is the type you will write most, so this section starts with how you write one down — quoting, escapes and interpolation — before moving on to what you can do with it.

Writing a String

A string literal is written between single or double quotes:

echo 'single quotes'
echo "double quotes"
single quotes
double quotes

The two styles are identical. They are not the two different things they are in some other languages. Both process escape sequences, both interpolate expressions, and both may span several lines. Nothing at all changes except which quote character ends the string:

var name = 'Zuri'

echo 'escape \t and interpolate ${name}'
echo "escape \t and interpolate ${name}"
escape 	 and interpolate Zuri
escape 	 and interpolate Zuri

Pick whichever avoids escaping. A string containing an apostrophe is easiest in double quotes; a string containing a quotation mark is easiest in single quotes:

echo "it's clearer this way"
echo 'she said "hello"'
it's clearer this way
she said "hello"

Most Zuri code, and all of the standard library, uses single quotes by default and switches to double quotes when the content calls for it.

Multi-line Strings

A literal may simply contain newlines. There is no separate triple-quoted form and no continuation character:

var message = 'Dear reader,

Thank you for the strings.'

echo message
Dear reader,

Thank you for the strings.

Everything between the quotes is part of the string, indentation included. When a long string is indented inside a function, the leading spaces on each line are in the string too, which is why it is advisable that code that builds indented text should assemble it with join() rather than writing one long literal.

Escape Sequences

A backslash starts an escape. The full set is:

EscapeProduces
\0the null character, U+0000
\athe alert (bell) character, U+0007
\bbackspace, U+0008
\thorizontal tab, U+0009
\nline feed, U+000A
\vvertical tab, U+000B
\fform feed, U+000C
\rcarriage return, U+000D
\\a single backslash
\'a single quote
\"a double quote
\xNNthe byte with hexadecimal value NN
\uNNNNthe code point U+NNNN
\UNNNNNNNNthe code point U+NNNNNNNN

The three numeric forms take exactly the number of hex digits shown — two, four and eight — and the digits may be upper or lower case:

echo 'hex byte: \x41'
echo 'code point: \u00e9'
echo 'astral: \U0001F600'
hex byte: A
code point: é
astral: 😀

\u covers everything up to U+FFFF. Anything above that — emoji, most historic scripts, mathematical alphanumerics — needs \U and its eight digits. Zero-pad to fill them.

Unknown Escapes Are Left Alone

A backslash followed by anything not in the table above is not an error, and it is not silently dropped either. The backslash and the character both survive:

echo 'a regex: \d+ and \w+'
echo 'a path: C:\temp\notes'
a regex: \d+ and \w+
a path: C:	emp
otes

The first line came through untouched: \d and \w are not escapes, so both backslashes survived, which is exactly why regular expressions can be written as plain strings.

The second line did not. \t is an escape, so C:\temp became C: followed by a tab and then emp; \n is an escape too, so \notes became a newline followed by otes. One path, two silent substitutions. The rule is easy to state and easy to forget: unknown escapes pass through, known ones do not, and nothing in the way you write them tells you which is which.

For Windows paths and regular expressions — the two places this bites — double the backslashes:

echo 'a path: C:\\temp\\notes'
a path: C:\temp\notes

Escaping the Quote You Chose

You only need to escape the quote character that would end the string:

echo 'it\'s escaped'
echo "she said \"hi\""
it's escaped
she said "hi"

Escaping the other quote is harmless but pointless, and because an unnecessary escape is an unknown escape, the backslash stays in the result:

echo 'this \" keeps its backslash'
this \" keeps its backslash

That is a good reason to choose the quote style that lets you write the text plainly and escape nothing at all.

Interpolation

${...} inside a string evaluates the expression between the braces and splices its text into the result:

var name = 'Ada'
var year = 1843

echo 'In ${year}, ${name} wrote the first program.'
In 1843, Ada wrote the first program.

This works in both quote styles, and it is the way most Zuri code builds a string. The section after next covers +, which does the same job with more punctuation.

Any Expression Fits

The braces take a whole expression, not just a variable name. Arithmetic, method calls, indexing, comparisons and conditionals all work:

var items = ['a', 'b', 'c']
var user = { name: 'ada', admin: true }

echo 'count: ${items.length()}'
echo 'first: ${items[0].upper()}'
echo 'math: ${(2 ** 10) - 24}'
echo 'role: ${user.admin ? 'admin' : 'user'}'
echo 'name: ${user.name.capitalize()}'
count: 3
first: A
math: 1000
role: admin
name: Ada

Notice the third line. The quotes inside ${...} are the same quote character that opened the string, and it still parses correctly, because the expression inside the braces is lexed as its own piece of code rather than as text. You do not need to switch quote styles or escape anything to put a string literal inside an interpolation.

Interpolations Nest

A string literal inside ${...} is a string literal like any other, which means it may contain interpolations of its own, to any depth:

var n = 1

echo 'outer ${ 'inner ${n + 1}' }'
echo '${ '${ '${2 ** 3}' }' }'
outer inner 2
8

This is not a curiosity. It is what makes a conditional inside a string able to produce formatted text rather than a bare word, which comes up constantly in messages meant for a person to read:

var user = { name: 'ada', unread: 3 }

echo 'Hi ${user.name.capitalize()}, you have ${user.unread > 0 ? '${user.unread} new ${user.unread == 1 ? 'message' : 'messages'}' : 'nothing new'}.'
Hi Ada, you have 3 new messages.

Three levels of interpolation are at work there: the outer message, the branch that chooses between a count and “nothing new”, and the branch that picks the singular or the plural. Each one is an ordinary string literal in an ordinary expression, and each one uses single quotes, because the expression inside ${...} is lexed as code and its quotes never collide with the ones around it.

Nesting is equally useful inside a callback, where the inner string is building one element of a larger result:

var items = ['a', 'b']

echo 'list: ${', '.join(items.map(@(i) => '<${i}>'))}'
list: <a>, <b>

Depth costs readability quickly. When a line like the message above stops being scannable, lift the inner pieces into local variables on the lines before it; the language is happy either way.

What Each Type Looks Like

Interpolation renders a value the same way echo does:

var settings = { a: 1 }

echo 'number: ${1 / 3}'
echo 'list: ${[1, 2]}'
echo 'dict: ${settings}'
echo 'nil: ${nil}'
echo 'bool: ${true}'
number: 0.3333333333333333
list: [1, 2]
dict: {a: 1}
nil: nil
bool: true

A dictionary literal written directly inside ${...} does not parse, because its closing brace runs into the interpolation’s closing brace:

echo 'dict: ${{ a: 1 }}'
SyntaxError: Expected '}' after dictionary

Put it in a variable first, as above. A list literal has no such problem, since ] and } are different characters.

There is one important exception, and it catches everyone once. Interpolating a class instance does not call its to_string() method:

class Point {

  @new(x, y) {
    self.x = x
    self.y = y
  }

  to_string() {
    return '(${self.x}, ${self.y})'
  }
}

var p = Point(3, 4)

echo 'implicit: ${p}'
echo 'explicit: ${p.to_string()}'
implicit: <instance of Point>
explicit: (3, 4)

Call the method yourself. This applies to 'text ' + p too. Interpolation and + never call a method on your behalf to produce text; if you want to_string() to run, write it. echo is the exception, through @to_string(), which Decorated Methods covers.

Writing a Literal ${

A backslash does not escape an interpolation. \${x} is an unknown escape, so the backslash survives and the interpolation still does not happen — which is rarely what anyone wants:

echo 'literal attempt: \${x}'
literal attempt: \${x}

To produce the two characters ${ followed by text, split the string so the $ and the { are never adjacent inside one literal:

echo 'shell syntax: $' + '{HOME}'
shell syntax: ${HOME}

A lone $ is not special, so only the exact sequence ${ needs this treatment:

echo 'price: $5.00'
echo 'total: $${12 + 8}'
price: $5.00
total: $20

Joining Strings Together

Concatenation with +

+ joins two strings end to end:

echo 'Hello, ' + 'world'
echo 'a' + 'b' + 'c'
Hello, world
abc

When one side is a string, the other side is converted to its text form first. This works in both directions, and for every built-in type:

echo 'n=' + 5
echo 5 + 'n'
echo 'pi is about ' + 3.14
echo 'flag: ' + true
echo 'missing: ' + nil
echo 'items: ' + [1, 2]
n=5
5n
pi is about 3.14
flag: true
missing: nil
items: [1, 2]

Read the second line again: 5 + 'n' gives '5n', not an error and not arithmetic. A string on either side turns the whole expression into concatenation. That is worth knowing when a value arrives from somewhere you do not control:

var quantity = '2'

echo quantity + 3
echo quantity.to_number() + 3
23
5

The first line is text, the second is arithmetic, and nothing in the expression tells you which you are getting. When a number must be a number, convert it at the point it enters your program rather than hoping.

The one type that does not convert usefully is a class instance:

class Point {

  @new(x, y) {
    self.x = x
    self.y = y
  }

  to_string() {
    return '(${self.x}, ${self.y})'
  }
}

echo 'at ' + Point(1, 2).to_string()
at (1, 2)

Concatenating the instance directly would produce <instance of Point>, for the same reason interpolation does: Zuri does not call to_string() for you.

Concatenation Versus Interpolation

Both produce the same string, so choose on readability:

var host = 'localhost'
var port = 8080

echo 'connecting to ' + host + ':' + port + '/health'
echo 'connecting to ${host}:${port}/health'
connecting to localhost:8080/health
connecting to localhost:8080/health

Interpolation wins whenever there is more than one hole to fill: the literal text stays in one piece, the quotes do not multiply, and there is no chance of losing a separator between two + signs. Reach for + when you are gluing exactly two pieces together, or when one of them is already a variable holding a complete string.

Repetition with *

* repeats a string a whole number of times:

echo 'ab' * 3
echo '-' * 20
echo '  ' * 2 + 'indented'
ababab
--------------------
    indented

That second line is the reason this operator earns its place. Separators, rules, indentation and padding are all one short expression instead of a loop.

The rules for the count are worth stating exactly:

  • A count of 1 returns the string unchanged.
  • A count of 0 returns the empty string.
  • A negative count also returns the empty string, rather than raising.
  • A fractional count is truncated toward zero, so * 2.7 repeats twice.
echo '[' + 'ab' * 1 + ']'
echo '[' + 'ab' * 0 + ']'
echo '[' + 'ab' * -1 + ']'
echo '[' + 'ab' * 2.7 + ']'
[ab]
[]
[]
[abab]

The negative case is the one to watch. '-' * (width - label.length()) silently produces nothing when the label is longer than the width, and no error tells you so. When the count is computed, clamp it:

def rule(width, label) {
  var padding = (width - label.length()).max(0)

  return label + '-' * padding
}

echo rule(10, 'ab')
echo rule(10, 'a much longer label')
ab--------
a much longer label

Unlike +, repetition does not commute. The string has to be on the left:

echo 3 * 'ab'
Unhandled TypeError: operator '*' not defined for call signature (number, string)

Building a String from Many Pieces

For more than a handful of pieces, neither operator is the right tool. Collect them in a list and join():

var names = ['ada', 'grace', 'alan']

echo ', '.join(names)
echo '\n'.join(names.map(@(n) => '- ' + n.capitalize()))
ada, grace, alan
- Ada
- Grace
- Alan

join() lives on the separator, not on the list, and it converts each element to text on the way — so a list of numbers joins as readily as a list of strings:

echo '-'.join([1, 2, 3])
1-2-3

Note also that join() puts the separator between elements, never at either end, which is exactly the behaviour a hand-written loop usually gets wrong on the last item.

Inspecting

var s = 'Hello, Zuri'

echo s.length()
echo s.is_empty()
echo s.index_of('Zuri')
echo s.count('l')
echo s.starts_with('Hello')
echo s.ends_with('Zuri')
echo s.contains('Zuri')
11
false
7
2
true
true
true

index_of() returns -1 when there is no match, and takes an optional second argument to start the search from.

last_index_of() searches from the other end, which is what you want whenever the interesting separator is the final one:

var path = 'src/vm/value.rs'

echo path.index_of('/')
echo path.last_index_of('/')
echo path[path.last_index_of('/') + 1, path.length()]
echo path.last_index_of('\\')
3
6
value.rs
-1

Both take a second argument, and in both it bounds where a match may begin. That makes the pair split a string at one index: for any n, index_of(str, n) finds the first match at or after n and last_index_of(str, n) the last match at or before it.

There is a family of character-class predicates, each true when every character in the string qualifies:

echo 'abc'.is_alpha()
echo '123'.is_number()
echo 'ABC'.is_upper()
echo '  '.is_space()
true
true
true
true

The full set is is_alpha, is_alnum, is_number, is_lower, is_upper, is_space, plus is_empty.

Changing Case

var s = 'Hello, Zuri'

echo s.upper()
echo s.lower()
echo s.capitalize()
echo s.title()
HELLO, ZURI
hello, zuri
Hello, zuri
Hello, Zuri

capitalize() uppercases the first character and lowercases the rest. title() does that to every word.

For comparing text from users, reach for case_fold() rather than lower(). Case folding is the Unicode operation defined for case-insensitive matching, and it handles cases simple lowercasing gets wrong:

echo 'Straße'.case_fold()
strasse

Trimming and Padding

echo '[' + '  pad  '.trim() + ']'
echo '[' + '  pad  '.ltrim() + ']'
echo '[' + '  pad  '.rtrim() + ']'
echo 'xxhixx'.trim('x')
echo '-=hi=-'.trim('-=')
[pad]
[pad  ]
[  pad]
hi
hi

With no argument, all three strip whitespace: spaces, tabs, newlines, carriage returns, vertical tabs and form feeds. Given a string, they strip every character in it, in any order, so '-=hi=-'.trim('-=') is hi.

echo 'x'.lpad(5, '.')
echo 'x'.rpad(5, '.')
....x
x....

lpad and rpad take a target width and an optional fill character that defaults to a space. A string already at or over the width comes back unchanged.

Splitting and Joining

echo 'Hello, Zuri'.split(', ')
echo ', '.join(['a', 'b', 'c'])
echo 'abc'.to_list()
echo 'line1\nline2'.lines()
[Hello, Zuri]
a, b, c
[a, b, c]
[line1, line2]

join lives on the separator, not on the list. to_list() splits into single characters. lines() splits on newlines and handles both \n and \r\n.

Replacing

echo 'Hello, Zuri'.replace('Zuri', 'World')
Hello, World

Every occurrence is replaced. The pattern may be a regular expression, and the replacement can refer back to capture groups with $1, $2 or ${1}:

echo 'John Smith'.replace('/(\w+) (\w+)/', '$2, $1')
Smith, John

The optional third argument turns regex handling off, so a pattern that would otherwise be read as a regex is matched literally, delimiters included:

echo 'a/b/c'.replace('/b/', 'X')
echo 'a/b/c'.replace('/b/', 'X', false)
a/X/c
aXc

When the replacement depends on what was matched, use replace_with(), which calls a function for each match:

echo 'a1b2'.replace_with('/[0-9]/', @(m) => '<' + m + '>')
a<1>b<2>

Regular Expressions

Zuri’s regular expressions are PCRE2, the same engine that powers regular expressions in PHP, R and a long list of other tools. It is the real library rather than a subset of it, so named groups, backreferences, lookahead, lookbehind, atomic groups, possessive quantifiers, conditionals and recursion all work exactly as PCRE2 documents them.

This section covers how a pattern is written in Zuri, which methods accept one, what each modifier does, and the handful of places where Zuri’s surface differs from what you may expect. It does not teach regular expression syntax itself; any PCRE or Perl reference applies unchanged.

Writing a Pattern

A pattern is an ordinary string whose first character is a non-word character, repeated to close the pattern, with any modifier letters after the closing delimiter:

'/[a-z]+/'
'/[a-z]+/i'
'#\d{3}#'
'!https?://\S+!i'

A “non-word character” is anything that is not a letter, a digit or an underscore. Forward slash is the conventional choice, and every example in this book uses it, but the delimiter is yours to pick — which is worth doing when the pattern itself is full of slashes:

echo 'see https://example.com/docs now'.match('!https?://\S+!')
{0: https://example.com/docs}

Because the pattern is a string, the backslashes in it are subject to Zuri’s own escape rules first. This is almost never a problem, because \d, \w, \s, \b and \S are not escape sequences and therefore pass through untouched. The exceptions are the letters that are escapes — \t, \n, \r, \f, \v, \a, \b, \0, \x and \u. Of those, \t and \n mean the same thing to both, so they are harmless; \b (word boundary in a pattern, backspace in a string) is the one that will catch you:

echo 'one two'.match('/\btwo\b/')
echo 'one two'.match('/\\btwo\\b/')
false
{0: two}

Double the backslash whenever the pattern needs \b.

A Plain String Is a Literal

Every method in this section also accepts a plain string, and treats it as a literal substring rather than a pattern:

echo 'a.b'.replace('.', 'X')
echo 'a.b'.replace('/./', 'X')
aXb
XXX

The first call replaced the single full stop. The second compiled . as a pattern, where it means “any character”, and replaced all three.

That distinction is decided purely by whether the string looks like a delimited pattern. When you have text from elsewhere that might accidentally look like one, replace() takes a third argument that turns pattern handling off:

echo 'a/b/c'.replace('/b/', 'X')
echo 'a/b/c'.replace('/b/', 'X', false)
a/X/c
aXc

With false, the delimiters are matched as characters like anything else.

Which Methods Take a Pattern

Five, and only five:

MethodWhat it does with the pattern
match(pattern)the first match and its groups, or false
matches(pattern)every match, grouped by capture group
replace(pattern, with, use_regex)replaces every match
replace_with(pattern, callback)replaces every match with a call
split(pattern)splits on every match

Everything else — contains(), index_of(), count(), starts_with(), ends_with(), trim() — takes a literal string only. A regex-looking argument passed to one of those is matched as the literal characters it contains:

echo 'a1b'.contains('/[0-9]/')
echo 'a1b'.index_of('/[0-9]/')
echo 'a1b'.count('/[0-9]/')
false
-1
0

All three answered about the six-character string /[0-9]/, which is not in a1b. When you need a pattern there, reach for match() instead.

match() and matches()

match() finds the first match and returns it as a dictionary under key 0. It returns false — not nil — when nothing matches:

echo 'one two'.match('/\w+/')
echo 'nope'.match('/zzz/')
echo typeof('nope'.match('/zzz/'))
{0: one}
false
bool

Both are falsy, so if s.match(p) { ... } reads correctly. An equality test against nil does not, and this is the single most common regex mistake in Zuri code.

The dictionary holds the pattern’s capture groups as well, each under its number, and a named group under its name too. A group that takes no part in the match is nil:

echo '2024-01'.match('/(?<year>\d+)-(\d+)/')
echo 'ab'.match('/([a-z]+)(\d+)?/')
{0: 2024-01, 1: 2024, 2: 01, year: 2024}
{0: ab, 1: ab, 2: nil}

matches() finds every match and groups the results by capture group. Key 0 holds the whole matches, key 1 the first group’s, and so on — each as a list, parallel across the keys:

echo 'a1b2c3'.matches('/[0-9]/')
echo 'a1b2'.matches('/([a-z])([0-9])/')
{0: [1, 2, 3]}
{0: [a1, b2], 1: [a, b], 2: [1, 2]}

In the second result, match zero is a1 with groups a and 1, and match one is b2 with groups b and 2. Read down the lists, not across.

When nothing matches, matches() returns {0: []} rather than false. The two methods differ here, so test matches() with is_empty() on key zero:

var found = 'nope'.matches('/zzz/')

echo found
echo found[0].is_empty()
{0: []}
true

Capture Groups in a Replacement

replace() refers to a capture group with $1, $2 and so on:

echo 'John Smith'.replace('/(\w+) (\w+)/', '$2, $1')
Smith, John

Do not write ${1}. The braced form is Zuri’s own string interpolation, and it is consumed before replace() ever sees it. The expression 2 evaluates to the number two, so the replacement string is already '2, 1' by the time it arrives:

echo 'John Smith'.replace('/(\w+) (\w+)/', '${2}, ${1}')
2, 1

Use the bare $1 form throughout. This is one of the few places where Zuri’s interpolation and another language’s syntax collide, and it fails quietly rather than loudly.

Named Groups

Named groups are supported in the pattern, because PCRE2 supports them:

echo 'John Smith'.matches('/(?<first>\w+) (?<last>\w+)/')
{0: [John Smith], 1: [John], 2: [Smith]}

The results are keyed by position, not by name. A name in a pattern documents the group and lets a backreference such as \k<first> refer to it; it does not change how the result is keyed. Count the groups to find the one you want.

Computing the Replacement

When the replacement depends on what matched, replace_with() calls a function for each match:

echo 'a1b2'.replace_with('/[0-9]/', @(m) => '<' + m + '>')
a<1>b<2>

The callback receives the whole match first, then one argument per capture group, then the offset of the match, then the whole subject string. Take only the arguments you need, since extra arguments are dropped:

echo 'aXbXc'.replace_with('/X/', @(match, offset) => '[${offset}]')
a[1]b[3]c

Splitting on a Pattern

split() accepts a pattern, which is how you split on “one or more of something” rather than on an exact separator:

echo 'a1b22c'.split('/[0-9]+/')
echo 'a, b,c ,  d'.split('/\s*,\s*/')
[a, b, c]
[a, b, c, d]

The Modifiers

Modifier letters go after the closing delimiter, in any order and any combination. /[a-z]+/mi is both multi-line and case-insensitive.

ModifierEffect
iCase-insensitive matching.
mMulti-line: ^ and $ match at every line boundary, not only at the start and end of the subject.
sDot-all: . matches a newline as well as everything else.
xExtended: unescaped whitespace in the pattern is ignored, and # starts a comment to end of line.
uUnicode properties: \d, \w, \s and friends become Unicode-aware instead of ASCII-only.
UUngreedy: quantifiers become lazy by default, and ? after one makes it greedy.
AAnchored: a match is only accepted if it begins exactly where the search started.
JAllow two capture groups in one pattern to share a name.

Each one in use:

echo 'HELLO'.match('/hello/i')
echo 'a\nb'.matches('/^./m')
echo 'a\nb'.match('/a.b/s')
echo 'abc'.match('/ a b c /x')
echo 'héllo'.matches('/\w/u')
echo 'aaa'.match('/a+?/U')
echo 'abc'.match('/abc/A')
echo 'xabc'.match('/abc/A')
{0: HELLO}
{0: [a, b]}
{0: a
b}
{0: abc}
{0: [h, é, l, l, o]}
{0: aaa}
{0: abc}
false

Read the last two together: the pattern matched when it was at the start of the subject and failed when it was not, which is what A is for. Without it, the second would have matched at offset one.

u is worth a second look as well. Without it, \w is ASCII-only and é is not a word character; with it, the Unicode properties apply and it is. Matching itself is always per-character rather than per-byte, so indexes and offsets are correct on multi-byte text regardless of u.

The One Exception to PCRE2 Compatibility

PCRE2’s D (dollar-endonly) modifier has no effect in Zuri. $ always matches before a trailing newline as well as at the absolute end of the subject, and there is no way to change that. The letter is accepted rather than rejected, so a pattern carrying it compiles and runs; it simply does not do anything.

The same is true of any other unrecognised modifier letter: it is accepted and ignored rather than raising. A typo in a modifier is therefore silent, which is worth remembering when a pattern behaves as though a flag you wrote is not set.

Everything else in the table above is supported and behaves exactly as PCRE2 documents it.

Invalid Patterns

A pattern that PCRE2 cannot compile raises, naming the pattern and the engine’s own explanation:

catch {
  echo 'text'.match('/(unclosed/')
} as e {
  echo e.message
}
invalid regular expression '(unclosed': PCRE2: error compiling pattern at offset 9: missing closing parenthesis

Patterns are compiled once and reused, so repeating the same pattern in a loop costs nothing after the first time through.

Conversion

echo '5'.to_number() + 1
echo 'ff'.to_number(16)
echo 'H'.ord()
echo 72.chr()
echo 'Hi'.to_bytes()
6
255
72
H
(48 69)

to_number() takes an optional base. ord() requires a single-character string and gives its code point; chr() on a number goes the other way.

to_number() Never Fails

This is the one conversion behaviour worth memorising. to_number() does not raise, and it does not produce NaN. It reads the first number written in the string, and text with no number in it at all becomes zero:

echo '42'.to_number()
echo '3.5'.to_number()
echo '-5'.to_number()
echo 'eighty'.to_number()
echo ''.to_number()
echo '12abc'.to_number()
echo '  7  '.to_number()
echo '96.3 of 31'.to_number()
42
3.5
-5
0
0
12
7
96.3

Look at the last three. What surrounds the number is ignored, so '12abc' is 12, the spaces around ' 7 ' are simply not part of it, and only the first number is read however many follow.

That cuts both ways. '3 apples' giving 3 is usually what was meant; '2026-09-16' giving 2026 usually is not. A number is taken to start at a digit, or at a sign or decimal point directly in front of one, so '-5' is negative five and 'a - 42', where the sign stands apart, is 42.

There is no error to catch and no sentinel to test, so a form field a user left blank and a form field they filled with '0' produce the same number, and so do 'eighty' and '0'. When the difference matters, check the text before converting it:

def to_count(text) {
  var trimmed = text.trim()

  if !trimmed.match('/^\d+$/') {
    raise ValueError('not a whole number: ${text}')
  }

  return trimmed.to_number()
}

echo to_count('  12  ')

catch {
  to_count('twelve')
} as e {
  echo e.message
}
12
not a whole number: twelve

is_number() is a narrower test than it looks — it is true only for a string of digits, so '3.5' and '-5' both fail it. A regular expression is the reliable check for anything beyond unsigned integers.

Walking a String

A string is iterable, one character at a time, and the same four forms available for a list work here.

for, One Character at a Time

for ch in 'hey' {
  echo ch
}
h
e
y

This is the form to reach for by default. Each ch is a one-character string, not a code point number — use ord() when you want the number.

for With Two Variables, for the Position

for index, ch in 'hey' {
  echo '${index}: ${ch}'
}
0: h
1: e
2: y

The key comes first and the value second, and for a string the key is the character position.

iter, When the Position Drives the Walk

var s = 'hey'

iter var i = s.length() - 1; i >= 0; i-- {
  echo s[i]
}
y
e
h

s[i] indexes by character, not by byte, so this is correct on text that is not ASCII. Use iter when you need to skip, step backwards, or look at s[i + 1] from inside the body.

each(), With a Function

'Hi'.each(@(ch, index) {
  echo '${index}:${ch}'
})
0:H
1:i

As everywhere else, the callback receives the value first and the index second — the opposite order to for.

each_line() and lines(), for Text in Lines

var doc = 'first\nsecond\nthird'

doc.each_line(@(line, index) {
  echo '${index}: ${line}'
})

echo doc.lines()
0: first
1: second
2: third
[first, second, third]

each_line() calls the function once per line; lines() hands you the whole list so you can for over it, filter it, or count it. Both split on \n and \r\n, so both read a file written on any platform.

Ordering and Comparison

== compares two strings by value, which is almost always what you want:

echo 'abc' == 'abc'
echo 'abc' == 'ABC'
true
false

The ordering operators are a different story. <, <=, > and >= are defined for numbers only, and comparing two strings with one raises a TypeError. Use compare(), which returns -1, 0 or 1:

echo 'abc'.compare('abd')
echo 'abc'.compare('abc')
echo 'abd'.compare('abc')
-1
0
1

That three-way result is exactly the shape a sort comparison wants, and it compares by code point, so it is stable and locale-independent.

Strings Are Immutable

Every method in this section returns a new string. Nothing mutates in place:

var original = 'hello'
var shouted = original.upper()

echo original
echo shouted
hello
HELLO

This is worth internalising, because it is the source of the single most common string mistake:

var name = '  ada  '

name.trim()
echo '[${name}]'

name = name.trim()
echo '[${name}]'
[  ada  ]
[ada]

The first trim() produced a trimmed string and threw it away. Assign the result, or chain onto it.

Immutability also means you can pass a string anywhere without copying it defensively. Nothing you hand a string to can change the one you still hold.

A Worked Example

Here is a small parser that puts most of this section together. It takes a block of key = value configuration text and produces a dictionary, skipping blank lines and comments, and tolerating whatever spacing the author used.

def parse_config(text) {
  var config = {}

  for line in text.lines() {
    var trimmed = line.trim()

    if trimmed.is_empty() or trimmed.starts_with('#') {
      continue
    }

    var split_at = trimmed.index_of('=')

    if split_at == -1 {
      continue
    }

    var key = trimmed[0, split_at].trim()
    var value = trimmed[split_at + 1, trimmed.length()].trim()

    config[key] = value
  }

  return config
}

var source = '# server settings
host = localhost
port   =   8080

# empty lines and comments are skipped
name = zuri app
'

var config = parse_config(source)

echo config.host
echo config.port.to_number() + 1
echo config.name
echo config.length()
localhost
8081
zuri app
3

Three things in there are worth naming. lines() handles both \n and \r\n, so the same code reads a file written on Windows. index_of() returning -1 is checked explicitly rather than relied on for truthiness, because -1 is falsy and so is 0 — and 0 is a legitimate position. And every value arrives as a string, because that is what text is; to_number() is how you leave.

Numbers

There is one numeric type, number, and it is a 64-bit IEEE-754 double. Integers and fractions are the same type, and the distinction you care about is whether a particular value happens to be integral.

Writing a Number

There are five ways to write a numeric literal, and they all produce the same type:

echo 42
echo 3.5
echo 0b1011
echo 0c17
echo 0xff
echo 6.02e23
echo 1.5e-3
42
3.5
11
15
255
602000000000000000000000
0.0015

0b is binary, 0c is octal, and 0x is hexadecimal. The letters in a hexadecimal literal may be either case, so 0xff and 0xFF are the same number. Note that octal is 0c, not the 0o some other languages use.

The exponent form takes e or E, and the exponent may be negative. 1e3 is 1000; the mantissa needs no decimal point. Note that the exponent is a way of writing the literal, not a property the value keeps: 6.02e23 prints as its full decimal expansion, because that is the number it is.

Digit Separators

An underscore between digits is ignored, which makes long numbers readable:

echo 1_000_000
echo 1_0.5_5
1000000
10.55

Separators work in decimal literals only. 0xdead_beef and 0b1010_1010 are syntax errors, not clever formatting.

Two Forms That Do Not Exist

A literal needs a digit on both sides of its decimal point. Neither of these parses:

echo .5
echo 5.

Write 0.5 and 5.0. The second case matters more than it looks, because 5. is also how a method call on a literal starts — which is the subject of the next-but-one section.

Methods, Not Functions

Everything you would reach into a math library for is a method on the number:

echo 2.sqrt()
echo 8.log2()
echo 100.log10()
echo 100.log()
echo 1.exp()
echo 2.cbrt()
1.4142135623730951
3
2
4.605170185988092
2.718281828459045
1.2599210498948732

log() is the natural logarithm. There is also log1p() and expm1() for the precision-sensitive forms near zero.

The full trigonometric set is there: sin, cos, tan, asin, acos, atan, atan2, and the hyperbolic sinh, cosh, tanh, asinh, acosh, atanh.

echo 1.atan2(1)
0.7853981633974483

Calling a Method on a Literal

A numeric literal takes a method directly. No parentheses, no temporary variable, in every base and with a decimal point or without:

echo 2.sqrt()
echo 3.7.round()
echo 255.hex()
echo 0xff.bin()
echo 1e3.int()
echo 2n.bits()
1.4142135623730951
4
ff
11111111
1000
2

There is exactly one case where you need parentheses, and it is a negative literal. A method call binds tighter than unary minus, so the minus applies to the result rather than to the number:

echo -3.abs()
echo (-3).abs()
-3
3

-3.abs() is -(3.abs()), which is -3. When the receiver is negative, parenthesise it — or put it in a variable, where the question disappears:

var n = -3

echo n.abs()
3

The same rule covers any expression you want to call a method on: wrap it, because the call would otherwise bind to the last term alone.

echo (1 / 0).is_inf()
echo (2 ** 10).hex()
true
400

Rounding

echo 2.5.ceil()
echo 2.5.floor()
echo 2.5.trunc()
echo 2.5.int()
echo (-2.5).int()
3
2
2
2
-2

trunc() and int() both drop the fractional part, rounding towards zero.

round() rounds half away from zero:

echo 2.5.round()
echo 3.5.round()
echo (-2.5).round()
3
4
-3

fixed() rounds to a number of decimal places and gives you a number back:

echo 2.567.fixed(2)
2.57

fraction() gives the digits after the decimal point as a whole number:

echo 2.567.fraction()
567

Sign, Magnitude and Comparison

echo (-5).sign()
echo (-3).abs()
echo 5.max(9)
echo 5.min(9)
-1
3
9
5

sign() is -1, 0 or 1.

Bases and Characters

echo 255.bin()
echo 255.oct()
echo 255.hex()
echo 65.chr()
11111111
377
ff
A

chr() turns a code point into a one-character string; 'A'.ord() goes back the other way.

The Special Values

echo (0 / 0).is_nan()
echo (1 / 0).is_inf()
echo 1.is_finite()
true
true
true

NaN is falsy, like zero, so var x = a / b or fallback replaces the NaN from a 0 / 0 with the fallback. Test with is_nan() when you need to tell the two apart.

NaN is also not equal to itself, as the standard requires, so x == x is a valid way to spot one.

Other Methods

factorial() for small integers, to_string(), to_bool() and to_bigint() for conversion.

echo 5.factorial()
echo 17.to_bigint()
120
17n

The math Module

Constants live in math, because a constant is not a method on anything:

import math

echo math.PI
echo math.E
echo math.Infinity
echo math.NaN
3.141592653589793
2.718281828459045
inf
NaN

It also carries LOG_2, LOG_10, LOG_2_E, LOG_10_E, ROOT_2, ROOT_3 and ROOT_HALF.

Bigints

When 2^53 is not enough, use a bigint. Write one with an n suffix:

var a = 2n ** 100n
echo a
echo a.bits()
1267650600228229401496703205376n
101

Bigints are arbitrary precision. They never overflow and never lose a digit.

They also never mix with numbers implicitly:

catch {
  echo 5n + 3
} as e {
  echo e.message
}
operator '+' not defined for call signature (bigint, number)

Convert explicitly, in whichever direction you need:

echo 5.to_bigint() * 2n
echo (2n ** 100n).to_number()
10n
1267650600228229400000000000000

Going to number is lossy once you are past 2^53, which is the whole reason bigints exist. Going the other way is exact.

The same separation holds in type annotations. bigint is a type name alongside number and int, and it accepts nothing else:

def scale(n: bigint, factor: number) {
  return n * factor.to_bigint()
}

echo scale(2n, 50)

catch {
  scale(2, 50)
} as e {
  echo e.message
}
100n
scale() expects parameter 'n' (argument 1) to be a bigint, got number

The number-theory methods are the reason bigints are worth having:

echo 100n.gcd(75n)
echo 100n.lcm(75n)
echo 2n.modpow(10n, 1000n)
echo 3n.modinv(11n)
echo 144n.sqrt()
25n
300n
24n
4n
12n

modpow(exponent, modulus) is the operation every public-key algorithm is built on, and it is computed without ever materialising the full power. modinv gives the modular multiplicative inverse.

There is also nth_root(), cbrt(), bit(), set_bit(), trailing_zeros(), is_zero(), is_even(), is_odd(), abs(), sign(), max(), min(), pow(), and bin()/oct()/hex().

A bigint is the right choice when exactness past 2^53 is the point: cryptography, currency in minor units, factorials, identifiers that must survive a round trip. For everything else — measurements, coordinates, counters, ratios — a number is the type you want.

Lists

A list is an ordered, growable sequence. It holds values of any type, including other lists, and it grows and shrinks as you work with it.

var items = [3, 1, 4, 1, 5]
var mixed = [1, 'two', [3], { four: 4 }]
var empty = []

echo items.length()
echo mixed[3].four
echo empty.is_empty()
5
4
true

Lists carry 39 methods, and they fall into six groups: reading, searching, adding and removing, ordering, transforming, and walking. This section takes them in that order. The one distinction to keep in mind throughout is whether a method mutates the list you called it on or returns a new one, because Zuri’s lists do both and the method names do not tell you which.

Reading

var l = [3, 1, 4, 1, 5]

echo l.length()
echo l.is_empty()
echo l[0]
echo l[-1]
echo l.first()
echo l.last()
echo l[1, 3]
5
false
3
5
3
5
[1, 4]

Indexing

An index counts from zero, and a negative index counts back from the end:

var l = ['a', 'b', 'c', 'd']

echo l[0]
echo l[3]
echo l[-1]
echo l[-4]
a
d
d
a

l[-1] is the last element, which saves writing l[l.length() - 1] everywhere.

Indexing out of range raises rather than returning nil:

var l = [1, 2, 3]

catch {
  echo l[99]
} as e {
  echo e.message
}
index 99 out of bounds (length 3)

get() is the form that can fall back, but only when you give it something to fall back to. get(i) with one argument raises exactly like l[i] does:

var l = [1, 2, 3]

echo l.get(1)
echo l.get(99, 'missing')

catch {
  echo l.get(99)
} as e {
  echo e.message
}
2
missing
list index 99 out of range at get()

Slicing

l[a, b] takes the elements from a up to but not including b, and gives you a new list:

var l = [1, 2, 3, 4, 5]

echo l[1, 3]
echo l[, 3]
echo l[3, ]
echo l[-2, ]
[2, 3]
[1, 2, 3]
[4, 5]
[4, 5]

Either bound may be left out: l[, b] starts at the beginning and l[a, ] runs to the end. Negative bounds count from the end, just as an index does.

A slice checks its bounds, exactly as an index does. Running past the end raises rather than returning what it can:

catch {
  echo [1, 2, 3][1, 99]
} as e {
  echo e.message
}
slice bounds 1..99 out of range (length 3)

The one bound that is always legal is length() itself, since a slice’s upper bound is exclusive. l[0, l.length()] is the whole list, and l[3, 3] on a three-element list is empty rather than an error.

var l = [1, 2, 3]

echo l[0, l.length()]
echo l[3, 3]
[1, 2, 3]
[]

This same rule governs string slicing, which is why Strings uses s[split_at + 1, s.length()] rather than a large number to mean “the rest”.

A slice is a copy. Changing one does not touch the other:

var original = [1, 2, 3]
var part = original[0, 2]

part[0] = 99

echo original
echo part
[1, 2, 3]
[99, 2]

Nested Lists

A list holds lists, and indexes chain:

var grid = [[1, 2], [3, 4], [5, 6]]

echo grid[1]
echo grid[1][0]
echo grid.length()
echo grid[1].length()
[3, 4]
3
3
2

grid.length() counts rows, not elements. There is no two-dimensional index; grid[1, 0] is a slice from row one to row zero, which is empty.

Combining Lists

Concatenation with +

+ joins two lists into a new one, leaving both operands alone:

var a = [1, 2]
var b = [3]

echo a + b
echo a
echo b
[1, 2, 3]
[1, 2]
[3]

That is the difference between + and extend(): a + b builds a third list, while a.extend(b) modifies a in place and returns it. Use + when you want to keep the originals, and extend() when you are accumulating.

Repetition with *

* repeats a list a whole number of times:

echo [1, 2] * 3
echo [0] * 5
[1, 2, 1, 2, 1, 2]
[0, 0, 0, 0, 0]

[0] * 5 is the idiom for a fixed-size list of a starting value, and it is worth knowing before you write the loop.

The count behaves exactly as it does for strings: zero and any negative number both produce an empty list, and a fractional count is truncated.

echo [1] * 0
echo [1] * -3
[]
[]

There is one trap. Repetition copies the elements, and when an element is itself a list, all the copies are the same list:

var grid = [[0, 0]] * 3

grid[0][0] = 9

echo grid
[[9, 0], [9, 0], [9, 0]]

One assignment changed every row, because there is only one row. Build nested structure with a loop, or with map():

var grid = 0..3.to_list().map(@(i) => [0, 0])

grid[0][0] = 9

echo grid
[[9, 0], [0, 0], [0, 0]]

Comparing Lists

== compares lists by value, element by element, however deeply they nest:

echo [1, 2] == [1, 2]
echo [1, [2, 3]] == [1, [2, 3]]
echo [1, 2] == [2, 1]
echo [1, 2] == [1, 2, 3]
true
true
false
false

Order matters and length matters. To compare two lists as sets, sort copies of them first, or use the set module from Chapter 13.

Searching

var l = [3, 1, 4, 1, 5]

echo l.contains(4)
echo l.index_of(1)
echo l.count(1)
true
1
2

index_of() gives -1 when the value is absent. last_index_of() searches from the other end:

var l = [3, 1, 4, 1, 5]

echo l.index_of(1)
echo l.last_index_of(1)
echo l.last_index_of(1, 2)
echo l.last_index_of(9)
1
3
1
-1

Both take a second argument bounding where a match may sit, so the pair splits the list at one index: index_of(x, n) finds the first match at or after n, last_index_of(x, n) the last one at or before it. Both compare by value, so a list of dictionaries can be searched with a dictionary literal.

Adding and Removing

These mutate the list in place:

MethodEffect
append(v)add to the end
insert(v, at)insert at a position
extend(other)append every element of another list
pop()remove and return the last element
shift()remove and return the first element
remove(v)remove the first element equal to v
remove_at(i)remove by position
delete(from, to)remove an inclusive range, return how many went
clear()empty it
var l = [3, 1, 4]

l.insert(9, 1)
echo l
echo l.shift()
echo l
[3, 9, 1, 4]
3
[9, 1, 4]
var l = [1, 2, 3, 4, 5]
echo l.delete(1, 3)
echo l
3
[1, 5]

Note that delete() takes a from and a to, both inclusive, and returns the number of elements removed rather than the list.

Ordering

var l = [3, 1, 2]

echo l.sort()
echo l
echo l.reverse()
echo l
[1, 2, 3]
[1, 2, 3]
[3, 2, 1]
[1, 2, 3]

Look closely at those two. sort() mutates and returns the list. reverse() returns a new list and leaves the original alone. That asymmetry is the single most common source of list bugs in Zuri code.

sort() takes no comparator. To sort by a computed key, decorate, sort, and undecorate:

var people = [{ name: 'Ada', age: 36 }, { name: 'Bob', age: 24 }]

var by_age = people
  .map(@(p) => [p.age, p.name])
  .sort()
  .map(@(pair) => pair[1])

echo by_age
[Bob, Ada]

Transforming

Every one of these returns a new list and leaves the receiver untouched:

var l = [1, 2, 3, 4]

echo l.map(@(x) => x * 2)
echo l.filter(@(x) => x > 2)
echo l.unique()
echo l.compact()
echo l.take(2)
echo l.clone()
[2, 4, 6, 8]
[3, 4]
[1, 2, 3, 4]
[1, 2, 3, 4]
[1, 2]
[1, 2, 3, 4]

compact() drops nil entries:

echo [1, nil, 2, nil].compact()
[1, 2]

reduce() folds the list down to one value. With no initial value it starts from the first element:

echo [1, 2, 3].reduce(@(acc, x) => acc + x)
echo [1, 2, 3].reduce(@(acc, x) => acc + x, 100)
6
106

partition() splits in one pass, matches first:

echo [1, 2, 3, 4].partition(@(x) => x % 2 == 0)
[[2, 4], [1, 3]]

Finding

var l = [1, 2, 3]

echo l.find(@(x) => x > 1)
echo l.find_index(@(x) => x > 1)
echo l.find_last(@(x) => x > 1)
echo l.find_last_index(@(x) => x > 1)
echo l.find_all(@(x) => x > 1)
2
1
3
2
[2, 3]

Asking About Every Element

echo [1, 2, 3].every(@(x) => x > 0)
echo [1, 2, 3].some(@(x) => x > 2)
true
true

some() stops at the first match; every() stops at the first failure.

Walking

There are four ways to visit every element, and they are not interchangeable. Pick by what you need in the body.

for, When You Want the Values

for value in ['a', 'b', 'c'] {
  echo value
}
a
b
c

This is the default. Reach for it whenever the position does not matter.

for With Two Variables, When You Want the Index Too

for index, value in ['a', 'b', 'c'] {
  echo '${index}: ${value}'
}
0: a
1: b
2: c

With two variables you get the key first and the value second, and for a list the key is the index. No call to length(), no manual counter.

iter, When You Need Control of the Position

var items = ['a', 'b', 'c']

iter var i = 0; i < items.length(); i++ {
  echo '${i}: ${items[i]}'
}
0: a
1: b
2: c

iter costs more typing and earns it only when the traversal is not one step forward per pass: walking backwards, walking in twos, comparing an element with its neighbour, or advancing the index from inside the body.

var items = ['a', 'b', 'c', 'd']

iter var i = items.length() - 1; i >= 0; i-- {
  echo items[i]
}

iter var i = 0; i < items.length(); i += 2 {
  echo items[i]
}
d
c
b
a
a
c

each(), When You Have a Function Already

[10, 20].each(@(value, index) {
  echo '${index}: ${value}'
})
0: 10
1: 20

each() takes a function, which makes it the one that composes: it chains onto a map() or a filter() without a temporary variable, and it accepts a function you already have by name.

Watch the argument order. each() hands the callback the value first and the index second, which is the opposite of for. Every list callback in this chapter follows the same rule — map, filter, find, some, every all take (value, index) — so it is one thing to remember rather than several. for is a loop over key-value pairs; each is a callback over values that happens to tell you where it is.

break and continue work inside for and iter. They do not exist inside an each() callback; return there ends that one call, not the walk. When you need to stop early, use a loop, or find() / some(), which stop on their own.

Combining

echo ['a', 'b'].zip([1, 2])
echo [1, 2, 3].zip_from([[4, 5, 6]])
[[a, 1], [b, 2]]
[[1, 4], [2, 5], [3, 6]]

zip() pairs with one other list. zip_from() takes a list of lists and zips across all of them at once.

Lists Are References

Assigning a list does not copy it:

var a = [1, 2]
var b = a
b.append(3)
echo a
[1, 2, 3]

Use clone() when you need an independent copy. The clone is shallow: nested lists inside it are still shared.

var original = [[1, 2], 3]
var copy = original.clone()

copy[1] = 99
copy[0].append(4)

echo original
echo copy
[[1, 2, 4], 3]
[[1, 2, 4], 99]

Replacing the top-level 3 affected only the copy. Appending to the nested list affected both, because both lists point at the same inner list. When you need a deep copy, copy the levels you care about yourself.

Mutating or Not: The Summary

This is the table to come back to.

Mutates the receiverReturns a new list
append, insert, extendmap, filter, find_all
pop, shift, remove, remove_atunique, compact, take
delete, clearreverse, clone, zip, zip_from
sortpartition

sort() is the one that catches people out: it is in the left column, and it also returns the list, so var sorted = items.sort() leaves items sorted too. If you need both orders, clone first:

var items = [3, 1, 2]
var sorted = items.clone().sort()

echo items
echo sorted
[3, 1, 2]
[1, 2, 3]

A Worked Example

A tiny report generator: take a list of records, drop the incomplete ones, group what is left, and print a summary. It uses filter, map, reduce, sort and each together, which is how they usually show up.

var sales = [
  { region: 'north', amount: 120 },
  { region: 'south', amount: 80 },
  { region: 'north', amount: 45 },
  { region: 'east', amount: nil },
  { region: 'south', amount: 200 },
]

def totals_by_region(records) {
  var totals = {}

  records
    .filter(@(r) => r.amount != nil)
    .each(@(r) {
      totals[r.region] = totals.get(r.region, 0) + r.amount
    })

  return totals
}

var totals = totals_by_region(sales)
var grand = totals.values().reduce(@(acc, n) => acc + n, 0)

totals
  .to_list()[0]
  .sort()
  .each(@(region) {
    echo '${region.rpad(6)} ${totals[region]}'
  })

echo 'total  ${grand}'
north  165
south  280
total  445

east is absent from the report, not present with a zero, because its one record was filtered out before any accumulating happened and nothing ever created the key. That is usually what you want from a report; when it is not, seed the dictionary with every region first.

Two other details are worth pulling out. totals.get(r.region, 0) supplies a starting value for a key that does not exist yet, which is what turns a dictionary into an accumulator. And totals.to_list()[0] takes the keys — to_list() returns keys and values as two parallel lists — which are then sorted so the report comes out in a stable order rather than in whatever order the records happened to arrive.

Dictionaries

A dictionary maps keys to values and remembers the order you inserted them.

var user = { name: 'Ada', age: 36 }
var empty = {}

A bare word key is a string, so { name: ... } and { 'name': ... } are the same. Keys may also be numbers:

echo { 1: 'one', 2.5: 'two-point-five' }
{1: one, 2.5: two-point-five}

A key that is a bare identifier is always taken as that literal name. To compute a key, make it an expression the parser cannot mistake for a name, which in practice means wrapping it in parentheses:

var key = 'dynamic'

echo { (key): 1 }
echo { 'a' + 'b': 2 }
{dynamic: 1}
{ab: 2}

Assigning through brackets is usually clearer:

var d = {}
d[key] = 3
echo d
{dynamic: 3}

Shorthand Keys

When a key and the variable holding its value have the same name, write it once:

var name = 'Ada'
var age = 36

echo { name, age }
{name: Ada, age: 36}

{ name, age } means exactly { name: name, age: age }. This comes up constantly in functions that build a result out of locals they have just computed, and Zuri code uses the shorthand whenever the names line up.

Nesting

A value may be another dictionary, or a list, to any depth. Access chains:

var config = {
  server: { host: 'localhost', port: 8080 },
  tags: ['web', 'internal'],
}

echo config.server.host
echo config['server']['port']
echo config.tags[0]
localhost
8080
web

Dot and bracket access mix freely, because they are the same operation. Use the dot when the key is a fixed identifier and brackets when it is computed, contains punctuation, or is a number.

Reading

Dot and bracket access are the same operation:

var user = { name: 'Ada', age: 36 }

echo user.name
echo user['name']
echo user.length()
echo user.is_empty()
echo user.keys()
echo user.values()
Ada
Ada
2
false
[name, age]
[Ada, 36]

Reading a key that is not there raises a PropertyError. get() is the safe form:

echo user.get('name')
echo user.get('nope', 'default')
echo user.contains('age')
Ada
default
true

With no fallback, get() on a missing key returns nil.

Writing

var user = { name: 'Ada' }

user.set('city', 'London')
user.age = 36
user['country'] = 'UK'

echo user
{name: Ada, city: London, age: 36, country: UK}

Assignment through a dot, a bracket or set() all do the same thing, and all three create the key if it does not exist. add() is a synonym for set().

To take a key back out:

echo user.remove('city')
echo user.contains('city')
London
false

remove() returns the value that was there.

clear() empties the dictionary.

Merging

var defaults = { host: 'localhost', port: 8080 }
var given = { port: 9000 }

defaults.extend(given)
echo defaults
{host: localhost, port: 9000}

extend() mutates the receiver and the right-hand side wins on conflicts. There is no + on dictionaries.

Transforming

var user = { name: 'Ada', age: 36, city: nil }

echo user.compact()
echo user.filter(@(value, key) => key == 'name')
echo user.some(@(value, key) => value == 36)
echo user.every(@(value, key) => value != nil)
echo user.reduce(@(acc, value, key) => acc + 1, 0)
{name: Ada, age: 36}
{name: Ada}
true
false
3

compact() drops entries whose value is nil.

Every callback here takes value, then key, the same order lists use.

Walking

Every form below visits the pairs in insertion order, which is the order they were first added, not the order they were last written to.

for With Two Variables, the Usual Form

var user = { name: 'Ada', age: 36 }

for key, value in user {
  echo '${key} = ${value}'
}
name = Ada
age = 36

Key first, value second. This is what you want almost every time.

for With One Variable Gives You the Values

var user = { name: 'Ada', age: 36 }

for value in user {
  echo value
}
Ada
36

This is worth stating plainly, because the equivalent loop in several other languages hands you the keys. In Zuri, one variable is always the value, whatever you are iterating.

iter Over the Keys, When You Need an Index

var user = { name: 'Ada', age: 36 }
var keys = user.keys()

iter var i = 0; i < keys.length(); i++ {
  echo '${i}. ${keys[i]} = ${user[keys[i]]}'
}
0. name = Ada
1. age = 36

A dictionary has no positional index of its own, so iter walks the key list instead. Reach for it when you need to number the output, or to look ahead to the next key.

each(), With a Function

var user = { name: 'Ada', age: 36 }

user.each(@(value, key) {
  echo '${key} -> ${value}'
})
name -> Ada
age -> 36

The callback receives the value first and the key second — the reverse of for, matching every other each() in the language.

Walking Just One Side

var user = { name: 'Ada', age: 36 }

for key in user.keys() {
  echo key
}

echo user.values()
name
age
[Ada, 36]

keys() and values() each return a plain list, so everything from Lists applies: sort them, filter them, reduce them. Sorting keys() is how you get output in a stable order regardless of insertion.

Converting

var user = { name: 'Ada', age: 36 }

echo user.to_list()
echo user.find_key('Ada')
echo user.clone()
[[name, age], [Ada, 36]]
name
{name: Ada, age: 36}

to_list() gives you two parallel lists, keys then values, not a list of pairs. find_key() searches by value and returns the first key holding it.

Dictionaries of Functions

A dictionary value can be a function, and that turns a dictionary into a dispatch table: a set of named behaviours you can choose between at runtime by looking a name up. It is the structure that replaces eval(), and it is worth knowing well.

def plus(a, b) {
  return a + b
}

var ops = {
  plus: plus,
  minus: @(a, b) => a - b,
  times: @(a, b) { return a * b },
}

All three forms are equivalent: a named function by name, an arrow function, and an anonymous function with a body.

Calling One

This is the part that surprises people, so it comes first. You cannot call one with a dot in a single step:

echo ops.plus(1, 2)
Unhandled TypeError: object of type dict does not define method 'plus'

A dot followed by a call looks for a method on the dictionary itself — length(), keys(), get() and the rest — not for a value stored under that key. Dictionaries have no method called plus, so the call fails.

There are two forms that do work. Read the function out first:

def plus(a, b) {
  return a + b
}

var ops = { plus: plus, minus: @(a, b) => a - b }

var chosen = ops.plus

echo chosen(1, 2)
3

Or index with brackets and call the result:

var ops = { plus: @(a, b) => a + b, minus: @(a, b) => a - b }

echo ops['minus'](5, 2)
3

The bracket form is the one to reach for when the key is computed, which is the whole point of a dispatch table:

var ops = {
  plus: @(a, b) => a + b,
  minus: @(a, b) => a - b,
  times: @(a, b) => a * b,
}

for name in ['plus', 'minus', 'times'] {
  echo '${name}: ${ops[name](6, 3)}'
}
plus: 9
minus: 3
times: 18

Beware of Names That Are Already Methods

A dictionary’s own methods take priority over its keys when you use a dot call. Storing a function under add, keys, get, length or any other method name produces a call that succeeds and does the wrong thing:

def plus(a, b) {
  return a + b
}

var d = { add: plus }

echo d.add(1, 2)
echo d.keys()
nil
[add, 1]

d.add(1, 2) called the dictionary’s own add(), which is a synonym for set() — so it stored the key 1 with the value 2 and returned nil. Nothing raised. The only sign is the extra key in the output.

Two habits avoid this entirely: use bracket calls for dispatch tables, and avoid method names as keys. Appendix E lists every name a dictionary already uses.

A Dispatch Table in Practice

The pattern is a table of handlers, a lookup with a fallback, and a call:

def _unknown(args) {
  return 'unknown command: ${args[0]}'
}

var commands = {
  greet: @(args) => 'hello, ${args[1]}',
  add: @(args) => args[1].to_number() + args[2].to_number(),
  version: @(args) => '1.0.0',
}

def run(line) {
  var args = line.split(' ')
  var handler = commands.get(args[0], _unknown)

  return handler(args)
}

echo run('greet ada')
echo run('add 2 40')
echo run('version')
echo run('explode')
hello, ada
42
1.0.0
unknown command: explode

commands.get(args[0], _unknown) is doing the work. A key that exists gives you its handler; one that does not gives you the fallback, so there is no separate contains() check and no branch for the error case.

The important property is that only the handlers in the table can ever run. A user typing anything at all selects one of four functions you wrote, or the fallback. That is the difference between a program that handles input and one that executes it, and it is why Zuri has no eval(). Chapter 21 covers the reasoning.

Storing Methods

zuri.reflect.bind_method() puts an instance’s method in a table with the instance still attached:

import zuri

class Counter {

  @new() {
    self.n = 0
  }

  increment() {
    self.n++
    return self.n
  }
}

var counter = Counter()

var actions = { up: zuri.reflect.bind_method(counter, 'increment') }

echo actions['up']()
echo actions['up']()
echo counter.n
1
2
2

The bound method keeps its receiver, so calling it through the table changes the counter it came from.

Comparing Dictionaries

== compares dictionaries by value, and order is not part of the comparison:

echo { a: 1 } == { a: 1 }
echo { a: 1 } == { a: 2 }
echo { a: 1, b: 2 } == { b: 2, a: 1 }
true
false
true

A dictionary remembers its insertion order for iteration, but two dictionaries with the same pairs in different orders are equal. That is the opposite of lists, where order is the whole point.

Dictionaries Are References

Like lists, assigning a dictionary shares it. clone() gives you a shallow copy.

Dictionaries and Objects

A dictionary is the right shape for data with a variable set of keys: configuration, a parsed JSON document, a set of HTTP headers. When the keys are fixed and there is behaviour attached to them, that is a class. Chapter 6 covers the difference.

A Worked Example

Dictionaries are the natural shape for counting, grouping and indexing. This example does all three over a list of log lines: it counts how often each level appears, groups the messages under their level, and builds an index from a request id back to the line that mentioned it.

var lines = [
  'INFO  req=a1 started',
  'WARN  req=a1 slow upstream',
  'INFO  req=b2 started',
  'ERROR req=b2 upstream refused',
  'INFO  req=a1 finished',
]

var counts = {}
var grouped = {}
var by_request = {}

for line in lines {
  var parts = line.split(' ').filter(@(p) => p != '')
  var level = parts[0]
  var request = parts[1].replace('req=', '', false)
  var message = ' '.join(parts[2, parts.length()])

  counts[level] = counts.get(level, 0) + 1

  if !grouped.contains(level) {
    grouped[level] = []
  }

  grouped[level].append(message)

  if !by_request.contains(request) {
    by_request[request] = []
  }

  by_request[request].append(level)
}

echo counts
echo grouped.ERROR
echo by_request.a1
echo by_request.get('zz', ['no such request'])
{INFO: 3, WARN: 1, ERROR: 1}
[upstream refused]
[INFO, WARN, INFO]
[no such request]

Four habits in there are worth taking away.

counts.get(level, 0) + 1 is the counting idiom. get() with a fallback means you never have to check whether the key exists before adding to it.

if !grouped.contains(level) { grouped[level] = [] } is the grouping idiom, and it has to be spelled out because get() would hand back a fresh default list each time rather than one you can keep appending to.

grouped.ERROR and by_request.a1 read keys with a dot, which works because both are valid identifiers. by_request['b2'] would be needed if the key were computed or awkward.

And the whole thing preserves order. counts came out INFO, WARN, ERROR — the order those levels were first seen, not alphabetical and not arbitrary.

Ranges and Iteration

Ranges

a..b builds a range value: the integers from a up to but not including b.

.. binds tighter than a method call, so a range takes a method directly with no parentheses around it:

echo 1..5.to_list()
echo 1..10.step(3).get_step()
[1, 2, 3, 4]
3

Parenthesise only when the bounds are themselves expressions, since .. takes primaries: (n * 2)..(n * 3).

A property access and a call are expressions too, so 0..self.size and 0..items.length() are (0..self).size and (0..items).length(). Write 0..(self.size) and 0..(items.length()) when the bound is the property rather than the range.

This is the price of the line above it. .. has to bind tighter than . for 1..10.step(3) to mean a range that steps, and once it does, there is no way for 0..self.step(5) to mean the instance’s own step instead. The parentheses say which one is meant, and the compiler cannot guess.

var r = 1..10

echo r
echo r.lower()
echo r.upper()
echo r.range()
1..10
1
10
9

range() is the distance between the bounds.

A range is a value, not loop syntax. Store it, pass it to a function, return it:

def page(number) {
  var size = 20
  return (number * size)..((number + 1) * size)
}

echo page(2)
40..60

Counting Down

Write the larger bound first and the range runs backwards:

for i in 10..1 {
  echo i
}
10
9
8
7
6
5
4
3
2

The rule is the same in both directions: the first bound is included, the second is not.

Stepping

for i in 1..10.step(3) {
  echo i
}
1
4
7

step() gives back a new range carrying that stride. get_step() reads it; the default is 1.

to_list() materialises a range into a list at stride one, ignoring any step you set:

echo 1..5.to_list()
[1, 2, 3, 4]

Membership

within() tests against the bounds inclusively on both sides, and it normalises direction, so 10..1.within(10) and 1..10.within(10) both answer the same:

var r = 1..10

echo r.within(5)
echo r.within(10)
true
true

That differs from iteration, which excludes the upper bound. When you want “would the loop visit this number”, compare against lower() and upper() yourself.

Walking a Range

for with one variable gives you the values, and with two it gives you the position first and the value second:

for value in 3..6 {
  echo value
}

for index, value in 3..6 {
  echo '${index}: ${value}'
}
3
4
5
0: 3
1: 4
2: 5

The index counts from zero regardless of where the range starts, which is what makes it useful: 3..6 yields values 3, 4, 5 at positions 0, 1, 2.

loop() is the callback form. It honours the step and the direction:

25..18.loop(@(i) { print('${i} ') })
print('\n')
25 24 23 22 21 20 19 

An iter loop needs no range at all — its three clauses already say everything a range says, and more, since the step can be any expression:

iter var i = 3; i < 6; i++ {
  echo i
}
3
4
5

Use a range with for when the bounds are the interesting part, and iter when the stepping is.

What Is Iterable

for ... in works on:

  • lists, giving index and value
  • dictionaries, giving key and value, in insertion order
  • strings, giving character position and character
  • bytes, giving index and the numeric byte
  • ranges, giving position and value
  • any class that defines @key() and @value()

is_iterable() answers the question for any value:

echo is_iterable([1, 2])
echo is_iterable('text')
echo is_iterable(42)
true
true
false

The Iterator Protocol

for is not magic. It desugars into a loop over two method calls.

Given for value in thing, Zuri evaluates thing once, then repeats:

  1. key = thing.@key(key), starting from nil
  2. stop if key is nil
  3. value = thing.@value(key)
  4. run the body

So @key(previous) answers “what comes after this one?”, returning nil when there is nothing left, and @value(key) answers “what is stored here?”.

Everything built in implements this. So can your own classes, which is what Chapter 6 shows.

Because the iterable is evaluated exactly once, this is safe:

for line in read_the_whole_file() {
  echo line
}

The function runs once, not once per line.

Functions

A function is a piece of behaviour with a name, a list of parameters, and a result. You have been calling them since Chapter 1; this chapter is about writing them.

Functions in Zuri are values. You can put one in a list, hand it to another function, return one from a function, store one in a dictionary, and call whatever comes back:

def double(n) {
  return n * 2
}

var operations = { twice: double, thrice: @(n) => n * 3 }

var chosen = operations.twice

echo chosen(5)
echo operations['thrice'](5)
echo [1, 2, 3].map(double)
10
15
[2, 4, 6]

That one property is what makes map, filter, reduce, every callback in the standard library, and every route handler in Chapter 15 possible.

Note the shape of those two calls. operations.twice reads the function out of the dictionary, and then you call what you read. Writing operations.twice(5) in one step does not work: a dot followed by a call looks for a method on the dictionary, and a dictionary has no method called twice. Either read it into a variable first, as above, or index with brackets and call the result: operations['twice'](5).

This chapter covers declaring a function, the spellings an anonymous one can take, how a closure captures the variables around it, and the optional type annotations that make the runtime check arguments for you.

Defining Functions

Declaration

def greet(name) {
  return 'Hello, ' + name
}

echo greet('Zuri')
Hello, Zuri

def, a name, a parameter list in parentheses, a block. There is no return type to declare, no forward declaration, and no separate signature.

A function with no return yields nil:

def silent() {
  echo 'working'
}

echo silent()
working
nil

There are no implicit returns. The last expression in a body is not its value; if you want something back, say return.

Arity Is Not Enforced

Call a function with fewer arguments than it declares and the missing ones arrive as nil. Call it with more and the extras are dropped:

def greet(name) {
  return 'Hello, ' + name
}

echo greet()
echo greet('Zuri', 'ignored')
Hello, nil
Hello, Zuri

This is not an oversight. It is the mechanism optional parameters are built on, since there is no default-value syntax in a parameter list. You write the default in the body instead:

def greet(name, greeting) {
  greeting = greeting or 'Hello'

  return '${greeting}, ${name}!'
}

echo greet('Ada')
echo greet('Ada', 'Hi')
Hello, Ada!
Hi, Ada!

The or Trap in Defaults

Because or does the work, a legitimately falsy argument is replaced by the default. In Zuri that set includes 0, NaN, '' and false:

def indent(text, width) {
  width = width or 2

  return ' ' * width + text
}

echo '[' + indent('a') + ']'
echo '[' + indent('a', 0) + ']'
[  a]
[  a]

The caller asked for zero indentation and got two. When a falsy value is a legitimate input, test for nil explicitly:

def indent(text, width) {
  if width == nil {
    width = 2
  }

  return ' ' * width + text
}

echo '[' + indent('a') + ']'
echo '[' + indent('a', 0) + ']'
[  a]
[a]

Use or for a default when every falsy value is genuinely absent — a name, a path, a list. Use == nil for numbers and booleans.

If you want the runtime to reject a missing argument rather than default it, annotate the parameter. Type Annotations covers that.

Variadic Functions

A final parameter prefixed with ... collects every remaining argument into a list:

def count(...items) {
  return items
}

echo count()
echo count(1)
echo count(1, 2, 3)
[]
[1]
[1, 2, 3]

Three facts about what it captures are worth being precise about.

It is always a list, including when nothing was passed. count() gives [], not nil, so items.length() is safe without a check.

It captures arguments, not contents. Passing a list passes one argument, which arrives as a list nested inside the variadic list:

def count(...items) {
  return items
}

echo count([1, 2])
echo count([1, 2], 3)
[[1, 2]]
[[1, 2], 3]

... marks a parameter, not an argument: count(...list) does not parse. To call a function with a list of arguments, use apply():

def count(...args) {
  return args.length()
}

echo count.apply([1, 2, 3])
3

When a function should accept a collection, though, take the list as an ordinary parameter and skip both.

Named parameters bind first. A variadic can follow named ones, and it takes whatever is left over:

def named(a, ...rest) {
  return [a, rest]
}

echo named()
echo named(1)
echo named(1, 2, 3)
[nil, []]
[1, []]
[1, [2, 3]]

The variadic must be last. def mid(a, ...rest, b) is a syntax error, because there is no rule that could decide how many arguments rest should keep.

A practical use is a function that formats an arbitrary number of pieces:

def log_line(level, ...parts) {
  return '[${level}] ' + ' '.join(parts)
}

echo log_line('warn', 'disk', 'almost', 'full')
echo log_line('info')
[warn] disk almost full
[info] 

Where a def Binds

A def scopes exactly like a var. Written at the top level of a file it binds a module-level name. Written anywhere else it binds a local of the block it appears in, and it goes away with that block:

def outer() {
  def helper() {
    return 'from helper'
  }

  return helper()
}

echo outer()

catch {
  helper()
} as e {
  echo e.message
}
from helper
undefined global 'helper'

A helper defined inside a function belongs to that function. Nothing outside can reach it, and nothing outside can be broken by it.

That extends to any block, not just a function body:

def choose(verbose) {
  if verbose {
    def describe(n) {
      return 'the number ${n}'
    }

    return describe(7)
  }

  return '7'
}

echo choose(true)
echo choose(false)
the number 7
7

Calling One Helper From Another

Helpers declared next to each other can call each other, in either direction:

def parity(n) {
  def even(k) {
    if k == 0 {
      return true
    }

    return odd(k - 1)
  }

  def odd(k) {
    if k == 0 {
      return false
    }

    return even(k - 1)
  }

  return even(n)
}

echo parity(10)
echo parity(7)
true
false

even mentions odd before odd is written, and that works because the compiler reserves a slot for every def in a block before it compiles any of them. What it does not do is move the declaration: until the def line itself runs, the slot holds nil, so calling a helper from above its declaration raises.

A helper can also call itself:

def factorial_of(n) {
  def factorial(k) {
    if k <= 1 {
      return 1
    }

    return k * factorial(k - 1)
  }

  return factorial(n)
}

echo factorial_of(6)
720

A Helper Captures What Surrounds It

A nested def is a closure over the locals around it, and it outlives the call that created it:

def make_counter(start) {
  var count = start

  def bump() {
    count++
    return count
  }

  return bump
}

var bump = make_counter(10)

echo bump()
echo bump()
11
12

An anonymous function bound to a var does the same job, and is the better choice when the helper is a one-liner or is being passed straight into something else:

def outer() {
  var helper = @() => 'private'

  return helper()
}

echo outer()
private

Pick whichever reads better. A def gives the function a real name in stack traces and can recurse without the extra var; a lambda is shorter.

One Declaration Per Name

Declaring the same function name twice in one scope is a compile error:

def pick() {
  return 'first'
}

def pick() {
  return 'second'
}
SyntaxError: multiple declaration for function 'pick' found
  --> /path/to/main.zu:5:5
  |
5 | def pick() {
  |     ^

A different parameter list does not make it a different function. Zuri has no overloading: one name, one function.

def render(value) {}
def render(value, width) {}
SyntaxError: multiple declaration for function 'render' found

When you want one name to handle several shapes of input, take the extra arguments as optional and branch in the body — which is what the greeting parameter above is doing.

The check is per scope, exactly like var’s. A function declared inside another is a separate declaration in a separate scope, so the compiler allows it, and it shadows the outer one for as long as its own scope lasts:

def render() {
  return 'top level'
}

def wrapper() {
  def render() {
    return 'inner'
  }

  return render()
}

echo wrapper()
echo render()
inner
top level

wrapper’s own render is a local of wrapper. It answers every call made inside that function and disappears when the call ends, leaving the module-level render untouched.

Classes follow the same rule for their methods:

class Duplicate {
  m() {}
  m() {}
}
SyntaxError: multiple declaration for method 'm' found in class 'Duplicate'

The REPL is the one exception. Retyping a def there replaces the previous one, because correcting something you just typed is what the prompt is for.

Declaration Order

A function must be declared before the top-level line that calls it:

echo later()

def later() {
  return 'nope'
}
Unhandled UndefinedError: undefined global 'later'

There is no hoisting. Declarations execute in order, like everything else.

Inside a function body the rule does not apply, because a name is resolved when the call runs rather than when it is compiled. That is what makes mutual recursion work with no forward declaration:

def is_even(n) {
  if n == 0 {
    return true
  }

  return is_odd(n - 1)
}

def is_odd(n) {
  if n == 0 {
    return false
  }

  return is_even(n - 1)
}

echo is_even(10)
echo is_odd(7)
true
true

is_even refers to is_odd before it exists. That is fine, because the reference is resolved when is_even(10) runs, by which time both declarations have executed.

Helpers nested inside a function get the same freedom by a different route: their slots are reserved together, before any of their bodies compile. Either way, what matters is that every declaration has run by the time the first call is made.

Functions as Values

A function is a value, so it can be passed, stored and returned:

def apply_twice(f, x) {
  return f(f(x))
}

def increment(n) {
  return n + 1
}

echo apply_twice(increment, 5)
echo apply_twice(@(s) => s + '!', 'hi')
7
hi!!

What a Function Knows About Itself

Every callable carries four methods:

def named(a, b, ...c) {
  return [a, b, c]
}

echo named.name()
echo named.arity()
echo named.is_variadic()
echo named.to_string()
named
3
true
<function named(3)>

arity() counts declared parameters, the variadic one included, so a variadic function’s arity is a minimum rather than a requirement.

An anonymous function is named after the order the compiler met it:

var f = @(x) => x

echo f.name()
@anon0

A method read off an instance keeps its receiver, and counts it:

class Greeter {

  hello() {
    return 'hi'
  }
}

var bound = Greeter().hello

echo bound()
echo bound.name()
echo bound.arity()
hi
hello
1

hello() declares no parameters, and arity() reports 1, because the instance is the first one.

call()

call() invokes a function with the arguments you give it:

def add(a, b) {
  return a + b
}

echo add.call(2, 3)
5

call() takes the arguments individually, exactly as a normal call does. It is not a spread: add.call([2, 3]) passes one argument, a list. Its use is calling something whose identity you only have as a value — a handler out of a dictionary, a method from zuri.reflect — in a place where the ordinary call syntax reads badly.

Testing Callability

def f() {}

echo is_callable(f)
echo is_callable(print)
echo is_callable(Error)
echo is_callable(42)

echo is_function(f)
echo is_function(Error)
true
true
true
false
true
false

is_callable() is true for anything you can put parentheses after, classes included — calling a class constructs an instance. is_function() is narrower: true for a def, an anonymous function and a built-in, false for a class.

A Worked Example

A small pipeline builder, using most of this section: functions as values, a variadic parameter, a default, and a returned closure.

def pipeline(...stages) {
  return @(input) {
    var value = input

    for stage in stages {
      value = stage(value)
    }

    return value
  }
}

def strip(text) {
  return text.trim()
}

def collapse_spaces(text) {
  return text.replace('/\s+/', ' ')
}

def truncate(text, limit) {
  limit = limit or 20

  if text.length() <= limit {
    return text
  }

  return text[0, limit - 1] + '…'
}

var tidy = pipeline(strip, collapse_spaces, @(t) => truncate(t, 12))

echo '[' + tidy('   too    many   spaces here   ') + ']'
echo '[' + tidy('  short  ') + ']'
echo '[' + pipeline()('untouched') + ']'
[too many sp…]
[short]
[untouched]

Four things to take from it.

pipeline() takes its stages variadically and returns a closure over them, so the returned function is a new function specialised to those stages.

pipeline() with no arguments still works, because a variadic parameter is an empty list rather than nil, and a for over an empty list runs zero times.

The third stage is wrapped in @(t) => truncate(t, 12) because truncate takes two parameters and a stage takes one. That wrapper is the thing a spread operator would otherwise be for.

And limit = limit or 20 is safe here precisely because 0 is not a sensible limit. Had it been, this would need the == nil form.

Anonymous Functions and Closures

The Spellings

An anonymous function has two halves you can vary independently: how it opens, and how its body is written. That gives a small grid rather than a list to memorise.

Opening. def is the keyword; @ is its shorthand. They are the same thing.

Body. A block in braces returns with return. An arrow returns the one expression after it.

var block_long = def(x) { return x * x }
var block_short = @(x) { return x * x }

var arrow_long = def(x) => x * x
var arrow_short = @(x) => x * x

echo [block_long(3), block_short(3), arrow_long(3), arrow_short(3)]
[9, 9, 9, 9]

When there are no parameters, the empty parentheses may be dropped as well:

var a = @() { return 1 }
var b = @() => 2
var c = def() { return 3 }
var d = def() => 4
var e = @ => 5
var f = @{ return 6 }
var g = def { return 7 }
var h = def => 8

echo [a(), b(), c(), d(), e(), f(), g(), h()]
[1, 2, 3, 4, 5, 6, 7, 8]

All eight are the same construct. Which to use:

  • @(x) => expr for a one-expression callback. This is what most Zuri code uses, and what the standard library is written in.
  • @(x) { ... } when the body needs more than one statement.
  • def(x) { ... } when the function is long enough that you want it to look like the declarations around it.
  • @ => expr for a thunk — a value computed on demand, with nothing passed in.

Compare the two common ones where they actually turn up:

echo ['  a ', ' b'].map(@(t) => t.trim())

echo [1, 2, 3].filter(@(n) {
  if n == 2 {
    return false
  }

  return n > 0
})
[a, b]
[1, 3]

The arrow form has no return because the expression is the return value. Adding one is a mistake the compiler will not catch: @(x) => return x does not parse, but @(x) { x } parses and returns nil.

Anonymous functions take variadic parameters and type annotations exactly like named ones:

var sum = @(...numbers) => numbers.reduce(@(a, b) => a + b, 0)
var doubled = @(n: number) => n * 2

echo sum(1, 2, 3)
echo doubled(4)

catch {
  doubled('four')
} as e {
  echo e.message.replace('/@anon\d+/', '@anonN')
}
6
8
@anonN() expects parameter 'n' (argument 1) to be a number, got string

The real message carries a number rather than the N shown here. An anonymous function is named @anonN in the order the compiler met it in the file, which is worth knowing when one turns up in a stack trace — and worth not depending on, since inserting another anonymous function above it renumbers everything below.

Closures

A function carries the variables it referenced from its surroundings, and keeps them alive after the enclosing function has returned:

def make_counter() {
  var n = 0

  return @() {
    n++
    return n
  }
}

var a = make_counter()
var b = make_counter()

echo a()
echo a()
echo b()
1
2
1

a and b each closed over their own n. Each call to make_counter() created a fresh one.

Capture Is by Reference

A closure captures the variable, not a snapshot of its value:

def demo() {
  var total = 0
  var add = @(n) { total += n }

  add(5)
  add(10)

  return total
}

echo demo()
15

This is what makes accumulator patterns work, and it is what makes loop variables surprising. If you want a per-iteration capture, declare a fresh variable inside the loop body:

var fns = []

iter var i = 0; i < 3; i++ {
  var captured = i
  fns.append(@() => captured)
}

echo fns.map(@(f) => f())
[0, 1, 2]

captured is a new variable on each pass, so each closure gets its own.

Closures over Parameters

Parameters are captured the same way, which is the basis of partial application:

def adder(amount) {
  return @(n) => n + amount
}

var add_five = adder(5)
var add_ten = adder(10)

echo add_five(1)
echo add_ten(1)
6
11

Recursion in an Anonymous Function

An anonymous function has no name to call itself by. Bind it first:

var factorial

factorial = @(n) {
  if n <= 1 {
    return 1
  }
  return n * factorial(n - 1)
}

echo factorial(5)
120

The variable is captured by reference, so by the time the body runs, factorial is bound.

Functions Compare by Identity

def f() {}

var g = f
echo f == g
echo f == @() {}
true
false

Two separately written functions are never equal, even with identical bodies.

A Worked Example

Closures are at their most useful when a function needs to remember something between calls without that something becoming a global. Here is a rate limiter: it hands back a function that answers “may I do this now?”, and keeps its own tally where nothing else can reach it.

def make_limiter(max_per_window: number) {
  var used = 0

  return {
    allow: @() {
      if used >= max_per_window {
        return false
      }

      used++
      return true
    },

    remaining: @() => max_per_window - used,

    reset: @() {
      used = 0
    },
  }
}

var limiter = make_limiter(2)
var allow = limiter.allow
var remaining = limiter.remaining
var reset = limiter.reset

echo allow()
echo allow()
echo allow()
echo remaining()

reset()
echo allow()
true
true
false
0
true

Three separate functions share one used, because all three closed over the same variable in the same call to make_limiter. A second call to make_limiter would produce a second, independent trio.

This is the closest Zuri gets to a private field without a class, and it is worth knowing for exactly that reason: there is no way for a caller to read or write used except through the three functions you gave them.

Type Annotations

Zuri is dynamically typed, and parameters may be annotated. An annotated parameter is checked at every call, and a mismatch raises a TypeError naming the parameter, its position and what actually arrived.

def repeat(text: string, times: number) {
  return text * times
}

echo repeat('ab', 3)

catch {
  repeat(42, 3)
} as e {
  echo e.message
}
ababab
repeat() expects parameter 'text' (argument 1) to be a string, got number

The Type Names

NameAccepts
anyanything, including nil
booltrue or false
numberany number
inta number with no fractional part
bigintan arbitrary-precision integer
stringa string
bytesa byte buffer
lista list
dicta dictionary
rangea range
filea file handle
functiona function or an anonymous function
callableanything callable, classes included
iterableanything for can walk
typea class
ClassNamean instance of that class or a subclass

Anything that is not one of those names is read as a class name.

number and bigint are separate, in annotations exactly as they are everywhere else. A parameter typed bigint rejects 10, and one typed number rejects 10n:

def factorial(n: bigint) {
  if n <= 1n {
    return 1n
  }

  return n * factorial(n - 1n)
}

echo factorial(25n)

catch {
  factorial(25)
} as e {
  echo e.message
}
15511210043330985984000000n
factorial() expects parameter 'n' (argument 1) to be a bigint, got number

Write bigint|number when a function genuinely takes either, and convert inside it with to_bigint().

Required by Default

A plain annotation makes the argument required. Omitting it is a TypeError, because the parameter arrives as nil:

def needs_number(x: number) {}

catch {
  needs_number()
} as e {
  echo e.message
}
needs_number() expects parameter 'x' (argument 1) to be a number, got nil

That is how you get argument checking out of a language that otherwise lets you call anything with anything.

Optional Parameters

Prefix the type with ? to allow nil:

def connect(host: string, port: ?number) {
  port = port or 8080
  echo '${host}:${port}'
}

connect('localhost')
connect('localhost', 9000)
localhost:8080
localhost:9000

?any is the same as no annotation at all, which is what an unannotated parameter gets.

Union Types

Separate alternatives with |:

def render(value: string|number|list) {
  echo value
}

render('text')
render(42)
render([1])

catch {
  render({ a: 1 })
} as e {
  echo e.message
}
text
42
[1]
render() expects parameter 'value' (argument 1) to be a string, a number, or a list, got dict

The ? goes at the front and applies to the whole union: ?string|number.

Classes and Subclasses

A class name accepts instances of that class and of anything that inherits from it:

class Point {
  @new(x, y) {
    self.x = x
    self.y = y
  }
}

def distance_from_origin(p: Point) {
  return (p.x * p.x + p.y * p.y).sqrt()
}

echo distance_from_origin(Point(3, 4))
5

This is what makes error handling readable, since every built-in error subclasses Error:

def report(e: Error) {
  echo '${e.type}: ${e.message}'
}

catch {
  raise ValueError('bad value')
} as e {
  report(e)
}
ValueError: bad value

Where Annotations Work

Annotations go on parameters: in a def, in a method, and in an anonymous function.

class Report {
  render(rows: list, title: ?string) {
    # ...
  }
}

var f = @(n: int) => n * 2

The same syntax goes on var declarations and on class fields:

var count: number = 0
var name: ?string

class Report {
  var rows: list = []
  var title: ?string
}

A variable annotation is a statement of intent. It is not checked at runtime, so the runtime will not stop you from putting a string in count. What it does is make the declaration self-documenting, and give static analysis tools something to work with: an annotated declaration tells a linter, an editor’s completion engine or a type checker exactly what belongs there, and lets them flag the assignment your eyes would otherwise have to catch.

Annotate declarations for the reader and the tooling. Annotate parameters for the runtime.

When to Annotate

Annotations are optional, and a program with none of them is perfectly ordinary Zuri. The question is where they earn their place.

Annotate a boundary. A function that receives data from outside your program — a request handler, a file parser, a public function in a module other people import — is where a wrong type first arrives. An annotation there turns a confusing failure deep in the call stack into a clear one at the door:

def parse_port(raw: string) {
  var port = raw.to_number()

  if port < 1 or port > 65535 {
    raise ValueError('port out of range: ${raw}')
  }

  return port
}

catch {
  parse_port(8080)
} as e {
  echo e.message
}

echo parse_port('8080')
parse_port() expects parameter 'raw' (argument 1) to be a string, got number
8080

Note the division of labour there. The annotation handles wrong type, and the explicit check handles wrong value. An annotation can never do the second job, because 70000 is a perfectly good number.

Annotate to replace a manual check. Any function that opens with if !is_string(x) { raise TypeError(...) } is spelling out by hand what an annotation says in one word, and the annotation produces a better message:

def shout(text: string) {
  return text.upper() + '!'
}

echo shout('hello')
HELLO!

Leave internal helpers alone if you prefer. A private function called from three places in the same file, all of which you can see, gains less. Annotate it if it documents something non-obvious; skip it if it does not.

One thing an annotation is not: a substitute for validation. text: string guarantees you have a string, not that the string is a valid email address, a well-formed date or a non-empty name. The validate module covers that job, and Chapter 13 introduces it.

Classes and Objects

A dictionary holds data. A class holds data and the behaviour that goes with it, under a name you can check for and inherit from.

class Temperature {

  @new(celsius) {
    self.celsius = celsius
  }

  fahrenheit() {
    return self.celsius * 9 / 5 + 32
  }

  to_string() {
    return '${self.celsius}C'
  }
}

var t = Temperature(100)

echo t.fahrenheit()
echo t.to_string()
212
100C

Three facts about Zuri classes shape everything in this chapter, and it is worth having them up front.

The constructor is @new. Methods whose names start with @ are decorated methods, and each one connects the class to a piece of the language’s own syntax: construction, arithmetic, comparison, iteration, JSON encoding. There is a fixed set of them, and Decorated Methods covers all of it.

Fields are reached through self. Inside a method, self is the instance. There is no bare balance that means self.balance, and no self parameter in the declaration either.

A class is sealed. The set of fields and methods is fixed when the class is declared. Nothing at runtime can add a new field to an instance or a new method to a class, and an attempt raises an error rather than quietly creating one. That rules out a family of bugs — a typo in a field name is an error, not a new field — and it rules out monkey-patching, which is a technique some languages rely on and Zuri does not offer.

Inheritance is single: a class has at most one parent, written with <. There are no interfaces, no mixins and no abstract keyword. The sections that follow show what you write instead.

Sealed does not mean unchangeable forever, though. A separate declaration written with > instead of < — a class extension — can add methods to a class after the fact, including one you did not write. Class Extensions covers it.

Defining a Class

class Account {
  var balance = 0
  var owner

  @new(owner, balance) {
    self.owner = owner
    self.balance = balance or 0
  }

  deposit(amount) {
    self.balance += amount
    return self
  }
}
  • class Name { ... } declares it. PascalCase is the convention.
  • var inside the body declares a field, with an optional default. A field with no default starts as nil.
  • A bare name(params) { ... } is a method. There is no def keyword on methods.
  • @new is the constructor.
  • self is the current instance.

Creating an Instance

Call the class. There is no new keyword:

var a = Account('Ada', 50)

echo a.owner
echo a.balance
Ada
50

If the class has no @new, calling it with no arguments gives you an instance with every field at its default.

The Constructor

@new is the constructor. It runs once, when the class is called, and its job is to put the instance into a usable state — which is what Account’s did above.

It Is Optional

A class with no @new is constructed with no arguments, and every field takes its declared default:

class Settings {
  var theme = 'dark'
  var retries
}

var s = Settings()

echo s.theme
echo s.retries
dark
nil

Extra arguments to a class with no @new are ignored rather than rejected, exactly as they are for a function.

It Declares Fields

self.x = value inside @new declares a field. This is the one place in the language where assignment creates a member, and it exists so a constructor does not have to repeat every field as a var line above it:

class Point {

  @new(x, y) {
    self.x = x
    self.y = y
    self.distance = (x * x + y * y).sqrt()
  }
}

var p = Point(3, 4)

echo p.distance
5

distance was never declared with var, and it is a real field.

That power belongs to @new’s own body and nowhere else. A constructor that delegates its setup to a helper must declare those fields:

class Delayed {
  var ready          # required: setup() cannot declare it

  @new() {
    self.setup()
  }

  setup() {
    self.ready = true
  }
}

echo Delayed().ready
true

Remove the var ready line and setup() raises undefined field 'ready'.

Arguments Are Not Checked Unless You Ask

@new is a function, so the usual rules apply: missing arguments arrive as nil, extra ones are dropped. A constructor that assumes it got something fails later, in a confusing place:

class Ctor {

  @new(x) {
    self.derived = x * 2
  }
}

Ctor()
Unhandled TypeError: operator '*' not defined for call signature (nil, float)

Annotate the parameter and the failure moves to the call, where it names the problem:

class Ctor {

  @new(x: number) {
    self.derived = x * 2
  }
}

echo Ctor(5).derived

catch {
  Ctor('five')
} as e {
  echo e.message
}
10
@new() expects parameter 'x' (argument 1) to be a number, got string

Constructors are the highest-value place in a program to annotate, because a badly built object goes wrong somewhere else entirely.

Validating in the Constructor

A constructor may raise. Nothing is returned to the caller, so an object that cannot be valid never exists:

class Port {

  @new(number: number) {
    if number < 1 or number > 65535 {
      raise ValueError('port out of range: ${number}')
    }

    self.number = number
  }
}

echo Port(8080).number

catch {
  Port(99999)
} as e {
  echo '${e.type}: ${e.message}'
}
8080
ValueError: port out of range: 99999

This is worth doing whenever a class has an invariant. Every other method can then assume it holds, instead of re-checking.

@new Cannot Return a Value

Calling a class always produces an instance of that class. A return in @new ends the constructor early; it does not change what the caller gets:

class Returns {

  @new() {
    self.v = 1

    return 'this is discarded'
  }
}

echo typeof(Returns())
Returns

If you want a call that may hand back something else — a cached instance, a subclass, nil on bad input — write a static factory method and call that instead.

Methods

A method is a name(params) { ... } declaration inside the class body. There is no def keyword on it:

class Account {

  @new(owner, balance) {
    self.owner = owner
    self.balance = balance
  }

  deposit(amount) {
    self.balance += amount
    return self
  }
}

var a = Account('Ada', 50)

a.deposit(25).deposit(25)

echo a.balance
100

deposit returns self, which is what makes the chain work. Returning self from a mutator is a common Zuri idiom.

Inside a method, self is required to reach a field or another method. There is no implicit receiver:

class Greeter {
  var name = 'world'

  greet() {
    return 'hello ' + self.name
  }
}

Writing name there would look for a local or a global, not a field.

Fields Are Declared, Not Discovered

The set of fields is fixed when the class is declared. Two things follow from that.

First, assigning to a field that was never declared is an error:

class Account {
  var balance = 0
}

var a = Account()

catch {
  a.nickname = 'rainy day'
} as e {
  echo e.message
}
undefined field 'nickname' on instance of 'Account'

Second, self.x = value inside @new does declare a field, as a convenience so that constructors do not have to repeat themselves:

class Point {
  @new(x, y) {
    self.x = x
    self.y = y
  }
}

echo Point(3, 4).x
3

That convenience applies to @new’s own body and nowhere else. A constructor that calls a helper to do its initialisation must declare those fields with var:

class A {
  var from_helper          # required

  @new() {
    self.setup()
  }

  setup() {
    self.from_helper = 2
  }
}

Without the var line, the assignment inside setup() raises undefined field 'from_helper'.

Static Members

static puts a field or method on the class rather than on each instance:

class Account {
  static var count = 0

  @new(owner) {
    self.owner = owner
    Account.count++
  }

  static open(owner) {
    return Account(owner)
  }
}

Account('Ada')
Account('Bob')

echo Account.count
echo Account.open('Carol').owner
2
Carol

A static method has no self:

SyntaxError: 'self' used outside of a method

Reach the class by name instead.

Static fields are the one mutable part of a class. The set of members is sealed; the values of static fields are not.

Duplicate Members Are Errors

Declaring the same method twice in one class does not silently keep the last one:

class B {
  m() {}
  m() {}
}
SyntaxError: multiple declaration for method 'm' found in class 'B'

The same applies to declaring the same class name twice in one module.

Instances and Dictionaries

dictclass
keysany, added at any timefixed at declaration
accessd.key or d['key']obj.field only
missing memberget() returns a fallbackerror
behaviournonemethods
cost of a readhash lookuparray index

Use a dictionary for data whose shape you learn at runtime: parsed JSON, HTTP headers, a config file. Use a class when the shape is known and there is behaviour to attach.

Inheritance

A class extends another with <:

class Shape {
  @new(name) {
    self.name = name
  }

  area() {
    raise NotImplementedError('${self.name} must define area()')
  }

  describe() {
    return '${self.name} with area ${self.area()}'
  }
}

class Circle < Shape {
  @new(radius) {
    parent('circle')
    self.radius = radius
  }

  area() {
    return 3.141592653589793 * self.radius ** 2
  }
}

echo Circle(1).describe()
circle with area 3.141592653589793

A subclass inherits every field and method. Single inheritance only; a class has at most one parent.

Note the direction of the arrow. class Circle < Shape creates a new class that borrows from Shape. Turning it around — class Anything > Shape — is a different declaration entirely: it adds methods to Shape itself and creates no class at all. See Class Extensions.

Naming the Parent

A class is a value like any other, so the parent does not have to be a bare name in the current file. Anything an expression starting with a name can reach works, which is what lets you build on a class another module owns:

import log

class MemoryTransport < log.Transport {
  @new() {
    self.lines = []
  }

  handle(record) {
    self.lines.append(record)
  }
}

echo instance_of(MemoryTransport(), log.Transport)
true

Property access, indexing and calls all work the same way, so a parent picked out of a registry (class Store < backends['redis']) or handed back by a function (class Store < chosen_backend()) is as valid as a name. The expression is evaluated once, where the class is declared.

The one rule is that it has to begin with a name. That is what keeps the { opening the body from ever being read as the start of a dictionary.

parent

parent means two related things.

parent(args) calls the parent’s constructor. That is the parent('circle') line in Circle above, and it comes first in @new, before the subclass sets up anything of its own. Calling it late means the parent’s constructor overwrites what you just assigned; not calling it at all means the parent’s fields are never initialised.

parent.method(args) calls the parent’s version of a method that this class has overridden:

class Square < Shape {
  @new(side) {
    parent('square')
    self.side = side
  }

  area() {
    return self.side ** 2
  }

  describe() {
    return 'a ' + parent.describe()
  }
}

echo Square(2).describe()
a square with area 4

Note what happened there. parent.describe() ran Shape’s describe, which calls self.area(), which dispatched back to Square’s area. A method always dispatches on the actual object, not on the class the code was written in.

Overriding

Redeclaring a method in a subclass replaces it. There is no keyword for it and no way to forbid it.

describe() above is the useful shape of this: a parent method written in terms of a method the child supplies. Shape.area() raises NotImplementedError, so a subclass that forgets to override it says so clearly:

catch {
  Shape('blob').area()
} as e {
  echo e.message
}
blob must define area()

That is Zuri’s abstract method. There is no abstract keyword because there does not need to be one.

A Subclass Without a Constructor

If a subclass declares no @new, it uses its parent’s:

class NoCtor < Shape {}

echo NoCtor('plain').name
plain

Testing Ancestry

instance_of() walks the whole chain:

echo instance_of(Circle(1), Shape)
echo instance_of(Circle(1), Square)
true
false

typeof() gives the most specific class name, as a string:

echo typeof(Circle(1))
Circle

The Depth Is Free

Inheritance chains cost nothing to walk at runtime. A method lookup on a class four levels deep is the same operation as one on a class with no parent, because every class’s method table is complete at declaration time. Write the hierarchy the design wants.

What You Write Instead of an Interface

Zuri has single inheritance and no interfaces, so the two patterns below do the jobs an interface would do elsewhere.

A base class that raises. When a base class needs every subclass to supply a method, declare it and raise:

class Shape {

  @new(name: string) {
    self.name = name
  }

  area() {
    raise NotImplementedError('${self.name} must define area()')
  }

  describe() {
    return '${self.name} has area ${self.area()}'
  }
}

class Circle < Shape {

  @new(radius: number) {
    parent('circle')
    self.radius = radius
  }

  area() {
    return 3.141592653589793 * self.radius ** 2
  }
}

class Blob < Shape {

  @new() {
    parent('blob')
  }
}

echo Circle(2).describe()

catch {
  echo Blob().describe()
} as e {
  echo '${e.type}: ${e.message}'
}
circle has area 12.566370614359172
NotImplementedError: blob must define area()

describe() calls self.area(), and self is the actual instance, so the subclass’s version runs. That is the whole of dynamic dispatch in Zuri: there is nothing to declare and nothing to mark virtual.

A parameter typed by the base class. An annotation naming a class accepts any subclass of it, which is how you say “anything that is a Shape”:

def total_area(shapes: list) {
  return shapes.reduce(@(sum, shape: Shape) => sum + shape.area(), 0)
}

Checking What Something Is

Two built-ins answer questions about an instance’s type, and they answer different questions.

instance_of(value, Class)

instance_of() asks “does this behave like a Class?”, and it walks the whole inheritance chain to answer:

class Shape {}
class Circle < Shape {}
class Square < Shape {}
class Ellipse < Circle {}

var c = Circle()

echo instance_of(c, Circle)
echo instance_of(c, Shape)
echo instance_of(c, Square)
echo instance_of(Ellipse(), Shape)
true
true
false
true

Ellipse is two levels below Shape and still answers true. There is no depth limit: the check follows parents until it finds a match or runs out. A sibling — Circle against Square — is false, because siblings share an ancestor rather than a lineage.

It never raises for the first argument. Anything that is not an instance of the class simply answers false, including values that are not instances at all:

class Shape {}

echo instance_of(42, Shape)
echo instance_of('circle', Shape)
echo instance_of(nil, Shape)
echo instance_of([1, 2], Shape)
false
false
false
false

That makes it safe to use as a guard on a value you know nothing about; no is_instance() check has to come first.

A class is not an instance of itself. Passing the class rather than an object gives false:

class Shape {}
class Circle < Shape {}

echo instance_of(Circle, Shape)
echo instance_of(Circle(), Shape)
false
true

This trips people up when a variable might hold either. Circle is a class; Circle() is an instance; only the second is “a Shape”.

The second argument must be a class, and here it does raise:

class Shape {}

catch {
  instance_of(Shape(), 'Shape')
} as e {
  echo e.message
}
instance_of() expects argument 2 to be a class, got string

The class name as a string is the common version of this mistake. Pass the class itself.

A class held in a variable works, because the check is on the value rather than on the spelling:

class Base {}
class Derived < Base {}

var Alias = Base

echo instance_of(Derived(), Alias)
true

So does a class reached through a module:

import set

echo instance_of(set.set([1, 2]), set.Set)
true

It Is How the Error Hierarchy Works

Every built-in error inherits from Error, so instance_of() is what lets one handler sort them:

echo instance_of(ValueError('bad'), Error)
echo instance_of(ValueError('bad'), TypeError)
true
false

That is the mechanism behind the “handle one kind, re-raise the rest” pattern in Error Handling, and behind mapping a domain error to an HTTP status in Chapter 25.

typeof(value)

typeof() asks a different question: “what exactly is this?”. On an instance it names the concrete class, and it never mentions ancestors:

class Shape {}
class Circle < Shape {}

echo typeof(Circle())
echo typeof(Shape())
echo typeof(Circle)
echo typeof(42)
echo typeof('text')
Circle
Shape
class
number
string

Note the third line. typeof() on the class itself answers class, not Circle — the class is a value of kind “class”, and its name is not what typeof() reports.

Which to Use

QuestionUse
can I treat this as a Shape?instance_of(x, Shape)
which exact class is this?typeof(x)
is this any instance at all?is_instance(x)

Reach for instance_of() by default. It is the one that respects inheritance, and code written with typeof(x) == 'Circle' breaks the day someone subclasses Circle — the subclass is a perfectly good Circle, and the string comparison says otherwise.

A Type Annotation Is the Same Check

Annotating a parameter with a class name performs exactly the instance_of() test, at every call, with a better error message:

class Shape {}
class Circle < Shape {}

def describe(shape: Shape) {
  return 'a shape of type ${typeof(shape)}'
}

echo describe(Circle())
echo describe(Shape())

catch {
  describe(42)
} as e {
  echo e.message
}
a shape of type Circle
a shape of type Shape
describe() expects parameter 'shape' (argument 1) to be a Shape, got number

Prefer the annotation when the answer decides whether the function should run at all, and instance_of() when the answer decides which branch to take.

Decorated Methods

A method whose name begins with @ is a decorated method. The runtime calls it for you when a piece of language syntax is applied to your instance. That is how a class hooks into construction, arithmetic, comparison and iteration without any special syntax of its own.

Decorated methods are ordinary methods. They can be inherited, overridden and called by name.

@new: Construction

Already covered. It runs when the class is called, receives the arguments, and is the one place self.x = value may declare a new field.

Arithmetic

Define the operator you want and it works on your instances:

class Vector {
  @new(x, y) {
    self.x = x
    self.y = y
  }

  @add(other) {
    return Vector(self.x + other.x, self.y + other.y)
  }

  @sub(other) {
    return Vector(self.x - other.x, self.y - other.y)
  }

  @mul(k) {
    return Vector(self.x * k, self.y * k)
  }

  @neg() {
    return Vector(-self.x, -self.y)
  }

  to_string() {
    return '(${self.x}, ${self.y})'
  }
}

var a = Vector(1, 2)
var b = Vector(3, 4)

echo (a + b).to_string()
echo (b - a).to_string()
echo (a * 3).to_string()
echo (-a).to_string()
(4, 6)
(2, 2)
(3, 6)
(-1, -2)

The full arithmetic set:

DecoratorOperator
@add+
@sub- (binary)
@mul*
@div/
@floordiv//
@mod%
@pow**
@neg- (unary)

Bitwise and Logic

DecoratorOperator
@and&
@or|
@xor^
@lshift<<
@rshift>>
@urshift>>>
@not~
class Flags {
  @new(bits) {
    self.bits = bits
  }

  @and(mask) {
    return self.bits & mask
  }

  @not() {
    return Flags(~self.bits)
  }
}

echo Flags(5) & 4
echo (~Flags(5)).bits
4
-6

@not is bound to ~, the bitwise complement. ! is logical negation and it is not overridable: an instance is always truthy, so !instance is always false.

Comparison

DecoratorOperator
@eq== and !=
@lt<
@lte<=
@gt>
@gte>=
class Version {
  @new(major, minor) {
    self.major = major
    self.minor = minor
  }

  @lt(other) {
    if self.major != other.major {
      return self.major < other.major
    }
    return self.minor < other.minor
  }

  @gt(other) {
    return other < self
  }
}

echo Version(1, 2) < Version(1, 10)
echo Version(2, 0) > Version(1, 10)
true
true

Each Operator Is Separate

There is no derivation between them. Defining @lt does not give you >, and defining @add does not give you += on the other side:

class Price {

  @new(cents) {
    self.cents = cents
  }

  @lt(other) {
    return self.cents < other.cents
  }
}

echo Price(250) < Price(500)

catch {
  echo Price(500) > Price(250)
} as e {
  echo e.message
}
true
operator '>' not defined for Price and Price

Define every operator you want to support. @gt is usually one line, as the Version example above shows.

The Other Operand Is Not Checked

A decorated method is an ordinary method, and its parameter is an ordinary parameter. Nothing guarantees the other side is the same class:

class Amount {

  @new(cents) {
    self.cents = cents
  }

  @add(other) {
    return Amount(self.cents + other.cents)
  }
}

catch {
  echo Amount(500) + 5
} as e {
  echo '${e.type}: ${e.message}'
}

catch {
  echo 5 + Amount(500)
} as e {
  echo '${e.type}: ${e.message}'
}
TypeError: cannot read property 'cents' on a number
TypeError: operator '+' not defined for call signature (number, Amount)

Read those two together. With the instance on the left, @add ran and failed inside your own method with a confusing message. With it on the right, the operator was never dispatched to your class at all — a decorated method only handles the case where its own instance is the left operand.

Annotate the parameter to fix the first message, and accept the second as the rule:

class Sum {

  @new(cents) {
    self.cents = cents
  }

  @add(other: Sum) {
    return Sum(self.cents + other.cents)
  }
}

catch {
  echo Sum(500) + 5
} as e {
  echo e.message
}
@add() expects parameter 'other' (argument 1) to be a Sum, got number

Equality

@eq defines == for your class, and != is always its negation:

class Point {
  @new(x, y) {
    self.x = x
    self.y = y
  }

  @eq(other) {
    if !instance_of(other, Point) {
      return false
    }

    return self.x == other.x and self.y == other.y
  }
}

var p = Point(1, 2)

echo p == Point(1, 2)
echo p != Point(1, 2)
echo p == Point(2, 1)
echo p == nil
true
false
false
false

@eq runs only when the value on the right is an object too: a string, a list, another instance and so on. p == nil, p == 5 and p == true compare the ordinary way without calling it, so a nil check stays a nil check whatever the class defines. It must return a bool; anything else raises a TypeError. Without @eq, two instances are equal only when they are the same object.

using matches through @eq as well, since it compares the way == does.

@eq decides == and nothing else. Lists and dictionaries compare the instances inside them by identity, and so do contains(), index_of() and dictionary keys:

class Point {
  @new(x, y) {
    self.x = x
    self.y = y
  }

  @eq(other) {
    if !instance_of(other, Point) {
      return false
    }

    return self.x == other.x and self.y == other.y
  }
}

var p = Point(1, 2)

echo [p] == [Point(1, 2)]
echo [p].contains(Point(1, 2))
echo [p].contains(p)
false
false
true

Iteration

@key and @value together make a class work with for ... in. They are the largest of the decorated methods to get right, so they have a section of their own: Making a Class Iterable.

DecoratorCalled bySignature
@keyfor ... in@key(previous)
@valuefor ... in@value(key)

@to_json

json.encode() calls @to_json() on an instance and encodes whatever it returns:

import json

class User {
  @new(name, password) {
    self.name = name
    self.password = password
  }

  @to_json() {
    return { name: self.name }
  }
}

echo json.encode(User('ada', 'hunter2'))
{"name":"ada"}

Without it, encoding an instance has nothing to work from. With it, you decide exactly what crosses the wire, which is the right place to leave a password behind.

@to_string

@to_string() decides what echo and print() show for an instance:

class Money {
  @new(cents) {
    self.cents = cents
  }

  @to_string() {
    return '$' + (self.cents / 100)
  }
}

var m = Money(500)

echo m
echo [m, Money(250)]
echo { total: m }
$5
[$5, $2.5]
{total: $5}

It applies wherever the instance sits in what is being shown, inside lists and dictionaries, as a key or as a value. Without it, an instance shows as <instance of Money>. It must return a string; anything else raises a TypeError.

to_string()

to_string() has no @ because it is not a decorator; it is a real method every value already has, and a class may override it to give its instances a plain string form.

Nothing calls it for you. String interpolation and + render an instance as <instance of Money>, so call it inside the interpolation. A class that has both usually builds one from the other:

class Money {
  @new(cents) {
    self.cents = cents
  }

  to_string() {
    return '$' + (self.cents / 100)
  }

  @to_string() {
    return '<Money ${self.to_string()}>'
  }
}

var m = Money(500)

echo m
echo 'cost: ${m.to_string()}'
<Money $5>
cost: $5

That is the split the standard library follows: to_string() is the value as text, and @to_string() is how the instance looks when you print it.

Keep both cheap and free of side effects. Error messages, logging and debugging all reach for them.

Encapsulation and Class Immutability

Private Members

A field or method whose name starts with _ is private. It can be reached through self or parent, and nowhere else:

class Box {
  var _items = []

  add(value) {
    self._items.append(value)
    return self
  }

  count() {
    return self._items.length()
  }
}

echo Box().add(1).add(2).count()
2

Reaching in from outside does not compile:

var b = Box()
echo b._items
SyntaxError: '_items' is private and can only be accessed via 'self' or 'parent'
  --> /path/to/main.zu:2:8
   |
 2 | echo b._items
   |        ^

This is a compile-time check, not a runtime one, so it costs nothing and cannot be worked around by computing the name.

A subclass can reach its parent’s private members, through self for fields and parent for methods:

class Base {
  var _secret = 'base secret'

  _internal() {
    return 'base internal'
  }
}

class Child < Base {
  reveal() {
    return self._secret + '/' + parent._internal()
  }
}

echo Child().reveal()
base secret/base internal

The same underscore convention governs modules: a module member whose name starts with _ cannot be imported by name. Chapter 8 covers that side of it.

Classes Are Sealed

Once a class is declared, its shape is final. You cannot add a field or a method to it, and you cannot add one to an instance:

class Account {
  var balance = 0
}

var a = Account()

catch {
  a.nickname = 'rainy day'
} as e {
  echo e.message
}
undefined field 'nickname' on instance of 'Account'

The practical consequence is that a misspelled field name is an error at the point you write it, rather than a new field that silently shadows the one you meant. Every field a class has is declared in one place, and that place is @new.

Static field values are mutable; the set of static fields is not.

Reflective Access

Four built-in functions read and write fields by name:

var b = Box()

echo hasprop(b, '_items')
echo getprop(b, '_items')
echo delprop(b, '_items')
echo getprop(b, '_items')
true
[]
true
nil

setprop(obj, name, value) writes; delprop resets the slot to nil. Both return false when the field does not exist on the class, because neither can create one.

These bypass the underscore rule, which is deliberate: they exist for serialisers, debuggers and test helpers, where reaching into an object is the entire point. Regular code should not use them.

The zuri module goes further, with a full reflection API over classes, functions and modules. Chapter 21 covers it.

Designing With Sealed Classes

Two habits follow from sealing.

Declare every field on the class, even the ones the constructor fills in. self.x = value inside @new will declare one for you, but a var line at the top of the body is documentation the next reader gets for free, and it is required the moment a helper method does the assigning instead.

Model optional state as a field holding nil, not as an absent field. There is no such thing as an absent field, so a var cached_result that starts nil is the shape you want.

A Worked Example

Privacy earns its place when a class has an invariant to protect. Here is a bounded history buffer: it keeps the last n entries and nothing else, and there is no way for a caller to break that from outside.

class History {
  var _entries = []
  var _limit = 0

  @new(limit: number) {
    if limit < 1 {
      raise ValueError('limit must be at least 1, got ${limit}')
    }

    self._limit = limit
  }

  record(entry) {
    self._entries.append(entry)

    if self._entries.length() > self._limit {
      self._entries.shift()
    }

    return self
  }

  # A copy, so a caller cannot append through the value we hand back.
  entries() {
    return self._entries.clone()
  }

  length() {
    return self._entries.length()
  }
}

var history = History(3)

history.record('a').record('b').record('c').record('d')

echo history.entries()
echo history.length()

var taken = history.entries()
taken.append('e')

echo history.entries()

catch {
  History(0)
} as e {
  echo '${e.type}: ${e.message}'
}
[b, c, d]
3
[b, c, d]
ValueError: limit must be at least 1, got 0

Four decisions are doing the work.

The list is private, so nothing outside can append to it and skip the trimming. history._entries.append('x') does not compile.

entries() returns a clone. Without it, the caller would hold the real list and could grow it past the limit — which is exactly what the fourth output line shows not happening. Handing out a private mutable collection is the most common way encapsulation leaks.

The invariant is established in the constructor. _limit is validated once, so record() never has to wonder whether it is sensible.

The class is sealed, so history.limit = 999 is an error rather than a second, ignored field sitting alongside _limit.

Making a Class Iterable

Any class can be walked with for ... in. It takes two decorated methods and no other ceremony: no interface to declare, no iterator object to return, and no state kept between passes.

This section covers the protocol, the two shapes it usually takes, what happens when a piece is missing, and the rules that keep a custom iterator well behaved.

The Protocol

for value in thing is not magic. Zuri evaluates thing once, then repeats three steps:

  1. key = thing.@key(key), starting from nil
  2. stop if key is nil
  3. value = thing.@value(key), and run the body

So the two methods answer two separate questions:

  • @key(previous) — “given the last position, what is the next one?” It receives nil on the first call and must return nil when there is nothing left.
  • @value(key) — “what is stored at this position?”

The loop never stores a cursor of its own. The key is the cursor, which is why @key gets the previous one back each time.

A Counting Example

class Countdown {

  @new(from) {
    self.from = from
  }

  @key(previous) {
    if previous == nil {
      return self.from
    }

    if previous <= 1 {
      return nil
    }

    return previous - 1
  }

  @value(key) {
    return key * 10
  }
}

for value in Countdown(3) {
  echo value
}
30
20
10

Trace it once and the protocol stops being abstract. @key(nil) returned 3, so the first value is @value(3), which is 30. @key(3) returned 2, then @key(2) returned 1, then @key(1) returned nil and the loop ended.

Note that the key and the value are genuinely different things here: the key counts down from three, the value is ten times it. Two variables in the loop give you both:

for key, value in Countdown(3) {
  echo '${key} -> ${value}'
}
3 -> 30
2 -> 20
1 -> 10

Wrapping a Collection

The more common case is a class holding a list, where the key is an index. The shape is always the same three checks:

class Stack {

  @new() {
    self.items = []
  }

  push(value) {
    self.items.append(value)
    return self
  }

  @key(previous) {
    if self.items.is_empty() {
      return nil
    }

    if previous == nil {
      return 0
    }

    if previous >= self.items.length() - 1 {
      return nil
    }

    return previous + 1
  }

  @value(key) {
    return self.items[key]
  }
}

var stack = Stack().push('a').push('b').push('c')

for value in stack {
  echo value
}

for index, value in stack {
  echo '${index}: ${value}'
}
a
b
c
0: a
1: b
2: c

Each of the three checks in @key earns its place:

The empty check comes first. Without it, @key(nil) would return 0 for an empty stack and @value(0) would index past the end.

previous == nil starts the walk, and 0 is the first index.

previous >= length() - 1 ends it. Using >= rather than == means a collection that shrank mid-loop still terminates.

An empty collection iterates zero times rather than failing:

for value in Stack() {
  echo 'never printed'
}

echo 'done'
done

What You Get for Free

Defining both methods makes is_iterable() answer true, and makes the class usable everywhere for is:

echo is_iterable(Stack().push('a'))
echo is_iterable(Countdown(1))
true
true

Both Are Required

for calls @key and then @value. Defining only one produces an error naming the one that is missing:

class OnlyKey {

  @key(previous) {
    return previous == nil ? 0 : nil
  }
}

catch {
  for value in OnlyKey() {
    echo value
  }
} as e {
  echo '${e.type}: ${e.message}'
}
PropertyError: undefined property '@value' on instance of 'OnlyKey'

A class with neither fails on @key instead:

class Plain {}

catch {
  for value in Plain() {
    echo value
  }
} as e {
  echo '${e.type}: ${e.message}'
}
PropertyError: undefined property '@key' on instance of 'Plain'

Rules Worth Knowing

A nil key ends the loop, always. That means a collection whose keys could legitimately be nil cannot use them as keys. Use indices and look the real key up in @value.

The instance is evaluated once. for x in build_thing() calls build_thing() a single time, before the first pass, so @key and @value are always called on the same object.

Neither method should mutate. They are called once per pass, in a loop you do not control, and a @key with a side effect is a loop whose behaviour depends on how many times something asked for the next key.

Iteration order is whatever @key says. There is no requirement to count upwards, or to visit everything; Countdown walks backwards, and a class could just as well skip, filter or repeat.

Class Extensions

An extension adds methods to a class that already exists. It is written with > rather than <, and it is a different thing from inheritance: inheritance creates a new class that borrows from an old one, while an extension modifies the old one in place.

class Account {

  @new(owner) {
    self.owner = owner
  }

  describe() {
    return 'account for ${self.owner}'
  }
}

class AccountExtras > Account {

  static shout(account) {
    return account.describe().upper()
  }

  static initials(account) {
    return account.owner[0]
  }
}

var a = Account('ada')

echo a.shout()
echo a.initials()
echo a.describe()
ACCOUNT FOR ADA
a
account for ada

Account gained two methods. There is no new class, no subclass, and no wrapper object: a is the same Account instance it was, and it now answers to shout() and initials() alongside the methods its own declaration gave it.

< Versus >

The two are easy to tell apart once you read the arrow as pointing at where the methods end up.

InheritanceExtension
Syntaxclass Child < Parentclass Anything > Target
Producesa new classnothing; it modifies Target
Affects existing instancesnoyes
May declare fieldsyesno
May declare @newyesno, it is not a field-holder
Methods areordinary methodsstatic, with the receiver as a parameter
self inside a methodthe instancenot available

The Rules

Four rules are enforced at compile time, and the error messages say exactly what is wrong.

The Name Is a Label

An extension’s own name is never bound to anything. It exists so the declaration reads like a declaration and so the error messages have something to say:

class AccountExtras > Account {
  static shout(account) { return 'hi' }
}

echo typeof(AccountExtras)
Unhandled UndefinedError: undefined global 'AccountExtras'

Name it after what it adds. AccountExtras, ShapeFormatting, NodeDebugging all read well; the name appears in no other code.

Every Method Must Be static

class Target {}

class Bad > Target {
  helper() {
    return 1
  }
}
SyntaxError: extension method 'helper' must be declared 'static' and receive the instance explicitly as their own first parameter if desired

static here does not mean the method ends up on the class rather than on instances. It means the method is compiled without an implicit receiver, so self does not exist inside it. The instance arrives as an ordinary argument instead.

No Fields

class Target {}

class Bad > Target {
  var cache = {}
}
SyntaxError: extension 'Bad' can only declare static methods but not fields

An extension adds behaviour, never state. A class’s set of fields is fixed when the class is declared, and nothing — including an extension — can add one afterwards. When your extension needs somewhere to keep something, keep it in the extending module, keyed by whatever identifies the instance.

The Target Must Exist

The target is evaluated when the extension is declared, so it has to be a name that is already bound to a class:

class Orphan > NoSuchClass {
  static m(x) {
    return 1
  }
}
Unhandled UndefinedError: undefined global 'NoSuchClass'

The name does not have to be a bare one. Anything an expression starting with an identifier reaches works, so a class another module owns can be named straight through that module:

import set

class SetExtras > set.Set {
  static summary(s) {
    return 'set of ${s.length()}'
  }
}

echo set.Set([1, 2, 3]).summary()
set of 3

Built-in types are not classes in scope, so string, list, number and the rest cannot be extended. Error and every class the standard library exports can be.

The Receiver

An extension method receives the instance as its first parameter. The name is yours to choose:

class Reading {

  @new(celsius) {
    self.celsius = celsius
  }
}

class ReadingConversion > Reading {

  static fahrenheit(reading) {
    return reading.celsius * 9 / 5 + 32
  }

  static scaled(reading, factor) {
    return reading.celsius * factor
  }
}

var r = Reading(100)

echo r.fahrenheit()
echo r.scaled(3)
212
300

r.scaled(3) passes r as reading and 3 as factor. Every argument at the call site shifts one place to the right of the receiver, exactly as it would for an ordinary method.

Because arity is not enforced, the receiver parameter is optional. A method that does not need the instance can simply not declare one:

class Widget {

  @new() {}
}

class WidgetVersion > Widget {

  static api_version() {
    return '1.0'
  }
}

echo Widget().api_version()
1.0

self Does Not Exist Here

class Thing {

  @new(n) {
    self.n = n
  }
}

class ThingBroken > Thing {

  static value(t) {
    return self.n
  }
}
SyntaxError: 'self' used outside of a method

Use the receiver parameter. t.n, not self.n.

Private Members Stay Private

An extension is outside code, and the underscore rule applies to it in full:

class Safe {
  var _hidden = 'secret'

  @new() {}
}

class Peek > Safe {

  static reveal(s) {
    return s._hidden
  }
}
SyntaxError: '_hidden' is private and can only be accessed via 'self' or 'parent'

This is the important limit on extensions. They can add behaviour built out of a class’s public surface, and they cannot reach inside it. An extension is not a way around encapsulation.

Replacing an Existing Method

An extension method with the same name as one the class already has replaces it:

class Greeter {

  @new(name) {
    self.name = name
  }

  greet() {
    return 'hello ${self.name}'
  }
}

var g = Greeter('ada')

echo g.greet()

class GreeterLoud > Greeter {

  static greet(greeter) {
    return 'HELLO ${greeter.name.upper()}'
  }
}

echo g.greet()
hello ada
HELLO ADA

Read that carefully. g was constructed before the extension was declared, and calling greet() on it after the declaration runs the new one. The replacement is not a shadow or a wrapper: the method table of the class itself changed, and every instance of it — past, present and future — sees the change.

This holds no matter how thoroughly the original has already been used. Build an eight-level tree, call the original count() four thousand times, then replace it:

var t = tree_with(8)
var total = 0

iter var i = 0; i < 4000; i++ {
  total = total + t.count()
}

echo total

class TreeNodeExt > TreeNode {

  static count(node) {
    return 999
  }
}

echo t.count()
2044000
999

Four thousand calls to the old method, and the next call runs the new one. Replacement is unconditional: there is no warm-up state, cached lookup or earlier result that can keep an old body alive past the extension.

Decorated Methods

An extension can add decorated methods, which is how you give operators, iteration or a string form to a class you did not write. The receiver still comes first, and the decorator’s own parameters follow it:

class Box {

  @new(n) {
    self.n = n
  }
}

class BoxOps > Box {

  static @add(a, b) {
    return Box(a.n + b.n)
  }

  static @key(box, previous) {
    return previous == nil ? 0 : nil
  }

  static @value(box, key) {
    return box.n
  }

  static to_string(box) {
    return 'Box(${box.n})'
  }
}

echo (Box(2) + Box(3)).n
echo Box(9).to_string()

for value in Box(7) {
  echo value
}
5
Box(9)
7

@add(a, b) gets the left operand as a and the right as b. @key(box, previous) gets the instance and then the previous key, which is the argument @key would normally take on its own.

They Are Instance Methods, Not Statics

static in the declaration describes how the method is compiled, not where it lands. The methods go onto instances:

class Item {

  @new() {}
}

class ItemExtras > Item {

  static label(item) {
    return 'an item'
  }
}

echo Item.label(Item())
Unhandled PropertyError: undefined static member 'label' on class 'Item'

Call it on the instance: Item().label().

Extensions and Inheritance

The two interact in one way worth knowing, and the rule is about when each declaration runs.

An extension changes the target’s method table. A subclass copies its parent’s methods when the subclass is declared. So a subclass sees an extension only if the extension came first:

class Vehicle {

  @new(wheels) {
    self.wheels = wheels
  }
}

class VehicleExtras > Vehicle {

  static doubled(vehicle) {
    return vehicle.wheels * 2
  }
}

class Car < Vehicle {

  @new() {
    parent(4)
  }
}

echo Vehicle(2).doubled()
echo Car().doubled()
4
8

Car was declared after the extension, so it inherited doubled. Move the class Car declaration above the extension and Car().doubled() raises undefined property 'doubled' on instance of 'Car', while Vehicle keeps it.

Declare extensions before the subclasses that should inherit them. In practice that means at the top of a module, or in a module imported at the top.

Extending a subclass never touches its parent:

class Animal {

  @new() {}
}

class Dog < Animal {

  @new() {}
}

class DogExtras > Dog {

  static speak(dog) {
    return 'woof'
  }
}

echo Dog().speak()

catch {
  echo Animal().speak()
} as e {
  echo e.message
}
woof
undefined property 'speak' on instance of 'Animal'

Extensions Are Global

This is the property that makes extensions powerful and the one that makes them worth using sparingly. An extension is not scoped to the module that declares it. It changes the class everywhere in the program, and it takes effect the moment the declaring module runs — which, for an imported module, is the moment it is imported.

Filename: shape.zu

class Shape {

  @new(name) {
    self.name = name
  }
}

Filename: extras.zu

import .shape { Shape }

class ShapeExtras > Shape {

  static label(s) {
    return 'shape: ${s.name}'
  }
}

Filename: main.zu

import .shape { Shape }

var s = Shape('circle')

catch {
  echo s.label()
} as e {
  echo 'before import: ${e.message}'
}

import .extras

echo s.label()
echo Shape('square').label()
before import: undefined property 'label' on instance of 'Shape'
shape: circle
shape: square

main.zu never mentions label’s definition. Importing extras — for any reason, including a reason unrelated to Shape — added a method to a class declared in a third file, and the instance created before the import gained it too.

Two consequences follow.

An import can change behaviour you did not ask it to change. A module that extends a class you use will alter it for your code as well, and nothing at your call site says so.

Two extensions of the same method collide silently. The last one to run wins, and the winner depends on import order:

class Slot {

  @new() {}
}

class First > Slot {

  static which(s) {
    return 'first'
  }
}

class Second > Slot {

  static which(s) {
    return 'second'
  }
}

echo Slot().which()
second

When to Use One

Extensions answer a question inheritance cannot: how do you add behaviour to a class whose instances are created somewhere you do not control? A subclass only helps if you are the one calling the constructor.

Good reasons:

Adding a view or a format to a domain class. to_string(), a @to_json, a summary() — presentation that does not belong in the model itself, added from the module that cares about it.

Giving a standard library class an operator or an iterator so it works with syntax it was not written for.

Instrumenting during debugging. Replacing a method with one that logs, then deleting the extension, is a diagnostic technique that costs nothing in the original file.

Reasons to think twice:

It is invisible at the call site. account.shout() gives no hint that shout lives in a different file from Account. A reader who greps the class body will not find it.

It is global. Everything above.

A plain function is usually enough. shout(account) is a function, it is obvious where it lives, and it cannot collide with anything. Reach for an extension when the thing genuinely needs to be a method — because syntax demands it, as with a decorated method, or because callers already have the instance and nothing else.

Error Handling

Things go wrong. A file is not there, a number arrives as text, a network peer stops answering. Zuri has one mechanism for all of it: an Error object, raised with raise and intercepted with catch. There is no try, no finally, and no checked exceptions.

This chapter covers raising, the three shapes of catch, what an error carries, the built-in hierarchy, writing your own, and the patterns that replace finally.

Raising

def withdraw(balance, amount) {
  if amount > balance {
    raise ValueError('cannot withdraw ${amount} from ${balance}')
  }
  return balance - amount
}

raise takes an instance of Error or any subclass, and nothing else. A bare string or number is a TypeError in its own right:

catch {
  raise 'something went wrong'
} as e {
  echo e.message
}
can only raise an Error or subclass, got a string

That rule is worth the small inconvenience: every value that travels through the error system has a type, a message and a stack trace, because there is no way to put anything else in.

An error that nobody catches ends the program with a message, a source excerpt and a stack trace:

Unhandled ValueError: cannot withdraw 100 from 50
  --> /path/to/main.zu:3

   1 | def withdraw(balance, amount) {
   2 |   if amount > balance {
>  3 |     raise ValueError('cannot withdraw ${amount} from ${balance}')
   4 |   }
   5 |   return balance - amount

Stack trace (most recent call last):
  at withdraw() /path/to/main.zu:3
  at @.script() /path/to/main.zu:8

Catching

catch {
  risky()
} as e {
  echo '${e.type}: ${e.message}'
}

The catch block runs. If anything inside it raises, execution jumps to the handler with the error bound to e. If nothing raises, the handler never runs.

There are three shapes, and each does something different.

catch { ... } with nothing after it swallows the error and carries on:

catch {
  raise Error('silent')
}
echo 'swallowed'
swallowed

Use this when failure genuinely does not matter, and nowhere else.

catch { ... } as e with no handler block binds the error to a variable that survives the statement. e is nil when nothing went wrong:

catch {
  raise Error('boom')
} as e

if e {
  echo e.message
}
boom

This is the shape to use when the recovery does not belong inside a handler, for example when you want to check several things in a row.

catch { ... } as e { ... } is the full form, and the one you will write most.

What an Error Carries

catch {
  raise ValueError('bad input')
} as e {
  echo e.type
  echo e.message
  echo e.stacktrace
}
ValueError
bad input
[/path/to/main.zu:2 -> @.script()]
  • message is the text.
  • type is the class name, as a string.
  • stacktrace is a list of frames, innermost first, each naming a file, a line and a function.

The stack trace is captured where the error was raised, not where it was caught, so it points at the origin no matter how many frames it travelled through:

def inner() {
  raise ValueError('deep')
}

def middle() {
  inner()
}

catch {
  middle()
} as e {
  echo e.stacktrace.length()
}
3

Three frames: inner, middle, and the script’s top level. Printing them is often the fastest way to answer “how did we get here?” in code you did not write.

The Built-in Errors

Every one of these is a class, and every one inherits from Error:

ClassRaised when
Errorthe base; a general failure
TypeErroran operation got the wrong type
ValueErrorthe type was right, the value was not
NumericErroran arithmetic operation failed
ArgumentErrorwrong number of arguments
NotImplementedErrora method that must be overridden was not
RangeErroran index or bound was out of range
AccessErrora permission or access check failed
AssertErroran assert failed
PropertyErrora member that does not exist was read
UndefinedErroran undefined name was read
ModuleNotFoundErroran import could not be resolved

Because they all inherit from Error, catching Error catches everything, and a parameter annotated Error accepts any of them.

Custom Errors

Subclass Error and carry whatever the caller needs:

class HttpError < Error {
  @new(message, status) {
    parent(message)
    self.type = 'HttpError'
    self.status = status
  }
}

catch {
  raise HttpError('not found', 404)
} as e {
  echo '${e.type} ${e.status}: ${e.message}'
  echo instance_of(e, Error)
}
HttpError 404: not found
true

Two things make this work well. Call parent(message) so the base constructor sets message and captures the stack trace. Set self.type so the class name shows up in logs and in the uncaught-error banner.

Deciding Which Error to Catch

catch catches everything inside its block. To handle one kind and let the others through, test and re-raise:

catch {
  load_config()
} as e {
  if !instance_of(e, ModuleNotFoundError) {
    raise e
  }
  echo 'no config, using defaults'
}

instance_of() walks the inheritance chain, so a test against Error matches everything and a test against HttpError matches only that branch.

There Is No finally

Code after the catch statement runs whether the block raised or not, because the handler either recovers or re-raises:

def cleanup_demo() {
  catch {
    raise Error('mid')
  } as e {
    echo 'caught'
  }

  echo 'always runs'
}

cleanup_demo()
caught
always runs

That covers the common case. When the handler re-raises, or when the block contains a return, the trailing code is skipped, so a resource that must be released either way goes in the handler as well:

var handle = file(path, 'w')

catch {
  write_everything(handle)
} as e {
  handle.close()
  raise e
}

handle.close()

The Bare as e Form Does It Once

The version above repeats handle.close(), once in the handler and once after the statement, because those are two different paths out. The third shape of catch — as e with no handler block — collapses them into one:

var handle = file(path, 'w')

catch {
  write_everything(handle)
} as e

handle.close()

if e {
  raise e
}

The difference is the missing { ... } after as e, and it changes the control flow rather than just the layout. With no handler there is nothing to jump into, so a raise inside the block is recorded in e and execution simply continues on the next line. Both paths — the one that raised and the one that did not — now run the same trailing code:

  fallthrough: closing
  fallthrough: re-raising
  caught: disk full

That gives you close() written once, running unconditionally, followed by an explicit decision about whether to re-raise. It is the closest thing Zuri has to finally, and it is assembled out of the ordinary pieces rather than being a separate construct:

LineJob
catch { ... } as erun it, record any failure, do not jump
handle.close()the cleanup, on every path
if e { raise e }pass the failure on, unchanged

Two things to be deliberate about.

e is nil when nothing went wrong, which is what makes if e { raise e } the whole of the decision. On the success path the cleanup runs and the function carries on normally.

Re-raising e itself keeps the original stack trace, pointing at the line that actually failed rather than at the raise you just wrote. Wrap it in a new error only when you have something to add, as Nesting and Re-raising shows.

Use the handler form when the failure needs handling. Use this form when it needs only cleaning up after.

return Inside catch

A return from inside a catch block or its handler leaves the enclosing function, exactly as it would anywhere else:

def find(items, needle) {
  catch {
    var i = items.index_of(needle)

    if i == -1 {
      raise ValueError('${needle} not in list')
    }

    return items[i]
  } as e {
    echo 'handled: ' + e.message
    return nil
  }
}

echo find([1, 2], 2)
echo find([1, 2], 9)
2
handled: 9 not in list
nil

Nesting

Handlers can raise, and an outer catch will see it:

catch {
  catch {
    raise Error('inner')
  } as inner {
    raise Error('outer: ' + inner.message)
  }
} as outer {
  echo outer.message
}
outer: inner

assert Versus raise

assert items.length() > 0, 'caller must pass a non-empty list'

assert raises an AssertError when its condition is falsy. The message is optional.

Draw the line like this. raise is for things that happen: a file that is not there, a number that does not parse, a network that times out. The caller is expected to handle it. assert is for things that cannot happen: an invariant the code itself is responsible for maintaining. A failed assertion means the program has a bug, not that the world was uncooperative.

Style

Keep the catch block small. The block should contain the operation that can fail and nothing else, so the handler is not accidentally catching a mistake somewhere further down:

Too wide — a failure inside render() is reported as a parse failure:

def show_wide(text) {
  catch {
    var data = json.decode(text)
    render(data)
  } as e {
    echo 'bad json'
  }
}

Right — only the call that can fail is inside the block:

def show_narrow(text) {
  var data

  catch {
    data = json.decode(text)
  } as e {
    echo 'bad json: ' + e.message
    return
  }

  render(data)
}

Note the var data outside the block. A catch block is a scope like any other, so a variable declared inside it is gone by the time the next statement runs.

Write messages that name the value:

raise ValueError('port must be between 1 and 65535, got ${port}')

The person reading that message is trying to work out what went wrong from one line of a log file. Give them the number.

Catching Inside a Loop

A catch inside a loop body handles one iteration and lets the rest carry on. This is the shape for processing a batch where individual items are allowed to fail:

def parse_positive(text) {
  if !text.match('/^\d+$/') {
    raise ValueError('not a number: ${text}')
  }

  var n = text.to_number()

  if n <= 0 {
    raise ValueError('must be positive: ${text}')
  }

  return n
}

var inputs = ['12', 'not a number', '30']
var total = 0
var rejected = []

for raw in inputs {
  catch {
    total += parse_positive(raw)
  } as e {
    rejected.append(raw)
  }
}

echo total
echo rejected
42
[not a number]

Put the catch outside the loop instead, and the first failure ends the whole loop — which is the right choice when one bad item makes the rest meaningless, and the wrong one when it does not. The placement of the block is the decision; there is no flag to set.

Nesting and Re-raising

A catch inside a handler works like any other, which is how you translate a low-level failure into one your caller understands:

import json

class ConfigError < Error {

  @new(message) {
    parent(message)
    self.type = 'ConfigError'
  }
}

def load(text) {
  catch {
    return json.decode(text)
  } as e {
    raise ConfigError('config is not valid JSON: ${e.message}')
  }
}

catch {
  load('{ broken')
} as e {
  echo '${e.type}: ${e.message}'
}

The caller now gets an error in its own vocabulary. Include the original message, as above, so the detail is not lost on the way up.

Re-raising the same error, rather than a new one, keeps the original stack trace pointing at the original line:

def only_handle_missing(work) {
  catch {
    return work()
  } as e {
    if !instance_of(e, ModuleNotFoundError) {
      raise e
    }

    return 'defaulted'
  }
}

echo only_handle_missing(@() => 'fine')

catch {
  only_handle_missing(@() { raise ValueError('not mine') })
} as e {
  echo '${e.type}: ${e.message}'
}
fine
ValueError: not mine

That is the pattern for “handle one kind and let everything else through”, and it is worth reaching for whenever a handler would otherwise swallow a bug along with the failure it meant to catch.

A Worked Example

Here is the whole chapter in one function: a loader that validates its input, distinguishes the failures a caller can act on from the ones it cannot, and closes what it opened on every path.

import json

class StoreError < Error {

  @new(message) {
    parent(message)
    self.type = 'StoreError'
  }
}

def read_records(path) {
  var handle = file(path)

  if !handle.exists() {
    raise StoreError('no store at ${path}')
  }

  var records

  catch {
    records = json.decode(handle.read())
  } as e {
    handle.close()
    raise StoreError('${path} is corrupt: ${e.message}')
  }

  handle.close()

  if !is_list(records) {
    raise StoreError('${path} should hold a list, found ${typeof(records)}')
  }

  return records
}

file('records.json', 'w').write('[{"id": 1}]')
echo read_records('records.json').length()

file('records.json', 'w').write('{ not json')

catch {
  read_records('records.json')
} as e {
  echo e.type
}

file('records.json').delete()

catch {
  read_records('records.json')
} as e {
  echo '${e.type}: ${e.message}'
}
1
StoreError
StoreError: no store at records.json

Four things in there are worth naming.

Every failure the caller might reasonably handle arrives as one error type, StoreError, so catch on the calling side needs one branch rather than three.

The message always names the path. A log line saying “file is corrupt” with no filename costs someone an hour.

handle.close() appears on both paths — once in the handler before the re-raise, once after the block. There is no finally to do it for you, and forgetting the one in the handler is the most common resource leak in Zuri code.

And the shape check at the end is a raise, not an assert. A file on disk containing the wrong thing is the world being uncooperative, not a bug in this function.

The Module System

A module is a .zu file. It runs at most once per program, it gets its own global namespace, and nothing it declares is visible anywhere else until someone imports it.

That is the whole model. The rest of this chapter is the syntax and the resolution rules.

Importing

import math
import os

echo math.PI
echo os.cwd()

import name binds the module under that name. Reach into it with a dot.

Picking Out Members

import math { PI, E }

echo PI

The named members are bound directly, and the module name itself is not.

{ * } brings in everything public:

import math { * }

echo PI
echo ROOT_2

Renaming

import http.websocket as ws

Use this when the natural name is long, or when it would collide.

as renames the module, not a member. There is no way to rename an individual name in a member list — import math { PI as pi } does not parse. When you want a different name for one imported thing, bind it yourself:

import math { PI }

var pi = PI

echo pi
3.141592653589793

Relative Imports

A path starting with . or .. is resolved against the directory of the file doing the importing, and is never searched for anywhere else.

import .helpers             # next to this file
import .models.user         # models/ next to this file, then user
import ..shared.config      # up one directory, then shared/, then config
import .. ..shared.config   # up two directories, then shared/, then config

Each .. climbs one directory, so .. .. climbs two and .. .. .. three. The dots may also be written together, ....shared.config, and it is the same import: a run of dots counts two for each directory up. zuri fmt leaves whichever way the path is written as it is.

Each segment resolves the same way a bare name does: name.zu is tried first, then name/index.zu. So import .helpers finds either helpers.zu or helpers/index.zu, and import .models.user finds models/user.zu or models/user/index.zu, with models itself being a directory either way.

Which one it finds is invisible at the import site, and that is what lets a module grow into a package: split helpers.zu into helpers/index.zu plus some siblings, and every import .helpers keeps working untouched.

Inside a package, a sibling is always import .sibling, never the full path from the project root. Writing import myapp.models.user from inside myapp/models/ sends the resolver out to the library search path and it will not find anything.

Exporting

Imports are local by default. If a.zu imports b.zu, a third file that imports a does not see b’s contents through it.

Prefix the path with @ to re-export:

import @.util { * }

Now everything util.zu exposed is part of this module’s public surface too. All three forms take the prefix:

import @.module            # the module itself is re-exported
import @.module { item }   # just that item
import @.module { * }      # everything

This is how a package’s index.zu assembles a public API out of several private files:

Filename: pkg/index.zu

import @.util { * }
import .sub.deep { deep_slug }

def hello() {
  return 'hello from pkg ' + VERSION
}

util’s members are re-exported. deep_slug is imported for this file’s own use and stays private.

Privacy

A member whose name starts with _ is private to its module and cannot be imported by name:

import .pkg.util { _secret }
SyntaxError: Cannot import private items from module
  --> /path/to/bad.zu:1:20
  |
1 | import .pkg.util { _secret }
  |                    ^

{ * } skips private members too. Prefix anything that is an implementation detail and the module system will keep it that way.

Packages

A directory with an index.zu is a package, and importing the directory runs its index.zu:

pkg/
  index.zu
  util.zu
  sub/
    deep.zu
import .pkg            # runs pkg/index.zu
import .pkg.util       # runs pkg/util.zu
import .pkg.sub.deep   # runs pkg/sub/deep.zu

The same rule applies to zuri run pkg on the command line, which is what makes a package runnable as well as importable.

How a Bare Name Is Resolved

import http, with no leading dot, is searched for in this order:

  1. .zuri/libs/http in the project
  2. $ZURI_ROOT/libs/http, or libs/http beside the zuri executable: the standard library
  3. a built-in native module named http
  4. $ZURI_HOME/libs/http, the packages installed for your user, which is ~/.zuri/libs unless ZURI_HOME moves it

At each step, http.zu is tried first and then http/index.zu.

The project is the nearest directory holding a project.toml, looking upwards from the script zuri run was given, or from the working directory for a command or the REPL. A program run from anywhere inside a project imports that project’s packages, and a program in no project uses .zuri/libs in the working directory. The project is found once, when the program starts, so changing directory part way through does not change where imports come from.

zuri install fills .zuri/libs, and Packages and Nyssa covers it in full. Step one is also what makes vendoring work. Dropping a file into .zuri/libs/ shadows a standard library module of the same name:

$ cat .zuri/libs/mylib.zu
def hi() { return 'from user libs' }

$ cat uses.zu
import mylib
echo mylib.hi()

$ zuri run uses.zu
from user libs

Modules Run Once

A module’s top level executes the first time it is imported, and never again. Every later import of the same file gets the same module object:

import .once
import .once as again
import .once
side effect ran

One line of output, three imports. Identity is by canonical filesystem path, so two different relative paths to the same file are the same module.

This makes a module’s top level the natural place for setup that must happen exactly once: opening a connection pool, reading a config file, registering handlers.

Circular Imports

A circular import works. A module is registered before its body runs, so when b.zu imports a.zu while a.zu is still loading, it gets the partially built module rather than looping forever:

Filename: a.zu

import .b

def from_a() {
  return 'a'
}

echo 'a loaded'

Filename: b.zu

import .a

echo 'b loaded'
b loaded
a loaded
a

The catch is visible in that output: b finished loading before a did, so anything b reads from a at its top level is not there yet. Reading it from inside a function is fine, because by then a has finished.

If a module fails while loading, it is dropped from the cache rather than left behind half built, so a later import genuinely retries.

Module Variables

Every module gets two names for free:

echo __file__
echo __root__

__file__ is this module’s own canonical path. __root__ is the entry file the program was started from, and it is the same in every module. Use them to locate files relative to your source rather than relative to whatever directory the user happened to run from:

import os

var templates = os.join_paths(os.dir_name(__file__), 'templates')

In the REPL both are defined, with placeholder values standing in for the file that does not exist:

%> __file__
@.repl
%> __root__
@.repl.root

They differ from each other there, so the __root__ == __file__ check below is false at the prompt — a REPL session is never the entry point of a program.

Running as a Program, Importing as a Module

Because __root__ is the entry file and __file__ is this one, comparing them tells a module whether it is the program being run or something being imported:

Filename: tool.zu

def add(a, b) {
  return a + b
}

if __root__ == __file__ {
  echo 'running as a program: ' + add(2, 3)
}
$ zuri run tool.zu
running as a program: 5
import .tool
echo tool.add(10, 20)
30

The same file is a clean, side-effect-free library when imported and a runnable command-line program when launched directly. Put the argument parsing and the entry point behind that check, and everything else above it.

These two names describe which file this is, so they are not exports. A wildcard import copies every public name out of the module it names, and __file__ and __root__ are deliberately excluded from that:

import math { * }

echo __file__ == __root__
true

That guarantee is what makes the __root__ == __file__ check above reliable in every file, including one that wildcard-imports a sibling.

Structuring a Project

A small program is one file. Past that, the shape that works is a package per area of responsibility, each with an index.zu that re-exports what is public:

myapp/
  index.zu          # import @.routes, import @.models, then start
  config.zu
  models/
    index.zu        # import @.user { * }, import @.task { * }
    user.zu
    task.zu
  routes/
    index.zu
    api.zu
    pages.zu
  storage/
    index.zu
    _json_store.zu  # private: the leading underscore says so

Run it with zuri run myapp. The capstone in Chapter 25 is laid out exactly this way.

A Package, End to End

Here is the smallest complete version of that shape. Three files, one package, one public entry point.

Filename: greet/english.zu

def hello(name) {
  return 'Hello, ${name}'
}

def _shout(text) {
  return text.upper()
}

Filename: greet/french.zu

def hello(name) {
  return 'Bonjour, ${name}'
}

Filename: greet/index.zu

import @.english
import @.french

def greet(name, language) {
  return language == 'fr' ? french.hello(name) : english.hello(name)
}

Filename: index.zu

import .greet

echo greet.greet('Ada', 'en')
echo greet.greet('Ada', 'fr')
echo greet.english.hello('Grace')
$ zuri run
Hello, Ada
Bonjour, Ada
Hello, Grace

Four things to take from it.

greet/index.zu is what import .greet loads. A directory with an index.zu is a package, and importing the directory runs that file.

The @ on import @.english is what re-exports it. Without it, greet.english would not be reachable from outside greet/index.zu, even though greet() itself would still work — which is often exactly what you want.

Both submodules define hello, and they do not collide. Each lives in its own namespace, reached through its own module name. That is the whole reason to use import @.english rather than import @.english { * } here.

_shout is unreachable from outside. Writing greet.english._shout(x) anywhere else is a compile error, not a runtime one:

SyntaxError: '_shout' is private and can only be accessed via 'self' or 'parent'

The leading underscore is the only declaration of privacy there is, and it is checked before the program runs.

Files and the Filesystem

Two things do the work here. The built-in file() function gives you a handle to one file. The os module covers everything else: directories, paths, globs, temporary files and the environment.

Opening a File

var handle = file('notes.txt')
var writer = file('notes.txt', 'w')

file(path) defaults to read mode. The second argument is the mode:

ModeMeaning
rread; the file must exist
wwrite; creates the file, truncates an existing one
aappend; writes always go to the end, creates the file
r+read and update; the file must exist
w+read and update; creates the file, does not truncate
a+read and append; creates the file
xwrite; creates the file, fails if the path already exists
x+read and write; creates the file, fails if the path already exists

Append b to any of them for binary mode: 'rb', 'wb', 'ab'.

x is the one to reach for when two programs might create the same file. The existence check and the creation happen in a single step, so exactly one of them succeeds and every other one fails, which is what a lock file needs.

Creating a handle does not touch the disk. Nothing happens until you read, write or call open().

Reading a Whole File

file('notes.txt', 'w').write('It works!')

echo file('notes.txt').read()
It works!

read() with no argument opens the file, reads all of it, and closes it again. That is the one-liner for “give me this file’s contents”, and it is what you want most of the time.

In text mode you get a string, decoded strictly as UTF-8. In binary mode you get bytes.

Reading in Chunks

read(length) reads at most that many bytes and leaves the handle open, so you can call it again. This is how you process a file too large to hold in memory at once:

file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')

var handle = file('notes.txt')
handle.open()

var chunks = 0

while true {
  var chunk = handle.read(6)

  if chunk.is_empty() {
    break
  }

  chunks++
}

handle.close()

echo chunks
3

Six bytes at a time is a demonstration; in real code the chunk is tens of kilobytes. The shape is what matters: open once, read until you get an empty result, close once.

Note that chunks fall wherever the byte count lands, not on line boundaries. A chunked reader that needs whole lines has to keep the tail of each chunk and join it to the front of the next.

Reading Lines

For text you want line by line, the simplest form reads the file and splits it:

file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')

for line in file('notes.txt').read().lines() {
  echo '[${line}]'
}
[alpha]
[beta]
[gamma]

lines() handles both \n and \r\n, and drops the trailing empty piece a final newline would otherwise produce.

Writing

file('notes.txt', 'w').write('It works!')

Like read(), write() opens the handle if it is closed, writes, flushes, and closes it again. One call, one complete file.

That auto-close has a consequence worth being precise about. Two consecutive write() calls on a closed handle in w mode each truncate the file, so only the last one survives. When you are writing more than once, open the handle yourself:

var handle = file('notes.txt', 'w')
handle.open()

handle.write('first line\n')
handle.write('second line\n')

handle.close()

echo file('notes.txt').read()
first line
second line

The explicit open() is what keeps the handle open across both writes. Without it, each write() would open, truncate, write and close, and only second line would survive.

An already-open handle is written to where it stands, so the sequence above does exactly what it reads like.

puts() writes without ever opening or closing. It requires an open handle, and it is the method to reach for inside a loop.

Reading and Writing at a Position

var handle = file('notes.txt')
handle.open()

handle.seek(6, 0)
echo handle.tell()
echo handle.read(4)

handle.close()

seek(offset, whence) takes 0 for the start of the file, 1 for the current position and 2 for the end. The io module names them:

import io

handle.seek(0, io.SEEK_SET)
handle.seek(-10, io.SEEK_END)

tell() reports the current offset.

Asking About a File

var handle = file('notes.txt')

echo handle.exists()
echo handle.path()
echo handle.abs_path()
echo handle.name()
echo handle.mode()
echo handle.is_open()
echo handle.is_closed()

stats() returns a dictionary of metadata about the file on disk:

file('notes.txt', 'w').write('alpha\nbeta\n')

var info = file('notes.txt').stats()

echo info.size
echo info.is_readable
echo info.keys()
11
true
[is_readable, is_writable, is_executable, is_symbolic, size, mode, dev, ino, nlink, uid, gid, mtime, atime, ctime, blocks, blksize]

size is in bytes. mtime, atime and ctime are epoch seconds, ready to hand to the date module. mode is the raw permission-and-type word, and the stat module is what turns it into an answer:

import stat

file('notes.txt').chmod(0c644)

var info = file('notes.txt').stats()

echo stat.S_ISREG(info.mode)
echo stat.S_ISDIR(info.mode)
echo stat.file_mode(info.mode)
true
false
-rw-r--r--

S_ISREG, S_ISDIR, S_ISLNK and the rest of the family each answer one question about the kind of entry. file_mode() renders the permission bits the way ls -l does.

Managing Files

file('notes.txt').copy('backup.txt')
file('backup.txt').rename('archive.txt')
file('archive.txt').delete()

There is also truncate(length), chmod(mode), set_times(access, modify) and symlink(target).

Always Close What You Opened

read() and write() clean up after themselves. Anything you opened with open() is yours to close, and the cleanest way to guarantee it is a catch that closes on the way out:

var handle = file(path, 'w')
handle.open()

catch {
  write_everything(handle)
} as e

handle.close()

if e {
  raise e
}

For something that has to be cleaned up however the program ends, rather than however one block ends, register it with os.at_exit():

import os

var scratch = file('scratch.txt', 'w')
scratch.open()

os.at_exit(@{
  scratch.close()
  scratch.delete()
})

scratch.write('working notes')
echo scratch.is_open()
true

Handlers run last registered first, and they run whether the program reached the end of its script, called os.exit(), or stopped on an uncaught error. A handler that raises is reported on standard error and the rest still run, so one failed cleanup cannot cancel the others.

Directories

import os

os.create_dir('sub/deep', nil, true)

The three arguments are the path, the permission bits, and whether to create intermediate directories. It returns false when the directory already existed.

echo os.dir_exists('sub')
echo os.is_dir('sub')
echo os.remove_dir('sub', true)

remove_dir’s second argument makes it recursive.

Listing and Globbing

Everything in this section works on a real tree, so build one first:

import os

os.create_dir('tree/sub', nil, true)

file('tree/top.txt', 'w').write('a')
file('tree/sub/nested.txt', 'w').write('b')
file('tree/sub/notes.md', 'w').write('c')

read_dir() lists one directory:

echo os.read_dir('tree')
echo os.read_dir('tree', true)
[., .., sub, top.txt]
[., .., sub, sub/nested.txt, sub/notes.md, top.txt]

Three things to notice. . and .. are included, so a loop over the result almost always wants to skip them. The recursive form returns nested entries as paths relative to the directory you asked about, not as bare names, which is what makes them usable directly. And entries come back sorted by name, with a directory’s contents following immediately after it, so the listing reads the same on every machine and filesystem.

glob() is usually what you actually want:

echo os.glob('*.txt', 'tree')
echo os.glob('**/*.txt', 'tree')
[top.txt]
[sub/nested.txt]

* matches within one path segment; ** matches across segments. The second argument is the base directory, and results come back relative to it.

Read those two results together, because the distinction catches people out. *.txt found the file at the top and not the nested one. **/*.txt found the nested one and not the top-level one, because **/ means “in a subdirectory”. Neither pattern finds both.

To match at every depth, glob ** and filter:

echo os.glob('**', 'tree')
echo os.glob('**', 'tree').filter(@(p) => p.ends_with('.txt'))
[sub, sub/nested.txt, sub/notes.md, top.txt]
[sub/nested.txt, top.txt]

** on its own matches every entry at every depth, directories included, which is why the filter is doing real work in the second line. glob() walks with read_dir() underneath, so matches arrive in that same sorted order.

Clean up when you are done:

echo os.remove_dir('tree', true)
true

Paths

Every path function is pure string manipulation except where noted:

echo os.join_paths('sub', 'a.txt')
echo os.base_name('sub/a.txt')
echo os.dir_name('sub/a.txt')
echo os.real_path('sub')
echo os.relative_path(os.cwd(), '/full/path/to/sub')
sub/a.txt
a.txt
sub
/home/you/project/sub
sub

real_path() resolves symlinks and requires the path to exist. abs_path() does not. expand_user() turns a leading ~ into the home directory. path_contains(base, candidate) answers whether one path is inside another, which is the check you need before serving a file a user named.

echo os.cwd()
echo os.home_dir()
os.change_dir('/some/where')

Temporary Files

echo os.temp_dir()

var path = os.create_temp_file('report-', '.csv')
var dir = os.create_temp_dir('build-')

Both create the thing and hand you its path. Clean them up yourself when you are done.

Environment Variables

echo os.get_env('HOME')
echo os.get_env('NOPE', 'fallback')

os.set_env('ZURI_BOOK', '1')
os.unset_env('ZURI_BOOK')

echo os.environ()
echo os.expand_vars('$HOME/projects')

get_env() takes a fallback. environ() gives the whole set as a dictionary.

Locating Files Relative to Your Code

The current working directory is wherever the user ran zuri from, which is not where your source lives. Use __file__:

import os

var HERE = os.dir_name(__file__)
var templates = os.join_paths(HERE, 'templates')

Doing this in every module that reads a file next to itself is the difference between a program that works and one that works only from the project root.

A Worked Example

Counting words across every text file in a directory tree, start to finish:

import os

def text_files(root) {
  return os.glob('**', root).filter(@(p) => p.ends_with('.txt'))
}

def word_count(root) {
  var counts = {}

  for path in text_files(root) {
    var text = file(os.join_paths(root, path)).read()

    for word in text.lower().split('/\W+/') {
      if word.is_empty() {
        continue
      }

      counts.set(word, counts.get(word, 0) + 1)
    }
  }

  return counts
}

os.create_dir('wc/sub', nil, true)

file('wc/a.txt', 'w').write('Hello world, hello!')
file('wc/sub/b.txt', 'w').write('World of Zuri.')
file('wc/sub/skip.md', 'w').write('not counted')

echo word_count('wc')

os.remove_dir('wc', true)
{hello: 2, world: 2, of: 1, zuri: 1}

Four things in that function are worth pointing at.

glob('**', root) returns paths relative to the root, so they have to be joined back onto it before they can be opened. Forgetting that join is the most common mistake in code that globs.

The .md file is absent from the result because text_files() filtered it out, which is the filter doing the job ** alone cannot.

split('/\W+/') is a regular expression, which is why punctuation does not end up in the keys — world, and hello! became world and hello. A plain split(' ') would have kept both.

And counts.get(word, 0) + 1 supplies the starting value for a key that does not exist yet. Without the fallback, the first sighting of every word would raise.

Where to Go Next

Binary files, byte streams and the io module’s in-memory files are Chapter 10. Reading a file over the network is Chapter 12. The complete list of methods a file handle carries is Appendix E.

Binary Data and Streams

Text is convenient. Protocols, file formats, images and checksums are not text, and this chapter is about the four tools Zuri gives you for them: the bytes type, the struct module, io.BytesIO, and compress.

The bytes Type

bytes is a mutable buffer of 8-bit values.

var b = bytes(3)
var c = bytes([1, 2, 3, 4, 5])
var d = 'Hi'.to_bytes()

echo b
echo c
echo d
(00 00 00)
(01 02 03 04 05)
(48 69)

bytes(n) allocates n zero bytes. bytes(list) builds one from numbers, each of which must be in 0..256. to_bytes() on a string gives you its UTF-8 encoding.

The hexadecimal-in-parentheses rendering is how bytes always prints, which makes it impossible to confuse one with a list at a glance.

Reading

var b = bytes([1, 2, 3, 4, 5])

echo b.length()
echo b[0]
echo b[1, 3]
echo b.first()
echo b.last()
echo b.get(2)
echo b.index_of(3)
echo b.last_index_of(3)
echo b.to_list()
echo 'Hi'.to_bytes().to_string()
5
1
(02 03)
1
5
3
2
2
[1, 2, 3, 4, 5]
Hi

Indexing gives a number. Slicing gives bytes. to_string() decodes as UTF-8, to_list() gives numbers.

index_of() and last_index_of() both return -1 when the byte is not there, and both take a second argument bounding where a match may sit, so they search the two halves either side of one index.

Slicing

b[a, b] takes the bytes from a up to but not including b, and returns a new byte stream:

var b = bytes([10, 20, 30, 40, 50])

echo b[1, 3]
echo b[, 3]
echo b[3, ]
echo b[-2, ]
(14 1e)
(0a 14 1e)
(28 32)
(28 32)

The rules are exactly the list’s. Either bound may be omitted: b[, n] starts at the beginning and b[n, ] runs to the end. Negative bounds count back from the end, so b[-2, ] is the last two bytes.

Note the difference between an index and a slice, because for bytes the two return different types:

var b = bytes([10, 20, 30])

echo b[0]
echo typeof(b[0])
echo b[0, 1]
echo typeof(b[0, 1])
10
number
(0a)
bytes

One index gives you the numeric value of that byte. A slice of length one gives you a byte stream containing it. Reaching for b[0] when you meant b[0, 1] is the most common slip here, and it shows up as a number where a bytes was expected rather than as an error at the slicing site.

Bounds Are Checked

A slice that runs past the end raises rather than returning what it can:

var b = bytes([10, 20, 30, 40, 50])

catch {
  echo b[1, 99]
} as e {
  echo '${e.type}: ${e.message}'
}
RangeError: slice bounds 1..99 out of range (length 5)

length() itself is always a legal upper bound, because the bound is exclusive, and an empty slice is legal rather than an error:

var b = bytes([10, 20, 30, 40, 50])

echo b[0, b.length()]
echo b[2, 2]
echo b[2, 2].is_empty()
(0a 14 1e 28 32)
()
true

An empty byte stream prints as ().

A Slice Is a Copy

Slicing allocates a new stream, so writing through one does not disturb the original:

var original = bytes([10, 20, 30])
var part = original[0, 2]

part[0] = 99

echo original
echo part
(0a 14 1e)
(63 14)

That matters when you are parsing a buffer. Pulling a header out with frame[0, 4] gives you something you can modify freely, and the frame you are still reading from is untouched. It also means slicing in a loop copies every time, so a parser that walks a large buffer should carry an offset and slice once per field rather than re-slicing the remainder each step.

Slice, Then Decode

The common shape when a buffer holds text with a known extent:

var b = bytes([72, 101, 108, 108, 111, 33])

echo b[0, 5].to_string()
echo b.to_string()
Hello
Hello!

to_string() decodes the whole stream it is called on, so the slice is what limits the extent. Slicing on a byte boundary in the middle of a multi-byte character produces a stream that is not valid UTF-8; decode whole units, or keep the tail for the next read.

There Is No Slice Assignment

A slice can be read but not written to:

var b = bytes([10, 20, 30])

b[0, 2] = bytes([1, 1])
SyntaxError: invalid assignment target

Assign to one index at a time, or rebuild the stream with extend(). A single index does accept assignment, and the value has to be a real byte:

var b = bytes([72, 101, 108, 108, 111, 33])

b[0] = 74

echo b.to_string()

catch {
  b[1] = 300
} as e {
  echo '${e.type}: ${e.message}'
}
Jello!
NumericError: bytes element must be an integer in 0..=255, got 300

Note the difference from bytes([300]), which wraps rather than raising. Construction is lenient; assignment is not.

Writing

var b = bytes([1, 2, 3])

b.append(4)
b.extend(bytes([9]))
echo b
echo b.pop()
echo b.reverse()
(01 02 03 04 09)
9
(04 03 02 01)

Unlike a list, bytes.reverse() mutates in place. So do append, extend, pop and remove.

dispose() releases the buffer’s memory immediately rather than waiting for the collector, which matters when you have just finished with something very large.

Splitting

split() takes a bytes delimiter, not a number:

echo bytes([1, 2, 3, 4, 5]).split(bytes([3]))
[(01 02), (04 05)]

Walking

A byte stream is iterable, and every form that works on a list works here. What you get out is always a number between 0 and 255, never a one-character string.

var b = bytes([72, 105])

for value in b {
  echo value
}
72
105

Two variables give you the index first and the value second:

var b = bytes([72, 105])

for index, value in b {
  echo '${index}: ${value}'
}
0: 72
1: 105

iter is the form to use when you are decoding a structure and the position drives the walk — reading a two-byte length, then skipping that many bytes, then reading the next field:

var b = bytes([72, 105, 33])

iter var i = 0; i < b.length(); i += 2 {
  echo '${i}: ${b[i]}'
}
0: 72
2: 33

And each() takes a function, with the value first and the index second as everywhere else:

bytes([72, 105]).each(@(value, index) {
  echo '${index}=${value}'
})
0=72
1=105

When you want the characters rather than the numbers, convert first: b.to_string() decodes the whole stream as UTF-8, and b.to_list() gives you the numbers as an ordinary list.

Binary Files

Append b to the file mode and reads give you bytes instead of a string, and writes accept bytes:

var out = file('data.bin', 'wb')
out.write(bytes([1, 2, 3]))
out.close()

echo file('data.bin', 'rb').read()
(01 02 03)

Reading a non-UTF-8 file without the b fails rather than silently producing replacement characters, which is the behaviour you want: a decoding failure is information.

struct: Packing and Unpacking

struct converts between Zuri values and a fixed binary layout. A format string is a sequence of /-separated fields, each CODE[COUNT][:NAME].

import struct

var header = struct.pack('N:magic/n:version/Z8:name/C:flags',
  0x5A555249, 1, 'task', 3)

echo header
echo header.length()
echo struct.unpack('N:magic/n:version/Z8:name/C:flags', header)
(5a 55 52 49 00 01 74 61 73 6b 00 00 00 00 03)
15
{magic: 1515541065, version: 1, name: task, flags: 3}

unpack() hands back a dictionary keyed by the :NAME you gave each field. Fields with no name get numbered:

echo struct.unpack('N/N', struct.pack('N2', 7, 8))
{1: 7, 2: 8}

A count greater than one under a single name numbers the keys:

echo struct.unpack('n3:vals', struct.pack('n3', 1, 2, 3))
{vals1: 1, vals2: 2, vals3: 3}

The Format Codes

Strings:

CodeSizeMeaning
acountNUL-padded string
Acountspace-padded string
ZcountNUL-padded and NUL-terminated, like C
h / Hceil(count/2)hex string, low or high nibble first

Integers. The letter tells you the width and the byte order:

CodeSizeMeaning
c / C1signed / unsigned 8-bit
?1boolean
s / S2signed / unsigned 16-bit, native order
n / v2unsigned 16-bit, big- / little-endian
i l / I L4signed / unsigned 32-bit, native order
N / V4unsigned 32-bit, big- / little-endian
q / Q8signed / unsigned 64-bit, native order
J / P8unsigned 64-bit, big- / little-endian
u / U16signed / unsigned 128-bit, little-endian

Floats:

CodeSizeMeaning
f / g / G4float: native / little / big endian
d / e / E8double: native / little / big endian
w / W2half-precision: little / big endian

Padding:

CodeMeaning
xwrite count NUL bytes, consumes no argument
Xback up count bytes
@seek to absolute position count

n, N and J are the network-order codes, which is what you want for almost every wire protocol.

A count of * means “the rest”.

Precision

Zuri numbers are doubles, so they hold integers exactly only up to 2^53. Every 64-bit and 128-bit code (q, Q, J, P, u, U) automatically produces a bigint when the unpacked value falls outside that range, rather than quietly losing digits. Packing accepts either.

The Rest of the Module

calcsize(format) gives the fixed byte size of a format with no * in it. pack_into(buffer, offset, format, ...) and unpack_from(format, buffer, offset) work on an existing buffer at a position, which is how you build one large frame without allocating and concatenating per field.

io.BytesIO

BytesIO is a file-shaped object backed by memory. It implements the same interface file() handles do, which means anything that takes a file takes a BytesIO too:

import io

var buffer = io.BytesIO(bytes(0), 'w')

buffer.write('hello ')
buffer.write('world')

echo buffer.source.to_string()
hello world

The constructor takes a bytes source and a mode. It has read, gets, write, puts, seek, tell, flush, close and stats, so a function written against files needs no changes to work in memory.

That is the useful part: test a function that writes a file without touching the disk, and parse an in-memory buffer with code written for a stream.

Compression

The compress module has a submodule per format, each with compress() and decompress():

import compress

var raw = ('the quick brown fox ' * 20).to_bytes()

echo raw.length()
echo compress.gzip.compress(raw).length()
echo compress.deflate.compress(raw).length()
echo compress.zstd.compress(raw).length()
echo compress.brotli.compress(raw).length()
echo compress.bzip2.compress(raw).length()
echo compress.lz4.compress(raw).length()
400
47
29
41
31
75
37

The formats are deflate, zlib, gzip, zstd, lz4, bzip2 and brotli. zlib is re-exported at the top level, so compress.compress() and compress.decompress() are the zlib pair:

echo compress.decompress(compress.compress(raw)).to_string() == raw.to_string()
true

Which to reach for: gzip when something else has to read it, zstd when you want the best ratio-to-speed trade, lz4 when speed is the only thing that matters, brotli for text you will serve over HTTP.

Archives

compress.tar and compress.zip read and write archives, and compress.checksum has crc32() and adler32():

import compress

echo compress.checksum.crc32('hello'.to_bytes())
907060870

Hashing and Encoding

base64 moves binary through text channels:

import base64

var encoded = base64.encode('hello'.to_bytes())

echo encoded
echo base64.decode(encoded).to_string()
aGVsbG8=
hello

hash covers the digest algorithms:

import hash

echo hash.sha256('hello')
echo hash.hmac_sha256('key', 'hello')
2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
9307b3b915efb5171ff14d8cb55fbcc798c6c0ef1456d66ded1a6aa723a58b7b

Every function takes an optional final as_bytes argument; pass true to get raw bytes instead of a hex string.

The algorithms are md2, md4, md5, sha1, sha224, sha256, sha384, sha512, the sha3 family, shake128/256, blake2b512, blake2s256, ripemd160, whirlpool and gost, each with an hmac_ counterpart, plus pbkdf2() for key derivation. hash.hash(algorithm, data) selects by name at runtime.

For password storage, use bcrypt rather than any of these. A fast hash is the wrong tool for a password, and Chapter 13 covers the right one.

A Worked Example: A Length-Prefixed Frame

Most binary protocols are a header followed by a payload. Here is the whole round trip:

import struct
import compress

var HEADER = 'N:length/C:compressed'

def encode_frame(payload) {
  var body = payload.to_bytes()
  var compressed = 0

  if body.length() > 128 {
    body = compress.gzip.compress(body)
    compressed = 1
  }

  var frame = struct.pack(HEADER, body.length(), compressed)
  frame.extend(body)

  return frame
}

def decode_frame(frame) {
  var size = struct.calcsize('N/C')
  var header = struct.unpack(HEADER, frame[0, size])
  var body = frame[size, size + header.length]

  if header.compressed == 1 {
    body = compress.gzip.decompress(body)
  }

  return body.to_string()
}

var frame = encode_frame('ping')
echo frame
echo decode_frame(frame)

echo decode_frame(encode_frame('the quick brown fox ' * 20)).length()
(00 00 00 04 00 70 69 6e 67)
ping
400

calcsize() tells you where the header ends, slicing gives you the two halves, and the compressed flag is one byte because a protocol that has to guess is a protocol that breaks.

Concurrency with Isolates

Zuri’s concurrency model has one rule: nothing is shared.

An isolate is a real operating-system thread running a complete, private VM. It has its own heap, its own garbage collector, its own register stack and its own module namespace. Two isolates never hold a pointer to the same object, which means there are no data races to reason about, no locks to take, and no stop-the-world pause in one thread caused by allocation in another.

Values cross between isolates by being copied. The copy is a faithful one: cycles are preserved, shared identity inside one payload is preserved, and class instances arrive as instances of the same class.

Spawning

import isolate

def square(n) {
  return n * n
}

var task = isolate.spawn(square, 12)

echo task.join()
144

spawn(fn, ...args) queues the call and returns an Isolate handle immediately. join() waits for it and gives you the result. Calling join() again returns the same value; it is not a one-shot.

What Can Be Spawned

Any function: a top-level def, a helper declared inside another function, a lambda, a bound method. What it captures travels with it, including the modules it uses:

import isolate
import json

def encode_all(items) {
  return items.map(@(item) => json.encode(item)).length()
}

echo isolate.spawn(encode_all, [{ a: 1 }, { b: 2 }]).join()
2

A module is not copied the way a list is. The isolate loads the same module for itself and uses its own copy, which is why this works at all: module top levels are declarations, they run once, and the result is cached. A module built out of live per-process state is the case to think twice about, since the isolate gets a fresh one rather than yours.

Two things still cannot be spawned. A native function or a class as the spawn target itself; wrap it in a def. And a resource handle such as an open socket or file is moved rather than copied, so the side that handed it over no longer has it.

Putting the worker in its own module stays the better structure once it grows past a few lines:

Filename: work.zu

import net

def serve(channel) {
  var listener = net.TcpStream()
  # ...
}

Filename: main.zu

import isolate
import .work

isolate.spawn(work.serve, channel)

That is the shape the rest of this book uses. A worker reached by name is resolved in the isolate’s own namespace, so nothing about it has to be reconstructed.

An import written inside the function body works too, and runs in the isolate:

import isolate

def make_id() {
  import uuid

  return uuid.v4().to_string().length()
}

echo isolate.spawn(make_id).join()
36

The arguments are copied, not shared. An isolate that mutates a list it was handed is mutating its own copy:

import isolate

def append_to(items) {
  items.append('from the isolate')
  return items.length()
}

var original = ['a']

echo isolate.spawn(append_to, original).join()
echo original
2
[a]

The isolate saw two elements. The caller still has one. That is the whole concurrency model in one example: there is no way to write a data race, because there is nothing shared to race over.

Running Many at Once

Spawn a list, join a list:

import isolate

def square(n) {
  return n * n
}

var tasks = [1, 2, 3, 4].map(@(n) => isolate.spawn(square, n))

echo tasks.map(@(task) => task.join())
[1, 4, 9, 16]

isolate.map() does exactly that in one call:

import isolate

def square(n) {
  return n * n
}

echo isolate.map(square, [1, 2, 3, 4, 5])
[1, 4, 9, 16, 25]

wait_all(tasks, timeout) waits for every task; wait_any(tasks, timeout) returns as soon as one finishes.

The Lifecycle of a Task

A spawned task moves through exactly two states, and four methods let you ask about it without blocking.

import isolate
import .work

var task = isolate.spawn(work.slow, 3)

echo task.try_join()
echo task.is_done()
echo task.status()

echo task.join()
echo task.status()
nil
false
pending
finished
done
MethodAnswersBlocks?
join(timeout)the resultyes
try_join()the result, or nil if not readyno
is_done()whether it has finishedno
status()'pending' or 'done'no
name()the name it was given, or nilno

try_join() returning nil is ambiguous when the task’s own result could be nil; pair it with is_done() when that matters.

join() may be called more than once. It is not a one-shot: the second call returns the same value immediately.

Naming a Task

spawn() leaves a task anonymous. spawn_named() gives it a name that shows up in diagnostics:

var named = isolate.spawn_named('importer', work.quick, 5)
var plain = isolate.spawn(work.quick, 5)

echo named.name()
echo plain.name()
importer
nil

Name anything long-running. A stuck program is far easier to diagnose when the task can say what it is.

Timeouts Are in Seconds

Every timeout in the isolate module is measured in seconds, and it takes a fraction:

import isolate

var empty = isolate.channel(1)
var start = time()

catch {
  empty.recv(0.3)
} as e {
  echo '${e.type} after about ${((time() - start) * 10).round() * 100}ms'
}
IsolateTimeoutError after about 300ms

This applies to join(), send(), recv(), select(), wait_any(), wait_all() and shutdown() alike.

It is worth stating loudly because the net module uses milliseconds for its own timeouts. socket.set_read_timeout(5000) is five seconds; channel.recv(5000) is an hour and twenty minutes. The two modules are easy to use in one program, and the mistake is silent — a timeout that never fires simply looks like a hang.

Omitting the timeout means “wait forever”, which is the right default when the other side is code you control and the wrong one when it is not.

Cancellation Is Cooperative

cancel() requests that a task stop. It does not kill anything.

What happens next depends on what the task is doing:

Blocked in a channel operation, a join(), a wait_* or a select() — the call is interrupted within roughly 50ms and raises IsolateCancelledError.

Running ordinary code — nothing is interrupted. The task’s own function must notice and return:

import isolate

def slow(seconds) {
  var start = time()

  while time() - start < seconds {
    if isolate.is_cancelled() {
      return 'stopped early'
    }
  }

  return 'ran to completion'
}

isolate.is_cancelled() is the module-level function a worker calls about itself. task.is_cancelled() is the method the spawner calls to ask whether it requested cancellation — and it answers true from the moment cancel() was called, whether or not the task noticed:

var task = isolate.spawn(work.slow, 3)

task.cancel()

echo task.join()
echo task.is_cancelled()
ran to completion
true

That output is not a contradiction. The worker in this example does not poll, so it finished normally; is_cancelled() reports the request, not the outcome. A loop with no cancellation check is a loop that cannot be stopped.

Put the check where the loop turns over, and make it cheap. Checking once per iteration of an outer loop is usually enough; checking inside the innermost arithmetic is not worth it.

What Can Cross, and What It Costs

Arguments in, results out, and channel traffic all cross the same way: everything is copied. There is no sharing and no reference that survives the boundary.

The Copy Is Faithful

A copy is not a shallow snapshot. Cycles survive, shared identity inside one payload survives, and a class instance arrives as an instance of the same class with its methods intact:

import isolate

class Point {

  @new(x) {
    self.x = x
  }

  doubled() {
    return self.x * 2
  }
}

def identity(value) {
  return value
}

# A list that contains itself.
var cyclic = [1]
cyclic.append(cyclic)

# Two slots holding one list.
var shared = [1]
var pair = [shared, shared]

echo isolate.spawn(identity, Point(3)).join().doubled()
echo typeof(isolate.spawn(identity, cyclic).join())

var back = isolate.spawn(identity, pair).join()
back[0].append(2)

echo back[1].length()
6
list
2

The last line is the one to notice. pair held the same list twice, and on the other side it still does: appending through back[0] is visible through back[1]. The copy preserved the shape of the sharing, not just the values.

What Cannot Cross

A handful of things are tied to the isolate that made them, and sending one raises rather than silently producing something broken:

import isolate

def identity(value) {
  return value
}

catch {
  isolate.spawn(identity, file('notes.txt'))
} as e {
  echo e.message.lines()[0]
}
cannot send a file across isolates; only nil, bool, number, string, bytes, bigint, range, list, dict, instance, class, bound method, function, module, and native-pointer values can cross

The message lists what can cross, which is the more useful half. An open file is the one you will meet: it is a handle onto a position in a descriptor this process owns, and there is no copy of that to hand over. Send the path instead and let the worker open it.

The Cost Model

Spawning is cheap. A task is a small descriptor holding the callee and a snapshot of its arguments, pushed onto a queue the pool drains, so queueing a hundred thousand of them is reasonable.

What is not free is the copy. Handing an isolate a large list copies the whole list, once per spawn. When a worker needs a lot of data, give it a description of the work — a path, a range of indices, a query — and let it do its own reading:

# Copies the file's contents into every worker.
isolate.map(work.process, files.map(@(p) => file(p).read()))

# Copies a short path into every worker instead.
isolate.map(work.read_and_process, files)

Errors Cross the Boundary

An uncaught error inside an isolate surfaces at the join() as an IsolateError, carrying the original error’s message and the stack trace from inside the worker:

import isolate

def boom() {
  raise ValueError('worker failed')
}

catch {
  isolate.spawn(boom).join()
} as e {
  echo e.type
  echo e.message.lines()[0]
}
IsolateError
ValueError: worker failed

Nothing is lost. You get the failure where you can do something about it, with enough information to find it. Note the shape of the message: the IsolateError wraps the worker’s own error rather than replacing it, so e.message still names ValueError and the line it came from.

An error in a task nobody joins has nowhere to surface. That is what scope(), further down, exists to prevent.

Channels

A channel is a bounded queue that crosses isolate boundaries. One side sends, the other receives, and the values are copied on the way across like everything else:

import isolate

def producer(channel, count) {
  iter var i = 0; i < count; i++ {
    channel.send(i)
  }

  channel.close()
}

var channel = isolate.channel(4)
var task = isolate.spawn(producer, channel, 3)

while true {
  var value = channel.recv()

  if value == nil and channel.is_closed() {
    break
  }

  echo value
}

task.join()
0
1
2

Read the loop condition carefully, because it is the part that is easy to get wrong. recv() returns nil both for “a nil was sent” and for “the channel is closed and drained”, so the test for the end of the stream is nil and is_closed(). Checking only for nil would stop early on a legitimate nil; checking only is_closed() would stop before the queue had drained.

channel(capacity) bounds the queue, which gives you backpressure for free: a producer that outruns its consumer blocks on send() rather than growing the queue without limit.

MethodBehaviour
send(value, timeout)blocks while the channel is full
recv(timeout)blocks while the channel is empty
try_recv()returns immediately, nil when empty
close()no more sends; pending receives drain, then return nil
is_closed()whether it has been closed
length()how many values are waiting

Every one of those behaviours is observable from a single isolate, which makes a channel easy to reason about before you introduce a second one:

import isolate

var c = isolate.channel(2)

c.send('a')
c.send('b')

echo c.length()
echo c.try_recv()

c.close()

echo c.is_closed()
echo c.recv()
echo c.recv()

catch {
  c.send('z')
} as e {
  echo '${e.type}: ${e.message}'
}
2
a
true
b
nil
IsolateError: cannot send on a closed channel

Read the last four lines together. After close(), the queue still drains: recv() returned the buffered b before it started returning nil. A closed, drained channel returns nil forever rather than raising, and only send() raises.

Closing twice is harmless. try_recv() on an empty channel returns nil immediately and never blocks.

Selecting Across Channels

select() waits on several channels at once and tells you which one produced a value:

import isolate

var c1 = isolate.channel(1)
var c2 = isolate.channel(1)

c2.send('from c2')

var picked = isolate.select([c1, c2], 1000)

echo picked[1]
from c2

It returns [channel, value], so you can tell where the value came from by comparing the first element against the channels you passed in. The second argument is a timeout in seconds, like every other timeout in this module — see Timeouts Are in Seconds, because it is the opposite of what the net module does.

This is the shape for a consumer fed by more than one producer — work on one channel, shutdown signals on another — without polling either.

Broadcast

A channel delivers each value to exactly one receiver. A broadcast delivers each value to every subscriber:

import isolate

var bus = isolate.broadcast(4)

var a = bus.subscribe()
var b = bus.subscribe()

bus.send('tick')

echo a.recv()
echo b.recv()
echo bus.subscriber_count()

bus.close()
tick
tick
2

subscribe() hands back an ordinary Channel, so everything in the channel table applies to it: recv(), try_recv(), length(), and the same backpressure.

A Subscriber Only Sees What Comes After It

There is no replay. A subscriber that arrives late has missed everything sent before it subscribed:

import isolate

var bus = isolate.broadcast(4)
var early = bus.subscribe()

bus.send('first')

var late = bus.subscribe()

bus.send('second')

echo 'early: ${early.recv()}, ${early.recv()}'
echo 'late: ${late.try_recv()}'
early: first, second
late: second

This matters when subscribers are set up concurrently with the producer. Subscribe everything before anything is sent, or accept that a consumer starting later begins mid-stream.

Unsubscribing, and Sending Into the Void

unsubscribe(channel) drops one subscriber. Sending with none at all is not an error — the value is simply discarded:

import isolate

var bus = isolate.broadcast(4)
var sub = bus.subscribe()

bus.unsubscribe(sub)

echo bus.subscriber_count()
echo bus.send('nobody is listening')
echo sub.try_recv()
0
nil
nil

A dropped subscriber stops receiving new values immediately, and its channel keeps whatever was already queued in it:

import isolate

var bus = isolate.broadcast(4)
var sub = bus.subscribe()

bus.send('queued before unsubscribe')
bus.unsubscribe(sub)

echo sub.try_recv()
queued before unsubscribe

So unsubscribing is not a way to discard what a consumer has not read yet; it only stops the flow.

Closing

close() closes the bus and every channel it handed out. Subscribers drain what they have and then receive nil; send() raises:

import isolate

var bus = isolate.broadcast(4)
var sub = bus.subscribe()

bus.close()

echo bus.is_closed()
echo sub.try_recv()

catch {
  bus.send('too late')
} as e {
  echo e.type
}
true
nil
IsolateError

Backpressure Applies Per Subscriber

The capacity you give broadcast() is the capacity of each subscriber’s channel. A subscriber that stops reading fills its own buffer, and once it is full the bus blocks on send() — one slow consumer holds up the producer and therefore everyone.

When a slow consumer must not be allowed to do that, give the bus a larger capacity, or have that consumer read into its own queue and fall behind on its own time.

Channel or Broadcast?

Use a channel when the work should be done once, by whichever worker is free. Use a broadcast when every consumer needs to see every event. Shutdown signals, configuration changes and progress events are broadcasts; jobs are channels.

Scopes

A scope owns its children and does not return until all of them have finished:

import isolate

def square(n) {
  return n * n
}

var result = isolate.scope(@(s) {
  s.spawn(square, 3)
  s.spawn(square, 4)

  return s.children().map(@(child) => child.join())
})

echo result
[9, 16]

scope(body) calls body with a scope object, waits for everything spawned through it, and returns whatever the body returned.

It Waits Whether or Not You Join

Joining inside the body is how you collect results. It is not how the waiting happens — that is the scope’s job either way:

import isolate

def square(n) {
  return n * n
}

echo isolate.scope(@(s) {
  s.spawn(square, 3)
  s.spawn_named('four', square, 4)

  return s.children().length()
})
2

Neither child was joined, and the scope still did not return until both had finished. children() gives you the Isolate objects in spawn order, and s.spawn_named() works exactly like the module-level one.

A Failure Stops the Group

This is the real reason to use a scope. A bare spawn() whose result is never joined swallows its own error in silence; a scope raises:

import isolate

def ok(n) {
  return n
}

def boom() {
  raise ValueError('worker failed')
}

catch {
  isolate.scope(@(s) {
    s.spawn(ok, 1)
    s.spawn(boom)

    return 'never returned'
  })
} as e {
  echo '${e.type}: ${e.message.lines()[0]}'
}
IsolateError: ValueError: worker failed

The body’s return value is discarded when a child failed, and the failure comes out of scope() itself. Reach for a scope whenever a piece of work fans out and must be complete — and correct — before the next step begins.

Waiting on Several Tasks

Three helpers cover the shapes that come up, and they return different things:

import isolate

def square(n) {
  return n * n
}

var tasks = [1, 2, 3].map(@(n) => isolate.spawn(square, n))

echo isolate.wait_all(tasks)
echo typeof(isolate.wait_any([isolate.spawn(square, 9)]))
echo isolate.map(square, [4, 5])
[1, 4, 9]
Isolate
[16, 25]

wait_all(tasks, timeout) returns the results, in the order the tasks were given, not the order they finished.

wait_any(tasks, timeout) returns the Isolate that finished first, not its result — you still call join() on it. That is what lets you tell which one won.

map(fn, items, timeout) spawns one task per item and collects the results, which is wait_all with the spawning done for you.

All three raise IsolateTimeoutError if the timeout passes, and all three propagate a worker’s failure:

import isolate

def halve(n) {
  if n == 0 {
    raise ValueError('cannot halve zero')
  }

  return n / 2
}

echo isolate.map(halve, [2, 4, 6])

catch {
  isolate.map(halve, [2, 0, 6])
} as e {
  echo '${e.type}: ${e.message.lines()[0]}'
}
[1, 2, 3]
IsolateError: ValueError: cannot halve zero

The whole call fails on the first failing element. When individual failures are acceptable, spawn and join yourself with a catch around each join(), as the worked example below does.

Isolates Can Spawn Isolates

There is no restriction on nesting. A worker may spawn its own tasks and join them, and they run on the same pool:

Filename: work.zu

import isolate

def double(n) {
  return n * 2
}

def nested(n) {
  return isolate.spawn(double, n).join()
}
$ zuri run main.zu
6

This is worth knowing mostly as a warning, which the next section covers: nested tasks that block on each other are the fastest way to exhaust the pool.

The Pool

Isolates do not each get a thread of their own. They run on a fixed pool of OS threads, sized to the machine’s core count by default:

import isolate

echo isolate.cpu_count() > 0
echo isolate.pool_size() == isolate.cpu_count()
true
true

Sizing It

configure(threads) changes the size, and it only works before the pool exists — which is to say, before the first spawn. It reports whether it did anything:

import isolate

echo isolate.configure(7)
echo isolate.pool_size()
true
7

Call it after a spawn and it returns false and changes nothing. It does not raise, so a configure() buried below some initialisation that already spawned will silently do nothing — put it at the very top of the entry file.

The size is a ceiling, not a head count. A thread is started only when an isolate is waiting and every thread already started is busy, so sizing the pool generously for a burst costs nothing while the burst is not happening. started_count() says how many threads the pool has started so far:

import isolate

isolate.configure(8)

echo isolate.started_count()
echo isolate.spawn(@() => 21 * 2).join()
echo isolate.started_count()
0
42
1

Why the Size Matters More Than It Looks

A pool sized to your core count can deadlock, and the failure looks like a hang rather than an error.

The mechanism: a task that blocks — on join(), on recv(), on a socket read — occupies its thread while it waits. If every thread is occupied by a task waiting for work that has no free thread to run on, nothing can ever progress.

That happens most easily with three patterns:

  • a server and a client that talk to each other in one process;
  • nested spawns where the parent joins the child;
  • a pipeline with more stages than threads.

Size the pool for the number of tasks that will be blocked at once, not for the number of cores. Threads that are blocked are not competing for CPU, so over-provisioning costs little; under-provisioning costs everything.

Watching It

import isolate

echo isolate.active_count() >= 0
echo isolate.queued_count() >= 0
echo isolate.is_shutdown()
true
true
false

active_count() is how many tasks are running; queued_count() is how many are waiting for a thread. A queued count that only grows is the signature of a starved pool.

shutdown(timeout) drains the pool and stops it. Most programs never call it; it is there for a long-lived process that wants to release its threads without exiting.

Memory While Waiting

A pool thread keeps its isolate’s heap from one task to the next, and the collector runs as a program allocates. A thread that waits allocates nothing, so a waiting thread collects on its own: once it has waited a second, for its next task or inside recv(), select(), join(), wait_any() or wait_all(), it collects its garbage and the memory that garbage held goes back to the system. The main thread does the same in those calls. A task that built and dropped a large structure leaves the process at its working size rather than its peak.

A wait of a second or less returns before that point, and a heap that has grown by less than 4 MB since its last such collection is left as it is, so a loop of short waits pays nothing for it.

The Errors

Three error types come out of this module, and they mean different things:

ErrorRaised when
IsolateErrora worker raised, or an operation is invalid — sending on a closed channel
IsolateTimeoutErrora timeout passed before the operation completed
IsolateCancelledErrora blocking call was interrupted by cancel()

All three inherit from Error, so catch on its own catches every one and instance_of() sorts them.

The distinction matters because they call for different responses. A timeout usually means retry or give up. A cancellation means shut down quietly. An IsolateError carrying a worker’s failure means something in your own code went wrong, and the message names it.

A Worked Example

Fan out a computation, collect the results, and report a failure without losing the rest:

import isolate

def classify(n) {
  if n < 0 {
    raise ValueError('negative input: ${n}')
  }

  if n < 2 {
    return 'small'
  }

  var divisor = 2

  while divisor * divisor <= n {
    if n % divisor == 0 {
      return 'composite'
    }

    divisor++
  }

  return 'prime'
}

def classify_all(numbers) {
  var tasks = numbers.map(@(n) => [n, isolate.spawn(classify, n)])
  var results = {}

  for pair in tasks {
    catch {
      results[pair[0]] = pair[1].join()
    } as e {
      results[pair[0]] = 'failed'
    }
  }

  return results
}

echo classify_all([1, 7, 9, -3, 97])
{1: small, 7: prime, 9: composite, -3: failed, 97: prime}

Three things are doing the work there.

Every task is spawned before any is joined. Spawning in one pass and joining in another is what makes the work overlap; joining inside the first loop would run them one after another and buy nothing.

The number is carried alongside its task, because the results come back in whatever order the pool produces them and a bare list of results would lose which input each belonged to.

And the catch is around one join(), so the one bad input becomes one failed entry rather than ending the batch. Move it outside the loop and the first failure takes the other four with it.

Choosing a Shape

Fan out, collect results. isolate.map(), or scope() when a failure must stop the group.

A pipeline. Channels between stages, each stage an isolate, each channel bounded so backpressure propagates all the way back to the source.

An event fan-out. A broadcast, one subscriber per consumer.

A server. http.serve() builds the whole pattern for you: one isolate per worker, a bounded backlog, connections handed out as they arrive. Chapter 15 covers it.

Nothing at all. Concurrency costs a copy at every boundary and a great deal of care at every join. A single-threaded loop that finishes in time is the better program.

Networking

The net module is the socket layer: TCP, UDP, unix domain sockets, TLS, DTLS, address parsing and polling. The http module sits on top of it and gives you a client and a server that speak HTTP/1.1 and HTTP/2.

Every example in this chapter runs a server in an isolate and a client in the main program, which is how you test network code without two terminals. Chapter 11 is the background.

TCP

A TcpStream is both ends. bind() and accept() make it a listener; connect() makes it a client. accept() blocks until someone connects, unless the listener was put in non-blocking mode, in which case it answers nil when nobody is waiting.

Filename: server.zu

import net

def echo_server(port_channel) {
  var listener = net.TcpStream()
  listener.bind('127.0.0.1:0')

  port_channel.send(listener.local_address())

  var client = listener.accept()
  var received = client.read_as_string()

  client.write_all('echo: ' + received)
  client.close()
  listener.close()
}

Filename: main.zu

import net
import isolate
import .server

var channel = isolate.channel(1)
var task = isolate.spawn(server.echo_server, channel)
var address = channel.recv()

var client = net.TcpStream()
client.connect(address)

client.write_all('hello')
client.shutdown(net.Shutdown.WRITE)

echo client.read_as_string()

client.close()
task.join()
echo: hello

Three things in that pair of files are doing real work.

The server lives in its own file. It has to: echo_server refers to net, and a function that references an imported module cannot be spawned from the file that did the importing. In server.zu, net is resolved inside the isolate instead of being captured from the caller. Chapter 11 covers the rule and the alternative.

Binding to port 0 asks the operating system for a free port, and local_address() reports which one it gave. That is the right way to write a test, and the right way to run several servers in one process. The port travels back to the main program through the channel, because the main program cannot know it in advance.

shutdown(net.Shutdown.WRITE) is not optional here. read_as_string() reads until the peer stops writing, so without a half-close from the client the server would still be waiting for more request while the client waits for a response — a deadlock that looks exactly like a hang. Shutdown has READ, WRITE and BOTH.

Reading

MethodBehaviour
read(length)up to length bytes, whatever has arrived
read_exact(length)exactly length bytes, waiting for them
read_all()everything until the peer closes, as bytes
read_as_string()the same, decoded as UTF-8
peek(length)look without consuming

read() returning fewer bytes than you asked for is normal, not an error. That is what read_exact() is for.

Writing

write(data) writes what it can and tells you how much. write_all(data) loops until everything is out, and is what you want almost always.

Options

socket.set_read_timeout(5000)
socket.set_write_timeout(5000)
socket.set_nodelay(true)
socket.set_non_blocking(true)
socket.set_ttl(64)

Timeouts here are milliseconds, which is the opposite of the isolate module’s seconds — easy to mix up in a program that uses both.

Set a read timeout on anything that talks to the network. A socket with no timeout and a peer that never answers is a thread that never comes back.

set_nodelay(true) disables Nagle’s algorithm, which is what you want for a request/response protocol where latency matters more than packet count.

UDP

import net

var a = net.UdpSocket()
a.bind('127.0.0.1:0')

var b = net.UdpSocket()
b.bind('127.0.0.1:0')
b.set_read_timeout(2000)

a.send_to('ping', b.local_address())

var pending = b.peek_from(1024)
echo pending.data.to_string()
echo pending.address.to_string()

echo b.receive_from(1024).to_string()

b.send_to('pong', pending.address)

a.set_read_timeout(2000)
echo a.receive_from(1024).to_string()
ping
127.0.0.1:<port>
ping
pong

The second line is written with a placeholder because the real one is not predictable: bind('127.0.0.1:0') asks the operating system for any free port, and it picks a different one every run.

receive_from(length) gives you the datagram’s bytes. To learn who sent it, peek_from(length) returns { data, address } without consuming the datagram, so the pattern for a responder is peek, read, reply.

A datagram larger than the buffer you give is truncated, and the rest is discarded. Size the buffer for your protocol’s largest message.

There is also connect() to fix a peer, set_broadcast(), and join_multicast_v4() / join_multicast_v6().

Unix Sockets

A unix domain socket is a path on the filesystem rather than an address on the network. It is how two programs on one machine usually talk when one of them is a server: databases, message brokers and system daemons nearly all listen on one.

There is no port to collide with and nothing to reach it from another host. And because the socket is a file, the filesystem decides who may connect — a socket in a directory only one user can enter is reachable by only that user, which is a stronger answer than binding to loopback and trusting everyone on the machine.

pair() gives two sockets already connected to each other, with no path and nothing left on disk:

import net

var ends = net.pair()

ends[0].write_all('ping')
echo ends[1].read_exact(4)

ends[1].write_all('pong')
echo ends[0].read_exact(4).to_string()

ends[0].close()
ends[1].close()
(70 69 6e 67)
pong

A server binds a path and accepts on it, exactly as a TcpStream binds an address:

import net

var server = net.UnixStream()
server.bind('/run/app.sock')

while true {
  var client = server.accept()
  client.write_all('hello\n')
  client.close()
}

Two things differ from TCP and both come from the socket being a file. bind() refuses a path that already exists, including one a crashed process left behind, so a server that expects to be restarted deletes a stale path first. And closing frees the descriptor without removing the file, because by then another process may have bound the same path and removing it would break them.

net.is_supported() is false on Windows, where these are not available. Every call raises there rather than the module being missing, so a program that can fall back to TCP tests it rather than catching an error.

Addresses

import net

var address = net.SocketAddr.parse('127.0.0.1:8080')

echo address.to_string()
127.0.0.1:8080

resolve() turns a name into addresses:

echo net.resolve('localhost:80').map(@(a) => a.to_string())
[[::1]:80, 127.0.0.1:80]

It returns a list because a name can have several addresses, and what comes back depends on the machine: a host with IPv6 configured answers for both families, one without gives only [127.0.0.1:80]. Never assume a position in that list, and never assume a family.

net.ip has IpAddress, Ipv4Address and Ipv6Address for parsing, comparing and classifying addresses, which is what you want before you trust an X-Forwarded-For header.

TLS

net.tls wraps an established TCP stream:

import net
import net.tls

var config = tls.TlsConfig()
config.set_cert_chain(certificate_pem, private_key_pem)

var listener = net.TcpStream()
listener.bind('127.0.0.1:0')

var client = listener.accept()
var secure = tls.TlsStream.accept(client, config)

echo secure.read_as_string()

A TlsStream has the same read and write interface a TcpStream does, so code written against one works against the other.

On the client side, tls.TlsStream.connect(socket, config, hostname) performs the handshake and verifies the certificate. config.add_ca_pem() adds a trust anchor, and config.require_client_cert(true) turns on mutual TLS. peer_certificate() gives you the other end’s certificate once the handshake is done.

net.dtls is the same thing over UDP.

Polling

net.poll waits on many sockets at once without a thread each. It is the right tool when you have thousands of mostly-idle connections and the wrong tool when you have a handful of busy ones, where an isolate per connection is simpler and faster.

Accepting without parking the thread

accept() on a blocking listener waits inside the runtime. Nothing else on that thread gets a turn while it waits: a signal trapped with os.on_signal() is not delivered, and a flag telling the server to stop is not read, until a connection happens to arrive. On an idle server that is never, which is why Ctrl+C on one can appear to do nothing at all.

net.Acceptor is the accept loop without that problem. It puts the listener into non-blocking mode and waits on a poller instead. next() hands over a connection when there is one and returns nil when its interval passes with nothing arriving, which is the loop’s chance to look at whatever else it has to look at.

import net

var listener = net.TcpStream()
listener.bind('127.0.0.1:0')

var acceptor = net.Acceptor(listener, 50)

# Nobody has connected yet, so the round comes back empty.
echo acceptor.next()

var client = net.TcpStream()
client.connect(listener.local_address())

var served = acceptor.wait_for_one()

echo served.peer_address().ip().to_string()

served.close()
client.close()
listener.close()
nil
127.0.0.1

The interval decides how soon a stopped server notices, not how soon a connection is served: one that arrives wakes the wait at once. Nothing is added while connections are arriving either, because next() tries accept() first and reaches for the poller only when the queue is empty.

That turns a server into something Ctrl+C can stop:

import net
import os

var listener = net.TcpStream()
listener.bind('127.0.0.1:8080')

var running = true

os.on_signal('INT', @() {
  running = false
  return true
})

var acceptor = net.Acceptor(listener)

while running {
  var client = acceptor.next()

  if client == nil {
    continue
  }

  serve(client)
}

listener.close()

The handler returns true to say it has taken responsibility for the signal. A handler that returns anything falsy declines, and the process then dies of the signal as it would have with nothing registered at all.

http accepts this way in both HttpServer.listen() and http.serve(), and so do the SMTP and IMAP servers in mail. A server built on any of them is already stoppable.

Above the Socket Layer

Everything so far has been bytes on a socket. Most programs want a protocol on top of that, and the standard library brings two of them.

http is a complete HTTP/1.1 and HTTP/2 client and server: routing, middleware, cookies, multipart uploads, static files, server-sent events and WebSockets. It is large enough to have its own chapter, and Chapter 15 is it. The one-line version:

import http

var response = http.get('https://example.com')

echo response.status
echo response.as_text()

net.tls wraps a TcpStream in TLS, as shown above, for protocols that are not HTTP — a mail client, a database driver, a custom binary protocol.

A rough guide to which layer you want:

You are writingReach for
a web API, or a client for onehttp
a browser-facing serverhttp
a client for an existing non-HTTP protocolnet.tcp, plus net.tls if it is encrypted
a protocol of your own designnet.tcp and struct
discovery, telemetry, gamesnet.udp
anything waiting on many sockets at oncenet.poll
a server that has to stop on Ctrl+Cnet.Acceptor

A Worked Example

A length-prefixed request/response protocol of the kind net is for: each message is a four-byte big-endian length followed by that many bytes of JSON. This is the pattern behind most binary protocols, and it is worth writing once by hand.

Filename: framed.zu

import struct

# Reads one frame, or returns nil once the peer has hung up.
#
# read_exact() raises rather than returning short when the stream ends
# mid-read, and a clean disconnect between frames looks exactly like
# that, so the end of the conversation arrives here as an error.
def read_frame(stream) {
  var header

  catch {
    header = stream.read_exact(4)
  } as e {
    return nil
  }

  var length = struct.unpack('N:size', header).size

  return stream.read_exact(length).to_string()
}

def write_frame(stream, payload) {
  stream.write_all(struct.pack('N', payload.length()))
  stream.write_all(payload)
}

Filename: server.zu

import net
import json
import .framed

def serve(port_channel) {
  var listener = net.TcpStream()
  listener.bind('127.0.0.1:0')

  port_channel.send(listener.local_address())

  var client = listener.accept()

  while true {
    var request = framed.read_frame(client)

    if request == nil {
      break
    }

    var parsed = json.decode(request)

    framed.write_frame(client, json.encode({ reply: parsed.message }))
  }

  client.close()
  listener.close()
}

Filename: main.zu

import net
import json
import isolate
import .framed
import .server

var channel = isolate.channel(1)
var task = isolate.spawn(server.serve, channel)

var client = net.TcpStream()
client.connect(channel.recv())
client.set_read_timeout(2000)

framed.write_frame(client, json.encode({ message: 'first' }))
echo framed.read_frame(client)

framed.write_frame(client, json.encode({ message: 'second' }))
echo framed.read_frame(client)

client.close()
task.join()
{"reply":"first"}
{"reply":"second"}

The framing is the whole point. TCP is a stream of bytes with no message boundaries in it: one write_all() may arrive as three reads, and three writes may arrive as one. read_exact(4) followed by read_exact(length) is what puts the boundaries back, and read() alone would not — it returns whatever has arrived, which is why the table above distinguishes the two.

The end of the conversation is the other thing worth studying. read_exact() raises when the stream ends before it has the bytes it was promised, and a peer that hangs up cleanly between frames produces exactly that. Catching it and returning nil is what turns “the connection closed” from a crash into the loop’s normal exit.

Note also that both ends share framed.zu. A protocol implemented twice, once per end, is a protocol that will eventually disagree with itself.

One last detail, easily missed: the reply key is reply, not echo. echo is a keyword, so it cannot be a bare dictionary key — { echo: x } is a syntax error. Quote it as { 'echo': x } if you need that exact name.

The Rest of the Module

net.poll answers “which of these sockets can I read right now?” without a thread per socket, which is how you serve many connections from one isolate. It takes a UnixStream alongside a TcpStream or a TlsStream. net.addr and net.ip parse, format and classify addresses — is_private(), is_loopback(), is_multicast() and the rest — which is what you want before trusting an address a client sent you. net.dtls is TLS over UDP.

Appendix F lists every submodule.

A Tour of the Standard Library

The standard library is part of the installation. There is nothing to add to a manifest, no package manager to run, and no dependency to resolve; import json works in a file you created ten seconds ago.

This chapter walks through what is there, grouped by the job you would reach for it to do, with a working example of each. It is a tour rather than a reference: Appendix F is the index, every module carries doc blocks in libs/, and the three largest modules have chapters of their own — Wire, HTTP and Imagine.

Read it once to learn what exists. The value of a tour like this is not remembering the details; it is recognising, six months from now, that the thing you are about to write by hand is already here.

Data Formats

json

import json

var data = { name: 'Ada', langs: ['zuri', 'rust'], active: true }

echo json.encode(data)
echo json.decode('{"a":1}').a
echo json.encode({ name: 'Ada' }, false)
{"name":"Ada","langs":["zuri","rust"],"active":true}
1
{
  "name": "Ada"
}

encode(value, compact, max_depth) defaults to compact. parse(path) reads and decodes a file; dump(value, file) writes one. A class that defines @to_json() controls its own encoding.

yaml

import yaml

echo yaml.parse('name: zuri\ntags:\n  - fast\n  - small')
{name: zuri, tags: [fast, small]}

Anchors, aliases, tags, multi-document streams and block scalars are all supported.

toml

import toml

echo toml.parse('[package]\nname = "zuri"\nversion = "1.0"\n')
{package: {name: zuri, version: 1.0}}

parse() and dump() treat a document as data. edit() treats it as a file somebody wrote: the Document it returns renders back byte for byte until it is changed, and a change disturbs only the line it lands on.

import toml

var doc = toml.edit('# what we ship\n[package]\nversion = "0.9.0"  # bump me\n')
doc.set('package.version', '1.0.0')

echo doc.to_string()
# what we ship
[package]
version = "1.0.0"  # bump me

That is what a program editing somebody else’s configuration file needs: the comment stayed, and so did the spacing on the line that changed.

csv

import csv

echo csv.parse('a,b\n1,2')
[[a, b], [1, 2]]

Reader and Writer stream large files, Dialect configures separators and quoting, and sniff_dialect() guesses from a sample.

struct

Binary layouts. Covered in Chapter 10.

base64, convert

import convert

echo convert.bytes_to_hex(bytes([255, 0]))
echo convert.to_base(255, 16)
echo convert.from_base('ff', 16)
ff00
ff
255

convert handles every base-to-base conversion you would otherwise write by hand, plus hex, binary, octal and unicode helpers.

Text and Markup

html

A WHATWG-conformant parser, a real DOM, and CSS selectors:

import html

var doc = html.parse('<ul><li class="a">one</li><li>two</li></ul>')

echo doc.query_selector('li.a').text_content()
echo doc.query_selector_all('li').length()
one
2

The DOM supports traversal, mutation and serialisation, which makes it a scraper, a templating backend and a sanitiser in one module.

wire

Templating, with directives expressed as HTML attributes rather than a second syntax layered over your markup:

import wire

echo wire.render_string('<p x-text="msg"></p>', { msg: 'hi' })
<p>hi</p>

Everything is escaped by default, and escaped correctly for where it sits: a value in an attribute, in a URL and in a <script> block are three different escapes, and Wire knows which is which because it parses your template as structure rather than text.

Chapter 14 is the full treatment.

url

import url

var parsed = url.parse('https://user@example.com:8443/a/b?q=1#top')

echo parsed.host
echo parsed.port
echo parsed.get_param('q')
example.com
8443
1

encode(), decode() and parse_query() handle percent-encoding.

mime

import mime

echo mime.detect_from_name('a.png')
image/png

detect(file) sniffs content rather than trusting the extension, which is the check you want on an upload.

colors

ANSI colour for terminal output, and conversion between every colour space you are likely to have a value in:

import colors

echo colors.hex_to_rgb('#ff8800')
echo colors.rgb_to_hex(255, 136, 0)
echo colors.rgb_to_ansi256(255, 136, 0)
[255, 136, 0, 1]
ff8800
214

colors.text(value, color, background) wraps a string in the escape codes:

import colors

echo colors.text('warning', colors.text_color.yellow)

The conversions are what make the same code work on a terminal that cannot do what you asked: a true-colour value becomes the nearest of 256, and 256 becomes the nearest of 16. hex(), rgb(), hsl(), hsv(), hwb(), cmyk() and xyz() each take a colour in that space and produce the escape sequence for it.

Time

date

import date

var d = date.date(2026, 9, 11, 8, 30, 0)

echo d.format('Y-m-d H:i:s')
echo d.format('l, jS F Y')
echo date.parse('2026-09-11').format('Y-m-d')
2026-09-11 08:30:00
Friday, 11th September 2026
2026-09-11

The format codes are single letters: Y four-digit year, m zero-padded month, d zero-padded day, H 24-hour, i minutes, s seconds, l weekday name, F month name, jS day with an ordinal suffix.

localtime() and gmtime() give the current time, from_time(seconds) converts a Unix timestamp, and the module carries a real IANA time zone database, so from_timezone('Europe/London', ...) does the right thing across a daylight-saving boundary.

Cryptography and Identity

hash

Digests and HMACs. Covered in Chapter 10.

bcrypt

Password hashing, which is a different problem from digesting:

import bcrypt

var stored = bcrypt.hash('secret')

echo bcrypt.compare('secret', stored)
echo bcrypt.get_rounds(stored)
true
10

Use this for passwords and hash for everything else. needs_rehash() tells you when a stored hash was made with a lower cost than you now require.

crypto

RSA signing and HKDF key derivation.

uuid

import uuid

echo uuid.v4().length()
echo uuid.is_valid(uuid.v7())
36
true

Versions 1, 3, 4, 5, 6, 7 and 8 are all there. v4 is the random one you usually want; v7 is time-ordered, which makes it a better database key.

jwt

Signing, verifying and decoding JSON Web Tokens, with a JWKS client for rotating keys.

Validation and Structure

validate

A fluent schema builder:

import validate

var schema = validate.schema({
  name: validate.required().string().max_length(10),
  age: validate.required().integer().min(0),
})

echo schema.check({ name: 'Ada', age: 36 })
echo schema.check({ name: 'a name that is far too long', age: 200.5 })
{valid: true, errors: []}
{valid: false, errors: [{field: name, message: The name field must not exceed 10 characters.}, {field: age, message: The age field must be an integer.}]}

check_or_raise() raises instead of returning. extend(), only() and except() build one schema from another, which is how a create schema and an update schema stay in sync.

types

Type predicates, one per type, as an alternative to the is_* built-ins:

import types

echo types.of(42)
echo types.int(42)
echo types.int(4.2)
echo types.digit('7')
echo types.alpha('a')
echo types.iterable([1])
echo types.instance(ValueError('x'), Error)
number
true
false
true
true
true
true

Each one answers a question and returns a boolean; none of them converts anything. types.int(4.2) is false because 4.2 is not an integer, not because it failed to become one.

types.of() is typeof(). digit(), alpha() and char() are the string-shape tests the built-ins do not cover, and instance(value, Class) walks the inheritance chain.

When you want conversion rather than a question, the methods on the value do it: to_number(), to_string(), to_bigint(), to_list(), to_bytes().

set

import set

var s = set.set([1, 2, 2, 3])
echo s.length()
3

Union, intersection, difference and subset tests, with insertion order preserved.

enum

import enum

var Color = enum.enum(['RED', 'GREEN'])
echo Color.RED
0

Pass a dictionary instead of a list to choose the values yourself.

array

Typed, fixed-width numeric arrays: Int8Array through Uint64Array, plus FloatArray and DoubleArray. They store values in their declared width rather than as doubles, which matters for memory and for talking to binary formats.

import array

var ints = array.Int32Array([1, 2, 3])
echo ints.length()
3

The System

os

Processes, the filesystem, paths and the environment. Covered in Chapter 9.

os.exec(command) runs a shell command and gives you its output. os.spawn(command, args, options) starts a process you can talk to. os.on_signal(name, handler) installs a signal handler. os.at_exit(handler) registers cleanup that runs however the program ends, and os.set_exit_code(code) decides the status it ends with without ending it there and then.

env

Configuration: a .env file read into the process environment, and values read back out already converted. Covered in Chapter 19.

import env

env.load()

var port = env.int('PORT', 8080)
var debug = env.bool('DEBUG', false)
var secret = env.require('SESSION_SECRET')

Names already set in the environment are left alone, so the file is a set of defaults and the deployment is what overrides them.

io

Standard streams, terminal control and in-memory files:

import io

var name = io.readline('Your name: ')
var secret = io.readline('Password: ', true)

echo io.stdout.is_tty()

io.TTY puts the terminal into raw mode, reads single keypresses and moves the cursor, which is what an interactive program needs. io.BytesIO is the in-memory file from Chapter 10.

io.capture(body) collects everything body writes to standard output instead of printing it, echo, print() and io.stdout alike:

import io

def greet(name) {
  echo 'Hello, ${name}!'
}

var out = io.capture(@{ greet('Ada') })

echo 'captured ${out.length()} characters'
captured 12 characters

Captures nest, and capture_begin()/capture_end() are the manual pair for when the body might raise and you want its output anyway. This is what lets Chapter 23 assert on what a function prints, and keep a passing test’s output out of the report.

test

Suites, matchers, mocks, snapshots and reports. Covered in Chapter 23.

import test { * }

describe('slug', @{
  it('lowercases and joins', @{
    expect(slug('Hello World')).to_be('hello-world')
  })
})

run()

stat

The S_IS* predicates over the mode word from file().stats(), plus a renderer for it:

import stat

file('notes.txt', 'w').write('x')
file('notes.txt').chmod(0c644)

var info = file('notes.txt').stats()

echo stat.S_ISREG(info.mode)
echo stat.S_ISDIR(info.mode)
echo stat.file_mode(info.mode)
true
false
-rw-r--r--

S_ISREG, S_ISDIR, S_ISLNK, S_ISCHR, S_ISBLK, S_ISFIFO and S_ISSOCK each answer one question about the kind of entry. S_IMODE(mode) strips the type bits and leaves the permissions; file_mode(mode) renders the whole thing the way ls -l does.

args

A command-line parser with subcommands, typed options, automatic --help and wrapped terminal output. Covered in Chapter 20.

import args

var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
parser.add_command('history', 'Show past greetings')

var parsed = parser.parse()

The help text is written from the same declarations that read the arguments, so the two cannot drift apart.

log

import log

log.info('server started')
log.error('connection refused')

A Logger binds structured fields, a child() logger inherits them, and transports send records to the console, a file or somewhere you write yourself.

isolate

Concurrency. Covered in Chapter 11.

ffi

Calling C and Rust. ffi loads shared libraries, reads their C headers or Rust sources as declarations, calls their functions with checked conversions, and turns Zuri functions into callbacks; it links static libraries into loadable ones too.

import ffi

var c = ffi.open(ffi.LIBC).declare('size_t strlen(const char *s);')

echo c.strlen('interop')
7

Covered in Chapter 26.

The Network

net

TCP, UDP, unix domain sockets, TLS, DTLS, addresses and polling. Covered in Chapter 12.

http

Client and server, HTTP/1.1 and HTTP/2, with routing, middleware, WebSockets, server-sent events, multipart uploads, static files and a reverse proxy. Introduced in Chapter 12, covered fully in Chapter 15, and used throughout Chapter 25.

rpc

JSON-RPC 2.0, the protocol for calling methods in another program, on both ends. A service answers calls with handlers, behind an http route or on a connection; a client calls a service over HTTP; and an endpoint on a WebSocket, a socket, a child process’s standard streams or a pipe between isolates both answers and calls, so either side can call the other at any time. rpc.serve() gives every connection to a listening socket an endpoint of its own, over TCP, TLS or a Unix domain socket.

import rpc
import isolate

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('add', @(params) => params[0] + params[1])
    .serve()
}, ends[1])

var client = rpc.endpoint(ends[0])

echo client.request('add', [2, 3])
client.close()
5

Covered in Chapter 28.

Databases

sql

One way to talk to a relational database, whichever one it is. sql defines what an adapter has to provide and supplies everything that is the same across engines: parameters, transactions and savepoints, cursors, pooling, introspection, and one error hierarchy. SQLite, PostgreSQL and MySQL adapters ship with it, and changing between them means changing the connection string.

import sql

var db = sql.open('sqlite://./app.db')
var id = db.insert('posts', { title: 'Hello' })

for post in db.query('select * from posts where id = ?', [id]) {
  echo post.title
}

Covered in Chapter 17.

Mail

mail

Messages and the three protocols that move them, written in Zuri from the socket up. mail.message() builds a message out of text, HTML and files and works out the MIME tree from what went in; mail.parse() reads one back. mail.smtp sends, mail.imap reads mail where it is kept, and mail.pop3 takes it away. Both server ends are here too: an SMTP server that decides what to accept through handlers of your own, and an IMAP server that answers out of a mail store, of which one keeps mail on disk in Maildir format and one keeps it in the process. mail.dkim signs outgoing mail and checks incoming mail.

import mail

mail.send('smtp://mail.example.com', mail.message({
  from: 'reports@example.com',
  to: 'ann@example.com',
  subject: 'Quarterly report',
  text: 'The numbers are in.',
}), { username: 'reports', password: secret })

Covered in Chapter 18.

Compression

compress

deflate, zlib, gzip, zstd, lz4, bzip2 and brotli, plus tar and zip archives and checksum for CRC32 and Adler-32. Covered in Chapter 10.

Graphics

imagine

Image creation and manipulation on an RGBA buffer: drawing primitives, text with a built-in stroke font, filters, colour-space conversion, and reading and writing the common formats.

import imagine { Image }

Image(400, 200, '#0f172a')
  .fill_circle(200, 100, 70, '#38bdf8')
  .circle(200, 100, 70, 'white', { thickness: 3 })
  .save('badge.png')

Almost every method returns an image, so operations chain. Image.open(path) decodes an existing file, with the format taken from the contents rather than the extension:

Image.open('photo.jpg')
  .thumbnail(400, 400)
  .save('thumb.webp')

Decoding and encoding are native; everything between them — filters, drawing, colour conversion — is ordinary Zuri you can read in libs/imagine and extend. Chapter 16 is the full treatment.

The Language Itself

zuri

Lexing, parsing, compiling and reflection, all reachable from Zuri code:

import zuri

echo zuri.tokenize('var x = 1').length() > 0
echo zuri.reflect.kind([1])
true
list

Chapter 21 is the full treatment.

math

The constants. Everything else is a method on number; see Chapter 4.

Finding the Rest

Every module’s source is in libs/, and every public function in it carries a doc block with its parameters, its defaults and its edge cases. Reading libs/set.zu is a faster way to learn set than any summary, and the standard library is written to be read.

Wire

Wire is Zuri’s built-in templating engine. Unlike most templating languages, Wire does not invent its own syntax on top of your markup. Every Wire feature is powered by attributes on ordinary HTML elements, which means anything a designer already knows about HTML transfers directly, and every valid HTML5 document is already a valid Wire template.

Under the hood, a Wire template is compiled once — not re-interpreted on every render — into a small instruction tree, using the exact same WHATWG-conformant parser that backs the html module. That has a consequence worth knowing up front: Wire understands your markup as structure, not as text. It knows that one interpolation sits inside a paragraph, another inside an href, and a third inside a <script> tag, and it escapes each one correctly for where it actually is. A value handed to a template can never turn into new markup by accident — that has to be requested explicitly, and Wire makes you say so out loud.

Introduction

Wire and Blade

If you have used Laravel’s Blade, PHP’s own answer to templating, a lot of Wire will feel familiar in spirit even though the syntax is different. Both engines compile templates rather than interpret them on every request, both let you extend a base layout and override named sections of it, and both escape everything by default so that printing a user’s name can never become a way for that user to run script in someone else’s browser.

Where Wire departs from Blade is in how it expresses control flow. Blade adds its own directive syntax on top of plain text (@if, @foreach, {{ }}) that a template author has to learn as a second language layered over HTML. Wire instead expresses everything as attributes on the HTML you were already going to write:

{{-- Blade --}}
@if ($user->isAdmin())
    <p>Welcome back, administrator.</p>
@endif
{{-- Wire --}}
<p x-if="user.is_admin">Welcome back, administrator.</p>

The practical benefit is that a Wire template can be handed to a designer who has never seen Zuri and they can still read it: it is HTML, with some attributes they can look up. It also means every Wire template validates as HTML5, can be opened directly in a browser to check its structure, and can be run through the html module’s own tools (html.format(), a linter, a selector query) without anything special-casing Wire’s own syntax.

Your First Template

Here is the smallest possible Wire template, rendered from a string:

import wire

echo wire.render_string('<p>Hello {{ name }}</p>', { name: 'Ada' })
# <p>Hello Ada</p>

{{ name }} is an interpolation: it evaluates the expression name against the variables you supplied and writes the result into the page, escaped for wherever it landed. Everything else in this guide is built out of that one idea, plus a handful of x- prefixed attributes that control whether and how many times an element is rendered.

Rendering Templates

Templates From Files

For anything beyond a one-off snippet, templates live in files. Build a Wire instance, point it at a directory, and render by path:

import wire

var view = wire.wire()
view.set_root('./views')

echo view.render('pages/home', { user, posts })

render()’s first argument is a path relative to the root. The .html extension is added automatically when the path as written names no file, so 'pages/home' finds pages/home.html; a path that already carries an extension ('pages/home.wire') is tried exactly as given first. set_extension() changes what gets tried when none is given.

Note A Wire instance is meant to be built once, configured, and reused for the life of your program — typically at startup, alongside however you already configure the rest of your application. Rendering does not mutate it, so it is safe to render from several places at once.

Templates From Strings

render_string() renders a template given directly as a string, without touching the filesystem. It behaves identically to render() in every other respect — the same directives, the same escaping, the same filters — and any x-include or x-extend inside the string still resolves against the configured root.

echo view.render_string('<p>{{ greeting }}</p>', { greeting: 'Hi there' })

The third argument names the source for error messages (it defaults to <source>), which is worth passing when the string came from somewhere with its own identity — a database row, a file you already had open for another reason:

view.render_string(row.body, { user }, 'cms:page:${row.id}')

render_string() is not cached, since there is no file path to key a cache entry on. Reach for render() for anything rendered more than once.

The Template Root

Every path — the one passed to render(), and the one written in every x-include and x-extend in every template — resolves inside the configured root directory, and nowhere else. This is not a convention; it is enforced by the loader on every single resolution, and it matters for security, not just organization.

view.set_root('./views')
view.root()
# '/home/you/project/views' — always the absolute path

The root does not have to exist yet. create_root() makes it, and reports whether it had to:

if view.create_root() {
  echo 'Created a fresh views/ directory.'
}

Wire never creates this directory on its own initiative — a typo in a root path should read as “template not found,” not silently produce an empty folder somewhere unexpected.

Displaying Data

You have already seen the basic form. {{ expression }} evaluates whatever is between the braces and writes the result into the page:

<h1>{{ post.title }}</h1>
<p>By {{ post.author.name }}</p>

expression is not limited to a bare variable name — it is a full expression, covered in its own section below — so all of the following are valid:

<p>{{ post.views > 1000 ? 'Popular' : 'New' }}</p>
<p>{{ post.tags|join(', ') }}</p>
<p>{{ user.nickname ?? user.name }}</p>

Escaping Data

Every interpolation is escaped for the specific place it lands, and this cannot be turned off from inside a template. If name holds <b>Ada</b>:

<p>Hello {{ name }}</p>

renders as:

<p>Hello &lt;b&gt;Ada&lt;/b&gt;</p>

which is exactly what you want when name came from a form field, a database column, or anywhere else outside your own control. Wire is not guessing at this — because it compiles a real parsed document rather than gluing strings together, it always knows precisely which kind of place an interpolation sits in, and it applies the escaping that place needs:

Where the interpolation isWhat happens
Text between tags&, <, > become entities
An ordinary attribute (title, alt, data-*, …)& and " become entities
href, src, action, and other URL attributesthe value’s URL scheme is checked too
<script>, or an on* event handler attributethe value is encoded as JSON
<style>anything outside a CSS-safe set is dropped

That is a materially stronger guarantee than “the special characters got replaced”: a value dropped into an href cannot smuggle in a javascript: scheme, and a value dropped into a <script> block cannot break out of its string literal, close the tag early, or open a comment — even though none of those things are <, >, or &.

Rendering Raw Markup

Escaping everything by default is right, but sometimes you genuinely have markup — the output of another render, a snippet you built and trust — and you want it written out as markup. The raw filter is how you say so:

<div class="article-body">{{ post.rendered_html|raw }}</div>

You can build the same value on the Zuri side and hand it to a template already marked as safe, with wire.safe():

import wire

view.render_string('<div>{{ body }}</div>', {
  body: wire.safe('<em>already-trusted markup</em>'),
})

Warning raw and wire.safe() are a promise that the value is safe to write unescaped at the place it lands. Applying either to anything a user submitted — a comment, a bio, a search query — reopens exactly the cross-site scripting hole the rest of Wire exists to close. Only reach for them on markup your own code produced.

If you need to render markup as text on purpose — showing someone a literal <script> tag in a code sample, say — that is what plain interpolation already does; there is nothing extra to opt into.

Wire and JavaScript Frameworks

Wire uses {{ }} for interpolation, the same delimiter many JavaScript templating libraries (Vue, Angular, Handlebars, Mustache) use for their own. If a Wire template also contains inline script that a browser-side framework is meant to interpret, escape the braces with a leading % so Wire leaves them alone:

<div id="app">
  <p>{{ user.name }}</p>          {{-- rendered by Wire, server-side --}}
  <p>%{{ message }}</p>           {{-- left as literal {{ message }}, for Vue --}}
</div>

%{{ renders as the literal text {{, and %{! does the same for template function calls. If you find yourself escaping braces constantly because most of a template belongs to a client-side framework, it may be worth keeping that section in its own file and serving it untouched rather than through Wire at all.

The Expression Language

Everything between {{ and }} — and everything given to x-if, x-for, x-attr, and the rest of the directives below — is written in Wire’s own small expression language. It is deliberately not a full embedded copy of Zuri: there is no assignment, no way to declare anything, and no way to reach a global variable. A template describes a page; the logic that decides what the page contains belongs in the code that calls render(), not in the template itself.

Literals

{{ 42 }}            {{-- a number --}}
{{ 3.14 }}           {{-- a decimal --}}
{{ 'a string' }}     {{-- single quotes --}}
{{ "a string" }}     {{-- or double, interchangeably --}}
{{ true }}           {{-- true / false / nil --}}
{{ [1, 2, 3] }}      {{-- a list literal --}}
{{ { id: 1, name } }} {{-- a dict literal; { name } is short for { name: name } --}}
{{ 0..pages }}       {{-- a range, exactly as Zuri writes one --}}

Looking Up Values

{{ user.name }}                 {{-- a dotted lookup --}}
{{ user.address.city }}         {{-- chains as deep as you like --}}
{{ items.0 }}                   {{-- a numeric key reads a list position --}}
{{ items[index] }}              {{-- a computed index --}}
{{ items[-1] }}                 {{-- negative counts from the end --}}
{{ items.length }}               {{-- a collection answers `length` by name --}}

Reading a variable that was never supplied gives nil rather than raising, and reading a key off nil gives nil too. That is what makes an optional value safe to reach through without a guard in front of it:

{{ user.profile.avatar_url }}

renders as nothing at all if user has no profile, instead of failing three levels down. Reach for x-if when the difference between “empty” and “genuinely missing” matters to what you show.

Note A name starting with an underscore can never be read from a template, the same way Zuri treats _field as private. If you find yourself wanting to read one, expose a public accessor from Zuri instead.

Operators

{{ price * quantity }}
{{ subtotal + tax }}
{{ stock - reserved }}
{{ total / count }}
{{ index % 2 }}

{{ a == b }}   {{ a != b }}
{{ a < b }}    {{ a <= b }}
{{ a > b }}    {{ a >= b }}

{{ 'admin' in user.roles }}
{{ 'admin' not in user.roles }}

{{ is_admin and is_active }}
{{ is_admin && is_active }}      {{-- && is the same as `and` --}}
{{ is_guest or is_banned }}
{{ is_guest || is_banned }}       {{-- || is the same as `or` --}}
{{ !is_active }}
{{ not is_active }}               {{-- ! is the same as `not` --}}

{{ stock > 0 ? 'In stock' : 'Sold out' }}
{{ nickname ?? name }}

A string on either side of + concatenates rather than raising:

{{ 'Hello, ' + user.name }}

?? and or look similar but answer different questions, and mixing them up is the single most common Wire mistake:

  • ?? asks “was this ever supplied?” and only falls back on nil.
  • or asks “is this worth showing?” using Wire’s own truthiness, and falls back on anything falsy — nil, false, an empty string, an empty collection, or the number 0.
{{ discount ?? 0 }}   {{-- a missing discount becomes 0 --}}
{{ discount or 0 }}   {{-- a discount that IS 0 also becomes 0, harmlessly here --}}

{{ stock_count ?? 'unknown' }}  {{-- 0 in stock still shows as 0 --}}
{{ stock_count or 'unknown' }}  {{-- 0 in stock is treated the same as never supplied --}}

Use ?? whenever zero is a legitimate value you want to keep, and or whenever you only want to show something when there is genuinely something to show.

Truthiness

x-if, x-not, and/or, !/not, and ? : all use Wire’s own notion of truthy and falsy, which differs from Zuri’s own rules in one place, deliberately, for template authoring:

ValueWire
nil, falsefalsy
0, NaNfalsy
any other number, including negative onestruthy
an empty stringfalsy
a non-empty stringtruthy
an empty list or dictfalsy
a non-empty list or dicttruthy
anything elsetruthy

The difference is the empty collection. Zuri treats [] and {} as truthy; Wire treats them as falsy, so x-if="results" correctly hides a section for a search that came back with nothing.

Calling Functions

A value registered as a global can be called directly:

<a href="{{ route('user.profile', user.id) }}">{{ user.name }}</a>

There is also an older, standalone spelling for calling a function that takes no arguments, kept from Wire’s very first version because it reads well on its own:

{! current_year !}

is exactly the same as writing:

{{ current_year() }}

Prefer {{ fn() }} for anything new; {! !} exists for templates that already use it and for the handful of cases — a page’s build stamp, a feature flag — where a function taking nothing at all reads a little cleaner without the parentheses.

Filters

A filter transforms the value on the left of a |. It is the same idea as a Unix pipe:

{{ name|upper }}
{{ price|round(2) }}
{{ post.body|truncate(150) }}

The value being filtered is always the filter’s first argument; anything written in parentheses follows it. truncate(150) above calls the truncate filter as truncate(post.body, 150).

Chaining Filters

Filters read left to right, each one’s output feeding the next:

{{ name|trim|title }}
{{ comment.body|strip_tags|truncate(200) }}

The = Argument Shorthand

For a filter that takes exactly one argument, name=value is shorthand for name(value):

{{ status|is='active' }}

is the same as:

{{ status|is('active') }}

This spelling exists for symmetry with Wire’s very first version and reads naturally for a short comparison; the parenthesised form is equally valid everywhere and is the only option once a filter needs more than one argument.

Available Filters

Escaping

FilterWhat it does
rawMarks the value as markup, skipping escaping entirely.
escape / eEscapes for a context other than the one the value is being written into. Takes 'text' (the default), 'attribute', 'url', 'script', or 'style'.

Text

FilterWhat it does
upperConverts to upper case.
lowerConverts to lower case.
titleTitle Cases Every Word.
capitalizeCapitalizes only the first letter, leaving the rest alone.
trimRemoves leading and trailing whitespace of every kind (space, tab, newline).
truncate(length, suffix?)Cuts to length characters, appending suffix (default '…') only if anything was actually cut.
replace(search, replacement?)Replaces every literal occurrence of search. Never a regular expression. replacement defaults to an empty string.
lpad(width, fill?)Pads on the left to width characters, with fill (default a space).
rpad(width, fill?)Pads on the right.
repeat(count)Repeats the value count times.
nl2brTurns line breaks into <br>, escaping the text first. Returns markup.
line_breaksAn alias for nl2br.
strip_tagsRemoves every HTML tag, keeping only the text — a real parse, not a pattern match.
slugLower-cased, hyphen-separated, safe for a URL segment.
url_encodePercent-encodes for use inside a URL.
jsonEncodes as JSON text.
json_script(id?)Wraps the value’s JSON encoding in a <script type="application/json">, optionally with an id, ready to be read back by a script on the page. Returns markup.

Numbers

FilterWhat it does
absRemoves the sign.
round(places?)Rounds to places decimal places (default 0), halves rounding away from zero.
floorRounds down to the nearest whole number.
ceilRounds up.
number_format(places?, point?, separator?)Groups thousands and fixes the decimal places, like 1,234,567.89. Pass point/separator to use another convention, e.g. 1.234.567,89.
filesize(binary?)A byte count written the way a person reads it — 1.4 MB by default, or 1.3 MiB with binary set to true.

Collections

FilterWhat it does
lengthHow many entries — works on a string, list, dict, or bytes. nil has length 0.
firstThe first entry, or nil if there is none.
lastThe last entry.
join(glue?)Joins entries into one string with glue (default '') between them.
sort(key?)Sorted ascending; key sorts a list of dicts or instances by one field.
reverseThe entries backwards, or a string reversed.
uniqueDuplicates removed, keeping the first of each.
keysA dict’s keys, in insertion order.
valuesA dict’s values, in insertion order.
slice(start, end?)The entries from start up to but not including end. Negative positions count from the end.
sum(key?)Adds the entries together; key sums one field of a list of dicts or instances.
split(separator?)Splits a string into a list. separator defaults to any run of whitespace.

Choice

FilterWhat it does
default(fallback) / altfallback when the value is falsy, otherwise the value.
emptyWhether the value has nothing in it. Unlike falsiness, a number is never empty — not even 0.
is(expected)Whether the value equals expected.
not(expected)Whether it differs.

Dates

FilterWhat it does
date(format?)Formats a date.Date, a Unix timestamp, or a parseable date string, using the same format directives as Date.format(). Defaults to 'Y-m-d H:i:s'.
<time datetime="{{ post.published_at|date('Y-m-d') }}">
  {{ post.published_at|date('jS F Y') }}
</time>

Writing Your Own Filter

See Custom Filters below.

Conditionals

x-if, x-elif, and x-else

x-if renders an element, and everything inside it, only when its expression is truthy:

<p x-if="user.is_admin">You have administrator access.</p>

If user.is_admin is falsy, the whole <p> — tag and contents — is left out of the page entirely. There is no empty element left behind.

Chain further conditions with x-elif, and close the chain with a plain x-else:

<p x-if="user.role == 'admin'">Administrator</p>
<p x-elif="user.role == 'staff'">Staff member</p>
<p x-elif="user.role == 'contributor'">Contributor</p>
<p x-else>Member</p>

Exactly one of these renders. Only whitespace and HTML comments are allowed between the elements in a chain — any real content in between ends it, and an x-elif or x-else with no x-if in front of it is rejected when the template compiles, not silently ignored:

{{-- this chain is broken by the text between the two elements --}}
<p x-if="a">A</p>
some text
<p x-else>not A</p> {{-- error: x-else has no x-if in front of it --}}

x-not

x-not is the plain inverse of x-if — it renders when its expression is falsy — and does not take part in a chain:

<div x-not="user.has_verified_email">
  <p>Please verify your email address.</p>
</div>

x-if="!condition" and x-not="condition" mean the same thing; x-not exists because it often reads more naturally for a guard clause.

Loops

x-for repeats an element once per entry of whatever its expression evaluates to — a list, a dict, a string, or a range:

<ul>
  <li x-for="posts" x-value="post">{{ post.title }}</li>
</ul>

Notice that the element itself repeats, not just its contents — the example above produces one whole <li> per post, not one <li> wrapping every post. x-value names the variable each iteration binds its current entry to; leave it out entirely if you do not need to refer to the entry by name.

An optional x-key binds the position (for a list) or the key (for a dict):

<tr x-for="users" x-key="id" x-value="user">
  <td>{{ id }}</td>
  <td>{{ user.name }}</td>
</tr>

Every kind of collection iterates naturally:

<li x-for="tags" x-value="tag">{{ tag }}</li>              {{-- a list --}}
<li x-for="scores" x-key="name" x-value="score">           {{-- a dict --}}
  {{ name }}: {{ score }}
</li>
<span x-for="word" x-value="letter">{{ letter }}</span>    {{-- a string, by character --}}
<option x-for="1..5" x-value="n">{{ n }}</option>          {{-- a range --}}

A missing or empty collection simply renders nothing — there is no need to guard a loop with an x-if first:

<li x-for="comments" x-value="comment">{{ comment.body }}</li>
{{-- renders nothing at all if `comments` is empty or was never supplied --}}

The loop Variable

Every iteration publishes a loop variable with the following fields:

FieldValue
loop.indexThe position, counting from 1.
loop.index0The position, counting from 0.
loop.firsttrue on the first pass.
loop.lasttrue on the last pass.
loop.lengthHow many entries there are in total.
loop.eventrue on the 2nd, 4th, 6th, … pass — even-numbered by loop.index.
loop.oddtrue on the 1st, 3rd, 5th, … pass.
loop.keyThe current key or index, whether or not x-key binds it too.
loop.valueThe current value, whether or not x-value binds it too.
loop.parentThe enclosing loop’s own loop, for a nested x-for.
<tr x-for="rows" x-value="row" x-attr="{ 'class': loop.odd ? 'zebra' : nil }">
  <td>{{ loop.index }}</td>
  <td>{{ row.name }}</td>
</tr>

Nesting Loops

x-loop renames the metadata variable a specific x-for publishes, which is what lets an inner loop’s own loop and an outer loop’s loop both be reached at once:

<table x-for="rows" x-loop="row" x-value="cells">
  <tr>
    <td x-for="cells" x-value="cell">
      {{ row.index }}, {{ loop.index }}: {{ cell }}
    </td>
  </tr>
</table>

Without x-loop, the inner loop’s own loop would simply shadow the outer one for the scope of the inner loop — reach for loop.parent instead if you would rather not rename anything:

<td x-for="cells" x-value="cell">
  outer position {{ loop.parent.index }}, inner position {{ loop.index }}
</td>

Looping Without a Wrapper Element

Sometimes you want to repeat several elements together without a real element wrapping them, or without introducing any element at all. Put the directive on a <template> instead of on the element you want repeated:

<select>
  <template x-for="countries" x-value="country">
    <option value="{{ country.code }}">{{ country.name }}</option>
  </template>
</select>

<template> is the one HTML element the parser lets through completely untouched, wherever it appears — inside a <head>, inside a <table>, inside a <select> — and Wire removes it from the output entirely, leaving only what was inside it, once per pass. This is also the way to apply x-if to a group of elements without picking one of them to carry the attribute:

<template x-if="user.is_admin">
  <a href="/admin">Dashboard</a>
  <a href="/admin/users">Users</a>
</template>

Content and Attributes

x-text

x-text replaces an element’s children with its expression, escaped as plain text — useful when the element already has other attributes and you would rather not write the value twice:

<p x-text="post.summary"></p>

is the same as:

<p>{{ post.summary }}</p>

Anything written inside the element in the source is discarded; it exists only to describe what would go there without JavaScript.

x-html

x-html is x-text’s unescaped counterpart — it replaces the element’s children with its expression’s value, written as markup rather than text:

<div x-html="post.rendered_body|raw"></div>

Note the |raw: x-html still expects a value marked safe, the same as an ordinary interpolation would. x-html only changes where the markup goes (replacing the whole element’s contents, rather than sitting inline in a text run) — it does not, on its own, turn escaping off. The same warning about untrusted input applies here just as much as it does to the raw filter.

x-attr

x-attr spreads a dictionary of names and values onto the element as attributes:

<input x-attr="{ type: 'text', name: field.name, required: field.is_required }">

Within that dictionary:

  • true produces a valueless attribute (required), exactly as writing required by hand would.
  • false and nil leave the attribute off entirely, rather than writing it with an empty or literal "false" value.
  • Anything else is written as the attribute’s value, escaped for whatever that attribute means (a URL attribute is checked as a URL, same as always).

This is what turns a boolean into a real HTML boolean attribute without a ternary in every place one is needed:

<button x-attr="{ disabled: !form.is_valid }">Submit</button>

A computed attribute takes over from one written directly on the same element, so you can set a sensible default and only override it when there is something to override:

<div class="card" x-attr="{ 'class': featured ? 'card card-featured' : nil }">

Comments

HTML’s own comment syntax is a server-side note in a Wire template. It never reaches the rendered page, and nothing inside one is evaluated — which is exactly what makes it safe to leave notes for other people maintaining the template, without those notes shipping to a browser:

<!-- TODO: replace this hard-coded banner once marketing sends the real copy -->
<div class="banner">Coming soon</div>

<!-- this variable is not evaluated: {{ some.internal.detail }} -->

If you genuinely want a comment to reach the browser — a conditional comment, a build stamp a deploy script checks for — turn comments back on for that Wire instance with set_comments(true).

Including Templates

<include path="..." /> renders another template in its place. It also has a directive spelling, <template x-include="...">, which means exactly the same thing — the pseudo-element forms in this guide (<include>, <extend>, <declare>, <define>) are sugar, rewritten into the directive form before the template is even parsed, which is what lets them work correctly inside a <head> or a <table> where an element the parser does not recognise would otherwise be moved or dropped.

<include path="partials/nav.html" />

is the same as:

<template x-include="partials/nav.html"></template>

The path is resolved inside the template root, the same as render()’s own path argument.

By default, an include sees every variable the page around it can see — it is not a separate scope:

{{-- page.html --}}
<include path="partials/greeting.html" />
{{-- partials/greeting.html --}}
<p>Hi, {{ user.name }}!</p>

renders correctly without user ever being passed explicitly to the partial.

Passing Data to an Include

x-with adds variables for the included template, alongside whatever it already inherits:

<include path="partials/badge.html" x-with="{ label: 'New', tone: 'green' }" />

only withholds the surrounding scope entirely, leaving the partial with only what x-with gave it — turning a plain partial into something closer to a real component with a defined interface:

<include path="components/price-tag.html" x-with="{ amount: item.price }" only />

Includes as Components

Anything written inside an <include> tag is handed to the included template as a named region called content, declared with <declare name="content">:

{{-- components/card.html --}}
<section class="card">
  <h3>{{ title }}</h3>
  <declare name="content"></declare>
</section>
{{-- the page --}}
<include path="components/card.html" x-with="{ title: 'Recent Orders' }">
  <p>{{ orders|length }} orders this week</p>
</include>

renders as:

<section class="card">
  <h3>Recent Orders</h3>
  <p>3 orders this week</p>
</section>

The content you write inside the <include> tag keeps the scope it was written in — the page’s, not the partial’s — which is exactly what lets {{ orders|length }} above read a variable from the page even though x-only was not used and the card component never mentions orders at all.

Computed Include Paths

A path is read as literal text, with {{ }} interpolation allowed inside it — not as an expression in its own right, so x-include="header" names a file called header, rather than reading a variable of that name:

<include path="themes/{{ current_theme }}/header.html" />

resolves a different file per render depending on current_theme, while everything after path="themes/" up to the next {{ is taken literally.

Template Inheritance

Includes are for small, reusable fragments — a navbar, a footer, a badge. For whole-page structure — the parts of a layout every page on your site shares — template inheritance is the better fit. Where an include composes fragments together, inheritance lets a base layout define the shape of a page once, and lets each page fill in only the parts that differ.

Defining a Layout

A base template marks the regions a page is allowed to fill in with <declare name="...">:

{{-- layouts/app.html --}}
<!DOCTYPE html>
<html>
  <head>
    <title>{{ title }}</title>
    <declare name="head"></declare>
  </head>
  <body>
    <nav><!-- shared navigation --></nav>
    <main>
      <declare name="content"></declare>
    </main>
    <footer>
      <declare name="footer">
        <p>&copy; {{ year }} My Company</p>
      </declare>
    </footer>
  </body>
</html>

Notice that footer has content already inside its <declare> tag. That becomes the default — what renders when a page does not define that region at all.

Extending a Layout

A page declares which layout it extends with <extend base="...">, and fills in regions with <define name="...">:

{{-- pages/dashboard.html --}}
<extend base="layouts/app.html">
  <define name="head">
    <meta name="description" content="Your account dashboard.">
  </define>
  <define name="content">
    <h1>Welcome back, {{ user.name }}</h1>
    <p>You have {{ notifications|length }} new notifications.</p>
  </define>
</extend>

Rendering pages/dashboard.html produces the layout’s full structure, with head and content replaced by what the page defined, and footer left exactly as the layout’s own default. Both the layout and the page are rendered against the same variables render() was given — there is no separate scope to pass anything through.

Everything inside an <extend> must be inside a <define>. The base template owns the page’s structure; a page’s job is only to fill in the regions it was offered, not to add markup of its own outside them. A template can extend exactly one base.

Default Slot Content

A region a page does not define keeps whatever the layout put inside its own <declare> tag:

{{-- pages/minimal.html --}}
<extend base="layouts/app.html">
  <define name="content">
    <p>Just this page's content — head and footer both use the layout's defaults.</p>
  </define>
</extend>

A <declare> tag with nothing inside it — like head in the layout above — simply renders empty when nothing defines it.

x-super

<super /> — or its directive spelling, <template x-super></template> — renders whatever the region it sits inside would have shown before this definition replaced it. That lets a page add to a section instead of fully restating it:

<extend base="layouts/app.html">
  <define name="footer">
    <super />
    <p><a href="/privacy">Privacy Policy</a></p>
  </define>
</extend>

renders the layout’s own copyright line, followed by the extra privacy link — without the page having to know or repeat what the layout’s default footer actually said.

Multi-Level Inheritance

A template that extends a base can itself declare regions of its own, letting a further template extend it. This is how a site with, say, a general layout and several page-type-specific layouts (a blog post, a product page) is usually structured:

{{-- layouts/app.html --}}
<html><body>
  <declare name="body"><p>default</p></declare>
</body></html>
{{-- layouts/article.html --}}
<extend base="layouts/app.html">
  <define name="body">
    <article>
      <declare name="article-content"></declare>
    </article>
  </define>
</extend>
{{-- pages/post.html --}}
<extend base="layouts/article.html">
  <define name="article-content">
    <h1>{{ post.title }}</h1>
    {{ post.body|raw }}
  </define>
</extend>

Rendering pages/post.html walks the whole chain: it extends layouts/article.html, which extends layouts/app.html, and the final page is assembled from all three.

Overriding a Definition

Defining the same region twice in one template is almost always an accident — two people editing the same file, a copy-paste that was never cleaned up — so Wire refuses it at compile time unless you say the replacement is deliberate with override:

<extend base="layouts/app.html">
  <define name="content"><p>First draft</p></define>
  <define name="content" override><p>Final version</p></define>
</extend>

Without override on the second one, this template fails to compile with a clear message naming the region and both locations, rather than silently keeping whichever definition happened to come last.

Security

Why Everything Is Escaped by Default

A templating engine’s job includes keeping a value that came from outside your program — a form submission, a query string, another user’s profile — from becoming markup, a script, or a link, unless you explicitly say it is safe to. Wire treats this as the default rather than an opt-in setting, for the same reason a seatbelt only works if putting it on is the ordinary thing to do rather than the thing you remember to do under pressure: the moment escaping is something you have to remember to turn on, the templates that skip it are the ones that get exploited.

Every interpolation is escaped for the exact place it lands — see the table earlier in this guide — and there is no template-level setting to disable that. The only way past it is the explicit, visible raw filter or x-html directive, which is exactly the friction you want between “displaying a value” and “trusting a value with unescaped markup.”

URLs Are Checked, Not Just Escaped

An attribute a browser reads as a URL — href, src, action, and several others — gets more than entity escaping. Its scheme is checked against an allowlist, and a disallowed scheme is replaced with about:blank rather than written through:

<a href="{{ profile_link }}">{{ user.name }}</a>

If profile_link were javascript:alert(document.cookie), the rendered href is about:blank, not the script URL. A URL with no scheme at all — every relative link, every absolute path, every protocol-relative URL — is always allowed, since none of those can name a handler in the first place.

The default allowlist covers http, https, mailto, tel, and a handful of others; it deliberately leaves out javascript, vbscript, data, and file. set_url_schemes() replaces the list for an application that genuinely needs another scheme — a custom app: handler, say.

srcset and ping, which each hold a list of URLs, have every entry in the list checked the same way, with descriptors (2x, 640w) and separators preserved exactly as written.

Scripts and Stylesheets

A value interpolated inside a <script> element, or inside an event handler attribute like onclick, is encoded as JSON rather than merely escaped — automatically, with no filter needed:

<script>
  var currentUser = {{ user }};
</script>

renders the whole user value as a JSON literal, with its own quotes included. That is worth internalising, because it changes how you write the surrounding script: do not wrap the interpolation in your own quotes.

{{-- correct: renders   var name = "Ada";   --}}
<script>var name = {{ user.name }};</script>

{{-- wrong: renders   var name = '"Ada"';   which is not what you meant --}}
<script>var name = '{{ user.name }}';</script>

A value inside a <style> block has anything outside a conservative, CSS-safe character set dropped rather than escaped — there is no escape sequence that would stop a < from being read by the HTML tokenizer looking for </style>, so the only sound answer is for the character not to be there at all.

The Template Root Is a Sandbox

Every path — whether it came from render()’s own argument, or from an x-include/x-extend written inside a template, even one built from an interpolated value — is resolved inside the configured root and refused if it resolves anywhere else, .. segments and symbolic links both included:

<include path="{{ theme }}/header.html" />

If theme ever came from something a request controls, an unbounded loader would turn this into a way to read arbitrary files the process can reach. Wire’s loader treats the root as a hard boundary instead: the worst a hostile value can do here is fail to find a template.

Extending Wire

Custom Filters

register_filter() adds a filter, or replaces one that already exists by that name. The value being filtered is always the first argument; anything the template passes after it follows:

view.register_filter('excerpt', @(value, words) {
  var count = words ?? 25
  return ' '.join(value.split('/\\s+/').take(count)) + '…'
})
<p>{{ post.body|excerpt(40) }}</p>

An argument the template leaves out arrives as nil, which is why words ?? 25 above supplies a sensible default rather than the filter insisting the template spell it out every time.

A filter returning a plain value has that value escaped like any other — the same as an ordinary interpolation. A filter that genuinely produces markup returns wire.safe(), and from that point on is responsible for what is inside it, exactly like raw and x-html are.

Globals

register_global() makes a value readable from every template without it being passed to render() explicitly:

view.register_global('site_name', 'Example Inc.')
view.register_global('route', @(name) {
  return '/' + name
})
<title>{{ page_title }} — {{ site_name }}</title>
<a href="{{ route('user.profile', user.id) }}">{{ user.name }}</a>

A variable of the same name passed to render() wins over a global — a global is a default, not a hard override, so a page can still shadow site_name for itself if it ever needs to.

Custom Elements

For the rare case the directives genuinely cannot express, register_element() claims an HTML tag name outright and hands every element of that name to a Zuri function instead of writing it out:

view.register_element('icon', @(view, element) {
  var name = element.attributes.get('name', 'dot')
  return wire.safe('<svg class="icon"><use href="#${name}"></use></svg>')
})
<icon name="star" />

The function receives the Wire instance and a dictionary describing the element — its tag name, its attributes (already rendered and escaped), its already-rendered children, and the scope it was written in. Return wire.safe(markup) to write markup, any other value to write it as escaped text, or nil to write nothing at all. Directives still work on a claimed element exactly as they read (<icon x-if="..." /> behaves the way it looks), and the whole mechanism exists for genuinely dynamic, code-driven markup that a partial and x-include cannot express — reach for an include first; it is easier for the next person to read and does not require them to go find the Zuri function behind it.

Configuration

Every setting can be passed to the Wire constructor at once, or set individually with its own method afterward:

var view = wire.wire({
  root: './views',
  extension: '.html',
  compact: true,
  comments: false,
  auto_reload: true,
  url_schemes: ['https', 'mailto'],
})
Option / methodDefaultWhat it controls
root / set_root(path)./templatesThe directory every template path resolves inside.
extension / set_extension(ext).htmlTried when a path names no file as written. Must start with ..
compact / set_compact(bool)falseDrops whitespace-only text between tags.
comments / set_comments(bool)falseWhether HTML comments reach the rendered output.
auto_reload / set_auto_reload(bool)trueWhether a cached template is checked against its file, and re-read if it changed, before each render.
url_schemes / set_url_schemes(list)see aboveWhich URL schemes an href/src/etc. is allowed to use.

root() reports the currently configured root as an absolute path, and create_root() makes the directory if it does not exist yet (returning whether it had to).

Compiling and Caching

render() compiles a template into its instruction tree the first time it is used, and keeps the compiled result for next time — parsing and directive resolution happen once per template, not once per request. Every subsequent render just walks that tree.

With auto_reload on (the default), each render checks the file’s modification time and size against what was cached, and recompiles if either changed. That is one filesystem stat per template per render — cheap, and exactly what you want while actively editing templates. Turn it off with set_auto_reload(false) once you are running under load and are not editing templates live; clear_cache() is then how a long-running process picks up a new deployment.

compile(path) compiles a template without rendering it, and returns the compiled result — or raises, if the template has a mistake in it. That makes it a good fit for a startup-time check across a whole directory of templates, so a broken one is caught before the first request that would have hit it:

for name in os.read_dir('./views', true) {
  view.compile(name)
}

Error Handling

Everything Wire raises is a wire.WireError, and every one of them carries the template path and the line/column of the tag responsible, so catching the base class is enough to handle any template failure the same way — turning it into a 500 page, logging it with context, whatever your application needs:

catch {
  echo view.render('pages/dashboard', { user })
} as e {
  if instance_of(e, wire.WireError) {
    log.error('template failed: ${e.reason} at ${e.location()}')
    echo view.render('errors/500')
  }
}

Three more specific errors, each a WireError, tell you what kind of thing went wrong:

ErrorRaised when
wire.TemplateSyntaxErrorA template does not compile: an unknown directive, a malformed expression, a broken inheritance chain. Always caught the first time a template is used, not on a later render.
wire.TemplateNotFoundErrorA path names no file, or resolves outside the template root.
wire.RenderErrorSomething that depends on the values a template was given: iterating something that cannot be iterated, a filter rejecting its input, an include chain that never bottoms out.

A variable that was simply never supplied is not one of these — that renders as empty and tests as falsy, by design, so an optional section can be written without a guard around every single field it touches.

Full Directive Reference

Every directive Wire understands. An x- prefixed attribute outside this list fails to compile — Wire owns that whole prefix, so a typo like x-fi is caught the moment the template is compiled rather than silently doing nothing.

DirectiveTakesMeaning
x-ifexpressionRenders only if truthy. Starts a chain.
x-elifexpressionContinues an x-if chain.
x-else(none)Closes an x-if chain.
x-notexpressionRenders only if falsy. Does not chain.
x-forexpressionRepeats the element once per entry.
x-valuenameBinds each iteration’s value.
x-keynameBinds each iteration’s key/index.
x-loopnameRenames the loop metadata variable.
x-textexpressionReplaces children with escaped text.
x-htmlexpressionReplaces children with unescaped markup.
x-attrexpression (dict)Spreads attributes onto the element.
x-includepathRenders another template in this element’s place.
x-withexpression (dict)Adds variables for an include.
x-only(none)Withholds the surrounding scope from an include.
x-extendpathThis template extends the named base.
x-slotnameDeclares a region an extending template may replace.
x-definenameReplaces a base template’s region.
x-override(none)Permits redefining a region already defined once.
x-super(none)Renders the definition this one replaces.

Five pseudo-elements are sugar for the directives above, rewritten before the template is parsed:

Pseudo-elementEquivalent to
<include path="p" /><template x-include="p"></template>
<extend base="p">…</extend><template x-extend="p">…</template>
<declare name="n">…</declare><template x-slot="n">…</template>
<define name="n">…</define><template x-define="n">…</template>
<super /><template x-super></template>

Cheat Sheet

{{-- Variables --}}
{{ name }}
{{ user.address.city }}
{{ items.0 }}
{{ items[index] }}

{{-- Filters --}}
{{ name|upper }}
{{ price|round(2) }}
{{ post.body|truncate(150) }}
{{ status|is='active' }}

{{-- Raw markup --}}
{{ trusted_html|raw }}

{{-- Conditionals --}}
<p x-if="condition">…</p>
<p x-elif="other">…</p>
<p x-else>…</p>
<p x-not="condition">…</p>

{{-- Loops --}}
<li x-for="items" x-value="item" x-key="i">{{ i }}: {{ item }}</li>
{{ loop.index }} {{ loop.first }} {{ loop.last }} {{ loop.length }}

{{-- Wrapper-free groups --}}
<template x-for="items" x-value="item">…</template>
<template x-if="condition">…</template>

{{-- Content and attributes --}}
<p x-text="value"></p>
<div x-html="trusted_html|raw"></div>
<input x-attr="{ type: 'text', disabled: !enabled }">

{{-- Comments (never rendered) --}}
<!-- a note for the next person -->

{{-- Includes --}}
<include path="partials/nav.html" />
<include path="components/card.html" x-with="{ title: 'X' }" only>
  content for the component's declared region
</include>

{{-- Inheritance --}}
<extend base="layouts/app.html">
  <define name="content">…</define>
  <define name="footer"><super />…</define>
</extend>

Further reading: the html module that Wire’s parser and serializer are built on, and the date module for the format directives the date filter accepts.

HTTP

The http module is Zuri’s HTTP stack: a client, a server, and the pieces both are built from. It speaks HTTP/1.1 and HTTP/2, over cleartext or TLS, and it is meant to face the internet directly.

That last part is a design decision rather than a boast. Most language runtimes ship an HTTP server that is fine for development and then expect a reverse proxy in front of it in production — something else terminates TLS, serves the static files, compresses the responses, enforces the request limits, and speaks HTTP/2 to the browser. Everything on that list is in this module, done properly, because a standard library that leaves them out is not really shipping a server.

Introduction

A First Request

import http

echo http.get('https://example.com').as_text()

get(), post(), put(), patch(), delete(), head(), options() and trace() all exist at module level and all go through one shared client, which keeps its connections open between calls. Request Methods covers each of them.

A First Server

import http

var server = http.server(3000)

server.get('/', @(request, response) {
  response.html('<h1>Hello</h1>')
})

server.get('/users/:id', @(request, response) {
  response.json({ id: request.param('id') })
})

server.listen()

Two things are already true of that server that are worth noticing: GET /users/42 matches the second route and hands the handler '42', and a HEAD / request is answered by the first route with the headers a GET would have produced and no body, because that is what HEAD means.

Blocks in this chapter that list several calls together — three ways to save, four filters, a family of methods — are reference listings, not programs. They show the shape of each call rather than a sequence you could run, and several would conflict if pasted into one file. Anything presented as a complete program on this page runs as written.

A Round Trip You Can Run

Most examples in this chapter call a host that does not exist, or start a server that never returns — neither of which you can paste into a terminal and watch. This one you can. It runs a real server on a real port and makes real requests against it, in two files.

The server goes in its own file, because a function that references an imported module cannot be spawned from the file that imported it. See Concurrency with Isolates for the rule.

Filename: service.zu

import http

def serve(port_channel) {
  var server = http.server(0, '127.0.0.1')

  server.get('/health', @(request, response) {
    response.json({ status: 'ok' })
  })

  server.post('/echo', @(request, response) {
    response.json({ heard: request.json_body().message }, 201)
  })

  # Bind first so the port is known, hand it back, then serve.
  server.bind()
  port_channel.send(server.socket.local_address().port())
  server.listen()
}

Filename: main.zu

import http
import isolate
import .service

var channel = isolate.channel(1)
var task = isolate.spawn(service.serve, channel)
var port = channel.recv()

var api = http.client('http://127.0.0.1:${port}')

echo api.get('/health').as_dict()
echo api.post('/echo', { message: 'hello' }).status

task.cancel()
$ zuri run main.zu
{status: ok}
201

Four details in there are worth carrying into your own code.

Port 0 asks the operating system for a free port. listen() would bind for you, but then nobody could ask which port it got, so the example calls bind() explicitly, reads the bound address, and only then starts accepting. That is the pattern for any test or any process running more than one server.

The port travels over a channel. The client cannot know it in advance, and the two halves are in different isolates with separate heaps, so a shared variable is not an option.

request.json_body() parses the request body, and response.json(value, status) writes the response. The names are not symmetrical because they are doing different jobs: one decodes what arrived, the other encodes and sets a status.

task.cancel() ends it. listen() runs until the server is closed, so without that the program would never exit.

Making Requests

The Shared Client

The module-level functions use a single HttpClient that lives for the life of the program. It pools connections, so a second request to a host it has already talked to skips the TCP handshake and, over HTTPS, the TLS handshake as well.

import http

var page = http.get('https://example.com')

echo page.status          # 200
echo page.headers.get('content-type')
echo page.as_text()

Reach that client with http.shared_client() when you want a setting to apply to every casual call in a program, and build your own for anything more specific.

Building a Client

var api = http.client('https://api.example.com', {
  headers: { 'Authorization': 'Bearer ' + token },
  read_timeout: 5000,
})

var me = api.get('/users/me').raise_for_status().as_dict()

A client built with a base URL prefixes it onto any target that is not already absolute, so the rest of the program addresses the service by path. Everything in the options dictionary is a field of HttpClient, plus headers, and every one of them can also be set afterwards:

api.user_agent = 'my-service/2.1'
api.max_redirects = 3
api.set_header('X-Client-Version', '2.1')

Note Build a client once and keep it. A fresh client per request throws away the connection pool, the cookie jar and the TLS configuration — which is to say, it throws away most of what a client is for.

Request Methods

Every method has a function of its own. They exist twice over, with identical signatures: on the module, where they go through the shared client, and on any HttpClient you build yourself.

The three methods that carry a body take it as their second argument; the rest take the options dictionary in that position.

CallSends a body
get(url, options)no
post(url, data, options)yes
put(url, data, options)yes
patch(url, data, options)yes
delete(url, options)no
head(url, options)no
options(url, options)no
trace(url, options)no
request(method, url, options)if options says so

Every argument after the first is optional, so the short forms all work:

import http

# Reading
http.get('https://example.com/items')
http.get('https://example.com/items', { query: { page: 2 } })

# Creating and updating
http.post('https://example.com/items', { name: 'Widget' })
http.put('https://example.com/items/1', { name: 'Widget', price: 9 })
http.patch('https://example.com/items/1', { price: 11 })

# Removing
http.delete('https://example.com/items/1')

# Asking about a resource without fetching it
var probe = http.head('https://example.com/large.iso')
echo probe.headers.get('content-length')

# Asking what a resource supports
echo http.options('https://example.com/items').headers.get('allow')

The same calls on a client of your own, which is what you want for anything that runs more than once:

var api = http.client('https://api.example.com')

api.get('/items')
api.post('/items', { name: 'Widget' })
api.delete('/items/1')

request() takes the method as an argument, which is how you send one that has no function of its own — a WebDAV PROPFIND, or anything else a service has invented:

api.request('PROPFIND', '/files/', {
  headers: { 'Depth': '1' },
  body: '<propfind xmlns="DAV:"><allprop/></propfind>',
  content_type: 'application/xml',
})

Method names are case-sensitive on the wire, and every registered one is uppercase; a lowercase name is uppercased for you.

Note head() does not follow redirects by default, and delete(), options() and trace() send no body. A DELETE that genuinely needs one — some APIs want a reason in the body — can use request('DELETE', url, { json: ... }).

Request Bodies

The second argument to post(), put() and patch() is the body, and what it is decides how it is sent:

api.post('/items', { name: 'Widget', price: 9 })   # JSON
api.post('/items', ['a', 'b'])                     # JSON
api.post('/items', 'raw text')                     # sent as-is
api.post('/items', file('photo.jpg', 'rb').read()) # sent as-is
api.post('/items', form_builder)                   # multipart

A dictionary or list becomes JSON with a matching Content-Type; a string or bytes is sent exactly as given, with no content type unless you name one.

For anything else, or to be explicit, name it in the options dictionary:

OptionSendsContent-Type
bodya string or bytes, exactly as givennone, unless content_type is set
jsonany value, JSON-encodedapplication/json
forma dictionaryapplication/x-www-form-urlencoded
multiparta MultipartBuildermultipart/form-data, boundary included
content_type—overrides whichever of the above applied
# A login form
api.post('/login', nil, {
  form: { username: 'ada', password: secret },
})

# An explicit content type over a raw body
api.put('/documents/1', nil, {
  body: markdown_source,
  content_type: 'text/markdown; charset=utf-8',
})

# A form whose field repeats
api.post('/search', nil, {
  form: { tag: ['new', 'featured'], q: 'zuri' },
})

Content-Length is set for you from the body, and cannot be overridden — two parties disagreeing about how long a body is, is a request-smuggling bug rather than a formatting choice.

Uploading Files

A file upload is a multipart/form-data body, which MultipartBuilder assembles:

import http

var form = http.MultipartBuilder()

form.add_field('title', 'Holiday')
form.add_field('album', '2026')
form.add_file('photo', 'beach.jpg', file('beach.jpg', 'rb').read(), 'image/jpeg')

var response = http.post('https://example.com/photos', form)

Passing the builder as the body is enough — the Content-Type, the boundary parameter and the Content-Length all follow from it.

MethodDoes
add_field(name, value)adds a plain form field; numbers and booleans are stringified
add_file(name, filename, content, content_type)adds a file part; content is bytes or a string, content_type defaults to application/octet-stream
content_type()the Content-Type header value, boundary included
boundary()the boundary string
build()the assembled body, as bytes
length()how many parts have been added

Several files under one field name is just several calls — that is what an <input type="file" multiple> sends:

for path in paths {
  form.add_file('attachments', os.base_name(path), file(path, 'rb').read())
}

A filename that is not plain ASCII is sent twice: once as a transliterated filename, and once as an RFC 5987 filename*, which is what lets the real name survive a recipient that only understands one of the two.

If you need the body separately — to sign it, to log its size, to send it somewhere this client is not going — build it yourself:

var body = form.build()

api.post('/photos', nil, {
  body,
  content_type: form.content_type(),
})

Note build() assembles the whole body in memory. For an upload large enough that this matters, send the file as a raw body with its own content type instead — multipart/form-data only earns its overhead when there are fields alongside the file.

The boundary is generated from the platform’s secure random source, not from a timestamp or a counter — a boundary a peer can predict is a boundary a peer can write into a field value to forge extra parts.

Reading a Response

var response = api.get('/items/1')

response.status              # 200
response.reason_phrase()     # 'OK'
response.version             # '1.1' or '2'
response.headers.get('etag')

response.as_text()           # the body, decoded as UTF-8
response.as_dict()           # the body, parsed as JSON
response.as_bytes()          # the body, untouched

response.is_ok()             # 2xx
response.is_redirect()       # 3xx
response.is_error()          # 4xx or 5xx

raise_for_status() turns a 4xx or 5xx into a StatusError and returns the response otherwise, so it can be used inline. The response stays reachable on the raised error, which matters because that is where an API usually explains what went wrong:

catch {
  var data = api.get('/items/1').raise_for_status().as_dict()
} as error {
  echo error.response.status
  echo error.response.as_text()
}

A compressed response body is decompressed automatically — gzip, deflate, brotli and zstd — and the Content-Encoding header is removed once it has been, so nothing downstream tries to decode it a second time.

Query Parameters and Headers

api.get('/search', {
  query: { q: 'zuri', page: 2, tag: ['new', 'featured'] },
  headers: { 'Accept-Language': 'en-GB' },
})

A list value repeats the parameter, which is how a query string carries more than one value under one name. Per-request headers are merged over the client’s own.

Authentication

api.get('/me', { auth: ['bearer', token] })
api.get('/me', { auth: ['basic', 'ada', 'lovelace'] })

Credentials passed this way are scoped to the origin you addressed. If the response is a redirect to a different scheme, host or port, they are dropped rather than followed — a redirect to a host of someone else’s choosing is otherwise a way to collect whatever Authorization header was in flight.

Redirects

Redirects are followed by default, up to max_redirects (10), and the number followed is on the response:

var final = http.get('https://example.com/old')

final.redirects    # how many were followed
final.responder    # the URL that finally answered

A 303, and in practice a 301 or 302, becomes a GET with no body when followed — that is what every browser and every other client does, and a server that meant otherwise should have sent 307 or 308, both of which preserve the method and body here.

Turn following off per request or per client:

http.get(url, { follow_redirects: false })

head() does not follow redirects by default, since the point of a HEAD is usually to inspect the very response a redirect would hide.

Cookies

A client with a cookie jar carries cookies between requests, so a login and the requests after it behave the way a browser would:

var session = http.client('https://example.com')
session.enable_cookies()

session.post('/login', nil, { form: { user: 'ada', password: secret } })
session.get('/dashboard')     # sends the session cookie

The jar applies RFC 6265’s matching rules: a cookie without a Domain is host-only, one with a Domain reaches subdomains, a path prefix matches below itself, and a Secure cookie is never sent over cleartext. A server trying to set a cookie for a domain it does not control is refused.

Timeouts and Retries

var api = http.client('https://api.example.com', {
  connect_timeout: 5000,
  read_timeout: 10000,
  write_timeout: 10000,
})

api.get('/slow', { timeout: 30000 })   # this one request

All timeouts are milliseconds. max_retries retries a failed request, but only when the method is idempotent — replaying a POST can mean charging a card twice, so GET, HEAD, PUT, DELETE, OPTIONS and TRACE are retried and nothing else is.

A pooled connection the far end closed while it was idle fails on the next write and looks exactly like a network failure. That specific case is retried once for an idempotent request even at the default max_retries of zero, because otherwise every connection a server reaps would surface as a spurious error.

Streaming a Response

A large response does not have to be held in memory:

var response = api.get('/exports/large.csv', { stream: true })
var out = file('large.csv', 'wb')

while true {
  var chunk = response.body_reader.read(65536)
  if chunk.length() == 0 {
    break
  }
  out.write(chunk)
}

out.close()
api.finish(response)

finish() hands the connection back to the pool once you are done with the body. Until then the connection is yours, since there is no way for the client to know when you have finished reading.

TLS and Certificates

Certificates are verified against the platform’s trust store by default. To talk to a service with an internal or self-signed certificate, trust its authority:

api.add_ca(file('/etc/ssl/internal-ca.pem').read())

There is also verify = false, which turns verification off completely. It exists for local development, and it makes the connection encrypted but unauthenticated — which is to say, trivially interceptable. add_ca() is the right answer everywhere else.

Proxies

A client goes through the proxy the environment names, the way command line tools do: HTTPS_PROXY for https addresses, HTTP_PROXY for http ones, and ALL_PROXY for either, each in upper or lower case. NO_PROXY lists the hosts reached directly, separated by commas: a name covers every name under it, an entry with a port covers only that port, and * covers everything. This machine’s own loopback addresses never go through a proxy.

A client of your own can name its proxy instead, or ignore the environment altogether:

var api = http.client('https://api.example.com', {
  proxy: 'http://ada:secret@proxy.internal:3128',
})

var direct = http.client('https://api.example.com', { trust_env: false })

An https request is tunnelled through the proxy with CONNECT, so the proxy carries the encrypted bytes and never sees what they say, and the certificate checked is the destination’s own. An http request is handed to the proxy whole. Credentials in the proxy address are sent to the proxy alone, as Proxy-Authorization. The proxy itself is reached over plain HTTP; an https:// proxy address is refused.

Serving Requests

Routing

A route pattern is made of literal segments, :name parameters, and an optional trailing catch-all:

PatternMatchesCaptures
/users/users
/users/:id/users/42id = '42'
/users/:id/posts/:post/users/4/posts/7id, post
/files/ + *path/files/css/site.csspath = 'css/site.css'
server.get('/users/:id', @(request, response) {
  response.json({ id: request.param('id') })
})

A literal segment always beats a parameter, and a parameter always beats a catch-all, so /users/new and /users/:id can both exist and the specific one wins — regardless of which was registered first. Matching runs over a trie, so a router with a thousand routes costs the same per request as one with ten.

get(), post(), put(), patch(), delete(), head() and options() register for one method. any() registers GET, POST, PUT, PATCH, DELETE and OPTIONS at once. handle() takes the method as an argument, which is how a method with no function of its own — PROPFIND, or anything a service has invented — gets a route:

server.handle('PROPFIND', '/files/' + '*path', list_properties)

Three things are answered without a handler:

  • HEAD falls back to the GET route for the same path, and the body is dropped on the way out.
  • OPTIONS answers 204 with an Allow header listing what the path actually accepts.
  • A path that exists for other methods answers 405, again with Allow — rather than a 404, which would be a lie.

Name a route to build URLs from it later:

server.get('/users/:id', show_user, 'user.show')

server.routes().url_for('user.show', { id: 42 })    # '/users/42'

The Request Object

server.post('/items', @(request, response) {
  request.method          # 'POST'
  request.path            # decoded and normalised
  request.target          # exactly as it arrived on the wire
  request.version         # '1.1' or '2'
  request.secure          # whether it came over TLS

  request.param('id')                 # a route parameter
  request.query_param('page', '1')    # a query parameter, with a default
  request.header('accept')            # a header field
  request.cookie('session')           # a cookie
  request.bearer_token()              # the token from an Authorization header

  request.text()          # the body as text
  request.json_body()     # the body parsed as JSON
  request.form()          # urlencoded or multipart fields
  request.file('avatar')  # an uploaded file
})

path is percent-decoded and then normalised, in that order — which is the only order that works, since %2e%2e%2f is ../ written to survive a naive check. Route on path; target is the raw form and matching on it is how directory traversal gets through.

An uploaded file is an UploadedFile:

var upload = request.file('avatar')

upload.filename       # what the client claimed
upload.safe_name()    # that, reduced to one path-safe segment
upload.content_type   # what the client declared
upload.size()
upload.save_to('/var/uploads/' + upload.safe_name())

filename and content_type are both attacker-controlled and say nothing about what the bytes are. safe_name() drops any directory component — including a Windows one — and reduces the rest to letters, digits, ., - and _.

Forms and File Uploads

A submitted form arrives as a request body, in one of two encodings, and both are read through the same methods. Which one a browser sends is decided by the enctype on the <form>: the default application/x-www-form-urlencoded for a form of plain fields, and multipart/form-data for one that carries a file.

server.post('/signup', @(request, response) {
  var email = request.form_field('email', '')
  var password = request.form_field('password', '')

  if email.is_empty() {
    response.json({ error: 'email is required' }, 422)
    return
  }

  create_account(email, password)
  response.redirect('/welcome', 303)
})
MethodReturns
form()every field as name -> value, keeping the first of a repeated name
form_all()every field as name -> [value, ...]
form_field(name, fallback)one field’s first value
files()uploaded files as name -> [UploadedFile, ...]
file(name)the first file uploaded under name, or nil

The body is parsed on first use and cached, so reading form() in a middleware and again in the handler costs one parse. A body that is neither encoding gives an empty dictionary rather than raising — use request.json_body() for a JSON body and request.text() for anything else.

A field a form can repeat — a set of checkboxes, a multi-select — needs form_all(), since form() keeps only the first value:

var tags = request.form_all().get('tags', [])

Note A redirect after a successful POST should be 303 See Other, as above. It is what stops a browser re-submitting the form when the user reloads the resulting page.

Receiving an upload

<form method="post" action="/avatar" enctype="multipart/form-data">
  <input type="text" name="caption">
  <input type="file" name="avatar">
  <button>Upload</button>
</form>
import os

server.post('/avatar', @(request, response) {
  var upload = request.file('avatar')

  if upload == nil {
    response.json({ error: 'no file was submitted' }, 400)
    return
  }

  if upload.size() > 2 * 1024 * 1024 {
    response.json({ error: 'the image must be under 2 MB' }, 413)
    return
  }

  var destination = os.join_paths('./uploads', upload.safe_name())
  upload.save_to(destination)

  response.json({
    saved: upload.safe_name(),
    caption: request.form_field('caption', ''),
    bytes: upload.size(),
  })
})

An UploadedFile carries:

MemberIs
filenamethe name the client claimed, exactly as sent
safe_name()that name reduced to one path-safe segment
content_typethe type the client declared, or application/octet-stream
contentthe bytes
size()how many of them
namethe form field it arrived under
headersevery header on that part of the body
save_to(path)writes the content, and returns the byte count
to_text()the content decoded as UTF-8

An <input type="file" multiple> sends several parts under one name, which is what files() returns a list for:

for upload in request.files().get('attachments', []) {
  upload.save_to(os.join_paths('./uploads', upload.safe_name()))
}

Two things a handler must not trust

The filename. It is attacker-controlled, and it routinely contains a full local path from a Windows client, .. segments from a hostile one, or a NUL byte meant to truncate a later check. Never join it to a path directly:

os.join_paths('./uploads', upload.filename)     # no
os.join_paths('./uploads', upload.safe_name())  # yes

safe_name() drops every directory component — both separators — and reduces the rest to letters, digits, ., - and _, returning 'unnamed' when nothing usable is left. Better still, name the file yourself and keep the client’s name as a label:

var stored = '${uuid.v4()}.jpg'
upload.save_to(os.join_paths('./uploads', stored))
record_upload(stored, upload.filename)

The content type. content_type is whatever the client wrote in the part header and says nothing about what the bytes are. A file claiming image/png may be anything at all. When the difference matters, look at the bytes — mime.detect_from_header() reads a real file’s leading bytes, so sniffing an upload means writing it somewhere first:

import mime
import os

var staged = os.join_paths(os.temp_dir(), uuid.v4())
upload.save_to(staged)

var actual = mime.detect_from_header(file(staged, 'rb'))

if actual != 'image/png' and actual != 'image/jpeg' {
  file(staged).delete()
  response.json({ error: 'that is not an image' }, 415)
  return
}

os.rename(staged, os.join_paths('./uploads', stored_name))

Size limits

The body is read into memory, bounded by the server’s max_body_size — 10 MiB by default, which is deliberately small:

server.max_body_size = 50 * 1024 * 1024   # accept uploads up to 50 MB

A request that announces a body over the limit is refused with 413 Content Too Large before a byte of it is read, which is the whole point of Content-Length. One that lies about its length is cut off at the limit and also refused. Either way the connection then closes, since a body that was only partly read leaves nothing safe to parse after it.

A client asking permission first with Expect: 100-continue gets its answer before it sends anything — the server checks the announced length against the limit and either says 100 Continue or refuses outright, so a rejected upload costs one round trip rather than a whole transfer.

For an upload too large to want in memory at all, take it as a raw body and stream it rather than as a form field. multipart/form-data earns its overhead only when there are fields alongside the file.

Request Validation

A request validates itself against a validate schema:

import http
import validate

var create_user = validate.schema({
  name:  validate.required().string().max_length(100),
  email: validate.required().string().email(),
  age:   validate.required().integer().gte(18).lte(120),
})

server.post('/users', @(request, response) {
  catch {
    var data = request.validate(create_user)
    response.json(create_account(data), 201)
  } as error {
    response.json({ errors: create_user.group_errors(error.errors) }, 422)
  }
})

validate() returns the input it checked, so the happy path is one line and the data you go on to use is the data that was validated. Failure raises the schema’s own validate.ValidationError, carrying an errors list of { field, message }; group_errors() turns that into a dictionary keyed by field, which is the shape most front ends want:

{
  "errors": {
    "email": ["The email field must be a valid email address."],
    "age": ["The age field must be greater than or equal to 18."]
  }
}

To branch rather than catch, validate the input yourself — there is no separate API for it:

var result = create_user.check(request.input())

if !result.valid {
  response.json({ errors: result.errors }, 422)
  return
}

The rules themselves — and there are around eighty of them, including cross-field ones like confirmed() and required_if() — belong to the validate module rather than to this one.

What gets validated

request.input() is the dictionary validate() checks. Three sources are merged, each overriding the one before it:

  1. route parameters, from the pattern that matched
  2. query string parameters
  3. the body — a JSON object’s keys, or the submitted form fields

So one schema covers POST /users with a JSON body, GET /users?… with a query string, and /users/:id with a route parameter, without the handler caring which arrived.

Take one source on its own by naming it:

request.validate(schema, 'body')     # only the body
request.validate(schema, 'query')    # only the query string
request.validate(schema, 'params')   # only the route parameters

Two things are deliberately left out of the merge:

  • Uploaded files. Nothing a schema can say about a file is expressible as a rule over its bytes; reach them with request.file() and check them as Forms and File Uploads describes.
  • A JSON body that is not an object. An array or a bare string has no names to merge, so it contributes nothing; read it with request.json_body().

A body that fails to parse as JSON also contributes nothing rather than raising, which leaves the schema’s own required rules to report what is missing. That is a better answer to a client than a parser message.

Values from the wire are strings

A query string and a urlencoded form carry text and nothing else. ?age=36 is the string '36', not the number 36, and that changes which rules hold:

RuleOn '36'Because
integer(), numeric(), gt(), gte(), lt(), lte()reads it as 36these coerce
size(), min(), max(), between()reads it as 2these measure size, which for a string is its character count

That is validate’s documented behaviour, not an accident of this module: min(8) on a password means eight characters. It only surprises when a schema written against a JSON body is later pointed at a query string.

Write a schema that has to serve both with the value rules:

age: validate.required().integer().gte(18).lte(120)   # both
age: validate.required().integer().between(18, 120)   # JSON bodies only

input() does not coerce anything on your behalf. It would have to guess, and a postcode of '01234' or a version of '1.0' silently becoming a number is worse than the rule you have to pick deliberately.

Repeated fields

A name may legally repeat in a query string or a form, so those sources arrive as name -> [values]. input() flattens a name carrying exactly one value to that value, and leaves a name carrying several as a list:

?tag=a           ->  { tag: 'a' }
?tag=a&tag=b     ->  { tag: ['a', 'b'] }

That is what lets a scalar rule see a scalar. A field that must always be a list, however many values arrived, is better read through form_all() or request.query directly and validated with validate’s .* wildcard.

Note The http module does not import validate. validate() takes any object with a check_or_raise() method and calls it, so a server that validates nothing never pays to load a schema engine — and a schema of your own, or from somewhere else, works just as well.

The Response Object

response.text('plain')                    # text/plain
response.html('<h1>hi</h1>')              # text/html
response.json({ ok: true })               # application/json
response.xml('<doc/>')                    # application/xml
response.write('more')                    # append, no content type

response.file('/var/www/report.pdf')      # streamed from disk
response.download('/var/www/report.pdf')  # ...as an attachment
response.render('pages/home', { user })   # a Wire template

response.redirect('/elsewhere')           # 302 by default
response.redirect('/elsewhere', 308)      # method-preserving

response.status = 201
response.header('X-Thing', 'value')
response.content_type('text/csv')
response.cache_for(3600)
response.no_cache()

Every writer sets a sensible Content-Type alongside the body, and each returns the response, so they chain.

file() streams from disk rather than reading the file into memory, which is what makes serving something larger than you would like to hold in memory a one-liner.

Middleware

A middleware takes (request, response, next) and decides whether the rest of the chain runs:

server.use(@(request, response, next) {
  var started = time()
  next()
  echo '${request.method} ${request.path} ${response.status} ' +
    '${(time() - started) * 1000}ms'
})

Not calling next() is how a middleware short-circuits, which is exactly what an authentication or rate-limiting layer wants:

server.use(@(request, response, next) {
  if request.header('x-api-key') != expected {
    response.json({ error: 'unauthorized' }, 401)
    return
  }
  next()
})

Middleware run in the order they were added, outermost first, so one added first wraps everything added after it.

Code after next() runs only when the rest of the chain returned. A handler that raises never comes back there, and the 500 the visitor is sent is decided later, by the server’s error handling. A middleware that reports on the response the client actually receives registers with response.on_finish() instead, which runs once the response is final, failures included:

server.use(@(request, response, next) {
  var started = time()

  response.on_finish(@{
    echo '${request.method} ${request.path} ${response.status} ' +
      '${(time() - started) * 1000}ms'
  })

  next()
})

Callbacks run once each, in the order they were registered, and one that raises is skipped rather than costing the client its response. middleware.logger() is built this way.

Anything a middleware wants to hand to the handler goes on request.context, which is a plain dictionary that exists for exactly that:

server.use(@(request, response, next) {
  request.context['started_at'] = time()
  next()
})

A middleware that wants to act on the response calls next() first and then works on what came back:

server.use(@(request, response, next) {
  next()
  response.header('X-Served-By', hostname)
})

The ones most services need are already written — see Built-in Middleware.

Errors

A handler that raises becomes a 500, and the exception message never reaches the client — an exception message routinely carries a file path, a query, or a fragment of the data being processed, and none of that belongs in a reply to whoever triggered it.

server.on_error(@(error, connection) {
  log.error('${error.message}')
})

server.error_handler(@(request, response, error) {
  response.status = 500
  response.render('errors/500')
})

server.not_found(@(request, response) {
  response.status = 404
  response.render('errors/404')
})

on_error() listeners see everything, including connection-level failures that never reached a handler. error_handler() produces the response body for a failed request; without one, the server sends a bare 500.

Static Files

server.serve_files('/static', './public', {
  cache_age: 86400,
  precompressed: true,
})

That single call brings with it the things a static file needs to actually behave well in front of a browser or a CDN:

  • ETag and Last-Modified, and 304 Not Modified for a conditional request that still holds
  • Range requests, answered with 206 and a Content-Range, and 416 for a range that falls outside the file
  • If-Match, If-Unmodified-Since, If-None-Match, If-Modified-Since and If-Range, evaluated in the order RFC 9110 lays down
  • Content-Type from the file extension
  • index.html for a request that names a directory
  • with precompressed, a .br or .gz sibling served in place of the original when the client accepts that coding

Every request path is percent-decoded, normalised, joined to the root, and then checked to still be inside the root — all three, because doing any two of them is not enough. Dotfiles are not served at all: .env, .git and .htpasswd all end up in a web root at some point in a project’s life, and none of them should ever go out.

For a single-page application, fallback serves the shell for any path that names no file, which is what makes client-side routes work on reload:

server.serve_files('/', './dist', { fallback: 'index.html' })

Compression

Responses are compressed on the way out, and it is on by default. A response is compressed when all of the following hold:

  • server.compression is on (it is),
  • the body is at least server.compression_min_size bytes (1024),
  • its media type is one that benefits — anything text/, plus JSON, JavaScript, XML, SVG, WebAssembly and the +json/+xml types,
  • the client’s Accept-Encoding accepts br or gzip,
  • nothing already set a Content-Encoding,
  • and the compressed body actually came out smaller.

Brotli is preferred over gzip when the client accepts both. Vary: Accept-Encoding is added either way, so a shared cache in front cannot serve a compressed body to a client that asked for none.

Turning it off, or moving the threshold:

server.compression = false          # off entirely
server.compression_min_size = 4096  # only bodies over 4 KiB

The list is an allowlist rather than a denylist on purpose. A JPEG, an MP4 or a zip is already compressed; running it through brotli spends CPU to make it slightly larger.

Note Responses built with response.file() — which includes everything serve_files() serves — are not compressed on the fly. They are streamed from disk rather than held in memory, and compressing them per request would mean reading them into memory to do it. Compress those ahead of time and let the static handler pick the compressed file up:

server.serve_files('/static', './public', { precompressed: true })

With that on, a request for site.css from a client that accepts brotli is answered with site.css.br if it exists, and with site.css.gz if that exists and gzip is accepted — at no CPU cost per request, and with the correct Content-Encoding and Vary.

On the client side there is nothing to configure: Accept-Encoding: gzip, br, deflate, zstd is sent by default and a compressed response body is decoded before you see it.

Streaming a Response Body

A response body can come from a callback instead of memory:

server.get('/export.csv', @(request, response) {
  response.content_type('text/csv')

  response.stream(@(writer) {
    writer.write('id,name\n')

    for row in rows {
      writer.write('${row.id},${row.name}\n')
      writer.flush()
    }
  })
})

Unless a Content-Length was set beforehand, the body is framed with chunked transfer encoding on HTTP/1.1 and as an ordinary DATA stream on HTTP/2 — the handler does not have to know which.

Setting Cookies

response.set_cookie('theme', 'dark', {
  max_age: 86400,
  secure: true,
  same_site: 'Strict',
})

response.clear_cookie('theme')

http_only defaults to true and same_site to 'Lax', which are what a cookie carrying anything sensitive should have; pass them explicitly to opt out. A cookie whose name carries the __Secure- or __Host- prefix has that prefix’s rules applied for it, rather than being sent in a form the browser will silently refuse to store.

For state that belongs to a visitor rather than to the browser, use a session, which keeps the state on the server and puts only an identifier in the cookie. Sessions covers it.

Content Negotiation

import http.negotiate

var type = negotiate.best_match(
  request.header('accept'),
  ['application/json', 'text/html']
)

best_match() picks what the client would most like out of what you can produce, honouring quality weights, with your own order deciding ties. request.accepts('text/html') and request.wants_json() cover the common cases.

Built-in Middleware

http.middleware holds the middleware nearly every public service ends up wanting. Each function returns a middleware, so it is called at the point you register it:

import http
import http.middleware

var server = http.server(3000)

server.use(middleware.request_id())
server.use(middleware.logger())
server.use(middleware.security_headers())

They are ordinary middleware with no special standing — read any of them as a worked example of writing your own.

None of them raises when it is built. That matters for the three that take a verifier — basic_auth(), bearer_auth() and jwt_auth() — where no verifier means the middleware is disabled: it passes every request straight through rather than refusing one, and rather than failing at registration.

# Authentication, off. Everything else in the chain is untouched.
server.use(middleware.jwt_auth(nil))

That is the right shape for a middleware, and it is deliberately useful: blanking a verifier switches its middleware off without unpicking the chain around it, which is what you want when bisecting a request that is failing somewhere in a stack of them.

Disabling is opt-in and never a fallback: a verifier that is present and refuses a request still refuses it.

logger()

Writes one line per request, after the response is finished.

server.use(middleware.logger())
127.0.0.1 - [Wed, 09 Sep 2026 20:09:17 GMT] "GET /static/big.txt HTTP/1.1" 200 320 3.07ms

That is the Common Log Format with the response time added, which every log analyser already understands. The size comes from Content-Length when the body is file-backed or streamed, so a static file logs its real size rather than zero.

OptionDefaultMeaning
sinkechoa function taking the finished line
formatthe format abovea function (request, response, milliseconds) returning the line
trust_proxyfalselog the forwarded client address rather than the peer
server.use(middleware.logger({
  sink: @(line) { log_file.puts(line + '\n') },
  trust_proxy: true,
}))

Structured logging is a format that returns JSON:

server.use(middleware.logger({
  format: @(request, response, elapsed) {
    return json.encode({
      method: request.method,
      path: request.path,
      status: response.status,
      ms: elapsed,
      id: request.context.get('request_id', nil),
    })
  },
}))

A request that raises is logged too, with the status the client was sent once the server had answered the failure: the 500 of the default handling, or whatever the error handler chose.

request_id()

Gives every request an identifier, so one request can be followed across a log file and across services.

server.use(middleware.request_id())

The identifier lands on request.context['request_id'] and in the response’s X-Request-Id. One supplied by the client is echoed rather than replaced, which is what makes a trace continue across a hop.

The single optional argument is the header name:

server.use(middleware.request_id('X-Correlation-Id'))

An identifier that came from outside is written into your logs, so it is capped at 64 characters and stripped of anything that is not plainly printable before it is trusted that far.

security_headers()

Adds the response headers a browser acts on to harden a page.

server.use(middleware.security_headers())
OptionDefaultHeader
content_type_options'nosniff'X-Content-Type-Options — stops the browser second-guessing a declared content type
frame_options'DENY'X-Frame-Options — refuses to be framed, which is what clickjacking needs
referrer_policy'strict-origin-when-cross-origin'Referrer-Policy — keeps paths and queries out of outbound referrers
hsts_max_age31536000Strict-Transport-Security
hsts_subdomainsfalseadds includeSubDomains
content_security_policynot setContent-Security-Policy
server.use(middleware.security_headers({
  content_security_policy: "default-src 'self'; img-src 'self' data:",
  hsts_subdomains: true,
  frame_options: 'SAMEORIGIN',
}))

Pass nil for any of them to leave that header off entirely.

Content-Security-Policy has no default because a wrong one breaks the page and a permissive one is theatre; it depends on what your application actually loads.

Strict-Transport-Security is only sent over HTTPS. Over cleartext it is at best ignored, and at worst a way to lock a client out of a host that has no certificate.

Every header is set with set_default, so a handler that set its own keeps it.

cors()

Answers CORS preflights and adds the headers a browser needs before it will let script read a response from another origin.

server.use(middleware.cors({
  origins: ['https://app.example.com'],
  credentials: true,
}))
OptionDefaultMeaning
origins['*']origins allowed, matched exactly; '*' allows any
methodsthe seven usual oneswhat a preflight is told is allowed
headersreflect what was askedrequest headers script may set
exposenoneresponse headers script may read
credentialsfalsewhether cookies and Authorization may be sent
max_age86400how long a preflight may be cached, in seconds

A request with no Origin is not a cross-origin request and passes straight through. An Origin that is not on the list gets no CORS headers at all — the browser then refuses to hand the response to script, which is the correct outcome; answering with an error instead would leak whether the resource exists.

Note A wildcard and credentials cannot be combined — every browser rejects that pairing. When credentials is set, the requesting origin is echoed back instead, which makes the origins list the only thing standing between an attacker’s page and an authenticated response. Name the origins explicitly whenever credentials are in play.

Vary: Origin is added whenever the answer depends on the request’s origin, so a shared cache cannot serve one origin’s response to another.

basic_auth()

Requires HTTP Basic authentication.

server.use(middleware.basic_auth(@(user, password) {
  return user == 'admin' and http.util.secure_equals(password, secret)
}, 'Admin area'))

The first argument is called with the username and password and returns whether they are acceptable. The second is the realm named in the challenge, and defaults to 'Restricted'.

On success the username is put on request.context['user']. On failure the response is 401 with a WWW-Authenticate header, which is what makes a browser show its credentials prompt.

Passing nil in place of the verifier disables the middleware entirely — see the note above.

Compare secrets with http.util.secure_equals() rather than ==: it compares in constant time, so a wrong guess and a nearly-right one take the same time and the comparison does not hand over the secret one character at a time.

To guard part of a site rather than all of it, wrap it:

var guard = middleware.basic_auth(check_credentials)

server.use(@(request, response, next) {
  if request.path.starts_with('/admin') {
    guard(request, response, next)
    return
  }
  next()
})

Note that the next() a guard calls on success continues the whole remaining chain, not just the handler. Two authentication middleware registered globally therefore both apply to every request; scope each of them by prefix, as above, when they are meant for different parts of the site.

bearer_auth()

Requires a bearer token — the usual shape for an API.

server.use(middleware.bearer_auth(@(token) {
  var claims = verify_jwt(token)
  return claims == nil ? false : claims.subject
}))

The verifier is called with the token and returns either something falsy to refuse it, or a value to attach — that value lands on request.context['user'], so returning the subject, the claims, or a whole user record all work.

Failure is 401 with a WWW-Authenticate: Bearer challenge and a JSON body. The realm is the optional second argument and defaults to 'api'. A nil verifier disables the middleware, as with basic_auth().

middleware.parse_bearer(header) and middleware.parse_basic(header) are exported separately for code that needs to read an Authorization header without installing a middleware at all, and request.bearer_token() reads the token straight off a request.

For a JSON Web Token specifically, jwt_auth() does the verification, the claim handling and the RFC 6750 challenges rather than leaving them to the verifier you write here.

jwt_auth()

Requires a valid JSON Web Token, verified by the jwt module.

import http.middleware
import jwt

server.use(middleware.jwt_auth(
  jwt.Verifier(secret, { algorithms: ['HS256'], audience: 'api' })
))

The first argument is either a jwt.Verifier or any function taking the token and returning its claims. The function form is what covers a key set resolved by kid, or anything else the jwt module can do that a fixed verifier cannot:

server.use(middleware.jwt_auth(@(token) {
  return jwt.verify_with_jwks(token, keys, { audience: 'api' })
}))
OptionDefaultMeaning
realm'api'the realm named in the challenge
optionalfalseattach the claims when a valid token is present, but do not refuse a request without one
scopesnonea list of scopes every token must carry

On success the claims land on request.context['claims'], and the sub claim — the usual place an issuer puts the account a token speaks for — on request.context['user']:

server.get('/me', @(request, response) {
  response.json({
    account: request.context.get('user', nil),
    issued_at: request.context.get('claims', {}).get('iat', nil),
  })
})

Failures follow RFC 6750 §3:

SituationAnswer
no token at all401, WWW-Authenticate: Bearer realm="api"
the token does not verify401, with error="invalid_token"
valid, but missing a required scope403, with error="insufficient_scope" and the scope that was needed

The distinction in that last row is the point of scopes: a token that failed to authenticate and a token that authenticated but is not allowed to do this are different problems, and answering both with 401 tells a client to go and get a new token when a new token will not help.

Note A challenge never says which check a token failed. An expired token and a forged one get exactly the same error_description, because which one it was is useful to your logs and useful to an attacker, and to nobody else. If you need the reason, log it from the verifier you passed in.

optional is for a route that behaves differently when it knows who is asking without requiring it — a public page that shows an edit button to its author:

server.use(middleware.jwt_auth(verifier, { optional: true }))

server.get('/posts/:id', @(request, response) {
  var viewer = request.context.get('user', nil)
  response.json(render_post(request.param('id'), viewer))
})

A token that is present but invalid is still refused under optional. Ignoring a bad token would let a client tamper with one and get the anonymous view rather than an error, which hides exactly the problem worth surfacing.

A nil verifier disables the middleware: every request passes through unauthenticated, and request.context['claims'] is simply never set. Nothing here raises when it is built, so registering it never needs a catch around it.

var verifier = production ? jwt.Verifier(secret, options) : nil

server.use(middleware.jwt_auth(verifier))

Like HttpRequest.validate(), this does not import the jwt module — the verifier is built by the caller, which keeps the token format the application’s business and means a server that authenticates nothing never pays to load it.

rate_limit()

Limits how many requests one client may make in a window of time.

server.use(middleware.rate_limit({ limit: 100, window: 60 }))
OptionDefaultMeaning
limit60requests allowed per window
window60the window, in seconds
keythe client addressa function of the request returning the bucket key
trust_proxyfalsederive the address from forwarding headers

Every response carries the current state, so a well-behaved client can back off before it is refused:

RateLimit-Limit: 100
RateLimit-Remaining: 87
RateLimit-Reset: 41

Over the limit is 429 Too Many Requests with a Retry-After.

Rate-limit per API key rather than per address by supplying a key:

server.use(middleware.rate_limit({
  limit: 1000,
  window: 3600,
  key: @(request) {
    return request.header('x-api-key', nil) or request.client_ip() or 'anonymous'
  },
}))

Note The counter lives in memory, so it is per worker: running with workers: 4 makes the effective limit four times limit. That is a deliberate trade — a shared counter would need shared state, and this is meant to blunt a runaway client rather than to meter billing. Divide limit by the worker count if the exact number matters, or keep the count somewhere both workers can see.

etag()

Computes a weak ETag over a finished response body and answers 304 Not Modified when the client already has that version.

server.use(middleware.etag())

The single optional argument is the smallest body worth tagging, in bytes; it defaults to 128, below which the validator costs more than the body it would save.

A handler that already set its own ETag is left alone — it knows something about the resource that hashing the bytes does not — and so are file-backed responses, which serve_files() already tags from the file’s size and modification time.

Because it works on the finished body, it pairs with anything: a rendered template, a JSON document, a generated report.

force_https()

Redirects every request that arrived over cleartext to the same URL over HTTPS.

server.use(middleware.force_https())
OptionDefaultMeaning
status308the redirect status; 308 preserves the method and body
portthe defaultthe HTTPS port, when it is not 443

This belongs on the cleartext listener, which usually exists only to perform this redirect:

# port 80: redirect and nothing else
var redirector = http.server(80, '0.0.0.0')
redirector.use(middleware.force_https())

Pair it with security_headers()’s Strict-Transport-Security on the TLS listener, so that after the first visit the browser stops making the cleartext request at all.

Ordering

Middleware run outermost first, so the order they are registered in is the order they wrap the request. A workable default:

server.use(middleware.request_id())        # so everything after can log it
server.use(middleware.logger())            # so it sees the final status
server.use(middleware.force_https())       # before any work is done
server.use(middleware.security_headers())
server.use(middleware.cors({ origins: allowed }))
server.use(middleware.rate_limit({ limit: 100 }))
server.use(middleware.jwt_auth(verifier))   # after the cheap refusals
server.use(middleware.etag())              # innermost: it needs the finished body

The reasoning behind each position: identifiers before anything that might log, logging outside everything so it records what actually happened, cheap refusals before expensive ones, and anything that inspects the response body innermost, where the body exists.

Sessions

A session is state that belongs to one visitor, kept on the server and found again by a cookie the browser sends back. Only the identifier travels, so a visitor can neither read what the session holds nor change it, and the cookie is worth nothing to anyone who cannot present the exact value that was issued.

import http

var server = http.server(3000)

server.use(http.session.session())

server.get('/', @(request, response) {
  var seen = request.session().get('seen', 0) + 1

  request.session().set('seen', seen)
  response.text('visit ${seen}')
})

server.listen()

session() is middleware. Register it once, above anything that reads a session, and every handler below it reaches its own through request.session().

Nothing Happens Until Something Uses It

A request that never touches its session costs nothing: no read from the store, no write, and no Set-Cookie. The record is created the first time something is written to the session, which is what keeps a crawler working through a public site from filling the store with empty sessions.

A request that only reads an existing session writes nothing back either, beyond moving the idle expiry along at most once every touch_interval seconds.

Signing In

server.post('/login', @(request, response) {
  var account = authenticate(request.form())

  if account == nil {
    response.status = 401
    response.html(render_login('Those details do not match.'))

    return
  }

  request.session().regenerate()
  request.session().set('account', account.id)

  response.redirect('/')
})

regenerate() is the line to get right. It gives the session a new identifier and destroys the record the old one named, keeping everything the session holds.

Without it, an attacker who can set a cookie in the victim’s browser beforehand — through a stray subdomain, an open redirect, a shared machine — knows the identifier the victim will be signed in under, and can simply use it afterwards. That is session fixation, and a new identifier is the whole of the defence. Call it whenever what the session means changes: signing in, elevating to administrator, a step-up authentication.

Signing out is destroy(), which removes the record and has the response expire the cookie:

server.post('/logout', @(request, response) {
  request.session().destroy()
  response.redirect('/')
})

clear() is the other one: it empties the session without ending it, keeping the identifier and the cookie.

Flash Messages

A handler that does the work and redirects cannot render the message saying what happened. A flash carries it to the page that can:

server.post('/posts', @(request, response) {
  create_post(request.form())

  request.session().flash('notice', 'Your post is up.')
  response.redirect('/posts')
})

server.get('/posts', @(request, response) {
  response.html(render(posts(), request.session().take_flash('notice')))
})

A flash is spent by the next request that touches the session at all, whether or not that request asks for this one. A page that looked at the session and did not read the message does not leave it for the page after; a request that never touched its session — a static file, an image — leaves it waiting.

Where Sessions Are Kept

Four things can hold a session, and swapping between them changes one line.

FileStore, one file per session. This is the default, because it needs no setup and is shared between the workers http.serve() starts. Given no directory it uses a private subdirectory of the platform’s temporary directory, created 0700:

server.use(http.session.session())

That directory is cleared on whatever schedule the platform keeps, and on most of them at every reboot, so name your own for anything that has to outlive the host:

server.use(http.session.session({
  store: http.session.FileStore('/var/lib/app/sessions'),
}))

A session file is a bearer credential in the same way the cookie is. The directory is created 0700 and each file 0600, and a directory that every user on the machine can reach is refused rather than used — which is why the temporary directory itself is never the default. Pass strict_permissions: false to accept one anyway.

SqlStore, one row per session. It lives in its own import, so a program using the default store never loads the sql module:

import http.session.sql { SqlStore }
import sql

var store = SqlStore(sql.pool('postgres://localhost/app'))
store.migrate()

server.use(http.session.session({ store }))

migrate() creates the table and its index if they are not there, and is safe to call on every start. It takes a sql.Connection or a sql.Pool; a server wants the pool.

MemoryStore, for a test. Nothing survives a restart, and nothing is shared between workers, so a browser whose next request lands on a different worker arrives with a session that worker has never heard of. It is the right store for a test and the wrong one for traffic.

Something of your own. A store is five methods, none of which sees an identifier or understands a payload:

class RedisStore < http.session.SessionStore {
  @new(client) {
    self._client = client
  }

  read(key) {
    return self._client.get('session:' + key)
  }

  write(key, payload, expires_at) {
    self._client.set_with_ttl('session:' + key, payload, (expires_at - time()).ceil())
  }

  destroy(key) {
    self._client.remove('session:' + key)
  }

  gc(now) {
    # Redis expires keys itself.
    return 0
  }
}

touch() has a working default built on read() and write(); override it where the backend can move an expiry on its own.

The key a store is handed is the SHA-256 of the identifier, not the identifier. Someone who reads the directory, the table, or a backup of either learns what is in the sessions but cannot resume one, because the value the browser presents is the preimage.

When a Session Ends

Two clocks, and a session ends at whichever runs out first:

OptionDefault
idle_timeout7200seconds of inactivity; rolls forward while the visitor is active
lifetime86400seconds the session may live however active it is

Either may be nil to remove that limit, but not both.

The absolute lifetime is what asks a tab left open overnight to sign in again.

touch_interval is how often a request that only read the session bothers to move the idle expiry. It is the difference between a store write on every request and a store write once a minute. It defaults to sixty seconds, or half the idle timeout where that is shorter, and one set by hand has to stay under idle_timeout — otherwise a session in constant use still expires, because nothing ever moves it.

Nothing schedules a sweep of expired sessions, so one rides along with ordinary traffic: gc_probability (default 0.01) is the chance that a write also sweeps. Set it to 0 where a cron job calls store.gc() instead.

OptionDefault
name'zuri_session'
path'/'
domainnilnil scopes the cookie to the exact host, which is the narrower choice
securenilnil follows the request’s own scheme
http_onlytrue
same_site'Lax''Strict', 'Lax' or 'None'
persistentfalsewhether the cookie outlives the browser

secure following the request is what lets development over cleartext work while production over TLS gets Secure without being told. A deployment behind a proxy that terminates TLS sets secure: true itself.

persistent: false sends a cookie that ends with the browser session, which is what a sign-in should normally do. persistent: true sends Max-Age instead, and moves it forward on every request.

An identifier is 32 bytes from the platform’s cryptographic generator, which is far beyond guessing. Signing adds nothing against that, and everything against volume: with a secret set, a cookie this server did not issue is thrown out after one HMAC, rather than after a read from disk or a query to the database.

import env

server.use(http.session.session({ secret: env.require('SESSION_SECRET') }))

Set it on anything facing the open internet, and give every worker the same secret — a cookie issued by one is otherwise refused by the next. Turning signing on refuses the cookies issued before it, so it signs everyone out once.

What a Session May Hold

Whatever JSON holds: strings, numbers, booleans, nil, lists and dictionaries of those. A class instance is not JSON, and storing one raises when the session is written.

Sessions are for identity and small state — who is signed in, which steps of a form are done, what to say on the next page. A payload over max_size (default 65536 bytes) raises SessionError; the answer to that is a row in a database with the session holding its key.

Across Workers

http.serve() runs each worker in its own isolate, so the store is built inside setup rather than handed in from outside:

# app.zu
import http
import http.session

def setup(server) {
  server.use(http.session.session({
    store: http.session.FileStore('/var/lib/app/sessions'),
    secret: os.get_env('SESSION_SECRET'),
  }))

  server.get('/', @(request, response) {
    response.text(request.session().get('account', 'nobody'))
  })
}

A sql connection belongs to the isolate that opened it, so a worker using SqlStore opens its own pool in setup too.

What Sessions Do Not Do

Two requests writing the same session at the same moment — parallel requests from one browser tab, usually — both succeed, and the one that finishes last is the one that survives. The file store writes through a rename, so a reader never sees half a session, and the SQL store writes in one statement; neither takes a lock. Do not use a session as a counter that several requests increment at once.

TLS

var server = http.server(443, '0.0.0.0')
server.load_certs('/etc/certs/site.crt', '/etc/certs/site.key')
server.listen()

Or from strings, with more control:

server.use_tls(cert_chain_pem, private_key_pem, {
  min_version: '1.2',
  client_ca: internal_ca_pem,
  require_client_cert: true,
})

cert_chain must be the server certificate followed by any intermediates. Leaving the intermediates out is the single most common TLS misconfiguration, and it fails only for the clients that do not happen to have them cached — which is to say, it fails in a way that looks fine from your own browser.

Note The certificate and key are parsed when the first handshake runs, not when they are set, so a malformed or mismatched pair surfaces as a failed connection rather than as an error from use_tls(). Make a request against the server as part of starting it if you want to find out early.

Turning on TLS also advertises h2 through ALPN, so a browser gets HTTP/2 without anything else being configured.

HTTP/2

HTTP/2 needs nothing turned on. HttpServer speaks it when ALPN negotiates h2 over TLS, and when a cleartext client opens with the HTTP/2 connection preface. HttpClient uses it whenever a server negotiates h2 for an HTTPS connection. Routes, middleware and handlers are the same code either way; request.version is '2' when it applies.

What that gets you is one connection carrying every request, HPACK header compression, and flow control — rather than the six-connection scramble HTTP/1.1 forces a browser into.

Set server.http2 = false to turn it off, which also stops h2 being advertised in ALPN.

http.h2 exposes the machinery underneath — the frame codec, HPACK and its Huffman coding — for anything that needs to speak the protocol directly.

Note Streams are multiplexed on the wire but handled in the order they complete: one connection is served by one isolate, so a request that arrives while another is being handled is read, buffered, and answered next rather than in parallel. That is a throughput property, not a correctness one, and http.serve() is how you get more than one connection served at a time.

WebSockets

import http.websocket as ws

server.get('/ws', @(request, response) {
  var socket = ws.accept(request, response, { protocols: ['chat'] })

  while socket.is_open() {
    var message = socket.receive()

    if message == nil or message.is_close() {
      break
    }

    socket.send('you said: ' + message.text())
  }
})

accept() completes the RFC 6455 handshake and takes the connection over; the server writes nothing more on it. Fragmented messages are reassembled, ping frames are answered, and a close frame from the peer is answered and then handed back so you can see the code and reason.

The client side is ws.connect():

var socket = ws.connect('wss://example.com/ws')

socket.send('hello')
echo socket.receive().text()
socket.close()

Client frames are masked with a key from the platform’s secure random source, as the RFC requires — the mask is what stops a hostile page from steering a proxy into caching a forged response.

Server-Sent Events

Where a WebSocket gives you a duplex connection, server-sent events give you a long-lived response body and nothing else — which is all a live feed of updates needs, and it survives proxies that would refuse an upgrade.

import http.sse

server.get('/events', @(request, response) {
  sse.stream(response, @(events) {
    for update in updates {
      events.send(update, 'update', update.id)
    }
  })
})

A browser’s EventSource reconnects on its own and sends back the last id it saw; sse.last_event_id(request) reads it, which is how a stream resumes where it left off instead of replaying from the start.

Reverse Proxying

var api = http.ReverseProxy('http://127.0.0.1:9000', {
  strip_prefix: '/api',
})

server.any('/api/' + '*path', @(request, response) {
  api.handle(request, response)
})

Hop-by-hop headers are stripped in both directions — including any field the message’s own Connection header names, which is how an endpoint declares an extra one. Forwarding a client-supplied Transfer-Encoding to an upstream that frames it differently is the classic request-smuggling setup, so this is not optional.

X-Forwarded-For, X-Forwarded-Proto, X-Forwarded-Host and RFC 7239’s Forwarded are added, and the response body is streamed rather than buffered.

Several upstreams get a LoadBalancer, which is round-robin with an upstream that fails taken out of rotation for a while rather than retried on every request:

var pool = http.LoadBalancer([
  'http://10.0.0.1:9000',
  'http://10.0.0.2:9000',
], { recovery_time: 30 })

Running in Production

Using More Than One Core

listen() serves connections on the calling isolate, one at a time. http.serve() runs the same pipeline across a pool of isolates: one accept loop hands each connection to whichever worker takes it next.

# app.zu
import http

def setup(server) {
  server.get('/', @(request, response) {
    response.text('hello')
  })

  server.serve_files('/static', './public', { cache_age: 86400 })
}
# main.zu
import http
import .app

http.serve(app.setup, {
  host: '0.0.0.0',
  port: 3000,
  workers: 8,
})

setup is called once inside each worker with that worker’s own HttpServer. It can live in the main script or in a module, and it can use whatever it imports; a module it reaches for is loaded again inside the worker rather than shared with it. Putting it in a module of its own, as above, is still the better shape once it registers more than a couple of routes.

Isolates share no memory, so anything a worker needs — a cache, a connection pool, a counter — is per worker. That is the trade the model makes: no locks and no shared-heap garbage collection pauses, in exchange for state that has to be either per worker or in something outside the process.

backlog bounds how many accepted connections may queue for a free worker. Bounding it is deliberate: an unbounded queue under load means accepting connections faster than they can be served and then answering all of them late, rather than letting the kernel’s own listen backlog apply back-pressure.

Limits

Every part of a request whose size a peer controls has a ceiling:

SettingDefaultBounds
max_line_size8 KiBthe request line, and each header line
max_header_size64 KiBthe header section
max_header_count100how many header fields
max_body_size10 MiBthe request body
header_timeout10show long the whole head may take to arrive
keep_alive_timeout5show long an idle connection is held
max_keep_alive_requests1000how many requests one connection may serve

header_timeout is the one that is easy to leave out and matters most: without it, a client sending one byte every few seconds holds a connection open indefinitely while never tripping any single read’s timeout. That is what slowloris is.

Raise max_body_size deliberately for a service that takes uploads, rather than discovering it by accident:

server.max_body_size = 100 * 1024 * 1024

A request that announces a body over the limit is refused with 413 before a byte of it is read, which is the whole point of the header.

Graceful Shutdown

import os

http.serve(app.setup, {
  port: 3000,
  on_ready: @(address, stop) {
    echo 'listening on ${address}'
    os.on_signal('INT', @() {
      stop()
      return true
    })
    os.on_signal('TERM', @() {
      stop()
      return true
    })
  },
})

stop() closes the listener, which is what breaks the blocking accept; serve() then waits for every worker to finish the connection it is on before returning. For a single-isolate server, close() does the same thing.

Behind Another Proxy

If something else really is in front — a CDN, a load balancer you control — tell the server, and tell it which addresses to believe:

server.trust_proxy = true
server.trusted_proxies = ['10.0.0.1', '10.0.0.2']

request.client_ip(true, server.trusted_proxies) then walks the forwarding chain in from the proxy end and returns the rightmost entry that is not itself a trusted proxy. Taking the leftmost entry — the common shortcut — hands an attacker whatever client address they care to claim, which matters the moment an address is used for rate limiting, allowlisting or an audit trail.

Left off, forwarding headers are ignored entirely and the peer is the client. That is the right default: those headers are request headers, which is to say anyone can write anything in them.

What the Module Refuses

Some of what this module does is refuse things, and it is worth being explicit about which, since each refusal is a request that some other implementation would have accepted:

  • A header field name with whitespace before its colon. Content-Length : 5 is read as a length by some parsers and as an unknown field by others, and that disagreement is a smuggled request.
  • Two Content-Length fields that disagree, for the same reason.
  • A Content-Length alongside a Transfer-Encoding.
  • A Transfer-Encoding whose last coding is not chunked.
  • An obsolete folded header line — RFC 9112 deprecated it, and intermediaries unfold it differently.
  • An HTTP/1.1 request with no Host, or with two.
  • A header value containing CR, LF or NUL, at the point it is set: a value that can inject a newline is a response-splitting bug wherever it eventually lands.
  • An uppercase header field name over HTTP/2, and any of the connection-specific fields there.
  • A static file path that resolves outside the root, before or after percent-decoding, and any dotfile.

Module Reference

ModuleWhat it holds
httpthe facade: get(), post(), server(), client(), serve()
http.statusstatus codes, reason phrases, and predicates
http.headersHeaders, field validation, canonical names
http.cookiesCookie, CookieJar, and both cookie header formats
http.sessionSession, SessionStore, FileStore, MemoryStore
http.session.sqlSqlStore, for sessions kept in a database
http.requestHttpRequest, query string encoding and decoding
http.responseHttpResponse
http.routerRouter, Route, RouteMatch
http.middlewareCORS, logging, security headers, auth, rate limiting
http.filesStaticFiles, byte ranges, validators
http.multipartMultipartBuilder, UploadedFile, the parser
http.negotiateAccept parsing and matching
http.bodymessage framing, chunked encoding, content codings
http.h1the HTTP/1.1 wire codec
http.h2HTTP/2: frames, HPACK, connections
http.websocketRFC 6455, client and server
http.sseserver-sent events
http.proxyReverseProxy, LoadBalancer
http.streamthe buffered connection both protocols read and write through
http.utildates, header parameters, percent coding, path normalisation
http.errorsthe error hierarchy

Two other standard library modules meet this one without it depending on either: HttpRequest.validate() takes a schema from validate, and middleware.jwt_auth() takes a verifier from jwt. Both are duck-typed, so neither module is loaded by a server that does not use it.

Every error this module raises descends from HttpError: ProtocolError for a malformed message, ConnectionError for a connection that failed, TimeoutError, TooLargeError, TooManyRedirectsError, StatusError, UnsupportedProtocolError and SessionError.

Imagine

The imagine module is Zuri’s image library: decoding, drawing, transforming and encoding raster images.

It covers the whole path an image takes through a program. A photograph arrives as an upload, gets checked, oriented, resized and sharpened, has a watermark composited onto it, and goes back out as a WebP. A chart is built from nothing but shapes and text. An avatar is cropped to a square, rounded off, and cached as a data URL. None of that needs anything outside the standard library.

Every image on this page is the output of the code beside it, produced by docs/book/src/imagine/figures.zu. Re-run that script and the figures follow whatever the module actually does.

Blocks in this chapter that list several calls together — three ways to save, four filters, a family of methods — are reference listings, not programs. They show the shape of each call rather than a sequence you could run, and several would conflict if pasted into one file. Anything presented as a complete program on this page runs as written.

Following Along

Most examples below operate on a file called photo.jpg. Use your own, or make one — the module can draw its own test subject:

import imagine { Image }

Image(320, 240, '#1e3a5f')
  .fill_circle(220, 70, 40, '#ffd166')
  .fill_rect(0, 170, 320, 70, '#2a9d8f')
  .fill_polygon([[40, 170], [110, 80], [180, 170]], '#264653')
  .save('photo.jpg')

echo Image.open('photo.jpg').size()
{width: 320, height: 240}

Everything from here on assumes that file exists in the working directory.

Introduction

A First Image

import imagine { Image }

Image.open('photo.jpg')
  .thumbnail(400, 400)
  .save('thumb.webp')

Three lines: decode, shrink, encode. The format on the way in comes from the file’s contents, and the format on the way out comes from the extension you saved it as.

Building an image from scratch works the same way.

import imagine { Image, Color }

Image(400, 200, '#0f172a')
  .fill_circle(200, 100, 70, '#38bdf8')
  .circle(200, 100, 70, 'white', { thickness: 3 })
  .save('badge.png')

Almost every method returns an image, so operations chain.

Copying and Mutation

There is one rule in this module that is worth learning before anything else, because everything else follows from it:

Operations that change the image’s size return a new image. Everything else changes the image in place.

resize(), crop(), rotate(), flip(), transpose() and pad() all leave the original untouched and hand back a new one. Drawing, filters and compositing all modify the image you called them on and return it so the chain continues.

var original = Image.open('photo.jpg')

var small = original.thumbnail(200, 200)   # original is untouched
small.grayscale()                          # small is now grey

echo original.size()                       # still the full size, still colour

That means a mixed chain reads correctly, and it means clone() is what you reach for when you want to keep an image before filtering it.

var greyed = photo.clone().grayscale()     # photo keeps its colour

Colours

Writing a Colour

Anywhere a colour is expected, five spellings are accepted. These are all the same red:

image.fill(Color(255, 0, 0))
image.fill('#ff0000')
image.fill('red')
image.fill(0xFF0000FF)
image.fill([255, 0, 0])

Hexadecimal strings take all four CSS lengths, with or without the leading #: #f00, #f00c, #ff0000, #ff0000cc. Names are the full CSS Color Level 4 list, matched ignoring case, spaces and hyphens, so 'Dark Sea Green' and 'darkseagreen' are the same colour.

Alpha runs 0 (transparent) to 255 (opaque), as in CSS and PNG.

The Color Class

Color is an immutable value type. Every method that would change one returns a new colour, so a colour held in a variable is safe to pass around.

import imagine { Color }

var brand = Color.hex('#4f46e5')

brand.r                    # 79
brand.to_hex()             # '#4f46e5'
brand.to_packed()          # 0x4F46E5FF

brand.with_alpha(128)      # half transparent
brand.fade(0.5)            # halves whatever alpha it already had
brand.lighten(20)          # 20 percentage points of HSL lightness
brand.darken(20)
brand.saturate(15)
brand.desaturate(15)
brand.rotate_hue(180)
brand.mix('white', 0.25)   # a quarter of the way to white
brand.invert()
brand.to_grayscale()

over() flattens a translucent colour against a background, which is what happens when it is drawn onto something solid:

Color(255, 0, 0, 128).over('white')    # '#ff7f7f'

Colour Spaces

Conversions come from the colors module, so the conventions are the same ones it and CSS use: hue in degrees, everything else in percentage points from 0 to 100.

Color.hsl(240, 100, 50)            # '#0000ff'
Color.hsv(120, 100, 100)           # '#00ff00'
Color.hwb(0, 0, 0)                 # '#ff0000'
Color.cmyk(0, 100, 100, 0)         # '#ff0000'
Color.lab(40.7, 50.6, -79.1)       # perceptually uniform
Color.xyz(0.2, 0.15, 0.7)

brand.to_hsl()      # { h: 243.4, s: 75.4, l: 58.6, a: 255 }
brand.to_hsv()
brand.to_hwb()
brand.to_cmyk()
brand.to_lab()
brand.to_xyz()

Lab is the right space for interpolating between two colours when the intermediate steps need to look evenly spaced to the eye. CMYK here is the naive conversion with no output profile, so it is right for generating colours and wrong for predicting a printing press.

Contrast and Accessibility

var background = Color.hex('#1e293b')

background.luminance()                  # 0 to 1, WCAG relative luminance
background.contrast_ratio('#ffffff')    # 1 to 21
background.is_dark()                    # true
background.best_contrast()              # white, since the background is dark

WCAG asks for a ratio of at least 4.5 for normal text and 3 for large text, so best_contrast() is the quick way to pick a legible foreground for a colour you did not choose:

var label = swatch.best_contrast('#ffffff', '#111111')

card.text(20, 20, name, font, label)

Loading and Saving

Opening an Image

var photo = Image.open('photo.jpg')          # from a path
var photo = Image.open(file('photo.jpg'))    # from an open file
var photo = Image.decode(upload)             # from bytes in memory

The format is detected from the contents rather than the name, so a mislabelled file still opens. When the contents cannot be identified — TGA carries no signature of its own — Image.open() falls back to the file extension.

Image.decode() takes options:

Image.decode(upload, { format: 'png' })   # fail unless it really is a PNG
Image.decode(raw, { orient: false })      # skip EXIF auto-rotation

Naming the format explicitly is worth doing when you already know it from a Content-Type or an extension: a file claiming to be a PNG that is really something else then fails loudly instead of being decoded as whatever it actually is.

Checking Before Decoding

An image costs four bytes per pixel once decoded, so a 6000x4000 photograph occupies 96 MB in memory however small its file was. For anything arriving from outside the program, read the header first:

import imagine
import imagine { Image }

var header = imagine.probe(upload)

if header == nil {
  raise Exception('that is not an image')
}

if header.width * header.height > 40000000 {
  raise Exception('image is too large')
}

var photo = imagine.decode(upload)

probe() returns {format, width, height} or nil. It reads only the header, which costs microseconds where decoding the same file costs tens of milliseconds. imagine.detect() is the same check when you only want the format name.

Saving

photo.save('out.png')
photo.save('out.jpg', { quality: 90 })
photo.save('out.dat', { format: 'webp' })   # extension overridden

The format comes from the extension unless format says otherwise. To get the bytes instead of a file:

var data = photo.encode('webp')
var data = photo.to_png()
var data = photo.to_jpeg(90)

encode() with no format uses whatever the image was decoded from, falling back to PNG for an image built in memory.

Format Reference

FormatReadWriteAlphaNotes
PNGyesyesyesLossless. The safe default.
JPEGyesyesnoLossy. Photographs only.
WebPyesyesyesWritten lossless, so smaller than PNG but larger than a lossy WebP.
AVIFnoyesyesSmallest files, slowest to encode.
GIFyesyes1-bit256 colours. The only animated format that can be written.
BMPyesyesyesUncompressed and enormous.
TIFFyesyesyesCommon in printing and scanning.
TGAyesyesyesNo signature; needs an extension or an explicit format.
QOIyesyesyesLossless, several times faster than PNG.
ICOyesyesyesIcon container.
PNMyesyesnoTrivially simple, trivially large.
WBMPyesyesnoOne bit per pixel.

Three entries need explaining.

AVIF is write-only: opening one raises DecodeError.

TGA has no magic number, so it can only be identified by its file extension or by naming the format outright.

WebP is written losslessly. Reading handles both lossy and lossless WebP, but the encoder here only writes lossless, so quality has no effect on it and a photograph saved as WebP will be larger than one saved by a tool that can write lossy WebP. For a photograph where size matters, JPEG or AVIF is the better target; WebP here is a smaller-than-PNG lossless format with alpha.

Never assume; ask:

import imagine

var can = imagine.capabilities()

echo can.decode.contains('png')
echo can.encode.contains('webp')
echo can.animated.contains('gif')
true
true
true

decode, encode and animated are each a list of format names, so contains() answers the question you actually have.

Encoding Options

OptionFormatsMeaning
qualityJPEG, AVIF1 to 100. Defaults to 85 for JPEG, 80 for AVIF.
compressionPNG'fast', 'default' or 'best'. All lossless.
speedAVIF, GIFAVIF 1-10, lower is slower and smaller. GIF 1-30, lower is slower and picks better colours; 15 by default.
backgroundJPEGWhat transparent pixels are flattened against. Defaults to white.
thresholdWBMPThe brightness cut, 0 to 255.

JPEG has no alpha channel, so transparency has to go somewhere on the way out. It is flattened against background rather than silently dropped, because dropping it turns transparent pixels black:

logo.save('logo.jpg', { background: '#ffffff' })

To see the result first, or to choose the colour once and keep it, do it explicitly:

logo.flatten('#ffffff').save('logo.jpg')

Data URLs

var url = icon.to_data_url('png')
# 'data:image/png;base64,iVBORw0...'

Base64 costs a third more bytes than the raw image, so this suits icons and small graphics rather than photographs.

Resizing and Reshaping

Choosing a Resize

There are six, and picking the right one is most of the work:

MethodAspect ratioEnlarges?Result
resize(w, h)not keptyesexactly w x h, possibly distorted
scale(factor)keptyesthe same image, scaled
thumbnail(w, h)keptnofits inside the box
fit(w, h)keptyesfits inside the box
cover(w, h)keptyesfills the box, cropping the overflow
contain(w, h)keptyesfits the box, padding the remainder

thumbnail() is the one you usually want for thumbnails, and the refusal to enlarge is the reason: scaling a small image up to fill a thumbnail box only makes it blurry. Use fit() when you do want it scaled up.

cover() and contain() both produce exactly the size asked for. They differ in what they sacrifice — cover() loses part of the image, contain() adds bars:

photo.cover(300, 300)                              # crops to a square
photo.cover(300, 300, { anchor: TOP })             # keeps the top, crops the bottom
photo.contain(300, 300, { background: 'white' })   # letterboxes instead

The same 320x180 image asked for a 120x120 result four ways:

thumbnail() keeps the proportions and does not fill the box. cover() fills it and loses the sides. contain() fills it and adds bars. resize() fills it by distorting the picture.

The anchors are TOP_LEFT, TOP, TOP_RIGHT, LEFT, CENTER, RIGHT, BOTTOM_LEFT, BOTTOM and BOTTOM_RIGHT. CENTER is the default, and TOP is what you want for photographs of people, where the face is rarely in the bottom third.

Resampling Filters

FilterSpeedUse
NEARESTfastestPixel art, QR codes, anything whose edges must stay hard.
BILINEARfastWhen speed matters more than sharpness.
BICUBICmoderateA good default for photographs.
GAUSSIANmoderateDeliberately soft; for noisy input, or before sharpening.
LANCZOSslowestThumbnails and any large reduction. The default.
photo.thumbnail(200, 200, LANCZOS)
sprite.scale(4, NEAREST)             # keeps pixel art crisp

A 16x16 sprite enlarged seven times over, so the differences are visible at all:

This is the case where NEAREST is right and everything else is wrong. Shrinking a photograph inverts that judgement entirely.

LANCZOS is the default because most resizing is shrinking, and that is where it earns its cost. It can produce faint ringing next to very high-contrast edges, which is the price of its sharpness; BICUBIC is the fallback when that shows.

Cropping, Padding and Trimming

photo.crop(100, 50, 400, 300)      # x, y, width, height
photo.pad(20)                      # 20 pixels on every side
photo.pad(10, 20, 10, 20, 'white') # top, right, bottom, left, colour
photo.trim()                       # remove a uniform border
photo.trim({ tolerance: 8 })       # allow for JPEG noise in the border

crop() requires its rectangle to fit inside the image and raises BoundsError otherwise. Clamping a crop that runs off the edge would hand back different dimensions than were asked for, which is a worse surprise than an error. Pad first when the region really is meant to extend past the edge.

trim() takes the border colour from the top-left pixel unless you name one. Scanned documents and screenshots almost always want a tolerance, because a “white” border from a lossy format is not exactly white.

Rotating and Mirroring

photo.rotate_90()                  # lossless
photo.rotate_180()
photo.rotate_270()
photo.rotate(37, 'white')          # resamples, grows the canvas

photo.flip_horizontal()
photo.flip_vertical()
photo.flip(FLIP_BOTH)
photo.transpose()                  # reflect across the main diagonal

Quarter turns are exact: every pixel is moved, none is resampled. Any other angle interpolates, and grows the canvas to hold the rotated corners with background filling the gaps.

EXIF Orientation

Phone cameras usually store the sensor’s orientation in the file rather than rotating the pixels, so a photograph whose bytes are sideways is meant to be displayed upright. Image.open() and Image.decode() handle this for you.

Image.open('photo.jpg')                       # upright
Image.open('photo.jpg', { orient: false })    # exactly as stored

All eight orientation values are handled, including the four mirrored ones that scanners and some front-facing cameras produce.

Filters

Filters change the image in place and return it, so they chain.

photo.grayscale().contrast(15).sharpen(0.5)

Tone and Exposure

photo.brightness(20)         # -255 to 255, added to each channel
photo.contrast(15)           # -100 to 100, around mid-grey
photo.gamma(1.2)             # above 1 lifts midtones, below 1 lowers them
photo.levels(20, 235)        # stretch this input range to full scale
photo.levels(20, 235, 1.1)   # ...with a midtone curve
photo.threshold(128)         # every pixel to black or white
photo.posterize(6)           # six evenly spaced steps per channel
photo.invert()               # a photographic negative
photo.opacity(0.5)           # scale the alpha channel

levels() is the single most useful correction for a flat or washed-out photograph. Everything at or below the black point becomes black, everything at or above the white point becomes white, and the range between is stretched to fill the scale.

Note the difference between posterize() and quantize(): posterize() spaces its levels evenly, quantize() picks the colours to suit the image, so a photograph survives far fewer of them.

photo.quantize(32)                # 32 well-chosen colours
photo.quantize(16, true)          # ...with dithering

auto_levels() does the same job as levels() without being told where the endpoints are; it is shown alongside the detail filters below.

Colour

photo.grayscale()
photo.sepia()
photo.saturate(1.4)               # 0 removes colour, 1 is unchanged
photo.hue_rotate(45)
photo.tint('#ff8800', 0.3)        # blend a flat colour in
photo.colorize('#4f46e5')         # monochrome in one colour
photo.duotone('#1e1b4b', '#fbbf24')
photo.flatten('white')            # composite over a colour, drop alpha

grayscale() weights the channels for perceived brightness, so a bright yellow comes out light and a deep blue comes out dark. That is different from saturate(0), which keeps HSL lightness and makes both mid-grey.

Blur, Sharpen and Detail

photo.blur(3)                # Gaussian; the argument is its sigma
photo.sharpen(0.8)
photo.smooth()               # a cheap, harsher 3x3 average
photo.emboss()
photo.edges()
photo.mean_removal()
photo.pixelate(12)

blur() is a true separable Gaussian, so a large radius costs linearly rather than quadratically. Its argument is the standard deviation: about two thirds of each pixel’s contribution falls within that distance, and the visible spread is roughly three times it.

Alpha is premultiplied for the duration of a blur, so blurring a shape on a transparent background does not drag a dark halo into its edge.

Writing Your Own Filter

Nearly every filter above is one of three primitives with different numbers in it, and all three are available directly.

A lookup table covers any per-channel tone curve. The table is built once whatever the image’s size, then applied at memory speed.

import imagine { filters }

# A gentle S-curve: more contrast, but the highlights survive.
var curve = filters.build_lut(@(value) {
  var t = value / 255
  return 255 * t * t * (3 - 2 * t)
})

photo.apply_lut(curve, curve, curve, nil)   # nil leaves alpha alone

A colour matrix covers anything that mixes channels — saturation, hue rotation, channel swaps, tinting. It is the 4x5 matrix SVG and CSS filters use, written row by row, with the last entry of each row a constant in 0-255 units.

# Swap the red and blue channels.
photo.apply_matrix([
  0, 0, 1, 0, 0,
  0, 1, 0, 0, 0,
  1, 0, 0, 0, 0,
  0, 0, 0, 1, 0,
])

Applying several matrices in a row is both slower and less accurate than combining them and applying the result once:

var both = filters.combine_matrices(
  filters.grayscale_matrix(),
  filters.saturation_matrix(1.2)
)

photo.apply_matrix(both)

A convolution kernel covers anything that reads a pixel’s neighbours.

photo.convolve([
  0, -1, 0,
  -1, 5, -1,
  0, -1, 0,
])

The kernel must be square with an odd side. The default divisor of 0 means “divide by the kernel’s own sum”, which is what nearly every published kernel expects.

One trap is worth knowing. Alpha goes through the kernel along with the colour channels, which is right for a blur and wrong for anything whose weights do not sum to one. A Laplacian over a uniformly opaque image sums to zero, which would make the whole result invisible:

photo.convolve(filters.edge_kernel(), { divisor: 1, keep_alpha: true })

The built-in edges(), emboss(), sharpen() and mean_removal() already do this.

The S-curve and the channel-swap matrix from this section, run against the same picture:

edge controls what happens off the image’s border: EDGE_CLAMP (the default) repeats the nearest edge pixel, EDGE_TRANSPARENT treats the outside as empty, and EDGE_WRAP wraps to the opposite side for images meant to tile.

Drawing

Shapes

Coordinates start at the top-left corner. Integer coordinates fall on pixel corners rather than centres, so a rectangle from (0, 0) to (10, 10) covers exactly the first ten pixels in each direction.

image.pixel(x, y, color)                        # one pixel, blended
image.line(x1, y1, x2, y2, color)
image.rect(x, y, w, h, color)                   # outline
image.fill_rect(x, y, w, h, color)
image.rounded_rect(x, y, w, h, radius, color)
image.fill_rounded_rect(x, y, w, h, radius, color)
image.circle(cx, cy, r, color)
image.fill_circle(cx, cy, r, color)
image.ellipse(cx, cy, rx, ry, color)
image.fill_ellipse(cx, cy, rx, ry, color)
image.arc(cx, cy, rx, ry, start, end, color)
image.pie(cx, cy, rx, ry, start, end, color)    # closed through the centre
image.chord(cx, cy, rx, ry, start, end, color)  # closed along the chord
image.bezier(x1, y1, cx1, cy1, cx2, cy2, x2, y2, color)

Angles are in degrees, measured clockwise from three o’clock, matching the direction the y axis runs.

Drawing outside the image is never an error; anything that falls outside is clipped away. That is what makes it safe to draw a shape that only partly overlaps.

Strokes and Anti-aliasing

image.thickness(4)         # applies to every subsequent stroke
image.antialias(false)     # hard edges

Both are settings on the image rather than per-call arguments, though a single call can override the thickness:

image.line(0, 0, 100, 100, 'black', { thickness: 8 })

Strokes are centred on the path, so a thickness of 4 puts 2 pixels on each side. Joints and the ends of an open path are rounded.

Anti-aliasing is on by default. Turn it off for output that has to be pixel-exact — barcodes, QR codes, anything that will be thresholded afterwards.

Stroke widths of 1, 3, 6 and 12:

Line Caps

How an open stroke finishes at its two ends is a separate choice from its width:

image.cap(CAP_ROUND)     # a half-disc. The default.
image.cap(CAP_SQUARE)    # a square, reaching the same distance
image.cap(CAP_BUTT)      # stops dead at the endpoint

The red marks are the exact coordinates the line was given. A round or square cap reaches half the stroke’s width past them; a butt cap does not.

CAP_BUTT is the one to reach for whenever the coordinates have to mean exactly what they say — segments meeting end to end, a scale bar of a known length, the pieces of a dashed line. CAP_SQUARE gives the same reach as round with a blunt finish.

A single call can override the surface’s setting:

image.line(40, 200, 40, 40, '#334155', { thickness: 6, cap: CAP_BUTT })

Joins between a path’s segments are always round, and closed outlines have no ends, so neither is affected by this.

Paths and Polygons

Every filled shape in this module becomes a polygon and goes through one scanline rasterizer, and every outline becomes the polygon around its stroke and goes through the same one. A shape not listed above can be drawn by supplying the points:

image.fill_polygon([[10, 10], [90, 30], [50, 80]], '#4f46e5')
image.polygon([[10, 10], [90, 30], [50, 80]], 'black')     # outline, closed
image.polyline([[10, 10], [90, 30], [50, 80]], 'black')    # not closed

Points may be a list of [x, y] pairs or a flat list of alternating values; both read naturally depending on where the points came from. The outline is closed for you, so the last point does not need to repeat the first.

Self-intersecting outlines are filled by the non-zero winding rule, which fills a five-pointed star solid. Pass { even_odd: true } for the other convention, which leaves its middle empty.

Filling Areas

image.fill('#0f172a')          # every pixel, ignoring the clip
image.clear()                  # every pixel to transparent
image.flood_fill(x, y, color)
image.flood_fill(x, y, color, 12)   # with a tolerance

Gradients get their own section below.

fill() replaces rather than blends, so filling with a transparent colour empties the image instead of leaving it unchanged.

Flood fill spreads four-connected — up, down, left and right, but not diagonally — through pixels within tolerance of the colour at the starting point. A tolerance of 0 spreads only through exactly equal pixels; on a photograph or anything anti-aliased you will want more.

Gradients

image.linear_gradient(x1, y1, x2, y2, stops)
image.radial_gradient(cx, cy, radius, stops)

A linear gradient runs along the vector from the first point to the second. Everything before that vector takes the first stop’s colour and everything past it takes the last stop’s, so a short vector across a large area gives a hard transition with flat bands on either side.

# top to bottom
image.linear_gradient(0, 0, 0, image.height(), ['#0f172a', '#334155'])

# diagonal, three colours
image.linear_gradient(0, 0, 400, 200, ['#4f46e5', '#f472b6', '#fbbf24'])

Stops are either bare colours, spaced evenly, or [offset, colour] pairs with the offset running 0 to 1:

image.linear_gradient(0, 0, 200, 0, [
  [0, 'black'],
  [0.25, 'red'],
  [1, 'white'],
])

By default a gradient replaces what is there. Pass { blend: true } to composite it instead, which is what makes overlays and vignettes work:

# a vignette over an existing photograph
photo.radial_gradient(
  photo.width() / 2,
  photo.height() / 2,
  photo.width() * 0.7,
  [[0, Color(0, 0, 0, 0)], [1, Color(0, 0, 0, 180)]],
  { blend: true }
)

# a caption scrim along the bottom
photo.linear_gradient(0, photo.height() - 120, 0, photo.height(), [
  Color(0, 0, 0, 0),
  Color(0, 0, 0, 200),
], { blend: true })

{ rect: {x, y, width, height} } confines the fill to a rectangle instead of covering the whole surface.

Both of those are the snippets above, run against the picture on the left.

Clipping

image.clip(20, 20, 100, 100)
image.fill_rect(0, 0, 500, 500, 'red')   # only the clip is painted
image.clear_clip()

A clip confines every drawing operation until it is cleared. It does not affect reading: get_pixel() sees the whole image either way.

Text

The Built-in Font

imagine ships no font file, so everything in the next section depends on what is installed on the machine and can fail. One font always works:

import imagine { Image, StrokeFont }

Image(300, 70, 'white')
  .text(16, 20, 'Always available', StrokeFont(28), '#0f172a')
  .save('label.png')

Font.builtin(size) is the same thing under a name you will find from Font.

It is not a font file. Every glyph is defined as geometry — centre-line strokes rather than filled outlines — in libs/imagine/strokefont.zu, and drawn through the same anti-aliased rasterizer as everything else. So it scales cleanly to any size, and its weight is a parameter rather than part of the design:

StrokeFont(24)              # regular
StrokeFont(24).weight(0.13) # bold
StrokeFont(24).weight(0.05) # light

Here is the whole thing:

What it is for. Labels, chart axes, watermarks, diagrams, placeholder text, and any output that has to work on a machine with no fonts installed. The look is geometric and single-weight, closer to a technical drawing than to a typeface, because that is what centre-line strokes give you honestly.

What it is not for. Body text, headlines, or anything where the shapes themselves matter. Load a real font for those.

Coverage is printable ASCII, from space through ~. Anything else draws the empty box a font uses for a glyph it does not have, so text in another script comes out visibly missing rather than silently blank.

A StrokeFont has the same interface as a Font — size(), metrics(), measure(), render(), wrap() — so anywhere a font is accepted, either works.

Loading a Font

import imagine { Font }

var font = Font.load('assets/Inter.ttf', 24)
var font = Font.from_bytes(embedded_font, 24)
var font = Font.system('DejaVu Sans', 24)
var font = Font.sans(24)

Font.load() reads a TrueType or OpenType file and caches the parsed face by path, so loading the same file twice in one process parses it once. Font.system() searches the platform’s font directories by family name, matching ignoring case, spaces and hyphens.

Font.sans() finds whichever common sans-serif font is installed. It is a convenience for scripts and tests, not something to rely on for output that must look the same everywhere: which font it lands on depends on the machine, and a container image with no fonts installed has none to find. It raises FontError in that case, pointing at Font.builtin(), which never depends on what is installed.

A Font is immutable and cheap to copy. size() returns the same face at another size, sharing the parsed data:

var title = font.size(32)
var body = font.size(14)

Drawing Text

image.text(20, 20, 'Hello', font, '#111111')

x and y are the top-left corner of the text’s box, not its baseline, because the corner is what you know when placing text in a layout. Pass { baseline: true } when you want y to mean the first line’s baseline instead.

A \n starts a new line.

card.text(24, 24, 'Quarterly report\n2026', title, '#111111', {
  align: ALIGN_CENTER,
  line_height: 1.4,
  tracking: 0.5,
})

align positions the lines against each other, not against the image. line_height is a multiplier on the font’s own recommended spacing, and tracking adds pixels between characters.

Measuring and Positioning

var box = image.text_size('Hello', font)
# { width: 58, height: 28, baseline: 22.3, lines: 1 }

Measuring costs a fraction of drawing, so it is the right way to lay text out before committing to it — centring, wrapping, or sizing a background to fit.

var box = card.text_size(label, font)

card.fill_rounded_rect(16, 16, box.width + 24, box.height + 16, 8, '#1e293b')
card.text(28, 24, label, font, 'white')

To position by anchor instead of coordinates:

card.place_text('SOLD OUT', font, 'white', CENTER)
card.place_text('v2.1', font, '#94a3b8', BOTTOM_RIGHT, { margin: 12 })

Wrapping

var body = Font.load('assets/Inter.ttf', 16)

page.text(40, 120, article, body, '#334155', { width: 520 })

The width option wraps the text to that many pixels before drawing it. text_size() takes the same option, so measuring wrapped text gives the box it will actually occupy.

To get the broken text itself — to store it, or to draw it in pieces — call wrap() on the font:

var lines = body.wrap(article, 520).split('\n')

echo '${lines.length()} lines'

Newlines already in the text are kept as paragraph breaks, and runs of spaces are collapsed. Words are kept whole where they can be; a single word too long for the width is broken between characters rather than allowed to overflow, since text spilling out of an image cannot be scrolled to.

What Text Layout Does Not Do

Glyphs are positioned by advance width with kerning applied. That covers Latin, Greek, Cyrillic and anything else written left to right without contextual shaping.

It does not do complex shaping. Arabic letters will not join, Indic clusters will not reorder, and ligatures are not substituted. Those need a shaping engine, and a standard library that pretended to do them would be worse than one that says plainly it does not.

Compositing

Drawing One Image Onto Another

photo.draw_image(logo, 20, 20)
photo.draw_image(logo, 20, 20, { opacity: 0.6 })
photo.place(logo, BOTTOM_RIGHT, { margin: 16, opacity: 0.5 })

The source is clipped to the destination, and its position may be negative, so a sprite hanging off the top-left corner draws correctly. Compositing an image onto itself works; it is copied first.

Blend Modes

photo.draw_image(texture, 0, 0, { blend: BLEND_MULTIPLY })
ModeEffect
BLEND_NORMALOrdinary alpha compositing. The default.
BLEND_MULTIPLYNever lighter than either input. Shadows.
BLEND_SCREENNever darker than either input. Glows.
BLEND_OVERLAYMultiply in the shadows, screen in the highlights.
BLEND_DARKEN / BLEND_LIGHTENKeep the darker or lighter channel.
BLEND_COLOR_DODGE / BLEND_COLOR_BURNStrong brighten or darken.
BLEND_HARD_LIGHT / BLEND_SOFT_LIGHTOverlay with the roles swapped; and a gentler version.
BLEND_DIFFERENCE / BLEND_EXCLUSIONAbsolute difference; and a lower-contrast variant.
BLEND_ADD / BLEND_SUBTRACTAdd or subtract, clamped.

The formulas are the ones in the CSS compositing specification, which is also what every image editor implements.

BLEND_DIFFERENCE makes a quick visual diff: identical images blended this way come out black.

A red circle drawn onto a gradient under eight of the modes:

var diff = before.clone().draw_image(after, 0, 0, { blend: BLEND_DIFFERENCE })

Drawing operations — lines, shapes, text — always use ordinary alpha compositing. To draw a shape under a blend mode, draw it on a layer and composite the layer.

Masks

var stencil = photo.layer()
stencil.fill_circle(200, 200, 150, 'white')

photo.mask(stencil)     # everything outside the circle becomes transparent

Each pixel keeps its colour and takes its transparency from the mask: opaque white shows the image through fully, black or transparent hides it, and greys give partial transparency. The mask’s own alpha counts, so a shape drawn on a transparent background works as a mask without being filled in first.

The mask must be the same size as the image.

Layers

layer() gives a transparent image of the same size. Building a composite out of layers is how you get effects that a single pass cannot, and it is the answer whenever you want a blend mode or an opacity applied to a group of operations rather than one:

var glow = photo.layer()

glow.fill_circle(200, 150, 80, '#fbbf24')
glow.blur(30)

photo.draw_image(glow, 0, 0, { blend: BLEND_SCREEN, opacity: 0.7 })

Inspecting an Image

photo.width()          # dimensions
photo.height()
photo.size()           # { width, height }
photo.bounds()         # { x: 0, y: 0, width, height }
photo.format()         # what it was decoded from, or nil
photo.is_opaque()      # walks the alpha channel
photo.info()           # all of the above, plus a pixel count

Beyond the shape, four methods read what is actually in the image.

histogram() counts how many pixels hold each value, as four 256-entry lists: red, green, blue and luma. Fully transparent pixels are skipped, since their colour is not visible. A histogram is what tells you an image is underexposed (everything bunched at the low end), flat (bunched in the middle), or clipped (a spike at 0 or 255).

auto_levels() is the automatic form of levels(): it finds the black and white points from the brightness histogram and stretches the range between them, ignoring a small fraction at each end so a handful of stray pixels cannot decide the result.

photo.auto_levels()
photo.auto_levels({ clip: 0.02, gamma: 1.1 })

average_color() returns the mean colour, weighting each pixel by its alpha so a mostly transparent image reports the colour of the part you can see.

dominant_colors() returns the colours occupying the most of the image, most common first. The image is shrunk and reduced to a small palette first, so it costs about the same whatever the original size.

var accent = photo.dominant_colors(1)[0]

page.fill(accent.darken(40))       # a placeholder while the photo loads

difference() scores how far two images are apart, from 0 (identical) to 1, which makes it usable as a rendering assertion:

if rendered.difference(expected) > 0.01 {
  raise Exception('the rendering changed')
}

Animation

import imagine { Animation }

var animation = Animation.open('loading.gif')

animation.length()      # frame count
animation.duration()    # milliseconds for one pass
animation.frame(0)      # one frame, as an Image
animation.frames()      # the live list

Frames come back already composited against whatever preceded them, so frame 5 can be used on its own without replaying the first four. The disposal and transparency rules an animated GIF is built from never surface.

Building and transforming:

var frames = []

iter var i = 0; i < 30; i++ {
  var frame = Image(200, 200, 'black')
  frame.fill_circle(100, 100, i * 3, '#38bdf8')
  frames.append(frame)
}

Animation(frames, 40, 0).save('pulse.gif')     # 40ms per frame, loops forever

animation
  .map(@(frame) {
    return frame.grayscale().blur(1)
  })
  .repeat(3)
  .save('out.gif')

map() copies each frame before the function sees it, so a filter chain works directly as the body and the original animation is left alone. That copy is why map() costs as much memory again as the animation itself; to filter in place, walk frames() and change each image directly.

Animated GIF and animated WebP can both be read, and only GIF can be written. A still image decodes as a one-frame animation rather than an error, so code handling both does not need to branch.

Every frame must be the same size, since animated formats have one canvas that each frame paints into.

Working With Pixels Directly

pixels() hands back the live buffer. It is 8-bit RGBA with straight (not premultiplied) alpha, laid out row by row with no padding, so pixel (x, y) begins at byte (y * width + x) * 4.

var buffer = image.pixels()
var total = buffer.length()

iter var at = 0; at < total; at += 4 {
  buffer[at] = 255 - buffer[at]        # invert red only
}

This is deliberate, and it is much faster than a method call per pixel, so an operation this module does not provide can still be written efficiently in Zuri. Annotating a function’s parameters pays off here:

def darken_edges(pixels: bytes, width: number, height: number) {
  iter var y = 0; y < height; y++ {
    iter var x = 0; x < width; x++ {
      var at = (y * width + x) * 4
      # ...
    }
  }
}

set_pixels() takes a buffer back, which is how you save a copy before a destructive filter and restore it afterwards:

var saved = image.pixels().clone()

image.blur(8)
image.set_pixels(saved)     # back where it started

Image.from_pixels() builds an image around a buffer that came from somewhere else entirely.

Errors

Failures fall into two groups. An argument of the wrong type raises TypeError, from the parameter’s own type declaration. Everything else descends from ImageError, so one catch covers it:

import imagine { Image, ImageError, DecodeError }

catch {
  var photo = Image.open(path)
} as e {
  if instance_of(e, DecodeError) {
    echo 'not a readable image'
  } else {
    echo 'something else went wrong: ${e.message}'
  }
}
ErrorRaised when
ImageErrorThe base class. Also raised directly for bad arguments.
DecodeErrorThe data is not an image, is a format this build cannot read, or is truncated.
EncodeErrorThe image cannot be written in the requested format.
FormatErrorA format name or file extension is not one this module knows.
BoundsErrorA rectangle, crop or resize falls outside the image, or a dimension is below 1.
FontErrorA font cannot be parsed, found, or laid out with.

Two kinds of failure sit outside that hierarchy on purpose.

Wrong argument type raises TypeError. Parameters declare their types, so the check happens at the boundary and the message names the parameter:

image.rotate('sideways')
# TypeError: rotate() expects parameter 'degrees' (argument 1)
#            to be a number, got string

A malformed colour raises ValueError, because that is what [[colors]] reports for it and relabelling would lose the distinction between “not a colour” and “not a string”:

Color.hex('nonsense')        # ValueError, from colors
Color.named('chartroose')    # ValueError, from colors
Color.hex(42)                # TypeError, from the type declaration

A value of the right type but the wrong range is still an ImageError subclass, since that is a judgement this module makes rather than a type the runtime can check:

Image(0, 100)                # BoundsError, not TypeError
filters.gamma_lut(-1)        # ImageError

DecodeError is the one to always be ready for, since anything arriving from outside the program can raise it.

Drawing operations do not raise BoundsError. A line running off the edge of the canvas is clipped, which is what every drawing API does; it is only the operations that must return an image of an exact size that have no sensible way to continue.

Performance and Memory

Four bytes per pixel, always. A 6000x4000 photograph is 96 MB decoded however small its file was, and a 100-frame 500x500 animation is 100 MB. probe() before decoding anything whose size you do not control.

Some rough guidance on what costs what:

  • Decoding and encoding dominate almost every pipeline. PNG at 'best' compression is several times slower than at 'default' for a few percent of size; AVIF is slower still.
  • Resizing costs roughly in proportion to the output size, so shrinking is cheap and enlarging is not. LANCZOS costs a few times NEAREST.
  • Blur is linear in its radius, not quadratic, because it is separable. A 3x3 convolve() is cheaper than blur(1), but by blur(5) the Gaussian has won by a wide margin.
  • Lookup tables and colour matrices are memory-bound and about as fast as touching every pixel can be. Chain as many as you like, but combine matrices with combine_matrices() rather than applying them one at a time — that is both faster and more accurate, since the intermediate result is never rounded back to 8 bits.
  • Drawing costs in proportion to the area covered, and anti-aliasing samples each pixel row four times over. Turning it off is a real saving on very large fills.

Resize before filtering whenever the result is going to be smaller anyway. Filtering a 24-megapixel photograph and then shrinking it to a thumbnail does the same visual work at forty times the cost.

Recipes

A thumbnail pipeline for uploads

import imagine
import imagine { Image, DecodeError, LANCZOS, TOP }

def make_thumbnail(upload) {
  var header = imagine.probe(upload)

  if header == nil {
    raise DecodeError('not an image')
  }

  if header.width * header.height > 50000000 {
    raise DecodeError('image too large')
  }

  return Image.decode(upload)
    .cover(400, 400, { anchor: TOP, filter: LANCZOS })
    .sharpen(0.4)
    .encode('webp')
}

A rounded avatar with a border

import imagine { Image }

def avatar(source, size) {
  var photo = Image.decode(source).cover(size, size)

  var stencil = photo.layer()
  stencil.fill_circle(size / 2, size / 2, size / 2 - 2, 'white')
  photo.mask(stencil)

  photo.circle(size / 2, size / 2, size / 2 - 2, '#e2e8f0', { thickness: 3 })

  return photo
}

A social card

import imagine { Image, Font, Color, ALIGN_LEFT }

def card(title, subtitle) {
  var image = Image(1200, 630, '#0f172a')
  var heading = Font.load('assets/Inter-Bold.ttf', 64)
  var body = heading.size(30)

  image.fill_rect(0, 0, 1200, 8, '#38bdf8')

  var box = image.text_size(title, heading, { align: ALIGN_LEFT })
  image.text(80, 200, title, heading, 'white', { align: ALIGN_LEFT })
  image.text(80, 200 + box.height + 24, subtitle, body, '#94a3b8')

  return image.to_png()
}

Serving a generated image over HTTP

import http
import imagine { Image, Font, CENTER }
import imagine.formats

var server = http.server(8000)
var font = Font.load('assets/Inter.ttf', 20)

server.get('/badge/{label}', @(request, response) {
  var image = Image(220, 60, '#1e293b')

  image.fill_rounded_rect(0, 0, 220, 60, 8, '#334155')
  image.place_text(request.param('label'), font, 'white', CENTER)

  response.content_type(formats.mime_for('png'))
  response.cache_for(86400)
  response.write(image.to_png())
})

server.listen()

Comparing two images

import imagine { BLEND_DIFFERENCE }

def differs(a, b) {
  if a.size() != b.size() {
    return true
  }

  var diff = a.clone().draw_image(b, 0, 0, { blend: BLEND_DIFFERENCE })
  var pixels = diff.pixels()
  var total = pixels.length()

  iter var at = 0; at < total; at += 4 {
    if pixels[at] > 8 or pixels[at + 1] > 8 or pixels[at + 2] > 8 {
      return true
    }
  }

  return false
}

Databases

The sql module is how a Zuri program talks to a relational database. It is one module rather than one per engine, and that is the whole design: sql defines what a database adapter has to provide, supplies everything that is the same whichever engine answers, and picks the adapter from the connection string.

Three adapters ship with it to support four major database engines — SQLite, PostgreSQL, MySQL and MariaDB. SQLite is a file-based database with no server to run and nothing to configure, which makes it the right choice for an application that ships with its data and for a test suite that wants a real database per run. PostgreSQL, MySQL, and MariaDB are servers, for everything that outgrows a file, and the MySQL adapter is the same adapter that drives MariaDB as well.

Changing from one to another means changing the connection string, and whatever SQL they genuinely spell differently.

Blocks on this page that list several calls together are reference listings, not programs: they show the shape of each call rather than a sequence to run. Anything presented as a complete program runs as written. The PostgreSQL and MySQL examples need a server, so they are shown rather than run.

Following Along

Most examples below use a small database of posts and their authors. This builds it:

import sql

var db = sql.open('sqlite://guide.db')

db.exec_script("
  drop table if exists posts;
  drop table if exists authors;

  create table authors (
    id integer primary key,
    name text not null unique
  );

  create table posts (
    id integer primary key,
    author_id integer not null references authors(id),
    title text not null,
    views integer not null default 0,
    published boolean not null default 0
  );
")

var ada = db.insert('authors', { name: 'Ada Lovelace' })
var grace = db.insert('authors', { name: 'Grace Hopper' })

db.insert_many('posts', [
  { author_id: ada, title: 'On Engines', views: 412, published: true },
  { author_id: ada, title: 'On Looms', views: 87, published: false },
  { author_id: grace, title: 'On Bugs', views: 1290, published: true },
  { author_id: grace, title: 'On Compilers', views: 640, published: true },
])

echo db.count('posts')
db.close()
4

Everything from here on assumes that file exists in the working directory.

Introduction

A database library usually ties a program to one engine. The calls are named after that engine, the placeholders are spelled its way, and the errors are its own numbers, so moving to another means rewriting every call site.

sql puts the parts that are the same in one place. A statement is written once:

import sql

var db = sql.open('sqlite://guide.db')

for post in db.query('select title, views from posts where published = ?', [true]) {
  echo '${post.title}: ${post.views}'
}

db.close()
On Engines: 412
On Bugs: 1290
On Compilers: 640

Point the first line at postgres://localhost/app or mysql://localhost/app and the rest runs unchanged. The ? becomes $1 on PostgreSQL because that is what PostgreSQL wants, and stays ? on MySQL because that is what MySQL wants. An insert asking for its new id gets a RETURNING clause on PostgreSQL, because PostgreSQL has no last insert id and the other two do. A duplicate key arrives as UniqueViolation from all three.

What it does not do is pretend the engines are the same. They disagree about auto-incrementing keys, about which functions exist, about a great deal of SQL. sql translates what can be translated and is explicit about the rest.

Connecting

open() takes a connection string and returns a Connection.

sql.open('sqlite://./app.db')          # a file
sql.open(':memory:')                   # a private database in memory
sql.open('./app.db')                   # a bare path is SQLite
sql.open('postgres://localhost/app')
sql.open('postgres://alice:secret@db.internal:5432/shop?sslmode=require')
sql.open('host=localhost dbname=app user=alice')
sql.open('mysql://localhost/app')
sql.open('mysql://alice:secret@db.internal:3306/shop?sslmode=verify')
sql.open('mariadb://localhost/app')

The scheme picks the adapter. A string with no scheme at all is taken as a SQLite path, since nothing else it could be.

Options can also be given as a dictionary, which is how anything an adapter accepts beyond the connection string is passed:

sql.open({
  driver: 'sqlite',
  path: './app.db',
  journal_mode: 'wal',
  busy_timeout: 10000,
})

or alongside a string, where they are merged over whatever it carried:

sql.open('sqlite://./app.db', { journal_mode: 'wal' })

A connection should be closed when it is finished with, and a catch block makes that certain even when the work between raises:

import sql

var db = sql.open(':memory:')

catch {
  db.exec('create table t (n integer)')
  db.exec('insert into t values (?)', [1])

  echo db.fetch_value('select n from t')
} as error {
  echo 'failed: ${error.message}'
}

db.close()
1

A connection belongs to the isolate that opened it and cannot be handed to another. An isolate that needs the database opens its own connection or its own pool.

Querying and Fetching

query() runs a statement and reads its whole result:

import sql

var db = sql.open('sqlite://guide.db')
var result = db.query('select title, views from posts order by views desc')

echo result.length()
echo result.first().title
echo result.column('title')

db.close()
4
On Bugs
[On Bugs, On Compilers, On Engines, On Looms]

A result is iterable, so the common case needs nothing else:

for post in db.query('select * from posts') {
  echo post.title
}

For the shapes that come up constantly there are shorter forms:

db.fetch_one(sql, params)             # the first row, or nil
db.fetch_all(sql, params)             # every row, as dictionaries
db.fetch_value(sql, params, fallback) # the first column of the first row
db.fetch_column(sql, params, column)  # one column's values, as a list
import sql

var db = sql.open('sqlite://guide.db')

echo db.fetch_value('select count(*) from posts')
echo db.fetch_one('select title from posts where views > ?', [1000]).title
echo db.fetch_column('select title from posts order by title limit 2')
echo db.fetch_value('select title from posts where views > ?', [99999], 'none')

db.close()
4
On Bugs
[On Bugs, On Compilers]
none

exec() is for a statement run for its effect rather than its rows:

import sql

var db = sql.open(':memory:')

db.exec('create table t (id integer primary key, n integer)')
var result = db.exec('insert into t (n) values (?)', [7])

echo result.rows_affected
echo result.last_insert_id

db.close()
1
1

And exec_script() runs several statements at once, which is what a schema file is:

db.exec_script(file('schema.sql').read())

Parameters

Values are bound by the engine, never pasted into the statement. That is what makes a value a value rather than a piece of SQL, and it is the whole of the defence against injection.

Write ? for positional parameters:

import sql

var db = sql.open('sqlite://guide.db')

echo db.fetch_column(
  'select title from posts where author_id = ? and published = ?',
  [1, true]
)

db.close()
[On Engines]

or :name for named ones, where order stops mattering and a name may be used more than once:

import sql

var db = sql.open('sqlite://guide.db')

echo db.fetch_column(
  'select title from posts where views > :floor and views < :ceiling',
  { floor: 100, ceiling: 1000 }
)

db.close()
[On Engines, On Compilers]

sql rewrites these into whatever the adapter wants: $1 and $2 for PostgreSQL, ?1 and ?2 for SQLite, and ? unchanged for MySQL, which already spells them that way. The scanner knows when a ? or a : is not a placeholder, so none of these are touched:

"a ? inside a string"      -- a string
"a ? in an identifier"     -- a quoted identifier
-- a ? in a line comment
/* a ? in a block comment */
$tag$ a ? in a dollar-quoted body $tag$
value::text                -- a PostgreSQL cast
arr[1:3]                   -- an array slice

PostgreSQL uses ? as a JSON operator, so a literal one is written ??:

db.query('select * from docs where data ?? ?', ['key'])

Never build a statement by joining strings around a value, even when the value looks safe. The two forms above cover every case where a value varies. Where an identifier varies, quote it through the driver rather than interpolating it:

var column = db.driver().quote_identifier(name)
db.query('select ${column} from posts')

Rows, Columns and Types

A row is a dictionary keyed by column name, which is what lets post.title read the way it does.

import sql

var db = sql.open('sqlite://guide.db')
var result = db.query('select id, title, views from posts order by id limit 1')

echo result.first()
echo result.columns
echo result.tuples()

db.close()
{id: 1, title: On Engines, views: 412}
[{name: id, type: integer}, {name: title, type: text}, {name: views, type: integer}]
[[1, On Engines, 412]]

Two columns of the same name collapse in a dictionary, which is worth knowing before selecting id from both sides of a join. Aliasing one of them is the fix; tuples() is the escape hatch when it cannot be.

Values correspond like this:

ZuriDatabase
nilNULL
boola boolean where the engine has one, otherwise 0 and 1
numberan integer where the value is whole, otherwise a float
biginta 64 bit integer, for values past what a double holds
stringtext
bytesa blob
date.Datea timestamp
sql.Timea TIME interval, on an engine that has one
list, dictJSON

Not every engine has every one of those. PostgreSQL has a real boolean and MySQL does not, so a true written to MySQL comes back as 1. db.supports() answers that sort of question without guessing, and each engine’s own section below says where it differs.

An integer too large for a double comes back as a bigint rather than silently rounded:

import sql

var db = sql.open(':memory:')
db.exec('create table t (n integer)')
db.exec('insert into t values (9223372036854775807)')

var n = db.fetch_value('select n from t')
echo typeof(n)
echo n

db.close()
bigint
9223372036854775807n

SQLite stores five things and remembers nothing about intent, so a boolean goes in as 1 and would come back as the number 1. What it does keep is the type each column was declared with, and that is what the adapter reads it back by:

import sql

var db = sql.open(':memory:')
db.exec('create table t (ok boolean, at datetime, doc json)')
db.exec('insert into t values (?, ?, ?)', [
  true, '2026-09-16T14:30:00.000000+00:00', '{"a":1}',
])

var row = db.fetch_one('select * from t')

echo typeof(row.ok)
echo typeof(row.at)
echo row.doc

db.close()
bool
Date
{a: 1}

A column with no declared type is an expression, and comes back exactly as it was stored.

Exact decimals

A number is a double, which holds 0.1 only approximately. For money that is not good enough, so sql.Decimal holds a value exactly:

import sql { Decimal }

var price = Decimal('19.99')
var tax = price.multiply(Decimal('0.20'))

echo price.to_string()
echo tax.to_string()
echo price.add(tax).to_string()
echo Decimal('0.1').add(Decimal('0.2')).to_string()
19.99
3.9980
23.9880
0.3

PostgreSQL’s numeric and MySQL’s DECIMAL columns read and write as Decimal automatically. SQLite has no exact decimal type, so store one as text or as an integer count of the smallest unit.

The CRUD Helpers

Four calls on a connection write no SQL at all. insert(), update(), delete() and find() take a table name and dictionaries, build the statement for whichever engine is on the other end, and bind every value as a parameter.

They are here because the statements they replace are the ones least worth writing by hand. They are mechanical, they differ between dialects in small ways that only show up in production, and a string built by concatenation is where an injection gets in. Anything harder than a flat list of conditions is still written as SQL, and the two mix freely on the same connection.

Inserting and ids

insert() takes a table and a dictionary, and returns the new row’s id:

import sql

var db = sql.open(':memory:')
db.exec('create table posts (id integer primary key, title text)')

echo db.insert('posts', { title: 'Hello' })
echo db.insert('posts', { title: 'World' })

db.close()
1
2

This is the call that hides the largest difference between the engines. SQLite and MySQL report the id of the row just inserted; PostgreSQL does not, and an insert that wants one has to ask with a RETURNING clause naming the primary key. MariaDB has both and uses RETURNING. insert() does whichever applies, and finds the key by asking the schema rather than assuming it is called id.

Where the key is not the column to return, name it:

db.insert('events', { name: 'started' }, { returning: 'uuid' })

Several rows at once

Several rows go in one statement, split so that no single statement binds more parameters than the engine allows:

import sql

var db = sql.open(':memory:')
db.exec('create table points (x integer, y integer)')

echo db.insert_many('points', [
  { x: 1, y: 2 },
  { x: 3, y: 4 },
])

db.close()
2

Every row has to name the same columns. A row naming a different set raises rather than being padded with nulls, because a missing column and a null column mean different things.

Finding rows

find() answers with a ResultSet, count() with a number, and find_one() with the first row as a dictionary:

import sql

var db = sql.open('sqlite://guide.db')

echo db.count('posts')
echo db.count('posts', { published: true })
echo db.count('posts', { author_id: [1, 2] })
echo db.count('posts', { author_id: [] })

echo db.find('posts', { published: true }, {
  columns: ['title'],
  order: ['views desc'],
  limit: 2,
}).column('title')

echo db.find_one('posts', { title: 'On Bugs' }).views

db.close()
4
3
4
0
[On Bugs, On Compilers]
1290

A find_one() that matches nothing answers nil. That is the answer, not a failure, so nothing is raised for it:

import sql

var db = sql.open('sqlite://guide.db')

echo db.find_one('posts', { title: 'On Bugs' }).views
echo db.find_one('posts', { title: 'Never Written' })

db.close()
1290
nil

What a filter can say

A condition’s value decides what it means. A plain value is equality, a list is IN, an empty list matches nothing, and nil is IS NULL rather than = NULL, which no row ever satisfies.

Several conditions in one dictionary are joined with AND. There is no OR, which is the first thing to write as SQL:

import sql

var db = sql.open('sqlite://guide.db')

echo db.count('posts', { published: true, author_id: 1 })
echo db.count('posts', { published: true, views: [87, 1290] })

db.close()
1
1

Passing nil as the whole filter matches every row. That has to be asked for rather than happening because a dictionary came out empty, which is what stands between a filter built from user input and an UPDATE with no WHERE.

Ordering, limits and pages

order is a list of column names, each optionally followed by asc or desc, applied in the order given:

import sql

var db = sql.open('sqlite://guide.db')

echo db.find('posts', nil, { order: ['views desc'] }).column('title')
echo db.find('posts', nil, { order: ['author_id asc', 'views desc'] }).column('title')

db.close()
[On Bugs, On Compilers, On Engines, On Looms]
[On Engines, On Looms, On Bugs, On Compilers]

limit and offset are whole numbers of rows, and together they are a page. A page past the end is empty rather than an error:

import sql

var db = sql.open('sqlite://guide.db')

def page(number, size) {
  return db.find('posts', nil, {
    columns: ['title'],
    order: ['views desc'],
    limit: size,
    offset: (number - 1) * size,
  }).column('title')
}

echo page(1, 2)
echo page(2, 2)
echo page(3, 2)

db.close()
[On Bugs, On Compilers]
[On Engines, On Looms]
[]

A limit and an offset go into the statement text rather than being bound, because not every engine allows a parameter in either place. That leaves them as the one part of a built statement that is not a parameter, so each is checked before it goes in:

import sql

var db = sql.open('sqlite://guide.db')

catch {
  db.find('posts', nil, { limit: 2.5 })
} as error {
  echo error.message
}

db.close()
a limit is a whole number of rows, not 2.5

Updating and deleting

update() and delete() return how many rows changed:

import sql

var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, n integer)')
db.insert_many('t', [{ n: 1 }, { n: 2 }, { n: 3 }])

echo db.update('t', { n: 0 }, { n: [1, 2] })
echo db.delete('t', { n: 0 })
echo db.count('t')

db.close()
2
2
1

They read a filter exactly as find() does, nil included: passing it changes or deletes every row.

Values that are not values

Where a value is not a value, sql.raw() marks a fragment to be used as written:

import sql

var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, views integer)')
db.insert('t', { views: 10 })

db.update('t', { views: sql.raw('views + 1') }, { id: 1 })
echo db.fetch_value('select views from t')

db.close()
11

A Raw is accepted everywhere a column or a value is, which is what makes an aggregate or a subquery reachable without leaving the helpers:

import sql

var db = sql.open('sqlite://guide.db')

echo db.find('posts', nil, {
  columns: [sql.raw('count(*) as n'), sql.raw('sum(views) as total')],
}).first()

echo db.count('posts', {
  author_id: sql.raw("(select id from authors where name = 'Ada Lovelace')"),
})

db.close()
{n: 4, total: 2429}
2

raw() is exactly as dangerous as it sounds. A fragment built from anything a user supplied is a SQL injection. Build them from literals, and keep values in the parameters where they belong.

Inside a transaction

A Transaction carries all of them, and they mean the same thing there. A read inside one sees what that transaction has written and nobody else has committed yet:

import sql

var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, n integer)')

db.transaction(@(tx) {
  tx.insert_many('t', [{ n: 1 }, { n: 2 }, { n: 3 }])

  echo tx.count('t')
  echo tx.find('t', nil, { order: ['n desc'] }).column('n')

  tx.update('t', { n: 9 }, { n: 1 })
  tx.delete('t', { n: 3 })
})

echo db.find('t', nil, { order: ['n asc'] }).column('n')

db.close()
3
[3, 2, 1]
[2, 9]

A rollback takes those rows with it, the reads included, and a transaction that has already committed refuses them the way it refuses everything else.

Where they stop

These calls cover a flat list of equality conditions and stop there. There is no join, no OR and no expression tree, because SQL is a better language for those than any chain of method calls. A query that wants one is a query() or a fetch_all() on the same connection, next to the helpers rather than instead of them.

Transactions

The closure form commits when the body returns and rolls back when it raises, and there is no path through it that leaves a transaction open:

import sql

var db = sql.open(':memory:')
db.exec('create table accounts (id integer primary key, balance integer)')
db.insert_many('accounts', [{ balance: 500 }, { balance: 0 }])

db.transaction(@(tx) {
  tx.update('accounts', { balance: sql.raw('balance - 100') }, { id: 1 })
  tx.update('accounts', { balance: sql.raw('balance + 100') }, { id: 2 })
})

echo db.fetch_column('select balance from accounts order by id')

db.close()
[400, 100]

A failure undoes the whole thing:

import sql

var db = sql.open(':memory:')
db.exec('create table t (n integer)')

catch {
  db.transaction(@(tx) {
    tx.exec('insert into t values (1)')
    raise Error('something went wrong')
  })
} as error {
  echo 'rolled back: ${error.message}'
}

echo db.count('t')
db.close()
rolled back: something went wrong
0

A transaction() called inside another becomes a savepoint, so a helper that opens one works the same whether it was called on its own or from inside a larger piece of work. Only the outermost commits, and a failure inside undoes just that inner piece:

import sql

var db = sql.open(':memory:')
db.exec('create table t (name text)')

db.transaction(@(outer) {
  outer.insert('t', { name: 'outer' })

  catch {
    db.transaction(@(inner) {
      inner.insert('t', { name: 'inner' })
      raise Error('inner failed')
    })
  } as _error {
    echo 'inner undone'
  }
})

echo db.fetch_column('select name from t')
db.close()
inner undone
[outer]

Schema changes are part of the transaction on SQLite and PostgreSQL, so a set of create table statements that fails half way leaves nothing behind. MySQL and MariaDB commit the open transaction at every create, alter and drop, and a transaction() whose body changes the schema there raises TransactionError when it comes to commit. db.supports('transactional_ddl') tells the two apart, so a program that migrates its own schema can run each change inside a transaction where that holds and on its own where it does not.

An isolation level can be asked for. Where an engine cannot provide one it says so rather than quietly giving something weaker:

db.transaction(@(tx) { ... }, sql.SERIALIZABLE)
sql.READ_UNCOMMITTEDMySQL honours it; PostgreSQL accepts it and gives read committed
sql.READ_COMMITTEDeverywhere
sql.REPEATABLE_READeverywhere, and MySQL’s default
sql.SERIALIZABLEeverywhere

begin() opens one to be committed or rolled back by hand, for a transaction whose lifetime is not a block. The closure form is safer and should be preferred.

Streaming Large Results

query() builds the whole result in memory, which is right for the hundreds of rows most queries return and wrong for the millions some do. stream() reads the same result a row at a time:

import sql

var db = sql.open('sqlite://guide.db')
var cursor = db.stream('select title from posts order by title')

for post in cursor {
  echo post.title
}

db.close()
On Bugs
On Compilers
On Engines
On Looms

Running to the end closes the cursor. A loop that stops early does not, so anything that might break has to close it:

var cursor = db.stream('select * from events')

for event in cursor {
  if done(event) {
    break
  }
}

cursor.close()

take(n) reads a batch at a time, for work that batches naturally:

import sql

var db = sql.open('sqlite://guide.db')
var cursor = db.stream('select title from posts order by title')

echo cursor.take(2).map(@(post) => post.title)
echo cursor.take(2).map(@(post) => post.title)
echo cursor.take(2)

db.close()
[On Bugs, On Compilers]
[On Engines, On Looms]
[]

On PostgreSQL this is a server-side portal and on MySQL a server-side cursor, both fetched in batches that { batch: n } sizes. On SQLite it is the engine’s own behaviour: rows are computed as they are asked for, so nothing is held but the current one.

A server-side cursor holds its connection while it is open, since the server is part way through answering. Running another statement on that connection before the cursor is read to the end or closed raises rather than letting the exchange fall out of step.

Prepared Statements

Compiling a statement is the expensive half of running one, and a statement run in a loop should pay for it once:

import sql

var db = sql.open(':memory:')
db.exec('create table points (x integer, y integer)')

var insert = db.prepare('insert into points (x, y) values (?, ?)')

for i in 0..(1000) {
  insert.exec([i, i * 2])
}

insert.close()

echo db.count('points')
db.close()
1000

The placeholders are translated once, when the statement is prepared, so the loop does no string work at all. A prepared statement carries the same query and fetch methods a connection does.

Connection Pools

Opening a connection is expensive: for PostgreSQL and MySQL it is a TCP connection, a TLS handshake and an authentication exchange before the first statement runs. A server that opened one per request would spend most of its time connecting.

import sql

var db = sql.pool('sqlite://guide.db', { max: 4 })

echo db.fetch_value('select count(*) from posts')
echo db.stats()

db.close()
4
{size: 1, idle: 1, in_use: 0, max: 4}

Used that way the pool takes a connection, runs the statement and gives it back. Where several statements have to run on the same connection, which a transaction requires, with_connection() holds one and releases it whatever happens:

db.with_connection(@(connection) {
  connection.transaction(@(tx) {
    tx.exec('...')
  })
})

transaction() on the pool does the same in one call.

Setting
maxmost connections to open. Ten by default
minhow many to open up front. None by default
idle_timeouthow long an idle connection is kept
max_lifetimehow long any connection is kept before being replaced
validate_on_acquirecheck a connection is alive before lending it
on_connectrun something on each new connection

A pool wants to be as small as the work allows. Connections are not free at the other end either, and a pool larger than the database can usefully serve turns a queue in the application into a queue in the database, where it is harder to see.

A pool belongs to the isolate that made it, which has two consequences. An isolate that needs database access makes its own. And there is no waiting when a pool is empty: nothing else can release a connection while the call is running, so running out is reported rather than waited on.

Errors

Every error is a sql.SqlError. Catch that to catch anything a database can do; the subclasses separate what is worth handling differently.

Error
└── SqlError
    ├── ConnectionError
    │   ├── AuthenticationError
    │   ├── TimeoutError
    │   └── ProtocolError
    ├── QueryError
    ├── IntegrityError
    │   ├── UniqueViolation
    │   ├── ForeignKeyViolation
    │   ├── NotNullViolation
    │   └── CheckViolation
    ├── TransactionError
    │   ├── SerializationError
    │   └── DeadlockError
    ├── PoolError
    │   └── PoolExhaustedError
    ├── NotSupportedError
    └── ClosedError

The classes are the same on every engine, which is the point. A unique constraint is SQLSTATE 23505 on PostgreSQL, error number 1062 on MySQL, and extended result code 2067 on SQLite; no one of those means anything to the others, and code matching on any of them would stop working the moment the adapter changed.

import sql

var db = sql.open(':memory:')
db.exec('create table users (id integer primary key, email text unique)')
db.insert('users', { email: 'ada@example.com' })

catch {
  db.insert('users', { email: 'ada@example.com' })
} as error {
  echo instance_of(error, sql.UniqueViolation)
  echo instance_of(error, sql.IntegrityError)
  echo instance_of(error, sql.SqlError)
  echo error.driver
}

db.close()
true
true
true
sqlite

The engine’s own account is kept rather than thrown away. code holds what the engine said, sqlstate the five character code where there is one, and query the statement that failed:

import sql

var db = sql.open(':memory:')
db.exec('create table users (id integer primary key, email text unique)')
db.insert('users', { email: 'ada@example.com' })

catch {
  db.insert('users', { email: 'ada@example.com' })
} as error {
  echo error.code
  echo error.sqlstate
  echo error.type
}

db.close()
2067
nil
UniqueViolation

SerializationError and DeadlockError are the two worth retrying: both mean the engine gave up on a transaction to keep its promises, and running the whole transaction again usually succeeds.

Schema Introspection

db.schema answers questions about what is in the database. The queries underneath are entirely different per engine, and the answers are the same shape.

Every method takes an optional schema to look in. PostgreSQL has schemas inside a database and resolves the default through search_path; MySQL calls a database a schema and has no layer above it, so naming one there names a database; SQLite has neither and ignores the argument.

import sql

var db = sql.open('sqlite://guide.db')

echo db.schema.tables()
echo db.schema.column_names('posts')
echo db.schema.primary_key('posts')
echo db.schema.has_column('posts', 'views')

db.close()
[authors, posts]
[id, author_id, title, views, published]
id
true

Each column is described the same way whichever engine answered:

import sql

var db = sql.open('sqlite://guide.db')

for column in db.schema.columns('posts') {
  echo '${column.name} ${column.type} ${column.nullable ? "null" : "not null"}'
}

db.close()
id integer null
author_id integer not null
title text not null
views integer not null
published boolean not null

indexes() and foreign_keys() describe the rest:

import sql

var db = sql.open('sqlite://guide.db')

echo db.schema.foreign_keys('posts')

db.close()
[{columns: [author_id], references_table: authors, references_columns: [id], on_update: NO ACTION, on_delete: NO ACTION}]

This is also what makes insert() portable: on an engine with no last insert id the insert needs a RETURNING clause, and the column to return is the table’s primary key, which is asked for here.

One caution when moving a schema between engines: MySQL parses a column level references clause and then ignores it, so a foreign key declared that way exists on the other two and not there. Declared at table level it exists on all three.

Switching Databases

The claim this module makes is that a program moves between engines by changing its connection string. Here is that claim, as a program:

import sql

def report(db) {
  db.exec_script('drop table if exists tally')
  db.exec_script(
    'create table tally (id ' + key_type(db) + ' primary key, word text not null)'
  )

  db.insert_many('tally', [{ word: 'alpha' }, { word: 'beta' }])

  var id = db.insert('tally', { word: 'gamma' })
  var found = db.fetch_column('select word from tally order by word')

  db.exec_script('drop table tally')

  return '${db.driver_name()}: ${found} (last id ${id})'
}

def key_type(db) {
  using db.driver_name() {
    when 'postgres' return 'serial'
    when 'mysql' return 'integer auto_increment'
    when 'mariadb' return 'integer auto_increment'
  }

  return 'integer'
}

var sqlite = sql.open(':memory:')

echo report(sqlite)
sqlite.close()
sqlite: [alpha, beta, gamma] (last id 3)

The same report() runs against PostgreSQL or MySQL by opening a different connection, and prints the same list.

Three things did have to be written twice, and they are the three that are genuinely different:

  • The connection string, which is the point.
  • An auto-incrementing key, which is not standard SQL. PostgreSQL spells it serial, MySQL auto_increment, SQLite integer primary key.
  • Any SQL only one of them has. sql does not parse or rewrite statements beyond their placeholders, so a PostgreSQL array operator, a MySQL on duplicate key update or a SQLite json_extract stays what it is.

db.supports() is how a program asks rather than assumes:

import sql

var db = sql.open(':memory:')

echo db.driver_name()
echo db.supports('returning')
echo db.supports('arrays')
echo db.supports('concurrent_writers')
echo db.capabilities().placeholder_style

db.close()
sqlite
true
false
false
indexed

SQLite Specifics

db.native() reaches the adapter underneath, where the engine’s own features live.

A SQLite database is a file, and :memory: is one that is not. An in-memory database belongs to the connection that opened it, so a second connection to :memory: is a second, empty database rather than another handle on the same data.

Foreign keys are enforced. SQLite leaves them unenforced by default, per connection, for compatibility with databases written before it had them; a database that declares them almost certainly means them, so the adapter turns them on. Pass foreign_keys: false to turn them back off.

Functions written in Zuri

A function registered on a connection runs inside the engine, once per row, and can be used anywhere an expression can:

import sql

var db = sql.open('sqlite://guide.db')

db.native().create_function('initials', 1, @(name) {
  return ''.join(name.split(' ').map(@(part) => part[0, 1]))
})

echo db.fetch_column('select initials(name) from authors order by name')

db.close()
[AL, GH]

An aggregate is written as a fold: step is called once per row with the accumulator and the row’s arguments and returns the next accumulator, and finish turns the last accumulator into the group’s value.

import sql

var db = sql.open('sqlite://guide.db')

db.native().create_aggregate('longest', 1, @(longest, title) {
  if longest == nil or title.length() > longest.length() {
    return title
  }

  return longest
}, @(longest) => longest)

echo db.fetch_value('select longest(title) from posts')

db.close()
On Compilers

A collation is an ordering for text, reached with collate:

import sql

var db = sql.open('sqlite://guide.db')

db.native().create_collation('bylength', @(a, b) => a.length() - b.length())

echo db.fetch_column('select title from posts order by title collate bylength')

db.close()
[On Bugs, On Looms, On Engines, On Compilers]

A function must not touch the connection it was registered on. SQLite is in the middle of a statement when it calls, and reentering would deadlock. An error raised inside one is carried out and re-raised with its own class intact once the statement has finished.

Blobs, backups and hooks

A large value read with a select arrives all at once. A blob handle opens one cell and reads windows of it:

var blob = db.native().blob('main', 'files', 'content', id, false)

var at = 0
while at < blob.length() {
  out.write(blob.read(at, 65536.min(blob.length() - at)))
  at += 65536
}

blob.close()

The cell has to exist and already be the right size: SQLite cannot grow a blob this way, which is what inserting zeroblob(n) is for.

Copying the file with the filesystem is only safe when nothing is writing. SQLite’s own backup copies page by page and notices when the source changes underneath:

db.native().backup_to('./snapshot.db')

And the hooks report what the engine is doing:

db.native().on_change(@(operation, database, table, rowid) { ... })
db.native().on_commit(@() => true)
db.native().on_rollback(@() { ... })
db.native().set_authorizer(@(action, first, second, database, trigger) { ... })
db.native().on_progress(1000, @() => keep_going())

PostgreSQL Specifics

The adapter speaks version 3 of the wire protocol, over TCP or over a unix domain socket, optionally under TLS.

A host beginning with a slash is a socket directory, which is how libpq spells it, and the socket inside is named after the port: /var/run/postgresql with port 5432 means /var/run/postgresql/.s.PGSQL.5432. A path that already names the socket is taken as given. TLS is neither offered nor wanted over a socket, since nothing sits in between, so sslmode is ignored there.

sql.open('postgres:///app?host=/var/run/postgresql')
sql.open({ driver: 'postgres', host: '/var/run/postgresql', database: 'app' })

sslmode chooses how TLS is used: disable never, prefer when the server offers it, and require always, failing when the server refuses. Only require verifies who is on the other end.

Statements go through the extended protocol, which compiles them on the server and binds values rather than substituting them. Values travel in the server’s own binary form wherever the adapter has a codec for the type, and as text otherwise, so a type it has never heard of still arrives readable rather than as bytes.

numeric columns arrive as sql.Decimal, arrays as lists, jsonb as dictionaries, and timestamps as date.Date.

Listening for notifications

LISTEN and NOTIFY are PostgreSQL’s own publish and subscribe. A connection that has run LISTEN channel receives a message whenever anything anywhere runs NOTIFY channel.

var listener = db.native().listener()

listener.listen('jobs')

while true {
  for message in listener.wait(nil) {
    handle(message.channel, message.payload)
  }
}

Notifications arrive between other messages, so a connection busy running statements collects them as it goes and poll() hands over whatever has accumulated.

Teaching it a type

A database that defines its own types can teach the adapter about them:

import sql.postgres { DEFAULT_REGISTRY }

DEFAULT_REGISTRY.register(oid, decoder, encoder)

Until it is taught, such a column arrives as text, which is correct if unexciting.

MySQL Specifics

The adapter speaks the client/server protocol directly, over TCP or over a unix domain socket, optionally under TLS, with no client library underneath it. MariaDB speaks the same protocol and the same adapter drives it.

A socket option names a socket, and so does a host beginning with a slash. Unlike PostgreSQL the path names the socket itself rather than the directory holding it:

sql.open({ driver: 'mysql', socket: '/var/run/mysqld/mysqld.sock', user: 'app' })
sql.open('socket=/var/run/mysqld/mysqld.sock user=app database=shop')
sql.open('mysql://app@localhost/shop?socket=/var/run/mysqld/mysqld.sock')

TLS is neither offered nor wanted there, since nothing sits in between. The connection does count as private, which is what lets caching_sha2_password and sha256_password send the password itself rather than encrypting it to the server’s public key.

A statement with no values is sent as text, which is one round trip. A statement with values is prepared, so the values travel in the server’s binary encoding instead of being written into the SQL, and the question of quoting never arises. Compiled statements are kept and reused, so a query run in a loop is compiled once.

import sql

var mysql = sql.driver('mysql')

echo mysql.capabilities().placeholder_style
echo mysql.capabilities().last_insert_id
echo mysql.capabilities().returning
echo mysql.quote_identifier('order by')
echo sql.driver('mariadb').capabilities().returning
question
true
false
`order by`
true

MySQL keeps ? as it is written, and reports the id of an inserted row rather than returning it, so insert() reads the reported id instead of adding a RETURNING clause. MariaDB differs in two ways that change what the layer above generates, which is why mariadb:// is a scheme of its own: it has RETURNING, and it has no JSON type.

What arrives from where

DECIMAL columns arrive as sql.Decimal, JSON as lists and dictionaries, DATE and DATETIME and TIMESTAMP as date.Date, and TIME as sql.Time.

MySQL has no boolean. BOOLEAN is another name for TINYINT(1), so a value written as true comes back as 1, and db.supports('booleans') is false to say so.

BLOB and TEXT share a type on the wire and are told apart only by the column’s collation, which the adapter reads: a binary column arrives as bytes and a text one as a string. BIT and GEOMETRY arrive as bytes, having no shape Zuri could represent without inventing one. SET arrives as the comma separated text the server stores rather than as a list, because a list written back would be encoded as JSON and the column would quietly stop matching.

A date MySQL considers absent is written as zeros, and 0000-00-00 is not a date any calendar has. It arrives as nil.

A TIME is a span, not a clock

A TIME column runs from -838:59:59.999999 to 838:59:59.999999. It is a duration rather than a point in the day, so it can be negative and can exceed twenty four hours, and neither of those would survive being read into a date.Date. sql.Time holds it:

import sql

var shift = sql.time(9, 30, 0)
var late = sql.time(0, 0, 30, 0, true)

echo shift.to_string()
echo shift.total_seconds()
echo late.to_string()
echo late.total_seconds()

# Two spans of the same length are equal however each was built.
echo sql.time(1, 30).equals(sql.time(0, 90))

# And the parts are kept as they were written.
echo sql.time(0, 90).minutes

echo sql.parse_time('-838:59:59.999999').to_string()
echo sql.time_from_seconds(-5400).to_string()
09:30:00
34200
-00:00:30
-30
true
90
-838:59:59.999999
-01:30:00

Time zones

A DATETIME carries no zone, and a TIMESTAMP is converted to and from whatever zone the session is in. So the session’s zone decides what a timestamp means, and leaving it to the server’s configuration would make the same database read differently from two machines.

The adapter sets the session to UTC when it connects. Every timestamp then arrives as UTC, and a date.Date written back is converted from whatever offset it carries. To leave the server’s own setting alone, pass time_zone as nil, or name a zone to use instead:

sql.open('mysql://localhost/app', { time_zone: nil })
sql.open('mysql://localhost/app', { time_zone: '+01:00' })

Authentication

caching_sha2_password, which is the default from MySQL 8.0 onwards, mysql_native_password, sha256_password, mysql_clear_password, and MariaDB’s client_ed25519.

Two of those need the password itself rather than a proof of it, the first time an account authenticates. Over TLS the password is sent as it is. Over a plain connection it is encrypted to a public key the server hands over first, so it is never readable in transit. mysql_clear_password has no such fallback and is refused outright on a connection that is not private, since it sends the password with nothing protecting it.

TLS

sslmode chooses how TLS is used:

ModeMeaning
disableNever.
preferWhen the server offers it, without checking who the server is. The default.
requireAlways, without checking who the server is.
verifyAlways, checking the server’s certificate and host name.

MySQL’s own spellings are accepted too. DISABLED, PREFERRED and REQUIRED mean what they say, and both VERIFY_CA and VERIFY_IDENTITY become verify. That makes VERIFY_CA stricter here than on the command line, where it checks the certificate but not the name: the difference can only cause a connection to be refused, never one to be wrongly trusted.

ssl_ca names a certificate authority to trust, and ssl_cert with ssl_key present a client certificate.

Compression

The protocol can compress everything after the handshake:

sql.open('mysql://localhost/app', { compression: 'zlib' })
sql.open('mysql://localhost/app', { compression: 'zstd' })

It is worth it for large results over a slow link and costs more than it saves on a local socket, so it is off unless asked for. A packet too small to benefit is sent uncompressed regardless, which the protocol allows for.

Sending a file

LOAD DATA LOCAL INFILE has the server name a file and the client send it. The file is read from the machine the program runs on, chosen by the server, so a hostile or compromised server could ask for anything the process can read. It is refused unless a reader is supplied:

import os

var db = sql.open('mysql://localhost/app', {
  local_infile: @(name) {
    return os.file(name).read()
  },
})

The reader decides what may be sent, which is where that decision belongs. Refusing still answers the server, so the connection stays usable afterwards rather than waiting for a file that will never come.

Statements that answer more than once

A stored procedure sends one result set per select inside it, then an OK packet. query() hands back the first and reads the rest, because leaving them unread would not lose them: it would give them to the next statement, which would then answer with this one’s rows.

var first = db.query('call two_results()')

A routine’s body has semicolons in it, and exec_script() splits a script on semicolons. So a create procedure goes through exec() as a single statement rather than through exec_script().

Reaching the adapter

db.native().flavor()          # 'mysql' or 'mariadb'
db.native().connection_id()   # what KILL names
db.native().parameters()      # what the server said about itself
db.native().reset()           # back to a fresh session

Writing Your Own Adapter

An adapter for another engine implements four things: a Driver that reads a connection string and opens a connection, a DriverConnection that runs statements, a DriverStatement, and a DriverCursor.

import sql

class DuckDbDriver < sql.Driver {
  name() { return 'duckdb' }
  schemes() { return ['duckdb'] }

  capabilities() {
    var caps = sql.default_capabilities()
    caps.set('placeholder_style', sql.QUESTION)
    caps.set('returning', true)

    return caps
  }

  parse_dsn(dsn) { ... }
  connect(options) { ... }
}

sql.register(DuckDbDriver())

var db = sql.open('duckdb://./app.duckdb')

Everything above the adapter comes for free: pooling, transactions and savepoints, the CRUD helpers, placeholder translation, cursors, and the error hierarchy. What the adapter supplies is the engine.

The base classes raise NotImplementedError for anything left out, so an unfinished adapter fails where the gap is. An engine that genuinely cannot do something raises NotSupportedError instead, which reads differently on purpose: the first is an unfinished adapter, the second is an honest limit.

What the Module Refuses

It is not an ORM. There is no model class, no lazy loading and no identity map. Rows are dictionaries.

It does not write your joins. The CRUD helpers cover a flat list of equality conditions. Anything past that is SQL.

It does not migrate schemas. exec_script() runs a schema file and db.schema reports what is there; deciding what to change and in what order is a separate problem.

It does not translate SQL. Placeholders are rewritten. Statements are not parsed, and a function only one engine has stays a function only one engine has.

It does not hide a difference by guessing. Where an engine cannot do what was asked, the answer is NotSupportedError naming the adapter and the feature, rather than an approximation nobody asked for.

Module Reference

The standard library reference documents every class and method. The shape of the module:

sql.open(dsn, options)opens a connection
sql.pool(dsn, options)opens a pool of them
sql.register(driver)adds an adapter
sql.drivers()what is registered
sql.raw(fragment)marks SQL to be used as written
sql.Decimal(text)an exact decimal
sql.time(h, m, s)a signed span, as a TIME column holds one
sql.parse_time(text)the same, read from text

On a Connection:

query, exec, exec_scriptrun a statement
fetch_one, fetch_all, fetch_value, fetch_columnread a result
stream, preparea cursor, and a compiled statement
insert, insert_many, update, delete, find, find_one, countthe four statements that are always the same
transaction, begin, in_transactiontransactions
schemaintrospection
driver, driver_name, capabilities, supportswhat this engine is
nativethe adapter underneath
ping, close, is_closedthe connection itself

On a Transaction:

query, execrun a statement
fetch_one, fetch_all, fetch_valueread a result
insert, insert_many, update, delete, find, find_one, countthe same helpers a connection has
commit, rollback, is_finishedending it
connection, nestedwhat it is running on, and whether it is a savepoint

Mail

The mail module is Zuri’s mail stack: the message format, the three protocols that move messages around, and the servers at the far end of two of them. It is written in Zuri from the socket up, and it speaks to real mail servers.

Mail is older than almost everything it runs on, and it shows. A message is text with a header block, defined in 1982 and extended ever since to carry anything that is not ASCII, which by now is most of what people send. Sending one is SMTP. Reading one where it is kept is IMAP. Taking one away is POP3. Three protocols, one format, and a great deal of accumulated history, almost all of which a program should not have to know about.

That is the design. A message is described by what is in it and the module works out the rest: which parts it needs, how they nest, which encoding each header and each body wants, and what has to be escaped so that a line of a single dot in the body does not end the message early. A program says who a message is from and what it says; it does not say Content-Transfer-Encoding.

Following Along

Most examples below build on one message. This is it:

import mail

var note = mail.message()
  .set_from('Ada Lovelace <ada@example.com>')
  .add_to('Grace Hopper <grace@example.com>')
  .set_subject('On Engines')
  .set_text('The analytical engine weaves algebraic patterns.')

echo note.subject()
echo note.sender().name
echo note.to()[0].address
On Engines
Ada Lovelace
grace@example.com

Nothing there touches a network. Building and reading messages is entirely separate from moving them, and a program that only needs to produce or parse one never opens a socket.

Introduction

There are three protocols and they do genuinely different things.

SMTP moves a message from wherever it was written to a server that will take responsibility for it. That is all it does. It has no notion of a folder, a read message, or a search; a message goes in and either is accepted or is not.

IMAP is for reading mail where it is kept. The mail stays on the server, the client works with it in place, and two devices looking at the same mailbox see the same thing. Almost everything a mail client does is IMAP.

POP3 is for taking mail away. There is a numbered list and there is removing things from it. It is the wrong protocol for reading mail on more than one device and the right one for a program whose job is to drain a mailbox into somewhere else.

The module covers the client end of all three and the server end of the two that have one worth having. There is no POP3 server here and there is not meant to be: a POP3 server is an IMAP server with almost everything taken away, and mail.imap is the one to run.

Building a Message

mail.message() starts one. What goes in it is said one call at a time, and each call hands the message back, so they read as one:

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_subject('On Engines')
  .set_text('The analytical engine weaves algebraic patterns.')

echo note.content_type().mime_type()
text/plain

A message starts with a Date, because the time it was written is the time it was written, and with MIME-Version. It gets its Message-ID when it is sent, because that is when the domain it belongs to is settled.

Addresses

Every address setter takes text, an Address, or a list of either:

import mail

var note = mail.message()
  .set_from('Ada Lovelace <ada@example.com>')
  .add_to(['grace@example.com', 'Charles Babbage <charles@example.com>'])
  .add_cc('archive@example.com')

echo note.to().map(@(person) => person.address)
echo note.to()[1].name
echo note.cc().length()
[grace@example.com, charles@example.com]
Charles Babbage
1

add_to() keeps whoever is already there, so it can be called in a loop. set_to() replaces them. The same pair exists for Cc and Bcc.

Bcc is removed from the message before it is handed to a server, after the recipients have been taken out of it. That is the whole point of a blind copy, and it is easy to get wrong by hand.

A display name with characters a header cannot carry is encoded, and one containing a comma is quoted, without being asked:

import mail

echo mail.Address('gruss@example.com', 'Grüße').to_string()
echo mail.Address('ada@example.com', 'Lovelace, Ada').to_string()
=?utf-8?B?R3LDvMOfZQ==?= <gruss@example.com>
"Lovelace, Ada" <ada@example.com>

Reading them back gives the name, not the encoding:

import mail

echo mail.parse_address('=?utf-8?B?R3LDvMOfZQ==?= <gruss@example.com>').name
echo mail.parse_address('"Lovelace, Ada" <ada@example.com>').name
Grüße
Lovelace, Ada

mail.parse_address_list() reads a whole header, including the comments and groups that RFC 5322 allows and nobody remembers:

import mail

var people = mail.parse_address_list(
  'Ada <ada@example.com> (the author), Engineers: grace@example.com, charles@example.com;'
)

echo people.map(@(person) => person.address)
[ada@example.com, grace@example.com, charles@example.com]

The group’s members come back alongside the rest, since a program sending mail cares who the recipients are and not how they were gathered. mail.parse_address_groups() keeps the grouping when that matters.

Text, HTML, or Both

Text alone is a plain message. HTML alone is an HTML message. Both together become a multipart/alternative, with the plain text first, because a reader that understands both is meant to take the last one it can display:

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_subject('On Engines')
  .set_text('The engine weaves algebraic patterns.')
  .set_html('<p>The engine weaves <em>algebraic patterns</em>.</p>')

echo note.content_type().mime_type()
echo note.parts().map(@(part) => part.content_type().mime_type())
multipart/alternative
[text/plain, text/html]

Setting the text again replaces it wherever in the tree it is, rather than adding a second copy. The shape is a consequence of the content, so it stays correct as the content changes.

Outgoing text is written as UTF-8, or as US-ASCII when that is all it needs. Naming any other character set raises, because the module cannot encode into one and a header claiming otherwise would be a lie.

Attachments

An attachment is a thing in its own right, built the way a message is and handed to attach():

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_subject('The figures')
  .set_text('Attached.')
  .attach(mail.attachment('name,total\nengines,41\n')
    .set_filename('figures.csv'))

echo note.content_type().mime_type()
echo note.attachments().map(@(part) => part.filename())
echo note.attachments()[0].content_type().mime_type()
multipart/mixed
[figures.csv]
text/csv

The media type came from the filename. The message became a multipart/mixed with whatever it held before as the first part; that happens once however many files are attached.

set_filename(name)the name to offer the file under
set_content_type(type)the media type, when the filename is not enough
set_disposition(kind)attachment by default, or inline
set_cid(id)the identifier HTML refers to it by
set_encoding(name)the transfer encoding, chosen from the data when absent
set_description(text)what some clients show beside the file

A file on disk is one call:

note.attach(mail.Attachment.from_file('reports/q3.pdf'))

The name and the media type both come from the path unless something says otherwise. A filename with characters a header cannot carry is encoded the way RFC 2231 says, split across continuations if it is long, and read back whole at the far end.

Images Inside the HTML

An image the message shows rather than offers is attached like any other file and referred to by an identifier:

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_subject('The logo')

var source = note.embed(
  mail.attachment(bytes([137, 80, 78, 71])).set_filename('logo.png')
)

note.set_html('<p>Our mark: <img src="${source}"></p>')

echo source.starts_with('cid:')
echo note.content_type().mime_type()
echo note.parts().map(@(part) => part.content_type().mime_type())
true
multipart/related
[text/html, image/png]

embed() returns the reference to point an <img> at and rearranges the message into the multipart/related that tells a reader the two belong together. The HTML ends up as the first part, which is what marks it as the one to display.

Headers of Your Own

Anything the setters do not cover goes through set_header():

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_header('X-Priority', '1')
  .set_header('List-Unsubscribe', '<https://example.com/unsubscribe>')

echo note.headers.get('X-Priority', nil)
1

The value is written exactly as given. A header that needs encoding needs it applied first; the setters that cover addresses and the subject do that for you, because those are the headers where getting it wrong is common.

add_header() keeps any header already there rather than replacing it, which is what Received needs.

Replies and Threads

A reply is an ordinary message with two more headers, and getting them right is what puts it in the same conversation in a mail client:

import mail

var original = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_subject('On Engines')
  .set_message_id('<first@example.com>')

var reply = mail.message()
  .set_from('grace@example.com')
  .add_to(original.sender())
  .set_subject('Re: ${original.subject()}')
  .set_in_reply_to(original.message_id())
  .set_references(original.references() + [original.message_id()])
  .set_text('Quite so.')

echo reply.headers.get('In-Reply-To', nil)
echo reply.references()
<first@example.com>
[<first@example.com>]

references() on the original gives whatever chain it already belonged to, so appending its own identifier extends the thread rather than starting a new one.

Reading a Message

mail.parse() takes the bytes and gives back the tree:

import mail

var note = mail.parse(
  'From: Ada Lovelace <ada@example.com>\r\n'
  + 'To: grace@example.com\r\n'
  + 'Subject: =?utf-8?Q?On_Engines_and_Gr=C3=BC=C3=9Fe?=\r\n'
  + 'Content-Type: text/plain; charset=utf-8\r\n'
  + '\r\n'
  + 'The engine weaves algebraic patterns.'
)

echo note.sender().name
echo note.subject()
echo note.text()
Ada Lovelace
On Engines and Grüße
The engine weaves algebraic patterns.

Nothing is decoded until something asks for it. Reading the subject of a message with a twenty megabyte attachment in it costs no more than reading the headers, which is what makes it reasonable to parse everything in a mailbox and look at only some of it.

The Tree

A part is a message too: it has headers and a body, and its body may be more parts. The same class covers both, so walking a message and reading a standalone one are the same code.

import mail

var note = mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_text('plain')
  .set_html('<p>html</p>')
  .attach(mail.attachment('some,data\n').set_filename('figures.csv'))

for part in note.walk() {
  echo part.content_type().mime_type()
}
multipart/mixed
multipart/alternative
text/plain
text/html
text/csv

walk() is the whole tree, outermost first. parts() is one level. find_part() is the first part of a given type anywhere in it.

Finding the Body

Most programs want the text, wherever it happens to be:

import mail

var note = mail.parse(mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_text('the plain version')
  .set_html('<p>the html version</p>')
  .to_bytes())

echo note.text_body()
echo note.html_body()
the plain version
<p>the html version</p>

Both return nil when the message has no such part, which is not the same as an empty one. An attachment that happens to be text/plain is not mistaken for the body: a part that says it is an attachment, or that carries a filename without saying either way, is left out.

Attachments Coming In

import mail

var note = mail.parse(mail.message()
  .set_from('ada@example.com')
  .add_to('grace@example.com')
  .set_text('Attached.')
  .attach(mail.attachment('name,total\nengines,41')
    .set_filename('figures.csv'))
  .to_bytes())

for part in note.attachments() {
  echo '${part.filename()} (${part.content_type().mime_type()})'
  echo part.body_bytes().to_string()
}
figures.csv (text/csv)
name,total
engines,41

body_bytes() decodes the transfer encoding, so base64 and quoted-printable both come back as the bytes that went in. text() goes one step further and applies the character set.

Character Sets

A header carrying anything outside ASCII carries it encoded, and there are two encodings it might have used. Reading a header through the module decodes both:

import mail.encoding

echo encoding.decode_words('=?utf-8?B?SGVsbG8=?= =?utf-8?B?IHdvcmxk?=')
echo encoding.decode_words('=?iso-8859-1?Q?caf=E9?=')
echo encoding.decode_words('this is not =?encoded? at all')
Hello world
café
this is not =?encoded? at all

Whitespace between two adjacent encoded words is dropped, which is what lets a long subject be split across several of them without a space appearing where none was written. Anything that is not a well-formed encoded word is left exactly as it is, including text that merely begins with =?.

Bodies say their character set in the Content-Type. The sets a message is likely to name are understood: UTF-8, US-ASCII, the ISO 8859 Latin sets 1 and 15, windows-1252, and UTF-16 in either byte order. One this does not know is read as UTF-8, which leaves whatever ASCII is in it intact rather than discarding the part that would have been readable.

import mail.encoding

echo encoding.decode_text(bytes([0x63, 0x61, 0x66, 0xe9]), 'iso-8859-1')
echo encoding.decode_text(bytes([0x93, 0x41, 0x94]), 'windows-1252')
café
“A”

Sending

mail.send() opens a connection, sends one message and closes it again:

import mail

mail.send('smtp://mail.example.com', mail.message()
  .set_from('reports@example.com')
  .add_to('ada@example.com')
  .set_subject('Quarterly report')
  .set_text('The numbers are in.'),
  { username: 'reports', password: secret })

That is the whole of sending mail for a program that sends one at a time.

A Client of Your Own

A program sending many wants a connection kept open across all of them, because opening one costs a handshake and an authentication exchange:

import mail.smtp { SmtpClient }

var server = SmtpClient.connect('smtp://mail.example.com', {
  username: 'reports',
  password: secret,
})

for note in queue {
  server.send(note, nil)
}

server.quit()

quit() says goodbye and closes. close() just closes, which is what to do when something has gone wrong and the conversation is no longer in a state the server would recognise.

schemeportwhat it means
smtp://587submission, with TLS negotiated over it
smtps://465TLS from the first byte

Port 25 is for one server relaying to another, not for a program submitting mail, and it is not a default here. Give it explicitly when that is genuinely what you are doing.

The Envelope

What the server is told and what the message says are two different things. The envelope comes from the message unless the options say otherwise: the sender from From, the recipients from To, Cc and Bcc together, with duplicates removed.

server.send(note, { from: 'bounces@example.com' })
server.send(note, { to: ['someone@example.com'] })

The first is what a mailing list does: the message says who wrote it, the envelope says where a bounce should go. The second delivers to somewhere the headers do not mention at all, which is how a blind copy actually works underneath.

An empty sender is the null path, which is what a bounce is sent from so that it cannot be bounced in turn:

server.send(bounce, { from: '' })

TLS

A client connects, reads what the server can do, negotiates TLS, and only then sends anything worth protecting. That is the default:

tlswhat happens
requirenegotiate TLS, and refuse to go on without it. The default.
prefernegotiate it when the server offers it
disabledo not ask

smtps://, imaps:// and pop3s:// handshake before the first byte instead, and then tls has nothing left to decide.

Whatever tls says, a mechanism that puts the password on the wire is never used on a connection that is not encrypted. A client offered nothing else raises rather than sending it. That is not configurable, and it is the one place the module refuses to do what it is told.

To trust a certificate the system does not, hand it a configuration:

import net.tls

var config = tls.TlsConfig()

config.add_ca_pem(file('internal-ca.pem').read())

var server = SmtpClient.connect('smtp://mail.internal', {
  username: 'reports',
  password: secret,
  tls_config: config,
})

Authenticating

Credentials go in the options, and the client picks the strongest mechanism both ends know:

SmtpClient.connect('smtp://mail.example.com', {
  username: 'reports',
  password: secret,
})

SmtpClient.connect('smtp://smtp.gmail.com', {
  username: 'reports@example.com',
  token: access_token,
})

A token picks a token mechanism. To force one rather than choosing, pass mechanisms; the choice is described under Authentication Mechanisms.

A username and password in the connection string work too, which is convenient for a string that came from configuration:

mail.send('smtp://reports:secret@mail.example.com', note, nil)

What the Server Can Do

echo server.capabilities().keys()
echo server.max_size()

A message larger than the server said it would take is refused before it is sent rather than after uploading it, which matters when the message is the reason the connection is slow.

Where the server offers CHUNKING, send() can hand the message over in pieces with nothing escaped:

server.send(note, { chunking: true })

Where it offers SIZE, 8BITMIME or DSN, those are used without being asked for. Where it does not, the message still goes.

When Sending Fails

A refusal partway through a transaction leaves the server holding half of one. send() abandons it before raising, so the connection is still usable for the next message rather than answering 503 to everything afterwards. That is worth knowing because doing it by hand is easy to forget:

for note in queue {
  catch {
    server.send(note, nil)
  } as error {
    failures.append([note, error])
  }
}

server.quit()

Every message after a failure still goes.

Reading Mail with IMAP

IMAP leaves the mail on the server. A client opens a mailbox, searches it, and fetches what it needs:

import mail.imap { ImapClient }

var inbox = ImapClient.connect('imaps://mail.example.com', {
  username: 'ada',
  password: secret,
})

inbox.select('INBOX')

for uid in inbox.search('UNSEEN', true) {
  var note = inbox.fetch_message(uid, true, false)

  echo '${note.sender().address}: ${note.subject()}'
}

inbox.logout()

Opening a Mailbox

var box = inbox.select('INBOX')

echo box.exists
echo box.uidvalidity
echo box.permanent_flags

select() opens a mailbox for reading and writing. examine() opens it read-only, which also means that reading a message does not mark it read. close_mailbox() closes it and removes anything marked deleted on the way out; unselect() closes it without doing that.

uidvalidity is how a server says its numbering has been reset. A client that remembers identifiers between sessions has to check it: if it has changed, every identifier it remembers means something else now.

To see what is there:

for box in inbox.list('', '*') {
  echo '${box.name} ${box.is_selectable() ? '' : '(container only)'}'
}

* matches anything including the separator between levels; % matches anything except it, which is what lists one level. status() asks what is in a mailbox without opening it, which is how a client shows unread counts for a dozen folders without selecting each one:

echo inbox.status('Archive', ['MESSAGES', 'UNSEEN'])

Searching

The criteria are IMAP’s own, and they read close to English:

inbox.search('UNSEEN', true)
inbox.search('FROM ada@example.com SINCE 1-Jan-2026', true)
inbox.search('SUBJECT engines LARGER 10000', true)
inbox.search('FLAGGED UNDELETED', true)

The second argument asks for unique identifiers rather than positions. A position is only good until something is removed from the mailbox and everything after it renumbers; an identifier outlives the connection and is what a program that runs twice should remember. Prefer true unless the numbers are being used immediately.

Fetching Less Than Everything

Fetching every message to show a list of them is the mistake IMAP exists to prevent:

for info in inbox.fetch('1:50', 'ENVELOPE FLAGS RFC822.SIZE', false) {
  echo '${info.envelope.subject} (${info.size} bytes)'
}

The envelope is the addresses and the date, parsed by the server. No body crossed the network at all.

what to ask forwhat comes back
ENVELOPEthe addresses, subject and date
FLAGSwhat has been done to the message
RFC822.SIZEhow large it is
INTERNALDATEwhen the server took it
BODYSTRUCTUREthe shape of the message, part by part
BODY.PEEK[HEADER]the header block
BODY.PEEK[]the whole message
BODY.PEEK[2]one part of it

BODYSTRUCTURE is the one worth knowing about. It reports what the message is made of without sending any of it, so a client can decide to fetch the text and leave a large attachment on the server:

var structure = inbox.fetch(uid, 'BODYSTRUCTURE', true)[0].structure

for part in structure.walk() {
  echo '${part.section} ${part.mime_type()} ${part.size}'
}

section is the number to ask for that part on its own.

fetch_headers() is the common case of asking for the header block, and fetch_message() the common case of asking for all of it:

inbox.fetch_message(uid, true, false)   # leaves it unread
inbox.fetch_message(uid, true, true)    # marks it read

Reading a message does not mark it read unless you say so. A program going through a mailbox should not change what a person sees when they next open it.

Flags

A flag is what has been done to a message:

inbox.add_flags(uid, ['\\Seen'], true)
inbox.remove_flags(uid, ['\\Flagged'], true)
inbox.mark_seen(uid, true)
inbox.mark_deleted(uid, true)

\Seen, \Answered, \Flagged, \Deleted and \Draft are the ones with a defined meaning. A server may allow others, and says which in the mailbox’s permanent_flags.

mark_deleted() only marks. Nothing goes until expunge():

inbox.mark_deleted(uid, true)
echo inbox.expunge()

expunge() returns the positions that went, highest first, because each removal renumbers everything after it and that is the order they have to be applied in.

Moving, Copying and Removing

inbox.copy(uid, 'Archive', true)
inbox.move(uid, 'Archive', true)

move() uses the server’s own MOVE where there is one, and otherwise does what MOVE was invented to replace: copy, mark deleted, expunge. Either way the message ends up in one place.

create(), delete(), rename(), subscribe() and unsubscribe() do what they say.

Putting a Message Back

append() adds a message to a mailbox without sending it anywhere, which is how a sent message gets into the Sent folder:

inbox.append('Sent', note, ['\\Seen'], nil)

The flags are the ones to file it under, and \Seen is the usual one for something the account itself wrote. The last argument is the internal date; the time of arrival is used when it is not given.

Waiting for New Mail

idle() waits for the server to say something rather than asking over and over:

inbox.on_event(@(response) {
  echo 'the server said ${response.name()}'
})

while true {
  inbox.idle(1500000)
}

A server may drop a connection that idles for too long, which is why the wait is bounded and the loop comes back around. Twenty-five minutes is the default and is what RFC 2177 recommends.

noop() does the same thing without waiting: it gives the server a chance to report anything that has changed, and keeps the connection from going idle at all.

Collecting Mail with POP3

POP3 hands the mail over and forgets it. There is no searching and no folders; there is a numbered list, and there is taking things off it:

import mail.pop3 { Pop3Client }

var mailbox = Pop3Client.connect('pop3s://mail.example.com', {
  username: 'ada',
  password: secret,
})

for entry in mailbox.list() {
  archive(mailbox.retrieve(entry.number))
  mailbox.delete(entry.number)
}

mailbox.quit()

list() gives every message with its size and, where the server offers one, an identifier that outlives the session. The number does not: it is only good until something is removed and the rest renumber.

Nothing is actually removed until quit(). delete() only marks and reset() unmarks everything. A connection that drops halfway leaves the mailbox exactly as it was, which is the protocol protecting you from a program that fails in the middle. It also means that closing without quit() is how to abandon a run.

top() fetches the headers and the first few lines of the body, which is enough to decide whether the rest is worth fetching:

var preview = mailbox.top(entry.number, 5)

Where the server’s greeting offers it, APOP is used in preference to sending the password, and where it offers SASL those mechanisms are preferred again. USER/PASS is the last resort and is only used over TLS.

Running an SMTP Server

SmtpServer knows the protocol and nothing about policy. What to accept is decided by handlers:

import mail.smtp { SmtpServer }

var server = SmtpServer({ port: 2525, hostname: 'mail.example.com' })

server.on_rcpt(@(session, recipient) {
  if !accounts.contains(recipient.local) {
    return { code: 550, message: 'no such user here', enhanced: '5.1.1' }
  }
})

server.on_data(@(session, raw) {
  for recipient in session.recipients {
    store.append(recipient.local, 'INBOX', raw, nil, nil)
  }
})

server.listen()

The Handlers

handlercalled withwhen
on_connectthe sessiona connection opens, before the greeting
on_auththe session and the credentialsa client authenticates
on_mailthe session and the sendera transaction starts
on_rcptthe session and a recipientfor each recipient
on_datathe session and the messagethe whole message has arrived
on_closethe sessionthe connection closes, however it closed
on_errorthe error and the sessionsomething inside the server failed

The session carries what is known so far: who connected, whether the connection is encrypted, who authenticated, the sender, and the recipients accepted up to now. session.state is an empty dictionary your handlers can put anything in, and it lives as long as the connection.

Refusing Properly

A handler that returns nothing accepts. One that returns a code and a message refuses with those:

server.on_mail(@(session, sender) {
  if blocklist.contains(sender.domain) {
    return { code: 550, message: 'not accepted from there', enhanced: '5.7.1' }
  }
})

A handler that raises is a failure on the server’s side, not a bad message, and the sender is told 451 and to try again later. That distinction matters: a database that is down should not turn into a bounce.

Refusing at on_rcpt is the useful one. It tells the sender immediately which address is wrong, while it is still connected and can do something about it, rather than accepting the message and generating a bounce to an address that may not exist either.

The null sender, which is what a bounce comes from, arrives at on_mail as an empty string rather than an address. Refusing to accept mail from it is how a server ends up unable to receive bounces, so handle it deliberately.

TLS and Authentication

server.use_tls(file('cert.pem').read(), file('key.pem').read())

That is what lets the server offer STARTTLS. Everything a client said before the handshake is discarded afterwards, including who it claimed to be, because none of it was protected.

Authentication needs one handler and, for the mechanisms that prove a password without sending it, a second:

server.on_auth(@(session, credentials) {
  if !accounts.verify(credentials.username, credentials.password) {
    return { code: 535, message: 'no' }
  }
})

server.on_password(@(username) {
  return accounts.password_of(username)
})

on_password is what makes CRAM-MD5 possible: checking that proof means working out the same one, which means knowing the password. A server that cannot produce one does not advertise the mechanism. Without on_auth the server advertises no authentication at all.

PLAIN and LOGIN are advertised only once the connection is encrypted. require_auth refuses mail from a client that has not authenticated, and require_tls refuses it on a connection that is not encrypted.

Limits

optiondefaultwhat it does
max_size35 MBthe largest message to take
max_recipients100recipients one message may have
timeout300000milliseconds a client may go quiet

A client that declares a size past the limit is refused at MAIL, before it uploads anything. One that does not declare it is refused when it goes past. Ten malformed commands in a row and the connection is dropped.

Running an IMAP Server

ImapServer runs the session state machine and answers every command out of a MailStore. What mail there is and where it lives is the store’s business:

import mail.imap { ImapServer, MaildirStore }

var store = MaildirStore('/var/mail')

store.add_account('ada', secret)

ImapServer({ port: 143 }, store).listen()

The server implements IMAP4rev1 along with UNSELECT, MOVE, IDLE, LITERAL+, SASL-IR and ID. It advertises exactly what it implements, so a client that reads the capability list and trusts it will not go wrong.

Where the Mail Lives

Two stores ship. MemoryStore keeps everything in the process and forgets it on exit, which is what a test wants and what a server embedded in something larger wants when the mail it holds is not the point:

import mail
import mail.imap { MemoryStore }

var store = MemoryStore(['INBOX', 'Archive'])

store.add_account('ada', 'secret')
store.append('ada', 'INBOX', mail.message()
  .set_from('grace@example.com')
  .add_to('ada@example.com')
  .set_subject('On Bugs')
  .set_text('Found one.')
  .to_bytes(), nil, nil)

echo store.mailboxes('ada')
echo store.messages('ada', 'INBOX').length()
echo store.messages('ada', 'INBOX')[0].uid
[Archive, INBOX]
1
1

MaildirStore is the same contract on disk.

Maildir

Each account is a directory, holding its INBOX directly and every other mailbox beside it under a leading dot, which is the Maildir++ arrangement every other mail tool understands. A mailbox written by this server can be read by anything else, and mail delivered by anything else turns up here.

Every mailbox is the three directories Maildir defines. A message is written into tmp, where nothing reads from, and only moved into place once it is whole, so a reader never sees half of one. new is where a delivery agent leaves mail nobody has looked at yet, and cur is where a message lives once a client has seen the mailbox, with its flags recorded in the filename after :2,. Windows does not allow a colon in a filename, so there the flags follow ;2, instead, as they do for mbsync. So an account on disk looks like this:

ada/
  cur/   new/   tmp/   zuri-uidlist
  .Archive/
    cur/   new/   tmp/   zuri-uidlist

That interoperability is the reason to choose it. A delivery agent can drop a message in and the server finds it on the next look, with no shared database and no protocol between them.

IMAP needs a message number that survives a rename and Maildir has no such thing, so each mailbox keeps a small index file beside its directories recording which file is which number. A message whose file has gone keeps its number retired rather than reused, so a client that remembered one is told the message is missing rather than handed a different one.

Mail is on disk; accounts are not:

store.add_account('ada', secret)

store.set_authenticator(@(username, password) {
  return accounts.verify(username, password)
})

Where an application keeps its passwords is the application’s business and not a mail library’s. A store with an authenticator can no longer produce a password, so the server stops advertising the mechanisms that need one.

A Store of Your Own

A store only has to answer the same calls. Nothing in the contract requires a file:

authenticate(username, password)are these the right credentials
password_of(username)the password, for the mechanisms that need it
has_passwords()whether it can produce one at all
mailboxes(username)what mailboxes there are
exists, create, remove, renamethe mailboxes themselves
messages(username, mailbox)everything in one, oldest first
append(username, mailbox, raw, flags, received)add a message
set_flags(username, mailbox, uid, flags)change one’s flags
expunge(username, mailbox)remove what is marked deleted
counters(username, mailbox)uidnext and uidvalidity

Subclass MailStore and implement them, and an ImapServer will serve whatever is behind it.

Serving More Than One Connection

Both servers serve one connection to the end before taking the next. That is the right shape for the protocols and the wrong shape for more than one client at a time. mail.pool puts a pool of isolates behind one listening socket:

import mail.pool
import .my_server

pool.serve(my_server.build, { host: '0.0.0.0', port: 143, workers: 8 })

build runs inside each worker to construct that worker’s own server, since isolates share nothing and each needs its own. It can be any function; keeping it in a module of its own is just a tidy place for a worker’s setup to live:

# my_server.zu
import mail.imap { ImapServer, MaildirStore }

def build() {
  var store = MaildirStore('/var/mail')

  store.set_authenticator(accounts.verify)

  return ImapServer({}, store)
}

An IMAP connection can be open for hours, so the pool is a ceiling on how many clients can be served at once rather than on how fast they are served. Size it for the number of clients, not the rate of requests.

pool.start() does everything except run the accept loop, which is what to use when the address has to be known before the first connection, as it does in a test.

Proving Where Mail Came From

A receiving server has no reason to believe a From header. DKIM is the domain signing the message on the way out, and the receiver checking the signature against a key published in that domain’s own DNS.

Signing

import mail.dkim { Signer }

var signer = Signer('example.com', 'default', private_key)

signer.sign(note)

The signature covers the body and the headers worth covering, and the header it adds goes at the top. It goes on last, once everything else about the message is settled: changing a signed header afterwards breaks it, and mail.send() sets the Message-ID and Date if they are absent, so let it or set them yourself before signing.

optiondefaultwhat it does
algorithmrsa-sha256or ed25519-sha256
canonicalisationrelaxed/relaxedheaders then body
headersa sensible setwhich headers to cover
expires_innoneseconds until the signature stops counting

relaxed forgives the whitespace and folding changes a mail server may make in passing. simple covers the bytes exactly, which means any change at all on the way breaks the signature. Both are implemented; relaxed/relaxed is what almost everything uses and is the default for that reason.

The public half goes in DNS under the selector:

default._domainkey.example.com  TXT  "v=DKIM1; k=rsa; p=MIIBIjANBg..."

Checking

import mail.dkim

for result in dkim.verify(incoming) {
  if result.valid {
    echo '${result.domain} takes responsibility for this'
  } else {
    echo '${result.domain}: ${result.reason}'
  }
}

One result per signature, in the order they appear. A message with no signatures gives an empty list, which is not a failure: it is a message nobody signed.

The key is looked up through net.resolver unless verify() is handed something else to look it up with, which is what a program with its own resolver or its own cache wants:

dkim.verify(incoming, @(name) {
  return my_dns.text_records(name)
})

A message that was parsed is checked against the bytes it was parsed from, which is the only thing a signature can be checked against. Changing a parsed message and checking it again reports on the message that arrived, not the one now in hand.

What a Signature Does Not Say

A valid signature says the domain vouches for the message. It does not say the message is wanted, that the From header matches the signing domain, or that the sender is who they claim to be to a human reader. Deciding what a signature from a given domain is worth is a separate question with a separate answer, and that answer is usually DMARC.

DNSSEC is not validated. Signatures are carried through when a server sends them and the dnssec option asks for them, but nothing here checks one. A program that needs a validated answer wants a validating resolver and a trusted path to it, which is what net.resolver’s tls option gives.

Authentication Mechanisms

All three protocols carry the same mechanisms, and mail.sasl holds them once rather than three times. A client is given credentials and picks the strongest thing both ends know:

SCRAM-SHA-256, SCRAM-SHA-1proves the password without sending it, and proves the server knew it too
CRAM-MD5proves it without sending it, and proves nothing about the server
XOAUTH2, OAUTHBEARERa bearer token, which is what the large providers want
PLAIN, LOGINthe password itself, so never without TLS
EXTERNALnothing: the client certificate already said who this is

The order is the order of that table. SCRAM is the one to use where a server offers it: the server stores something derived from the password rather than the password, the client never sends it, and the exchange ends with the server proving it knew it too, which is what stops a server that simply says yes to everything.

Channel binding, the -PLUS form of SCRAM, is not offered. It needs a value out of the TLS session that net.tls does not expose.

To force a mechanism rather than choosing:

SmtpClient.connect(url, {
  username: 'reports',
  password: secret,
  mechanisms: ['SCRAM-SHA-256'],
})

Errors

Every error is a MailError. Catching that catches everything the module raises on its own.

MessageErrorthe bytes are not a message, or are one that contradicts itself
ProtocolErrorthe server said something the protocol does not allow
ConnectionClosedthe connection went away mid-conversation
AuthenticationErrorthe credentials were refused, or nothing usable was offered
StateErrora command that makes no sense where it was issued
MailboxErrora mailbox that does not exist, or cannot be created
SmtpErrorthe SMTP server refused, with a code
ImapErrorthe IMAP server answered NO or BAD
Pop3Errorthe POP3 server answered -ERR

The one distinction worth building a mail program around is SMTP’s:

catch {
  mail.send(url, note, credentials)
} as error {
  if instance_of(error, mail.SmtpTransientError) {
    queue.retry(note)
  } else if instance_of(error, mail.SmtpPermanentError) {
    queue.bounce(note, error.code)
  } else {
    raise error
  }
}

A 4xx reply means the server could not take the message now and the sender should try again later. A 5xx means it will not take the message and trying again changes nothing. Queue the first and bounce the second; treating them alike is how a mail queue either loses mail or sends it forty times.

SmtpError carries code, the three digit reply, and enhanced, the finer grained code from RFC 3463 when the server sends one. Neither is meant to be matched on beyond its first digit.

ImapError carries status, which is NO when the server understood the command and refused it, and BAD when it did not understand it at all. BAD points at the client rather than at the request.

What the Module Refuses

It will not send a password over an unencrypted connection. Not with an option, not with a flag. A client offered only mechanisms that would do that raises instead. tls: 'disable' turns off negotiating TLS; it does not turn off this.

It will not write a character set it cannot encode. Asking for a body in ISO 8859-1 raises rather than writing UTF-8 bytes under a header that claims otherwise.

It does not validate DNSSEC. See What a Signature Does Not Say.

It does not implement a POP3 server. A POP3 server is an IMAP server with almost everything removed, and running mail.imap is the better answer.

It does not decide whether mail is wanted. There is no spam filtering, no reputation, and no policy. DKIM tells you who signed something; what to do about that is yours.

Module Reference

The standard library reference documents every class and method. The shape of the module:

mail.message()starts a message
mail.parse(data)reads one
mail.send(url, note, options)sends one, connection and all
mail.smtp(url, options)a connection to a server that sends mail
mail.imap(url, options)a connection to a server that stores it
mail.pop3(url, options)a connection to a server that hands it over
mail.parse_address(text)one address
mail.parse_address_list(text)every address in a header
mail.format_addresses(people)the other direction
mail.attachment(data)starts an attachment
mail.Attachment.from_file(path)one read from disk

On an Attachment:

set_filename, set_content_type, set_encodingwhat it is
set_disposition, set_cid, set_descriptionhow it is presented
filename, content_type, disposition, cid, size, datareading them back
to_partthe message part it becomes

On a Message:

set_from, set_sender, set_reply_towho it is from
add_to, add_cc, add_bcc, set_to, set_cc, set_bccwho it is for
set_subject, set_text, set_htmlwhat it says
attach, embedwhat it carries, as an Attachment
set_header, add_header, set_date, set_message_idanything else
set_in_reply_to, set_referenceswhich conversation it belongs to
sender, to, cc, bcc, reply_to, recipientsreading them back
subject, text, text_body, html_bodyreading what it says
parts, walk, find_part, attachmentsthe tree
to_bytes, to_stringwriting it out

The protocol modules:

mail.smtpSmtpClient, SmtpServer, SmtpSession, Reply
mail.imapImapClient, ImapServer, Mailbox, Envelope, BodyPart, MessageInfo
mail.imapMailStore, MaildirStore, MemoryStore, StoredMessage
mail.pop3Pop3Client, Entry
mail.dkimSigner, Signature, Result, verify, is_signed
mail.saslof, choose, and a class per mechanism
mail.poolserve, start, Cluster

And the pieces underneath, for a program that needs them directly:

mail.addressparse, parse_list, parse_groups, format_list
mail.headersHeaders, fold, unfold
mail.encodingthe transfer encodings, encoded words and parameters
mail.contentContentType, ContentDisposition
mail.streamLineStream, connect, start_tls, endpoint
mail.imap.parserthe IMAP grammar, for reading a response by hand

Configuration and the Environment

Every program that talks to anything else needs to be told where it is. A database address, an API key, a port to bind, a flag that turns verbose logging on for one afternoon. None of that belongs in the source: it changes between your laptop and the server, and some of it must never reach a repository at all.

The answer the industry settled on is the process environment. Orchestrators set it, CI runners set it, shells set it, and every language can read it. What none of them solve is the first machine in the chain: yours, where there is no orchestrator and nobody wants to prefix every run with eight assignments.

The env module closes that gap. It reads a .env file into the process environment at startup, and it reads values back out already converted to the number, boolean or list your program is going to use. The file is a development convenience; the environment is the interface. A program written against env does not know or care which one supplied a value, which is exactly what lets it run unchanged in both places.

The First Line

One call, as early in the program as you can put it:

import env

file('.env.demo', 'w').write('PORT=8080\nGREETING=hello\n')

env.load('.env.demo')

echo env.get('GREETING')
echo env.int('PORT')
hello
8080

env.load() with no argument reads .env in the current working directory, which is what almost every program wants. The examples in this chapter write their file first and name it, so that you can run each one as it stands.

The file is optional. If it is not there, load() leaves the environment exactly as it found it and the program carries on with whatever the shell already provided. That is deliberate: the same binary runs on a laptop with a .env file and on a server without one.

Writing a .env File

One name per line, a =, and a value:

HOST=127.0.0.1
PORT=8080
DATABASE_URL=postgres://app@localhost/app

A name is a letter or underscore followed by letters, digits and underscores, which is the same rule a shell applies. Anything else stops the load with a ParseError rather than being silently dropped, because a name no shell could ever export is a typo every time.

A leading export is allowed and ignored, so the same file can be fed to source in a terminal when you want the variables in your shell too.

Quoting

An unquoted value is the text up to the end of the line, with the whitespace trimmed off both ends. Wrap it in quotes when that is not what you want:

import env

var values = env.parse(
  'BARE=  plain text  \n' +
  'COMMENTED=value # not part of it\n' +
  'PASSWORD=hunter#2\n' +
  "RAW='no \\n escape, no expansion'\n" +
  'COOKED="caf\\u00e9"\n'
)

for name, value in values {
  echo '${name} -> [${value}]'
}
BARE -> [plain text]
COMMENTED -> [value]
PASSWORD -> [hunter#2]
RAW -> [no \n escape, no expansion]
COOKED -> [café]

The three quoting forms differ in exactly one way each:

WrittenMeans
KEY=valuetrimmed, ends at a comment or the line
KEY='value'every character as written, nothing resolved
KEY=`value`the same, for values containing both other quotes
KEY="value"backslash escapes resolved, $ references expanded

Inside double quotes, \n, \r, \t, \f, \v, \b, \a, \e and \0 mean what they do in Zuri, and \xHH, \uHHHH and \u{H...} name a codepoint. A backslash before anything else keeps both characters, so "C:\Users\me" is the path you meant and not a lesson in escaping.

Reach for single quotes whenever a value is a secret. A generated password containing $ or \ goes through untouched, and nothing in it can accidentally name another variable.

Comments

A # starts a comment, either on its own line or after a value. Inside an unquoted value it only does so when it begins the value or follows a space, which is why PASSWORD=hunter#2 above kept its #. Where that rule is too subtle to rely on, quote the value and the question does not arise.

Values That Span Lines

A quoted value runs until its closing quote, newlines included. This is how a private key goes in a file:

SIGNING_KEY="-----BEGIN PRIVATE KEY-----
MIIEvQIBADANBgkqhkiG9w0BAQEFAASCBKcwggSjAgEAAoIBAQC7...
-----END PRIVATE KEY-----"

Single quotes work the same way and resolve nothing, which is usually the better choice for key material.

What Is Already Set Wins

A name that is already set in the process environment is left alone. The file fills in the rest:

import env
import os

os.set_env('DATABASE_URL', 'postgres://production/app', true)

file('.env.demo', 'w').write(
  'DATABASE_URL=postgres://localhost/app\n' +
  'CACHE_TTL=300\n'
)

var result = env.load('.env.demo')

echo os.get_env('DATABASE_URL')
echo result.applied
echo result.skipped
postgres://production/app
[CACHE_TTL]
[DATABASE_URL]

This is the whole point of the default. The file carries what a developer needs to run the program at all; production sets the real values through the shell, the orchestrator or the CI runner, and the file never has to know that it did.

load() returns a Result that accounts for every name and every file, so a program never has to guess whether its configuration arrived. applied is what the load set, skipped is what was already set, values is everything the file defined either way, and files and missing say which paths were read and which were not there.

Overriding

Where the file genuinely should win, say so:

import env
import os

os.set_env('LOG_LEVEL', 'warn', true)

file('.env.demo', 'w').write('LOG_LEVEL=debug\n')

env.loader().path('.env.demo').override().load()

echo os.get_env('LOG_LEVEL')
debug

Empty Means Unconfigured

Throughout the module, a variable set to the empty string counts as unset. FLAG= in a file, an empty shell variable, and a name nobody ever set all mean the same thing, because that is the one thing anybody means by any of them. A load will fill in an empty name, has() is false for it, require() raises on it, and every typed reader returns its default.

os.get_env() is the unfiltered view for the rare case that needs to tell an empty value apart from an absent one.

References Between Values

A value may refer to another value, or to the environment around it, with the syntax a shell uses:

HOST=localhost
PUBLIC_URL=http://${HOST}:${PORT:-8080}
import env

file('.env.demo', 'w').write(
  'HOST=localhost\n' +
  'PUBLIC_URL=http://$' + '{HOST}:$' + '{PORT:-8080}\n'
)

var loaded = env.load('.env.demo')

echo loaded.values.PUBLIC_URL
http://localhost:8080

The split string in that example is a Zuri detail, not a .env one: a ${ written inside a Zuri literal is an interpolation, so building .env text in source means keeping the $ and the { apart. A file you type in an editor has no such problem, as the dosini block above shows.

A reference resolves to the value the name will hold once the load has finished. That single rule is what makes PORT=${PORT:-8080} read the way it looks: whatever the environment already says, and 8080 when it says nothing. It holds no matter which line of the file the reference sits on, and it changes with override() exactly as precedence does.

A definition may also reach past itself to the value it is replacing, which is how a path gets extended rather than clobbered:

PATH=${PATH}:/opt/app/bin

Fallbacks and Demands

The three modifiers are POSIX’s, and the leading colon on each widens the test from “is it set” to “is it set to something”:

WrittenMeans
${NAME:-fallback}the value, or fallback when it is unset or empty
${NAME-fallback}the value, or fallback when it is unset
${NAME:+instead}instead when the value is set and not empty, otherwise nothing
${NAME+instead}instead when the value is set, otherwise nothing
${NAME:?reason}the value, or a MissingVariable carrying reason
${NAME?reason}the same, counting an empty value as set

The text after a modifier is itself expanded, so a fallback may have a fallback:

CACHE_DIR=${XDG_CACHE_HOME:-${HOME}/.cache}/app

:? is worth knowing. It turns the file itself into the place where a deployment’s requirements are written down, and the failure arrives at startup with the reason attached:

SESSION_SECRET=${SESSION_SECRET:?the deploy must supply a session secret}

Literal Dollar Signs

\$ is a dollar sign and nothing else. A $ that does not name anything, as in PRICE=$5.00, stands for itself already. A value in single quotes or backticks is never expanded at all, so a secret containing ${ needs no thought:

API_KEY='sk_live_${not_a_reference}'

Reading Values Back

An environment variable is always text, and almost nothing in a program wants text. The typed readers convert it and refuse anything that is not what it claims to be:

import env

file('.env.demo', 'w').write(
  'PORT=8080\n' +
  'DEBUG=yes\n' +
  'REQUEST_TIMEOUT=2.5\n' +
  'CORS_ORIGINS= https://a.example , https://b.example \n' +
  'SENTRY_DSN=\n'
)

env.load('.env.demo')

echo env.int('PORT', 3000)
echo env.bool('DEBUG', false)
echo env.float('REQUEST_TIMEOUT', 1)
echo env.list('CORS_ORIGINS', ',', [])
echo env.get('SENTRY_DSN', 'not configured')
echo env.has('SENTRY_DSN')
8080
true
2.5
[https://a.example, https://b.example]
not configured
false
get(name, default)the text, or the default
require(name, reason)the text, or a MissingVariable
has(name)whether it is configured at all
int(name, default)a decimal integer, optionally signed
float(name, default)a number, exponents included
bool(name, default)1, true, yes, y, on and their opposites
list(name, separator, default)split, trimmed, empty items dropped

A value that is set but not convertible is a ValueError naming the variable, not a silent fall back to the default. PORT=eighty is a mistake in the configuration, and the moment to hear about it is startup rather than the first request that needed a port.

These read the process environment, not the file. They answer the same whether a value came from .env or from the shell, which is what lets one program run in both places without a branch anywhere in it.

require() deserves its own line in most programs. A setting the code cannot invent a default for should stop the program at the top, with its own name in the message:

var secret = env.require('SESSION_SECRET', 'set SESSION_SECRET in .env')

Building a Load

env.loader() returns a Loader when the one-line form is not enough. Every setting returns the loader, so a whole configuration is one expression, and nothing about the loader changes when it runs, so the same one can be kept and used again.

Layering Files

Sources are read in the order they are added, and the last one to define a name is the one that defines it:

import env

file('.env.demo', 'w').write('HOST=localhost\nPORT=8080\n')
file('.env.local.demo', 'w').write('HOST=127.0.0.1\n')

var result = env.loader()
  .path('.env.demo')
  .path('.env.local.demo')
  .path('.env.missing.demo')
  .read()

echo result.values
echo result.files
echo result.missing
{HOST: 127.0.0.1, PORT: 8080}
[.env.demo, .env.local.demo]
[.env.missing.demo]

That pairing is the useful one: .env holds what the team shares and is worth committing as .env.example, and .env.local holds what one machine does differently and is not. A path may begin with ~, which expands to the home directory.

Text Instead of a File

source() adds text rather than a path, so configuration that arrived over the network or out of a secret store goes through exactly the same parsing, expansion and precedence as a file on disk:

env.loader()
  .path('.env')
  .source(vault.fetch('app/production'))
  .load()

Demanding the File

A file that is not there is not an error by default, and its path is recorded in Result.missing. Where the file is genuinely part of the deployment, say so and let it fail loudly:

env.loader().path('/etc/app/env').required().load()

Looking Without Loading

read() does everything load() does except the last step. Nothing is written to the process environment, and $ references still resolve the way they would have, so what comes back is exactly what load() would have set. It is how a tool inspects a file, and how a test checks one without changing the process it is running in.

expand(false) turns expansion off entirely, for a file of opaque secrets where every value should be taken exactly as written.

Writing a File Back Out

stringify() is parse() backwards, for the program that generates a .env file rather than reading one:

import env

print(env.stringify({
  HOST: '127.0.0.1',
  PORT: 8080,
  GREETING: 'hello there',
  DEBUG: false,
}))
HOST=127.0.0.1
PORT=8080
GREETING="hello there"
DEBUG=false

Values are written bare where that is unambiguous and double-quoted where it is not, and a $ inside a value is escaped so that loading the file back does not expand it. The output round-trips: feeding it to parse() returns the same names and the same values.

Numbers, bigints, booleans and bytes are converted to their text. Anything else is a TypeError, because guessing what a list should look like in an environment file is how configuration goes wrong quietly.

When It Goes Wrong

Every error the module raises is an EnvError, and each subclass carries the detail a message alone cannot:

import env

catch {
  env.parse('12FACTOR=yes')
} as error {
  echo '${error.type}: ${error.message}'
  echo 'line ${error.line}, column ${error.column}'
}

catch {
  env.loader().path('.env.nowhere').required().load()
} as error {
  echo '${error.type}: ${error.message}'
}

catch {
  env.require('NOTHING_SET_THIS')
} as error {
  echo '${error.type}: ${error.message}'
}
ParseError: expected a variable name
line 1, column 1
MissingFile: no environment file at .env.nowhere
MissingVariable: NOTHING_SET_THIS is not set
ClassRaised whenCarries
ParseErrora source is not a valid environment fileline, column
MissingFilea required file is not therepath
MissingVariablea demanded variable is not configuredname

Catching EnvError catches all three, and a ValueError from a typed reader is the ordinary one, so it goes wherever the rest of your validation failures go.

Keeping Secrets Out of Git

A .env file holds the values that differ between one deployment and the next, which is to say it holds the secrets. Three lines of housekeeping and the subject never comes up again:

$ echo '.env' >> .gitignore
$ echo '.env.local' >> .gitignore
$ cp .env .env.example    # then replace every real value

Commit .env.example with every name in it and no real value. It is the only documentation of what the program needs that cannot go stale, because the day someone adds a setting without adding it there is the day a new checkout stops working.

Resist the pull towards a .env.production checked in beside a .env.staging. Configuration varies per deploy, not per named environment, and the moment there are two files somebody edits the wrong one. Use one file per machine and let required() and :? say out loud what that machine still owes you.

Module Reference

The whole surface:

load(path)read a file into the process environment
loader()a Loader, for a load that needs more
parse(source)text to names and values, nothing else
stringify(values)names and values back to text
is_name(name)whether a name is a legal variable name

Reading values back:

get, require, hastext, or the absence of it
int, float, bool, listtext converted, or a ValueError

The classes:

Loaderpath, paths, source, override, expand, required, read, load
Resultvalues, applied, skipped, files, missing
EnvErrorParseError, MissingFile, MissingVariable

And the piece underneath, for a program that needs it directly:

env.expandexpand(value, resolve), the $ reference syntax on its own

Parsing the Command Line

A program started from a shell is handed a list of strings and nothing else. Everything a user meant by -v, --output report.csv or commit -m "fix the thing" has to be recovered from that list, and the recovering is where command-line programs quietly rot. The first version reads os.args and checks a few positions. The second adds a flag, and the checks become a chain of conditions. By the fourth nobody can say what -vo does without running it, the help text has drifted from the code that reads the flags, and a typo silently means something else instead of saying so.

The args module takes the declaration instead. You say which options exist, what type each one carries and which are required; it works out what the user typed, converts it, validates it, writes the help text from the same declarations, and refuses anything that does not fit.

The First Parser

Three lines of declaration and a call:

import args

var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })

var parsed = parser.parse(['--name', 'Ada'])

echo parsed.options.name
Ada

parse() with no argument reads the real command line, which is what a program does. Everywhere in this chapter it is given an explicit list instead, because that is also how you test one, and because a book example cannot rely on how you invoked it.

What parse() Returns

One dictionary with three keys, always:

import args

var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_index('source', 'The file to read')
parser.add_command('check', 'Check it over')

echo parser.parse(['-v', 'check', 'notes.txt'])
{options: {verbose: true}, command: {name: check, value: nil}, indexes: [notes.txt]}

options holds every option that was supplied or has a default. command is nil or the sub-command that was named. indexes holds the positional arguments in the order they were declared. The three are independent: a program can use one of them and ignore the rest.

A declared positional that nobody filled and that has no default keeps its place in indexes as a nil, so the argument after it is still found at the index it was declared at rather than sliding down one.

Options

An option is declared by its long name. The short name is optional, and so is everything else:

import args

var parser = args.Parser('tool')
parser.add_option('output', 'Where to write', { short_name: 'o', type: args.STRING })
parser.add_option('force', 'Overwrite without asking', { short_name: 'f' })

echo parser.parse(['-o', 'out.csv', '-f']).options
{output: out.csv, force: true}

--output and -o are the same option.

Attaching a Value

A long option takes its value either as the next argument or attached with =. The two spellings mean the same thing:

import args

def parser() {
  var p = args.Parser('tool')
  p.add_option('output', 'Where to write', { short_name: 'o', type: args.STRING })
  return p
}

echo parser().parse(['--output', 'out.csv']).options.output
echo parser().parse(['--output=out.csv']).options.output
out.csv
out.csv

The split is on the first =, so --output=a=b writes to a=b rather than losing half the name, and --output= is an explicit empty string rather than a missing value. The attached form is also the only way to pass a value that begins with a dash, because in --output -5 the parser reads -5 as an option:

import args

var parser = args.Parser('seek')
parser.add_option('offset', 'Where to start', { type: args.INT })

echo parser.parse(['--offset=-5']).options.offset
-5

A short option takes its value as the next argument only: -o out.csv, never -oout.csv and never -o=out.csv. A short token is a bundle of single-character flags, and neither = nor the letters of a value are flags.

Types

The type decides what the string becomes and whether a value is taken at all.

ConstantThe option takesBecomes
args.NONEnothingtrue when present
args.STRINGa valuethe string itself
args.INTa valuea whole number
args.NUMBERa valuea number, fractions allowed
args.BOOLa valuea boolean
args.LISTa value, repeatablea list of strings
args.CHOICEa value from a fixed setthe string, or what it maps to
args.OPTIONALa value, if one is therethe string, or true

NONE is the default and is the plain flag. BOOL is the one that takes an explicit answer, and it accepts every spelling a user is likely to reach for:

import args

def parser() {
  var p = args.Parser('tool')
  p.add_option('colour', 'Use colour', { short_name: 'c', type: args.BOOL })
  return p
}

echo parser().parse(['-c', 'yes']).options.colour
echo parser().parse(['-c', 'off']).options.colour
echo parser().parse(['-c', '0']).options.colour
true
false
false

1, true, yes, y and on all mean true; 0, false, no, n and off all mean false.

Defaults and Absence

value is what the option is worth when nobody supplies it:

import args

var parser = args.Parser('tool')
parser.add_option('count', 'How many', { short_name: 'c', type: args.INT, value: 1 })
parser.add_option('verbose', 'Say more', { short_name: 'v' })

echo parser.parse([]).options
echo parser.parse(['-c', '5', '-v']).options
{count: 1}
{count: 5, verbose: true}

An option with no default and no value on the command line is not in the dictionary at all. That is the difference between a flag that was left off and one that was set to false, and it is why verbose is missing from the first line rather than sitting there as false. Use options.contains('verbose') to ask, or give the option a default and stop having to.

Requiring an Option

import args

var parser = args.Parser('tool')
parser.add_option('output', 'Where to write', {
  short_name: 'o',
  type: args.STRING,
  required: true,
})

echo parser.parse(['-o', 'out.csv']).options.output
out.csv

Leave it off and the program stops with error: required option --output is missing, the usage line, and exit status 1. Nothing else runs.

Restricting the Value

CHOICE with a list accepts only what is in the list:

import args

var parser = args.Parser('tool')
parser.add_option('level', 'How loud', {
  type: args.CHOICE,
  choices: ['quiet', 'normal', 'loud'],
})

echo parser.parse(['--level', 'loud']).options.level
loud

Anything else stops the program with error: --level expects one of {'quiet', 'normal', 'loud'}, got "shouty", which names the offender and lists the alternatives without you writing either.

Give it a dictionary instead and the user types the key while the program receives the value. This is how a short spelling on the command line becomes a meaningful value inside:

import args

var parser = args.Parser('tool')
parser.add_option('mode', 'What to do', {
  short_name: 'm',
  type: args.CHOICE,
  choices: { r: 'read', w: 'write', rw: 'read-write' },
})

echo parser.parse(['-m', 'rw']).options.mode
read-write

Collecting Repeats

LIST accumulates every time the option appears:

import args

var parser = args.Parser('tool')
parser.add_option('tag', 'Tag to apply', { short_name: 't', type: args.LIST })

echo parser.parse(['-t', 'urgent', '-t', 'docs', '--tag', 'review']).options.tag
[urgent, docs, review]

Retiring an Option

An option you no longer want but cannot remove yet keeps working and says so on stderr:

import args

var parser = args.Parser('tool')
parser.add_option('old-name', 'Use --name instead', {
  type: args.STRING,
  deprecated: true,
})

echo parser.parse(['--old-name', 'Ada']).options['old-name']
Ada

The warning goes to stderr, so it reaches the person running the program without contaminating output that something else is reading.

Positional Arguments

A positional argument is declared with add_index, and they fill in the order they were declared:

import args

var parser = args.Parser('copy')
parser.add_index('source', 'The file to read', { required: true })
parser.add_index('dest', 'Where to write it', { value: 'out.txt' })

echo parser.parse(['in.csv']).indexes
echo parser.parse(['in.csv', 'report.csv']).indexes
[in.csv, out.txt]
[in.csv, report.csv]

They take type, choices, value, required and metavar, the same as options do. A positional that is neither declared nor expected is an error rather than something quietly ignored, which is what makes a mistyped flag fail loudly instead of being swallowed as a filename.

Sub-commands

A program with several jobs gives each one a name and its own options:

import args

var parser = args.Parser('notes')
parser.add_option('verbose', 'Say more', { short_name: 'v' })

parser.add_command('list', 'Show every note')
parser.add_command('remove', 'Delete a note').
  add_option('force', 'Do not ask first', { short_name: 'f' })

echo parser.parse(['-v', 'list'])
echo parser.parse(['remove', '-f'])
{options: {verbose: true}, command: {name: list, value: nil}, indexes: []}
{options: {force: true}, command: {name: remove, value: nil}, indexes: []}

add_command returns the command, so its options chain straight off it. A command’s own options land in the same options dictionary as the global ones; command.name is what tells you which job was asked for.

Order matters, and it is the order every tool of this shape uses: the parser’s own options come before the command name, and the command’s options come after it. notes -v list works and notes list -v does not, because after list the parser is reading list’s options and -v is not one of them.

A Value of Its Own

A command can take one value directly, the way git commit -m takes a message:

import args

var parser = args.Parser('notes')
parser.add_command('add', 'Write a note', {
  type: args.STRING,
  metavar: 'text',
}).add_option('pin', 'Keep it at the top', { short_name: 'p' })

echo parser.parse(['add', 'buy milk', '-p'])
{options: {pin: true}, command: {name: add, value: buy milk}, indexes: []}

The value comes immediately after the command name, before the command’s own options. metavar is the word the help text shows in place of the value, so the usage line reads notes add <text> rather than notes add <value>.

Running Something Directly

A command can carry the function that implements it, called after parsing with the options and the command’s value:

import args

var parser = args.Parser('notes')

parser.add_command('add', 'Write a note', {
  type: args.STRING,
  metavar: 'text',
  action: @(options, value) {
    var prefix = options.contains('pin') ? '[pinned] ' : ''
    echo prefix + value
  },
}).add_option('pin', 'Keep it at the top', { short_name: 'p' })

parser.parse(['add', 'buy milk', '-p'])
[pinned] buy milk

This is worth reaching for once there are more than two or three commands, because the alternative is a chain of comparisons on command.name that has to be kept in step with the declarations by hand.

The Shapes a Command Line Can Take

Users type things the parser was not asked about directly. These are the conventions it honours.

Bundling

Several flags behind one dash:

import args

var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_option('quiet', 'Say less', { short_name: 'q' })

echo parser.parse(['-vq']).options
{verbose: true, quiet: true}

Every character in a bundle has to name an option. -vqz is an error rather than -vq with the z quietly dropped, which is the difference between a typo you find now and one you find after the run.

Abbreviation

A long option may be shortened to any prefix that is still unambiguous:

import args

var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })

echo parser.parse(['--verb']).options
{verbose: true}

Set parser.allow_abbrev = false to turn this off, which is worth doing for a program whose option set is still growing: a prefix that is unambiguous today stops being so the day somebody adds --version, and a script written against the short form breaks.

End of Options

A bare -- ends option parsing. Everything after it is positional, even when it starts with a dash:

import args

var parser = args.Parser('run')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_index('command', 'What to run')
parser.add_index('argument', 'What to pass it')

echo parser.parse(['-v', '--', 'grep', '-i']).indexes
[grep, -i]

This is how a program passes arguments through to another one without having to know what they mean.

Arguments From a File

A token beginning with @ is replaced by the contents of that file, one argument per line:

import args

file('greet.args', 'w').write('--name\nAda\n--count\n3\n')

var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { type: args.STRING })
parser.add_option('count', 'How many times', { type: args.INT, value: 1 })

echo parser.parse(['@greet.args']).options
echo parser.parse(['@greet.args', '--name', 'Grace']).options
{name: Ada, count: 3}
{name: Grace, count: 3}

The file is expanded in place, so anything after it still wins. This is the escape hatch for a command line that has outgrown what a shell will comfortably hold, and for arguments a program would rather not put where ps can read them. Set parser.allow_atfile = false to turn it off for a program that should treat a leading @ as ordinary text.

Help

The help text is written from the declarations, so it cannot drift away from what the program accepts:

import args

var parser = args.Parser('greet')
parser.set_terminal_width(68)
parser.description = 'Greet somebody, once or several times.'
parser.epilog = 'Set NO_COLOR to turn off colour.'

parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
parser.add_option('count', 'How many times', { short_name: 'c', type: args.INT, value: 1 })
parser.add_index('output', 'Where to write the greeting')

parser.add_command('history', 'Show past greetings')

parser.help()
Usage: greet [OPTIONS] [COMMAND] [output]

  Greet somebody, once or several times.

POSITIONAL ARGUMENTS:
  [output]  Where to write the greeting

OPTIONS:
  -h, --help           Show this help message and exit
  -n, --name <name>    Who to greet
  -c, --count <count>  How many times (default: 1)

COMMANDS:
  history  Show past greetings

Set NO_COLOR to turn off colour.

Run "greet --help [COMMAND]" for help on a specific command.

-h and --help are declared for you and handled wherever they appear, including after a command, where they describe that command instead of the whole program. help() is the same thing called directly; both print and exit 0.

Colour is used when stdout is a terminal and NO_COLOR is unset, so a run whose output is piped or redirected gets plain text rather than escape codes. set_terminal_width() fixes the wrapping width, which is what the example above does so the output is the same on every terminal; left alone, the parser asks the terminal and falls back to COLUMNS, then to 80.

A parser whose command line matched nothing at all prints its help and carries on. An option with a default counts as a match, so a parser that defaults anything never does this; pass false as the second argument to args.Parser to turn the behaviour off outright.

When It Does Not Fit

Every failure follows the same shape: a line on stderr beginning error:, the usage text, and exit status 1.

What happenedWhat it says
an option nobody declaredunknown option: --colour
a value-taking option at the endoption --name expects <name>
a required option left offrequired option --output is missing
a value outside choices--level expects one of {'quiet', 'loud'}, got "shouty"
a required positional left offrequired positional argument <source> is missing
an argument that fits nothingunexpected argument: report.csv

None of these raise. A command-line program that has been handed something it cannot use has nothing useful left to do, and unwinding a stack trace into a user’s terminal tells them less than one line does. The errors that do raise are the ones in the declarations — a duplicate option name, a short name already taken, a choices that is neither a list nor a dictionary — because those are bugs in the program rather than mistakes by its user, and they raise ArgsError or a TypeError at the add_option that caused them.

Testing a Command Line

parse() takes an explicit list, and that is the seam:

import args

def build() {
  var parser = args.Parser('greet')
  parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
  parser.add_option('count', 'How many times', { short_name: 'c', type: args.INT, value: 1 })
  return parser
}

var parsed = build().parse(['-n', 'Ada', '-c', '3'])

assert parsed.options.name == 'Ada', 'the name is read'
assert parsed.options.count == 3, 'the count is coerced to a number'
assert build().parse([]).options.count == 1, 'the default applies'

echo 'all good'
all good

Build the parser in a function so each test gets a clean one; a parser carries the results of the last parse, and sharing one between tests makes them depend on their order. The paths that end in os.exit() — help, and every error above — need a real subprocess to observe, which is what os.exec() is for; Chapter 23 covers the rest of testing.

Module Reference

Building a parser:

Parser(name, default_help)a new parser
add_option(name, help, opts)an option, global or on a command
add_command(name, help, opts)a sub-command, returned for chaining
add_index(name, help, opts)a positional argument
parse(custom_args)read the command line, or a list
help()print the help text and exit 0
set_terminal_width(width)fix the wrapping width

Properties you can set after construction:

description, epilogprose above and below the options
allow_abbrevunambiguous long-option prefixes, default true
allow_atfile@file expansion, default true
terminal_widththe wrapping width directly

Keys opts understands:

short_namethe single-character form
typeone of the type constants
valuewhat it is worth when absent
choicesa list of allowed values, or a map from key to value
requiredrefuse to run without it
metavarthe word help shows in place of the value
deprecatedwarn on stderr when it is used
actionon a command, the function to run after parsing

The type constants:

NONE, STRING, INT, NUMBERa flag, and the three plain values
BOOL, LIST, CHOICE, OPTIONALan answer, a repeat, a fixed set, a maybe

Metaprogramming and Reflection

Zuri’s own compiler is available to Zuri programs. The zuri module hands you the lexer, the parser and the bytecode compiler as ordinary functions, alongside a reflection API over live functions, classes, modules and instances.

That is an unusual amount of the language to expose, so it is worth being clear about what it is for. This is the machinery behind documentation generators, linters, plugin loaders, serialisers, debuggers and editor tooling — programs whose subject is other programs. It is not machinery for ordinary code, and the chapter ends with the one thing it deliberately cannot do.

Two Halves That Cannot Answer for Each Other

The module divides cleanly, and confusing the halves is the most common mistake:

ReflectionThe compiler API
Subjectobjects that exist right nowsource text
Entry pointszuri.reflect.*tokenize(), parse(), compile()
Can tell youa function’s arity, a class’s fields, an instance’s valuesa doc block, a comment, a line number, what the compiler emitted
Cannot tell youanything about the source it came fromanything about a running value

A function’s arity is a runtime fact: the function object carries it, and reflection reads it. That same function’s doc block is a source fact: it exists only in the file, and only the parser can find it. There is no call that crosses the gap, which is why a documentation generator parses files rather than importing them.

One Shape for Everything

tokenize(), parse() and compile() all return the same shape of result: a flat list of small, uniformly tagged records.

FunctionReturnsTagPayload
tokenize(source)Token listkindvalue, text
parse(source)Node listkindfields
compile(source)Instr listopfields

There is one Token class, one Node class and one Instr class — not one class per token kind, grammar rule or opcode. The grammar has around fifty shapes and the instruction set around sixty opcodes; a class for each would be a hundred-odd near-identical classes to keep in step with the compiler forever.

The cost of that choice is real: nothing tells you which fields a given kind carries. You cannot autocomplete your way to a Binary node’s left, op and right. So the module leans on a different habit — run it and look:

import zuri

echo zuri.parse('var total = a + 1')[0].dump()
Stmt@1:1
  statement:
    Var@1:1
      name: 'total'
      name_span: 1:5-1:10
      value:
        Binary@1:13
          left:
            Identifier@1:13
              name: 'a'
          op: 'Plus'
          right:
            Integer@1:17
              value: 1
      type_hint: nil
      is_constant: false

Everything about that node is in front of you: it is a Stmt wrapping a Var, the initialiser is a Binary whose op is the string 'Plus', and the literal 1 is an Integer node rather than a generic literal. Each header names where that node starts, and name_span says exactly where the name total is. Nothing had to be looked up.

dump() is the fastest way to learn any node’s shape, and it is the first thing to reach for whenever you are unsure. zuri.dump_file(path) does the same for a whole file.

Runtime Reflection

Reflection answers questions about a value you are holding. Every entry point is under zuri.reflect.

What Kind of Thing Is This?

import zuri

def add(a, b) {
  return a + b
}

class Marker {}

echo zuri.reflect.kind(add)
echo zuri.reflect.kind(Marker)
echo zuri.reflect.kind(Marker())
echo zuri.reflect.kind(42)
echo zuri.reflect.kind(print)
function
class
instance
number
function

kind() is typeof()’s more literal cousin. Note the last line: a built-in native function reports 'function', as does a closure and a bound method, because from the language’s point of view all three are interchangeably callable.

Functions

import zuri

def add(a, b) {
  return a + b
}

var info = zuri.reflect.function_info(add)

echo info.name
echo info.arity
echo info.variadic
echo info.is_method
echo info.owning_class_name
add
2
false
false
nil

function_info() returns { name, arity, variadic, is_method, owning_class_name, source_path }. The last one is the only field that reaches back toward the source, and it is a path, not content — to read the function’s doc block you still have to parse that file.

Classes

import zuri

class Shape {
  static var sides = 0

  @new(name) {
    self.name = name
  }

  area() {
    return 0
  }

  _secret() {
    return 'hidden'
  }

  @to_json() {
    return { name: self.name }
  }
}

var info = zuri.reflect.class_info(Shape)

echo info.name
echo info.superclass_name
echo info.fields
echo info.statics
echo info.methods.keys()
Shape
nil
[name]
[sides]
[@new, area, @to_json, _secret]

Four things in that output are worth pausing on.

fields lists name, which was never declared with var — it was created by self.name = name inside @new. Reflection reports the class’s real shape, not just what the var lines said.

statics is separate from fields, because a static belongs to the class and a field belongs to each instance.

Decorated methods appear under their @ names. @new and @to_json are in methods alongside ordinary ones.

_secret is listed. Reflection sees private members; more on that below.

Each entry in methods is a full function_info, so info.methods.area.arity works without another call.

superclass_name is a string or nil, not a class:

import zuri

class Base {}
class Derived < Base {}

echo zuri.reflect.class_info(Derived).superclass_name
echo zuri.reflect.class_info(Base).superclass_name
Base
nil

Instances

import zuri

class Point {

  @new(x, y) {
    self.x = x
    self.y = y
  }

  distance() {
    return (self.x ** 2 + self.y ** 2).sqrt()
  }
}

var p = Point(3, 4)

echo zuri.reflect.get_props(p)
echo zuri.reflect.has_prop(p, 'x')
echo zuri.reflect.get_prop(p, 'x')

zuri.reflect.set_prop(p, 'x', 10)
echo p.x
[x, y]
true
3
10

get_props() lists the field names an instance actually has. get_prop(), set_prop(), has_prop() and del_prop() read and write them by computed name — the reflective equivalents of the getprop family of built-ins from Chapter 6.

They cannot create a field. A class is sealed, so set_prop() on a name the class never declared returns false and changes nothing.

Methods by Name

Three calls turn a method name into something callable:

import zuri

class Greeter {

  @new(name) {
    self.name = name
  }

  greet() {
    return 'hello ${self.name}'
  }

  @to_json() {
    return { name: self.name }
  }
}

var g = Greeter('ada')

echo zuri.reflect.has_method(g, 'greet')
echo zuri.reflect.has_decorator(g, 'to_json')
echo typeof(zuri.reflect.get_method(g, 'greet'))
echo zuri.reflect.bind_method(g, 'greet')()
true
true
function
hello ada

The distinction between the last two matters. get_method() returns the unbound function; calling it needs the instance supplied yourself. bind_method() returns it already attached to the instance, so it can be stored, passed around and called with nothing extra — which is what makes a dispatch table of methods possible:

import zuri

class Counter {

  @new() {
    self.n = 0
  }

  up() {
    self.n++
    return self.n
  }

  down() {
    self.n--
    return self.n
  }
}

var counter = Counter()

var actions = {
  up: zuri.reflect.bind_method(counter, 'up'),
  down: zuri.reflect.bind_method(counter, 'down'),
}

echo actions['up']()
echo actions['up']()
echo actions['down']()
echo counter.n
1
2
1
1

has_decorator() and get_decorator() do the same for @-prefixed methods, taking the name without the @.

Reflection Sees Private Members

Everything above ignores the leading-underscore rule:

import zuri

class Vault {
  var _combination = '1234'

  @new() {}
}

echo zuri.reflect.get_props(Vault())
echo zuri.reflect.get_prop(Vault(), '_combination')
[_combination]
1234

This is deliberate, and it is the same decision getprop() makes. A serialiser has to see every field or it writes an incomplete record; a debugger has to see every field or it shows a lie. Hiding them would defeat the only purpose these functions have.

The rule to take from it: reflection is infrastructure, not a way around encapsulation. Code that reaches for get_prop() to read a private field it could not otherwise read has not found a loophole; it has written something the next reader will not expect.

Modules

import zuri
import math

var info = zuri.reflect.module_info(math)

echo info.name
echo info.members.contains('PI')
math
true

module_info() reports a module’s name, its path and the members it exports. That is enough for a plugin loader: import a directory of modules, ask each whether it has a known function, and call the ones that do.

The Collector

The garbage collector runs on its own as a program allocates. gc() runs a full collection on the spot:

import zuri

var scratch = [1, 2, 3]
scratch = nil

zuri.reflect.gc()
echo 'collected'
collected

Every object nothing reaches is freed before gc() returns, and anything that releases a resource when collected releases it then: an ffi pointer taken over with own() runs its destructor. The memory it frees goes back to the system before it returns too. No program needs gc() to stay correct. It is for the moments timing matters, such as releasing native resources at a known point, or a test checking what collection does. A full collection visits every live object, so calling it in a loop is slow.

Tokens

tokenize() is the first stage: text in, a flat list of lexical tokens out.

import zuri

for token in zuri.tokenize('var x = 1 # note') {
  echo '${token.kind} at ${token.line}:${token.column} -> ${token.text}'
}
Var at 1:1 -> var
Identifier at 1:5 -> x
Equal at 1:7 -> =
Integer at 1:9 -> 1
Comment at 1:11 -> # note
Eof at 1:17 -> 

Nothing is filtered. Comments, doc blocks and newlines are all real tokens, and the list always ends with one Eof. The parser’s grammar skips trivia when it runs; tokenize() reports everything the lexer saw, which is exactly what a formatter or a syntax highlighter needs.

Each Token carries:

FieldWhat it is
kindthe lexer’s own variant name — 'Identifier', 'Plus', 'Comment'
line, column1-indexed position; column counts characters, not bytes
start, endcharacter offsets spanning the token’s exact text
textthe exact source text, delimiters included
valuethe payload, where there is one — a literal’s value, an identifier’s name, a comment’s content

start and end are what make edits possible: they let you recover or replace a token’s original text without re-lexing, which is how a rename tool or an automatic formatter works.

is_trivia() is true for comments and doc blocks, so filtering them is one call:

import zuri

var tokens = zuri.tokenize('var x = 1 # note')

echo tokens.filter(@(t) => t.is_trivia()).map(@(t) => t.kind)
echo tokens.filter(@(t) => !t.is_trivia()).length()
[Comment]
5

Lexing Never Raises

A malformed input does not throw. It produces a token of kind 'Error' carrying what went wrong:

import zuri

var tokens = zuri.tokenize("var s = 'unterminated")

echo tokens.map(@(t) => t.kind)
echo tokens.find(@(t) => t.kind == 'Error') != nil
[Var, Identifier, Equal, Error, Eof]
true

tokenize() is therefore total over any input at all, which is what makes it safe to point at a file you did not write, or at a half-typed buffer in an editor. Neither parse() nor compile() has that property — both raise on bad input, because neither can produce a meaningful result from it.

The Syntax Tree

parse() is the second stage: tokens become a tree.

import zuri

var tree = zuri.parse('def double(n) { return n * 2 }')

echo tree[0].kind
echo tree[0].fields.keys()
echo tree[0].fields.name
echo tree[0].line
Function
[name, name_span, parameters, body, is_variadic]
double
1

A Node has a kind, a position and a fields dictionary. The fields hold more nodes, plain lists, plain dictionaries or scalars and nothing else, so walking the tree is uniform no matter which node you are looking at.

Where Each Node Is

A node’s position covers the whole of the source it was read from. start and end are character offsets, the first character and the one just past the last, so slicing the source with them gives back the node’s own text. line and col are where it starts and end_line and end_col where it ends, all counted from 1:

import zuri

var source = 'def double(n) {\n  return n * 2\n}'
var function = zuri.parse(source)[0]
var product = function.fields.body.fields.statements[0].fields.value

echo source[product.start, product.end]
echo '${function.line}:${function.col} to ${function.end_line}:${function.end_col}'
echo source[function.fields.name_span.start, function.fields.name_span.end]
n * 2
1:1 to 3:2
double

A declaration starts at its keyword, so the function starts at def, and its name has a name_span of its own. That is everything a tool needs to underline a mistake, jump to a definition or select a whole statement.

Walking It

walk_nodes(tree, visitor) visits every node, depth first:

import zuri

var tree = zuri.parse('def double(n) { return n * 2 }')
var kinds = []

zuri.walk_nodes(tree, @(node) {
  kinds.append(node.kind)
})

echo kinds
[Function, Argument, TypeHint, Any, Block, Return, Binary, Identifier, Integer]

That output repays a second look, because it shows how much the parser makes explicit. The single unannotated parameter n still produced an Argument wrapping a TypeHint wrapping an Any — the absence of an annotation is represented in the tree, not omitted from it. A tool walking this never has to special-case “no type was written”.

The visitor may also be a dictionary from kind to handler, in which case only matching nodes are called:

import zuri

var tree = zuri.parse('def a() {}
def b() {}
var c = 1')
var names = []

zuri.walk_nodes(tree, {
  Function: @(node) {
    names.append(node.fields.name)
  },
})

echo names
[a, b]

find_nodes(tree, kind) is the shortcut when you want one kind and nothing else:

import zuri

var tree = zuri.parse('def double(n) { return n * 2 }')

echo zuri.find_nodes(tree, 'Return').length()
echo zuri.find_nodes(tree, 'Binary')[0].fields.op
1
Multiply

Comments Survive

This is the property that makes documentation tooling possible. Comments and doc blocks are kept in the tree, in their original position, as their own nodes:

import zuri

var source = "# a note
def add(a, b) {
  return a + b
}
"

echo zuri.parse(source).map(@(n) => n.kind)
[Comment, Function]

A doc block sits as a sibling immediately before whatever it documents, so pairing them is one pass with one variable of state — which is exactly what the worked example below does.

Most parsers throw comments away. Keeping them is what separates a parser you can build a formatter or a documentation generator on from one you can only build an interpreter on.

Bytecode

compile() is the third stage: the tree becomes VM instructions.

import zuri

for instr in zuri.compile('var a = 1 + 2') {
  echo '${instr.op} from line ${instr.line}'
}
LoadConst from line 0
AddImm from line 1
SetGlobal from line 1
LoadNil from line 1
Return from line 1

An Instr has an op, a line and a fields dictionary — the same shape a Node has — and walk_instrs() and find_instrs() mirror the AST walkers exactly.

Notice AddImm. The compiler folded the constant 2 into the add instruction rather than loading it separately. That is the sort of question only the bytecode can answer, and reading it is the most direct way to find out what the compiler actually made of something:

import zuri

echo zuri.compile('var a = 2 * 3 ** 2').map(@(i) => i.op)
[LoadConst, LoadConst, LoadConst, Pow, Mul, SetGlobal, LoadNil, Return]

The power comes before the multiply, which is ** binding tighter than *. One line of bytecode settles an argument that reading the expression does not.

Resolved Operands

Several opcodes carry only a raw index into the compiler’s constant pool, which on its own tells you nothing. Rather than make every caller fetch and correlate a constants table, each such field arrives with its resolved value alongside the index:

import zuri

var load = zuri.find_instrs(zuri.compile('var a = 42'), 'LoadConst')[0]

echo load.fields.keys()
echo load.fields.value
[dst, const_idx, value]
42

const_idx is the raw index, value is what it points at, and dst is the register the result lands in. The same pairing appears on GetGlobal, GetField, Invoke and the rest.

A Closure instruction’s resolved value is a nested function prototype ({ name, arity, variadic, instructions }) whose own instructions are wrapped the same way, recursively — so compiling a file with functions in it exposes their bodies too, not only the code around them.

Line Numbers Are Statement-Grained

Every instruction carries the source line it came from, but never a column. The compiler’s own tracking is line-only. A Token and a Node both have real column information from the lexer; an Instr does not.

Checking Source Without Running It

parse() and compile() raise on the first sign of bad source, which is right for a tool that needs the whole program. A tool working on source as it is being written, half a statement at a time, needs the opposite: everything that can be read, and an exact account of what is wrong.

parse_partial() reads as much as it can. Where a statement fails, the parser reports it, skips to where the next statement starts and carries on, so one mistake costs one statement and is reported once:

import zuri

var result = zuri.parse_partial('var a = 1
var = 2
var c = 3')

echo result.is_clean()
echo result.errors
echo result.nodes.map(@(n) => n.fields.statement.fields.name)
false
[2:5: Variable name expected.]
[a, , c]

Each error is a Diagnostic: a message and the same six-part position a node has, so it can be underlined exactly. The statement that failed is still in the tree, its missing name an empty string sitting where the name belongs.

check() goes one stage further. Source that parses goes on to the compiler, which catches what the grammar cannot see, such as a break with no loop around it or a name declared twice in one scope:

import zuri

var source = 'def total(items) {
  var sum = 0
  for item in items {
    sum += item
  }
  break
}'

for problem in zuri.check(source) {
  echo problem
  echo source[problem.start, problem.end]
}
6:3: 'break' used outside of a loop
break

An empty list from check() means the source compiles. Neither function runs anything or raises on bad source, so both can be pointed at whatever an editor’s buffer holds at the time.

Reading a File

Each of these has a _file counterpart that reads the path first:

Source stringFile
tokenize(source)tokenize_file(path)
parse(source)parse_file(path)
compile(source)compile_file(path)
parse_partial(source)parse_partial_file(path)
check(source)check_file(path)

There is a fourth with no string equivalent: dump_file(path) reads, parses and returns the dump() of every top-level node. It is usually the fastest possible answer to “what is actually in this file”.

A Worked Example: A Documentation Extractor

Here is the pattern the book’s own reference appendices are built on. It takes source text and returns every documented function in it, pairing each declaration with the doc block above it.

The key fact is that zuri.parse() keeps comments in the tree. A doc block comes back as its own DocBlock node, sitting as a sibling immediately before whatever it documents. zuri.doc.attach() walks a list of siblings and pairs each one with the doc block right above it, and zuri.doc reads the inside of each block into its prose and its @tags:

import zuri

def documented_functions(text) {
  var found = []

  for attached in zuri.doc.attach(zuri.parse(text)) {
    var node = attached.node

    if node.kind == 'Function' and attached.documented {
      var params = node.fields.parameters.map(@(p) => p.fields.name)

      found.append({ name: node.fields.name, params, doc: attached.block })
    }
  }

  return found
}

var source = "/**\n * Adds two numbers.\n *\n * @param number a\n * @param number b\n */\n" +
  "def add(a: number, b: number) {\n  return a + b\n}\n\n" +
  "def undocumented(x) {\n  return x\n}\n"

for fn in documented_functions(source) {
  var takes = fn.doc.all('param').map(@(tag) => '${tag.name}: ${tag.type}')

  echo '${fn.name}(${', '.join(fn.params)})'
  echo '  ' + fn.doc.summary()
  echo '  takes ' + ', '.join(takes)
}
add(a, b)
  Adds two numbers.
  takes a: number, b: number

undocumented is absent from the output: attach() hands it back with documented false, because no doc block sits above it.

Three things make this approach worth preferring over scanning the text yourself.

The parser knows what a declaration is. A function named def_handler, a def inside a string literal, a doc block inside a block comment — all of them fool a text scan and none of them fool the parser.

Parameter names and types come from the real nodes. A parameter’s type_hint node carries its types and whether it is nullable, so a generated signature matches what the runtime will actually enforce rather than what the source happened to look like.

It cannot drift. When the grammar gains something, the parser gains it too, and a tool built this way keeps working.

The tags are read for you. zuri.doc joins a tag’s wrapped lines, reads its type and the name it documents, and treats @return as @returns. It is the same reading the book’s own reference is generated with, so a block reads the same here as it does there.

A Second Example: A Linter

The other half of the module’s use is checking rather than generating. Here is a rule of the kind a team accumulates: flag every function whose parameter list is longer than some limit.

import zuri

def long_signatures(source, limit) {
  var offenders = []

  zuri.walk_nodes(zuri.parse(source), {
    Function: @(node) {
      var count = node.fields.parameters.length()

      if count > limit {
        offenders.append({ name: node.fields.name, line: node.line, count })
      }
    },
  })

  return offenders
}

var source = "def small(a, b) {}\n" +
  "def large(a, b, c, d, e) {}\n" +
  "def also_large(a, b, c, d) {}\n"

for problem in long_signatures(source, 3) {
  echo 'line ${problem.line}: ${problem.name}() takes ${problem.count}'
}
line 2: large() takes 5
line 3: also_large() takes 4

Three properties make this worth doing with the parser rather than with a regular expression over the text.

It cannot be fooled by text that looks like code. A function named def_handler, the word def inside a string literal, a commented-out declaration — none of them are Function nodes, so none of them are counted.

It reports the real line. node.line comes from the lexer, so the message points at the declaration whatever the formatting around it.

It keeps working. When the grammar gains something, the parser gains it too, and a rule written this way does not quietly stop matching.

The dictionary form of the visitor is doing real work here: only Function nodes reach the handler, so there is no if node.kind == ... and no chance of matching a kind you did not mean to.

Serialising Without a @to_json

Reflection covers the case where you would otherwise write the same method on every class:

import zuri
import json

class User {

  @new(name, age) {
    self.name = name
    self.age = age
  }
}

class Product {

  @new(title, price) {
    self.title = title
    self.price = price
  }
}

def to_record(instance) {
  var record = {}

  for field in zuri.reflect.get_props(instance) {
    record[field] = zuri.reflect.get_prop(instance, field)
  }

  return record
}

echo json.encode(to_record(User('ada', 36)))
echo json.encode(to_record(Product('desk', 120)))
{"name":"ada","age":36}
{"title":"desk","price":120}

One function, every class, no per-class method to keep in step with the fields. Note that this reads private fields too, which is right for a debugging dump and wrong for an API response — for the latter, filter on the leading underscore, or write a real @to_json and let the class decide what it exposes.

What Else This Is Good For

Plugin systems. Load modules from a directory, ask each whether it has the function your host expects, and call the ones that do. module_info() answers the question without importing blindly and hoping.

Editor tooling. tokenize(), parse_partial() and check() never raise, so they can run against a half-typed buffer. Every token and every node carries start and end, so a rename or a reformat can rewrite exact spans, and every problem check() finds can be underlined where it is.

Understanding the compiler. zuri.compile() on a snippet is faster than reading the compiler’s source, and it can never go out of date with the compiler you are actually running.

Answering questions about this book. Appendices D and E are generated from libs/ with exactly the techniques above, and the audit that checks those stubs is written the same way.

What This Is Not Good For

There is no eval(). You can compile source to bytecode and inspect it; you cannot execute a string as code.

Zuri leaves it out for security. eval() erases the line between data and code, and every program that has one eventually runs a string it did not mean to: a form field, a query parameter, a config value, a webhook payload. The moment any of those reaches an eval(), whoever supplied it is running code with your program’s full permissions. It reads your files, opens your sockets and sends your secrets anywhere it likes. This is the single most damaging vulnerability class in dynamic languages, and it keeps happening because the dangerous call looks harmless in review: one function, three characters of input, and nothing in the language warns you.

So Zuri does not provide one, and every job eval() is usually reached for has a better answer here:

Instead ofUse
evaluating a user-supplied expressiona dictionary of handlers keyed by the input
calling a method whose name you computedzuri.reflect.bind_method()
varying behaviour at runtimepass a function in
turning text into datajson.decode()

Each of those does the job with a fixed, auditable set of things that can happen. That is the difference between a program that handles input and a program that obeys it.

Classes are sealed, so reflection cannot add a method to one at runtime. Patterns from other languages that rely on monkey-patching do not translate, and the alternative is the one the language wants: express the variation as a subclass, or as a function you pass in.

Performance and the JIT

Zuri runs your bytecode in an interpreter until a piece of it gets hot, then compiles that piece to machine code and runs it there instead. Two compilers share the work, one quick and one thorough, and a hot function passes through both. The tiering happens on its own, on background threads, with no annotations and no flags.

Here is what that is worth on a plain recursive Fibonacci, measured on one idle machine:

def fib(n) {
  if n < 2 {
    return n
  }
  return fib(n - 1) + fib(n - 2)
}

var start = time()
echo fib(27)
echo 'took ${((time() - start) * 1000).round()}ms'
$ zuri run fib.zu
196418
took 14ms

$ ZURI_JIT=0 zuri run fib.zu
196418
took 47ms

Over three times, for a function that does nothing but add and compare, and you write nothing to get it. The rest of this chapter explains how the compilers work and covers the handful of cases where what you write decides whether you get it.

The Two Compilers

Kebbi compiles first. It turns a function’s bytecode into machine code one instruction at a time, keeps values in machine registers, and proves what it can about types before it compiles: a parameter annotated list is a list, and a counter that starts at 0 and steps by 1 is a whole number. Kebbi compiles quickly, and its code runs several times faster than the interpreter.

Bayelsa compiles second. It builds the whole function as a graph of typed values, then optimizes that graph before it generates any machine code. It carries a fact proven once to every later use, moves checks that cannot change out of loops, keeps loop counters as plain integers, builds small functions into the functions that call them, and bets on what a function has done so far wherever nothing proves it. Bayelsa takes longer to compile, and its code is the fastest Zuri produces.

Both compile with Cranelift, and both run on background threads. A program never waits for a compile: it carries on in the interpreter, or in the code it already has, until the new code is ready.

How a Function Moves Between Them

  1. The interpreter. Every function starts here. It counts its calls and its loop turns, and records the kind of value each operation meets.
  2. Kebbi, profiling. Once a function is warm, Kebbi compiles it with the counting and recording built in. The function is fast from here on, and still watching itself.
  3. Bayelsa. Once the profiling code has done enough work, Bayelsa compiles the function from what it recorded. A loop running at that moment moves into the new code at its next turn.

A function every operation of which has already run in the interpreter skips the profiling step and goes straight to Bayelsa, since profiling would only record what the interpreter already knows.

When a Bet Fails

Bayelsa compiles what it has seen. A loop that only ever added whole numbers does integer arithmetic; a module constant that held 10 is compared as the integer 10; an index that only ever met bytes reads a byte. Each bet is checked where it is made, and a failed check hands the frame back to the interpreter at that exact instruction, with every variable as the compiled code left it. This is a deoptimisation.

The place that failed is remembered. The next compile of the function makes no bet there, so a function deoptimises at a given place once, not every time round.

What Stays in Kebbi

Bayelsa leaves a function to Kebbi in two cases:

  • The function contains a catch. Kebbi compiles it, handlers included.
  • The function’s hot operations are ones Kebbi handles inline and Bayelsa would hand to the runtime: arithmetic on values that were not always numbers, method calls on strings, and indexing lists and strings whose kind nothing settles. Kebbi’s code is the faster of the two there.

A function left in Kebbi is rebuilt without its profiling and runs at Kebbi’s full speed.

How Tiering Works

Every function counts its own calls, and every loop counts its own back-edges. When a counter crosses that function’s threshold, a compile job goes to a background thread; when the machine code comes back, later calls use it.

The threshold is not a fixed number. It scales with the function’s size:

threshold = K / sqrt(instruction_count)

A large function gets a low threshold, because it is doing more work per call and the compilation pays for itself sooner. A three-instruction getter gets a high one, because compiling it is speculative and most tiny functions never run enough to earn it back.

Loops get their own, lower threshold through on-stack replacement. A loop inside a function that is only ever called once still compiles, and execution jumps from the interpreter into the middle of the compiled version without unwinding anything. That is what makes a one-shot batch script fast.

You can watch it happen:

$ ZURI_JIT_LOG=1 zuri run fib.zu
[jit] 'fib' goes straight to Bayelsa
[jit] compiled 'fib' in Bayelsa (13 bytecode ops, 0 osr point(s))
196418

fib goes straight to Bayelsa because every one of its operations ran in the interpreter before it warmed up. With Bayelsa off, Kebbi compiles it instead, and the line lists what Kebbi proved:

$ ZURI_JIT_BAYELSA=0 ZURI_JIT_LOG=1 zuri run fib.zu
[jit] compiled 'fib' in Kebbi (13 bytecode ops, 0 osr point(s), speculative_params=0x1, speculative_regs=0x0, speculative_lists=0x0, speculative_ints=0x1)
196418

A function Bayelsa leaves to Kebbi gets a line saying why:

[jit] 'report' stays in Kebbi: catch at ip 5

And you can see what did and did not make it:

$ ZURI_JIT_COVERAGE=1 zuri run fib.zu
=== jit coverage ===
status	calls	threshold	ops	osr	name
compiled	62009	169	11	0	fib
cold	0	56	100	0	@.script
...

calls against threshold tells you how close something came.

Keep Errors for Errors

A function with a catch compiles like any other. While no error comes through, a catch costs next to nothing: the handler is registered and dropped around the body, and the body runs as compiled code.

What costs is an error that is actually raised. A raise hands the frame back to the interpreter at that instruction, and a caught error resumes at its handler in the interpreter too. The function returns to compiled code the next time its loop comes round, or the next time it is called. For a genuine failure that is the right trade. For an outcome a loop meets every few iterations, it is not:

def parse_or_raise(i) {
  if i % 10 == 0 {
    raise ValueError('not a digit')
  }
  return i % 10
}

def parse_or_nil(i) {
  if i % 10 == 0 {
    return nil
  }
  return i % 10
}

def with_raise(n) {
  var total = 0
  iter var i = 0; i < n; i++ {
    catch {
      total += parse_or_raise(i)
    } as e {
      total += 0
    }
  }
  return total
}

def with_nil(n) {
  var total = 0
  iter var i = 0; i < n; i++ {
    var digit = parse_or_nil(i)
    if digit != nil {
      total += digit
    }
  }
  return total
}

def timed(label, work) {
  var start = time()
  work()
  echo '${label}: ${((time() - start) * 1000).round()}ms'
}

timed('raise and catch', @() {
  iter var r = 0; r < 200; r++ {
    with_raise(1000)
  }
})

timed('return nil     ', @() {
  iter var r = 0; r < 200; r++ {
    with_nil(1000)
  }
})

Two hundred rounds of a thousand iterations, one in ten of them failing:

VariantTime
raise, caught in the loop73ms
nil returned and checked11ms

When failing is an expected outcome, input that might not parse or a key that might be missing, return a value that says so and test it. Keep raise for what the caller cannot reasonably carry on from. A raise that never fires, a guard clause at the top of a function, costs the compiled path nothing.

Type Annotations Are a Performance Feature

The JIT speculates on the types it has seen. When it has been told instead, it does not need to guess, and the guard disappears.

def sum_untyped(values) {
  var total = 0
  iter var i = 0; i < values.length(); i++ {
    total += values[i]
  }
  return total
}

def sum_typed(values: list) {
  var total = 0
  iter var i = 0; i < values.length(); i++ {
    total += values[i]
  }
  return total
}

Twenty thousand rounds over a thousand-element list:

VariantTime
values: list85ms
no annotation150ms

Running them in both orders gives the same answer, which is the check that tells you it is the annotation and not the warm-up.

Annotate the parameters of any function on a hot path. It costs one word per parameter and it documents the function at the same time.

Keep Types Stable

Speculation works by assuming the future looks like the past. A variable that holds a number on ten thousand iterations and a string on the ten thousand and first causes a deoptimisation: the compiled code bails to the interpreter, and the function may be recompiled with a weaker assumption.

One deopt is cheap. A loop that deopts every iteration is slower than never compiling at all.

In practice this means:

  • One variable, one kind of value.
  • A list of numbers, not a list of numbers-and-sometimes-strings.
  • A field that starts nil and later holds a number is two types. Start it at 0.

That last one is the common case. var count on a class field is a nil until the constructor runs, and every read of it in compiled code has to handle both. var count = 0 does not.

Where to Put a Hot Loop

A per-element loop belongs in a typed free function, not in a method reading a field:

Slower, because every iteration reads self.pixels back through a field guard:

class Canvas {
  var pixels = bytes(0)

  brighten(amount) {
    iter var i = 0; i < self.pixels.length(); i++ {
      self.pixels[i] = self.pixels[i] + amount
    }
  }
}

Faster, because the loop sees plain locals whose types are declared:

def _brighten(pixels: bytes, amount: number) {
  iter var i = 0; i < pixels.length(); i++ {
    pixels[i] = pixels[i] + amount
  }
}

class Canvas {
  var pixels = bytes(0)

  brighten(amount) {
    _brighten(self.pixels, amount)
  }
}

The standard library’s imagine module is written this way throughout, and the difference on a per-pixel loop is measured in multiples, not percent.

Reading the Bytecode

When you want to know what the compiler actually did, ask it:

import zuri

echo zuri.compile('var a = 1 + 2').map(@(i) => i.op)
[LoadConst, AddImm, SetGlobal, LoadNil, Return]

AddImm rather than a separate load tells you the constant was folded into the instruction. This is the fastest way to check whether a rewrite did what you hoped, and it never goes stale.

Measuring Honestly

Every number in this chapter is one measurement on one machine. Treat them as ratios rather than as figures to reproduce: yours will differ with your processor, your build and what else the machine is doing. Five rules make the difference between a measurement and a guess.

Warm up before you time. The first few hundred iterations run interpreted, and if your benchmark is short, that is all you measured.

Run each variant on its own. Two variants in one process share warm-up state and a heap. Order them both ways; if the answer changes, you measured the order.

Watch the machine. A laptop under sustained load throttles, and a second run is not comparable to the first. Check the load before the timed run, not before the build that precedes it.

Change one thing. A “fix” that touches three sites has three possible explanations for its effect, and at least one of them is usually a regression hiding behind the other two.

Measure the real program. A microbenchmark of one function tells you about that function in isolation, with a warm cache and no competing allocation. ZURI_JIT_COVERAGE=1 on the actual workload tells you which function to look at in the first place, which is nearly always the more valuable answer.

Timing Something Yourself

time() returns epoch seconds with microsecond resolution, so a timing harness is four lines:

def timed(label, work) {
  var start = time()
  var result = work()

  echo '${label}: ${((time() - start) * 1000).round()}ms'

  return result
}

var squares = timed('build a list', @() {
  var out = []

  iter var i = 0; i < 200000; i++ {
    out.append(i * i)
  }

  return out
})

echo squares.length()

The elapsed line will differ every run; the length will not. Print both, so a change in the second tells you the harness broke rather than the code getting faster.

Choosing a Mode

Both compilers run by default. ZURI_JIT_BAYELSA=0 turns Bayelsa off, and every hot function then stays in Kebbi.

Keep the default for anything that runs long enough for its speed to matter: servers, batch jobs, numeric work, anything whose hot loops run for more than a moment. That is where Bayelsa’s code repays its compile many times over.

Turn Bayelsa off when:

  • The program is short. A script that finishes in a fraction of a second spends a real share of its life compiling a second time, and the faster code arrives too late to pay for itself.
  • Cores are scarce. Bayelsa’s compiles take more processor time than Kebbi’s. On a machine or container held to one core, or with every core busy with isolates, that time comes out of the program’s own.
  • Start-up is the workload. A command-line tool run over and over, a few milliseconds each time, gains nothing from code that is faster on its thousandth iteration.
  • Timings have to hold from the first run. Kebbi reaches its speed sooner and stays there. Under Bayelsa a function speeds up once more, part way through a run.

Measure both. The switch is one variable, and running the real workload each way settles the question for that workload.

The Environment Variables

These exist for measurement and debugging. Ordinary programs need none of them.

VariableEffect
ZURI_JIT=0disable the JIT entirely
ZURI_JIT_BAYELSA=0compile with Kebbi alone
ZURI_JIT_LOG=1one line per compilation attempt
ZURI_JIT_LOG_IR=1dump the Cranelift IR, and Bayelsa’s own before it
ZURI_JIT_COVERAGE=1a table of what compiled, at exit
ZURI_JIT_NO_SPECIALIZATION=1compile, but do not speculate on types
ZURI_JIT_THREADS=nbackground compiler threads
ZURI_JIT_CALL_K, ZURI_JIT_OSR_Kthe warm-up curve constants
ZURI_JIT_TIERUP_Khow much profiling work comes before Bayelsa
ZURI_GC_LOG=1garbage collector activity
ZURI_GC_NURSERY_MB=nthe young generation’s starting and smallest size, in megabytes
ZURI_OPCODE_PROFILE=1interpreter opcode histogram
MIMALLOC_PURGE_DELAY=nmilliseconds freed memory waits before going back to the system; 10 unless set

ZURI_JIT=0 is the most useful one. Running a benchmark with and without it tells you immediately whether your hot path is being compiled at all, which is the first question to ask when something is slower than it should be.

When to Stop

The interpreter is fast and the JIT is automatic. Most Zuri code needs no performance work at all, and the code that does usually needs exactly one of the three things in this chapter: return a value instead of raising for an expected outcome, annotate a parameter, or stop putting two types in one variable.

Reach for anything more exotic only after ZURI_JIT_COVERAGE has told you which function is actually the problem.

Testing

A test is a program that runs your program and says whether it did the right thing. The test module gives you the pieces: a way to name and group tests, a way to state what you expected, and a report that tells you which expectation failed and where.

Nothing needs installing. import test and you have it.

Your First Test

import test { * }

def subtotal(items) {
  return items.reduce(@(total, item) {
    return total + item.price * item.quantity
  }, 0)
}

describe('subtotal', @{

  it('is zero for an empty cart', @{
    expect(subtotal([])).to_be(0)
  })

  it('multiplies price by quantity', @{
    expect(subtotal([{ price: 250, quantity: 3 }])).to_be(750)
  })

  it('adds every line together', @{
    expect(subtotal([
      { price: 250, quantity: 3 },
      { price: 100, quantity: 1 },
    ])).to_be(850)
  })

})

Run it the way you run anything else:

$ zuri run cart.zu

  subtotal
    ✓ is zero for an empty cart
    ✓ multiplies price by quantity
    ✓ adds every line together

   PASS

  3 passed  •  3 total
  suites 1   time 1ms

Three things are worth noticing straight away.

import test { * }. A test file wants describe, it, expect and the rest in scope, not behind a module name. import test on its own works too, and gives you test.describe, test.expect and so on; use that from code that is not itself a test file.

Nothing says “now run”. describe and it do not run anything; they build a tree, and the tree is run once the file has finished declaring it, through os.at_exit(). That separation is what lets the framework count the tests before the first one starts, focus on one of them, filter by name, and run them in a random order. run() exists for when you want to pass options or read the result, and Controlling a Run comes back to it.

@{ ... } is a function. it('name', @{ ... }) hands the framework a body to call later, which is why it can decide not to.

When It Fails

Change 850 to 900 and run it again:

  subtotal
    ✓ is zero for an empty cart
    ✓ multiplies price by quantity
    ✗ adds every line together

  Failures

  1) subtotal › adds every line together

     Expected 850 to be 900

     + expected  900
     - received  850

     at cart.zu:23 in @anon3

       21 |       { price: 250, quantity: 3 },
       22 |       { price: 100, quantity: 1 },
     > 23 |     ])).to_be(900)
       24 |   })
       25 |

   FAIL

  1 failed  •  2 passed  •  3 total
  suites 1   time 4ms

The message, the two values, the line, and the code around it. The process exits with status 1, which is what a CI system reads.

Grouping

describe nests as deeply as the thing you are describing:

describe('Cart', @{

  describe('subtotal()', @{
    it('is zero when empty', @{ ... })
  })

  describe('add()', @{
    it('appends a line', @{ ... })
    it('merges a duplicate SKU', @{ ... })
  })

})

The nesting shows up in the report and in the name a failure is filed under: Cart > add() > merges a duplicate SKU. Choose names that read as a sentence when joined like that, because that is how you will read them.

Expectations

expect(value) gives you an object with every matcher on it.

expect(total).to_be(850)
expect(cart).to_have_length(2)
expect(user).to_match_object({ name: 'Ada' })
expect(items).not.to_be_empty()

.not negates any matcher. There is one implementation behind both directions, so to_contain and not.to_contain can never disagree about what containing means.

Matchers chain, because each returns the object it was called on:

expect(port).to_be_int().to_be_between(1024, 65535)

A second argument names the value, which earns its keep the moment the number alone would not say which number it was:

expect(response.status, 'status').to_be(200)
Expected status (404) to be 200

Choosing Between to_be and to_equal

to_be is Zuri’s ==. That already compares lists, dictionaries and bytes by their contents, so it is the right matcher for most things:

expect([1, 2]).to_be([1, 2])          # passes
expect({ a: 1 }).to_be({ a: 1 })      # passes

It compares two instances by identity, though, unless their class defines @eq, because that is what == does:

expect(Point(1, 2)).to_be(Point(1, 2))      # fails without @eq: two objects
expect(Point(1, 2)).to_equal(Point(1, 2))   # passes: same contents

to_equal is the structural one. It walks an instance field by field, treats NaN as equal to itself, and survives a value that contains itself. Reach for it whenever the objects are built fresh on both sides of the comparison.

The Matchers

GroupMatchers
Equalityto_be, to_equal, to_be_same_as, to_be_close_to, to_be_within, to_match_object
Truthinessto_be_true, to_be_false, to_be_truthy, to_be_falsy, to_be_nil, to_be_defined
Typesto_be_a, to_be_string, to_be_number, to_be_int, to_be_float, to_be_bigint, to_be_bool, to_be_list, to_be_dict, to_be_bytes, to_be_function, to_be_callable, to_be_iterable, to_be_class, to_be_instance, to_be_instance_of
Numbersto_be_greater_than, to_be_greater_than_or_equal, to_be_less_than, to_be_less_than_or_equal, to_be_between, to_be_positive, to_be_negative, to_be_zero, to_be_divisible_by, to_be_even, to_be_odd, to_be_nan, to_be_finite, to_be_infinite
Textto_contain, to_contain_ignoring_case, to_equal_ignoring_case, to_start_with, to_end_with, to_match, to_be_blank
Collectionsto_have_length, to_be_empty, to_contain_equal, to_contain_all, to_contain_any, to_contain_none, to_contain_exactly, to_have_key, to_have_keys, to_have_value, to_have_property, to_be_sorted, to_have_unique_items
Errorsto_raise, to_raise_instance_of, to_raise_with_message, to_not_raise
Mocksto_have_been_called, to_have_been_called_times, to_have_been_called_with, to_have_been_last_called_with, to_have_been_nth_called_with, to_have_returned, to_have_returned_with, to_have_raised
Outputto_print, to_print_exactly, to_print_nothing
Snapshotsto_match_snapshot
Anything elseto_satisfy

Each one’s exact behaviour, and what it refuses, is in its doc block.

A few are worth calling out.

to_be_close_to is how you compare anything that has been through floating-point arithmetic. expect(0.1 + 0.2).to_be(0.3) fails; the sum is 0.30000000000000004.

expect(0.1 + 0.2).to_be_close_to(0.3)

to_raise* takes a function, not a value, because the framework has to be the one to call it:

expect(@{ parse('') }).to_raise_instance_of(ValueError)
expect(@{ parse('{}') }).to_not_raise()

to_have_property walks a dotted path through dictionaries, instances and list indices alike, which is what a decoded response usually is all the way down:

expect(payload).to_have_property('data.items.0.id', 7)

to_satisfy is the escape hatch when nothing else fits:

expect(port).to_satisfy(@(n) { return n % 2 == 0 }, 'to be an even port')

Failing on Purpose

using response.status {
  when 200 handle_ok()
  when 404 handle_missing()
  default fail('unexpected status ${response.status}')
}

And when the assertions live inside a callback that might never run, say how many you expect:

it('reports every error', @{
  assertions(2)

  validate(bad_input, @(error) {
    expect(error.field).to_be_string()
  })
})

If the callback ran once instead of twice, the test fails with Expected 2 assertions but 1 ran rather than passing on a technicality. has_assertions() is the looser form: at least one.

Setup and Teardown

import test { * }

describe('Session', @{

  var store = nil

  before_all(@{
    echo 'connecting'
  })

  after_all(@{
    echo 'disconnecting'
  })

  before_each(@{
    store = { rows: [] }
  })

  it('starts empty', @{
    expect(store.rows).to_be_empty()
  })

  it('is a fresh store every time', @{
    store.rows.append('a')
    expect(store.rows).to_have_length(1)
  })

})
$ zuri run session.zu
connecting
  Session
    ✓ starts empty
    ✓ is a fresh store every time
disconnecting

   PASS

  2 passed  •  2 total
  suites 1   time 2ms

The rules:

  • before_all runs once, immediately before the first test in its suite that is actually going to run. A suite everything was filtered out of never connects to anything.
  • after_all runs after the last one, and only if before_all ran. It runs when before_all raised as well, since a setup that failed part way may already hold a connection or a server, so write it to cope with whatever the setup got as far as.
  • before_each runs outermost suite first, after_each innermost first, so teardown undoes setup in the order it was done.
  • after_each runs even when the test failed, which is exactly when you need it to.

A hook that raises is reported as a hook failure against the test it was preparing, rather than taking the run down. A before_all that raises fails every test in its suite, because none of them got the setup they were written against.

Focusing, Skipping and Todo

While you are working on one thing:

it_only('the one I am fixing', @{ ... })
describe_only('the area I am in', @{ ... })

As soon as anything is marked only, everything else is skipped and still listed, so you cannot forget it is on.

it_skip('known broken, see #412', @{ ... })
describe_skip('the old API', @{ ... })

it_todo('handle an empty payload')

A todo is a test with no body. it('name') with nothing after it means the same thing. It counts towards nothing and shows up in the report as a reminder that survives being committed.

Two options change what a failure means:

it('reproduces issue 412', @{ ... }, { failing: true })
it('talks to a flaky endpoint', @{ ... }, { retries: 2 })

failing: true passes when the body fails, and fails when the body passes, so the day someone fixes the bug the test tells you. retries runs the body again on failure, and a test that passes on a later attempt is reported as flaky rather than quietly green.

One Test, Many Inputs

import test { * }

def slug(title) {
  return title.lower().replace('/[^a-z0-9]+/', '-').trim('-')
}

it_each([
  ['Hello World', 'hello-world'],
  ['  Spaced  Out  ', 'spaced-out'],
  ['Zuri 1.0!', 'zuri-1-0'],
], 'turns $0 into $1', @(title, expected) {
  expect(slug(title)).to_be(expected)
})
$ zuri run slug.zu
  ✓ turns 'Hello World' into 'hello-world'
  ✓ turns '  Spaced  Out  ' into 'spaced-out'
  ✓ turns 'Zuri 1.0!' into 'zuri-1-0'

   PASS

  3 passed  •  3 total
  suites 0   time 3ms

Each row becomes a separate test with its own name and its own place in the report, so one bad row does not hide the others. $0, $1 and so on stand for the row’s values and $# for the row number. describe_each does the same for whole suites, which is how you run one set of tests against several implementations of the same interface.

Test Doubles

mock() gives you a function that records how it was called and does whatever you tell it to.

import test { * }

def retry(operation, attempts) {
  iter var attempt = 1; attempt <= attempts; attempt++ {
    catch {
      return operation()
    } as error {
      if attempt == attempts {
        raise error
      }
    }
  }
}

describe('retry', @{

  it('returns the first success', @{
    var operation = mock()
    operation.returns('ok')

    expect(retry(operation.fn, 3)).to_be('ok')
    expect(operation).to_have_been_called_times(1)
  })

  it('tries again after a failure', @{
    var operation = mock()
    operation.raises_once(Error('connection reset'))
    operation.returns('ok')

    expect(retry(operation.fn, 3)).to_be('ok')
    expect(operation).to_have_been_called_times(2)
  })

  it('gives up eventually', @{
    var operation = mock()
    operation.raises(Error('connection reset'))

    expect(@{ retry(operation.fn, 3) }).to_raise_with_message('connection reset')
    expect(operation).to_have_been_called_times(3)
  })

})

mock() returns a Mock, and mock.fn is the plain function you hand to the code under test. The matchers accept either, so expect(operation) and expect(operation.fn) mean the same thing.

The behaviour is scripted in front-to-back order: everything queued with a _once suffix runs first, one call each, and then the standing behaviour takes over.

MethodEffect
returns(v) / returns_once(v)return v
raises(e) / raises_once(e)raise e
implements(fn) / implements_once(fn)run fn with the real arguments
reset()forget the calls
clear_behaviour()forget the behaviour

Spying on Something That Already Exists

spy_on swaps a function out in place and gives you a Mock that both records and stands in for it:

import test { * }

def checkout(cart, gateway) {
  var total = cart.reduce(@(sum, line) { return sum + line }, 0)
  return gateway['charge'](total)
}

it('charges the cart total once', @{
  var gateway = { charge: @(amount) { return 'live-charge' } }
  var charge = spy_on(gateway, 'charge')
  charge.returns('receipt-1')

  expect(checkout([250, 100], gateway)).to_be('receipt-1')
  expect(charge).to_have_been_called_times(1)
  expect(charge).to_have_been_called_with(350)
})

A spy calls through to the original unless you tell it otherwise, and is restored automatically after the test that installed it, even one that failed halfway through. Nothing has to be undone by hand.

It works on a dictionary entry and on an instance property that holds a function. It does not work on a class method or a module function: Zuri classes are immutable once declared, and a module’s members cannot be assigned from outside it. Code you want to substitute takes its collaborators as arguments or holds them in properties, which is the shape worth designing for anyway.

Snapshots

For a value too big to write out by hand, record it once and compare against the recording from then on:

def invoice(customer, lines) {
  return {
    customer,
    lines,
    total: lines.reduce(@(sum, line) { return sum + line.amount }, 0),
  }
}

it('builds the document', @{
  expect(invoice('Ada', [{ label: 'Design', amount: 4200 }])).to_match_snapshot()
})

The first run writes the file and passes:

$ zuri run invoice.zu

  invoice
    ✓ builds the document

   PASS

  1 passed  •  1 total
  suites 1   time 2ms
  snapshots 1 written

It lands beside the test file, in __snapshots__:

# Zuri snapshot file v1

=== invoice > builds the document 1 ===
  {
    customer: 'Ada',
    lines: [
      {
        amount: 4200,
        label: 'Design'
      }
    ],
    total: 4200
  }

Commit it. Reviewing the change to that file in a pull request is the entire value of the technique: an unexplained diff there is exactly the thing worth noticing.

When the value changes, you get the diff:

  1) invoice › builds the document

     Snapshot 'invoice > builds the document 1' no longer matches

       + expected  - received

       {
         customer: 'Ada',
         lines: [
           {
     -       amount: 5200,
     +       amount: 4200,
             label: 'Design'
           }
         ],
     -   total: 5200
     +   total: 4200
       }

If the new value is right, rewrite the snapshots:

$ ZURI_UPDATE_SNAPSHOTS=1 zuri run invoice.zu

That also deletes entries nothing asks for any more. And on CI, where CI is set in the environment, writing a brand new snapshot is a failure rather than a silent pass, because a snapshot nobody has looked at asserts nothing.

Testing What Something Prints

expect(@{ greet('Ada') }).to_print('Hello, Ada')
expect(@{ quiet_mode() }).to_print_nothing()

Or take the output and assert on it yourself:

var printed = capture_output(@{ report(rows) })
expect(printed.lines()).to_have_length(4)

This works because the runtime can redirect everything Zuri writes to standard output. The same mechanism is why a passing test’s output does not clutter the report: it is captured and shown only when the test fails. run({ verbose: true }) shows it either way, and run({ capture: false }) turns it off entirely.

More Than One File

A real project has a directory of them:

project/
  tests/
    cart.zu
    pricing.zu

zuri test runs the lot:

$ zuri test

  zuri test  2 files in tests

   PASS   cart.zu  188ms  2 tests

   FAIL   pricing.zu  281ms  1 test
        ✗ pricing > applies the discount
          Expected 90 to be 100
          at tests/pricing.zu:4

  2 files  •  1 with failures
  1 failed  •  2 passed  •  3 total
  time 476ms

There is nothing to write for this. The command finds the tests directory you ran it from, and ends the process with 1 when anything failed, so a CI job needs nothing added to it either.

Each file runs in a process of its own. That is not an implementation detail you can ignore, because it is what you are buying:

  • A file that loops forever is killed, and the rest still run. --timeout 30s says how long is too long.
  • A file that crashes, or calls os.exit() halfway through, is reported as a file that never reported rather than taking the run with it.
  • Global state, a module loaded for its side effect, a changed working directory: none of it leaks from one file into the next.

Files are reported in the order they were discovered whatever order they finish in, so --jobs 4 makes a big suite faster without making the report move around.

A test file needs nothing special to be conducted. It declares its tests, and is equally runnable on its own.

Running One File

Name it, with or without its .zu:

$ zuri test pricing
$ zuri test pricing.zu
$ zuri test tests/pricing.zu

All three run tests/pricing.zu. A name is looked for under tests first, so a test keeps its own name even when something else in the project shares it, and a name that is nowhere to be found under that directory is looked for once more by filename alone anywhere beneath it, so a file in a subdirectory answers to its own name.

A directory works too, wherever it sits:

$ zuri test tests/api
$ zuri test packages/store/tests

The Flags

FlagWhat it does
-j, --jobs <count>How many files to run at once. auto is one per CPU. Default 1.
-t, --timeout <duration>How long to give a single file before killing it. 500ms, 30s, 2m, 1h, or a bare number of milliseconds. Default: no limit.
-b, --bail [count]Stop after this many failing files. On its own, stop at the first.
-m, --match <pattern...>Filename patterns to run, instead of *.zu.
-i, --ignore <pattern...>Filename patterns to skip, instead of _*, .* and index.zu.
--no-recursiveOnly the files directly in the directory.
-e, --env <assignment...>Extra environment for every test process, as KEY=VALUE.
-l, --listPrint the files that would run, one to a line, and stop.
$ zuri test --jobs auto --timeout 30s
$ zuri test --bail
$ zuri test --match '*_test.zu' '*_spec.zu'

--bail and the two pattern flags take as many words as follow them, so put the file you are naming ahead of them, or close them with --:

$ zuri test pricing --bail
$ zuri test --match '*_test.zu' -- pricing

A count for --bail has to be attached, --bail=3, for the same reason: a count written as a separate word is indistinguishable from the file.

Conducting a Suite Yourself

conduct() is the function zuri test is built on, and a project that wants the run under its own control can call it directly. A directory handed to zuri run runs its index.zu, so one file makes the suite zuri run tests:

Filename: tests/index.zu

import os
import test

test.conduct(os.dir_name(__file__), { jobs: 4, timeout: 30000 })

It takes the same choices the flags do, as a dictionary, and returns the run rather than only reporting it. conduct leaves index.zu out of discovery, and never runs the script that called it either, so the index cannot end up running itself.

Controlling a Run

Call run() yourself when you want to change how the run behaves, or to read the result. It takes an options dictionary, and calling it takes over from the automatic run, so nothing happens twice.

The options you will reach for:

OptionWhat it does
filterrun only tests whose full name contains this, or matches it as a regular expression
bailstop after this many failures
shuffle / seedrun in a random order, and reproduce that order later
reporterspec, dot, tap, junit, json, ndjson, silent
exitexit the process with the run’s status when it finishes
run({ filter: 'subtotal', bail: 1 })

filter and reporter can also come from ZURI_TEST_FILTER and ZURI_TEST_REPORTER, which is what lets one command be pointed at one test without editing the file.

Shuffling is the one worth turning on deliberately. Tests that only pass because an earlier test left something behind are a real and common problem, and running them in a different order every time is how you find out:

run({ shuffle: true })

The seed is printed with the summary, and passing it back reproduces that exact order.

On CI

Nothing to add. A test file exits 0 when everything passed and 1 when it did not, whether or not it called run().

For a CI system that wants a machine-readable report, { reporter: 'junit' } writes JUnit XML to standard output and { reporter: 'tap' } writes TAP version 14.

Writing Your Own Reporter

The runner knows nothing about output. It walks the tree and calls methods on a reporter, and every method has a do-nothing default:

import test { * }
import test.reporter { Reporter }

class Quiet < Reporter {
  test_finished(one) {
    if one.status == 'failed' {
      echo one.full_name()
    }
  }
}

describe('a suite', @{
  it('passes quietly', @{ expect(1).to_be(1) })
  it('fails loudly', @{ expect(1).to_be(2) })
})

run({ reporter: Quiet(), exit: false })

The objects handed to it are the same ones the built-in reporters see: Case, Suite, Failure and Summary, documented in test.result.

Two Things to Know

A timeout on a test is measured after the body returns. Zuri runs synchronously, so a test runs to completion and is then timed. That catches one that got too slow; it does not rescue one that hangs. A hang needs a process boundary, which is exactly what conduct puts around each file, and why its timeout is the enforcing one.

Everything in one file shares one interpreter. Reset what you change in before_each, and reach for shuffle to find out whether you missed anything.

Where to Go Next

The module’s own documentation carries the full matcher list with each one’s edge cases, every option run() and conduct() accept, and the snapshot file format. Appendix F lists the submodules. Chapter 24 is what to do once a test has told you something is wrong.

Debugging

Something is not doing what you expected. This chapter is about closing that gap: reading what the runtime tells you, recognising the handful of messages that account for most confusion, and the techniques that turn a vague “it’s broken” into a line number.

Reading an Error

An uncaught error prints three things: what went wrong, where, and how you got there.

Unhandled ValueError: bottomed out
  --> /path/to/main.zu:3

  1 | def recurse(n) {
  2 |   if n <= 0 {
> 3 |     raise ValueError('bottomed out')
  4 |   }
  5 |   recurse(n - 1)

Stack trace (most recent call last):
  at recurse() /path/to/main.zu:3
  at recurse() /path/to/main.zu:5
  ... 19 more frames ...
  at recurse() /path/to/main.zu:5

The type and message come first. Then the source around the failure, with the offending line marked. Then the call stack, innermost first.

A deep stack is truncated in the middle, because the top and the bottom are the parts that tell you anything: the top is where it broke, the bottom is where you started, and two hundred identical recursive frames in between are noise.

The process exits with status 1.

The Messages You Will Actually See

Every message below is real output. Run this and you get all of them:

class Config {

  @new() {
    self.name = 'default'
  }
}

def needs_string(s: string) {
  return s
}

def show(label, work) {
  catch {
    work()
  } as e {
    echo '${label}: ${e.type} — ${e.message}'
  }
}

show('name never declared', @() => undeclared_name)
show('key not in dict', @() { var d = { a: 1 }; return d['b'] })
show('field not on class', @() { var c = Config(); c.nope = 1 })
show('property not on class', @() { var c = Config(); return c.missing })
show('method on nil', @() { var x = nil; return x.f() })
show('operator on nil', @() => nil + 1)
show('index past the end', @() => [1][9])
show('wrong argument type', @() => needs_string(5))
name never declared: UndefinedError — undefined global 'undeclared_name'
key not in dict: PropertyError — undefined key 'b' in dict
field not on class: PropertyError — undefined field 'nope' on instance of 'Config'
property not on class: PropertyError — undefined property 'missing' on instance of 'Config'
method on nil: TypeError — object of type nil does not define method 'f'
operator on nil: TypeError — operator '+' not defined for call signature (nil, number)
index past the end: RangeError — index 9 out of bounds (length 1)
wrong argument type: TypeError — needs_string() expects parameter 's' (argument 1) to be a string, got number

What each one usually means in practice:

MessageWhat to look for
undefined global 'x'a typo, or a def that appears below the top-level line calling it
undefined key 'x' in dicta key that is genuinely absent; get(key, fallback) is the fix when absence is legal
undefined field 'x' on instance of 'C'a typo in a field name — classes are sealed, so this cannot create one
undefined property 'x' on instance of 'C'reading a field or method the class never declared
undefined member 'x' on module mthe module did not export it, or the export needs an @
object of type nil does not define method 'f'something upstream returned nil
operator '+' not defined for call signature (nil, number)the same, one step earlier
'x' is private and can only be accessed via 'self' or 'parent'a leading underscore, reached from outside
'x' is already declared in this scopetwo vars of one name in one block
cannot assign to constant 'x'writing to a local const
module 'x' could not be foundusually a missing leading . on a sibling import

The two nil messages are worth internalising. Zuri never tells you where a nil came from, because by the time it causes trouble the value has already been passed along. When you see one, stop looking at the line that failed and look at whatever produced the value.

Techniques

Echo the Value and Its Type

Dynamic typing means the surprise is almost always that something is not what you assumed it was:

var value = '42'

echo typeof(value)
echo value
echo value + 1
echo value.to_number() + 1
string
42
421
43

One line of typeof() would have saved the third line’s confusion.

Give a Class @to_string()

echo shows an instance through its class’s @to_string(), and as <instance of Point> when the class has none:

class Point {

  @new(x, y) {
    self.x = x
    self.y = y
  }

  @to_string() {
    return '(${self.x}, ${self.y})'
  }
}

var p = Point(1, 2)

echo p
echo [p, Point(3, 4)]
(1, 2)
[(1, 2), (3, 4)]

Define @to_string() on any class you expect to look at while debugging. The five minutes it costs are repaid the first time you print a list of them.

Encode Nested Data Instead of Echoing It

A deep dictionary printed by echo is one unreadable line. json.encode() with compact off indents it:

import json

var request = {
  method: 'POST',
  headers: { accept: 'application/json' },
  body: { title: 'write chapter 19', tags: ['docs', 'zuri'] },
}

echo json.encode(request, false)
{
  "method": "POST",
  "headers": {
    "accept": "application/json"
  },
  "body": {
    "title": "write chapter 19",
    "tags": [
      "docs",
      "zuri"
    ]
  }
}

Add Context on the Way Up

An error raised deep in a call chain says what failed, not what you were doing at the time. Catch it, say what you were doing, and re-raise:

def parse_port(raw) {
  if !raw.match('/^\d+$/') {
    raise ValueError('not a number')
  }

  return raw.to_number()
}

def load_settings(source, raw) {
  catch {
    return { port: parse_port(raw) }
  } as e {
    raise ValueError('while reading ${source}: ${e.message}')
  }
}

catch {
  load_settings('config.json', 'eighty')
} as e {
  echo e.message
}
while reading config.json: not a number

“not a number” is a fact. “while reading config.json: not a number” is a fact you can act on.

The trace is a list on the error, so a handler can log it without letting the program die:

def inner() {
  raise ValueError('deep')
}

def outer() {
  inner()
}

catch {
  outer()
} as e {
  for frame in e.stacktrace {
    echo frame
  }
}
/path/to/main.zu:2 -> inner()
/path/to/main.zu:6 -> outer()
/path/to/main.zu:10 -> @.script()

Ask the Compiler What It Made of Your Code

When an expression does not behave the way you read it, the bytecode settles the argument:

import zuri

echo zuri.compile('var a = -2 ** 2').map(@(i) => i.op)
[LoadConst, LoadConst, Pow, Neg, SetGlobal, LoadNil, Return]

Read the order: the power happens before the negation. ** binds tighter than unary minus, so -2 ** 2 is -(2 ** 2), which is -4, and not the 4 a reader expecting (-2) ** 2 would get. The bytecode settles it in one line. Chapter 21 covers zuri.compile() and zuri.parse() properly.

Narrow It With assert

An assert is a claim you can leave in the code:

def average(numbers) {
  assert !numbers.is_empty(), 'average() needs at least one number'

  return numbers.reduce(@(a, b) => a + b, 0) / numbers.length()
}

echo average([2, 4, 6])

catch {
  average([])
} as e {
  echo '${e.type}: ${e.message}'
}
4
AssertError: average() needs at least one number

Without it, average([]) would have returned NaN and the problem would have surfaced somewhere else entirely, in a value that looks like a number.

Watch the Collector

ZURI_GC_LOG=1 reports garbage collector activity, which is the one to reach for when memory rather than logic is the question:

$ ZURI_GC_LOG=1 zuri run main.zu

Chapter 22 lists the rest of the runtime’s diagnostic switches alongside what each one measures.

The Traps Worth Knowing by Heart

These are the behaviours that produce a wrong answer rather than an error, which makes them far more expensive to find.

Zero is falsy. var n = count or 10 turns a real 0 into 10, and if index is false for the first position and true for -1, “not found”. Compare explicitly.

to_number() returns 0 for text it cannot parse. 'eighty', '' and '12abc' all become 0, and so does ' 7 ' with its spaces. There is no NaN and no error to catch, so a bad input silently becomes a valid zero. Validate the text before converting it.

[] and {} are truthy. if items is true for an empty list. Use is_empty().

sort() mutates; reverse() does not. var s = items.sort() leaves items sorted as well. Clone first when you need both orders.

A method call on a string result was discarded. name.trim() does nothing on its own; strings are immutable, so you must assign the result.

A nested def is local. A function declared inside another is scoped to it, like a var, and nothing outside that scope can call it.

x++ evaluates to the new value. Unlike C and JavaScript.

Structuring Code You Can Reason About

Three habits pay for themselves the first time something breaks.

Separate the decision from the effect. A function that reads a file, parses it and decides something is three functions. Split them, and the parsing and the decision can both be exercised without a filesystem in the way.

Pass dependencies in. A function that calls time() behaves differently every second. One that takes a timestamp behaves identically every time you call it with the same number, which means you can reproduce a failure instead of waiting for it.

Return values instead of printing them. echo inside a function is invisible to its caller and useless to anything that wants to check the result. Return the string and let the caller decide.

A Full-Stack Task Board

We are going to build a real web application: a shared task board with three columns, a browser interface, and a JSON API over the same data.

It is about four hundred lines of Zuri, and it uses almost everything this book has covered. Classes and inheritance for the domain model. Custom errors for validation. The module system for structure. Files and JSON for persistence. The HTTP server for routing, middleware and static files. Wire for server-rendered HTML. os for configuration and signals. log for output.

No dependencies. Nothing to install. zuri run taskboard and it runs.

What It Does

  • Three columns: todo, doing, done.
  • A page at / showing the board, with forms to add, move and delete.
  • A JSON API under /api doing the same things, for anything that is not a browser.
  • A board.json file holding the state, written atomically.
  • One error handler turning domain errors into the right status code, for both the HTML and the JSON side.

How the Chapter Is Organised

Each section builds one layer, from the inside out:

  1. Laying Out the Project: the directory structure and why it is shaped this way.
  2. The Storage Layer: the JSON file and the board that sits on top of it.
  3. Validation and the Domain Model: the Task class, which owns every rule about what a task is.
  4. The JSON API: six routes over the board.
  5. Server-Rendered Pages with Wire: the templates, the forms and the redirect-after-post pattern.
  6. Middleware, Logging and Errors: the two pieces of cross-cutting behaviour every request goes through.
  7. Running It for Real: configuration, signals and what changes when you want more than one core.
  8. Testing the Board: a suite over all three layers, and what the layering bought.

Read it in order. Each section assumes the previous one exists.

Laying Out the Project

taskboard/
  index.zu              entry point
  app.zu                wiring: board + templates + server
  config.zu             every setting, with defaults
  models/
    index.zu
    task.zu             the Task class and its rules
  storage/
    index.zu            the Board
    _json_store.zu      private: reading and writing the file
  routes/
    index.zu
    api.zu              the JSON routes
    pages.zu            the HTML routes
  templates/
    layout.html
    board.html
  static/
    app.css
  tests/
    task.zu             one file per layer
    board.zu
    api.zu

Four decisions are worth explaining, because they are the ones that make the rest of the code short.

Before them, the shape itself. Every arrow in this application points one way:

index.zu  ->  app.zu  ->  routes/  ->  storage/  ->  models/

routes knows about storage; storage knows about models; models knows about nothing but config. Nothing points back the other way, which is what makes each layer readable on its own and testable without the ones above it. Testing the Board is where that second half is cashed in.

config.zu sits outside that chain — everything may read it, and it reads nothing. That is the one module allowed to be depended on from anywhere, and it earns the exemption by containing no behaviour at all.

One Package per Responsibility

models, storage and routes are packages: directories with an index.zu. The index.zu decides what is public:

Filename: models/index.zu

import @.task { * }

That one line re-exports everything task.zu declares, so the rest of the application writes import .models { Task } and never needs to know there is a task.zu at all. Splitting task.zu into two files later changes that one line and nothing else.

The @ is what makes it a re-export rather than a private import. Without it, models/index.zu could use Task itself but nothing outside could reach it through models — the import would be local, and import .models { Task } elsewhere would fail. That distinction is the subject of Exporting, and this is the single most common place to get it wrong.

The index.zu is therefore a deliberate, editable list of what a package offers, rather than an accident of which files happen to exist.

The Underscore Is Load-Bearing

storage/_json_store.zu starts with an underscore, so no module outside storage can import it. The compiler enforces that:

SyntaxError: Cannot import private items from module

The point is not secrecy. It is that swapping the JSON file for a database means rewriting one file, and the compiler guarantees nothing else reached past the Board to touch it. Not “we checked and nothing does” — nothing can, and an attempt does not compile.

That guarantee is what turns a convention into a boundary. A comment saying “internal, do not use” is advice; a leading underscore is enforced.

app.zu and index.zu Are Separate

app.zu builds a fully configured server and returns it. index.zu starts one:

Filename: index.zu

import log
import os

import @.app
import .config

if __root__ == __file__ {
  var server = app.build()

  os.on_signal('INT', @() {
    log.info('shutting down')
    os.exit(0)
  })

  log.info('task board on http://${config.HOST}:${config.PORT}')
  server.listen()
}

Three things are happening in that short file.

if __root__ == __file__ is the whole trick. zuri run taskboard runs the directory, finds index.zu, and the two are equal, so the body runs and the server starts. import .taskboard from another program leaves them different — __root__ is that program’s entry file — so nothing starts, and the importer gets app.build() to use however it likes. One file, both jobs, and Chapter 8 covers the idiom.

The signal handler is installed before serving. listen() does not return, so anything that needs to happen on the way out has to be arranged first. os.on_signal('INT', ...) is what turns Ctrl+C into an orderly exit rather than a killed process.

The log line comes before listen(), so the address is printed the moment the process is ready rather than after it stops. It is one line, and it is the difference between “did it start?” and knowing.

Note what index.zu does not do. It builds nothing, configures nothing and knows nothing about boards, templates or routes. Every one of those decisions is in app.zu, which is why the whole of Middleware, Logging and Errors can walk through one function and cover the entire assembly.

Configuration Has Defaults

Filename: config.zu

import os

var HERE = os.dir_name(__file__)

var PORT = os.get_env('PORT', '8000').to_number()
var HOST = os.get_env('HOST', '127.0.0.1')
var DATA_DIR = os.get_env('DATA_DIR', os.join_paths(HERE, 'data'))

var TEMPLATE_DIR = os.join_paths(HERE, 'templates')
var STATIC_DIR = os.join_paths(HERE, 'static')

var COLUMNS = ['todo', 'doing', 'done']

Four things to notice.

Every environment lookup has a fallback, so the application runs with no configuration at all. git clone, zuri run taskboard, and it works. A program that requires six environment variables before it will start is a program nobody tries.

PORT is converted with to_number(). Environment variables are always strings, and '8000' is not a port a socket will accept. The conversion is here, once, rather than at the place the port is used.

HERE is derived from __file__, not from os.cwd(). The templates and the stylesheet live next to the source, so the person running the program is not required to be standing in the right directory. This is the rule from Chapter 9, and a web application is where ignoring it hurts most: the program starts, serves a page, and fails to find a template only once someone requests it.

COLUMNS is here because it is the one piece of knowledge the model, the storage layer and the templates all share. _clean_column() validates against it, Board.columns() and Board.summary() iterate it, and the template renders one section per entry. Adding a fourth column is a one-line change in this file, and every one of those follows.

That is the test for whether something belongs in config.zu: not “is it a setting”, but “would changing it otherwise mean editing several files consistently?”

The Storage Layer

Two files. One knows about the disk and nothing about tasks. The other knows about tasks and nothing about the disk.

The File

Filename: storage/_json_store.zu

import json
import os

/**
 * Reads every stored task dictionary.
 *
 * A missing file is an empty board, not an error: a fresh install has
 * nothing saved yet and should start cleanly. A file that exists but
 * cannot be parsed IS an error, because silently discarding someone's
 * data is worse than refusing to start.
 *
 * @param string path
 * @returns list[dict]
 * @throws Error if the file exists and does not contain a JSON list.
 */
def read_all(path: string) {
  var handle = file(path)

  if !handle.exists() {
    return []
  }

  var decoded

  catch {
    decoded = json.decode(handle.read())
  } as error {
    raise Error('${path} is not readable as JSON: ' + error.message)
  }

  if !is_list(decoded) {
    raise Error('${path} should contain a JSON list, found ' + typeof(decoded))
  }

  return decoded
}

/**
 * Replaces the stored board with `records`.
 *
 * Writes to a temporary file beside the target and renames it into
 * place, so a write interrupted halfway leaves the previous board
 * intact rather than a truncated file.
 *
 * @param string path
 * @param list records
 */
def write_all(path: string, records: list) {
  var directory = os.dir_name(path)

  if !os.dir_exists(directory) {
    os.create_dir(directory, nil, true)
  }

  var temporary = path + '.tmp'
  var handle = file(temporary, 'w')

  handle.open()
  handle.write(json.encode(records, false))
  handle.close()

  file(temporary).rename(path)
}

Four decisions in fifty lines.

A missing file is not an error. A first run has nothing saved, and the program should start. A file that exists but cannot be parsed is an error, because the alternative is overwriting whatever was in there with an empty board.

The error message names the path. The person reading it is looking at one line of a log.

The write is atomic. Writing into a temporary file and renaming it into place means a process killed mid-write leaves the previous board intact, because a rename within a directory either happens or does not.

The handle is opened explicitly. write() on a closed handle opens, writes and closes again, so two write() calls in w mode would each truncate the file. open() first, then write, then close(). That is the rule from Chapter 9, and this is exactly where it bites.

The Board

storage/index.zu is the layer above. It holds Task objects, answers questions about them, and persists through _json_store — and it is the only thing in the application that knows a file is involved at all.

Rather than read it as one hundred lines, here it is a piece at a time.

The Error It Raises

Filename: storage/index.zu

import os

import ..config
import ..models { Task, from_dict }
import ._json_store

/**
 * Raised when a task id does not name a task on the board.
 */
class NotFoundError < Error {

  @new(message) {
    parent(message)
    self.type = 'NotFoundError'
  }
}

A custom error class, four lines, and it earns them. Every layer above can say instance_of(error, NotFoundError) instead of matching on a message string, and the middleware in Middleware, Logging and Errors turns exactly this class into a 404. Had get() raised a plain Error('not found'), that mapping would be a substring search.

Note the two lines inside @new. parent(message) lets the base constructor set message and capture the stack trace; self.type is what makes the class name show up in logs and in the uncaught-error banner. Both are the pattern from Chapter 7.

The Fields and the Constructor

/**
 * A board backed by a JSON file.
 */
class Board {

  /** Every task, newest last. */
  var tasks = []

  /** The file this board reads from and writes to. */
  var path

  /**
   * @param ?string directory: defaults to `config.DATA_DIR`
   */
  @new(directory) {
    self.path = os.join_paths(directory or config.DATA_DIR, 'board.json')
    self.tasks = _json_store.read_all(self.path).map(@(record) => from_dict(record))
  }

Two fields, declared with var even though @new assigns both. The constructor could have declared them implicitly; writing them out means the class’s shape is visible at the top without reading the constructor, and it is required the moment any other method assigns to them.

The constructor does two things and no more. It works out where the file is, and it loads what is in it.

directory or config.DATA_DIR is the optional-parameter idiom from Chapter 5. It is safe here precisely because a directory is never legitimately '' or 0 — the falsy-default trap does not apply to paths.

The map on the second line is the boundary between two worlds. read_all() returns plain dictionaries, because that is what JSON is; from_dict() turns each into a real Task. Everything above this line deals in objects, everything below it deals in dictionaries, and this is the single place they meet.

Reading the Board

  all() {
    return self.tasks
  }

  in_column(column: string) {
    return self.tasks.filter(@(task) => task.column == column)
  }

all() is a one-liner, and it hands back the real list rather than a copy — deliberate, because every caller in this application only reads it. If that changed, this is the line that would need clone().

in_column() is filter() and nothing else. It is a method rather than a loop at each call site so that “which column is this task in” is decided in one place.

Shaping It for the Page

  /**
   * The board grouped for display: one entry per configured column,
   * in board order, each with its own tasks.
   *
   * @returns list[dict]: `{ name, tasks }`
   */
  columns() {
    return config.COLUMNS.map(@(name) {
      var tasks = self.in_column(name).map(@(task) => task.to_view())

      return { name, tasks }
    })
  }

This is the method the HTML template consumes, and three decisions inside it are worth naming.

It iterates config.COLUMNS, not the tasks. That means a column with no tasks still appears, as an empty column — which is what a board should look like. Grouping by walking the tasks instead would silently drop empty columns and produce them in whatever order the data happened to be in.

It returns to_view() results, not Task objects. A view carries the formatted date and the is_done flag already computed, so the template never has to. Server-Rendered Pages explains why that matters for templates specifically.

{ name, tasks } uses the shorthand. Both keys match the variables holding them, so writing { name: name, tasks: tasks } would be noise.

Finding One

  /**
   * @throws NotFoundError if no task has that id.
   */
  get(id: string) {
    var found = self.tasks.find(@(task) => task.id == id)

    if found == nil {
      raise NotFoundError('no task with id ${id}')
    }

    return found
  }

The important decision is that get() raises rather than returning nil.

That choice propagates. Every caller either has a real task or has an error, so no route handler contains if task == nil. The alternative — returning nil — would put that check in five places and guarantee one of them was eventually forgotten, producing a nil field access somewhere far away from the cause.

The message includes the id, because the person reading the log has the request and nothing else.

Changing It

  add(title, options) {
    var task = Task(title, options)

    self.tasks.append(task)
    self._save()

    return task
  }

  update(id: string, changes: dict) {
    var task = self.get(id).update(changes)

    self._save()

    return task
  }

  remove(id: string) {
    var task = self.get(id)

    self.tasks.remove_at(self.tasks.index_of(task))
    self._save()

    return task
  }

All three follow one shape: change the in-memory list, then save, then return what changed.

update() and remove() both start by calling get(), which means a bad id raises NotFoundError before anything is modified. There is no path that half-applies a change.

Each returns the affected task rather than nothing, which is what lets the JSON API answer with the created or updated record without a second lookup.

remove_at(index_of(task)) rather than remove(task) is deliberate: list remove() compares by value, and two tasks with identical fields would make that ambiguous. Removing by position removes exactly the object get() found.

Counting, and Saving

  summary() {
    var counts = {}

    for name in config.COLUMNS {
      counts.set(name, self.in_column(name).length())
    }

    return counts
  }

  _save() {
    _json_store.write_all(self.path, self.tasks.map(@(task) => task.to_dict()))
  }
}

summary() walks the configured columns for the same reason columns() does: an empty column should report 0, not be missing.

_save() is the other side of the constructor’s map. Objects go out as dictionaries via to_dict(), exactly as they came in through from_dict(). The two are a matched pair, and a field added to Task needs to appear in both or it will not survive a restart.

_save() is private, and that is the class’s most important property. There is no public method that changes the board without persisting, because add, update and remove each call it and nothing else can. A caller cannot forget to save, because a caller is never given the choice.

The Trade It Makes

The whole board is held in memory and rewritten in full after every change.

For a board a team can read on one screen, that is the right trade: the code is simple enough to hold in your head, and a full rewrite of a few kilobytes is immaterial. It is also the first assumption to revisit if this ever has to hold a hundred thousand tasks, at which point the answer is a real database rather than a cleverer file format.

Concurrency

Board holds the whole file in memory. Two processes writing the same board.json would each overwrite the other’s changes, and the atomic rename does not fix that: it guarantees the file is never half-written, not that two writers agree.

This application runs one server process, so the question does not arise. Running It for Real is where it does.

Validation and the Domain Model

Every rule about what a task is lives in one class. Nothing above it re-checks a title, and nothing below it stores a task that broke a rule.

Filename: models/task.zu

import date
import uuid

import ..config

/**
 * Raised when a task cannot be built from the data given.
 */
class TaskError < Error {
  @new(message, field) {
    parent(message)

    self.type = 'TaskError'
    self.field = field
  }
}

TaskError carries a field as well as a message. That one extra value is what lets the API answer {"error": "...", "field": "title"}, which is what lets a form highlight the input that is wrong. A custom error class exists precisely so it can carry more than a string.

The Class

/**
 * A single task on the board.
 */
class Task {

  /** The task's stable identifier, a UUID v7 so ids sort by age. */
  var id

  /** One line describing the work. Required, at most 120 characters. */
  var title

  /** Free-form detail. Optional, defaults to an empty string. */
  var notes = ''

  /** Which column the task sits in. One of `config.COLUMNS`. */
  var column = 'todo'

  /** Unix timestamp, in seconds, of when the task was created. */
  var created_at

  /**
   * @param string title
   * @param ?dict options: `notes`, `column`, `id`, `created_at`
   * @throws TaskError if the title is empty, too long, or the column
   *    is not one this board has.
   */
  @new(title, options) {
    options = options or {}

    self.id = options.get('id', nil) or uuid.v7()
    self.title = _clean_title(title)
    self.notes = options.get('notes', nil) or ''
    self.column = _clean_column(options.get('column', nil) or 'todo')
    self.created_at = options.get('created_at', nil) or time()
  }

Every field is declared with var and documented, even the ones the constructor fills in. Classes are sealed, so the list is the whole truth about what a task holds, and writing it out means the next reader gets that truth without reading the constructor.

The constructor takes a required title and an options dictionary for everything else. That shape is worth copying: a positional argument for the thing that is always there, and named options for the rest, so a call site never reads Task('x', nil, nil, 'todo', nil).

uuid.v7() rather than v4(), because v7 embeds a timestamp, so ids sort by age. When your identifier is going to end up as a key in something ordered, that is free value.

Mutation

  /**
   * Moves the task to another column.
   *
   * @param string column
   * @returns Task: this task, so calls chain.
   * @throws TaskError if the column is not one this board has.
   */
  move_to(column) {
    self.column = _clean_column(column)
    return self
  }

  /**
   * Applies a partial update. Only the keys present in `changes` are
   * touched; anything absent keeps its current value.
   *
   * @param dict changes: any of `title`, `notes`, `column`
   * @returns Task: this task, so calls chain.
   * @throws TaskError on an invalid title or column.
   */
  update(changes) {
    if changes.contains('title') {
      self.title = _clean_title(changes.title)
    }

    if changes.contains('notes') {
      self.notes = changes.notes or ''
    }

    if changes.contains('column') {
      self.column = _clean_column(changes.column)
    }

    return self
  }

update() uses contains() rather than truthiness. That is the whole difference between a partial update that works and one that does not: a PATCH body of {"notes": ""} means “clear the notes”, and if changes.notes would read that as “no change requested”.

Both return self, so board.get(id).move_to('done') reads as one thought.

Conversion

  /**
   * The task as a plain dictionary, which is what the store writes
   * and what `from_dict()` reads back.
   */
  to_dict() {
    return {
      id: self.id,
      title: self.title,
      notes: self.notes,
      column: self.column,
      created_at: self.created_at,
    }
  }

  /**
   * The same dictionary, plus the derived fields a template or an API
   * client wants and should not have to compute.
   */
  to_view() {
    var view = self.to_dict()

    view.set('created_on', date.from_time(self.created_at).format('M j, Y'))
    view.set('is_done', self.column == 'done')

    return view
  }

  /**
   * What `json.encode()` uses, so a Task can be handed straight to a
   * JSON response with no conversion at the call site.
   */
  @to_json() {
    return self.to_dict()
  }

  @to_string() {
    return 'Task(${self.id}, ${self.column}, ${self.title})'
  }
}

Two representations, on purpose. to_dict() is what gets stored, and it holds exactly what from_dict() needs to rebuild the task. to_view() is what gets displayed, and it adds things that are derived rather than stored: a formatted date, a boolean the template can branch on.

Keeping them apart means a change to the display format never changes the file format.

@to_json() means a Task can be passed straight to response.json(). @to_string() means echo shows something useful when you look at a task while debugging.

Rebuilding From Storage

/**
 * Rebuilds a task from the dictionary `to_dict()` produced.
 *
 * @param dict data
 * @returns Task
 * @throws TaskError if the stored data is not a valid task.
 */
def from_dict(data: dict) {
  return Task(data.get('title', nil), {
    id: data.get('id', nil),
    notes: data.get('notes', nil),
    column: data.get('column', nil),
    created_at: data.get('created_at', nil),
  })
}

This is four lines with one important property: it goes through the same constructor as everything else.

It would have been shorter to assign the fields directly. Doing it this way means a hand-edited board.json containing an unknown column, or a task with no title, is rejected when the board loads rather than becoming a task that nothing can render and nothing can fix.

data.get(key, nil) rather than data.key throughout, because a stored record written by an older version of the program may be missing a key entirely. get() with a fallback turns that into nil, which the constructor’s own defaults then handle; data.column would raise undefined key.

It is also the exact inverse of to_dict(). The two are a matched pair, and a field added to Task has to appear in both or it will not survive a restart — the kind of bug that only shows up the second time you run the program.

Where the Rules Live

def _clean_title(title) {
  if !is_string(title) {
    raise TaskError('a task needs a title', 'title')
  }

  var cleaned = title.trim()

  if cleaned.is_empty() {
    raise TaskError('a task needs a title', 'title')
  }

  if cleaned.length() > 120 {
    raise TaskError('a title must be 120 characters or fewer', 'title')
  }

  return cleaned
}

Read the order of those three checks, because it is the whole function.

The type check comes first. A PATCH body of {"title": 42} reaches this function as a number, and 42.trim() would be a TypeError about methods rather than a TaskError about titles. Checking first means the caller gets an error in the application’s own vocabulary.

The trim happens before the emptiness check, so ' ' is rejected. A title of three spaces is not a title, and checking title.is_empty() on the raw input would have accepted it.

The length check happens after the trim, so trailing whitespace does not count against the limit.

And the function returns the cleaned value. It is not a validator that answers yes or no; it is a normaliser that either produces a good value or raises. That is why the constructor can write self.title = _clean_title(title) with nothing around it.

def _clean_column(column) {
  if !config.COLUMNS.contains(column) {
    raise TaskError(
      'unknown column "${column}", expected one of ' + ', '.join(config.COLUMNS),
      'column'
    )
  }

  return column
}

The message names both what was wrong and what would have been right. unknown column "backlog", expected one of todo, doing, done tells the caller how to fix it; invalid column does not.

Note that the valid set comes from config.COLUMNS rather than being written out here. Adding a column to the board is a one-line change in one file, and this check, columns(), and summary() all follow it.

Two Call Sites Each

Both helpers are module-private — the leading underscore means no other file can reach them — and each is called from exactly two places: the constructor and update().

That is the property the whole chapter is built on. There is no path into a Task that skips them. Not from_dict(), which goes through the constructor. Not move_to(), which calls _clean_column() itself. Not a route handler, which cannot reach the private function at all.

Every layer above can therefore assume a Task it is holding is valid, which is why no route handler in The JSON API re-checks a title.

Why Not the validate Module?

Zuri has a schema validator, and for a form with fifteen fields it is exactly right:

import validate

var schema = validate.schema({
  title: validate.required().string().max_length(120),
  column: validate.required().string(),
})

Here the rules are three lines of Zuri that also normalise (the trim()), produce a domain error with a field on it, and live next to the data they constrain. A schema would be a second place to look.

Use validate when the shape of the input is the problem. Use methods on the class when the rules are part of what the thing is.

The JSON API

Six routes. Every one of them reads or writes through the board and returns JSON, and not one of them handles an error.

The Routes, One at a Time

The Shape of the File

Filename: routes/api.zu

import ..models { TaskError }
import ..storage { NotFoundError }

/**
 * Registers every `/api` route on `server`, reading and writing
 * through `board`.
 *
 * @param HttpServer server
 * @param Board board
 */
def register(server, board) {

register(server, board) takes the server and the board rather than importing them. That is what lets the same routes run against a different board — a temporary one in a test, a seeded one in a demo — and it keeps this file free of any decision about where the data lives.

Listing, With an Optional Filter

  server.get('/api/tasks', @(request, response) {
    var column = request.query_param('column', nil)
    var tasks = column == nil ? board.all() : board.in_column(column)

    response.json({
      tasks: tasks.map(@(task) => task.to_view()),
      summary: board.summary(),
    })
  })

One route serving two questions: every task, or every task in one column.

query_param('column', nil) supplies the fallback, and the test is column == nil rather than if column. That distinction matters: a query string of ?column= produces an empty string, which is falsy but present — and an empty column name should be rejected by in_column(), not silently treated as “no filter”.

The response carries summary alongside tasks because a client rendering a board wants both, and making it ask twice would be two requests for one screen.

Fetching One

  server.get('/api/tasks/:id', @(request, response) {
    response.json(board.get(request.param('id')).to_view())
  })

One line, and it can fail. board.get() raises NotFoundError for an unknown id, and this handler does nothing about it — which is the subject of the section below.

Creating

  server.post('/api/tasks', @(request, response) {
    var body = request.json_body() or {}

    var task = board.add(body.get('title', nil), {
      notes: body.get('notes', nil),
      column: body.get('column', nil),
    })

    response.json(task.to_view(), 201)
  })

request.json_body() returns nil for a request with no body, or a body that is not JSON at all. or {} turns that into an empty dictionary, so body.get('title', nil) works either way and a missing title becomes the model’s problem — which is where the message about it already lives.

The 201 is the second argument to response.json(). A created resource is not a 200, and saying so is one character of effort.

task.to_view() rather than task, for the reason below.

Updating and Deleting

  server.patch('/api/tasks/:id', @(request, response) {
    var body = request.json_body() or {}

    response.json(board.update(request.param('id'), body).to_view())
  })

  server.delete('/api/tasks/:id', @(request, response) {
    board.remove(request.param('id'))
    response.json({ deleted: true })
  })

  server.get('/api/summary', @(request, response) {
    response.json(board.summary())
  })
}

The PATCH handler passes the decoded body straight to board.update() with no filtering. That is safe because update() only looks at the three keys it knows — title, notes, column — using contains(), so a body containing {"id": "hacked"} changes nothing. The whitelist lives in the model, once, rather than in every route that accepts a body.

DELETE returns a body rather than a bare 204, because a client that parses every response as JSON should not have to special-case one route.

Handlers Do Not Handle Errors

This is the chapter’s real point, and it is easiest to see by counting: six handlers, zero catch blocks.

board.get() raises NotFoundError for an unknown id. board.add() and board.update() raise TaskError for a bad title or an unknown column. Not one handler catches either, and every one of them is one to five lines as a direct result.

The middleware in the next section catches both and turns them into a 404 and a 422. There is exactly one place in this application that knows which domain error means which status code, and it is not in a route.

Consider the alternative for a moment. Six handlers each wrapping their board call in a catch, each deciding on a status, each formatting an error body — thirty-odd lines of duplication, and a seventh route added next month that gets one of them subtly wrong.

Note also the two imports at the top of the file. TaskError and NotFoundError are imported and never mentioned again in this file; they are there because the module that catches them needs them re-exported through routes. That is the module system doing its job, and Chapter 8 covers why the @ matters there.

to_view() at the Boundary

Every response sends to_view(), never the Task itself.

That single habit means the API’s shape is a decision Task makes in one method, rather than something that leaks out of however the object happens to be laid out. Add a private field to Task tomorrow and the API returns exactly what it returned yesterday.

It is also why the responses carry created_on and is_done, which are not stored anywhere — to_view() computes them, so every client gets a formatted date and a boolean without doing the work itself.

The Routes in Practice

$ curl -s localhost:8000/api/tasks -H 'content-type: application/json' \
    -d '{"title":"write the capstone","notes":"chapter 17"}'
{"id":"01a08e15-e758-741d-b368-d5a988cb90cd","title":"write the capstone",
 "notes":"chapter 17","column":"todo","created_at":1789090195.9,
 "created_on":"Sep 11, 2026","is_done":false}

$ curl -s localhost:8000/api/tasks
{"tasks":[...],"summary":{"todo":1,"doing":0,"done":0}}

$ curl -s -X PATCH localhost:8000/api/tasks/01a08e15-... \
    -H 'content-type: application/json' -d '{"column":"doing"}'
{"id":"01a08e15-...","column":"doing",...}

$ curl -s localhost:8000/api/tasks?column=doing
{"tasks":[...],"summary":{"todo":0,"doing":1,"done":0}}

$ curl -s localhost:8000/api/tasks -H 'content-type: application/json' -d '{"title":""}'
{"error":"a task needs a title","field":"title"}

$ curl -s localhost:8000/api/tasks/does-not-exist
{"error":"no task with id does-not-exist"}

The last two are the interesting ones. A 422 with a field, and a 404 with a message, from handlers that said nothing at all about status codes.

Server-Rendered Pages with Wire

The browser side is two templates and four routes. Wire’s full reference is its own chapter; this section uses the parts a real page needs.

The Layout

Filename: templates/layout.html

<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <meta name="viewport" content="width=device-width, initial-scale=1">
    <title x-text="title">Task Board</title>
    <link rel="stylesheet" href="/static/app.css">
  </head>
  <body>
    <header class="masthead">
      <h1>Task Board</h1>
      <p class="tagline" x-text="tagline"></p>
    </header>

    <main>
      <template x-slot="content">
        <p>Nothing here yet.</p>
      </template>
    </main>

    <footer>
      <p>Served by Zuri.</p>
    </footer>
  </body>
</html>

This is a valid HTML5 document. Open it in a browser and you get a page with a heading and a paragraph. That is the whole point of Wire’s design: the directives are attributes, so the template is still the thing it renders.

x-slot="content" declares a region an extending template may replace. The content inside it is the default, used when nothing replaces it.

The Page

Filename: templates/board.html

<extend base="layout.html">
  <define name="content">
    <form class="new-task" method="post" action="/tasks">
      <input name="title" placeholder="What needs doing?" maxlength="120" required>
      <select name="column">
        <option x-for="columns" x-value="column"
                x-attr="{ value: column.name }" x-text="column.name"></option>
      </select>
      <button type="submit">Add</button>
    </form>

    <p class="error" x-if="error" x-text="error"></p>

    <div class="board">
      <section class="column" x-for="columns" x-value="column">
        <h2>
          <span x-text="column.name"></span>
          <span class="count">{{ column.tasks|length }}</span>
        </h2>

        <p class="empty" x-not="column.tasks">Nothing here.</p>

        <article class="task" x-for="column.tasks" x-value="task">
          <h3 x-text="task.title"></h3>
          <p class="notes" x-if="task.notes" x-text="task.notes"></p>
          <p class="meta">Added <span x-text="task.created_on"></span></p>

          <form method="post" x-attr="{ action: '/tasks/' + task.id + '/move' }">
            <select name="column">
              <option x-for="columns" x-value="target"
                      x-attr="{ value: target.name }" x-text="target.name"></option>
            </select>
            <button type="submit">Move</button>
          </form>

          <form method="post" x-attr="{ action: '/tasks/' + task.id + '/delete' }">
            <button type="submit" class="danger">Delete</button>
          </form>
        </article>
      </section>
    </div>
  </define>
</extend>

Six directives carry the whole page.

x-for with x-value repeats an element once per entry, binding each one to a name. The nested x-for="column.tasks" inside x-for="columns" is an ordinary nested loop, and the inner one can still see columns from the outer scope, which is how the move dropdown lists every column.

x-text replaces an element’s children with escaped text. Nothing a task’s title contains can become markup.

x-attr takes a dictionary and spreads it onto the element, which is how a form’s action gets built from a task id.

x-if / x-not render an element conditionally. x-not="column.tasks" shows the “Nothing here.” line when the list is empty, because in Wire an empty list is falsy. That is not Zuri’s rule, where [] is truthy; Wire uses its own, friendlier one for template authoring.

{{ column.tasks|length }} is an interpolation through a filter. length is one of Wire’s built-in filters, and it works on strings, lists, dictionaries and bytes.

The Routes

routes/pages.zu registers four routes: one that renders the board, and three that handle the forms on it. Here it is a piece at a time.

The Shape of the File

Filename: routes/pages.zu

import ..config

/**
 * Registers the page and form routes on `server`.
 *
 * @param HttpServer server
 * @param Board board
 * @param Wire view: the template engine to render through
 */
def register(server, board, view) {

The whole file is one register() function that takes its dependencies as arguments. Nothing here reaches for a global board or a global template engine, which is what makes the routes testable: hand register() a board backed by a temporary directory and the same code runs against it.

This is the “pass dependencies in” habit from Debugging, applied at the layer where it costs nothing and buys the most.

Rendering the Board

  server.get('/', @(request, response) {
    response.html(view.render('board.html', {
      title: 'Task Board',
      tagline: _tagline(board),
      columns: board.columns(),
      error: request.query_param('error', nil),
    }))
  })

One route, one call, no logic. Everything it hands the template is either a constant or a method call on the board — there is no loop, no formatting and no branching in this handler, because all of that already happened somewhere better suited to it.

board.columns() does the grouping. _tagline() does the counting. to_view(), back in the domain model, did the date formatting. By the time the template runs, every value it needs is sitting in front of it.

request.query_param('error', nil) is the other half of the redirect pattern below: a failed form redirects with ?error=..., and this is where that message comes back in to be rendered.

Handling a Form

  server.post('/tasks', @(request, response) {
    var form = request.form()

    catch {
      board.add(form.get('title', nil), { column: form.get('column', nil) })
    } as error {
      response.redirect('/?error=' + _escape(error.message))
      return
    }

    response.redirect('/')
  })

All three form handlers follow this exact shape, so it is worth reading once carefully.

request.form() parses a URL-encoded body into a dictionary. form.get('title', nil) rather than form.title, because a browser can post a body with any fields at all — or none — and form.title would raise undefined key on a request that simply omitted it. With get(), a missing field arrives as nil, and _clean_title() turns that into a proper TaskError.

The catch wraps only the board call. response.redirect('/') on the success path sits outside it, so a mistake in the redirect is not reported as a validation failure. That is the “keep the catch block small” rule from Chapter 7.

The return inside the handler is what stops execution falling through to the success redirect. Without it, a failed add would issue two redirects.

The Other Two

  server.post('/tasks/:id/move', @(request, response) {
    catch {
      board.update(request.param('id'), {
        column: request.form().get('column', nil),
      })
    } as error {
      response.redirect('/?error=' + _escape(error.message))
      return
    }

    response.redirect('/')
  })

  server.post('/tasks/:id/delete', @(request, response) {
    catch {
      board.remove(request.param('id'))
    } as error {
      response.redirect('/?error=' + _escape(error.message))
      return
    }

    response.redirect('/')
  })
}

:id in the path is a route parameter, and request.param('id') reads it.

Notice what is not in these handlers. Neither checks that the id names a real task — board.get() raises NotFoundError and the catch picks it up. Neither validates the column — _clean_column() does. Neither checks that the task exists before deleting it — remove() calls get() first.

That is the payoff for putting the rules in the domain model. A route handler is four lines because there is nothing left for it to do.

The Two Helpers

def _tagline(board) {
  var counts = board.summary()

  return '${counts.todo} to do, ${counts.doing} in progress, ${counts.done} done'
}

def _escape(message) {
  return message.replace('/[^a-zA-Z0-9 .,-]/', '').replace(' ', '+', false)
}

_tagline() turns the board’s counts into the line under the heading. It lives here rather than in Board because it is a presentation decision: another front end would word it differently, and the board should not have an opinion.

_escape() is doing something more careful than it looks. The message is about to be put into a URL, so it strips everything that is not a letter, digit, space or basic punctuation, then turns spaces into +.

Two details in that one line. The first replace() uses a regular expression; the second passes false as the third argument to turn pattern handling off, so the single space is matched literally rather than as a pattern. And stripping rather than percent-encoding is the deliberate choice: this is a message we generated, not user input echoed back, so a conservative allow-list is simpler than encoding and cannot produce a malformed URL.

Redirect After Post

Every form handler ends in a redirect rather than rendering a page. That is the post/redirect/get pattern, and it exists because a browser that rendered a page in response to a POST will re-submit that POST when the user presses refresh. Adding a task twice because someone hit F5 is not a bug you want to explain.

The failure path redirects too, carrying the message as a query parameter that the next GET renders into the error banner. So both outcomes leave the browser sitting on a plain GET /, which is refreshable, bookmarkable and safe to go back to.

Why These Handlers Catch and the API’s Do Not

The API handlers in The JSON API let errors propagate, because the middleware turns them into status codes. These catch, because a browser submitting a form does not want a 422 page — it wants the board back with a message on it.

Same domain errors, two presentations, and the choice is made at the layer that knows which kind of client it is talking to. Neither the Board nor the Task has to know that a browser is involved.

The Stylesheet

server.serve_files('/static', config.STATIC_DIR) mounts the directory. The static-file handler deals with content types, ETags, conditional requests and range requests on its own, so app.css is a plain file with nothing around it:

$ curl -s -o /dev/null -w '%{http_code} %{content_type}\n' localhost:8000/static/app.css
200 text/css

What It Renders

<!DOCTYPE html><html lang="en"><head>
    <meta charset="utf-8">
    <meta name="viewport" content="width=device-width, initial-scale=1">
    <title>Task Board</title>
    <link rel="stylesheet" href="/static/app.css">
  </head>
  <body>
    <header class="masthead">
      <h1>Task Board</h1>
      <p class="tagline">1 to do, 0 in progress, 0 done</p>
    </header>
    ...
    <div class="board">
      <section class="column">
        <h2>
          <span>todo</span>
          <span class="count">1</span>
        </h2>

        <article class="task">
          <h3>write the capstone</h3>
          <p class="notes">chapter 17</p>
          <p class="meta">Added <span>Sep 11, 2026</span></p>

          <form method="post" action="/tasks/01a08e15-e758-741d-b368-d5a988cb90cd/move">
            ...

Server-rendered HTML, no JavaScript, and every value escaped by the engine rather than by the person who wrote the template.

Middleware, Logging and Errors

Two pieces of behaviour belong to every request and to no handler: logging what happened, and turning a domain error into a status code. Both are middleware.

Filename: app.zu

import http
import log
import wire

import .config
import .models { TaskError }
import .routes { api, pages }
import .storage { Board, NotFoundError }

/**
 * Builds the application.
 *
 * @param ?number port: defaults to `config.PORT`; pass `0` to let the
 *    operating system choose a free one.
 * @param ?string data_dir: defaults to `config.DATA_DIR`
 * @returns HttpServer: bound to nothing yet; the caller decides how to
 *    serve it.
 */
def build(port, data_dir) {
  var board = Board(data_dir)

  var view = wire.wire({
    root: config.TEMPLATE_DIR,
    auto_reload: true,
  })

  var server = http.HttpServer(port == nil ? config.PORT : port, config.HOST)

  server.max_body_size = 64 * 1024

  _install_middleware(server)

  api.register(server, board)
  pages.register(server, board, view)

  server.get('/health', @(request, response) {
    response.json({ ok: true, tasks: board.all().length() })
  })

  server.serve_files('/static', config.STATIC_DIR)

  return server
}

That one function is the whole assembly, and the order of its lines is the order of the application.

It builds the board first, because everything else needs it. Passing data_dir through rather than letting Board find it means a test, or a second instance, can point at a different directory without touching config.

auto_reload: true on the template engine re-reads a template when the file changes, which is what you want while writing one. It is also the first line to reconsider before a deployment, since it costs a filesystem check per render.

port == nil ? config.PORT : port is not port or config.PORT. That distinction is load-bearing here: port 0 is a legitimate value meaning “let the operating system pick a free one”, and 0 is falsy, so the or form would silently turn it into the configured port. This is the negative-and-zero trap from Chapter 3 in a place where it would really bite.

max_body_size is set explicitly. A server that accepts a request body of any size is a server anyone can exhaust with a single request. 64 KiB is generous for a form that carries a title and a column.

Middleware is installed before the routes. Middleware runs outermost first, so anything registered here wraps every route added afterwards — including /health and the static files below.

The routes are registered by handing each module what it needs. api.register(server, board) and pages.register(server, board, view) are the same shape: a function, given its dependencies, that attaches things to the server. The page routes get the template engine; the API routes do not, because they never render one.

/health is defined here rather than in a route module, because it is about the process rather than about tasks. It reports a task count so that a monitoring check proves the board actually loaded, not merely that the socket answers.

serve_files() comes last, mapping /static onto a directory. It is last because more specific routes should be registered before a catch-all prefix.

And build() returns the server without binding a socket. Nothing here listens. That separation is what lets the entry point decide how to serve it — one process, a pool of isolates, or a test harness that never listens at all — and it is what makes the next section possible.

The Middleware

/**
 * Request logging, and one error handler that turns every exception
 * the routes can raise into the right status code. Handlers below it
 * are then free to raise and say nothing about HTTP.
 */
def _install_middleware(server) {
  server.use(@(request, response, next) {
    var started = time()

    next()

    var elapsed = ((time() - started) * 1000).round()
    log.info('${request.method} ${request.path} -> ${response.status} (${elapsed}ms)')
  })

  server.use(@(request, response, next) {
    catch {
      next()
    } as error {
      _render_error(request, response, error)
    }
  })
}

A middleware takes three arguments: the request, the response, and a next function. next() is where everything below it runs — the other middleware, and eventually the route handler itself.

That one fact explains the shape of both functions here. Code written before the next() call happens on the way in; code written after it happens on the way out, once the handler has finished and the response is populated. A middleware is therefore a pair of moments, not a single step, and the call in the middle is the seam between them.

Reading the Logger

  server.use(@(request, response, next) {
    var started = time()

    next()

    var elapsed = ((time() - started) * 1000).round()
    log.info('${request.method} ${request.path} -> ${response.status} (${elapsed}ms)')
  })

The timestamp is taken on the way in, before anything else runs. The log line is written on the way out, which is the only point at which response.status is known — on the way in, nothing has decided it yet.

That is the whole reason timing middleware works: the same function body runs at both ends of the request, with a local variable surviving in between.

Reading the Error Handler

  server.use(@(request, response, next) {
    catch {
      next()
    } as error {
      _render_error(request, response, error)
    }
  })

This one has nothing before next() and nothing after it. All its work is in the handler, and what it catches is everything that raised anywhere below it — a route, the board, the model, the store.

That is the mechanism the whole application leans on. A catch around next() catches the routes, because next() is the routes.

Why the Logger Comes First

Middleware run outermost first, in the order they were added. So the logger wraps the error handler, which wraps the routes:

logger        starts the clock
  errors        catches whatever escapes
    routes        handles the request
  errors        turns an error into a status
logger        writes the line, with that status

Read the order bottom to top on the way out and the reason becomes obvious: a request that failed still gets logged, and it gets logged with the status the error handler chose rather than with whatever the response held when the exception was raised.

Swap the two and a failing request would be logged before the error handler had set a status — or not at all, if the exception escaped the logger first.

The Module Already Has One

http.middleware.logger() is a ready-made request logger, and in a real application it is what you would reach for:

import http.middleware

server.use(middleware.logger())

It writes the Common Log Format extended with the response time, which every log analyser already parses, and it takes a sink to send lines somewhere other than standard output, a format to build the line yourself, and trust_proxy to log the forwarded client address instead of the peer.

The nine lines above are written out by hand here because a middleware you have read the whole of teaches more than one you called. Once you can see what next() does, swap in the real one.

One Place That Knows About Status Codes

def _render_error(request, response, error) {
  var status = 500
  var body = { error: 'internal error' }

  if instance_of(error, NotFoundError) {
    status = 404
    body = { error: error.message }
  } else if instance_of(error, TaskError) {
    status = 422
    body = { error: error.message, field: error.field }
  } else {
    log.error('unhandled: ' + error.message)
  }

  if request.path.starts_with('/api') or request.wants_json() {
    response.json(body, status)
    return
  }

  response.html('<h1>' + status + '</h1><p>' + body.error + '</p>', status)
}

This is the payoff for defining NotFoundError and TaskError as real classes instead of raising Error('not found') everywhere. instance_of() maps a domain error to a status code, once.

Three details matter.

Unknown errors are a 500 with a generic body, and the real message goes to the log. An error message can contain a path, a query, or a fragment of a file. It goes where operators can read it, not where users can.

The response shape follows the client. A request under /api, or one whose Accept header asks for JSON, gets JSON. Everything else gets HTML. request.wants_json() reads the header for you.

A handler that has already committed a response is not overwritten, because the error handler only ever runs when next() raised.

What It Looks Like

2026-09-11T02:29:55+01:00 INFO [taskboard]: task board on http://127.0.0.1:8000
2026-09-11T02:29:55+01:00 INFO [taskboard]: POST /api/tasks -> 201 (15ms)
2026-09-11T02:29:55+01:00 INFO [taskboard]: GET / -> 200 (409ms)
2026-09-11T02:30:02+01:00 INFO [taskboard]: POST /api/tasks -> 422 (3ms)
2026-09-11T02:30:02+01:00 INFO [taskboard]: GET /api/tasks/nope -> 404 (0ms)
2026-09-11T02:30:03+01:00 INFO [taskboard]: GET /static/app.css -> 200 (1ms)

log puts the timestamp, the level and the module name on every line and colours it when the output is a terminal. The module name comes from where the call was made, so a message from routes/api.zu says so without being told.

The 409ms on that first GET / is Wire compiling the two templates. Every render after it walks the cached instruction tree instead, which is the 1-3ms the later lines show.

Middleware Worth Adding

http.middleware has ready-made pieces for the things every application eventually wants: CORS, compression, rate limiting, and authentication. Each is a function you pass to server.use(), in the position you want it in the chain.

Running It for Real

$ zuri run taskboard
2026-09-11T02:29:55+01:00 INFO [taskboard]: task board on http://127.0.0.1:8000

Open http://127.0.0.1:8000 and the board is there. Add a task, move it, delete it. Stop the server with Ctrl+C and start it again, and everything is still there.

The Entry Point

Filename: index.zu

import log
import os

import @.app
import .config

if __root__ == __file__ {
  var server = app.build()

  os.on_signal('INT', @() {
    log.info('shutting down')
    os.exit(0)
  })

  log.info('task board on http://${config.HOST}:${config.PORT}')
  server.listen()
}

import @.app re-exports app, so another program can import .taskboard { app } and build a server of its own. The __root__ == __file__ check means importing this package starts nothing.

os.on_signal('INT', ...) catches Ctrl+C. This application saves after every change, so there is nothing to flush, and the handler is here because the place to put “finish what you were doing” is obvious once the hook exists.

Configuration

Every setting comes from the environment with a default:

$ PORT=9000 HOST=0.0.0.0 DATA_DIR=/var/lib/taskboard zuri run taskboard

Running it with no environment at all works too, which is what makes the first run painless.

Exercising It

$ curl -s localhost:8000/health
{"ok":true,"tasks":0}

$ curl -s -X POST -H 'content-type: application/json' \
    -d '{"title":"write the capstone","notes":"chapter 17"}' \
    localhost:8000/api/tasks
{"id":"01a08e15-...","title":"write the capstone","notes":"chapter 17",
 "column":"todo","created_at":1789090195.9,"created_on":"Sep 11, 2026",
 "is_done":false}

$ curl -s -X POST -d 'title=review+it&column=doing' localhost:8000/tasks -o /dev/null -w '%{http_code} -> %{redirect_url}\n'
302 -> http://localhost:8000/

$ curl -s -X POST -d 'title=&column=todo' localhost:8000/tasks -o /dev/null -w '%{redirect_url}\n'
http://localhost:8000/?error=a+task+needs+a+title

$ curl -s localhost:8000/api/summary
{"todo":0,"doing":1,"done":1}

The HTML form and the JSON API reach the same board through the same validation, and each reports failure in the form its own client understands.

What Is Stored

$ cat taskboard/data/board.json
[
  {
    "id": "01a08e15-e758-741d-b368-d5a988cb90cd",
    "title": "write the capstone",
    "notes": "chapter 17",
    "column": "todo",
    "created_at": 1789090195.9
  }
]

Readable, editable, and greppable. An invalid column typed in by hand is rejected when the board next loads, because from_dict() goes through the same constructor everything else does.

The Single-Process Assumption

server.listen() serves one connection at a time on one isolate. For a team board that is more than enough: the slowest thing in a request is Wire’s first compile, and everything after it is single-digit milliseconds.

Zuri can serve across cores, and http.serve() is how. It takes a setup function, calls it once inside each worker with that worker’s own server, and runs the pool:

import http
import .app

# In app.zu, alongside build():
#
#   def setup(server) {
#     ...register the same routes and middleware here...
#   }

http.serve(app.setup, { port: 8000, workers: 4 })

Doing it here would change one thing that matters: each isolate has its own heap, so each worker would have its own Board, its own copy of every task, and its own idea of what board.json should contain. Two workers saving at once would each write a complete file, and one of them would win.

The board would need to stop being in-memory state. The choices are the usual ones:

  • A database. Move _json_store.zu to a real store with transactions. Nothing above it changes; that is why it is one private file.
  • One owner. Keep the board in a single isolate and have the workers talk to it over a channel. Every mutation becomes a message, and one isolate serialises them.
  • A lock. Guard the file with an advisory lock and re-read before every write. Simplest to add, and it makes every request pay for the file.

Which one is right depends on how many people are using the board, and none of them is worth doing before the answer is “more than this can handle.” The design already leaves the door open, and that is the part to get right early.

What This Application Used

Almost all of it.

ChapterWhat the application uses it for
3, 4every line
5anonymous handlers, closures over board, typed parameters
6Task, Board, custom errors, @to_json, to_string
7TaskError, NotFoundError, catch at the form boundary
8packages, index.zu, @ re-export, _ privacy, __file__
9file(), atomic rename, os.join_paths, os.get_env
12the HTTP server, routing, middleware, static files
13json, uuid, date, log
14Wire: layout, slots, loops, filters, escaping
16the request log, and the error shapes that make it readable
19the suite in Testing the Board

What it did not use is as informative. There are no isolates, because one process is enough. There is no validate, because the rules belong to the model. There is no binary handling, no compression, no reflection. Those are all available, and reaching for them here would have made the application longer without making it better.

Where to Take It

The natural next steps, roughly in order of how much they teach:

Authentication. bcrypt for password hashing, a signed cookie for the session, a middleware that rejects anything unauthenticated. The middleware slot is already there.

Live updates. http.sse pushes an event to the browser when the board changes, so two people looking at it see each other’s edits.

Per-user boards. One board.json per user, a Board cache keyed by user id, and path becoming a parameter rather than a default.

A database. Replace _json_store.zu, change nothing else, and watch the underscore earn its keep.

Search and filtering. filter() over the board, exposed as query parameters, with the same code serving the HTML page and the API.

Every one of those is a change to one layer. That is what the layout was for, and Testing the Board is how you make any of them without holding your breath.

Testing the Board

The application works. This section is about keeping it working, and it is the last thing the layering from Laying Out the Project pays for.

Recall the shape:

index.zu  ->  app.zu  ->  routes/  ->  storage/  ->  models/

That is also the order to test in. models depends on nothing but config, so its tests need nothing. storage depends on models and a file, so its tests need a directory. routes depends on storage and a server, and takes both as arguments, so its tests need neither a socket nor a real board.

Each layer is tested against the layer below it as it really is, and against the layer above it not at all.

Where the Tests Live

taskboard/
  index.zu
  app.zu
  config.zu
  models/
  storage/
  routes/
  tests/
    task.zu
    board.zu
    api.zu
$ zuri test

That is the whole arrangement. The application is zuri run taskboard and its tests are zuri test, and neither needs a file name remembering. The command exits 1 when anything failed, which is all CI needs.

Each test file is an ordinary script: it declares its tests and stops. Each one runs in a process of its own, which matters here for one concrete reason: every one of these files is going to create a Board, and a Board is a file on disk. One process per file means one file’s leftovers can never reach another’s.

The Model

models/task.zu is pure. No file, no clock you care about, no network. Its tests are the fastest and the ones worth writing first, because every rule about what a task is lives there and nothing above re-checks any of them.

Filename: tests/task.zu

import test { * }

import ..models { Task, TaskError, from_dict }

describe('Task', @{

  describe('construction', @{

    it('needs a title', @{
      expect(@{ Task(nil) }).to_raise_instance_of(TaskError)
      expect(@{ Task('') }).to_raise_instance_of(TaskError)
      expect(@{ Task('   ') }).to_raise_instance_of(TaskError)
    })

    it('says which field was wrong', @{
      catch {
        Task(nil)
      } as error {
        expect(error.field).to_be('title')
      }
    })

    it('refuses a title over 120 characters', @{
      expect(@{ Task('x' * 121) }).to_raise_instance_of(TaskError)
      expect(@{ Task('x' * 120) }).to_not_raise()
    })

    it('defaults everything but the title', @{
      var task = Task('Write the tests')

      expect(task.notes).to_be('')
      expect(task.column).to_be('todo')
      expect(task.id).to_be_string()
      expect(task.created_at).to_be_number()
    })

    it('sorts by id, because v7 embeds the time', @{
      var first = Task('one')
      var second = Task('two')

      expect([first.id, second.id]).to_be_sorted()
    })

  })

  describe('update()', @{

    it('touches only the keys it was given', @{
      var task = Task('Write the tests', { notes: 'keep me' })

      task.update({ column: 'doing' })

      expect(task.column).to_be('doing')
      expect(task.notes).to_be('keep me')
      expect(task.title).to_be('Write the tests')
    })

    it('clears a value when the key is present and empty', @{
      var task = Task('Write the tests', { notes: 'remove me' })

      task.update({ notes: '' })

      expect(task.notes).to_be('')
    })

    it('refuses a column the board does not have', @{
      var task = Task('Write the tests')

      expect(@{ task.update({ column: 'sideways' }) }).to_raise_instance_of(TaskError)
      expect(task.column).to_be('todo')
    })

  })

  describe('round-tripping', @{

    it('rebuilds a task from what it stored', @{
      var original = Task('Write the tests', { notes: 'a note', column: 'doing' })

      expect(from_dict(original.to_dict())).to_equal(original)
    })

    it('rejects a stored record that broke a rule', @{
      expect(@{ from_dict({ id: 'x', title: '', column: 'todo' }) })
        .to_raise_instance_of(TaskError)
    })

    it('survives a record written before a field existed', @{
      var task = from_dict({ title: 'Write the tests' })

      expect(task.column).to_be('todo')
      expect(task.notes).to_be('')
    })

  })

})

Four of those are worth pointing at.

expect(@{ task.update(...) }).to_raise_instance_of(TaskError) followed by expect(task.column).to_be('todo'). Two assertions, because there are two claims: that it refused, and that it refused before changing anything. A validator that raises after assigning is a bug you only find by checking the second one.

to_equal, not to_be, for the round trip. from_dict() builds a new Task, so identity is never going to match. The whole question is whether the contents survived.

expect([first.id, second.id]).to_be_sorted() is the cheapest possible way to state what uuid.v7() was chosen for. If someone ever changes that line to v4(), this is the test that says so.

The record with no column and no notes. That is the older-version-of-the-program case that data.get(key, nil) exists to handle. Writing the test is what stops the next person from simplifying get() into .column and finding out on someone’s real board.json.

The Storage Layer

A Board is a file, so its tests need a directory of their own. One per test, created and removed by the hooks:

Filename: tests/board.zu

import test { * }
import os

import ..models { Task }
import ..storage { Board, NotFoundError }

describe('Board', @{

  var directory = nil
  var board = nil

  before_each(@{
    directory = os.create_temp_dir('taskboard-test')
    board = Board(directory)
  })

  after_each(@{
    os.remove_dir(directory, true)
  })

  it('starts empty', @{
    expect(board.all()).to_be_empty()
    expect(board.summary()).to_match_object({ todo: 0, doing: 0, done: 0 })
  })

  it('keeps what it was given', @{
    board.add('Write the tests')

    expect(board.all()).to_have_length(1)
    expect(board.all()[0].title).to_be('Write the tests')
  })

  it('survives a restart', @{
    var created = board.add('Write the tests', { notes: 'a note' })

    expect(Board(directory).get(created.id)).to_equal(created)
  })

  it('raises for an id it does not have', @{
    expect(@{ board.get('no-such-id') }).to_raise_instance_of(NotFoundError)
  })

  it('filters by column', @{
    board.add('one')
    var moved = board.add('two')
    board.update(moved.id, { column: 'doing' })

    expect(board.in_column('todo')).to_have_length(1)
    expect(board.in_column('doing')).to_have_length(1)
    expect(board.in_column('done')).to_be_empty()
  })

  it('counts what it holds', @{
    board.add('one')
    board.add('two')

    expect(board.summary()).to_match_object({ todo: 2 })
  })

  it('removes', @{
    var created = board.add('Write the tests')

    board.remove(created.id)

    expect(board.all()).to_be_empty()
    expect(@{ board.get(created.id) }).to_raise_instance_of(NotFoundError)
  })

})

The one that earns its place is survives a restart. It constructs a second Board over the same directory and asks it for the task the first one created. That is the only test here that exercises _json_store, to_dict() and from_dict() together, and it is the test that fails the day someone adds a field to Task and forgets one half of the pair.

after_each removes the directory whether the test passed or not, so a failing test leaves nothing behind for the next one to trip over. That is the property to rely on: teardown that only runs on success is teardown you cannot trust.

Note that none of these tests read board.json themselves. They ask the Board what it holds. The file format is storage’s business, and a test that parsed it would fail the day the format changed for a reason that has nothing to do with what the test was checking.

The Routes

register(server, board) takes its two collaborators as arguments. That was presented in The JSON API as being about seeding a demo board; here is the other half of what it buys.

A handler needs three things: something to register on, a request, and a response. Stand in for all three:

Filename: tests/api.zu

import test { * }
import os

import ..routes { register_api }
import ..storage { Board }

class FakeServer {
  var routes = {}

  @new() {
    self.routes = {}
  }

  get(path, handler) {
    self.routes.set('GET ${path}', handler)
  }

  post(path, handler) {
    self.routes.set('POST ${path}', handler)
  }

  patch(path, handler) {
    self.routes.set('PATCH ${path}', handler)
  }

  delete(path, handler) {
    self.routes.set('DELETE ${path}', handler)
  }
}

class FakeRequest {
  var params = {}
  var query = {}
  var body = nil

  @new(options) {
    options = options or {}

    self.params = options.get('params', {})
    self.query = options.get('query', {})
    self.body = options.get('body', nil)
  }

  param(name) {
    return self.params.get(name, nil)
  }

  query_param(name, fallback) {
    return self.query.get(name, fallback)
  }

  json_body() {
    return self.body
  }
}

class FakeResponse {
  var body = nil
  var status = 200

  json(body, status) {
    self.body = body
    self.status = status or 200
  }
}

Three small classes and the routes become ordinary functions:

describe('the JSON API', @{

  var directory = nil
  var server = nil
  var board = nil

  before_each(@{
    directory = os.create_temp_dir('taskboard-test')
    board = Board(directory)
    server = FakeServer()

    register_api(server, board)
  })

  after_each(@{
    os.remove_dir(directory, true)
  })

  def call(route, options) {
    var response = FakeResponse()
    server.routes[route](FakeRequest(options), response)

    return response
  }

  it('registers every route', @{
    expect(server.routes).to_have_keys([
      'GET /api/tasks',
      'GET /api/tasks/:id',
      'POST /api/tasks',
      'PATCH /api/tasks/:id',
      'DELETE /api/tasks/:id',
      'GET /api/summary',
    ])
  })

  it('answers 201 when it creates a task', @{
    var response = call('POST /api/tasks', { body: { title: 'Write the tests' } })

    expect(response.status).to_be(201)
    expect(response.body).to_match_object({ title: 'Write the tests', column: 'todo' })
    expect(board.all()).to_have_length(1)
  })

  it('lets the model reject a bad body', @{
    expect(@{ call('POST /api/tasks', { body: {} }) }).to_raise_with_message('title')
  })

  it('treats a missing body as an empty one', @{
    expect(@{ call('POST /api/tasks', {}) }).to_raise_with_message('title')
  })

  it('lists everything, with a summary', @{
    board.add('one')

    var response = call('GET /api/tasks', {})

    expect(response.body).to_have_keys(['tasks', 'summary'])
    expect(response.body.tasks).to_have_length(1)
  })

  it('filters by column when asked', @{
    board.add('one')

    expect(call('GET /api/tasks', { query: { column: 'done' } }).body.tasks).to_be_empty()
  })

  it('sends the view, not the stored task', @{
    var created = board.add('Write the tests')

    var response = call('GET /api/tasks/:id', { params: { id: created.id } })

    expect(response.body).to_have_keys(['created_on', 'is_done'])
  })

  it('ignores a key the model does not own', @{
    var created = board.add('Write the tests')

    call('PATCH /api/tasks/:id', {
      params: { id: created.id },
      body: { id: 'hacked', column: 'doing' },
    })

    expect(board.get(created.id).column).to_be('doing')
    expect(board.get(created.id).id).to_be(created.id)
  })

})

Several things worth saying about that.

The board is real. Only the HTTP machinery is faked. A fake board would have meant writing down what board.add() returns, and the test would then pass forever afterwards regardless of what board.add() actually did. Faking is for the things that are slow, remote or awkward, and a temporary file is none of those.

lets the model reject a bad body. From Handlers Do Not Handle Errors: the handler does not catch anything, so an invalid title comes back out of the handler as a TaskError. That is the behaviour the middleware relies on, and this asserts it directly rather than through the middleware.

ignores a key the model does not own is the whitelist claim from update(), tested where an attacker would aim it. One line of test for a property the code gets for free from contains(), and the day someone “simplifies” that method it fails.

Use a class, not a dictionary, for a fake. A dictionary already has get, add, set and keys of its own, so fake.get('id') calls the dictionary’s method rather than yours. A small class gives you the names you meant.

What to Test Next

Two layers are deliberately not tested above.

The pages. routes/pages.zu renders templates. The same fake server works, and the assertion becomes expect(response.body).to_contain(...) against the rendered HTML, or to_match_snapshot(), which records the whole page once and watches it from then on:

it('renders the board', @{
  expect(call('GET /', {}).body).to_match_snapshot()
})

That records the page in tests/__snapshots__/pages.zu.snap. Reviewing the diff when it changes is the point; a template edit that alters more of the page than you intended shows up there and nowhere else.

The whole thing. Starting the real server on port 0, making real requests with the http client, and shutting it down in after_all gives you one test that covers routing, middleware, templates and storage together. Write a handful of those, not a hundred: they are the slowest tests you own and the ones that break for reasons unrelated to what they were checking.

What This Bought

Run it:

$ zuri test

  zuri test  3 files in tests

   PASS   api.zu  312ms  8 tests
   PASS   board.zu  198ms  7 tests
   PASS   task.zu  164ms  10 tests

  3 files
  25 passed  •  25 total
  time 674ms

Twenty-five tests, under a second, no server and no network. That is a direct consequence of the layering: every arrow points one way, and every layer takes its collaborators as arguments rather than importing them.

The layering was justified in Laying Out the Project on the grounds that it makes each layer readable on its own. This is the other half of the claim, and it is the half you feel every day.

Foreign Functions: C and Rust

The ffi module calls code written in C and in Rust. It loads a shared library, describes the functions in it, and calls them with ordinary Zuri values; it turns Zuri functions into C function pointers so a library can call back; and it reads the declarations a header or a crate already contains, so the description is written once, by the people who wrote the library.

Every value that crosses is converted and checked against its C type. An integer that does not fit is a RangeError at the call, not a truncated argument inside the library. Memory the module allocates knows its size, and a read past its end is refused before it happens. Where a check cannot be made, because C handed back an address with no extent attached, the module says so plainly rather than guessing.

A C library, a Rust cdylib, and a static library of either are all reachable, and the layout of every record, packed, over-aligned or full of bitfields, is the one the platform’s own compiler would give it.

Blocks on this page that list several calls together are reference listings, not programs: they show the shape of each call rather than a sequence to run. Anything presented as a complete program runs as written. The programs that call the C runtime run on Linux as shown; the ones that need a library of your own are shown rather than run.

Following Along

The C runtime is on every machine, so most examples below call it. ffi.LIBC names it for the platform the program is running on, and ffi.LIBM names the maths library:

import ffi

var libc = ffi.open(ffi.LIBC)
var strlen = libc.function('strlen', ffi.size_t, [ffi.string])

echo strlen('Hello, C')
8

Some sections use a small library of your own. Save this as geometry.c:

#include <math.h>
#include <stdlib.h>

typedef struct { double x, y; } point;

double distance(point a, point b) {
  return hypot(a.x - b.x, a.y - b.y);
}

point midpoint(const point *a, const point *b) {
  point m = { (a->x + b->x) / 2, (a->y + b->y) / 2 };
  return m;
}

void scale_all(point *points, size_t n, double k) {
  for (size_t i = 0; i < n; i++) {
    points[i].x *= k;
    points[i].y *= k;
  }
}

and build it as a shared library with the C compiler:

cc -shared -fPIC geometry.c -o libgeometry.so -lm       # Linux
cc -dynamiclib geometry.c -o libgeometry.dylib          # macOS
cl /LD geometry.c                                       # Windows

The Rust sections use a crate built with crate-type = ["cdylib"], shown where it is used.

Introduction

Every language that grows up eventually needs code it did not write: a database engine, a codec, a cryptography library, a scientific kernel, a system call the standard library has no wrapper for. The code exists, it is fast and it is tested, and it speaks the C calling convention, which is the one convention every language on every platform agrees on. Rust libraries speak it too, through extern "C".

Calling across that boundary is a matter of describing, exactly, what the other side expects: how wide each integer is, where each member of a struct sits, which register a value travels in. The machine does not check any of it. A description that is off by one byte corrupts memory silently, and the bug shows up somewhere unrelated, later.

ffi makes that description a value the program can inspect, builds it from the source the library’s authors already wrote, lays out every type the way the platform’s compiler does, and checks every value against it at the moment it crosses:

import ffi

var c = ffi.open(ffi.LIBC).declare('
  int abs(int n);
  typedef struct { int quot; int rem; } div_t;
  div_t div(int numerator, int denominator);
')

echo c.abs(-7)
echo c.div(17, 5)
7
{quot: 3, rem: 2}

The declarations are the ones <stdlib.h> contains. The struct came back as a dictionary. And the conversion rules held on the way in:

c.abs(3000000000)
RangeError: abs() argument 1 'n': 3000000000 is out of range for 'int', which holds -2147483648 to 2147483647

Loading a Library

ffi.open() loads a shared library and returns a Library. It takes a path, a file name, or a bare name the platform’s conventions complete:

var sqlite = ffi.open('sqlite3')                    # libsqlite3.so, libsqlite3.dylib, sqlite3.dll
var local = ffi.open('./build/libgeometry.so')      # a path, as given
var vendored = ffi.open('geometry', { paths: ['vendor/lib'] })

A bare name is tried in the directories given as paths first, then in the platform’s own search order. On Linux, a library’s unversioned libname.so is usually only installed with its development package, so ffi.open() also asks the dynamic loader’s cache for the versioned file the runtime package ships, such as libsqlite3.so.0. On macOS the name is tried as a .dylib and as a framework.

ffi.find() answers the same question without loading anything:

import ffi

echo ffi.find('zuri_has_no_such_library')
nil

With no name at all, ffi.open() returns the running process, whose symbols include the C runtime and everything already loaded globally.

Two options change how a library is loaded. lazy: true resolves its symbols as they are first used instead of all at once. global: true makes its symbols visible to libraries loaded after it and to the process handle, which a plugin that expects its host’s symbols needs. Both are Unix concepts and have no effect on Windows, where every loaded module is searched when the process is asked for a symbol.

A library that cannot be found or loaded raises LoadError, carrying the loader’s own explanation: a missing dependency, a library for another architecture, a file that is not a library.

Symbols

A Library answers whether it exports a name, and where:

import ffi

var libc = ffi.open(ffi.LIBC)

echo libc.has('strlen')
echo libc.has('zuri_has_no_such_symbol')
echo libc.symbol('zuri_has_no_such_symbol')
true
false
nil

symbol() returns a Pointer to the symbol. On glibc, name@VERSION asks for a particular version of a versioned symbol, for the rare program that must pin one.

Closing

close() stops a library being used: every function bound from it raises LoadError from then on. The code itself is unloaded once nothing refers to the library any longer, functions bound from it included, so a closed library is never unloaded out from under a call in progress.

Calling a Function

A function is bound by name, return type and parameter types, and comes back as an ordinary Zuri function:

import ffi

var libm = ffi.open(ffi.LIBM)
var pow = libm.function('pow', ffi.double, [ffi.double, ffi.double])

echo pow(2, 10)
echo pow.name()
echo pow.arity()
1024
pow
2

It can be stored, passed to map(), spawned onto an isolate, and called like any other function. The number of arguments is checked like any function’s:

pow(2)
ArgumentError: 'pow' expects 2 arguments, got 1

ffi.describe() reports what a foreign function is, and ffi.is_foreign() tells a foreign function from a Zuri one:

import ffi

var abs = ffi.open(ffi.LIBC).function('abs', ffi.int, [ffi.int])
var about = ffi.describe(abs)

echo about.name
echo about.type.name()
echo ffi.is_foreign(abs)
echo ffi.is_foreign(@(x) => x)
abs
int (*)(int)
true
false

Binding one function at a time suits a handful. For more, declarations read a header, as the next sections show.

Types

Every C type is a Type. The module exports the built-in ones:

ZuriCSize
ffi.voidvoid
ffi.boolbool1
ffi.char, ffi.schar, ffi.ucharchar, signed char, unsigned char1
ffi.short, ffi.ushortshort, unsigned short2
ffi.int, ffi.uintint, unsigned int4
ffi.long, ffi.ulonglong, unsigned long8, or 4 on Windows
ffi.longlong, ffi.ulonglonglong long, unsigned long long8
ffi.int8 … ffi.uint64int8_t … uint64_t1 to 8
ffi.int128, ffi.uint128__int128, unsigned __int12816
ffi.size_t, ffi.ssize_t, ffi.ptrdiff_tthe same8
ffi.intptr_t, ffi.uintptr_tthe same8
ffi.wchar_t, ffi.char16_t, ffi.char32_tthe same4, 2 on Windows; 2; 4
ffi.float, ffi.doublethe same4, 8
ffi.longdoublelong double16, or 8 on Windows and Apple Arm
ffi.complex_float, ffi.complex_doublefloat _Complex, double _Complex8, 16
ffi.ptrvoid *8
ffi.stringconst char *, as text8
ffi.wstringconst wchar_t *, as text8

and Rust’s names for the same types: ffi.i8 through ffi.u128, ffi.isize, ffi.usize, ffi.f32, ffi.f64, and ffi.rust_char for Rust’s char. A Rust name and its C counterpart are the same type:

import ffi

echo ffi.i32.equals(ffi.int32)
echo ffi.int32.equals(ffi.int)
echo ffi.usize.equals(ffi.size_t)
true
true
true

A type knows its size, alignment and kind on this platform:

import ffi

echo ffi.int.size()
echo ffi.double.align()
echo ffi.uint16.kind()
echo ffi.string.kind()
4
8
int
string

Building types

Pointers, arrays, const and function types are built from other types:

import ffi

var names = ffi.pointer(ffi.char).array(4)
var compare = ffi.function_type(ffi.int, [ffi.ptr, ffi.ptr])

echo names.name()
echo names.size()
echo compare.pointer().name()
echo ffi.char.as_const().pointer().name()
char *[4]
32
int (*)(void *, void *)
const char *

ffi.type() reads a C spelling of a built-in type, which is often the shortest way to write one:

import ffi

echo ffi.type('unsigned long long').equals(ffi.ulonglong)
echo ffi.type('void (*)(int, const char *)').kind()
true
function pointer

Structs, unions and enums are built member by member; Structs, Unions and Arrays and Enums and Constants cover them.

Numbers at the Boundary

A Zuri number is a double, which holds every integer up to 2^53 exactly and no further. C integer types reach 2^64 and, with __int128, 2^128. So integers cross the boundary this way:

  • Going in, a number must be a whole number that fits the type, and a bigint may be given for any integer type. A fraction, a string, or a bool for an integer is a TypeError; a value outside the type’s range is a RangeError. Nothing is ever truncated or wrapped.
  • Coming out, an integer is a number when it lies within 2^53 of zero, and a bigint beyond that, so no value is ever rounded.
import ffi

var strtoull = ffi.open(ffi.LIBC).function('strtoull', ffi.uint64,
  [ffi.string, ffi.ptr, ffi.int])

echo strtoull('42', nil, 10)
echo strtoull('18446744073709551615', nil, 10)
echo typeof(strtoull('18446744073709551615', nil, 10))
42
18446744073709551615n
bigint

A type that can be either is the program’s to handle: is_bigint(value), or comparing against a bigint, which compares correctly with a number too.

Floating-point types take any number. long double is wider than a double on some platforms, 80-bit extended precision on x86-64 Linux and macOS and 128-bit quad precision on Arm Linux; going in is exact, and coming out rounds to the nearest double, as C’s own conversion does.

bool takes and gives a Zuri bool, and a C character type takes a number or a one-character string:

import ffi

var toupper = ffi.open(ffi.LIBC).function('toupper', ffi.int, [ffi.int])

echo toupper('q'.ord())
echo toupper(97).chr()
81
A

A complex number is a list of two, [real, imaginary], and passing one by value is available wherever the C compiler has _Complex, which is everywhere but Windows.

Declaring from C

ffi.declare() reads C declarations and returns a Declarations: a set of types, functions, variables and constants that is not tied to any library. Binding it to one produces a namespace:

import ffi

var api = ffi.declare('
  typedef struct { long quot; long rem; } ldiv_t;
  ldiv_t ldiv(long numerator, long denominator);
  long labs(long n);
')

var c = api.bind(ffi.open(ffi.LIBC))

echo c.labs(-12)
echo c.ldiv(100, 7)
echo c.ldiv_t.size()
12
{quot: 14, rem: 2}
16

Library.declare() does both steps at once, and is what most programs use.

The namespace holds a function for each declared function, a Pointer for each declared variable, each constant, and each type that has a name of its own. It is read the way a module is. A tagged type with no typedef, such as struct stat, is reached through the declarations instead: api.type('struct stat').

A declared function that the library does not export is a SymbolError naming every one that is missing. allow_missing: true binds the rest and leaves those out, for a header describing several versions of a library.

What is read

Everything a header says about an API:

  • typedef, struct, union and enum, including forward declarations, self-referencing records, nested records, anonymous members, bitfields and flexible array members;
  • function prototypes, variadic ones included, and function pointer types however deeply they nest;
  • extern variables;
  • _Static_assert, which is evaluated, so a header that checks its own assumptions checks them here too;
  • __attribute__((packed)), aligned(n), ms_abi and sysv_abi, __declspec(align(n)), _Alignas, and asm labels that give a function a different symbol name; every other attribute is read past;
  • extern "C" { ... }, as headers shared with C++ write it.

A function with a body, as an inline helper in a header has one, is read and not bound, because the library need not export it.

import ffi

var api = ffi.declare('
  struct node;
  typedef struct node node;
  struct node { int value; node *next; };

  typedef struct {
    struct { double x, y; } origin;
    union { int id; char tag[8]; };
  } shape;

  _Static_assert(sizeof(shape) == 24, "shape is 24 bytes");
')

echo api.type('node').size()
echo api.type('shape').offset_of('tag')
echo api.type('shape').has_field('id')
16
16
true

The preprocessor

Real header text is full of preprocessor lines, so enough of the preprocessor runs for it to read as written:

  • #define of a value is expanded wherever the name appears and becomes a constant: numbers in any base with any suffix, floating-point numbers, character constants, strings, adjacent strings joined, and expressions over other constants. A value cast to a pointer type, the way a header spells a sentinel such as ((void *) -1), is a Pointer of that type at that address, and nil when the address is zero. A macro can use a type declared after it, as C allows, since a macro means something only where it is used.
  • #if, #ifdef, #ifndef, #elif, #else and #endif are evaluated, with defined(), against the macros the target platform’s compiler predefines: __linux__, __APPLE__, _WIN32, __x86_64__, __aarch64__, __LP64__ and the rest. __has_include() and the other __has_ queries answer no.
  • #pragma pack applies to the records declared after it, exactly where it is written.
  • #undef removes a macro, and #error stops with its message.
  • Including a C standard header is accepted, because everything it declares is already known; <stdio.h> also declares FILE.
import ffi

var api = ffi.declare('
  #include <stdint.h>
  #define VERSION_MAJOR 2
  #define VERSION_MINOR 7
  #define VERSION ((VERSION_MAJOR << 8) | VERSION_MINOR)
  #define NAME "geo" "metry"

  #if VERSION >= 0x200
  typedef struct { uint32_t flags; double scale; } options;
  #else
  typedef struct { uint32_t flags; } options;
  #endif

  #pragma pack(push, 1)
  typedef struct { uint8_t kind; uint32_t length; } header;
  #pragma pack(pop)
')

echo api.constant('VERSION')
echo api.constant('NAME')
echo api.type('options').size()
echo api.type('header').size()
519
geometry
16
5

Two things are refused, each with the line and column of the problem. Including any other file, because declarations are read as given, never fetched: paste in the ones the program needs. And using a function-like macro, which would need the preprocessor’s full expansion rules; define a function-like macro and it is recorded, and only using one is an error.

Text returns

C cannot say who owns a returned pointer, but its convention is clear enough to follow: a function returning const char * hands back text it keeps, and one returning char * usually hands over memory. So a declared function returning const char * returns a string, nil for a null pointer, and one returning char * returns a Pointer, which the program reads and releases. The same holds for const wchar_t *.

import ffi

var c = ffi.open(ffi.LIBC).declare('
  int setenv(const char *name, const char *value, int overwrite);
  const char *getenv(const char *name);
')

c.setenv('ZURI_FFI_DEMO', 'hello', 1)
echo c.getenv('ZURI_FFI_DEMO')
echo c.getenv('ZURI_FFI_NOT_SET')
hello
nil

getenv is declared char *getenv(const char *) in the real header; writing it as const char * here is the program saying it will not free the result, which is true.

Growing a set, and sharing one

Sources are added to a set one after another, and each sees what came before. One set can include another, whose types it can then use without owning them, which is how two libraries share one set of common types:

import ffi

var common = ffi.declare('typedef struct { double x, y; } point;')

var shapes = ffi.declarations()
  .include(common)
  .declare('typedef struct { point from, to; } segment;')
  .declare('double length(segment s);')

echo shapes.type('segment').size()
echo shapes.functions().length.signature().params[0].name()
32
segment

types(), functions(), variables() and constants() list what a set holds, and type() resolves any C spelling against it: shapes.type('segment *[2]').

Declaring from Rust

A Rust library is reached through the C ABI it exports: extern "C" functions, and the #[repr(C)] types they take. ffi.declare_rust() reads that surface as a crate writes it, so its source can be handed over as it stands. Given this lib.rs, built with crate-type = ["cdylib"]:

#![allow(unused)]
fn main() {
use std::ffi::{CStr, CString, c_char};

#[repr(C)]
pub struct Point {
    pub x: f64,
    pub y: f64,
}

#[repr(C)]
pub enum Shape {
    Circle { radius: f64 },
    Rect { w: f64, h: f64 },
}

#[unsafe(no_mangle)]
pub extern "C" fn area(shape: Shape) -> f64 {
    match shape {
        Shape::Circle { radius } => std::f64::consts::PI * radius * radius,
        Shape::Rect { w, h } => w * h,
    }
}

#[unsafe(no_mangle)]
pub extern "C" fn midpoint(a: &Point, b: &Point) -> Point {
    Point { x: (a.x + b.x) / 2.0, y: (a.y + b.y) / 2.0 }
}

#[unsafe(no_mangle)]
pub unsafe extern "C" fn greet(name: *const c_char) -> *mut c_char {
    let name = unsafe { CStr::from_ptr(name) }.to_string_lossy();
    CString::new(format!("hello, {name}")).unwrap().into_raw()
}

#[unsafe(no_mangle)]
pub unsafe extern "C" fn free_greeting(text: *mut c_char) {
    drop(unsafe { CString::from_raw(text) });
}
}

the Zuri side reads the same file:

import ffi

var source = file('shapes/src/lib.rs').read()
var shapes = ffi.open('shapes', { paths: ['shapes/target/release'] }).declare_rust(source)

echo shapes.area({ variant: 'Rect', w: 2, h: 3 })
echo shapes.midpoint({ x: 0, y: 0 }, { x: 4, y: 2 })

var greeting = shapes.greet('zuri').own(shapes.free_greeting)
echo greeting.read_string()
6
{x: 2, y: 1}
hello, zuri

What is read, and what is skipped:

  • functions in extern "C" blocks (and unsafe extern "C" ones), with #[link_name] for a different symbol;
  • extern "C" fn definitions marked #[no_mangle] or #[export_name], with their bodies skipped; one with neither is refused, because it is exported under a mangled name nothing can look up;
  • #[repr(C)], #[repr(C, packed)], #[repr(C, align(N))] and #[repr(transparent)] structs, tuple structs and unions;
  • enums with #[repr(C)] or an integer #[repr], with and without fields;
  • type aliases, const items and statics;
  • #[cfg(...)], evaluated for the platform the program is running on;
  • everything else, from use lines to impl blocks, macros and private functions, is read past.

A type may be used before it is declared, as Rust allows, and a struct without #[repr(C)] can only be used through a pointer, because Rust’s own layout is unspecified. Rust in Depth covers how each Rust type crosses.

C and Rust declarations go in the same set, and each sees the other’s types, which suits a crate that ships a C header beside its source.

Pointers and Memory

A Pointer is an address with, optionally, the type of what it points at. Memory comes from ffi.alloc(), zeroed and typed:

import ffi

var numbers = ffi.alloc(ffi.int, 4)
numbers.set(0, 10)
numbers.set(3, 40)

echo numbers.get(0)
echo numbers.to_list(4)
echo numbers.add(3).get()
echo numbers.size()
10
[10, 0, 0, 40]
40
16

get() and set() read and write elements of the pointer’s type, as C’s ptr[i] does, and add() steps by elements, as C’s ptr + i does. read() and write() take a type of their own and a byte offset, and work on any pointer:

import ffi

var header = ffi.alloc_bytes(16)
header.write(ffi.uint32, 3405691582)
header.write(ffi.double, 2.5, 8)

echo header.read(ffi.uint32)
echo header.read(ffi.double, 8)
echo header.read(ffi.uint8)
echo header.cast(ffi.uint16).get(1)
3405691582
2.5
190
51966

Memory is little-endian on every platform the module runs on, which the third line shows. cast() gives a pointer a type, or takes it away.

Checked access

Memory the module allocated knows its extent, and every access through a pointer into it is checked against that extent before it happens:

import ffi

var numbers = ffi.alloc(ffi.int, 4)

catch {
  numbers.get(4)
} as e {
  echo e.type
  echo e.message
}
PointerError
4 bytes at offset 16 fall outside the 16-byte block this pointer belongs to

A write that fails to convert writes nothing, so a record is never left half updated. Freed memory refuses every access, and freeing it twice is an error rather than a double free.

A pointer C handed back carries no extent, because C did not say how much memory is behind it. Only a null pointer is caught through one; how much is there is whatever the library’s documentation says.

Pointers as arguments

A pointer parameter accepts more than a Pointer, each for the duration of the call:

PassedBecomes
nila null pointer
a Pointerits address
a stringa NUL-terminated copy, for a pointer to characters
bytesthe bytes’ own storage, with nothing copied
a lista temporary array of the target type, written back after the call
a dictionarya temporary record, written back after the call
a Zuri functiona callback, for a function pointer
a Callbackits code
a foreign functionits address

A list or dictionary is written back only when the pointer is not to a const type, so it works as an out-parameter:

import ffi

var frexp = ffi.open(ffi.LIBM).function('frexp', ffi.double,
  [ffi.double, ffi.pointer(ffi.int)])

var exponent = [0]
echo frexp(48, exponent)
echo exponent[0]
0.75
6

Bytes pass their own storage, so a C function that fills a buffer fills the bytes directly:

import ffi

var memset = ffi.open(ffi.LIBC).function('memset', ffi.ptr,
  [ffi.ptr, ffi.int, ffi.size_t])

var buffer = bytes(6)
memset(buffer, 65, 4)
echo buffer
(41 41 41 41 00 00)

A bytes value passed to a call must not be resized by a callback while the call is running; its storage is only guaranteed to stay where it is for as long as nothing changes its length.

Other allocations

ffi.alloc_bytes(size, align) allocates untyped memory. ffi.malloc() allocates from the C allocator, for memory that will be handed to C code which frees it with free(); it is never freed when the pointer is collected. ffi.at(address, type) makes a pointer from a known address. copy_from(), fill() and compare() are memmove, memset and memcmp over checked memory.

Strings and Text

A string passed where a pointer to characters is expected becomes a NUL-terminated copy for the duration of the call, in the encoding the character type implies: UTF-8 for char, UTF-16 for char16_t, UTF-32 for char32_t, and wchar_t’s own, which is UTF-16 on Windows and UTF-32 elsewhere.

ffi.string and ffi.wstring are pointer types that also convert back: a function declared to return one returns a Zuri string, and nil for a null pointer.

import ffi

var libc = ffi.open(ffi.LIBC)
var strstr = libc.function('strstr', ffi.string, [ffi.string, ffi.string])

echo strstr('needle in a haystack', 'hay')
echo strstr('needle in a haystack', 'pin')
haystack
nil

Text that has to outlive a call, because it is stored in a struct or kept by the library, goes in memory of its own with ffi.alloc_string():

import ffi

var text = ffi.alloc_string('naïve')

echo text.read_string()
echo text.read_bytes(6)
echo ffi.alloc_string('naïve', 'utf-16').read_string(nil, 'utf-16')
naïve
(6e 61 c3 af 76 65)
naïve

read_string() reads up to the terminator, or exactly a given number of code units, in any of the four encodings, and raises ValueError for text that is not valid in its encoding. write_string() writes text and its terminator. When a read has no terminator inside a block of known size, it is refused rather than run past the end.

Structs, Unions and Arrays

A record is built member by member, and laid out the way the platform’s C compiler lays it out:

import ffi

var Header = ffi.struct('Header')
  .add_field('tag', ffi.uint8)
  .add_field('length', ffi.uint32)
  .add_field('checksum', ffi.uint64)

echo Header.size()
echo Header.offset_of('length')
echo Header.offset_of('checksum')
16
4
8

It is sealed the first time anything needs its size, including passing a value of it, and cannot change after that. set_packed() packs it as #pragma pack does, set_align() raises its alignment, and add_field() takes a minimum alignment for one member, as _Alignas gives:

import ffi

var Packed = ffi.struct('Packed').add_field('tag', ffi.uint8).add_field('length', ffi.uint32).set_packed()
var Wide = ffi.struct('Wide').add_field('tag', ffi.uint8).set_align(64)

echo Packed.size()
echo Wide.size()
5
64

Values

A record value is a dictionary of its members, going in and coming out. A member left out of a dictionary going in is zero. Nested records are nested dictionaries, and array members are lists, except that an array of char reads as the string in it and an array of other bytes reads as bytes:

import ffi

var Item = ffi.struct('Item')
  .add_field('name', ffi.char.array(8))
  .add_field('counts', ffi.int.array(3))

var item = ffi.alloc(Item)
item.set(0, { name: 'widget', counts: [1, 2] })

echo item.get()
echo item.get_field('counts')
{name: widget, counts: [1, 2, 0]}
[1, 2, 0]

Through a pointer, get_field() and set_field() read and write one member, as ptr->name does, and field() points at one, as &ptr->name does. fields() lists the members with their offsets.

Bitfields

add_bitfield() adds a member of a given width. Bitfields follow the platform’s rules for packing them, GCC’s and Clang’s on Linux and macOS and MSVC’s on Windows, including a zero-width bitfield ending a storage unit:

import ffi

var Flags = ffi.struct('Flags')
  .add_bitfield('ready', ffi.uint, 1)
  .add_bitfield('mode', ffi.uint, 3)
  .add_bitfield('delta', ffi.int, 4)

var flags = ffi.alloc(Flags)
flags.set_field('mode', 5)
flags.set_field('delta', -3)

echo flags.get()
echo Flags.fields()[2].bit_offset

catch {
  flags.set_field('mode', 8)
} as e {
  echo e.message
}
{ready: 0, mode: 5, delta: -3}
4
8 does not fit in the 3-bit field 'mode', which holds 0 to 7

Unions

A union value going in is a dictionary holding one member, the one to store. Coming out, it holds every member, each read from the same bytes, because only the program knows which one is meaningful:

import ffi

var Bits = ffi.union('Bits').add_field('f', ffi.float).add_field('u', ffi.uint32)
var value = ffi.alloc(Bits)
value.set(0, { f: 1 })

echo value.get()
{f: 1, u: 1065353216}

A union read back and passed in again is accepted as it is, because its members agree; members that disagree are a TypeError.

By value

Records pass to and return from functions by value, whatever their shape. Which registers a record travels in depends on its size and on the types of its members, differently on every platform, and those rules are applied to the record’s real layout, packed and over-aligned records and bitfields included:

var geometry = ffi.open('geometry').declare('
  typedef struct { double x, y; } point;
  double distance(point a, point b);
  point midpoint(const point *a, const point *b);
  void scale_all(point *points, size_t n, double k);
')

echo geometry.distance({ x: 0, y: 0 }, { x: 3, y: 4 })
echo geometry.midpoint({ x: 0, y: 0 }, { x: 4, y: 2 })

var points = [{ x: 1, y: 1 }, { x: 2, y: 3 }]
geometry.scale_all(points, 2, 10)
echo points
5
{x: 2, y: 1}
[{x: 10, y: 10}, {x: 20, y: 30}]

midpoint takes pointers to records, and the dictionaries went in as temporary records. scale_all takes a pointer to an array of them, and a list of dictionaries went in as a temporary array and was written back.

Enums and Constants

An enum is an integer type with named constants. A value of one is a number, and anywhere one goes in, a constant’s name may go instead:

import ffi

var Level = ffi.enum('Level')
  .add_constant('LOW', 1)
  .add_constant('HIGH', 10)

echo Level.constants()
echo Level.value('HIGH')
echo Level.size()

var level = ffi.alloc(Level)
level.set(0, 'HIGH')
echo level.get()
{LOW: 1, HIGH: 10}
10
4
10

Declared enums come with their constants, and a declared enum is stored as int unless its values need more, or it says otherwise (enum x : uint8_t), or GCC’s packed attribute asks for the smallest type. Every enum constant, #define constant and Rust const is a constant of the declarations, and a member of the namespace they bind to.

Callbacks

A Zuri function passed where a function pointer is expected becomes a C function pointer for the length of that call:

import ffi

var c = ffi.open(ffi.LIBC).declare('
  void qsort(void *base, size_t count, size_t size,
             int (*compare)(const void *, const void *));
')

var numbers = ffi.alloc(ffi.int, 5)
for i, n in [42, 7, 19, 3, 25] {
  numbers.set(i, n)
}

c.qsort(numbers, 5, 4, @(a, b) {
  return a.cast(ffi.int).get() - b.cast(ffi.int).get()
})

echo numbers.to_list(5)
[3, 7, 19, 25, 42]

The arguments arrive converted from their C types, pointers as Pointers, records as dictionaries, and the return value is converted to the callback’s return type, a small integer widened to a full register the way a C compiler returns one.

Lasting callbacks

A library that keeps a function pointer and calls it later, a signal handler, an event callback, a logging hook, needs a callback that outlives the call that handed it over. ffi.callback() makes one:

var on_event = ffi.callback(@(code) {
  echo 'event ${code}'
}, ffi.function_type(ffi.void, [ffi.int]))

library.set_handler(on_event)

A lasting callback lives until release(). Nothing else frees it, because nothing can know when the library has finished with the pointer; a callback that is never released lives as long as the program. After release() it can no longer be passed, and C must not call it again.

Errors inside a callback

A Zuri error cannot unwind through C frames; the C code in between expects to finish. So an error raised inside a callback is trapped, C gets a zero back (or the callback’s error_value), and the error is raised again, as it was, the moment the C function returns:

import ffi

var c = ffi.open(ffi.LIBC).declare('
  void qsort(void *base, size_t count, size_t size,
             int (*compare)(const void *, const void *));
')

var numbers = ffi.alloc(ffi.int, 3)

catch {
  c.qsort(numbers, 3, 4, @(a, b) {
    raise ValueError('comparison refused')
  })
} as e {
  echo '${e.type}: ${e.message}'
}
ValueError: comparison refused

The first error is the one raised; later calls into the callback during the same C call see the error value.

Callbacks and Threads

C libraries run threads of their own, and those threads call callbacks. A Zuri isolate runs on one thread, so a call from any other thread is posted to the isolate that made the callback, and the calling thread waits until the isolate has run it. The isolate answers at its next safepoint, the same points where it checks for signals, or at once when it is inside a foreign call or in ffi.serve().

That covers a library whose threads call back while the program does something else. It does not cover a function that blocks the isolate until its own threads have finished calling back: the isolate would be waiting for the function, the function for its threads, and the threads for the isolate. ffi.threaded() gives such a function a variant that runs on a helper thread while the isolate keeps answering:

import ffi

var c = ffi.open(ffi.LIBC).declare('
  typedef unsigned long pthread_t;
  int pthread_create(pthread_t *thread, const void *attributes,
                     void *(*start)(void *), void *argument);
  int pthread_join(pthread_t thread, void **result);
')

var seen = []
var start = ffi.callback(@(argument) {
  seen.append(argument.address())
  return nil
}, ffi.function_type(ffi.ptr, [ffi.ptr]))

var thread = ffi.alloc(ffi.ulong)
c.pthread_create(thread, nil, start, ffi.at(42))

var join = ffi.threaded(c.pthread_join)
join(thread.get(), nil)

echo seen
start.release()
[42]

The callback ran on the isolate’s own thread, in the middle of the threaded call. A program with nothing else to do while it waits for calls from another thread can wait in ffi.serve(timeout), which answers whatever is posted and returns how many it answered.

Variadic Functions

A variadic function, printf and its family, takes its fixed parameters by type and any number after them. Each argument past the fixed ones travels as the type its value suggests:

ValueTravels as
a whole number that fits an intint
a larger whole numberlong long
any other numberdouble
a bigintlong long, or unsigned long long when it needs to be
a boolint
a stringconst char *
a pointer, bytes or nilvoid *

Anything else is given its type with Type.of(), and C’s promotions still apply on top: a float travels as a double, and a type narrower than int as an int.

import ffi

var snprintf = ffi.open(ffi.LIBC).function('snprintf', ffi.int,
  [ffi.pointer(ffi.char), ffi.size_t, ffi.string], { variadic: true })

var buffer = ffi.alloc(ffi.char, 64)

snprintf(buffer, 64, '%s has %d items at %.2f each', 'cart', 3, 4.5)
echo buffer.read_string()

snprintf(buffer, 64, '%.1f, %ld, %c', ffi.double.of(2), ffi.long.of(-7), ffi.char.of('z'))
echo buffer.read_string()
cart has 3 items at 4.50 each
2.0, -7, z

The second call shows why of() exists: %.1f with a bare 2 would pass an int, and printf would read a double that was never there.

Errors and errno

The module’s own errors all descend from FfiError:

ClassRaised when
FfiErrora type cannot be used as asked: a record with no size passed by value, a signature libffi cannot call
LoadErrora library cannot be found or loaded, or a function is called after its library was closed
SymbolErrora library does not export a symbol
DeclarationErrorC or Rust source cannot be read; line and column point at the problem
PointerErrormemory would be accessed out of bounds, through null, after being freed, or freed twice
CallbackErrora released callback is used
LinkErrora static library cannot be linked

A value that does not convert to its C type raises the prelude’s TypeError or RangeError, the same errors any function raises for a bad argument, and the message names the function, the argument and the type.

errno

C reports failure through errno, which anything that runs afterwards may change. So errno is cleared immediately before every foreign call and read immediately after it, and ffi.errno() reports what the most recent call on this isolate left there:

import ffi

var strtol = ffi.open(ffi.LIBC).function('strtol', ffi.long,
  [ffi.string, ffi.ptr, ffi.int])

echo strtol('123', nil, 10)
echo ffi.errno()

strtol('99999999999999999999999', nil, 10)
echo ffi.errno()
123
0
34

34 is ERANGE. On Windows, ffi.last_error() reports GetLastError() the same way; it is always zero elsewhere. ffi.set_errno() sets errno for the rare function that reads it.

Ownership and Lifetimes

Three kinds of memory cross the boundary, and each has one owner.

Memory from ffi.alloc() belongs to the program. It is freed when the last pointer into it is collected, or earlier with free(). Pointers made from it with add(), offset(), cast() and field() share it and keep it alive. The one rule is that a pointer must stay reachable for as long as C holds the address; memory C was given and Zuri forgot is memory freed under C’s feet.

Memory from ffi.malloc() belongs to whoever frees it. It is never freed on collection, because the usual reason to allocate it is to hand it to C code that frees it with free(). free() releases it otherwise.

Memory C allocated belongs to C until the program takes it over with own(), naming the function that releases it, or with no function for memory the C allocator’s free() releases:

import ffi

var c = ffi.open(ffi.LIBC).declare('
  char *strdup(const char *text);
  void free(void *pointer);
')

var copy = c.strdup('owned by Zuri now').own(c.free)
echo copy.read_string()
echo copy.is_owned()

copy.free()
echo copy.is_freed()
owned by Zuri now
true
true

An owned pointer is released when it is collected, or with free(); every pointer derived from it refuses access afterwards. A destructor is any foreign function of one pointer: sqlite3_close, png_destroy, CString::from_raw wrapped in an extern "C" function. Only a pointer to the start of an allocation can free it.

The collector runs a destructor in the middle of its own work, where no Zuri code can run, so a destructor that calls a callback during a collection gets zero back, or the callback’s error_value, every time. free() runs the destructor from the program, where callbacks work as they do anywhere else.

Static Libraries

A static library is object code waiting for a linker; nothing can load it at run time as it stands. ffi.link() hands it to the platform’s linker, which links every object in it into a shared library, and loads that:

var geometry = ffi.link('build/libgeometry.a', { libraries: ['m'] })
var rust = ffi.link('shapes/target/release/libshapes.a')

The linker is the C compiler, cc or whatever $CC names, on Linux and macOS, and MSVC’s link.exe on Windows, found through the Visual Studio installation. The result is cached under a name derived from the archives’ contents and the options, in ffi.default_link_cache() or the cache option’s directory, so the linker runs once for a given input and every later program start loads the cached library.

A Rust staticlib carries the Rust standard library with it, which in turn needs a handful of system libraries. An archive holding Rust code is recognised and linked against them without being asked.

On Windows, a static library’s functions are not marked for export, so the linked DLL exports every symbol the archives define with an unmangled name, or exactly the ones listed in the exports option.

libraries, search_paths and flags pass further libraries, their directories and raw arguments to the linker, and linker replaces it. A failure raises LinkError carrying the linker’s own output.

Isolates

Types, pointers, libraries, foreign functions, callbacks and declarations all cross to another isolate, and none of them is moved: both sides keep a working handle on the same thing. A type is a description, a library a handle the loader shares across threads, and a pointer an address, so memory one isolate writes, another reads:

import ffi
import isolate

def fill(memory, value) {
  memory.set(0, value)
  return memory.get(0)
}

var shared = ffi.alloc(ffi.int, 1)
echo isolate.spawn(fill, shared, 99).join()
echo shared.get(0)
99
99

Whatever the memory holds is shared without any synchronisation, as it is between C threads. A callback belongs to the isolate that made it: passed elsewhere, its calls still run on that isolate.

Rust in Depth

Every Rust type that has a defined C ABI crosses, and the rest are refused with the reason.

RustCrosses as
i8 … u128, isize, usizeintegers, 128-bit ones by value included
f32, f64numbers
boola bool
chara one-character string; checked to be a Unicode scalar value
*const T, *mut Ta Pointer, nil for null
&T, &mut T, NonNull<T>, Box<T>a Pointer; nil is refused going in
Option<&T>, Option<NonNull<T>>, Option<Box<T>>a Pointer or nil
extern "C" fn(...)a function pointer: a Zuri function or a Callback in, a callable out
Option<extern "C" fn(...)>the same, or nil
NonZeroU32 and the other NonZero typesa number; zero is refused
Option<NonZeroU32>a number, or nil for None
[T; N]a list
#[repr(C)] structa dictionary
#[repr(transparent)] structits field
#[repr(C)] or #[repr(u8)] enum without fieldsa number, or a variant name going in
#[repr(C)], #[repr(C, u8)] or #[repr(u8)] enum with fieldsa dictionary with a variant key
MaybeUninit<T>, ManuallyDrop<T>, Cell<T>as T
PhantomData<T>nothing; it takes no space

An enum with fields is laid out by the rules those representations define. A value names its variant, and its fields are the rest of the dictionary; a tuple variant’s fields are '0', '1' and so on, and a variant without fields may be passed as just its name:

shapes.area({ variant: 'Circle', radius: 2 })
shapes.area({ variant: 'Rect', w: 2, h: 3 })
tokens.value({ variant: 'Number', '0': 42 })
tokens.value('Plus')

A slice crosses as the pointer and the length Rust’s own FFI convention pairs it into; ffi.slice(type) builds that #[repr(C)] struct:

var Slice = ffi.slice(ffi.i32)
var values = ffi.alloc(ffi.i32, 3)
lib.sum_slice({ ptr: values, len: 3 })

Refused, each with the reason: &str, &[T], String, Vec, and every other type without a stable layout; trait objects; tuples; generic types and functions; Option of a type without a niche; a function without #[no_mangle] or #[export_name]; and the extern "Rust" ABI.

A Rust function declared extern "C" aborts the process if it panics, which is Rust’s own rule for that ABI, and a panic never reaches Zuri. extern "C-unwind" functions are called the same way; a panic unwinding out of one likewise ends the process, because nothing can unwind safely through a Zuri frame.

Platform Differences

The module runs on 64-bit little-endian platforms, x86-64 and Arm, on Linux, macOS and Windows, and follows each one’s C compiler:

Linux x86-64Linux ArmmacOS x86-64macOS ArmWindows x86-64
long88884
charsignedunsignedsignedsignedsigned
wchar_t4, signed4, unsigned4, signed4, signed2, unsigned
long double80-bit128-bit80-bit64-bit64-bit
_Complex by valueyesyesyesyesno
enum past intwider typewider typewider typewider typeint, truncated
bitfield rulesGCCGCCClangClangMSVC

ffi.platform() reports these for the running platform. A program that passes records between platforms through files or sockets gets the same layouts the C compiler on each one produces, which is what the C code there expects.

Calling conventions follow the platform too: System V on Linux and macOS x86-64, AAPCS64 on Arm, with Apple’s variations on it, and the Microsoft x64 convention on Windows. On x86-64, a function type or Library.function() can ask for abi: 'win64', and on Unix abi: 'sysv64', for a function compiled for the other convention; declarations read ms_abi and sysv_abi attributes and Rust’s extern "win64" and extern "sysv64".

128-bit integers cross by value under every one of them, in both directions. The Microsoft x64 convention passes one by reference and returns it in a vector register, as rustc, Clang and GCC compile it, so a callback returning i128 works there as it does everywhere else.

What Is Checked

A foreign call runs code the module cannot see into, so the guarantees stop where C begins. Inside them:

  • every value is converted to its declared type, with integers checked against their range and every other kind against its type;
  • memory the module allocated is bounds-checked and refuses use after free(); freeing twice is an error;
  • a null pointer is caught before it is read or written through;
  • a failed write changes nothing;
  • a Zuri error inside a callback never unwinds through C, and a Rust panic inside the module never reaches C;
  • a callback called from any thread runs on its own isolate.

Beyond them, what C does is C’s. A description that disagrees with the library, a pointer read past what the library says is there, a callback C calls after it was released, bytes resized by a callback while C writes to them: each of these is undefined behaviour in C, and it stays so here. Reading declarations from the library’s own header or crate, rather than writing them by hand, is the surest way to keep the first of these away.

What the Module Refuses

It does not run the C preprocessor in full. Macros that define values and conditional blocks are evaluated; including other files and expanding function-like macros are not, and are refused where they appear.

It does not compile C. Static libraries are linked by the platform’s own linker, which needs a C toolchain on the machine that links them.

It does not guess ownership. A pointer C returns is not freed until the program says how, with own().

It does not free a callback on its own. A callback lives until release(), because only the program knows when the library is done with it.

It does not unwind through foreign frames. Errors are trapped and raised again once C has returned; a panic crossing the boundary ends the process, as it does in Rust.

Module Reference

The standard library reference documents every class and method. The shape of the module:

ffi.open(name, options)loads a library, or the process
ffi.find(name, paths)where a library would load from
ffi.link(archives, options)links static libraries and loads the result
ffi.declare(source), ffi.declare_rust(source), ffi.declarations()sets of declarations
ffi.struct(name), ffi.union(name), ffi.enum(name, type)records and enums, built by hand
ffi.pointer(type, options), ffi.array(type, length), ffi.function_type(returns, params, options), ffi.slice(type), ffi.type(spelling)other types
ffi.alloc(type, count), ffi.alloc_bytes(size), ffi.malloc(size), ffi.alloc_string(text), ffi.at(address, type)memory
ffi.function(pointer, type), ffi.threaded(function), ffi.describe(function), ffi.is_foreign(value)foreign functions
ffi.callback(function, type, options), ffi.serve(timeout)callbacks
ffi.errno(), ffi.set_errno(n), ffi.last_error()error codes
ffi.platform(), ffi.default_link_cache()facts about the platform
ffi.LIBC, ffi.LIBMthe C runtime and maths libraries

On a Library: function, variable, symbol, has, declare, declare_rust, bind, path, close, is_closed.

On a Pointer: get, set, read, write, get_field, set_field, field, read_string, write_string, read_bytes, write_bytes, to_list, add, offset, cast, copy_from, fill, compare, own, free, is_owned, is_freed, address, is_null, type, size, equals.

On a Type: name, kind, size, align, is_const, target, length, signature, pointer, array, as_const, of, equals; on a StructType or UnionType also add_field, add_bitfield, set_packed, set_align, fields, offset_of, has_field, variants; on an EnumType, add_constant, constants, value.

On a Declarations: declare, declare_rust, include, type, constant, constants, types, functions, variables, bind.

On a Callback: pointer, type, release, is_released.

Packages and Nyssa

A package is a Zuri project other projects can use. Zuri installs, publishes and serves them itself: the commands ship with the runtime, and so does Nyssa, the repository they talk to. There is nothing else to install.

$ zuri install http-extra
Resolving dependencies
Installing http-extra 1.5.0
Installing json-schema 1.2.0
  + http-extra 1.5.0
  + json-schema 1.2.0
Installed 2 packages.
import http_extra

This chapter covers the whole of it:

  1. Projects and Versions: what project.toml says, how packages are named, and how version ranges read.
  2. Installing Packages: adding, updating and removing dependencies, the lockfile, and where packages can come from.
  3. Publishing Packages: accounts, tokens, and putting a version on a registry.
  4. Commands From Packages: packages that add zuri commands, and installing tools for your user.
  5. Bundles and Upgrades: shipping a program to machines without Zuri, and keeping Zuri itself up to date.
  6. Running Nyssa: hosting a repository for a team, a company, or the public.

The Commands

CommandWhat it does
zuri initstarts a project
zuri installadds packages, or installs everything a project declares
zuri uninstallremoves packages and whatever only they needed
zuri updatemoves packages to newer versions
zuri restoreinstalls exactly what the lockfile names
zuri infodescribes the project’s packages, or one on a registry
zuri searchfinds packages on a registry
zuri accountsigns in, and manages tokens
zuri publishpublishes a version
zuri yankstops a version from being chosen
zuri ownermanages who may publish a package
zuri cleanfrees the space downloads and installs take
zuri bundlepackages a program with a runtime
zuri upgradereplaces this Zuri with a newer release
zuri serveruns a Nyssa repository

Every one of them answers --help, and every one that changes a project answers --dry-run with exactly what it would do.

What Holds It Together

  • A project is a directory with a project.toml. Every command works on the project around the directory it runs in, found by looking upwards, the same way imports find it.
  • Packages install into the project. They land in .zuri/libs, which import searches before the standard library. Two projects on one machine never share, or fight over, an installed package.
  • One version of each package. Every requirement in the project is satisfied at once or the install stops and says which requirements clash. Nothing is installed twice at two versions.
  • The lockfile is the record. project.lock pins every package to an exact version and the checksum of what was downloaded, and zuri restore reproduces it anywhere.
  • Nothing half done. An install is staged beside .zuri/libs and swapped in whole, so a failure or an interrupted command leaves the project as it was.

Projects and Versions

Starting a Project

$ zuri init weather
$ cd weather
$ zuri run
Hello, world!

zuri init writes a project that runs and a test that passes, with project.toml describing it. zuri init --help lists what it asks and what it can be told up front.

project.toml

[project]
name = "weather"
version = "0.1.0"
description = "Shows the weather."
authors = ["Ada Lovelace <ada@example.com>"]
license = "MIT"
readme = "README.md"
homepage = "https://example.com/weather"
repository = "https://github.com/example/weather"
keywords = ["weather", "cli"]
zuri = ">=0.1"

[dependencies]
http-extra = "^1.4"

[dev-dependencies]
fixtures = "^0.3"
KeyMeaning
namewhat the package is called, and what it is imported as
versionthe version publishing sends, a semantic version
descriptionone sentence, shown in search results
authors, license, readme, homepage, repository, keywordsshown on the package’s page; license is an SPDX expression
zurithe versions of Zuri the package works with, as a range
include, excludewhich files publishing sends, as globs

[dependencies] is what the project needs to run. [dev-dependencies] is what working on it needs, such as test fixtures, and is never required of a project that depends on this one. Both are managed by the package commands, which keep every comment and blank line the file already has.

Four more sections come up later in the chapter: [registries] in Installing Packages, [install] and [hooks] in Install Scripts, and [bundle] in Bundles and Upgrades.

Names

A package name is lowercase letters, digits, hyphens and underscores, starting with a letter and ending with a letter or digit, at most 64 characters. It is imported with every hyphen made an underscore:

PackageImport
http-extraimport http_extra
json-schemaimport json_schema
orm_liteimport orm_lite

Because of that, two names that differ only in hyphens and underscores are the same package, and a registry refuses the second. A name the standard library already uses, such as json or http, cannot be a package, since the package could never be imported past the standard library module.

Versions

A version is major.minor.patch, as Semantic Versioning lays it out:

  • patch for fixes that change nothing anyone relies on,
  • minor for additions that break nothing,
  • major for anything that can break code written against the version before.

A pre-release, such as 2.0.0-rc.1, sorts below its release, and build metadata after a + is ignored when comparing.

Ranges

A dependency says which versions it accepts:

RangeAccepts
1.4.2 or ^1.4.2>=1.4.2, <2.0.0: anything compatible
^0.4.2>=0.4.2, <0.5.0: below 1.0, a minor change may break
^0.0.4exactly 0.0.4
~1.4.2>=1.4.2, <1.5.0: patches only
=1.4.2exactly 1.4.2
1.4>=1.4.0, <2.0.0, a caret range like any bare version
1.4.*, 1.4.x>=1.4.0, <1.5.0
*any release
>=1.2, <1.8both at once; a comma or a space joins comparators
^1.2 || ^2.1either

A bare version is a caret range, so the common case takes compatible updates without saying so. = pins.

A pre-release is only chosen for a range that names a pre-release of the same major.minor.patch itself: ^2.0.0-rc.1 accepts 2.0.0-rc.3, and ^1.4 accepts no pre-release at all. Nobody who asked for releases is handed an unfinished version.

Installing Packages

Adding a Dependency

$ zuri install http-extra
Resolving dependencies
Installing http-extra 1.5.0
Installing json-schema 1.2.0
  + http-extra 1.5.0
  + json-schema 1.2.0
Installed 2 packages.

The newest version is taken and recorded in project.toml as a caret range, http-extra = "^1.5.0". Name a range to choose otherwise, pass --exact to record =1.5.0, or --dev to add a development dependency:

zuri install http-extra@^1.4
zuri install http-extra --exact
zuri install fixtures --dev

Everything the package needs is installed with it, and everything lands in .zuri/libs under its import name:

weather/
├── project.toml
├── project.lock
└── .zuri/
    ├── installed.toml
    └── libs/
        ├── http_extra/
        └── json_schema/

zuri init ignores .zuri in git, apart from .zuri/cmds, so installed packages are never committed. project.toml and project.lock are, and together they reproduce the directory.

The Lockfile

project.lock records the exact version of every package the project was resolved to, direct or not, where it came from, and the checksum of what was downloaded:

version = 1

[[package]]
name = "http-extra"
version = "1.5.0"
source = "registry+https://pub.zurilang.org"
checksum = "sha256:c5b2e20f517379696623b2bb9af8b9f5459a81f62f1aaf3273d9eebbbca3a8af"
dependencies = ["json-schema"]

[[package]]
name = "json-schema"
version = "1.2.0"
source = "registry+https://pub.zurilang.org"
checksum = "sha256:c92464109e86305502e8acaca519d1f7cb2f515a3ab12f3aa26bd1ec67e66b55"
dependencies = []

zuri restore installs exactly that, and is what a fresh checkout, a teammate or a build server runs:

zuri restore
zuri restore --frozen
zuri restore --production

A package whose download does not match its recorded checksum is refused. --frozen refuses to go on when the lockfile no longer matches project.toml instead of resolving again, which is what a build that must install exactly what was reviewed wants. --production leaves out what only development dependencies need.

zuri install with no package named does the same as zuri restore after a change to project.toml: it resolves only what changed and keeps every other package at its locked version.

When Requirements Clash

Every package gets exactly one version, chosen so that every requirement in the project holds at once. When no choice does, the install stops, changes nothing, and explains the clash step by step:

$ zuri install report-kit
Resolving dependencies
install: no set of versions satisfies every requirement:

(1) Because http-extra 1.5.0 depends on json-schema ^1.2 and http-extra 1.4.2 depends on json-schema ^1.2, http-extra requires json-schema ^1.2.
(2) Because report-kit 1.0.0 depends on json-schema ^2.0 and http-extra requires json-schema ^1.2 (1), report-kit 1.0.0 cannot be used with http-extra.
(3) Because report-kit 1.0.0 cannot be used with http-extra (2) and weather depends on http-extra 1.4, report-kit 1.0.0 cannot be used.
    Because report-kit 1.0.0 cannot be used (3) and weather depends on report-kit *, version solving failed.

Read from the bottom: report-kit needs json-schema 2, http-extra needs json-schema 1, and a project cannot have both. The way out is a newer http-extra that accepts json-schema 2, when there is one, or doing without one of the two.

Updating

$ zuri update --dry-run
Resolving dependencies
  package      current  wanted  latest
  json-schema  1.2.0    1.2.0   2.0.0 (breaking)

current is what is installed, wanted the newest the declared range allows, and latest the newest there is, marked when moving to it can break code written against the current one.

zuri update                        # everything, within its range
zuri update http-extra             # one package; the rest stay locked
zuri update json-schema --latest   # move the range itself
zuri update --check                # exit 1 when anything is out of date

--latest rewrites the range in project.toml to the newest release, keeping an exact range exact, and says so for every move to a new major version.

Removing

$ zuri uninstall http-extra
Resolving dependencies
  - http-extra 1.5.0
Removed 1 package.

Whatever was installed only because the package needed it goes too. A package something else still needs stays, and zuri uninstall says what needs it.

Looking Around

$ zuri info
weather 0.1.0
Shows the weather

  package      declared  locked  installed
  http-extra   1.4       1.5.0   1.5.0
  json-schema  -         1.2.0   1.2.0

$ zuri info --tree
weather
└── http-extra 1.5.0
    └── json-schema 1.2.0

A package whose declared range, locked version and installed version do not agree is marked: missing, drifted, unlocked, extraneous, or, with --verify, modified when its files changed since it was installed.

The registry answers questions too:

$ zuri search json
  json-schema  2.0.0  Validates data against JSON Schema drafts 4 to 2020-12.
  report-kit   1.0.0  Builds reports from JSON data.

2 packages on https://pub.zurilang.org, page 1 of 1

$ zuri info http-extra
http-extra 1.5.0
Helpers for building HTTP services: routing, sessions and rate limits.

  license     MIT
  owners      ada
  downloads   0
  depends on  json-schema ^1.2

versions: 1.5.0, 1.4.2, 1.0.0 (yanked)

Git and Local Packages

A package does not have to be on a registry:

zuri install --tag v1.2.0 --git https://example.com/tools.git
zuri install --branch main --git git@example.com:acme/tools.git
zuri install --path ../shared
[dependencies]
tools = { git = "https://example.com/tools.git", tag = "v1.2.0" }
shared = { path = "../shared" }

The package’s name comes from its own project.toml. A git dependency is locked to the commit its tag, branch or revision pointed at and to the checksum of what that commit packs to, so zuri restore installs the same files long after the branch has moved. zuri update follows a branch to its newest commit. A path dependency is copied afresh each time the project is installed, which suits a package being worked on beside the project that uses it.

A published package can only depend on registry packages, since a git or path dependency means nothing on the machine of whoever installs it.

Other Registries

Packages come from the default registry, https://pub.zurilang.org, unless something says otherwise:

zuri install internal-tools --registry https://packages.example.com

The dependency records the registry it came from, so everyone who installs the project gets it from the same place. A project that uses a registry often names it once:

[registries]
company = "https://packages.example.com"

[dependencies]
internal-tools = { version = "^2", registry = "company" }

Aliases can also live in $ZURI_HOME/config.toml for your own use in every project. default there, or ZURI_REGISTRY in the environment, changes the registry used when nothing names one; a project that names its own default keeps it.

A package never falls back from one registry to another. A package asked for from one registry is only ever installed from that registry, so a package with the same name elsewhere can never stand in for it.

Install Scripts

A package may name scripts to run around its installation:

[hooks]
post-install = "scripts/setup.zu"
pre-uninstall = "scripts/teardown.zu"

A script can do anything the person running zuri install can, so a dependency’s scripts run only when the project allows that package by name:

[install]
allow-hooks = ["native-sqlite"]

The project’s own scripts always run. --allow-hooks allows more for one run, and --no-hooks runs none. A script runs with zuri run, from the package’s directory, with no shell in between, and must finish within ten minutes with status 0; its output goes to a log under $ZURI_HOME/logs, which a failure names. ZURI_PACKAGE_NAME, ZURI_PACKAGE_VERSION, ZURI_PACKAGE_DIR and ZURI_PROJECT_DIR tell it where it is. A failing script undoes the whole install.

Offline, Caches and Space

Downloads are kept in a cache, checked by checksum, and shared by every project on the machine. --offline installs from the cache alone, and fails naming what is missing rather than reaching the network.

zuri clean                  # the cache and the script logs
zuri clean --libs           # this project's installed packages
zuri clean --older-than 30d

Everything zuri clean removes comes back by itself, downloaded again or restored from the lockfile, the next time it is needed.

The Environment

VariableWhat it does
ZURI_HOMEwhere your packages, settings and tokens live; ~/.zuri by default
ZURI_CACHEwhere downloads are kept; $ZURI_HOME/cache by default
ZURI_REGISTRYthe default registry, by alias or address
ZURI_TOKEN, ZURI_TOKEN_<ALIAS>a registry token, before the saved one
HTTPS_PROXY, HTTP_PROXY, NO_PROXYthe proxy every download goes through
NO_COLORplain output

Publishing Packages

An Account

Publishing needs an account on the registry. Installing needs none.

zuri account create

That asks for a username, an email address and a password, creates the account, and signs the command line in. It also prints a recovery key, once: the only way back into the account if the password is lost. An account can be made on the registry’s website just as well, and signed in to afterwards:

$ zuri account login
$ zuri account whoami
ada on https://pub.zurilang.org

Signing in exchanges the password for a token, and the token is what every later command sends. The password is never stored. Tokens are kept one per registry in $ZURI_HOME/credentials.toml, readable by you alone.

Tokens

$ zuri account token ci --scopes publish --days 90
$ zuri account tokens
  id                name                 scopes                          expires
  74fed0927f53cc20  ci                   publish                         2026-10-28
  544af66ae6171536  zuri account create  publish, yank, owners, account  2027-09-28
$ zuri account revoke 74fed0927f53cc20

A token carries scopes that bound what it can do:

ScopeAllows
publishpublishing new versions
yankyanking and restoring versions
ownersadding and removing a package’s owners
accountissuing and revoking tokens

A token is shown once, when it is issued, and the registry keeps only a digest of it. It lasts 365 days unless --days says otherwise, and zuri account logout revokes the saved one and forgets it.

A build server signs in with a token rather than a password, read from standard input so it never lands in the shell history or the process list, or taken from the environment:

echo "$PUBLISH_TOKEN" | zuri account login --token-stdin
ZURI_TOKEN="$PUBLISH_TOKEN" zuri publish

Give it a token with the publish scope alone, and nothing it leaks can yank, change owners or issue more tokens.

What Goes In

A package is the project’s files, packed. Inside a git repository that is what git tracks or would track; outside one, every file below the project. .git and .zuri never go. include and exclude in [project] narrow or widen that with globs, * within one directory and ** across any number:

[project]
exclude = ["docs/drafts/**", "*.log"]

Files that look like they hold secrets, such as .env, private keys and credential files, are left out, and the publish stops and names them. A file that is meant to go is named in include.

See exactly what would be sent first. This project leaves its tests out with exclude = ["tests"]:

$ zuri publish --dry-run
Would publish weather 0.1.0

  53 B   .gitattributes
  217 B  .gitignore
  461 B  README.md
  541 B  app/index.zu
  376 B  index.zu
  206 B  project.toml

6 files, 1.8 KiB unpacked, 1.6 KiB packed
sha256:0a6b340e81682bde7e4364e2209333ba32c1f3852ff10558bf513550e2758b30

The archive is built the same way every time: sorted entries, fixed times and owners, normalised permissions. The same files always make the same bytes and the same checksum, on any machine.

Publishing a Version

zuri publish

Before anything is sent, the project must have a valid name and version, git must have no uncommitted changes (--allow-dirty goes ahead anyway), and the package must hold at most 20,000 files that unpack to at most 256 MiB. A license that does not read as an SPDX expression is a warning. The registry then checks the archive again for itself, and refuses anything unsafe to unpack.

A published version never changes. Publishing the same files again is reported as already published; publishing different files under a version that exists is refused. To fix a version, publish the next one.

The first account to publish a name owns it.

Yanking

zuri yank http-extra@1.4.1
zuri yank http-extra@1.4.1 --undo

A yanked version stays on the registry, and every project whose lockfile already names it keeps installing it, so yanking never breaks a build. It is only left out when versions are chosen anew. Yank a version with a serious bug; do not yank it to hide that it existed.

Owners

zuri owner list http-extra
zuri owner add http-extra grace
zuri owner remove http-extra ada

Every owner may publish, yank and change the owners. A package always keeps at least one.

Another Registry

publish, yank, owner and account work with the default registry unless --registry names another, by address or by alias:

zuri account login --registry company
zuri publish --registry company

A dependency from another registry must say which in project.toml, so that whoever installs the package finds it. A dependency naming no registry comes from the registry the package itself is published on.

Commands From Packages

A package can add commands to zuri, the way the runtime’s own are written: a cmds directory in the package, one command per .zu file or directory, as Appendix I describes.

lint-tools/
├── project.toml
├── index.zu
└── cmds/
    └── lint/
        └── index.zu

Installed into a project, the package’s commands work from anywhere inside it and show up in zuri --help, marked with the package they come from:

$ zuri install lint-tools --dev
$ zuri lint
$ zuri --help
...
PACKAGE COMMANDS:
  lint       Check the project for the mistakes CI rejects. (from lint-tools)

The runtime’s commands come first and a project’s own .zuri/cmds next, so a package can never replace either. Two installed packages providing the same command is refused when the second is installed.

Tools for Your User

--global installs into $ZURI_HOME instead of a project, which is how a tool you use everywhere is installed:

zuri install lint-tools --global
zuri uninstall lint-tools --global
zuri update --global

A globally installed package’s commands work from any directory. Each one also gets a launcher in $ZURI_HOME/bin, so with that directory on your PATH, the command runs on its own:

export PATH="$HOME/.zuri/bin:$PATH"
lint

zuri install --global says when the directory is missing from PATH. Globally installed packages are importable from any script too, after the project’s packages and the standard library.

Bundles and Upgrades

Bundling a Program

zuri bundle packages a project with a runtime into something that runs on a machine where Zuri is not installed:

$ zuri bundle --format exe
$ ./dist/weather-0.1.0-x86_64-unknown-linux-gnu
Hello, world!

A bundle holds the runtime renamed after the program, the standard library it was built with, and the project with the packages it needs in production. Running it starts the project’s index.zu, whatever directory it is started from, and its arguments reach the program in os.args. It uses nothing installed on the machine it runs on, neither a ZURI_ROOT nor anything in a ZURI_HOME, so it behaves the same everywhere.

FormatWhat it makes
archivea .tar.gz of the bundle directory, or a .zip for Windows; the default
dirthe bundle directory itself
exeone file, which unpacks itself into the user’s cache the first time it runs and starts from there afterwards
appa macOS application

Bundles go in dist in the project, named <name>-<version>-<platform>, unless --output and --name say otherwise.

What Goes In

The lockfile has to match project.toml, and every package it names, apart from development dependencies, has to be installed at its locked version, so a bundle never ships packages nobody resolved. zuri restore puts either right. A project with no dependencies needs no lockfile at all. The project’s files are chosen the same way publishing chooses them, so include and exclude apply here too.

Other Platforms

zuri bundle --target aarch64-apple-darwin --target x86_64-pc-windows-msvc
zuri bundle --target all

A bundle for another platform is built with the release of this same Zuri version for that platform, downloaded once, checked against its published checksum, and kept in $ZURI_HOME/runtimes. --runtime names an unpacked runtime to use instead. The platforms are:

  • x86_64-unknown-linux-gnu
  • aarch64-unknown-linux-gnu
  • x86_64-apple-darwin
  • aarch64-apple-darwin
  • x86_64-pc-windows-msvc

A single-file bundle for macOS is built on any machine. Its payload goes inside the executable’s image rather than after it, and the result is signed ad hoc, which is all an Apple silicon Mac needs to run it.

Signing for Distribution

A program downloaded onto a Mac runs only when it is signed with a Developer ID and notarized, and notarization requires the hardened runtime. Under the hardened runtime, Zuri’s JIT needs the com.apple.security.cs.allow-jit entitlement to create executable memory, so sign with an entitlements file that grants it:

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
  <key>com.apple.security.cs.allow-jit</key>
  <true/>
</dict>
</plist>
$ codesign --force --options runtime --entitlements entitlements.plist \
    --sign "Developer ID Application: Example Ltd" dist/weather-0.1.0-aarch64-apple-darwin

--force replaces the ad hoc signature. Sign an app bundle the same way, naming the .app directory in place of the executable.

macOS Applications

[bundle]
name = "Weather"
identifier = "com.example.weather"
icon = "assets/weather.icns"

An app bundle needs identifier, in reverse domain form. name is what Finder shows, and icon an .icns file in the project.

Upgrading Zuri

$ zuri upgrade --check
$ zuri upgrade

--check says whether a newer release exists, and exits 1 when one does. zuri upgrade downloads the release for this platform, checks it against its published checksum, unpacks it beside the installation, and runs it to confirm the version it reports. Only then are the executable, the standard library and the shipped commands swapped in together, and if any part of that fails, all of it is put back.

A release with no published checksum is refused, and so is an installation you cannot write to, before anything is downloaded. --version picks a release, --prerelease considers pre-releases, and moving to an older release needs --allow-downgrade. Releases come from the project’s GitHub releases; ZURI_RELEASES_URL points at another listing in the same shape, such as a mirror, and GITHUB_TOKEN is sent when it is set.

Running Nyssa

Nyssa is the package repository, and every Zuri installation can run one: the registry every package command talks to, and a website for finding packages, reading their documentation, and managing an account and its tokens.

$ zuri serve
Nyssa is serving http://127.0.0.1:3000
listening on 127.0.0.1:3000, storage in /home/ada/.zuri/nyssa

That is a working repository, with nothing to set up first. Point the package commands at it:

zuri account create --registry http://127.0.0.1:3000
zuri publish --registry http://127.0.0.1:3000

Settings

Every setting is a flag, a NYSSA_ environment variable, or a line in nyssa.toml in the storage directory, in that order of precedence:

FlagVariableDefault
--hostNYSSA_HOST127.0.0.1
--portNYSSA_PORT3000
--storageNYSSA_STORAGE$ZURI_HOME/nyssa
--configNYSSA_CONFIGnyssa.toml in the storage directory
--public-urlNYSSA_PUBLIC_URLhttp://<host>:<port>
--workersNYSSA_WORKERSone per CPU core, up to 8
--databaseNYSSA_DATABASESQLite in the storage directory
--signupNYSSA_SIGNUPopen
--max-archive-sizeNYSSA_MAX_ARCHIVE_SIZE10485760 bytes
--trust-proxyNYSSA_TRUST_PROXYoff
--read-onlyNYSSA_READ_ONLYoff
--tls-cert, --tls-keyNYSSA_TLS_CERT, NYSSA_TLS_KEYnone
NYSSA_NAMENyssa, the name the site shows
NYSSA_MAIL_URL, NYSSA_MAIL_USERNAME, NYSSA_MAIL_PASSWORD, NYSSA_MAIL_FROMnone
[server]
host = "0.0.0.0"
port = 8080
public_url = "https://packages.example.com"
workers = 4
trust_proxy = true

[database]
url = "postgres://nyssa@db.internal/nyssa"

[registry]
name = "Example Packages"
signup = "closed"
max_archive_size = 20971520

[mail]
url = "smtp://mail.example.com:587"
username = "nyssa"
from = "Example Packages <packages@example.com>"

public_url is the address people reach the repository at, which links and emails use and which package commands name it by. Passwords belong in the environment, never in a flag, where they would show in the process list: NYSSA_DATABASE for a connection string that carries one, and NYSSA_MAIL_PASSWORD.

The Database

SQLite in the storage directory needs nothing set up, and serves a repository for a team or a company comfortably. PostgreSQL, MySQL and MariaDB serve the same site through the same connection strings the sql module opens:

NYSSA_DATABASE=postgres://nyssa:secret@db.internal/nyssa zuri serve
NYSSA_DATABASE=mysql://nyssa:secret@db.internal:3306/nyssa zuri serve

The schema is kept up to date by numbered migrations, applied when the repository starts, so starting a newer Zuri against an older database brings it up to date. zuri serve migrate does the same and exits, for preparing a database ahead of a deployment. On PostgreSQL and SQLite each migration runs in a transaction; MySQL and MariaDB commit at every schema change, so there each migration runs on its own and is recorded once it has finished.

Published archives are stored under the storage directory by checksum, so the database and that directory are what a backup has to hold.

Serving It Safely

Passwords and tokens cross the network whenever anyone signs in or publishes, so a repository reachable beyond one machine is served over HTTPS, one of two ways:

zuri serve --host 0.0.0.0 --tls-cert fullchain.pem --tls-key privkey.pem
zuri serve --trust-proxy

The second is for a repository behind a proxy that terminates TLS, such as nginx or a load balancer, and makes it believe the client addresses the proxy forwards. Listening beyond this machine with neither prints a warning.

Accounts

With signup = "open", anyone may create an account. A private repository closes sign-up and creates accounts itself:

$ zuri serve admin create grace
Email: grace@example.com
Password:
Password again:
Created grace. Their recovery key, shown only this once:
CommandWhat it does
zuri serve admin create <username>creates an account, whether or not sign-up is open
zuri serve admin promote <username>makes the account an administrator
zuri serve admin demote <username>makes it an ordinary publisher again
zuri serve admin suspend <username>suspends it and revokes every token it holds
zuri serve admin restore <username>lifts a suspension

An administrator may yank any version and change the owners of any package, which is how a malicious release is stopped or an abandoned package handed on. Publishing a new version stays with the package’s owners. Every change anyone makes is recorded in the audit log, with who made it and from where.

With mail configured, a new account confirms its email address before it can publish, and a lost password can be reset by email. Without mail, the recovery key an account is given is the way back in.

Looking After It

zuri serve check             # every stored archive against its checksum
zuri serve backup nyssa.db   # a SQLite database, while it is in use
zuri serve --read-only       # browse and install, nothing published

A PostgreSQL or MySQL database is backed up with its own tools, such as pg_dump or mysqldump. Read-only mode keeps a repository serving during maintenance or a migration elsewhere.

Running It as a Service

On Linux, systemd keeps it running:

[Unit]
Description=Nyssa package repository
After=network.target

[Service]
User=nyssa
Environment=NYSSA_STORAGE=/var/lib/nyssa
Environment=NYSSA_PUBLIC_URL=https://packages.example.com
EnvironmentFile=/etc/nyssa/secrets.env
ExecStart=/usr/local/bin/zuri serve --host 127.0.0.1 --port 3000 --trust-proxy
Restart=on-failure

[Install]
WantedBy=multi-user.target

secrets.env holds NYSSA_DATABASE and NYSSA_MAIL_PASSWORD, readable by the service’s user alone. Stopping the service lets every request already running finish first.

How It Protects Itself

  • Passwords are hashed with Argon2id. Tokens and recovery keys are kept only as digests, and every token carries scopes and an expiry.
  • Five failed sign-ins lock an account for fifteen minutes, counted in the database so every worker sees them, and each worker allows an address 30 attempts a minute to sign in, sign up or recover.
  • Every form carries a token only the site’s own pages have, pages are sent with a content security policy that runs no scripts, and README HTML is sanitised before it is shown.
  • An archive is checked before it is stored: its checksum, its size, every path in it, and that its project.toml names the package and version it was published as. Anything that could not be unpacked safely is refused.
  • A name that differs from a published one only by case, hyphens or underscores is refused, so no package can pass for another.
  • A failure answers the visitor with a plain error page, and the whole of it, with its stack, goes to standard error for whoever runs the repository.

The API

The package commands speak version 1 of a JSON API, which any client can use. Every failure is { "error": { "code", "message" } } with a status that describes it, and the code is stable.

RequestWhat it does
GET /api/v1/configwhat the repository is and accepts
GET /api/v1/index/:nameevery version of a package, with checksums and dependencies
GET /api/v1/packages/:namea package as its page shows it
GET /api/v1/packages/:name/:version/archiveone archive
GET /api/v1/search?q=packages matching a query, with sort, page and per_page
PUT /api/v1/packagespublishes a version, as a multipart upload
POST, DELETE /api/v1/packages/:name/:version/yankyanks a version, and restores it
GET, PUT /api/v1/packages/:name/ownersthe owners, and adding one
DELETE /api/v1/packages/:name/owners/:usernameremoves an owner
POST /api/v1/accountscreates an account
POST, GET /api/v1/tokenssigns in or issues a token, and lists tokens
DELETE /api/v1/tokens/:idrevokes a token, current being the one sent
GET /api/v1/methe account behind the token
GET /healthzanswers 200 while the repository is up

A request that changes anything sends its token as Authorization: Bearer nys_....

JSON-RPC

The rpc module is JSON-RPC 2.0, on both ends and over everything it travels on: as the body of an HTTP request, over a WebSocket, a TCP or TLS socket, a Unix domain socket, a pipe to a child process, or any other stream of bytes. It is written in Zuri on top of json, http, isolate and net, and it holds to the specification exactly.

JSON-RPC is a remote procedure call protocol, and a small one. One program sends the name of a method and its parameters; the other runs it and sends back the result or an error. That is very nearly all of it, and the smallness is the point. The protocol says nothing about what the methods are, how the two programs are connected, or which of them is in charge, so it fits wherever two programs need to call each other. Services call each other with it over HTTP. Blockchain nodes publish their whole API through it, over HTTP for calls and over WebSockets for what they push to subscribers. Mining pools, wallets, trading systems, editors and their tools, and plugin hosts speak it over sockets and pipes.

The module is built in the same spirit. A service answers messages and nothing else, so the same service answers whatever carried the message there; an HTTP route, a WebSocket and a socket server are each a few lines around it.

Following Along

A service says which methods it answers, and answers one message at a time:

import rpc

var calculator = rpc.Service()
  .on_request('add', @(params) => params[0] + params[1])

echo calculator.answer('{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [2, 3]}')
{"jsonrpc":"2.0","id":1,"result":5}

The message went in as JSON text and the answer came out the same way. Nothing touched a network. Everything else in this chapter is about how the message gets to a service and how its answer gets back: as the body of an HTTP request, over a WebSocket or a socket, or between two isolates of one program.

Introduction

JSON-RPC has three kinds of message, each a small JSON object whose jsonrpc member is "2.0".

A request names a method, carries its params, and carries an id. The other side must answer it, and the answer carries the same id.

A notification is a request without an id. Nothing answers it, not even when it fails. It is for things the sender wants done but has no need to hear back about: a log line, a progress report, a change the other side should know of.

A response carries the id of the request it answers, and exactly one of result and error.

Parameters are either a list, matched to the method’s parameters by position, or a dictionary, matched by name. Nothing else is allowed: a call to square(4) sends [4], never 4.

Messages can also travel together as a batch, a JSON array of them. The requests in a batch are answered together, as an array of responses, each matched to its request by id.

The protocol has no transport of its own, and two kinds carry it. In an exchange, such as an HTTP request and its response, one message or batch goes each way and that is the end of it. On a connection, such as a WebSocket or a socket, messages flow both ways for as long as it stays open, and neither side is the client: either may send requests and notifications at any time, including while it is in the middle of answering one. On a stream of bytes, both ends must also agree on where one message stops and the next begins. That agreement is the framing.

The module has a layer for each of these:

LayerWhat it holds
messagesRequest, Notification, Response, RpcError, and encode(), decode() and read()
servicesService, and the Context its handlers are given
HTTPhttp_handler() to serve a service, and HttpClient to call one
framingHeaderFraming, LineFraming and MessageFraming
transportsStdioTransport, SocketTransport, ProcessTransport, WebSocketTransport, ChannelTransport and pipe()
endpointsEndpoint, a service on a connection
serversserve(), an endpoint for every connection to a listening socket

import rpc reaches all of it.

Messages

Requests and Notifications

A message is built from what it says, and rpc.encode() writes it as JSON-RPC sends it:

import rpc

var call = rpc.Request('add', [2, 3], 1)
var note = rpc.Notification('log', { level: 'info', text: 'started' })

echo call
echo rpc.encode(call)
echo rpc.encode(note)
Request(add #1)
{"jsonrpc":"2.0","id":1,"method":"add","params":[2,3]}
{"jsonrpc":"2.0","method":"log","params":{"level":"info","text":"started"}}

An id is a string or a number, and is whatever the sender chooses to match the answer by. An endpoint numbers its own requests from 1.

params of nil leaves the member out of the message, which the specification allows for a method that takes nothing. Anything other than a list, a dictionary or nil is refused when the message is built, before it can be sent:

import rpc

catch {
  rpc.Request('square', 4, 1)
} as error {
  echo error.message
}
params must be a list, a dictionary or nil, not number

Responses

Response.success() answers a request with its result, and Response.failure() with an error:

import rpc

var done = rpc.Response.success(1, 5)
var missing = rpc.RpcError(rpc.METHOD_NOT_FOUND, 'Method not found: mul')
var failed = rpc.Response.failure(2, missing)

echo rpc.encode(done)
echo rpc.encode(failed)
echo failed.is_error()
{"jsonrpc":"2.0","id":1,"result":5}
{"jsonrpc":"2.0","id":2,"error":{"code":-32601,"message":"Method not found: mul"}}
true

A response’s id is nil only when it answers a message whose id could not be read, such as one that was not JSON at all. A result of nil is a result like any other, and is sent as null.

Errors and Their Codes

An error is an RpcError: a code, a short message, and optional data carrying whatever else the other side needs to know.

import rpc

var error = rpc.RpcError(rpc.INVALID_PARAMS, 'a name is required', { field: 'name' })

echo error.code
echo error.message
echo error.to_dict()
-32602
a name is required
{code: -32602, message: a name is required, data: {field: name}}

The specification reserves the codes from -32768 to -32000, and defines these:

ConstantCodeMeaning
PARSE_ERROR-32700the message is not valid JSON
INVALID_REQUEST-32600the JSON is not a valid message
METHOD_NOT_FOUND-32601the method does not exist
INVALID_PARAMS-32602the method cannot take these parameters
INTERNAL_ERROR-32603the method failed while it ran
SERVER_ERROR_MIN to SERVER_ERROR_MAX-32099 to -32000errors an implementation defines for itself

An application’s own errors take any code outside the reserved range. Pick them once, write them down, and keep them stable: the code is what the other side’s program matches on, and the message is for the person reading the log.

RpcError is an Error, so it is raised and caught like any other. Raised from a request handler, it is the answer.

Reading and Writing JSON

rpc.decode() reads the JSON text of a message, and checks it against every rule of the specification:

import rpc

var text = '{"jsonrpc": "2.0", "id": "a7", "method": "subtract", ' +
  '"params": {"minuend": 42, "subtrahend": 23}}'
var call = rpc.decode(text)

echo call
echo call.params.minuend - call.params.subtrahend
Request(subtract #a7)
19

What it reads is a Request, a Notification or a Response, told apart the way the specification tells them apart: a message with a method and an id is a request, one with a method and no id is a notification, and one with a result or an error is a response.

Anything that breaks a rule is refused with an RpcError carrying the code it would be answered with:

import rpc

var attempts = [
  '{"jsonrpc": "2.0", "method": 1}',
  '{"jsonrpc": "1.0", "id": 1, "method": "ping"}',
  '{"jsonrpc": "2.0", "id": 1, "method": "ping", "params": 1}',
  '{"jsonrpc": "2.0", "id": 1, "result": 1, "error": null}',
  '{"jsonrpc": "2.0", "id": 1',
]

for text in attempts {
  catch {
    rpc.decode(text)
  } as error {
    echo '${error.code} ${error.message}'
  }
}
-32600 Invalid Request: method must be a string
-32600 Invalid Request: the jsonrpc member must be '2.0'
-32600 Invalid Request: params must be an array or an object
-32600 Invalid Request: a response needs exactly one of result and error
-32700 Parse error: json.decode(): expected ',' or '}' in object

rpc.read() does the same for a value already decoded from JSON, for a program that got the JSON from somewhere that decodes it already. rpc.encode() is the other direction, for any message or list of them.

Batches

A list of messages is encoded as a batch:

import rpc

echo rpc.encode([
  rpc.Request('add', [1, 2], 1),
  rpc.Notification('log', ['adding']),
])
[{"jsonrpc":"2.0","id":1,"method":"add","params":[1,2]},{"jsonrpc":"2.0","method":"log","params":["adding"]}]

A batch decodes to a list. Each entry in it is read on its own, so one bad entry does not spoil the rest: in its place is the RpcError saying what is wrong with it, ready to be answered.

import rpc

var batch = rpc.decode('[' +
  '{"jsonrpc": "2.0", "id": 1, "method": "sum", "params": [1, 2]},' +
  '{"jsonrpc": "2.0", "method": "notify_hello"},' +
  '{"foo": "boo"}' +
']')

for entry in batch {
  if instance_of(entry, rpc.RpcError) {
    echo 'refused: ${entry.message}'
  } else {
    echo entry
  }
}
Request(sum #1)
Notification(notify_hello)
refused: Invalid Request: the jsonrpc member must be '2.0'

An empty batch, [], is not a message at all, and decoding one raises INVALID_REQUEST.

Services

A Service holds the handler for every method a program answers. answer() takes one message, or a batch, as JSON text or as its bytes, and returns the JSON text of the answer, or nil when there is nothing to answer.

Answering Requests

on_request() gives a method its handler. The handler is called with the request’s params and a Context, and what it returns is the result sent back:

import rpc

var service = rpc.Service()
  .on_request('add', @(params) => params[0] + params[1])
  .on_request('whoami', @(params, context) => '${context.method} #${context.id}')

echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [2, 3]}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "whoami"}')
echo service.answer('{"jsonrpc": "2.0", "id": 3, "method": "mul", "params": [2, 3]}')
{"jsonrpc":"2.0","id":1,"result":5}
{"jsonrpc":"2.0","id":2,"result":"whoami #2"}
{"jsonrpc":"2.0","id":3,"error":{"code":-32601,"message":"Method not found: mul"}}

The Context names the request’s id and method, the endpoint handling it, and the request it arrived in, which for a message from HTTP is the HTTP request. A handler that has no use for the context leaves it out, as add does. A request for a method with no handler is answered with METHOD_NOT_FOUND, as mul is.

A handler is replaced by calling on_request() again with the same method. Method names that start with rpc. are reserved by the specification, and on_request() refuses them.

A message that is not valid UTF-8, or not valid JSON, is answered with PARSE_ERROR and an id of nil, since its id cannot be read.

Failing Properly

A handler fails a request by raising. An RpcError is sent back as it is, code, message, data and all. Any other error is sent back as an INTERNAL_ERROR carrying the error’s message:

import rpc

var service = rpc.Service()
  .on_request('divide', @(params) {
    if params.length() != 2 or params[1] == 0 {
      raise rpc.RpcError(rpc.INVALID_PARAMS, 'divide takes a number and a non-zero divisor')
    }

    return params[0] / params[1]
  })
  .on_request('save', @(params) {
    raise Error('the disk is full')
  })

echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "divide", "params": [1, 0]}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "save", "params": ["notes"]}')
{"jsonrpc":"2.0","id":1,"error":{"code":-32602,"message":"divide takes a number and a non-zero divisor"}}
{"jsonrpc":"2.0","id":2,"error":{"code":-32603,"message":"the disk is full"}}

An error that is not an RpcError also goes to the service’s on_error() handler, since it is a failure in the program rather than a refusal the program meant to make. A handler that must not let its failures’ messages reach the other side catches them and raises an RpcError of its own instead.

On the calling side, an error answer raises from request() as the RpcError the other side sent.

Notifications

on_notification() gives a notification its handler. It is called the same way, with the params and a Context, and whatever it returns is ignored, since nothing is sent back:

import rpc

var seen = []
var service = rpc.Service()
  .on_notification('log', @(params) {
    seen.append(params.text)
  })

echo service.answer('{"jsonrpc": "2.0", "method": "log", "params": {"text": "started"}}')
echo seen
nil
[started]

A notification for a method with no handler is dropped, and a notification handler that raises sends nothing back either: the error goes to on_error(). That is the specification’s rule, and it is deliberate. The sender asked not to be answered.

Methods Without a Handler

on_unhandled() sets a fallback for every request and notification whose method has no handler of its own. It is called like any other handler, and context.method says which method it is standing in for:

import rpc

var service = rpc.Service()
  .on_unhandled(@(params, context) {
    if context.method.starts_with('legacy.') {
      return 'retired: ${context.method}'
    }

    raise rpc.RpcError(rpc.METHOD_NOT_FOUND, 'Method not found: ${context.method}')
  })

echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "legacy.export"}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "export"}')
{"jsonrpc":"2.0","id":1,"result":"retired: legacy.export"}
{"jsonrpc":"2.0","id":2,"error":{"code":-32601,"message":"Method not found: export"}}

It is the place for a family of methods answered the same way, for a proxy that passes calls on, and for logging what a client asks for that nothing answers. Raising METHOD_NOT_FOUND from it refuses a method exactly as a service without a fallback would. Methods whose names start with rpc. never reach it.

Batches and Their Limit

A batch is answered as a batch. Each request in it gets its answer, in place, and each notification gets none; a batch of nothing but notifications has no answer at all:

import rpc

var service = rpc.Service()
  .on_request('add', @(params) => params[0] + params[1])

echo service.answer('[' +
  '{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [1, 2]},' +
  '{"jsonrpc": "2.0", "method": "add", "params": [3, 4]},' +
  '{"jsonrpc": "2.0", "id": 3, "method": "sub"}' +
']')
[{"jsonrpc":"2.0","id":1,"result":3},{"jsonrpc":"2.0","id":3,"error":{"code":-32601,"message":"Method not found: sub"}}]

One message holding a million requests is one message, and a service that took it whole would do a million calls’ work for it. A batch of more than batch_limit() messages, 1000 unless set otherwise, is refused whole, before any of it runs:

import rpc

var service = rpc.Service()
  .set_batch_limit(2)
  .on_request('ping', @() => 'pong')

var call = '{"jsonrpc": "2.0", "id": 1, "method": "ping"}'

echo service.answer('[${call}, ${call}, ${call}]')
{"jsonrpc":"2.0","id":null,"error":{"code":-32600,"message":"Invalid Request: a batch of 3 messages is over the limit of 2"}}

set_batch_limit(nil) takes a batch of any size, for a service that only ever hears from programs it trusts.

JSON-RPC Over HTTP

Over HTTP, each message, or batch, is the body of a POST, and its answer is the body of the response. Every exchange stands on its own: nothing is framed, and no connection is kept between calls beyond what HTTP keeps for itself. It is how services call each other, and how most of the JSON-RPC in the world is spoken.

Serving a Service

rpc.http_handler() turns a service into a route handler for an http server. Here the server runs on an isolate of its own and the client calls it:

import rpc
import isolate

def serve_calculator(ready) {
  import http
  import rpc

  var calculator = rpc.Service()
    .on_request('add', @(params) => params[0] + params[1])

  var server = http.server(0, '127.0.0.1')

  server.post('/rpc', rpc.http_handler(calculator))
  server.bind()
  ready.send(server.socket.local_address().port())
  server.listen()
}

var ready = isolate.channel(1)
var server = isolate.spawn(serve_calculator, ready)
var calculator = rpc.HttpClient('http://127.0.0.1:${ready.recv()}/rpc')

echo calculator.request('add', [2, 3])
echo calculator.request_batch([['add', [1, 2]], ['add', [3, 4]]])

server.cancel()
5
[3, 7]

In production the service is built in the setup each http worker runs, and the server spreads its requests across a pool of isolates:

import http
import rpc

def setup(server) {
  var api = rpc.Service()
    .on_request('add', @(params) => params[0] + params[1])

  server.post('/rpc', rpc.http_handler(api))
}

http.serve(setup, { host: '0.0.0.0', port: 8545 })

Everything http offers a route applies to this one: middleware, TLS, HTTP/2, compression, body limits, and the rest of Chapter 15.

A request that arrives over HTTP is answered in the response to that HTTP request, so its handler answers by returning: context.defer() raises there, since there is no connection to send a later answer on.

Calling a Service

rpc.HttpClient calls a service at a URL. Each call is one POST:

import rpc

var node = rpc.HttpClient('https://node.example.com', {
  headers: { Authorization: 'Bearer ${token}' },
  timeout: 10,
})

echo node.request('eth_blockNumber', [])
node.notify('log', ['checked the block number'])

var answers = node.request_batch([
  ['eth_blockNumber', []],
  ['eth_gasPrice', []],
])

request() returns the result or raises the RpcError the service answered with. notify() expects nothing back. request_batch() sends every call in one POST and returns the answers in the order of the calls, whatever order the server sent them in, with each failed call’s RpcError in its place.

The options are headers, sent with every call, timeout, the most seconds to wait for an answer, and client, the http.HttpClient to send through, for its proxy, TLS and connection settings. Each call can also take a timeout of its own.

Statuses

The statuses on the wire follow common practice:

StatusWhen
200an answer, whether it is a result or a JSON-RPC error
204nothing to answer: a notification, or a batch of them
405a method other than POST, with an Allow: POST header
415a body that is not application/json in UTF-8

A JSON-RPC error never changes the status: METHOD_NOT_FOUND is a 200 whose body says so. The http server’s own limits apply before the service sees anything, so a body over its max_body_size is a 413. Register the handler with server.any() and every method other than POST gets its 405; with server.post(), they get the server’s 404.

On the calling side, any other status raises an RpcHttpError carrying the status and the body. Some servers send a JSON-RPC error with a status other than 200, as older JSON-RPC over HTTP drafts asked for; when the body of such a response is a JSON-RPC answer, HttpClient reads it as one, and the error raises as the RpcError it is.

Who Is Calling

Every handler gets the HTTP request it came in as context.request, so it can read headers, the client’s address, or whatever a middleware left in the request’s context. Authentication belongs in a middleware, which refuses a request before the service sees it:

def setup(server) {
  var api = rpc.Service()
    .on_request('balance', @(params, context) {
      return accounts.balance(context.request.context.user)
    })

  server.use(@(request, response, next) {
    var user = tokens.user_for(request.bearer_token())

    if user == nil {
      response.text('who are you?\n', 401)
      return
    }

    request.context.user = user
    next()
  })

  server.post('/rpc', rpc.http_handler(api))
}

Framing

A stream of bytes has no edges. A program reading a socket gets bytes in whatever pieces the network delivers them, and two messages may arrive in one piece, or one message in ten. A framing says where each message ends, so the reader can put them back together.

Both framings work the same way. feed() takes the next bytes, in whatever pieces they come, and returns every message they complete. frame() goes the other way, turning a message’s text into the bytes to send.

Content-Length Headers

HeaderFraming puts a short header block in front of each message, giving its length in bytes:

import rpc

var framed = rpc.HeaderFraming().frame('{"jsonrpc":"2.0","method":"ping"}')

echo framed.to_string().split('\r\n')
[Content-Length: 33, , {"jsonrpc":"2.0","method":"ping"}]

The header block is ASCII: lines of Name: value, each ended by \r\n, then a blank line, then exactly as many bytes as Content-Length says. The length counts bytes rather than characters, so a message in any language frames correctly.

Reading it back, the pieces can be split anywhere at all:

import rpc

var framing = rpc.HeaderFraming()
var stream = framing.frame('{"a":1}') + framing.frame('{"b":2}')

var first = framing.feed(stream[0, 30])
var second = framing.feed(stream[30, stream.length()])

echo first.map(@(m) => m.to_string())
echo second.map(@(m) => m.to_string())
echo framing.pending()
[{"a":1}]
[{"b":2}]
0

The first piece held all of one message and the start of the next. feed() handed back the whole one and kept the rest, and pending() says how many bytes it is keeping. The second piece finished the second message.

Header names are read without regard to case. Content-Length is required. Content-Type may be given, and if it names a charset, the charset must be UTF-8. Any other header is read past and ignored.

One Message to a Line

LineFraming ends each message with a line break instead, the framing known as newline-delimited JSON:

import rpc

var framing = rpc.LineFraming()
var lines = framing.feed('{"a":1}\n{"b":'.to_bytes())

echo lines.length()
echo framing.feed('2}\r\n'.to_bytes())[0].to_string()
echo framing.frame('{"c":3}').to_string().trim()
1
{"b":2}
{"c":3}

A line may end in \n or \r\n, and an empty line is skipped. json.encode() never writes a raw line break, since a line break inside a string is escaped, so a line is always exactly one message. A peer that pretty-prints its JSON over several lines cannot use this framing.

One Message to Each Read

MessageFraming frames nothing. It is for a transport that keeps messages apart itself, a WebSocket above all, whose every read returns exactly one message: feed() takes each piece as one whole message, and frame() hands the text back as it is.

import rpc

var framing = rpc.MessageFraming()

echo framing.feed('{"a":1}'.to_bytes())[0].to_string()
echo framing.frame('{"b":2}').to_string()
{"a":1}
{"b":2}

A transport that wants it says so with a framing() method of its own, and an endpoint over it uses that framing unless told otherwise. rpc.websocket() does exactly that.

Choosing a Framing

Both ends must use the same framing, so the choice is usually made by whatever is on the other end. An endpoint uses its transport’s own framing when it has one, and HeaderFraming otherwise, unless told something else:

var endpoint = rpc.endpoint(transport).set_framing(rpc.LineFraming())

Where the choice is yours, HeaderFraming is the sturdier of the two: a reader knows how much is coming before it arrives, and refuses a message that is too large without reading it.

Limits

Every framing refuses a message larger than max_size(), 64 MiB by default, and HeaderFraming refuses a header block longer than 8 KiB. Without them, a peer could make a reader hold any amount of memory by never finishing a message.

import rpc

var framing = rpc.HeaderFraming().set_max_size(1024)

catch {
  framing.feed('Content-Length: 4096\r\n\r\n'.to_bytes())
} as error {
  echo error.message
}
a message of 4096 bytes is over the 1024-byte limit

The message is refused as soon as its header arrives, before any of it is read.

A framing that cannot read its stream raises RpcFramingError. Every reason is the same in one respect: nothing after it in the stream can be found, because the reader no longer knows where the next message starts. The connection is over at that point.

Transports

A transport carries the bytes. It is any object with three methods:

  • read(max) returns the next bytes that arrive, at most max of them, waiting until at least one does. Once the stream has ended, it returns empty bytes.
  • write(data) sends all of data.
  • close() ends the stream in the direction it writes.

A transport may also have can_listen(), true when another isolate can read it while the one that made it writes to it. That is what listen() needs; without the method, the answer is false.

Standard Streams

rpc.stdio() talks over the program’s own standard input and output. It is the transport of a program that another program starts and talks to: the parent writes to the child’s stdin and reads its stdout.

# calculator.zu
import io
import rpc

rpc.endpoint(rpc.stdio())
  .on_request('add', @(params) => params[0] + params[1])
  .on_notification('log', @(params) {
    io.stderr.write('${params[0]}\n')
  })
  .serve()

Stdout belongs to the protocol. Anything else the program prints there lands in the middle of the stream and breaks it for the other side, so a program serving over stdio sends everything else to stderr, as this one does with its log. Every write the transport makes is flushed at once.

Sockets

rpc.socket() talks over a connected TcpStream, UnixStream or TlsStream from net:

import net
import rpc
import isolate

var ready = isolate.channel()

var server = isolate.spawn(@(ready) {
  import net
  import rpc

  var listener = net.TcpStream()
  listener.bind('127.0.0.1:0')
  ready.send(listener.local_address().to_string())

  var connection = listener.accept()

  rpc.endpoint(rpc.socket(connection))
    .on_request('add', @(params) => params[0] + params[1])
    .serve()

  connection.close()
  listener.close()
}, ready)

var stream = net.TcpStream()
stream.connect(ready.recv())

var client = rpc.endpoint(rpc.socket(stream))

echo client.request('add', [20, 22])
client.close()
server.join()
42

The server binds to port 0, so the system picks a free one, and sends the address back over a channel for the client to connect to. It serves the one connection it accepts, and its serve() returns when the client closes its end.

A socket is held by one isolate at a time, so a socket endpoint reads with serve() and request() rather than listen().

This server serves one connection and ends. A server for many clients at once is rpc.serve(), which gives every connection an endpoint of its own.

Child Processes

rpc.process() talks to a child process over its standard streams. Spawn it with stdin and stdout both 'pipe':

import os
import rpc

var child = os.spawn('zuri', ['run', 'calculator.zu'], {
  stdin: 'pipe',
  stdout: 'pipe',
})
var calculator = rpc.endpoint(rpc.process(child))

echo calculator.request('add', [2, 3])  # 5
calculator.notify('log', ['added two numbers'])

calculator.close()
child.wait()

Closing the endpoint closes the child’s stdin. The calculator above reads that as the end of the connection, its serve() returns, and the program ends, which is what wait() waits for. A child’s streams are held by the isolate that spawned it, so this endpoint reads with serve() and request() too.

WebSockets

rpc.websocket() talks over a WebSocket from http.websocket, either one a route accepted or one websocket.connect() opened. Each JSON-RPC message is one WebSocket message, so an endpoint over it uses a MessageFraming without being told to.

A WebSocket is a connection in both directions, so either side calls the other whenever it likes. That is what makes it the transport of choice for a server that pushes to its clients, such as a node streaming new blocks to whoever subscribed:

import http.websocket
import rpc
import isolate

def serve_shop(ready) {
  import http
  import http.websocket
  import rpc

  var server = http.server(0, '127.0.0.1')

  server.get('/ws', @(request, response) {
    var socket = websocket.accept(request, response)

    rpc.endpoint(rpc.websocket(socket))
      .on_request('total', @(params, context) {
        var price = context.endpoint.request('price_of', [params.item])

        return price * params.count
      })
      .serve()
  })

  server.bind()
  ready.send(server.socket.local_address().port())
  server.listen()
}

var ready = isolate.channel(1)
var shop = isolate.spawn(serve_shop, ready)
var socket = websocket.connect('ws://127.0.0.1:${ready.recv()}/ws')

var client = rpc.endpoint(rpc.websocket(socket))
  .on_request('price_of', @(params) => params[0] == 'tea' ? 3 : 5)

echo client.request('total', { item: 'tea', count: 4 })

client.close()
shop.cancel()
12

The route accepts the WebSocket and serves an endpoint on it for as long as the client stays connected. Asked for a total, the server asks the client for a price before it answers. A WebSocket is held by one isolate at a time, so an endpoint over one reads with serve() and request().

Between Isolates

rpc.pipe() returns two transports joined to each other: what one writes, the other reads. Each end is a ChannelTransport over a pair of isolate channels, and channels cross isolates, so either end can be handed to another isolate and used from there.

import rpc

var ends = rpc.pipe()

ends[0].write('ping'.to_bytes())
echo ends[1].read(4096).to_string()
ping

A pipe is the natural way to give one part of a program a JSON-RPC interface to another, and the easiest way to test an endpoint: the server and its test talk exactly as they would over a socket, with nothing listening on the network.

A Transport of Your Own

Anything with the three methods is a transport. This one wraps another and counts what it sends:

import rpc
import isolate

class Counted {

  @new(inner) {
    self.inner = inner
    self.sent = 0
  }

  read(max) {
    return self.inner.read(max)
  }

  write(data) {
    self.sent += data.length()
    self.inner.write(data)
  }

  close() {
    self.inner.close()
  }
}

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('add', @(params) => params[0] + params[1])
    .serve()
}, ends[1])

var counted = Counted(ends[0])
var client = rpc.endpoint(counted)

client.request('add', [1, 2])
client.close()

echo '${counted.sent} bytes sent'
76 bytes sent

Endpoints

An endpoint is a service on a connection. rpc.endpoint(transport) makes one, and everything in Services holds for it: its handlers, its fallback, its batch limit and its errors. The connection adds the other direction. An endpoint calls the other side, and a handler on one can answer later than it returns.

Calling the Other Side

request() sends a request and waits for its answer:

var sum = client.request('add', [2, 3])

It returns the result, raises the RpcError the other side answered with, and raises RpcClosedError if the connection ends first. While it waits, it handles everything else that arrives, exactly as serve() would: other answers, notifications, and requests from the other side.

notify() sends a notification, and returns as soon as it is sent:

client.notify('log', { level: 'info', text: 'started' })

send_request() sends a request and returns at once, with the request’s id. The answer goes to a callback when it arrives:

import rpc
import isolate

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('add', @(params) => params[0] + params[1])
    .on_request('fail', @(params) {
      raise rpc.RpcError(-32001, 'no')
    })
    .serve()
}, ends[1])

var client = rpc.endpoint(ends[0])

client.send_request('add', [1, 2], @(result, error) {
  echo 'add: ${result}'
})
client.send_request('fail', nil, @(result, error) {
  echo 'fail: ${error.message}'
})

echo 'waiting'
echo client.request('add', [3, 4])
client.close()
waiting
add: 3
fail: no
7

A callback is called with the result and nil, or with nil and the error. Callbacks run on the isolate that owns the endpoint, while it is reading: here, while request() waited for its own answer, the two earlier answers arrived first and their callbacks ran. When the connection ends with a callback still waiting, it is called with an RpcClosedError.

Sending a Batch

request_batch() sends several requests as one batch and waits for every answer. Each call is a pair of a method and its params, and the answers come back in the order of the calls, whatever order the other side sent them in:

import rpc
import isolate

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('add', @(params) => params[0] + params[1])
    .serve()
}, ends[1])

var client = rpc.endpoint(ends[0])

var answers = client.request_batch([
  ['add', [1, 2]],
  ['multiply', [3, 4]],
  ['add', [5, 6]],
])

for answer in answers {
  if instance_of(answer, rpc.RpcError) {
    echo 'failed: ${answer.message}'
  } else {
    echo answer
  }
}

client.close()
3
failed: Method not found: multiply
11

A batch can partly succeed, so a failed call does not raise: its place in the list holds the RpcError it was answered with. Only an ending connection or a timeout raises, since then no answer can be trusted to come.

Answering Later

A handler normally answers by returning. Sometimes the answer comes from work still to be done: a job on another isolate, a reply from a third program, an event that has not happened yet. The handler then calls context.defer(), keeps the context, and answers through it with reply() or fail() when the answer is ready:

import rpc

var ends = rpc.pipe()
var waiting = []

var server = rpc.endpoint(ends[0])
  .set_framing(rpc.LineFraming())
  .on_request('next_job', @(params, context) {
    waiting.append(context.defer())
  })
  .on_notification('add_job', @(params) {
    for context in waiting {
      context.reply(params)
    }

    waiting = []
  })

server.handle({ jsonrpc: '2.0', id: 1, method: 'next_job' })
server.handle({ jsonrpc: '2.0', id: 2, method: 'next_job' })
server.handle({ jsonrpc: '2.0', method: 'add_job', params: { file: 'cat.png' } })

echo ends[1].read(4096).to_string().trim()
echo ends[1].read(4096).to_string().trim()
{"jsonrpc":"2.0","id":1,"result":{"file":"cat.png"}}
{"jsonrpc":"2.0","id":2,"result":{"file":"cat.png"}}

Once a request is deferred, what its handler returns is ignored. It is answered exactly once: answering it a second time, or answering a request that was never deferred, raises ValueError. A handler that raises after deferring fails the request with that error, unless it was already answered. A request deferred from inside a batch is answered on its own, after the batch’s other answers.

Calls in Both Directions

Either side can call the other at any time, including from inside a handler. Here the server, asked for an order’s total, asks the client for a price first:

import rpc
import isolate

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('total', @(params, context) {
      var price = context.endpoint.request('price_of', [params.item])

      return price * params.count
    })
    .serve()
}, ends[1])

var prices = { tea: 3, cake: 5 }
var client = rpc.endpoint(ends[0])
  .on_request('price_of', @(params) => prices[params[0]])

echo client.request('total', { item: 'tea', count: 4 })
client.close()
12

The client’s request('total') is waiting when the server’s request for price_of arrives, and it answers that while it waits, then goes on waiting for its own answer. Nothing has to be arranged for this: every wait handles what arrives.

Reading the Connection

An endpoint reads its transport in one of two ways, and every handler runs on the isolate that owns the endpoint either way.

Serving

serve() reads and answers messages until the connection ends, stop() is called, or the endpoint is closed. It is the whole program for a server that does nothing but answer, as every server so far in this chapter has been.

A message that is not valid UTF-8, or not valid JSON, is answered with PARSE_ERROR and an id of nil, since its id cannot be read. A stream whose framing cannot be read raises RpcFramingError out of serve().

Listening While Doing Other Work

A program that waits on other things as well cannot sit inside serve(). listen() reads the transport on an isolate of its own, decodes each message there, and returns a Channel of what arrives. The program takes items from it when it is ready, alongside its other channels, and passes each to dispatch():

import rpc
import isolate

var ends = rpc.pipe()

var caller = isolate.spawn(@(transport) {
  import rpc

  var client = rpc.endpoint(transport)
  var squares = [3, 4, 5].map(@(n) => client.request('square', [n]))

  client.close()

  return squares
}, ends[0])

# The work runs on a worker of its own, and its results come back on
# `done` whenever they are ready.
var jobs = isolate.channel()
var done = isolate.channel()

isolate.spawn(@(jobs, done) {
  var job = jobs.recv()

  while job != nil {
    done.send({ id: job.id, result: job.value * job.value })
    job = jobs.recv()
  }
}, jobs, done)

var waiting = {}
var server = rpc.endpoint(ends[1])
  .on_request('square', @(params, context) {
    waiting[context.id] = context.defer()
    jobs.send({ id: context.id, value: params[0] })
  })

var inbox = server.listen()

while !server.is_closed() {
  var ready = isolate.select([inbox, done])

  if ready[0] == inbox {
    server.dispatch(ready[1])
  } else {
    waiting[ready[1].id].reply(ready[1].result)
    waiting.remove(ready[1].id)
  }
}

jobs.close()
echo caller.join()
[9, 16, 25]

The server defers each request and hands the work to a worker. Its loop waits on both the connection and the worker’s results, and whichever is ready first is dealt with first, so the endpoint never stops answering while work is under way.

The reader isolate does the framing and the JSON decoding, so a large message costs the owning isolate nothing until it is dispatched. Everything is still sent from the owning isolate, and every handler still runs there.

Only a transport another isolate can read can be listened to. That is stdio and a pipe; a socket, a WebSocket and a child process are held by one isolate at a time, and listen() on one raises ValueError. Set the framing before calling listen(), since the reader isolate takes the framing with it.

serve() and request() work on a listening endpoint too, taking what arrives from the same channel.

Reading It Yourself

A program that reads the transport itself, such as one polling many connections at once, hands each read to feed(). It handles every message the bytes complete, exactly as serve() would have, and empty bytes end the connection. rpc.serve() drives every connection it holds this way.

Timeouts

request() and request_batch() take a timeout in seconds:

catch {
  var report = client.request('build_report', nil, 30)
} as error {
  if instance_of(error, rpc.RpcTimeoutError) {
    echo 'gave up on the report'
  }
}

A timeout needs the endpoint to be listening. An endpoint reading its transport itself is inside that transport’s read while it waits, and waits as long as the read does. Over a socket, the socket’s own set_read_timeout() is the limit, and a read that runs out raises the socket’s error.

An answer that arrives after its request timed out has nothing waiting for it, and goes to on_error().

Stopping and Closing

stop() makes serve() return once the message it is handling is done, which lets a handler end the conversation:

server.on_request('shutdown', @(params, context) {
  context.endpoint.stop()
  return 'bye'
})

close() closes the transport. Nothing more can be sent, and sending raises RpcClosedError; every send_request() callback still waiting is called with an RpcClosedError. is_closed() is true once the endpoint is closed or the connection has ended from the other side.

Serving Many Connections

rpc.serve() binds a listening socket and gives every connection accepted on it an endpoint of its own, spread across a pool of worker isolates. Each worker polls the connections it holds rather than blocking on one, so a quiet client costs a descriptor and not a thread, and one worker serves hundreds of long-lived connections side by side.

import isolate
import net
import rpc

isolate.configure(6)

def setup(endpoint, peer) {
  endpoint.on_request('add', @(params) => params[0] + params[1])
}

def run(ready) {
  import rpc

  rpc.serve(setup, {
    port: 0,
    workers: 2,
    framing: 'line',
    max_connections: 1,
    on_ready: @(address, stop) {
      ready.send(address.to_string())
    },
  })
}

var ready = isolate.channel(1)
var server = isolate.spawn(run, ready)

var stream = net.TcpStream()
stream.connect(ready.recv())

var client = rpc.endpoint(rpc.socket(stream)).set_framing(rpc.LineFraming())

echo client.request('add', [20, 22])
client.close()
server.join()
42

setup runs inside a worker for every connection, with the connection’s endpoint and the client’s address, and gives the endpoint its handlers. A worker shares nothing with the others, so setup is a function of a module, or one that uses nothing but its own imports, and builds whatever a connection needs. Each endpoint is a full peer: its handlers can call the client back, and it can notify the client whenever it has something to say.

OptionMeaningDefault
hostthe address to listen on'127.0.0.1'
portthe port to listen on; 0 picks a free one8000
patha Unix domain socket to listen on, in place of host and portnil
workershow many worker isolates serve connectionsthe number of CPUs
backloghow many accepted connections may wait for a workerworkers * 4
framing'header' for Content-Length headers, 'line' for one message to a line'header'
max_message_sizethe largest message accepted, in bytes64 MiB
max_connections_per_workerthe most connections one worker holds256
idle_timeoutseconds a connection may stay silent before it is closednil, never
read_timeoutseconds a read may wait once a message has begun30
write_timeoutseconds a write may wait30
cert_chain, private_keyPEM strings that put every connection behind TLSnil
on_readycalled once bound, with the address and a stop functionnil
max_connectionsstop accepting after this manynil, never

on_ready is called on the isolate that called serve(), once the socket is bound. The stop it is given stops accepting connections; every worker then finishes the message it is handling, closes its connections and ends, and serve() returns. That is what a signal handler calls to shut a server down cleanly. Reaching max_connections also stops accepting, but serves the connections already accepted until each closes, which is what the example above relies on.

Each worker holds a thread of the isolate pool for as long as the server runs. When serve() is the first thing in a program to use an isolate it sizes the pool itself; otherwise, call isolate.configure() first thing, as the example does, with room for the workers and whatever else the program runs.

With cert_chain and private_key, every connection is TLS, and a client that fails the handshake is closed without reaching a worker’s handlers. With path, the server listens on a Unix domain socket, the way local daemons offer an API to the programs on their machine; the socket file is removed when the server stops, and Unix domain sockets need a platform that has them, which net.unix.is_supported() answers.

A handler that calls its client with request() holds its worker until the answer arrives, and every other connection on that worker waits with it. For a call that can take a while, send_request() with a callback keeps the worker free.

Watching the Conversation

on_trace() sees the text of every message, as it is sent and as it arrives:

import rpc
import isolate

var ends = rpc.pipe()

isolate.spawn(@(transport) {
  import rpc

  rpc.endpoint(transport)
    .on_request('add', @(params) => params[0] + params[1])
    .serve()
}, ends[1])

var client = rpc.endpoint(ends[0])
  .on_trace(@(direction, text) {
    echo '${direction}: ${text}'
  })

client.request('add', [1, 2])
client.close()
out: {"jsonrpc":"2.0","id":1,"method":"add","params":[1,2]}
in: {"jsonrpc":"2.0","id":1,"result":3}

on_error() sees every problem the other side is not told about: an error a notification handler raised, an error other than an RpcError a request handler raised, an error a callback raised, a response nothing was waiting for, and a connection that failed. It is called with the error and the message it concerns, as a dictionary, or nil when no one message is to blame:

server.on_error(@(error, message) {
  io.stderr.write('${error.type}: ${error.message}\n')
})

Without an on_error() handler, these are dropped. A server that leaves it unset has no way to learn its own handlers are failing, so set it.

Errors

Every error the module raises is one of these:

RpcErrora JSON-RPC error, with a code: the other side’s answer, or a message that broke the rules
RpcHttpErroran HTTP status with no JSON-RPC answer, with the status and the body
RpcFramingErrorthe stream’s framing cannot be read, and nothing after it can be either
RpcClosedErrorthe connection ended before an answer, or the endpoint has been closed
RpcTimeoutErrora request’s timeout ran out before its answer
ValueErrora mistake in the calling program: params that are not structured, a reserved method name, a request answered twice

The first is the one a program handles as part of its work. The next four are about the transport, and the last is a bug to fix.

What the Module Refuses

It will not send params that are not a list or a dictionary. The specification allows nothing else, and a message that breaks it is refused when it is built, not when the other side receives it.

It will not accept a message that breaks the specification. A missing or wrong jsonrpc member, a method that is not a string, an id that is not a string, a number or null, a response with both a result and an error: each is refused with the code the specification gives it, never guessed at.

It will not read text that is not UTF-8. JSON is UTF-8, and a message that is not is answered with PARSE_ERROR rather than read with its bad bytes replaced. A header block must be ASCII, and an HTTP body that declares another charset is refused with 415.

It will not answer a response. A malformed response goes to on_error(), never back to the peer, so two endpoints can never fall into answering each other’s errors forever.

It will not let a handler claim a reserved name. Method names that start with rpc. belong to the specification, and neither on_request(), on_notification() nor on_unhandled() will answer them.

It will not do unbounded work for one message. A batch over the service’s limit is refused whole, and a framing refuses an oversized message from its header, before reading a byte of it.

It will not defer what has nowhere to be answered. A request that arrived over HTTP is answered in its HTTP response, and defer() on it raises rather than leaving the client waiting for an answer that cannot come.

It will not listen to a transport only one isolate can hold. A socket, a WebSocket or a child process is read where it is held, and listen() on one raises rather than reading it from somewhere it cannot be.

Module Reference

The standard library reference documents every class and method. The shape of the module:

rpc.Service()a service: handlers, and answer()
rpc.http_handler(service)an http route handler answering with a service
rpc.HttpClient(url, options)a client calling a service over HTTP
rpc.endpoint(transport)an Endpoint over a transport
rpc.serve(setup, options)an endpoint for every connection to a listening socket
rpc.stdio()the program’s own standard streams
rpc.socket(stream)a connected net stream
rpc.process(child)a child process’s standard streams
rpc.websocket(socket)a WebSocket from http.websocket
rpc.pipe()two transports joined to each other
rpc.encode(message)a message, or a list of them, as JSON text
rpc.decode(text)JSON text as a message, or a batch of them
rpc.read(value)a decoded value as a message, or a batch of them

The messages:

Request(method, params, id)a call that expects an answer
Notification(method, params)a call that expects none
Response(id, result, error)an answer, also made by Response.success() and Response.failure()
RpcError(code, message, data)an error, with to_dict()
PARSE_ERROR, INVALID_REQUEST, METHOD_NOT_FOUNDthe codes the specification defines
INVALID_PARAMS, INTERNAL_ERRORthe rest of them
SERVER_ERROR_MIN, SERVER_ERROR_MAXthe range set aside for implementations

On a Service, and so on an Endpoint:

on_request, on_notification, on_unhandledwhat it answers
on_error, on_tracewhat it reports
set_batch_limit, batch_limitthe most messages a batch may hold, DEFAULT_BATCH_LIMIT by default
answerone message or batch in, its answer out

On an Endpoint as well:

set_framinghow its messages are told apart
request, request_batch, send_request, notifycalling the other side
send, handlesending and handling messages as they are
serve, stopreading on the isolate that calls it
listen, dispatchreading on an isolate of its own
feedhandling what the program read itself
close, is_closedending the connection

On an HttpClient: request, request_batch, notify and send.

On a Context:

endpoint, id, method, requestwhat is being handled, where, and what it arrived in
is_notificationwhether anything is answered
defer, is_deferredtaking over answering
reply, fail, is_answeredanswering a deferred request

The framings, HeaderFraming, LineFraming and MessageFraming:

feed(data)the messages the next bytes complete
frame(message)the bytes that send a message
set_max_size(size), max_size()the largest message accepted
pending()how many bytes are held for an unfinished message

The transports, StdioTransport, SocketTransport, ProcessTransport, WebSocketTransport and ChannelTransport, each have read(max), write(data), close() and can_listen(), and WebSocketTransport has framing().

Editor Support

Zuri ships a language server. zuri lsp speaks the Language Server Protocol, so any editor that speaks it too gets completion, navigation, diagnostics as you type, refactoring, formatting and test running for Zuri, from the same server that ships with the runtime. The server is written in Zuri, on the zuri module’s own parser and compiler, so what it reports is what the runtime would do with the same code.

Visual Studio Code has an extension that sets everything up. Neovim, Helix, Sublime Text and Emacs each need a few lines of configuration, shown at the end of this chapter.

What the Server Does

Writing Code

Completion offers what can be written at the cursor:

  • after a ., the members of the receiver’s type: an instance’s fields and methods, inherited ones and the ones extensions add, a class’s statics, a module’s exports, a built-in type’s methods, and the keys of a dictionary written out where its variable was declared;
  • elsewhere, the names in scope, the built-in globals, and the keywords that can start what is being written;
  • after import, module paths, and between an import’s braces, what the module exports;
  • at the start of a class member, @new and the operator decorators the class does not have yet, as method skeletons;
  • after the : of a type hint, the type names;
  • inside a doc block, its tags after @, and after @param the names of the parameters it documents.

A name another module exports is offered with the import it needs: a module of the standard library or a package as import json and json.encode, and a module of the project as import .shapes { Point }.

Signature help follows a call while its arguments are written. It marks the parameter the cursor is on and shows its documentation, follows a constructor to its @new, and stays on a variadic parameter for every argument it takes.

Hover shows a declaration as a line of Zuri, its documentation laid out the way the standard library reference lays it out, the type a variable is inferred to hold, the value of a constant, and a module’s overview.

Inlay hints name the parameter each literal argument goes to, and, when turned on, the type a variable is inferred to hold.

Finding Your Way

Go to definition follows an imported name back to the module that declares it, and a built-in to the stub that documents it. Go to declaration stops at the import that brings a name into the file, and go to type definition goes to the class of the value a name holds. Go to implementation lists the methods of subclasses that override a method, or the subclasses of a class.

Find references and the highlights of the current name go by what each name resolves to, never by its spelling, so a distance method of one class is never confused with another class’s.

The call hierarchy shows what calls a function and what it calls; the type hierarchy, the classes above and below a class. Each file has an outline, and the workspace symbol search finds a declaration by the letters of its name in order, so hrq finds HttpRequest. Each import links to the file it loads, and each class, function and method shows how often it is used.

Diagnostics

Every error the server reports is either an error the compiler reports or a failure the runtime is certain to raise when the code runs:

  • a syntax error, with the recovered rest of the file still checked;
  • an error from the compiler, such as break outside a loop;
  • an import of a module that does not exist, or of a name its module does not have;
  • module.name for a name the module does not have;
  • a name nothing declares;
  • self.name = value outside @new, for a field no class in the chain declares;
  • a literal argument of the wrong type, or a missing one, for a typed parameter, with the runtime’s own message.

A member that an instance of a fully known class does not have is a warning, since an extension the server has not read could add it. Unused imports and declarations, and code after a return, raise, break or continue, are shown faded. A use of anything documented @deprecated is struck through.

The server never guesses. A value whose type cannot be told, a module whose wildcard imports lead somewhere unread, and a class whose superclass cannot be found are never diagnosed. Arity is not checked either, because Zuri does not enforce it.

Quick fixes add the import a missing name needs, correct a misspelt name or member, remove an unused import, and declare a field set outside @new.

Changing Code

Rename changes a declaration and every use of it across the workspace. A { name } dictionary entry keeps its key and becomes { name: renamed }. A rename is refused, with the reason, when the new name is a keyword, when it is already declared where the declaration or a use of it would see it, when it would make private something used from outside, and for what the standard library or a package declares. A member used through a receiver whose type cannot be told is left alone, and the editor is told where.

Extract and inline refactorings are offered only where they keep what the program does:

  • Extract variable moves an expression into a var just before its statement, never out of a loop’s condition, the right of and or or, a branch of ?:, a when label or an else if condition, and never past something with a side effect that runs first. Every identical copy in the block can share the variable when the expression has no side effect.
  • Extract function turns whole statements of one block into a function. The outer locals they use become its parameters, and what they set and is read afterwards comes back: one value directly, several in a dictionary. Statements that use self or parent become a private method.
  • Inline variable replaces each use of a var set once and never again with its value.
  • Inline function replaces a call of a function whose body is a single return with that expression, and removes the function once every call is inlined.

Formatting runs zuri fmt on a file or a selection and sends back only the lines that differ. A file with a syntax error is left as it is.

Renaming or moving a file updates the relative imports that reach it, in the files that import it and in the moved file itself.

Running Tests

The server finds the tests every file of the workspace declares with the test module: each describe and it, and their _only, _skip, _todo and _each forms, whose name is written as a string. A lens above each one runs it, and one above the first runs the whole file.

A test runs as zuri test would run it: zuri run on its file, in a process of its own, from the root of the project. Each result reaches the editor as it happens, and each failed assertion is shown on the line it failed at until the file changes or its tests run again.

Highlighting

Semantic highlighting colours each name by what it resolves to: a class, a function, a method, a parameter, a variable, a field, a module or a decorator, with whatever the standard library defines marked as such. It works with any editor theme that colours semantic tokens.

How the Server Reads a Project

The server reads every .zu file of the workspace folders the editor opens, in the background, on isolates of its own, so it answers while it reads. Inside a git work tree it reads what git tracks and what git would track, so .gitignore is honoured; anything under .git, .zuri or node_modules is left out, and so is whatever zuri.exclude names.

Imports resolve exactly as the runtime resolves them: a relative import against the importing file, then the project’s .zuri/libs, then the standard library, then the native modules, then the packages installed for the user in ZURI_HOME/libs. The standard library is read from the Zuri installation the server belongs to, or from zuri.root, or from ZURI_ROOT. A module the workspace imports is read the first time it is needed.

Settings

The server reads its settings from the zuri section of the editor’s configuration, as nested objects or dotted names, and follows changes as they are made.

SettingDefaultWhat it does
root''The installation whose standard library is read; empty uses ZURI_ROOT, or the installation beside zuri
exclude[]Globs, relative to a workspace folder, of files never read
index.workers2Isolates that read the workspace in the background
diagnostics.enabletrueWhether problems are reported at all
diagnostics.scope'openFiles''openFiles', or 'workspace' for every file of the workspace
diagnostics.delay300Milliseconds after a change before a file is checked
diagnostics.unusedHintstrueWhether unused declarations are shown faded
inlayHints.parameterNames'literals'Which arguments are named: 'none', 'literals' or 'all'
inlayHints.variableTypesfalseWhether inferred variable types are shown
codeLens.referencestrueWhether declarations show how often they are used
completion.autoImporttrueWhether names from modules not imported yet are offered
completion.callParenthesesfalseWhether completing a function writes its parentheses
testing.codeLenstrueWhether suites and tests show lenses that run them

Starting the Server

Editors start the server themselves, over its standard input and output:

zuri lsp [--stdio] [--log <path>] [--log-level <level>]

--stdio names the only way the server talks, and is accepted because editors pass it. --log appends the server’s log to a file as well as sending it to the editor. --log-level sets how much is logged: error, warn, info (the default), debug, or trace, which also writes every message the editor and the server exchange to the log file.

The server exits with 0 when the editor shuts it down in order, and with 1 when the connection ends without that, or when the editor that started it has exited.

Visual Studio Code

Install the Zuri extension from the marketplace, or from a downloaded package:

code --install-extension zuri-vscode-0.2.0.vsix

The extension starts the server for every workspace with a .zu file. Its settings are the server’s, under zuri., plus two of its own: zuri.path, the zuri executable to run, and zuri.trace.server, which records the messages between the editor and the server in the Zuri Language Server output. Tests appear in the Test Explorer, and the commands Zuri: Restart Language Server, Zuri: Show Language Server Output and Zuri: Check Zuri Installation are in the command palette.

Neovim

Neovim 0.11 configures language servers itself:

vim.filetype.add({ extension = { zu = 'zuri' } })

vim.lsp.config('zuri', {
  cmd = { 'zuri', 'lsp' },
  filetypes = { 'zuri' },
  root_markers = { 'project.toml', '.git' },
  settings = {
    zuri = {
      inlayHints = { variableTypes = true },
    },
  },
})

vim.lsp.enable('zuri')

Inlay hints are shown with vim.lsp.inlay_hint.enable(), and a code lens runs with vim.lsp.codelens.run().

Helix

Add the language and its server to languages.toml:

[[language]]
name = "zuri"
scope = "source.zuri"
file-types = ["zu"]
roots = ["project.toml"]
comment-token = "#"
block-comment-tokens = { start = "/*", end = "*/" }
indent = { tab-width = 2, unit = "  " }
language-servers = ["zuri-lsp"]

[language-server.zuri-lsp]
command = "zuri"
args = ["lsp"]

[language-server.zuri-lsp.config.zuri]
inlayHints = { parameterNames = "literals" }

Sublime Text

With the LSP package installed, add a client to its settings:

{
  "clients": {
    "zuri": {
      "enabled": true,
      "command": ["zuri", "lsp"],
      "selector": "source.zuri",
      "settings": {
        "zuri.diagnostics.scope": "openFiles"
      }
    }
  }
}

The selector names the syntax Zuri files open with. The Visual Studio Code extension’s grammar, zuri.tmLanguage.json, is a TextMate grammar whose scope is source.zuri; PackageDev’s Convert command turns it into a .tmLanguage file for Sublime Text.

Emacs

Eglot, built into Emacs 29, needs a mode for Zuri files and the command that starts the server:

(define-derived-mode zuri-mode prog-mode "Zuri"
  "A major mode for Zuri source."
  (setq-local comment-start "# "))

(add-to-list 'auto-mode-alist '("\\.zu\\'" . zuri-mode))

(with-eval-after-load 'eglot
  (add-to-list 'eglot-server-programs '(zuri-mode "zuri" "lsp")))

(add-hook 'zuri-mode-hook #'eglot-ensure)

Settings go in eglot-workspace-configuration:

(setq-default eglot-workspace-configuration
              '(:zuri (:inlayHints (:variableTypes t))))

When Something Goes Wrong

The server sends what it logs at its --log-level and above, info by default, to the editor’s log for language servers, and while the editor asks for a trace, everything down to debug as well. Starting it with --log and --log-level debug writes all of it to a file; --log-level trace adds every message exchanged.

When nothing works at all, the editor cannot start zuri: running zuri --version in a terminal shows whether it is on PATH, and an editor that takes a path, as zuri.path does in Visual Studio Code, can be given the executable directly.

When the standard library cannot be found, the server says so in its log. Setting zuri.root, or ZURI_ROOT for the editor, to the Zuri installation fixes it.

Appendix A: Keywords

Zuri reserves thirty-one words. None of them can be used as a variable, function, class or parameter name.

KeywordWhat it doesCovered in
andlogical conjunction, short-circuitingOperators
asbinds an error in catch, or renames an importErrors, Modules
assertraises AssertError when its condition is falsyControl Flow
breakleaves the innermost loopControl Flow
catchruns a block, intercepting anything it raisesErrors
classdeclares a class, or with > an extension to oneClasses, Extensions
constdeclares a name that cannot be reassignedVariables
continueskips to the next iterationControl Flow
defdeclares a function, named or anonymousFunctions
defaultthe fall-through branch of a usingControl Flow
dobegins a do/while loopControl Flow
echoprints a value and a newlineHello, World!
elsethe alternative branch of an ifControl Flow
falsethe boolean falseData Types
foriterates over anything iterableControl Flow
ifconditional branchControl Flow
importloads a moduleModules
inseparates a for loop’s variables from its iterableControl Flow
iterthe counting loopControl Flow
nilthe absence of a valueData Types
orlogical disjunction, short-circuitingOperators
parentthe superclass constructor, or a superclass methodInheritance
raiseraises an errorErrors
returnleaves the current functionControl Flow
selfthe current instance, inside a methodClasses
staticputs a field or method on the class, not the instanceClasses
truethe boolean trueData Types
usingmulti-way branch on one subjectControl Flow
vardeclares a variableVariables
whenone branch of a usingControl Flow
whilethe conditional loopControl Flow

Names That Are Not Keywords

These are ordinary globals, not reserved words, so nothing stops you from shadowing one. Doing so is a good way to confuse the next reader:

time      sum       bytes     file      instance_of  typeof
delprop   getprop   hasprop   setprop   id           print
rand      is_bigint is_bool   is_bytes  is_callable  is_class
is_dict   is_file   is_function is_instance is_int    is_iterable
is_list   is_number is_object is_string

The built-in error classes are also globals: Error, TypeError, ValueError, NumericError, ArgumentError, NotImplementedError, RangeError, AccessError, AssertError, PropertyError, UndefinedError and ModuleNotFoundError.

Reserved by Convention

Two module-level names are provided by the runtime rather than declared by you: __file__ and __root__. See Modules.

Names beginning with $ are never produced by the lexer, which is how the compiler synthesises loop variables that cannot collide with yours.

A leading underscore marks something private, both for class members and for module members. The compiler enforces it in both cases.

Appendix B: Operators and Precedence

Precedence

Tightest first. Operators on the same row bind equally and associate left to right, except **, which associates right to left.

LevelOperatorsNotes
1literals, (...), [...], {...}, self, parent
2..binds primaries only
3. () []member, call, index and slice
4++ --postfix only
5**right-associative; the exponent may carry a unary operator
6! - ~unary
7* / // %
8+ -
9<< >> >>>
10&
11^
12|
13< <= > >= == !=
14and
15or
16? :
17= and every compound assignment

Consequences worth remembering:

  • 2 ** 3 ** 2 is 512. ** is right-associative.
  • 2 * 3 ** 2 is 18. ** outranks *.
  • -2 ** 2 is -4. ** binds tighter than unary minus; write (-2) ** 2 to raise a negative number.
  • 2 ** -1 is 0.5. The exponent may carry its own sign.
  • 1 + 2..5 is 1 + (2..5). Parenthesise ranges with computed endpoints.

Arithmetic

OperatorMeaningDecorator
+addition; string and list concatenation@add
-subtraction@sub
*multiplication; string and list repetition@mul
/division, always floating point@div
//floor division, rounds toward negative infinity@floordiv
%remainder, keeps the sign of the left operand@mod
**exponentiation@pow
- (unary)negation@neg

Comparison

OperatorMeaningDecorator
==equal, by value for numbers, strings, lists, dicts; by identity otherwise@eq
!=not equal, the negation of ==@eq
<less than, numbers only@lt
<=less than or equal@lte
>greater than@gt
>=greater than or equal@gte

@eq runs only when both operands are objects, so x == nil and x == 5 never call it.

Logic

OperatorMeaning
andboth truthy; returns the operand, not a bool
oreither truthy; returns the operand, not a bool
!logical negation, always a bool
? :conditional expression

Bitwise

OperatorMeaningDecorator
&and@and
|or@or
^xor@xor
~complement@not
<<left shift@lshift
>>arithmetic right shift@rshift
>>>logical right shift, zero filling@urshift

Assignment

OperatorEquivalent to
=assignment
+= -= *= /= //= **= %=x = x op y
&= |= ^= ~= <<= >>= >>>=x = x op y
++ --increment, decrement; postfix only, evaluates to the new value

Access

SyntaxMeaning
x.namemember of an object, dictionary or module
x[key]index with a computed key
x[a, b]slice, from a up to but not including b
x[, b]slice from the start
x[a, ]slice to the end
a..ba range value
...xvariadic parameter, in a parameter list only

Negative indices count back from the end, for strings, lists and bytes.

Truthiness

Falsy: false, nil, 0, 0.0, -0.0, NaN, 0n, '', and bytes(0).

Truthy: everything else, including every negative number, [], {} and '0'.

Operators Zuri Does Not Have

No in for membership; use contains(). No ?. optional chaining. No ?? null coalescing; or covers it, with the truthiness caveat above. No comma operator. No prefix ++/--.

Appendix C: Decorated Methods

A method whose name begins with @ is called by the runtime when a piece of syntax is applied to an instance of its class. They are ordinary methods otherwise: inherited, overridable, and callable by name.

See Decorated Methods for the guided treatment, and Class Extensions for adding one to a class you did not write.

Construction

DecoratorCalled bySignature
@newClassName(...)@new(...args)

@new is the only place self.x = value may declare a field that was not declared with var.

Arithmetic

DecoratorOperatorSignature
@adda + b@add(other)
@suba - b@sub(other)
@mula * b@mul(other)
@diva / b@div(other)
@floordiva // b@floordiv(other)
@moda % b@mod(other)
@powa ** b@pow(other)
@neg-a@neg()

Bitwise

DecoratorOperatorSignature
@anda & b@and(other)
@ora | b@or(other)
@xora ^ b@xor(other)
@not~a@not()
@lshifta << b@lshift(other)
@rshifta >> b@rshift(other)
@urshifta >>> b@urshift(other)

@not is bound to ~, the bitwise complement. ! is logical negation and is not overridable: an instance is always truthy, so !instance is always false.

Comparison

DecoratorOperatorSignature
@eqa == b, a != b@eq(other)
@lta < b@lt(other)
@ltea <= b@lte(other)
@gta > b@gt(other)
@gtea >= b@gte(other)

!= is the negation of @eq. @eq runs only when the right operand is an object too, so x == nil never calls it, and it must return a bool. Lists, dictionaries, contains() and index_of() compare instances by identity whether or not the class defines it.

Iteration

DecoratorCalled bySignature
@keyfor ... in@key(previous)
@valuefor ... in@value(key)

@key(previous) receives the previous key, starting from nil, and returns the next one or nil when the sequence is finished. @value(key) returns what is stored at that key.

Defining both is what makes is_iterable() return true for the class.

Display

DecoratorCalled bySignature
@to_stringecho, print()@to_string()

@to_string() returns the text shown for the instance, including when it sits inside a list or dictionary being shown, as a key or as a value. It must return a string; anything else raises a TypeError. Without it, an instance shows as <instance of ClassName>.

Serialisation

DecoratorCalled bySignature
@to_jsonjson.encode()@to_json()

Returns whatever should be encoded in the instance’s place, which is where you decide what does and does not cross the wire.

Not a Decorator

to_string() has no @. It is a real method every value already carries, and a class may override it.

Nothing calls it implicitly. String interpolation and + render an instance as <instance of ClassName>, and echo and print() use @to_string().

Resolution

The left operand decides. a + b looks for @add on a’s class only; if a is a number and b is your instance, the operation is a TypeError rather than a call to b’s @add.

An operator with no matching decorator raises a TypeError naming the exact signature:

operator '+' not defined for call signature (nil, number)

Appendix D: Built-in Functions

These functions are available in every file with no import. They are the parts of the language that happen to be spelled as calls rather than as syntax.

FunctionReturnsSummary
time()numberReturns the current epoch time to the microseconds resolution.
sum(...values: list)numberCalculates the sum of all the elements passed as arguments.
bytes(x: number|list)bytes|anyIf x is a number, this function returns a new bytes object with length x having all its bytes set to 0x0.
file(path: string, mode: ?string)fileReturns an open file handle to the file specified in the path in the specified mode.
instance_of(x, y)booleanReturns true if x is an instance of the given class y or false otherwise.
typeof(x)stringReturns the type of the given value as a string.
delprop(object: instance, name: string)voidDeletes the property name from the given instance of object.
getprop(object: instance, name: string)any|nilReturns the value of the property name from the given instance of object.
hasprop(object: instance, name: string)booleanReturns true if the property name exists in the given instance of object.
setprop(obj: instance, prop: string, value)booleanSets the value of the object’s property with the matching name to the given value.
id(x)numberReturns the unique identifier of value x within the system.
print(...values: list)voidPrints the given arguments to standard output.
rand(x: ?number, y: ?number)numberIf no argument is given, returns a random number between 0 and 1.
is_bigint(x)booleanReturns true if x is a bigint or false otherwise.
is_bool(x)booleanReturns true if x is a boolean or false otherwise.
is_callable(x)booleanReturns true if x is a callable or false otherwise.
is_class(x)booleanReturns true if x is a class or false otherwise.
is_dict(x)booleanReturns true if x is a dictionary or false otherwise.
is_function(x)booleanReturns true if x is a function or false otherwise.
is_instance(x)booleanReturns true if x is an instance of any class or false otherwise.
is_int(x)booleanReturns true if x is an integer or false otherwise.
is_list(x)booleanReturns true if x is a list or false otherwise.
is_number(x)booleanReturns true if x is a number or false otherwise.
is_object(x)booleanReturns true if x is an object or false otherwise.
is_string(x)booleanReturns true if x is a string or false otherwise.
is_bytes(x)booleanReturns true if x is bytes or false otherwise.
is_file(x)booleanReturns true if x is a file or false otherwise.
is_iterable(x)booleanReturns true if x is an iterable object or false otherwise.

time()

time() -> number

Returns the current epoch time to the microseconds resolution.

Example:

%> time()
1686787200.123456

The time is returned as a floating point number where the integer part represents the number of seconds since the epoch and the fractional part represents the microseconds.

Returns number

Note: The epoch time is the number of seconds that have elapsed since January 1, 1970 (midnight UTC/GMT).

sum()

sum(...values: list) -> number

Calculates the sum of all the elements passed as arguments. Returns 0 when no argument is passed in.

Example:

%> math.sum([1, 2, [3, 4, [5, 6]]])
21

Parameters

  • values (...number)

Returns number

bytes()

bytes(x: number|list) -> bytes|any

If x is a number, this function returns a new bytes object with length x having all its bytes set to 0x0.

If x is a list, it returns a new bytes object whose contents are the bytes specified in the list.

%> bytes(5)
(00 00 00 00 00)
%> bytes([65, 66, 67, 68, 69])
(41 42 43 44 45)

Parameters

  • x (number|list) — The number or list to convert to bytes.

Returns bytes|any

Note: If x is a list, then the list must only contain valid bytes which can be any number between 0 and 255.

file()

file(path: string, mode: ?string) -> file

Returns an open file handle to the file specified in the path in the specified mode. If the mode is not specified, the file will be opened in the read only mode.

Valid modes include:

%> file('sample.txt', 'r')
<file at sample.txt in mode r>
%> file('sample.txt', 'w')
<file at sample.txt in mode w>
%> file('sample.txt', 'a')
<file at sample.txt in mode a>
%> file('sample.txt', 'r+')
<file at sample.txt in mode r+>
%> file('sample.txt', 'w+')
<file at sample.txt in mode w+>
%> file('sample.txt', 'a+')
<file at sample.txt in mode a+>
%> file('sample.lock', 'x')
<file at sample.lock in mode x>

x and x+ create the file and fail when anything already exists at the path. The check and the creation are one step, so of two programs creating the same path in x mode, exactly one succeeds. That makes it the mode for lock files. The failure comes on the first read, write or open(), since creating the handle touches nothing.

Parameters

  • path (string) — The path to the file to open.
  • mode (?string) — The mode to open the file in.

Returns file

instance_of()

instance_of(x, y) -> boolean

Returns true if x is an instance of the given class y or false otherwise.

Parameters

  • x (any) — The value to check.
  • y (class) — The class to check for.

Returns boolean

typeof()

typeof(x) -> string

Returns the type of the given value as a string.

Parameters

  • x (any) — The value to check.

Returns string

delprop()

delprop(object: instance, name: string) -> void

Deletes the property name from the given instance of object.

Parameters

  • object (instance) — The instance to delete the property from.
  • name (string) — The name of the property to delete.

Returns void

getprop()

getprop(object: instance, name: string) -> any|nil

Returns the value of the property name from the given instance of object. If the object has no such property, nil is returned.

Parameters

  • object (instance) — The instance to get the property from.
  • name (string) — The name of the property to get.

Returns any|nil

hasprop()

hasprop(object: instance, name: string) -> boolean

Returns true if the property name exists in the given instance of object. If the object has no such property, false is returned.

Parameters

  • object (instance) — The instance to check for the property.
  • name (string) — The name of the property to check.

Returns boolean

setprop()

setprop(obj: instance, prop: string, value) -> boolean

Sets the value of the object’s property with the matching name to the given value. If the property already exists, it overwrites it and returns true, otherwise it returns false.

Parameters

  • obj (instance) — The object to set the property of.
  • prop (string) — The property to set.
  • value (any) — The value to set the property to.

Returns boolean

id()

id(x) -> number

Returns the unique identifier of value x within the system. This value is also equivalent to the current address of object x in memory.

Parameters

  • x (any) — The value to get the identifier of.

Returns number

print()

print(...values: list) -> void

Prints the given arguments to standard output.

Unlike echo (which always appends a newline and only ever prints one value), print() writes every argument back-to-back with no separator and no trailing newline. It also critically writes a bytes object as RAW bytes rather than its Display text. That raw-byte path is what lets a script stream binary output (e.g. a PBM/PNG image body one scanline at a time).

Parameters

  • values (...any) — Any number of arguments to print

Returns void

Note: In the REPL, it also appends a newline at the end.

rand()

rand(x: ?number, y: ?number) -> number

If no argument is given, returns a random number between 0 and 1. If x is given, returns a random number between 0 and x. If y is given, returns a random number between x and y.

Parameters

  • x (?number) — The lower bound of the random number.
  • y (?number) — The upper bound of the random number.

Returns number

is_bigint()

is_bigint(x) -> boolean

Returns true if x is a bigint or false otherwise. A bigint is a distinct type from number created either with the n literal suffix (123n) or by an operation whose result overflows what a regular number can represent exactly. is_number(x) and is_int(x) are both false for a bigint even though it holds an integer value; check is_bigint(x) separately when a value might be either.

Parameters

  • x (any) — The value to check.

Returns boolean

is_bool()

is_bool(x) -> boolean

Returns true if x is a boolean or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_callable()

is_callable(x) -> boolean

Returns true if x is a callable or false otherwise. Callables includes classes, functions, methods and closures.

Parameters

  • x (any) — The value to check.

Returns boolean

is_class()

is_class(x) -> boolean

Returns true if x is a class or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_dict()

is_dict(x) -> boolean

Returns true if x is a dictionary or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_function()

is_function(x) -> boolean

Returns true if x is a function or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_instance()

is_instance(x) -> boolean

Returns true if x is an instance of any class or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_int()

is_int(x) -> boolean

Returns true if x is an integer or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_list()

is_list(x) -> boolean

Returns true if x is a list or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_number()

is_number(x) -> boolean

Returns true if x is a number or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_object()

is_object(x) -> boolean

Returns true if x is an object or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_string()

is_string(x) -> boolean

Returns true if x is a string or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_bytes()

is_bytes(x) -> boolean

Returns true if x is bytes or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_file()

is_file(x) -> boolean

Returns true if x is a file or false otherwise.

Parameters

  • x (any) — The value to check.

Returns boolean

is_iterable()

is_iterable(x) -> boolean

Returns true if x is an iterable object or false otherwise. Iterables includes lists, dictionaries, strings, bytes, and instances of any class that defines both @key() and @value() decorator functions.

Parameters

  • x (any) — The value to check.

Returns boolean

Module Variables

Every file also has these variables, which describe the file itself rather than anything it declares.

__file__

The path of the file it is read in. In the file a program was started from, it is the path the program was started with; in an imported module, it is the module’s full path.

Type string

__root__

The path of the file the program was started from, the same in every module of one run and in every isolate, so __root__ == __file__ holds in that one file alone. That is how a file serves both as a module others import and as a program of its own:

def main() {
  echo 'started directly'
}

if __root__ == __file__ {
  main()
}

__root__ and __file__ are both set when a file runs; the REPL, which runs no file, defines neither.

Type string

Appendix E: Built-in Type Methods

Every method on every built-in type, with its signature, what it returns, and its edge cases. This is the reference; the chapters in Text, Numbers and Collections are the introduction.

Methods are called with a dot, on the value itself:

echo 'zuri'.upper()
echo 255.hex()
echo [3, 1, 2].sort()
ZURI
ff
[1, 2, 3]
TypeMethodsPage
string41String Methods
number43Number Methods
bigint26Bigint Methods
bool1Boolean Methods
list40List Methods
dict21Dictionary Methods
range9Range Methods
bytes25Bytes Methods
file27File Methods
function6Function Methods

What Is Not Listed Here

@key and @value exist on every iterable built-in type, which is what makes for ... in work on them. They are documented as a protocol in Appendix C rather than repeated on every page.

to_string() exists on every value, nil included.

Class instances carry whatever their class declares, plus to_string(). See Classes and Objects.

Module members are not methods. See Appendix F.

Conventions in These Pages

A parameter written name: ?type is optional. A parameter written ...name is variadic. A parameter with no type shown takes any value.

A method that mutates its receiver says so. Where a type has both, the distinction matters: list.sort() mutates and returns the list, while list.reverse() returns a new list and leaves the original alone.

String Methods

Every method on the built-in string type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
length()numberReturns the length of a string.
upper()stringReturns a copy of the string with all the cased characters converted to uppercase.
lower()stringReturn a copy of the string with all the cased characters converted to lowercase.
is_alpha()booleanReturns true if all the characters in the string are all alphabets and the string is not empty., otherwise returns false.
is_alnum()booleanReturns true if all the characters in the string are either alphabets or numbers and the string is not empty, otherwise returns false.
is_number()Returns true if all the characters in the string are all digits and the string is not empty, otherwise returns false.
is_lower()booleanReturns true if at least one character in the string is cased, all cased characters are lower cased and the string is not empty.
is_upper()booleanReturns true if at least one character in the string is cased, all cased characters are upper cased and the string is not empty.
is_space()booleanReturns true if there are only whitespace characters in the string and the string is not empty.
ord()numberReturns the Unicode code point of the string, which must be exactly one character long.
trim(chars: ?string)stringReturns a copy of the string with characters stripped from both ends.
ltrim(chars: ?string)stringReturns a copy of the string with characters stripped from its start only.
rtrim(chars: ?string)stringReturns a copy of the string with characters stripped from its end only.
join(string: string)stringReturns a string which is a concatenation of the items in the iterable using the string as the separator.
split(delimiter: string)listReturns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found.
index_of(str: string, start_index: ?number)numberReturns the index position of the first occurrence of the string str in the string string.
last_index_of(str: string, end_index: ?number)numberReturns the index position of the last occurrence of the string str in the string string, searching from the end.
starts_with(str: string)booleanReturns true if the string begins with the string or character specified in str, otherwise it returns false.
ends_with(str: string)booleanReturns true if the string ends with the string or character specified in str, otherwise it returns false.
count(str: string)numberReturns the number of non-overlapping occurrences of the substring str in the string.
to_number(base)numberReturns the first numeric value contained in the string if any exists or 0 if the string contains no numeric value.
to_bigint(base)bigintReturns the integer value of the string as a bigint, or 0n if the string does not spell one.
to_list()listReturns a list whose elements consists of every character contained in the string in order of appearance.
to_bytes()bytesReturns the content of the string as a stream of bytes.
lpad(width: number, fill: ?string)stringReturns the string left justified in a string of length width.
rpad(width: number, fill: ?string)stringReturns the string right justified in a string of length width.
match(str: string)boolean|dictionaryIf the string str is a regular string, this method returns true if the string contains a substring str.
matches(reg: string)dictionaryReturns a dictionary containing every match of the given regular expression reg in the source string.
replace(str: string, replacement: string, use_regex: ?bool)stringReturns a copy of the string with all occurrences or matches of str replaced by the replacement string.
replace_with(regex: string, callback: function)stringReturns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred.
ascii()stringReinterprets the string as a raw byte view: each byte of its UTF-8 encoding becomes its own character (a codepoint between 0 and 255, i.e.
case_fold()stringReturns a copy of the string case-folded for case-insensitive comparison, using full Unicode case folding rather than plain lowercasing.
compare(other: string)numberCompares the string with another string.
is_empty()booleanReturns true if the string is empty, false otherwise.
contains(str: string)booleanReturns true if the string contains the specified substring, false otherwise.
lines()listReturns the lines of the string as an list as it would be if split on newline characters.
each_line(callback: function)voidIterates over each line of the string, calling the provided callback function with the line and its index.
each(callback: function)voidIterates over each character of the string, calling the provided callback function with the character and its index.
capitalize()stringReturns a new string with the first character capitalized and the rest in lowercase.
title()stringReturns a new string with each word capitalized.
to_string()stringReturns the string itself.

length()

length() -> number

Returns the length of a string. Note that this method is UTF-8 compatible and will return the UTF-8 length for the string if the string contains UTF-8 characters whether written directly or via the \u or \U escapes.

For example:

%> 'This is a pretty long string'.length()
28
%> 'उनका एक समय'.length()
11
%> 'This text mixes English and 粵語'.length()
30

Returns number

upper()

upper() -> string

Returns a copy of the string with all the cased characters converted to uppercase. Note that the result of this method may return false when tested with is_upper() of the string contains Unicode characters that are not case folded.

For example:

%> 'zuri'.upper()
'ZURI'

Returns string

lower()

lower() -> string

Return a copy of the string with all the cased characters converted to lowercase.

For example:

%> 'Zuri Is Bae'.lower()
'zuri is bae'

Returns string

is_alpha()

is_alpha() -> boolean

Returns true if all the characters in the string are all alphabets and the string is not empty., otherwise returns false.

For example:

%> 'abracadabra'.is_alpha()
true
%> 'my tooth aches'.is_alpha()
false
%> ''.is_alpha()
false

Returns boolean

is_alnum()

is_alnum() -> boolean

Returns true if all the characters in the string are either alphabets or numbers and the string is not empty, otherwise returns false. This method is the same as string.is_alpha() or string.is_number().

For example:

%> '3Idiots'.is_alnum()
true
%> 'Three Idiots'.is_alnum()
false
%> '3 Idiots'.is_alnum()
false
%> '3'.is_alnum()
true
%> 'idiots'.is_alnum()
true
%> ''.is_alnum()
false

Returns boolean

is_number()

is_number()

Returns true if all the characters in the string are all digits and the string is not empty, otherwise returns false.

For example:

%> '123.5'.is_number()
false
%> '1970'.is_number()
true
%> '1980s'.is_number()
false

is_lower()

is_lower() -> boolean

Returns true if at least one character in the string is cased, all cased characters are lower cased and the string is not empty. Otherwise, it returns false.

For example:

%> 'all'.is_lower()
true
%> 'all...123'.is_lower()
true
%> 'All...123'.is_lower()
false
%> ''.is_lower()
false

Returns boolean

is_upper()

is_upper() -> boolean

Returns true if at least one character in the string is cased, all cased characters are upper cased and the string is not empty. Otherwise, it returns false.

For example:

%> 'ALL'.is_upper()
true
%> 'ALL...123'.is_upper()
true
%> 'All...123'.is_upper()
false
%> ''.is_upper()
false

Returns boolean

is_space()

is_space() -> boolean

Returns true if there are only whitespace characters in the string and the string is not empty. Otherwise, it returns empty.

For example:

%> '.     '.is_space()
false
%> '\r\n'.is_space()
true
%> '\t  '.is_space()
true

Returns boolean

ord()

ord() -> number

Returns the Unicode code point of the string, which must be exactly one character long.

%> 'A'.ord()
65
%> 'AB'.ord()
Unhandled Error: ord() must be called on a single character, got AB
StackTrace:
  <repl>:1 -> @.script()

Returns number

Raises Error if the string is not exactly one character long.

trim()

trim(chars: ?string) -> string

Returns a copy of the string with characters stripped from both ends.

With no argument, whitespace is stripped: space, tab (\t), line feed (\n), vertical tab, form feed and carriage return (\r). Other Unicode spaces, such as a no-break space, are kept.

Given chars, every character in it is stripped instead, in any order and any number of times, until a character not in chars is reached at each end. chars is a set of characters, not a prefix or suffix: 'xyax'.trim('xy') is 'a'. An empty chars strips nothing.

The string itself is never changed. A string with nothing to strip comes back as an equal copy.

For example:

%> '  example  '.trim()
'example'
%> '\t example \r\n'.trim()
'example'
%> '  example  '.trim('e')
'  example  '
%> 'example'.trim('e')
'xampl'
%> '--==example==--'.trim('-=')
'example'

Parameters

  • chars (?string) — The characters to strip (Default = whitespace).

Returns string

ltrim()

ltrim(chars: ?string) -> string

Returns a copy of the string with characters stripped from its start only. The characters stripped are chosen exactly as they are for trim(): whitespace when chars is not given, or every character of chars when it is.

For example:

%> '  example  '.ltrim()
'example  '
%> 'example'.ltrim('e')
'xample'
%> '0012'.ltrim('0')
'12'

Parameters

  • chars (?string) — The characters to strip (Default = whitespace).

Returns string

rtrim()

rtrim(chars: ?string) -> string

Returns a copy of the string with characters stripped from its end only. The characters stripped are chosen exactly as they are for trim(): whitespace when chars is not given, or every character of chars when it is.

For example:

%> '  example  '.rtrim()
'  example'
%> 'example'.rtrim('e')
'exampl'
%> 'line\r\n'.rtrim('\r\n')
'line'

Parameters

  • chars (?string) — The characters to strip (Default = whitespace).

Returns string

join()

join(string: string) -> string

Returns a string which is a concatenation of the items in the iterable using the string as the separator. If the iterable contains just one item or the string is empty, the original element is returned. If the iterable contains non-string items, the items are converted to their string representation before joining.

Bytes are the only non supported iterables.

For example:

%> ','.join(['ok', 1, true])
'ok,1,true'
%> '--'.join('name')
'n--a--m--e'
%> ','.join('a')
'a'

Parameters

  • string (string) — The string to join the items in the iterable.

Returns string

split()

split(delimiter: string) -> list

Returns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found.

If the delimiter is an empty string, the resultant list will contain the individual characters of the string in the order in which they appear in the original string. Consecutive delimiters are not grouped together and are deemed to delimit empty strings. Splitting an empty string with a specified separator returns an empty list.

This method has full UTF-8 support.

For example:

%> 'name'.split('')
[n, a, m, e]
%> '1<>2<>3'.split('<>')
[1, , 2, , 3]
%> '1,2,3'.split(',')
[1, 2, 3]
%> ''.split(',')
[]
%> '地点'.split('')
[地, 点]
%> 'who is in the garden'.split('/\s/')
[who, is, in, the, garden]

Parameters

  • delimiter (string) — The delimiter to use the split the string.

Returns list

index_of()

index_of(str: string, start_index: ?number) -> number

Returns the index position of the first occurrence of the string str in the string string. If the str cannot be found anywhere in string, it returns -1. If the start_index parameter is given, it will start scanning from the given index.

For example:

%> 'hello, world'.index_of(' ')
6
%> 'hello, world'.index_of('e')
1
%> 'hello, world'.index_of('q')
-1
%> 'hello, world'.index_of('o')
4
%> 'hello, world'.index_of('o', 5)  # next index of `o` starting from index 5.
8

Parameters

  • str (string) — The string to search for.
  • start_index (?number) — The index to start the search from.

Returns number

last_index_of()

last_index_of(str: string, end_index: ?number) -> number

Returns the index position of the last occurrence of the string str in the string string, searching from the end. If str cannot be found anywhere in string, it returns -1.

If the end_index parameter is given, only a match that begins at or before that index counts. That is the same thing index_of()’s own second parameter bounds, so for any index n, index_of(str, n) and last_index_of(str, n) are the first and last matches of the two halves n splits the string into.

An empty str returns -1, matching index_of().

For example:

%> 'hello, world'.last_index_of('o')
8
%> 'hello, world'.last_index_of('l')
10
%> 'hello, world'.last_index_of('q')
-1
%> 'hello, world'.last_index_of('o', 7)  # last `o` starting at or before index 7.
4

Splitting a path on its final separator is the usual reason to reach for it:

%> var path = 'a/b/c'
%> path.last_index_of('/')
3
%> path[path.last_index_of('/') + 1, path.length()]
'c'

Parameters

  • str (string) — The string to search for.
  • end_index (?number) — The highest index a match may start at.

Returns number

starts_with()

starts_with(str: string) -> boolean

Returns true if the string begins with the string or character specified in str, otherwise it returns false.

For example:

%> 'hello, world'.starts_with('hello')
true
%> 'hello, world'.starts_with('hellios')
false

Parameters

  • str (string) — The string to search for.

Returns boolean

ends_with()

ends_with(str: string) -> boolean

Returns true if the string ends with the string or character specified in str, otherwise it returns false.

For example:

%> 'gumtree'.ends_with('tree')
true
%> 'gumtree'.ends_with('mree')
false

Parameters

  • str (string) — The string to search for.

Returns boolean

count()

count(str: string) -> number

Returns the number of non-overlapping occurrences of the substring str in the string.

For those coming from Python who may consider this method similar to Python’s own, this method differs in that it does not allow specifying a start and end region for the operation. Zuri considers this unnecessary as the same can be accomplished by slicing the string.

For example:

%> 'Hallelujah'.count('l')
3
%> 'ding dong'.count('ng')
2
%> 'ding dong'[2,7].count('ng') # setting region to search for counts - 'ng do'
1

Parameters

  • str (string) — The string to search for.

Returns number

to_number()

to_number(base) -> number

Returns the first numeric value contained in the string if any exists or 0 if the string contains no numeric value. Floating numbers that have the same value as their integer counterparts will return the integer value.

For example:

%> '123.0 hell'.to_number()
123
%> '427 and 12'.to_number()
427
%> '96.3 of 31'.to_number()
96.3
%> 'error'.to_number()
0

Parameters

  • base (number) — The base the digits are in, from 2 to 36. Defaults to 10. A fractional part is only read in base 10, since no other base spells one.

Returns number

Raises RangeError if base is outside 2 to 36.

to_bigint()

to_bigint(base) -> bigint

Returns the integer value of the string as a bigint, or 0n if the string does not spell one.

This is to_number() for integers too large to be a number. A number is exact only up to 2^53; past that, digits are lost, and an id or a BIGINT UNSIGNED read from a database routinely runs past it. Every digit survives here however long the run.

%> '9007199254740993'.to_bigint()
9007199254740993n
%> '9007199254740993'.to_number()   # rounded down by one
9007199254740992
%> '-42'.to_bigint()
-42n
%> 'ff'.to_bigint(16)
255n
%> 'row 427 of 12'.to_bigint()
427n
%> 'error'.to_bigint()
0n

The number is found exactly as to_number() finds it: the first one written in the string, with any text around it ignored. No fractional part is read, since this produces an integer, so '12.5'.to_bigint() is 12n.

Parameters

  • base (number) — The base the digits are in, from 2 to 36. Defaults to 10.

Returns bigint

Raises RangeError if base is outside 2 to 36.

to_list()

to_list() -> list

Returns a list whose elements consists of every character contained in the string in order of appearance. Characters that repeat in the string will have different entries in the same index as they appear in the string.

For example:

%> 'Zuri'.to_list()
[Z, u, r, i]
%> 'Plantation'.to_list()
[P, l, a, n, t, a, t, i, o, n]

Returns list

to_bytes()

to_bytes() -> bytes

Returns the content of the string as a stream of bytes.

The Zuri REPL may truncate long bytes data when printing to console/terminal.

For example:

%> 'Zuri'.to_bytes()
(42 6c 61 64 65)
%> 'Plantation'.to_bytes()
(50 6c 61 6e 74 61 74 69 6f 6e)

Returns bytes

lpad()

lpad(width: number, fill: ?string) -> string

Returns the string left justified in a string of length width. Padding is done using the specified character fill if given of a space (' ') if a fill is not specified. The original string is returned if width is less than string.length().

For example:

%> 'cat'.lpad(5)
'  cat'
%> 'cat'.lpad(5, '-')
'--cat'
%> 'cat'.lpad(2, '-')
'cat'

Parameters

  • width (number) — The length of the string after padding.
  • fill (?string) — The character to use for padding.

Returns string

rpad()

rpad(width: number, fill: ?string) -> string

Returns the string right justified in a string of length width. Padding is done using the specified character fill if given of a space (' ') if a fill is not specified. The original string is returned if width is less than string.length().

For example:

%> 'Hmm'.rpad(6)
'Hmm   '
%> 'Hmm'.rpad(6, '.')
'Hmm...'
%> 'Hmm'.rpad(3, '.')
'Hmm'

Parameters

  • width (number) — The length of the string after padding.
  • fill (?string) — The character to use for padding.

Returns string

match()

match(str: string) -> boolean|dictionary

If the string str is a regular string, this method returns true if the string contains a substring str. Otherwise, it returns false.

If the string str contains a valid regular expression (we’ll get to that shortly below), it returns false if a match for the regex str cannot be found in the string. Otherwise, it returns a dictionary of the first match: the whole match under key 0, each capture group under its number, and a named group under its name as well. A group that takes no part in the match is nil. Groups that share a name, under the J modifier, give the name to the one that took part.

If the offset argument is specified, it becomes the offset in the string at which to start matching.

For example:

%> 'gorilla'.match('go')      # regular string match
true
%> 'gorilla'.match('gox')     # regular string non-match
false
%> 'gorilla'.match('/?gox/')  # regular expression match
{0: go}
%> 'gorilla'.match('/gox\d/') # regular expression non-match
false
%> '2024-01'.match('/(?<year>\d+)-(\d+)/')
{0: 2024-01, 1: 2024, 2: 01, year: 2024}

Parameters

  • str (string) — The string to match.

Returns boolean|dictionary

matches()

matches(reg: string) -> dictionary

Returns a dictionary containing every match of the given regular expression reg in the source string. If no match is found, an empty dictionary is returned.

If the offset argument is specified, it becomes the offset in the string at which to start matching.

For example:

%> '123 dollars'.matches('/[a-z]+|\d+/')
{0: [123, dollars]}
%> 'who is in the garden'.matches('/\w+/')
{0: [who, is, in, the, garden]}

Parameters

  • reg (string) — The regular expression to match.

Returns dictionary

replace()

replace(str: string, replacement: string, use_regex: ?bool) -> string

Returns a copy of the string with all occurrences or matches of str replaced by the replacement string.

In the replacement string, if str is a regular expression, then capture groups can be referenced using the syntax $index. Taking as an example, capture group 0 contains the entire match and can be used in the replacement string as $0.

To escape the $ sign in the replacement string, use the double backslashes (\\).

For example:

%> 'lady friend'.replace('d', 'z')  # non-regex
'lazy frienz'
%> 'John is 26 years old'.replace('/(\d+)/', '1$1') # regex example
'John is 126 years old'
%> 'John is 26 years old'.replace('/(\d+)/', '1\\$2')
'John is 1$2 years old'

Parameters

  • str (string) — The string to match.
  • replacement (string) — The replacement string.
  • use_regex (?bool) — Whether to use the regular expression or the string string as the match string (default = true).

Returns string

Note: When the third parameter use_regex is set to false, str will never be treated as a regular expression even if it contains a valid regular expression.

replace_with()

replace_with(regex: string, callback: function) -> string

Returns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred.

The callback function is defined as follows:

def replacer(match, p1, p2, /* …, */ pN, offset, string) {
  return replacement
}

The arguments to the function are as follows:

  • match: The matched substring. (Corresponds to $0.)

  • p1, p2, …, pN: The nth string found by a capture group (including named capturing groups) corresponds to $1, $2, etc. For example, if the pattern is /(\a+)(\b+)/, then p1 is the match for \a+, and p2 is the match for \b+. If the group is part of a disjunction (e.g. "abc".replace_with('/(a)|(b)/', replacer)), the unmatched alternative will be nil.

  • offset: The offset of the matched substring within the whole string being examined. For example, if the whole string was 'abcd', and the matched substring was 'bc', then this argument will be 1.

  • string: The whole string being examined.

The exact number of arguments depends on how many capture groups are contained in the regex.

For example:

%> echo 'name'.replace_with('/m/', @(match, offset) {
..   return match + '-'
.. })
'nam-e'

Below is another example that uses a capture group:

%> var text = 'all is well'
%> 
%> echo text.replace_with('/([a-z]+)/', @(match, val) {
..   if val == 'is' return 'is not'
..   return 'will be'
.. })
'will be is not will be'

Parameters

  • regex (string) — The regular expression to match.
  • callback (function) — The callback function to invoke for each match.

Returns string

ascii()

ascii() -> string

Reinterprets the string as a raw byte view: each byte of its UTF-8 encoding becomes its own character (a codepoint between 0 and 255, i.e. a Latin-1-style one-byte-per-character mapping), rather than the decoded sequence of Unicode characters that length(), each(), and indexing otherwise operate on.

The result is still a valid string (every codepoint between 0 and 255 is valid UTF-8), so it can be used anywhere a normal string can. It just no longer round-trips back through the original multi-byte characters if the string had any, and its length() now reports the original BYTE count of the string rather than its original CHARACTER count.

This is meant for the rare case where code needs to walk a string byte-for-byte instead of character-by-character, e.g. one that originated from a byte stream where the bytes were never meant to be decoded as Unicode at all.

%> 'café'.length()
4
%> 'café'.ascii().length()
5

Returns string

case_fold()

case_fold() -> string

Returns a copy of the string case-folded for case-insensitive comparison, using full Unicode case folding rather than plain lowercasing. This matters for characters whose fold is not just their lowercase form: for example, the German ß folds to ss.

Two strings that are considered equal ignoring case will always produce identical output from case_fold(), which makes it the correct method to use for case-insensitive comparisons; lower() is not a substitute for it.

%> 'HELLO World'.case_fold()
'hello world'
%> 'Straße'.case_fold()
'strasse'

Returns string

compare()

compare(other: string) -> number

Compares the string with another string.

Parameters

  • other (string) — The other string to compare with.

Returns number — A negative number if the string is less than the other string. - Zero if the strings are equal. - A positive number if the string is greater than the other string.

Raises Error if the other string is not a string.

is_empty()

is_empty() -> boolean

Returns true if the string is empty, false otherwise.

Returns boolean

contains()

contains(str: string) -> boolean

Returns true if the string contains the specified substring, false otherwise.

Parameters

  • str (string) — The substring to search for.

Returns boolean

lines()

lines() -> list

Returns the lines of the string as an list as it would be if split on newline characters.

Returns list

each_line()

each_line(callback: function) -> void

Iterates over each line of the string, calling the provided callback function with the line and its index.

Parameters

  • callback (function) — A function that takes two arguments: the line and its index.

Returns void

Raises Error if the callback is not a function.

each()

each(callback: function) -> void

Iterates over each character of the string, calling the provided callback function with the character and its index.

Parameters

  • callback (function) — A function that takes two arguments: the character and its index.

Returns void

Raises Error if the callback is not a function.

capitalize()

capitalize() -> string

Returns a new string with the first character capitalized and the rest in lowercase.

Returns string

title()

title() -> string

Returns a new string with each word capitalized.

Returns string

to_string()

to_string() -> string

Returns the string itself.

Returns string

Number Methods

Every method on the built-in number type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
to_string()stringReturns the string representation of the number.
to_bool()booleanConverts the number to a boolean, by the same rule if uses.
to_bigint()bigintConverts the number to a bigint, the counterpart to bigint.to_number().
abs()numberReturns the absolute value of the number.
chr()stringReturns the Unicode character whose code point is equal to the number.
bin()stringConverts the number to its binary string representation.
hex()stringConverts the number to its hexadecimal string representation.
oct()stringConverts the number to its octal string representation.
int()numberTruncates the number down to its integer part, discarding anything after the decimal point.
max(other: number)numberReturns the larger of the number and other.
min(other: number)numberReturns the smaller of the number and other.
factorial()numberReturns the factorial of the number, i.e.
sin()numberReturns the sine of the number, taken to be in radians.
cos()numberReturns the cosine of the number, taken to be in radians.
tan()numberReturns the tangent of the number, taken to be in radians.
sinh()numberReturns the hyperbolic sine of the number.
cosh()numberReturns the hyperbolic cosine of the number.
tanh()numberReturns the hyperbolic tangent of the number.
asin()numberReturns the arcsine (inverse sine) of the number, in radians.
acos()numberReturns the arccosine (inverse cosine) of the number, in radians.
atan()numberReturns the arctangent (inverse tangent) of the number, in radians.
atan2(x: number)numberReturns the four-quadrant arctangent of the number and x, in radians.
asinh()numberReturns the inverse hyperbolic sine of the number.
acosh()numberReturns the inverse hyperbolic cosine of the number.
atanh()numberReturns the inverse hyperbolic tangent of the number.
exp()numberReturns e (Euler’s number) raised to the power of the number.
expm1()numberReturns e raised to the power of the number, minus 1.
log()numberReturns the natural logarithm (base e) of the number.
log2()numberReturns the base-2 logarithm of the number.
log10()numberReturns the base-10 logarithm of the number.
log1p()numberReturns the natural logarithm of 1 plus the number.
cbrt()numberReturns the cube root of the number.
sqrt()numberReturns the square root of the number.
sign()numberReturns the sign of the number: 1 if it is positive, -1 if it is negative, and 0 (with its own original sign preserved) if it is zero.
ceil()numberReturns the smallest whole number greater than or equal to the number.
round()numberRounds the number to the nearest whole number.
floor()numberReturns the largest whole number less than or equal to the number.
is_nan()booleanReturns true if the number is NaN (not a number, e.g.
is_inf()booleanReturns true if the number is positive or negative infinity, false otherwise.
is_finite()booleanReturns true if the number is neither infinite nor NaN, false otherwise.
trunc()numberTruncates the number towards zero, discarding anything after the decimal point.
fraction()numberReturns the digits after the number’s decimal point, read as a whole number rather than a fraction.
fixed(n)Returns the number rounded to n decimal places, with a half rounding away from zero the same way round() does.

to_string()

to_string() -> string

Returns the string representation of the number.

%> 5.to_string()
'5'

Returns string

to_bool()

to_bool() -> boolean

Converts the number to a boolean, by the same rule if uses. Zero (either sign) and NaN are false, and every other number, negative ones included, is true.

%> 5.to_bool()
true
%> 0.to_bool()
false
%> (-5).to_bool()
true
%> (0 / 0).to_bool()
false

Returns boolean

to_bigint()

to_bigint() -> bigint

Converts the number to a bigint, the counterpart to bigint.to_number().

%> 12345.to_bigint()
12345n
%> 2.to_bigint() ** 100.to_bigint()
1267650600228229401496703205376n

Returns bigint

Raises RangeError if the number is not an exact integer.

Note: Only an exact integer has a bigint form, so a fractional number, Infinity and NaN all raise rather than being rounded or clamped. Numbers above 2^53 are already imprecise as doubles, so converting one yields the exact integer the double holds, not the decimal literal it was written as.

abs()

abs() -> number

Returns the absolute value of the number.

%> (-5).abs()
5
%> 5.abs()
5

Returns number

chr()

chr() -> string

Returns the Unicode character whose code point is equal to the number.

%> 65.chr()
'A'

Returns string

bin()

bin() -> string

Converts the number to its binary string representation. The number is truncated to an integer first.

%> 10.bin()
'1010'

Returns string

Note: A negative number always returns '0'; there is no signed or two’s-complement form.

hex()

hex() -> string

Converts the number to its hexadecimal string representation. The number is truncated to an integer first.

%> 255.hex()
'ff'

Returns string

Note: A negative number always returns '0'; there is no signed or two’s-complement form.

oct()

oct() -> string

Converts the number to its octal string representation. The number is truncated to an integer first.

%> 8.oct()
'10'

Returns string

Note: A negative number always returns '0'; there is no signed or two’s-complement form.

int()

int() -> number

Truncates the number down to its integer part, discarding anything after the decimal point. Unlike floor(), this rounds towards zero rather than towards negative infinity, so the result for a negative number differs from floor().

%> 3.9.int()
3
%> (-3.9).int()
-3

Returns number

max()

max(other: number) -> number

Returns the larger of the number and other.

A NaN never wins: when one of the two is NaN, the other is returned, and only two NaNs give NaN. -0 counts as smaller than 0, so (-0.0).max(0) is 0 whichever side the zeros are on.

%> 5.max(9)
9
%> 9.max(5)
9
%> (0 / 0).max(3)
3
%> (-0.0).max(0)
0

Parameters

  • other (number) — The number to compare against.

Returns number

min()

min(other: number) -> number

Returns the smaller of the number and other.

A NaN never wins: when one of the two is NaN, the other is returned, and only two NaNs give NaN. -0 counts as smaller than 0, so 0.min(-0.0) is -0 whichever side the zeros are on.

%> 5.min(9)
5
%> 9.min(5)
5
%> (0 / 0).min(3)
3
%> 0.min(-0.0)
-0

Parameters

  • other (number) — The number to compare against.

Returns number

factorial()

factorial() -> number

Returns the factorial of the number, i.e. the product of every positive integer less than or equal to it. 0.factorial() is 1, matching the standard mathematical definition.

%> 5.factorial()
120
%> 0.factorial()
1

Returns number

Raises Error if the number is negative or not a whole number.

sin()

sin() -> number

Returns the sine of the number, taken to be in radians.

Returns number

cos()

cos() -> number

Returns the cosine of the number, taken to be in radians.

Returns number

tan()

tan() -> number

Returns the tangent of the number, taken to be in radians.

Returns number

sinh()

sinh() -> number

Returns the hyperbolic sine of the number.

Returns number

cosh()

cosh() -> number

Returns the hyperbolic cosine of the number.

Returns number

tanh()

tanh() -> number

Returns the hyperbolic tangent of the number.

Returns number

asin()

asin() -> number

Returns the arcsine (inverse sine) of the number, in radians.

Returns number

Note: Only defined for a receiver between -1 and 1 inclusive; outside that range, this returns NaN rather than raising an error.

acos()

acos() -> number

Returns the arccosine (inverse cosine) of the number, in radians.

Returns number

Note: Only defined for a receiver between -1 and 1 inclusive; outside that range, this returns NaN rather than raising an error.

atan()

atan() -> number

Returns the arctangent (inverse tangent) of the number, in radians.

Returns number

atan2()

atan2(x: number) -> number

Returns the four-quadrant arctangent of the number and x, in radians. The receiver is treated as the y-coordinate and x as the x-coordinate, matching the conventional atan2(y, x) signature: y.atan2(x).

%> 1.0.atan2(1.0)
0.7853981633974483

Parameters

  • x (number) — The x-coordinate.

Returns number

asinh()

asinh() -> number

Returns the inverse hyperbolic sine of the number.

Returns number

acosh()

acosh() -> number

Returns the inverse hyperbolic cosine of the number.

Returns number

Note: Only defined for a receiver greater than or equal to 1; below that, this returns NaN rather than raising an error.

atanh()

atanh() -> number

Returns the inverse hyperbolic tangent of the number.

Returns number

Note: Only defined for a receiver between -1 and 1 exclusive; outside that range, this returns NaN rather than raising an error.

exp()

exp() -> number

Returns e (Euler’s number) raised to the power of the number.

Returns number

expm1()

expm1() -> number

Returns e raised to the power of the number, minus 1. For a number close to zero, this is more numerically accurate than computing n.exp() - 1 directly.

Returns number

log()

log() -> number

Returns the natural logarithm (base e) of the number.

%> 1.0.log()
0

Returns number

log2()

log2() -> number

Returns the base-2 logarithm of the number.

%> 8.0.log2()
3

Returns number

log10()

log10() -> number

Returns the base-10 logarithm of the number.

%> 100.0.log10()
2

Returns number

log1p()

log1p() -> number

Returns the natural logarithm of 1 plus the number. For a number close to zero, this is more numerically accurate than computing (1 + n).log() directly.

Returns number

cbrt()

cbrt() -> number

Returns the cube root of the number.

%> 27.0.cbrt()
3

Returns number

sqrt()

sqrt() -> number

Returns the square root of the number.

%> 16.sqrt()
4

Returns number

Note: For a negative number this returns NaN rather than raising an error; there is no bigint-style promotion into complex numbers. Check is_nan() on the result, or the sign of the receiver beforehand, if that distinction matters to the caller.

sign()

sign() -> number

Returns the sign of the number: 1 if it is positive, -1 if it is negative, and 0 (with its own original sign preserved) if it is zero.

%> 7.sign()
1
%> (-7).sign()
-1
%> 0.sign()
0

Returns number

ceil()

ceil() -> number

Returns the smallest whole number greater than or equal to the number.

%> 3.14159.ceil()
4

Returns number

round()

round() -> number

Rounds the number to the nearest whole number. A value exactly halfway between two whole numbers rounds away from zero.

%> 3.14159.round()
3
%> 3.6.round()
4

Returns number

floor()

floor() -> number

Returns the largest whole number less than or equal to the number.

%> 3.14159.floor()
3

Returns number

is_nan()

is_nan() -> boolean

Returns true if the number is NaN (not a number, e.g. the result of 0/0), false otherwise.

Returns boolean

is_inf()

is_inf() -> boolean

Returns true if the number is positive or negative infinity, false otherwise.

Returns boolean

is_finite()

is_finite() -> boolean

Returns true if the number is neither infinite nor NaN, false otherwise.

Returns boolean

trunc()

trunc() -> number

Truncates the number towards zero, discarding anything after the decimal point. For values that fit in a 64-bit integer this matches int(); unlike int(), trunc() stays a floating-point result rather than going through an integer cast, so it does not overflow for numbers larger than a 64-bit integer can hold.

%> (-3.9).trunc()
-3

Returns number

fraction()

fraction() -> number

Returns the digits after the number’s decimal point, read as a whole number rather than a fraction. Note that this is NOT the same as (n - n.int()): 1.92.fraction() is 92, not 0.92.

%> 1.92.fraction()
92
%> 1.5.fraction()
5
%> 5.fraction()
0

Returns number

fixed()

fixed(n)

Returns the number rounded to n decimal places, with a half rounding away from zero the same way round() does.

A number already shorter than n places is returned unchanged, and so are NaN and the infinities. Beyond 17 places an f64 has no digits left to round, so a larger n behaves as 17.

%> 1.554576852757686786786.fixed(9)
1.554576853
%> 1.554576852757686786786.fixed(8)
1.55457685
%> 1.554576852757686786786.fixed(1)
1.6
%> 1.554576852757686786786.fixed(0)
2
%> (-2.5).fixed(0)
-3

Bigint Methods

Every method on the built-in bigint type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
to_string(radix)stringReturns the decimal digits of the bigint, with a leading - when it is negative and no trailing n.
to_number()numberConverts the bigint to a number.
to_bool()booleanConverts the bigint to a boolean, by the same rule if uses.
to_bytes(order)bytesReturns the two’s-complement byte representation, which carries the sign and so round-trips back to the same value.
bin()stringReturns the base-2 digits, equivalent to to_string(2).
hex()stringReturns the base-16 digits in lowercase, equivalent to to_string(16).
oct()stringReturns the base-8 digits, equivalent to to_string(8).
abs()bigintReturns the absolute value.
sign()numberReturns the sign as a plain number: 1 when positive, -1 when negative and 0 for zero.
max(other)bigintReturns the larger of the two bigints.
min(other)bigintReturns the smaller of the two bigints.
pow(exponent)bigintRaises the bigint to exponent, the method form of **.
sqrt()bigintReturns the integer square root, truncated towards zero, so 145n.sqrt() is 12n rather than 12.04....
cbrt()bigintReturns the integer cube root, truncated towards zero.
nth_root(n)bigintReturns the integer nth root, truncated towards zero.
gcd(other)bigintReturns the greatest common divisor of the two bigints.
lcm(other)bigintReturns the least common multiple of the two bigints.
modpow(exponent, modulus)bigintReturns (self ** exponent) % modulus without ever building the full power, which is what makes it usable for the huge exponents cryptography needs.
modinv(modulus)bigint|nilReturns the modular multiplicative inverse: the x solving self * x == 1 (mod modulus).
bits()numberReturns how many bits the magnitude occupies, ignoring the sign.
bit(index)booleanReturns whether the bit at index is set, counting from the least significant bit at index 0.
set_bit(index, value)bigintReturns a new bigint with the bit at index set or cleared.
trailing_zeros()number|nilReturns the count of least-significant zero bits, which is the largest power of two dividing the bigint.
is_zero()booleanReturns whether the bigint is zero.
is_even()booleanReturns whether the bigint is even.
is_odd()booleanReturns whether the bigint is odd.

to_string()

to_string(radix) -> string

Returns the decimal digits of the bigint, with a leading - when it is negative and no trailing n. Pass radix to render in another base instead, using lowercase letters for digit values above nine.

%> 255n.to_string()
'255'
%> 255n.to_string(16)
'ff'
%> (-255n).to_string(16)
'-ff'

Parameters

  • radix (number) — base to render in, from 2 to 36. Defaults to 10.

Returns string

Raises RangeError if radix is outside 2 to 36.

Note: The n that echo and string interpolation show is part of the repr, not the conversion; to_string() never includes it.

to_number()

to_number() -> number

Converts the bigint to a number.

%> 6n.to_number()
6

Returns number

Note: This is lossy for anything past 2^53: the result is the nearest double, and a value past the double range becomes inf or -inf rather than wrapping or reading as zero. Check bits() beforehand when that matters.

to_bool()

to_bool() -> boolean

Converts the bigint to a boolean, by the same rule if uses. 0n is false, and every other bigint, negative ones included, is true.

%> 5n.to_bool()
true
%> 0n.to_bool()
false
%> (-5n).to_bool()
true

Returns boolean

to_bytes()

to_bytes(order) -> bytes

Returns the two’s-complement byte representation, which carries the sign and so round-trips back to the same value. The result is the shortest byte string that can hold it, and is never empty: zero is a single 0x00 byte.

%> 258n.to_bytes()
(01 02)
%> 258n.to_bytes('little')
(02 01)
%> (-1n).to_bytes()
(ff)

Parameters

  • order (string) — 'big' or 'little'. Defaults to 'big'.

Returns bytes

Raises RangeError if order is neither 'big' nor 'little'.

bin()

bin() -> string

Returns the base-2 digits, equivalent to to_string(2).

%> 255n.bin()
'11111111'

Returns string

Note: This is the sign-and-magnitude form, so a negative bigint comes back with a leading - rather than as two’s complement. Use to_bytes() for the two’s-complement view.

hex()

hex() -> string

Returns the base-16 digits in lowercase, equivalent to to_string(16).

%> 255n.hex()
'ff'

Returns string

oct()

oct() -> string

Returns the base-8 digits, equivalent to to_string(8).

%> 255n.oct()
'377'

Returns string

abs()

abs() -> bigint

Returns the absolute value.

%> (-5n).abs()
5n

Returns bigint

sign()

sign() -> number

Returns the sign as a plain number: 1 when positive, -1 when negative and 0 for zero.

%> (-9n).sign()
-1
%> 0n.sign()
0

Returns number

max()

max(other) -> bigint

Returns the larger of the two bigints.

%> 3n.max(7n)
7n

Parameters

  • other (bigint)

Returns bigint

Raises TypeError if other is not a bigint.

min()

min(other) -> bigint

Returns the smaller of the two bigints.

%> 3n.min(7n)
3n

Parameters

  • other (bigint)

Returns bigint

Raises TypeError if other is not a bigint.

pow()

pow(exponent) -> bigint

Raises the bigint to exponent, the method form of **.

%> 2n.pow(100)
1267650600228229401496703205376n

Parameters

  • exponent (number|bigint) — a integer from 0 to 2^32 - 1. Negative exponents have no integral answer and are rejected rather than truncated to zero.

Returns bigint

Raises RangeError if exponent is negative, fractional or too large.

sqrt()

sqrt() -> bigint

Returns the integer square root, truncated towards zero, so 145n.sqrt() is 12n rather than 12.04....

%> 144n.sqrt()
12n
%> 145n.sqrt()
12n

Returns bigint

Raises RangeError if the bigint is negative.

cbrt()

cbrt() -> bigint

Returns the integer cube root, truncated towards zero. Negatives are fine here, unlike sqrt().

%> (-27n).cbrt()
-3n

Returns bigint

nth_root()

nth_root(n) -> bigint

Returns the integer nth root, truncated towards zero.

%> 1000000n.nth_root(3)
100n

Parameters

  • n (number|bigint) — a integer from 1 to 2^32 - 1.

Returns bigint

Raises RangeError if n is zero, negative, fractional or too large, or if n is even and the bigint is negative.

gcd()

gcd(other) -> bigint

Returns the greatest common divisor of the two bigints. The result is always non-negative regardless of either sign, and 0n.gcd(0n) is 0n.

%> 48n.gcd(18n)
6n

Parameters

  • other (bigint)

Returns bigint

Raises TypeError if other is not a bigint.

lcm()

lcm(other) -> bigint

Returns the least common multiple of the two bigints. The result is always non-negative, and is 0n when either side is zero.

%> 48n.lcm(18n)
144n

Parameters

  • other (bigint)

Returns bigint

Raises TypeError if other is not a bigint.

modpow()

modpow(exponent, modulus) -> bigint

Returns (self ** exponent) % modulus without ever building the full power, which is what makes it usable for the huge exponents cryptography needs.

%> 4n.modpow(13n, 497n)
445n

Parameters

  • exponent (bigint)
  • modulus (bigint) — must not be zero.

Returns bigint

Raises TypeError if either argument is not a bigint.

Raises RangeError if modulus is zero, or if exponent is negative and no modular inverse exists.

Note: The remainder is floored rather than truncated, so the result carries the sign of modulus, not of the receiver. A negative exponent is allowed only when the receiver is invertible modulo modulus.

modinv()

modinv(modulus) -> bigint|nil

Returns the modular multiplicative inverse: the x solving self * x == 1 (mod modulus).

%> 3n.modinv(11n)
4n
%> 4n.modinv(8n)
nil

Parameters

  • modulus (bigint) — must not be zero.

Returns bigint|nil

Raises TypeError if modulus is not a bigint.

Raises RangeError if modulus is zero.

Note: Returns nil rather than raising when the receiver and modulus are not coprime, since having no inverse is an ordinary answer and not a caller mistake. The result carries the sign of modulus.

bits()

bits() -> number

Returns how many bits the magnitude occupies, ignoring the sign. Zero occupies none.

%> 255n.bits()
8
%> 0n.bits()
0

Returns number

bit()

bit(index) -> boolean

Returns whether the bit at index is set, counting from the least significant bit at index 0.

%> 5n.bit(0)
true
%> 5n.bit(1)
false

Parameters

  • index (number) — a non-negative integer.

Returns boolean

Raises RangeError if index is negative or fractional.

Note: The bigint is read as two’s complement, so a negative receiver reports true for every index above its magnitude rather than running out of bits.

set_bit()

set_bit(index, value) -> bigint

Returns a new bigint with the bit at index set or cleared. The receiver is left untouched.

%> 5n.set_bit(1, true)
7n

Parameters

  • index (number) — a non-negative integer.
  • value (boolean)

Returns bigint

Raises RangeError if index is negative or fractional.

trailing_zeros()

trailing_zeros() -> number|nil

Returns the count of least-significant zero bits, which is the largest power of two dividing the bigint.

%> 40n.trailing_zeros()
3
%> 0n.trailing_zeros()
nil

Returns number|nil

Note: Returns nil for zero, which has no largest such power and would otherwise have to report an arbitrary number.

is_zero()

is_zero() -> boolean

Returns whether the bigint is zero.

%> 0n.is_zero()
true

Returns boolean

is_even()

is_even() -> boolean

Returns whether the bigint is even. Zero is even.

%> 4n.is_even()
true

Returns boolean

is_odd()

is_odd() -> boolean

Returns whether the bigint is odd.

%> 5n.is_odd()
true

Returns boolean

Boolean Methods

Every method on the built-in bool type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
to_string()stringReturns the string representation of the boolean.

to_string()

to_string() -> string

Returns the string representation of the boolean.

%> true.to_string()
'true'
%> false.to_string()
'false'

Returns string

List Methods

Every method on the built-in list type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
length()numberReturns the number of items in the list.
append(value)listAdds the given value x to the end of the list.
clear()Removes all items from the list.
clone()listReturns a new list containing all items from the list.
count(value)numberReturns the number of times item x occurs in the list.
extend(list: list)listUpdates the content of the list by appending all the contents of list x to the end of the original list in exact order.
index_of(value, start_index: ?int)numberReturns the zero-based index of the first occurrence of the value x in the list starting from the given start_index or -1 if the list does not contain the value x.
last_index_of(value, end_index: ?int)numberReturns the zero-based index of the last occurrence of the value x in the list, searching from the end, or -1 if the list does not contain the value x.
insert(value, index: int)listInserts the item x into the list at the specified index.
pop()anyRemoves the last item in a list and returns the value of that item.
shift(count: ?int)anyRemoved the specified count of items from the beginning of the list and returns it.
remove_at(index: int)anyRemoves the item at the specified index in the list and returns it.
remove(value)anyRemoves the first occurrence of item x from the list.
reverse()listReturns a new list containing the items in the original list in reverse order.
sort(comparator: ?function)listSorts the items in the list in-place and returns the sorted list.
contains(value)booleanReturns true if the list contains the item x or false otherwise.
delete(start: int, end: int)numberDeletes a range of items from the list starting from the start to the end limit and returns the number of items removed.
first()anyReturns the first item in the list or nil if the list is empty.
last()anyReturns the last item in the list or nil if the list is empty.
is_empty()booleanReturns true if the list is empty or false otherwise.
take(n: int)listReturns a new list containing the first n items in the list or a new copy of the list if n greater than or equals to the list.length().
get(index: int)anyReturns the value at the specified index in the list.
compact()listReturns a new list containing the items in the original list but with all nil values removed.
unique()listReturns a new list containing the unique values from the original list.
zip(...lists: list)listReturns a list that contains the items in the original list merged with corresponding items from the individual arguments.
zip_from(list: list)listThe same as list.zip() except that instead of accepting an arbitrary list or arguments, it accepts a single list that should contain other lists.
to_dict()dictReturns a number indexed dictionary representing the list.
each(callback: function)Iterates over each element in the list, calling the provided callback function with the current element as an argument.
map(callback: function)listCreates a new list populated with the results of calling a provided function on every element in the calling list.
filter(callback: function)listCreates a new list with all elements that pass the test implemented by the provided function.
reduce(callback: function, initial)anyApplies a function against an accumulator and each element in the list (from left to right) to reduce it to a single value and returns the accumulated result of the callback function.
some(callback: function)booleanTests whether at least one element in the list passes the test implemented by the provided function.
every(callback: function)booleanTests whether all elements in the list pass the test implemented by the provided function.
find(callback: function)anyReturns the value of the first element in the list that satisfies the provided testing function.
find_index(callback: function)numberReturns the index of the first element in the list that satisfies the provided testing function.
find_last(callback: function)anyReturns the value of the last element in the list that satisfies the provided testing function.
find_last_index(callback: function)numberReturns the index of the last element in the list that satisfies the provided testing function.
find_all(callback: function)listReturns a new list containing all elements of the calling list that satisfy the provided testing function.
partition(callback: function)listReturns an list containing two lists: the first with elements that satisfy the provided testing function, and the second with elements that do not satisfy the testing function.
to_string()stringReturns the string representation of the list.

length()

length() -> number

Returns the number of items in the list.

For example:

%> ['A', 'B', 'C'].length()
3

Returns number

append()

append(value) -> list

Adds the given value x to the end of the list.

For example:

%> var a = [1,2,3]
%> a.append(4)
%> a
[1, 2, 3, 4]

Parameters

  • value (any)

Returns list

clear()

clear()

Removes all items from the list.

For example:

%> var a = [1,2,3,4,5]
%> a
[1, 2, 3, 4, 5]
%> a.clear()
%> a
[]

clone()

clone() -> list

Returns a new list containing all items from the list. The new list is a shallow copy of the original list. This is equivalent to list[,].

For example:

%> var a = [1, 2, 3]
%> var b = a.clone()
%> a.append(4)
%> a
[1, 2, 3, 4]
%> b
[1, 2, 3]

Returns list

count()

count(value) -> number

Returns the number of times item x occurs in the list.

For example:

%> [1, 2, 1, 3, 2, 1, 1].count(1)
4

Parameters

  • value (any)

Returns number

extend()

extend(list: list) -> list

Updates the content of the list by appending all the contents of list x to the end of the original list in exact order. This is equivalent to list + x.

For example:

%> var a = [1, 2, 3]
%> var b = [4, 5, 6]
%> a.extend(b)
%> a
[1, 2, 3, 4, 5, 6]
%> b
[4, 5, 6]

Parameters

  • list (list)

Returns list

index_of()

index_of(value, start_index: ?int) -> number

Returns the zero-based index of the first occurrence of the value x in the list starting from the given start_index or -1 if the list does not contain the value x.

For example:

%> [1,2].index_of(3)
-1
%> [4,5,6,5].index_of(5)
1
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].index_of('a')
0
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].index_of('a', 1)
3

Parameters

  • value (any)
  • start_index (?int)

Returns number

last_index_of()

last_index_of(value, end_index: ?int) -> number

Returns the zero-based index of the last occurrence of the value x in the list, searching from the end, or -1 if the list does not contain the value x.

If end_index is given, only a match at or before that index counts. That is the same position index_of()’s own second parameter bounds, so for any index n, index_of(x, n) and last_index_of(x, n) are the first and last matches of the two halves n splits the list into.

Values are compared the way index_of() compares them, by value rather than by identity, so two separate dictionaries holding the same entries match each other.

For example:

%> [1,2].last_index_of(3)
-1
%> [4,5,6,5].last_index_of(5)
3
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].last_index_of('a')
5
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].last_index_of('a', 4)
3

Parameters

  • value (any)
  • end_index (?int)

Returns number

insert()

insert(value, index: int) -> list

Inserts the item x into the list at the specified index. By specifying an index of zero (list.insert(x, 0)), one can prepend the list and list.insert(x, list.length()) is equivalent to list.append(x). If the index specified is greater than list.length(), the list will be padded with nil up till the index preceding the specified index.

For example:

%> var a = [1,2,3]
%> a.insert(4, 0)
%> a
[4, 1, 2, 3]
%> a.insert(5, a.length())
%> a
[4, 1, 2, 3, 5]
%> a.insert(6, 3)
%> a
[4, 1, 2, 6, 3, 5]
%> a.insert(7, 11)
%> a
[4, 1, 2, 6, 3, 5, nil, nil, nil, nil, nil, 7]

Parameters

  • value (any)
  • index (int)

Returns list

pop()

pop() -> any

Removes the last item in a list and returns the value of that item.

For example:

%> var a = [4, 5, 6]
%> a.pop()
6
%> a
[4, 5]

Returns any

shift()

shift(count: ?int) -> any

Removed the specified count of items from the beginning of the list and returns it. If count is not specified, count defaults to 1. If one item is shifted, the method returns that item. If more than one item is shifted, the method returns a list containing the shifted items.

The square brackets ([]) around the count: number in the method definition indicates that the parameter is optional and does not mean you have to type the square brackets.

If the number of items required to be shifted exceeds the size of the list, the list is cleared and nil is returned.

For example:

%> var a = [9, 8, 7, 6, 5, 4, 3, 2, 1, 0]
%> a.shift()
9
%> a
[8, 7, 6, 5, 4, 3, 2, 1, 0]
%> a.shift(3)
[8, 7, 6]
%> a
[5, 4, 3, 2, 1, 0]
%> a.shift(10)
%> a
[]

Parameters

  • count (?int)

Returns any

remove_at()

remove_at(index: int) -> any

Removes the item at the specified index in the list and returns it. If the index is less than 0 or greater than list.length() - 1, an Error is raised.

For example:

%> var a = [1, 2, 3, 4, 5]
%> a.remove_at(3)
4
%> a
[1, 2, 3, 5]
%> a.remove_at(6)
Unhandled Error: list index 6 out of range at remove_at()
  StackTrace:
    <repl>:1 -> @.script()
%> a.remove_at(-1)
Unhandled Error: list index -1 out of range at remove_at()
  StackTrace:
    <repl>:1 -> @.script()

Parameters

  • index (int)

Returns any

remove()

remove(value) -> any

Removes the first occurrence of item x from the list.

For example:

%> var a = ['Kirk', 'Tasha', 'Emily', 'Kirk']
%> a.remove('Kirk')
%> a
[Tasha, Emily, Kirk]

Notice that only the first occurrence of Kirk was removed.

Parameters

  • value (any)

Returns any

reverse()

reverse() -> list

Returns a new list containing the items in the original list in reverse order.

For example:

%> var a = ['apple', 'mango', 'banana', 'orange', 'peach']
%> a.reverse()
[peach, orange, banana, mango, apple]

Returns list

sort()

sort(comparator: ?function) -> list

Sorts the items in the list in-place and returns the sorted list. Sorting in Lists follows are strict set of precedence based on the object type. The order for sorting is as follows in ascending orders:

nil, boolean, numbers, strings, ranges, lists, dictionaries, file, bytes, functions, classes and modules.

When the corresponding items in the list are of the same type, they are sorted based on their respective values according to the type. For example, the number 5 is less than 8 and as such will appear first in the sort.

For example:

%> var a  = ['A', 5, false, nil, [21, 13, 46]]
%> a.sort()
%> a
[nil, false, 5, A, [13, 21, 46]]

Notice how the boolean value precedes the number and how the number in turn precedes the string and the strings in turn, precedes the list in the result. Also, note that the items of the inner list is sorted.

Sorting by something else

Pass a comparator to decide the order yourself. It is given two items and returns a negative number to put the first one first, a positive number to put the second one first, and zero to leave them as they are:

%> [3, 1, 2].sort(@(a, b) => b - a)
[3, 2, 1]
%> ['pear', 'fig', 'banana'].sort(@(a, b) => a.length() - b.length())
['fig', 'pear', 'banana']

The sort is stable, so items the comparator calls equal keep the order they were already in. That is what lets a list be sorted by one thing and then another to order by both:

people.sort(@(a, b) => a.name.compare(b.name))
people.sort(@(a, b) => a.age - b.age)

leaves people of the same age in name order.

Parameters

  • comparator (function)

Returns list

Note: A comparator sorts only the list it is given. The inner lists that sort() sorts on its own are left alone, since only the comparator knows what the order is meant to be.

Note: A comparator that contradicts itself produces some order rather than an error; there is no arrangement that satisfies it.

contains()

contains(value) -> boolean

Returns true if the list contains the item x or false otherwise.

For example:

%>  ['dog', 'cat', 'wolf', 'tiger'].contains('cat')
true
%>  ['dog', 'cat', 'wolf', 'tiger'].contains('giraffe')
false

Parameters

  • value (any)

Returns boolean

delete()

delete(start: int, end: int) -> number

Deletes a range of items from the list starting from the start to the end limit and returns the number of items removed. If the start and end are the same, this will be equivalent to list.remove_at(start).

For example:

%> var a = [1, 2, 3, 4, 5, 6, 7, 8, 9]
%> a.delete(3, 6)
4
%> a
[1, 2, 3, 8, 9]
%> a.delete(1,1)  # equal start and end
1
%> a
[1, 3, 8, 9]

Parameters

  • start (int)
  • end (int)

Returns number

first()

first() -> any

Returns the first item in the list or nil if the list is empty.

For example:

%> ['c', 'd', 'a', 'b'].first()
'c'

Returns any

last()

last() -> any

Returns the last item in the list or nil if the list is empty.

For example:

%> ['c', 'd', 'a', 'b'].last()
'b'

Returns any

is_empty()

is_empty() -> boolean

Returns true if the list is empty or false otherwise.

For example:

%> [1, 2].is_empty()
false
%> [].is_empty()
true

Returns boolean

take()

take(n: int) -> list

Returns a new list containing the first n items in the list or a new copy of the list if n greater than or equals to the list.length(). If n < 0, returns list.take(list.length() - n).

For example:

%> var a = [10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20]
%> a.take(4)
[10, 11, 12, 13]
%> a.take(11) # taking more than the size of the list
[10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20]
%> a.take(-5)   # taking n < 0
[10, 11, 12, 13, 14, 15]

Parameters

  • n (int)

Returns list

get()

get(index: int) -> any

Returns the value at the specified index in the list. If index is outside the boundary of the list indexes (0..(list.length() - 1)), an Error is thrown. This method is equivalent to list[index].

For example:

%> [13, 14, 15, 16].get(1)
14
%> [13, 14, 15, 16].get(6)
Unhandled Error: list index 6 out of range at get()
  StackTrace:
    <repl>:1 -> @.script()

Parameters

  • index (int)

Returns any

compact()

compact() -> list

Returns a new list containing the items in the original list but with all nil values removed.

For example:

%> [21, nil, 14, 'age', nil, nil, [], 11].compact()
[21, 14, age, [], 11]

Returns list

unique()

unique() -> list

Returns a new list containing the unique values from the original list.

For example:

%> [1, 1, 3, 5].unique()
[1, 3, 5]

Returns list

zip()

zip(...lists: list) -> list

Returns a list that contains the items in the original list merged with corresponding items from the individual arguments. This generates a list of length equal to the length of the original argument.

If the size of any of the arguments is less than the size of the original list, it’s corresponding entry will be nil.

For example:

%> var a = [4, 5, 6]
%> var b = [7, 8, 9]
%> [1, 2, 3].zip(a, b)
[[1, 4, 7], [2, 5, 8], [3, 6, 9]]
%> [1, 2].zip(a, b)
[[1, 4, 7], [2, 5, 8]]
%> a.zip([1, 2], [8])
[[4, 1, 8], [5, 2, nil], [6, nil, nil]]
%> [1, 2].zip([3])
[[1, 3], [2, nil]]
%> [1].zip([10, 11], [12, 13, 14])
[[1, 10, 12]]
%> [[1, 2], [3]].zip(a, b)
[[[1, 2], 4, 7], [[3], 5, 8]]

Parameters

  • lists (...list)

Returns list

zip_from()

zip_from(list: list) -> list

The same as list.zip() except that instead of accepting an arbitrary list or arguments, it accepts a single list that should contain other lists.

For example:

%> [1, 2].zip_from([[3, 4]])
[[1, 3], [2, 4]]

Parameters

  • list (list)

Returns list

to_dict()

to_dict() -> dict

Returns a number indexed dictionary representing the list.

For example:

%> ['English', 'French', 'Spanish'].to_dict()
{0: English, 1: French, 2: Spanish}

Returns dict

each()

each(callback: function)

Iterates over each element in the list, calling the provided callback function with the current element as an argument.

Example:

['A', 'B', 'C'].each(@(r) {
  echo r
})

# Output: A B C

The each method does not return a new list; it simply executes the callback for each element. If you want to create a new list based on the original, consider using the map method instead.

Parameters

  • callback (function) — The function to execute for each element in the list.

Raises Error if the callback is not a function.

map()

map(callback: function) -> list

Creates a new list populated with the results of calling a provided function on every element in the calling list.

Example:

echo [1, 2, 3].map(@(x) {
  return x * 2
})

# Output: [2, 4, 6]

Parameters

  • callback (function) — The function to execute on each element in the list. It receives the current element and its index as arguments.

Returns list

Raises Error if the callback is not a function.

filter()

filter(callback: function) -> list

Creates a new list with all elements that pass the test implemented by the provided function.

It returns a new list with the elements that pass the test. If no elements pass the test, an empty list will be returned.

Example:

echo [1, 2, 3].filter(@(x) {
  return x % 2 == 0
})

# Output: [2]

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns list

Raises Error if the callback is not a function.

reduce()

reduce(callback: function, initial) -> any

Applies a function against an accumulator and each element in the list (from left to right) to reduce it to a single value and returns the accumulated result of the callback function.

Example:

echo [1, 2, 3].reduce(@(acc, x) {
  return acc + x
})

# Output: 6

Parameters

  • callback (function) — The function to execute on each element in the list. It receives the current element and its index as arguments.
  • initial — The initial value to use as the accumulator. If no initial value is provided, the first element of the list will be used as the initial accumulator, and the iteration will start from the second element.

Returns any

Raises Error if the callback is not a function.

some()

some(callback: function) -> boolean

Tests whether at least one element in the list passes the test implemented by the provided function.

Example:

echo [1, 2, 3].some(@(x) {
  return x % 2 == 0
})

# Output: true

The some method returns true if the callback function returns a truthy value for at least one element in the list. If the callback function returns a falsy value for all elements, some will return false. If the list is empty, some will return false by default.

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns boolean

Raises Error if the callback is not a function.

every()

every(callback: function) -> boolean

Tests whether all elements in the list pass the test implemented by the provided function.

Example:

echo [1, 2, 3].every(@(x) {
  return x > 0
})

# Output: true

The every method returns true if the callback function returns a truthy value for every element in the list. If the callback function returns a falsy value for any element, every will return false. If the list is empty, every will return true by default.

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns boolean

Raises Error if the callback is not a function.

find()

find(callback: function) -> any

Returns the value of the first element in the list that satisfies the provided testing function. If no elements satisfy the testing function, find returns nil.

Example:

echo [1, 2, 3].find(@(x) {
  return x % 2 == 0
})

# Output: 2

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns any — The first element that satisfies it, or nil.

Raises Error if the callback is not a function.

find_index()

find_index(callback: function) -> number

Returns the index of the first element in the list that satisfies the provided testing function. If no elements satisfy the testing function, find_index returns -1.

Example:

echo [1, 2, 3].find_index(@(x) {
  return x % 2 == 0
})

# Output: 1

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns number — The index of the first element that satisfies it, or -1.

Raises Error if the callback is not a function.

find_last()

find_last(callback: function) -> any

Returns the value of the last element in the list that satisfies the provided testing function. If no elements satisfy the testing function, find_last returns nil.

Example:

echo [1, 2, 3].find_last(@(x) {
  return x % 2 == 0
})

# Output: 2

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns any — The last element that satisfies it, or nil.

Raises Error if the callback is not a function.

find_last_index()

find_last_index(callback: function) -> number

Returns the index of the last element in the list that satisfies the provided testing function. If no elements satisfy the testing function, find_last_index returns -1.

Example:

echo [1, 2, 3].find_last_index(@(x) {
  return x % 2 == 0
})

# Output: 1

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns number — The index of the last element that satisfies it, or -1.

Raises Error if the callback is not a function.

find_all()

find_all(callback: function) -> list

Returns a new list containing all elements of the calling list that satisfy the provided testing function.

Example:

echo [1, 2, 3].find_all(@(x) {
  return x % 2 == 0
})

# Output: [2]

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns list

Raises Error if the callback is not a function.

partition()

partition(callback: function) -> list

Returns an list containing two lists: the first with elements that satisfy the provided testing function, and the second with elements that do not satisfy the testing function.

Example:

echo [1, 2, 3].partition(@(x) {
  return x % 2 == 0
})

# Output: [[2], [1, 3]]

Parameters

  • callback (function) — The function to test each element of the list. It receives the current element and its index as arguments.

Returns list

Raises Error if the callback is not a function.

to_string()

to_string() -> string

Returns the string representation of the list.

%> [1, 'two', 3].to_string()
'[1, two, 3]'

Returns string

Dictionary Methods

Every method on the built-in dict type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
length()numberReturns the length of the dictionary.
add(key: string, value)Adds a new key-value pair to the dictionary with the given key and value.
set(key: string, value)Sets the value of the given key to the given value in the dictionary.
clear()Clears the content of the dictionary.
clone()dictReturns a new dictionary which is a deep copy of the original dictionary.
compact()dictReturns a new dictionary that contains every key-value pair in the original dictionary except for keys whose associated value is nil.
contains(key: string)booleanReturns true if any of the keys in the dictionary is equal to x, false otherwise.
extend(dict: dict)Adds all key-value pairs in dictionary x to the original dictionary.
get(key: string, default_value)any|nilReturns the value of the given key in the dictionary.
keys()listReturns a list containing the keys in the dictionary.
values()listReturns a list containing the value of all keys in the dictionary.
remove(key)any|nilRemoves a given key and it’s corresponding value from the dictionary and returns the value of the key.
is_empty()booleanReturns true if the dictionary is empty, otherwise returns false.
find_key(value)string|nilReturns the key whose value is equal to x in the dictionary or nil if no key has the value x.
to_list()listReturns a list that contains a list of key and a list of values from the dictionary.
each(callback: function)voidIterates over each key-value pair in the dictionary, calling the provided callback function with the value and key as arguments.
filter(callback: function)dictCreates a new dictionary containing only the key-value pairs for which the provided callback function returns true.
some(callback: function)booleanTests whether at least one key-value pair in the dictionary passes the test implemented by the provided callback function.
every(callback: function)booleanTests whether all key-value pairs in the dictionary pass the test implemented by the provided callback function.
reduce(callback: function, initial)anyReduces the dictionary to a single value by iteratively combining each key-value pair using the provided callback function.
to_string()stringReturns the string representation of the dictionary.

length()

length() -> number

Returns the length of the dictionary. The length of a Zuri dictionary is equal to the number of keys it contains. i.e. dict.length() == dict.keys().length().

For example:

%> {name: 'Zuri', version: 1}.length()
2

Returns number

add()

add(key: string, value)

Adds a new key-value pair to the dictionary with the given key and value.

For example:

%> var dict = {}
%> dict.add('name', 'Zuri')
%> dict
{name: Zuri}

Parameters

  • key (string)
  • value (any)

set()

set(key: string, value)

Sets the value of the given key to the given value in the dictionary. If there is no existing entry for the key in the dictionary, a new entry will be added.

For example:

%> dict.set('name', 'New Zuri')
%> dict
{name: New Zuri}
%> dict.set('version', 1)
%> dict
{name: New Zuri, version: 1}

@note: dict.set(x, y) is equivalent to the following Zuri code.

%> if dict.contains(x) {
..   dict[x] = 1
.. } else {
..   dict.add(x, 1)
.. }

Parameters

  • key (string)
  • value (any)

clear()

clear()

Clears the content of the dictionary.

For example:

%> var a = {name: 'Zuri'}
%> a
{name: Zuri}
%> a.clear()
%> a
{}

clone()

clone() -> dict

Returns a new dictionary which is a deep copy of the original dictionary.

For example:

%> var new_dict = dict.clone()
%> new_dict
{name: New Zuri, version: 1}

Returns dict

compact()

compact() -> dict

Returns a new dictionary that contains every key-value pair in the original dictionary except for keys whose associated value is nil.

For example:

%> var dict2 = {name: 'James', age: 20, address: nil, country: nil}
%> dict2.compact()
{name: James, age: 20}

Returns dict

contains()

contains(key: string) -> boolean

Returns true if any of the keys in the dictionary is equal to x, false otherwise.

For example:

%> dict2.contains('name')
true
%> dict2.contains('street')
false

Parameters

  • key (string)

Returns boolean

extend()

extend(dict: dict)

Adds all key-value pairs in dictionary x to the original dictionary.

For example:

%> var dict = {name: 'Zuri'}
%> dict.extend({version: 1})
%> dict
{name: Zuri, version: 1}

Parameters

  • dict (dict)

get()

get(key: string, default_value) -> any|nil

Returns the value of the given key in the dictionary. If the given key is not defined in the dictionary and the default value is given, the default value will be returned. Otherwise, nil is returned.

For example:

%> dict.get('version')   # value exists
1
%> dict.get('age')   # value does not exist
%> dict.get('age', 6)   # value does not exist, but default is given
6
%> dict.get('version', 1.1)   # value exists and default is given
1

Parameters

  • key (string)
  • default_value (any|nil)

Returns any|nil

keys()

keys() -> list

Returns a list containing the keys in the dictionary.

For example:

%> dict.keys()
[name, version]

Returns list

values()

values() -> list

Returns a list containing the value of all keys in the dictionary.

For example:

%> dict.values()
[Zuri, 1]

Returns list

remove()

remove(key) -> any|nil

Removes a given key and it’s corresponding value from the dictionary and returns the value of the key.

For example:

%> dict = {username: 'james', email: 'a@b.c', active: true}
%> dict.remove('active')
true
%> dict
{username: james, email: a@b.c}

Parameters

  • key (string)

Returns any|nil

is_empty()

is_empty() -> boolean

Returns true if the dictionary is empty, otherwise returns false.

For example:

%> dict.is_empty()
false
%> {}.is_empty()
true

Returns boolean

find_key()

find_key(value) -> string|nil

Returns the key whose value is equal to x in the dictionary or nil if no key has the value x.

For example:

%> dict.find_key('james')
'username'
%> dict.find_key('camel')

Parameters

  • value (any)

Returns string|nil

to_list()

to_list() -> list

Returns a list that contains a list of key and a list of values from the dictionary.

For example:

%> var dict = {username: 'james', email: 'a@b.c'}
%> dict.to_list()
[[username, email], [james, a@b.c]]

Returns list

each()

each(callback: function) -> void

Iterates over each key-value pair in the dictionary, calling the provided callback function with the value and key as arguments.

Example:

var myDict = {a: 1, b: 2, c: 3}
myDict.each(@(value, key) {
  echo '${key}: ${value}'
})

# Output:
# a: 1
# b: 2
# c: 3

Parameters

  • callback (function) — The function to call for each key-value pair.

Returns void

Raises Error If the callback is not a function.

filter()

filter(callback: function) -> dict

Creates a new dictionary containing only the key-value pairs for which the provided callback function returns true. The callback function is called with the value and key as arguments.

Example:

var myDict = {a: 1, b: 2, c: 3}
var filteredDict = myDict.filter(@(value, key) {
  return value > 1
})
echo filteredDict

# Output: {'b': 2, 'c': 3}

Parameters

  • callback (function) — The function to test each key-value pair. It should return true to keep the pair, or false to exclude it.

Returns dict

Raises Error If the callback is not a function.

some()

some(callback: function) -> boolean

Tests whether at least one key-value pair in the dictionary passes the test implemented by the provided callback function. The callback function is called with the value and key as arguments. The method returns true if the callback returns true for any key-value pair, otherwise it returns false.

Example:

var myDict = {a: 1, b: 2, c: 3}
var hasGreaterThanTwo = myDict.some(@(value, key) {
  return value > 2
})
echo hasGreaterThanTwo

# Output: true

Parameters

  • callback (function) — The function to test each key-value pair. It should return true to indicate a passing pair, or false to indicate a failing pair.

Returns boolean

Raises Error If the callback is not a function.

every()

every(callback: function) -> boolean

Tests whether all key-value pairs in the dictionary pass the test implemented by the provided callback function. The callback function is called with the value and key as arguments. The method returns true if the callback returns true for every key-value pair, otherwise it returns false.

Example:

var myDict = {a: 1, b: 2, c: 3}
var allGreaterThanZero = myDict.every(@(value, key) {
  return value > 0
})
echo allGreaterThanZero

# Output: true

Parameters

  • callback (function) — The function to test each key-value pair. It should return true to indicate a passing pair, or false to indicate a failing pair.

Returns boolean

Raises Error If the callback is not a function.

reduce()

reduce(callback: function, initial) -> any

Reduces the dictionary to a single value by iteratively combining each key-value pair using the provided callback function. The callback function is called with the accumulator, value, key, and the dictionary itself as arguments. The method returns the final accumulated value after processing all key-value pairs in the dictionary.

Example:

var myDict = {a: 1, b: 2, c: 3}
var sum = myDict.reduce(@(accumulator, value, key) {
  return accumulator + value
}, 0)
echo sum

# Output: 6

Parameters

  • callback (function) — The function to execute on each key-value pair in the dictionary. It should return the updated accumulator value after processing the pair.
  • initial (any) — The initial value to use as the first argument to the first call of the callback function.

Returns any

Raises Error If the callback is not a function.

to_string()

to_string() -> string

Returns the string representation of the dictionary.

%> {a: 1, b: 2}.to_string()
'{a: 1, b: 2}'

Returns string

Range Methods

Every method on the built-in range type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
lower()numberReturns the lower limit of the range.
upper()numberReturns the upper limit of the range.
range()numberReturns a number equal to the numbers between the range.
within(value: number)booleanReturns true if the given number falls somewhere within the or false otherwise.
step(size: int)rangeSets the step size of the range.
get_step()numberReturns the step size of the range.
loop(callback: function)voidIterates over each number in the range, calling the provided callback function with the number, its index.
to_list()listReturns the range as a list of its individual numbers, stepping from the lower limit to the upper limit (exclusive), or in reverse when the range descends.
to_string()stringReturns the string representation of the range.

lower()

lower() -> number

Returns the lower limit of the range.

For example:

%> (10..100).lower()
10

Returns number

upper()

upper() -> number

Returns the upper limit of the range.

For example:

%> (20..30).upper()
30

Returns number

range()

range() -> number

Returns a number equal to the numbers between the range.

For example:

%> (21..93).range()
72

The result of stays the same irrespective of the direction of the range. For example, swapping the upper and lower limit of our previous still returns the same result.

%> (21..93).range()
72

Returns number

within()

within(value: number) -> boolean

Returns true if the given number falls somewhere within the or false otherwise.

For example:

%> (93..21).within(103)
false
%> (93..21).within(57)
true

Parameters

  • value (number)

Returns boolean

step()

step(size: int) -> range

Sets the step size of the range.

For example:

%> var a = (10..100).step(20)
%> a
<range 10..100, step=20>
%> for i in a {
..   echo i
.. }
10
30
50
70
90

Parameters

  • size (int) — The step size of the range.

Returns range

get_step()

get_step() -> number

Returns the step size of the range.

Returns number

loop()

loop(callback: function) -> void

Iterates over each number in the range, calling the provided callback function with the number, its index.

Example:

var r = 0..5 # 0, 1, 2, 3, 4
r.loop(@(num, index) {
  echo 'Number at index ${index}: ${num}'
})

# Output:
# Number at index 0: 0
# Number at index 1: 1
# Number at index 2: 2
# Number at index 3: 3
# Number at index 4: 4

Parameters

  • callback (function) — A function that takes two arguments: the number, its index.

Returns void

Raises Error if the callback is not a function.

to_list()

to_list() -> list

Returns the range as a list of its individual numbers, stepping from the lower limit to the upper limit (exclusive), or in reverse when the range descends.

%> (1..5).to_list()
[1, 2, 3, 4]
%> (5..1).to_list()
[5, 4, 3, 2]

Returns list

to_string()

to_string() -> string

Returns the string representation of the range.

%> (1..5).to_string()
'1..5'

Returns string

Bytes Methods

Every method on the built-in bytes type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
length()numberReturns the number of bytes in the byte stream.
is_empty()booleanReturns true if the byte stream holds no bytes at all, and false otherwise.
append(n: int)bytesAdds an item to the top of a byte stream.
clone()bytesReturns a deep clone of the byte stream.
extend(n: bytes)bytesExtends the byte stream with the bytes from the given byte stream.
index_of(byte: int, start_index: ?number)numberReturns the index of the first occurrence of the given byte in the byte stream.
last_index_of(byte: int, end_index: ?number)numberReturns the index of the last occurrence of the given byte in the byte stream, searching from the end, or -1 when the byte is not there.
pop()numberRemoves the last item in a byte stream and returns it.
remove(index: number)bytesRemoves the item at the specified index in the byte stream and return the previous value at the specified index.
reverse()bytesReverses the items in the byte stream.
first()numberReturns the first item in the byte stream or nil if the byte stream is empty.
last()numberReturns the last item in the byte stream or nil if the byte stream is empty.
get(index: number)numberReturns the item at the specified index in the byte stream.
take(n: int)bytesReturns a new byte stream containing the first n items in the bytes or a new copy of the bytes if n greater than or equals to the bytes.length().
split(delimiter: bytes)listSplits the content of a byte stream based on the specified delimiter.
dispose()Due to the nature of byte stream and their use-case (especially streaming data), it is easy for the system memory to get filled up with data in the byte stream.
is_alpha()booleanReturns true if the byte stream only contains alpha characters, false otherwise.
is_alnum()booleanReturns true if the byte stream only contains alpha characters and numbers, false otherwise.
is_number()booleanReturns true if the byte stream only contains numbers, false otherwise.
is_lower(n)booleanReturns true if the byte stream only contains lower case characters, false otherwise.
is_upper()booleanReturns true if the byte stream only contains upper case characters, false otherwise.
is_space()booleanReturns true if the byte stream only contains space characters, false otherwise.
to_list()listReturns the byte stream as a list of bytes.
to_string()stringReturns the byte stream as a string, reading it as UTF-8.
each(callback: function)voidIterates over each byte of the bytes object, calling the provided callback function with the byte and its index.

length()

length() -> number

Returns the number of bytes in the byte stream.

%> bytes([25, 57]).length()
2

Returns number

is_empty()

is_empty() -> boolean

Returns true if the byte stream holds no bytes at all, and false otherwise. Equivalent to testing length() == 0, and unaffected by whatever the bytes happen to contain: a stream of zero bytes is not empty.

%> bytes(0).is_empty()
true
%> bytes(3).is_empty()
false
%> 'hi'.to_bytes().is_empty()
false

Returns boolean

append()

append(n: int) -> bytes

Adds an item to the top of a byte stream.

For example,

%> var a = bytes([0x40, 0x75])
%> a.append(0x16)
%> echo a
(40 75 16)

Parameters

  • n (int) — The byte to add.

Returns bytes

clone()

clone() -> bytes

Returns a deep clone of the byte stream.

For example,

%> bytes([19, 11]).clone()
(13 b)

Returns bytes

extend()

extend(n: bytes) -> bytes

Extends the byte stream with the bytes from the given byte stream.

For example,

%> var a = bytes([33, 91, 126])
%> var b = bytes([119, 42])
%> a
(21 5b 7e)
%> b
(77 2a)
%> a.extend(b)
(21 5b 7e 77 2a)
%> a
(21 5b 7e 77 2a)

Parameters

  • n (bytes) — The byte stream to extend with.

Returns bytes

Note: extend() is an in-place action so the original byte stream will be modified.

index_of()

index_of(byte: int, start_index: ?number) -> number

Returns the index of the first occurrence of the given byte in the byte stream.

%> bytes([25, 57, 25]).index_of(57)
1
%> bytes([25, 57, 25]).index_of(25, 1)
2

Parameters

  • byte (int) — The byte to search for.
  • start_index (?number) — The index to start the search from. Defaults to 0.

Returns number

last_index_of()

last_index_of(byte: int, end_index: ?number) -> number

Returns the index of the last occurrence of the given byte in the byte stream, searching from the end, or -1 when the byte is not there.

If end_index is given, only a match at or before that index counts. That is the same position index_of()’s own second parameter bounds, so for any index n, index_of(b, n) and last_index_of(b, n) are the first and last occurrences in the two halves n splits the stream into.

%> bytes([25, 57, 25]).last_index_of(25)
2
%> bytes([25, 57, 25]).last_index_of(25, 1)
0

Parameters

  • byte (int) — The byte to search for.
  • end_index (?number) — The highest index a match may sit at. Defaults to the last byte.

Returns number

pop()

pop() -> number

Removes the last item in a byte stream and returns it.

%> var a = bytes([79, 43, 9])
%> a.pop()
9
%> a
(4f 2b)

Returns number

remove()

remove(index: number) -> bytes

Removes the item at the specified index in the byte stream and return the previous value at the specified index.

%> var a = bytes([25, 57, 25])
%> a.remove(1)
57
%> a
(25 25)

Parameters

  • index (number) — The index to remove.

Returns bytes

reverse()

reverse() -> bytes

Reverses the items in the byte stream.

%> bytes([5, 4, 3, 2, 1]).reverse()
(1 2 3 4 5)

Returns bytes

first()

first() -> number

Returns the first item in the byte stream or nil if the byte stream is empty.

%> bytes([25, 57, 42]).first()
25

Returns number

last()

last() -> number

Returns the last item in the byte stream or nil if the byte stream is empty.

%> bytes([25, 57, 42]).last()
42

Returns number

get()

get(index: number) -> number

Returns the item at the specified index in the byte stream.

Parameters

  • index (number) — The index to get the item from.

Returns number

take()

take(n: int) -> bytes

Returns a new byte stream containing the first n items in the bytes or a new copy of the bytes if n greater than or equals to the bytes.length(). If n < 0, returns bytes.take(bytes.length() - n).

For example:

%> var a = bytes([10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20])
%> a.take(4)
(0a 0b 0c 0d)
%> a.take(11) # taking more than the size of the bytes
(0a 0b 0c 0d 0e 0f 10 11 12 13 14)
%> a.take(-5)   # taking n < 0
(0a 0b 0c 0d 0e 0f)

Parameters

  • n (int)

Returns bytes

split()

split(delimiter: bytes) -> list

Splits the content of a byte stream based on the specified delimiter.

For example,

%> bytes(0).split(bytes(0))
[]
%> echo 'test'.to_bytes().split(bytes(0))
[(74), (65), (73), (74)]

Parameters

  • delimiter (bytes) — The delimiter to split on.

Returns list

dispose()

dispose()

Due to the nature of byte stream and their use-case (especially streaming data), it is easy for the system memory to get filled up with data in the byte stream. The method allows users to reset a byte stream and empty it.

This method allows a fine-grained control on manual memory management of byte stream.

For example,

%> var a = bytes([13, 36])
%> a.dispose()
%> a
()

is_alpha()

is_alpha() -> boolean

Returns true if the byte stream only contains alpha characters, false otherwise.

%> bytes([65, 66, 67]).is_alpha()
true
%> bytes([65, 66, 67, 128]).is_alpha()
false

Returns boolean

is_alnum()

is_alnum() -> boolean

Returns true if the byte stream only contains alpha characters and numbers, false otherwise.

%> bytes([65, 66, 67, 48, 49, 50]).is_alnum()
true
%> bytes([65, 66, 67, 48, 49, 50, 8]).is_alnum()
false

Returns boolean

is_number()

is_number() -> boolean

Returns true if the byte stream only contains numbers, false otherwise.

%> bytes([48, 49, 50]).is_number()
true
%> bytes([48, 49, 50, 68]).is_number()
false

Returns boolean

is_lower()

is_lower(n) -> boolean

Returns true if the byte stream only contains lower case characters, false otherwise.

%> bytes([97, 98, 99]).is_lower()
true
%> bytes([97, 98, 99, 68]).is_lower()
false

Returns boolean

is_upper()

is_upper() -> boolean

Returns true if the byte stream only contains upper case characters, false otherwise.

%> bytes([65, 66, 67]).is_upper()
true
%> bytes([65, 66, 67, 98]).is_upper()
false

Returns boolean

is_space()

is_space() -> boolean

Returns true if the byte stream only contains space characters, false otherwise.

%> bytes([32, 32, 32]).is_space()
true
%> bytes([32, 32, 32, 68]).is_space()
false

Returns boolean

to_list()

to_list() -> list

Returns the byte stream as a list of bytes.

%> bytes([0x31, 0x55, 0xe9, 0x21]).to_list()
[49, 85, 233, 33]

Returns list

to_string()

to_string() -> string

Returns the byte stream as a string, reading it as UTF-8. Each sequence that is not valid UTF-8 reads as U+FFFD, the replacement character, so the string of such bytes does not encode back to the same bytes.

%> bytes([65, 66, 67, 68, 69]).to_string()
'ABCDE'

Returns string

each()

each(callback: function) -> void

Iterates over each byte of the bytes object, calling the provided callback function with the byte and its index.

Example:

var data = bytes([0x48, 0x65, 0x6C, 0x6C, 0x6F]) # "Hello" in bytes
data.each(def(byte, index) {
  echo 'Byte at index ${index}: ${byte}'
})

# Output:
# Byte at index 0: 72
# Byte at index 1: 101
# Byte at index 2: 108
# Byte at index 3: 108
# Byte at index 4: 111

Parameters

  • callback (function) — A function that takes two arguments: the byte and its index.

Returns void

Raises Error if the callback is not a function.

File Methods

Every method on the built-in file type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
exists()booleanReturns true if a file exists or false otherwise.
close()voidCloses the stream to an opened file.
open()voidOpens the stream to a file for the operation originally specified on the file object during creation.
read(length: ?int)string|bytesReads the content of an opened file up to the specified length and returns it as string or bytes if the file was opened in the binary mode.
gets(length: ?int)string|bytesSame as read(), but doesn’t open or close the file automatically.
write(data: string|bytes)string|bytesWrites a string or bytes to an opened file at the current insertion point.
puts(data: string|bytes)string|bytesSame as write(), but doesn’t open or close the file automatically.
number()intReturns the integer file descriptor number that is used by the underlying implementation to request I/O operations from the operating system.
is_tty()booleanReturns true if the file is connected to a TTY like device or false otherwise.
is_open()booleanReturns true if the file is open for reading or writing and false otherwise.
is_closed()booleanReturns true if the file is closed for reading or writing and false otherwise.
flush()voidFlushes the buffer held by a file.
stats()dictReturns the statistics or details of a file.
symlink()booleanCreates a symbolic link for the original file at the specified path.
delete()booleanDeletes a file.
rename(new_name: string)booleanRenames a file to to new_name.
path()stringReturns the path to the file.
abs_path()stringReturns the absolute path to the file.
copy(path: string)booleanCopies a file from the path specified in the original file to the given path.
truncate(length: ?number)booleanTruncates the entire file if length is not given or truncates the file such that only length number of bytes is left in it.
chmod(mode: int)booleanChanges the permission on the file to the one specified in the number given.
set_times(atime: number, mtime: number)booleanSets the last access time and last modified time of the file.
seek(offset: number, seek_type: int)booleanSets the position of a file reader or writer in a file.
tell()numberReturns the current position of the reader/writer in a file.
mode()stringReturns the mode in which the current file was opened.
name()stringReturns the name of the current file.
to_string()stringReturns the file handle as a string, naming its path and the mode it was opened in.

exists()

exists() -> boolean

Returns true if a file exists or false otherwise.

For example:

%> file('sample.txt').exists()
true

Returns boolean

close()

close() -> void

Closes the stream to an opened file. You’ll rarely ever need to call this method yourself in most use cases.

For example:

%> var f = file('sample.txt')
%> f.close()

Returns void

open()

open() -> void

Opens the stream to a file for the operation originally specified on the file object during creation. You may need to call this method after a call to read() if the length isn’t specified or write() if you wish to read or write again as the file will already be closed.

For example:

%> f.open()

Returns void

read()

read(length: ?int) -> string|bytes

Reads the content of an opened file up to the specified length and returns it as string or bytes if the file was opened in the binary mode. If the length is not specified, the file will be read to the end.

In text mode the bytes read must be valid UTF-8; anything else raises rather than being silently replaced. Open the file in a binary mode ('rb') to read arbitrary bytes instead. Note that io.stdin is already binary.

This method requires that the file be opened in the read mode (default mode) or a mode that supports reading. If you aren’t reading the full length of the file, you’ll need to call the close() method to free the file for further reading, otherwise, the close() method will be automatically called for you.

An example has been given above.

Parameters

  • length (?int)

Returns string|bytes

gets()

gets(length: ?int) -> string|bytes

Same as read(), but doesn’t open or close the file automatically.

Parameters

  • length (?int)

Returns string|bytes

write()

write(data: string|bytes) -> string|bytes

Writes a string or bytes to an opened file at the current insertion point. When the file is opened with the a mode enabled, write will always start from the end of the file. If the seek() method has been previously called, write will begin from the seeked position, otherwise it will start at the beginning of the file.

An example has been given above.

Parameters

  • data (string|bytes)

Returns string|bytes

puts()

puts(data: string|bytes) -> string|bytes

Same as write(), but doesn’t open or close the file automatically.

Parameters

  • data (string|bytes)

Returns string|bytes

number()

number() -> int

Returns the integer file descriptor number that is used by the underlying implementation to request I/O operations from the operating system. This can be very useful for low-level interfaces that uses or act as file descriptors.

For example:

%> file('sample.txt').number()
6

A standard stream reports the descriptor it is named by, 0, 1 or 2, rather than the private duplicate the runtime holds open for it, so number() is what tells io.stdout apart from io.stderr.

Returns int

Note: -1 on a platform with no file descriptors, and on any handle that is not currently open.

is_tty()

is_tty() -> boolean

Returns true if the file is connected to a TTY like device or false otherwise.

For example:

%> file('sample.txt').is_tty()
false
%> import io
%> io.stdout.is_tty()   # io.stdin is a file...
true

Returns boolean

is_open()

is_open() -> boolean

Returns true if the file is open for reading or writing and false otherwise.

@note: std files are always open.

For example:

%> file('sample.txt').is_open()
true

Returns boolean

is_closed()

is_closed() -> boolean

Returns true if the file is closed for reading or writing and false otherwise.

For example:

%> file('sample.txt').is_closed()
false

Returns boolean

flush()

flush() -> void

Flushes the buffer held by a file. This could be useful for writable files as file writes are buffered.

For example:

%> w.flush()

Returns void

stats()

stats() -> dict

Returns the statistics or details of a file.

For example:

%> file('sample.txt').stats()
{is_readable: true, is_writable: true, is_executable: false,
is_symbolic: false, size: 72, mode: 33188, dev: 16777230,
ino: 4865113, nlink: 1, uid: 501, gid: 20, mtime: 1631395239,
atime: 1631395271, ctime: 1631395239, blocks: 8, blksize: 4096}

Every key above is present on every platform, so reading size or mtime needs no check of which one you are on.

Returns dict

Note: Windows keeps a different set of facts about a file, and the ones it has no answer for read as 0: dev, ino, uid and gid, with nlink always 1. mode is assembled from the file’s type and its read-only attribute, so it carries the right file-type bits and either 0o444 or 0o666, widened by 0o111 for a directory or a name PATHEXT says the shell would run. ctime is the file’s creation time there, Windows having no equivalent of a Unix inode-change time.

symlink() -> boolean

Creates a symbolic link for the original file at the specified path.

For example:

%> file('sample.txt').symlink('sample2.txt')
true

Returns boolean

Note: Windows decides at creation time whether a link stands for a file or a directory, so the original is inspected first; one pointing at something that does not exist yet is made as a file link. Creating any symbolic link there is privileged, and fails unless the machine is in developer mode or the process is elevated.

delete()

delete() -> boolean

Deletes a file.

For example:

%> file('test-2.zu').delete()
true

Returns boolean

Note: If the file is opened by one or more processes or threads outside of the current process or thread, the file will not be deleted until the last process frees it.

Note: This method throws Error on failure.

rename()

rename(new_name: string) -> boolean

Renames a file to to new_name. The new name can be a full path in another location in which case the file will be moved.

For example:

%> file('sample copy.txt').rename('sample-2.txt')
true

Parameters

  • new_name (string)

Returns boolean

Note: The new name cannot be empty

Note: This method throws Error on failure.

path()

path() -> string

Returns the path to the file.

For example:

%> file('sample.txt').path()
'sample.txt'

Returns string

abs_path()

abs_path() -> string

Returns the absolute path to the file.

For example:

%> file('sample.txt').abs_path()
'C:\Users\username\zuri-docs\sample.txt'

Returns string

copy()

copy(path: string) -> boolean

Copies a file from the path specified in the original file to the given path.

For example:

%> file('./sample.txt').copy('samp.txt')
true

Parameters

  • new_name (string)

Returns boolean

truncate()

truncate(length: ?number) -> boolean

Truncates the entire file if length is not given or truncates the file such that only length number of bytes is left in it.

For example:

%> file('./samp.txt').truncate()
true

Parameters

  • length (?number)

Returns boolean

chmod()

chmod(mode: int) -> boolean

Changes the permission on the file to the one specified in the number given.

@note: The number is required to be an octal number. e.g. 0c755

For example:

%> file('sample.txt').chmod(0c755)
true

Parameters

  • mode (int)

Returns boolean

Note: Windows stores one read-only attribute where Unix stores nine permission bits, so the owner-write bit decides it and the rest are dropped: 0o755 and 0o700 are the same instruction there. A mode with no owner-write bit marks the file read-only.

set_times()

set_times(atime: number, mtime: number) -> boolean

Sets the last access time and last modified time of the file.

@note: Time is expected in UTC seconds
@note: set argument -1 to leave the current value.

For example:

%> file('sample.txt').set_times(time(), time())
true
%> file('sample.txt').stats()
{is_readable: true, is_writable: true, is_executable: true,
is_symbolic: false, size: 72, mode: 33261, dev: 16777230,
ino: 4865113, nlink: 1, uid: 501, gid: 20, mtime: 1631477099,
atime: 1631477100, ctime: 1631477099, blocks: 8, blksize: 4096}

Parameters

  • atime (number)
  • mtime (number)

Returns boolean

seek()

seek(offset: number, seek_type: int) -> boolean

Sets the position of a file reader or writer in a file. The position must be within the range of the file size. seek_type must be on of SEEK_SET, SEEK_CUR or SEEK_END from the io package.

For example:

%> f.seek(5, io.SEEK_SET)
true

Parameters

  • offset (number)
  • seek_type (int)

Returns boolean

tell()

tell() -> number

Returns the current position of the reader/writer in a file.

For example:

%> import io
%> var f = file('sample.txt')
%> f.seek(5, io.SEEK_SET)
true
%> f.tell()
5

Returns number

mode()

mode() -> string

Returns the mode in which the current file was opened.

For example:

%> file('sample.txt').mode()
'r'

Returns string

name()

name() -> string

Returns the name of the current file.

For example:

%> file('./sample.txt').name()
'sample.txt'

Returns string

to_string()

to_string() -> string

Returns the file handle as a string, naming its path and the mode it was opened in.

%> file('sample.txt', 'w').to_string()
'<file at sample.txt in mode w>'

Returns string

Function Methods

Every method on the built-in function type, with its signature, what it returns, and the cases where it does something other than the obvious thing.

MethodReturnsSummary
name()stringThe function’s declared name.
arity()numberHow many parameters the function declares, counting a variadic one as a single parameter.
is_variadic()boolWhether the last parameter is variadic (...name).
call(...args: list)anyCalls the function with the given arguments and returns its result.
apply(args: list)anyCalls the function with the arguments in a list, and returns its result.
to_string()stringThe function rendered for display, as <function NAME(ARITY)>, with a trailing ... on the arity when the function is variadic.

name()

name() -> string

The function’s declared name. An anonymous function is named @anon followed by a number, counted in the order the compiler met it; a bound method reports the method’s own name, not its class’s.

%> def named(a, b, ...c) {}
%> named.name()
'named'
%> print.name()
'print'

Returns string

arity()

arity() -> number

How many parameters the function declares, counting a variadic one as a single parameter.

%> def named(a, b, ...c) {}
%> named.arity()
3

A method read off an instance counts the instance as its first parameter, so a method declaring two parameters reports 3. A foreign function from ffi has no receiver and reports its C parameter list.

Returns number

is_variadic()

is_variadic() -> bool

Whether the last parameter is variadic (...name).

%> def named(a, b, ...c) {}
%> named.is_variadic()
true

Returns bool

call()

call(...args: list) -> any

Calls the function with the given arguments and returns its result. The same as calling it directly; useful when the function is held in a variable and the call site reads better spelled out.

%> def add(a, b) { return a + b }
%> add.call(2, 3)
5

Parameters

  • args (...any) — The arguments to call with.

Returns any — Whatever the function returns.

Raises anything the called function raises.

apply()

apply(args: list) -> any

Calls the function with the arguments in a list, and returns its result. call() takes them written out; this one takes them in a list.

%> def add(a, b) { return a + b }
%> add.apply([2, 3])
5
%> def collect(first, ...rest) { return [first, rest] }
%> collect.apply([1, 2, 3])
[1, [2, 3]]

A list shorter than the function’s arity leaves the remaining parameters nil, exactly as calling it directly with too few arguments does; a longer one overflows into a variadic parameter, or is discarded when there is none.

Parameters

  • args (list) — The arguments to call with, in order.

Returns any — Whatever the function returns.

Raises TypeError when args is not a list, and anything the called function raises.

to_string()

to_string() -> string

The function rendered for display, as <function NAME(ARITY)>, with a trailing ... on the arity when the function is variadic.

%> def named(a, b, ...c) {}
%> named.to_string()
'<function named(3...)>'

Returns string

Appendix F: The Standard Library Index

Every module in the standard library, what it is for, and where the book covers it. All of them are reachable with a bare import, with no package manager and no dependency to add.

Each module’s source is in libs/, and every public function in it carries a doc block stating its parameters, its defaults and its edge cases.

Data and Serialisation

ModuleWhat it is forBook
jsonencode, decode, read and write JSON13
yamlparse YAML, with anchors, tags and multi-document streams13
tomlparse and write TOML, and edit one without disturbing its layout13
csvread and write CSV, with dialect detection13
structpack and unpack binary layouts10
base64Base64 encode and decode10
convertbase, hex, binary, octal and unicode conversions13
compressdeflate, zlib, gzip, zstd, lz4, bzip2, brotli, tar, zip, checksums10

Text and Markup

ModuleWhat it is forBook
htmla WHATWG-conformant parser, a DOM and CSS selectors13
wiretemplating, with directives as HTML attributes14
urlparse, build and percent-encode URLs13
mimedetect a media type from a name or from content13
colorsANSI terminal colour, with graceful degradation13

Time

ModuleWhat it is forBook
datedates, times, formatting, parsing and IANA time zones13

Cryptography and Identity

ModuleWhat it is forBook
hashdigests and HMACs, plus PBKDF210
bcryptpassword hashing13
cryptoRSA signing and HKDF13
uuidUUID versions 1, 3, 4, 5, 6, 7 and 813
jwtsign, verify and decode JSON Web Tokens, with JWKS13

Structure and Validation

ModuleWhat it is forBook
validatea fluent schema builder13
typeschecked coercion between types13
setan ordered set with the usual algebra13
enumnamed constants from a list or a dictionary13
arraytyped fixed-width numeric arrays, Int8 through Double13

The System

ModuleWhat it is forBook
osprocesses, filesystem, paths, environment, signals9
enva .env file into the environment, and typed values back out19
iostandard streams, the terminal, and in-memory files9, 10
statthe S_IS* predicates over a file mode13
argsa command-line parser with subcommands and --help20
loglevelled, structured logging with pluggable transports13
isolateOS-thread concurrency, channels and broadcasts11
fficalling C and Rust libraries, callbacks, and linking static libraries26

Testing

ModuleWhat it is forBook
testsuites, matchers, test doubles, snapshots and reports23

Databases

ModuleWhat it is forBook
sqlone contract for every relational database, with SQLite, PostgreSQL and MySQL adapters17

Mail

ModuleWhat it is forBook
mailmessages, SMTP, IMAP and POP3, both server ends, and DKIM18

The Network

ModuleWhat it is forBook
netTCP, UDP, unix sockets, TLS, DTLS, addresses and polling12
httpan HTTP/1.1 and HTTP/2 client and server15
rpcJSON-RPC 2.0 over HTTP, WebSockets, sockets and streams, both ends28

Graphics

ModuleWhat it is forBook
imaginedecode, draw, filter and encode images16

Numbers

ModuleWhat it is forBook
maththe mathematical constants4

Metaprogramming

ModuleWhat it is forBook
zurilexing, parsing, compiling and runtime reflection21

Packages and Their Submodules

Most of the larger modules are packages: a directory whose index.zu re-exports what its parts make public. import http reaches almost all of http without naming a submodule. Import a submodule directly when you want only that part of it, or when a name would otherwise collide.

array

One module per element width, each exporting a single class.

SubmoduleClassElement
array.int8Int8Arraysigned 8-bit
array.uint8UInt8Arrayunsigned 8-bit
array.int16Int16Arraysigned 16-bit
array.uint16UInt16Arrayunsigned 16-bit
array.int32Int32Arraysigned 32-bit
array.uint32UInt32Arrayunsigned 32-bit
array.int64Int64Arraysigned 64-bit
array.uint64UInt64Arrayunsigned 64-bit
array.floatFloatArray32-bit float
array.doubleDoubleArray64-bit float

compress

SubmoduleWhat it is for
compress.deflateraw DEFLATE streams
compress.zlibDEFLATE with a zlib header
compress.gzipDEFLATE with a gzip header and trailer
compress.zstdZstandard, levels 1 to 22
compress.lz4LZ4, block and frame formats
compress.bzip2bzip2
compress.brotliBrotli
compress.tarreading and writing TAR archives
compress.zipreading and writing ZIP archives
compress.checksumCRC32, CRC32C, Adler-32

ffi

SubmoduleWhat it is for
ffi.errorsevery error the module raises, under FfiError
ffi.typesType, and the struct, union and enum builders
ffi.pointerPointer: native memory and access through it
ffi.libraryLibrary: a loaded shared library
ffi.declareDeclarations: C and Rust source read into types and signatures
ffi.callbackCallback: a Zuri function behind a C function pointer

html

SubmoduleWhat it is for
html.tokenizerthe WHATWG tokenizer: text in, tokens out
html.parsertree construction: tokens in, a document out
html.nodethe document tree and everything you can do to it
html.selectorfinding nodes with CSS selectors
html.serializewriting a document back out, minified or pretty
html.entitiesnamed character references, both directions
html.elementsthe element tables tree construction consults
html.namespacesthe five namespace URIs the parser deals in

http

SubmoduleWhat it is for
http.clientHttpClient, the request side
http.serverHttpServer, the listening side
http.workerserving across several isolates
http.routermatching a method and path to a handler
http.middlewareCORS, access logging, security headers, and the rest
http.requestthe request object handlers receive
http.responsethe response object handlers return
http.headersHeaders, with the field-name rules of RFC 9110
http.cookiesCookie and CookieJar
http.sessionserver-side sessions, and the stores that keep them
http.session.sqlkeeping sessions in a relational database
http.bodyreading and writing message bodies
http.multipartmultipart/form-data, including file uploads
http.filesserving files from disk, with ranges and caching
http.streamchunked and streaming transfers
http.sseserver-sent events
http.websocketthe WebSocket protocol
http.proxyforwarding requests to another server
http.negotiateparsing Accept-style headers
http.statusthe IANA status codes and their reason phrases
http.h1the HTTP/1.1 wire format
http.errorsthe module’s error hierarchy

imagine

SubmoduleWhat it is for
imagine.imageImage, the pixel buffer everything else operates on
imagine.canvasdrawing: lines, shapes, fills, text
imagine.colorColor, and conversion between colour spaces
imagine.filtersblur, sharpen, convolution, and the rest
imagine.fontloading and measuring fonts
imagine.strokefontthe built-in stroke font, with no file to load
imagine.formatsdecoding and encoding PNG, JPEG, GIF, WebP and more
imagine.animationmulti-frame images
imagine.constantsthe named constants the module understands
imagine.errorsthe module’s error hierarchy

io

SubmoduleWhat it is for
io.bytesioBytesIO, a file-shaped object backed by memory
io.ttyterminal control: raw mode, size, cursor

isolate

SubmoduleWhat it is for
isolate.channelbounded multi-producer, multi-consumer queues
isolate.broadcastone-to-many publish and subscribe
isolate.errorIsolateError

jwt

SubmoduleWhat it is for
jwt.coreencode(), decode(), sign(), verify()
jwt.signerSigner, a reusable configured signer
jwt.verifierVerifier, a reusable configured verifier
jwt.tokenthe Token object a complete decode returns
jwt.jwksresolving a signing key from a JSON Web Key Set
jwt.codecalgorithm identifiers and the low-level encoding
jwt.errorsthe module’s error hierarchy

log

SubmoduleWhat it is for
log.loggerthe module-level info(), warn(), error() and friends
log.levelthe LogLevel enum and the default level
log.transportTransport, the base class every sink extends
log.consoleConsoleTransport, the default
log.fileFileTransport, with size-based rotation
log.dispatchconfiguring which transports receive what

net

SubmoduleWhat it is for
net.tcpTcpSocket and TcpStream
net.udpUdpSocket
net.unixUnixStream, over a path rather than an address
net.tlsTLS over a TCP stream
net.dtlsDTLS over a UDP socket
net.ipparsing, formatting and classifying IP addresses
net.addrSocketAddrV4 and SocketAddrV6
net.pollasking which of a set of sockets is ready

os

SubmoduleWhat it is for
os.pathjoining, resolving and comparing path strings
os.fsdirectories, permissions, symlinks, globbing
os.envreading, writing and listing environment variables
os.processprocess identity, subprocesses, signals
os.systemfacts about the process, the runtime and the machine
os.tempfilethe temporary directory, and scratch files in it

rpc

SubmoduleWhat it is for
rpc.messageRequest, Notification, Response, and reading and writing them as JSON
rpc.errorRpcError, the codes the specification defines, and the module’s other errors
rpc.serviceService and the Context its handlers are given
rpc.httphttp_handler() and HttpClient, JSON-RPC over HTTP
rpc.framingHeaderFraming, LineFraming and MessageFraming, telling messages apart
rpc.transportthe transports, WebSockets among them, and pipe() between isolates
rpc.endpointEndpoint, a service on a connection
rpc.serverserve(), an endpoint for every connection to a listening socket

sql

SubmoduleWhat it is for
sql.driverthe contract an adapter implements, and the capability flags
sql.errorsevery error a database raises, under one root
sql.paramsrewriting ? and :name into whatever an engine wants
sql.typeshow Zuri values and database values correspond
sql.decimalDecimal, for a column a float must not hold
sql.resultResultSet and ExecResult
sql.cursorreading a result a row at a time
sql.statementa statement compiled once and run many times
sql.transactionTransaction, and the savepoints inside it
sql.connectionthe Connection a program holds
sql.crudbuilding the four statements that are always the same
sql.schemaasking a database what is in it
sql.poolkeeping connections open and lending them out
sql.sqlitethe SQLite adapter, and its blobs, backups and hooks
sql.postgresthe PostgreSQL adapter, and LISTEN/NOTIFY
sql.mysqlthe MySQL and MariaDB adapter

mail

SubmoduleWhat it is for
mail.errorsevery error the mail stack raises, under one root
mail.addressreading and writing the addresses in a header
mail.headersthe header block, in order and without regard to case
mail.encodingthe encodings a header and a body use
mail.contentContent-Type and Content-Disposition
mail.messagea message, its MIME tree, and building one
mail.dkimsigning a message and checking a signature
mail.saslthe authentication mechanisms all three protocols share
mail.streama line-oriented connection, and negotiating TLS over one
mail.smtpthe sending and receiving ends of SMTP
mail.imapthe client and server ends of IMAP, and where mail is kept
mail.imap.parserthe IMAP grammar
mail.imap.storeMailStore, MaildirStore and MemoryStore
mail.pop3the client end of POP3
mail.poolrunning a mail server on more than one connection at once

test

SubmoduleWhat it is for
test.expectExpect and every matcher on it
test.runnercollecting the declarations and running them
test.reporterReporter, and the seven built-in ones
test.resultCase, Suite, Failure and Summary
test.mockMock, mock() and spy_on()
test.snapshotthe snapshot store and its file format
test.conductdiscovering and running a directory of test files
test.diffstructural equality, and rendering what differs
test.formatrendering any value for a failure message
test.sourcereading a stack trace back to the failing line
test.styleterminal colour, symbols and width
test.errorAssertionError and TestSetupError
test.contextwhat is true while one test is running

validate

SubmoduleWhat it is for
validate.validatorsthe one-line entry points, one per rule
validate.validatorValidator, the fluent builder
validate.schemaSchema, validating a whole dictionary at once
validate.ruleRule, the base class custom rules extend
validate.rulesevery built-in rule

wire

SubmoduleWhat it is for
wire.compileturning a parsed template into an instruction tree
wire.renderwalking a compiled template and writing the page
wire.expressionthe language between {{ and }}
wire.filtersthe filters every template starts with
wire.escapecontext-aware escaping
wire.loaderresolving the path in an x-include
wire.normalizerewriting the pseudo elements before parsing
wire.valueshow a template reads the values it is given
wire.constantsthe directive names Wire reserves
wire.errorsthe module’s errors, and the locations they carry

zuri

SubmoduleWhat it is for
zuri.tokentokenize(), and the Token type it returns
zuri.astparse(), and the Node type it returns
zuri.compilecompile(), and the Instr type it returns
zuri.reflectinspecting a live function, class, module or instance

Shadowing a Module

A file in ./.zuri/libs/ shadows a standard library module of the same name, because that directory is searched first. See The Module System.

Appendix G: The Error Hierarchy

Every error in Zuri is an instance of a class, and every one of them inherits from Error. They are declared in ordinary Zuri and go through the same class machinery user code does, which is why subclassing one behaves exactly like subclassing anything else.

Error
├── TypeError
├── ValueError
├── NumericError
├── ArgumentError
├── NotImplementedError
├── RangeError
├── AccessError
├── AssertError
├── PropertyError
├── UndefinedError
└── ModuleNotFoundError

Fields

Every error carries three:

FieldWhat it holds
messagethe text, defaulting to 'An unexpected error has occurred'
typethe class name, as a string
stacktracea list of frames, innermost first
catch {
  raise ValueError('bad input')
} as e {
  echo e.type
  echo e.message
  echo e.stacktrace
}
ValueError
bad input
[/path/to/main.zu:2 -> @.script()]

When Each One Is Raised

ClassRaised when
Errorthe base class; a general failure with nothing more specific to say
TypeErroran operation received the wrong type: an undefined operator signature, a method on nil, an annotated parameter given the wrong thing
ValueErrorthe type was right and the value was not
NumericErroran arithmetic operation failed
ArgumentErrora call passed the wrong number of arguments to a native function
NotImplementedErrora method meant to be overridden was not
RangeErroran index or a bound fell outside what the value allows
AccessErrora permission or access check failed
AssertErroran assert condition was falsy
PropertyErrora member that does not exist was read: a missing dictionary key, an undeclared field, a module member that was not exported
UndefinedErroran undefined global was read
ModuleNotFoundErroran import could not be resolved

Catching

catch catches everything inside its block. To handle one kind and let the rest through, test and re-raise:

catch {
  load_config()
} as e {
  if !instance_of(e, ModuleNotFoundError) {
    raise e
  }

  echo 'no config, using defaults'
}

instance_of() walks the whole chain, so a test against Error matches everything.

A parameter annotated Error accepts any of them, which is the readable way to write a handler:

def report(e: Error) {
  echo '${e.type}: ${e.message}'
}

Subclassing

class HttpError < Error {
  @new(message, status) {
    parent(message)

    self.type = 'HttpError'
    self.status = status
  }
}

Two things make this work well. Call parent(message) so the base constructor sets message and the stack trace is captured. Set self.type so the class name appears in logs and in the uncaught-error banner.

Carry whatever the handler needs. An error class exists precisely so it can hold more than a string; the capstone’s TaskError carries the name of the field that failed validation, and that is what lets an API answer {"error": "...", "field": "title"}.

Uncaught

An error nobody catches prints its type, its message, the source around the failure, and the stack trace, then exits 1:

Unhandled ValueError: bottomed out
  --> /path/to/main.zu:3

  1 | def recurse(n) {
  2 |   if n <= 0 {
> 3 |     raise ValueError("bottomed out")
  4 |   }
  5 |   recurse(n - 1)

Stack trace (most recent call last):
  at recurse() /path/to/main.zu:3
  at recurse() /path/to/main.zu:5
  ... 19 more frames ...
  at recurse() /path/to/main.zu:5

A deep stack is truncated in the middle. The top and the bottom are the parts that tell you anything.

There Is No finally

Code after a catch statement runs whether the block raised or not, because the handler either recovers or re-raises. See Error Handling for the patterns that replace it.

Appendix H: Coming From Another Language

The places Zuri will surprise you are the places it looks most familiar. This is that list.

Everyone

You expectZuri does
x++ as a statement onlyx++ is postfix only, and in an expression it evaluates to the new value. ++x does not parse.
[] to be falsy[] and {} are truthy. Use is_empty().
finallyThere is none. Code after the catch statement runs either way.
tryThe keyword is catch, and it takes the block that might fail: catch { ... } as e { ... }.
new Thing()Call the class: Thing().
switch/case with fall-throughusing/when, first match only, no break.
a main functionA file’s top level is the program.
declaration hoistingNone. A def must appear above the top-level line that calls it.
overloading by arityNone. Two defs of one name in one scope is a compile error.
reopening a classclass Ext > Target adds methods to an existing class, globally. See Extensions.
string ordering with << is numbers only. Use compare(), which returns -1, 0 or 1.
x in collectionNo membership operator. Use contains().
?. and ??Neither exists. or covers the common case, with the truthiness caveat above.
an eval()There is none, deliberately. See Metaprogramming.

From Python

  • Blocks are braces, not indentation, and every control-flow body takes one.
  • def declares functions and class declares classes, but there is no self parameter: self is implicit inside a method and required for every field access.
  • The constructor is @new, not __init__. Dunder methods are @-prefixed decorated methods: @add, @lt, @key, @to_json.
  • __eq__ is @eq, and != is always its negation. It runs only when the other side is an object too, so x == nil never calls it. Without it, instances compare by identity.
  • No list comprehensions. map(), filter() and reduce() are methods on the list.
  • len(x) is x.length(). str(x) is x.to_string(). int(x) is x.int() or x.to_number().
  • Slicing is s[a, b], with a comma, not s[a:b].
  • elif is else if.
  • Modules run once and are cached, as in Python. Circular imports behave the same way, and have the same caveat.
  • if __name__ == '__main__' is if __root__ == __file__.

From JavaScript

  • var is block-scoped and behaves like let. const prevents rebinding and is enforced in local scopes.
  • == does no coercion. '1' == 1 is false. There is no ===.
  • Arrow functions are @(x) => x * 2 or def(x) => x * 2. The @ is the common spelling.
  • There is no this rebinding to worry about. self is the instance, always.
  • Objects and dictionaries are the same thing, and a class is not one. Classes are sealed: you cannot add a property to an instance, though a class Ext > Target declaration can add a method to the class.
  • null and undefined are both nil.
  • No async/await and no event loop. Concurrency is isolates: real threads with separate heaps, communicating by copying.
  • JSON.stringify is json.encode, and compact defaults to true.
  • Template literals are '${expr}', in ordinary single or double quotes.

From Ruby

  • No implicit returns. A function without return yields nil.
  • No blocks or yield. Pass an anonymous function.
  • nil is the only nil-like value, and false is separate from it.
  • Methods do not end in ? or !. Predicates are named is_* and mutation is documented rather than punctuated.
  • each hands the callback value first, index second. for hands you key first, value second.
  • There is no method_missing. Adding methods to an existing class is possible, but through an explicit class Ext > Target declaration rather than by reopening the class.
  • Modules are files, not a language construct. There is no include or extend.

From Go

  • Dynamically typed, with optional annotations on parameters that are checked at every call.
  • Errors are raised and caught, not returned. There is no err != nil pattern.
  • Isolates are not goroutines. They are OS threads with separate heaps, and values crossing between them are copied. That is the whole concurrency model, and it is why there are no mutexes.
  • Channels are bounded queues and behave the way you expect, select() included.
  • No interfaces. A parameter typed Error accepts any subclass, and that is the whole of the polymorphism story alongside inheritance.
  • defer has no equivalent. Close what you opened, on both paths.

From Java or C#

  • No static typing, no generics, no interfaces, no packages-as-namespaces. A module is a file.
  • Single inheritance, and no abstract keyword. A base method that raises NotImplementedError is the idiom.
  • public/private is a leading underscore, enforced at compile time for both class members and module members.
  • There is no overloading. One name, one method, and a second declaration of either is a compile error.
  • equals() is @eq, and == calls it.
  • toString() is to_string(), and nothing calls it for you. What echo shows for an instance comes from @to_string().

From C

  • Numbers are doubles. There is no integer type, and / never truncates. // is floor division.
  • % keeps the sign of the left operand, as in C. // rounds toward negative infinity, which C’s / does not.
  • No pointers, no manual memory management. A generational collector owns the heap.
  • bytes is the buffer type, and struct is how you read and write binary layouts.
  • switch is using, with no fall-through.

Things That Will Save You an Hour

Interpolation does not call to_string(). '${thing}' is <instance of Thing>. Call the method. echo shows an instance through @to_string(), a separate method.

list.sort() mutates and returns; list.reverse() does neither to the original. That asymmetry is the most common list bug in Zuri code.

A def scopes like a var. At the top level of a file it binds a module-level name; anywhere else it is a local of the block it is written in, and it goes away with that block.

import is local by default. If your module imports something and a third file cannot see it through you, add the @.

A same-directory import is import .sibling, not the full path from the project root.

A conditional expression breaks across lines either way. ? and : may each end a line or begin the next, so cond ? on one line and a line starting with ? or : both parse.

Appendix I: Custom Commands

zuri fmt, zuri init and zuri test are Zuri programs. The runtime ships them in a cmds directory beside itself, and when the first word after zuri is not run, it looks that word up there and runs the script it finds. A project adds commands of its own the same way, in its .zuri/cmds directory, and they are run, listed and documented exactly like the ones that ship.

This is the place for the scripts a project keeps running by hand: a release checklist, a data import, a code generator. As a command, each one is found by name from anywhere in the checkout’s root, shows up in zuri --help with a line saying what it does, and parses its own arguments with the same args module as everything else.

Where Commands Are Found

zuri <name> looks in these places, in this order, and runs the first match:

WhereWhat it holds
$ZURI_ROOT/cmds, or cmds beside the executable when ZURI_ROOT is unsetthe commands the runtime ships
.zuri/cmds in the projectthe project’s own commands
cmds in each package in the project’s .zuri/libscommands the project’s packages provide
cmds in each package in $ZURI_HOME/libscommands the packages installed for your user provide

The project is the nearest directory above the working directory that holds a project.toml, so a project’s commands run from anywhere inside it. With no project, .zuri in the working directory is used.

The shipped commands come first, so a project cannot replace one: a project command named test is never run, and zuri --help leaves it out of the listing. A project’s own commands come before anything a package provides, and a project’s packages before your user’s.

Two packages in the same place providing the same command is refused rather than settled by chance. zuri <name> then names both packages and runs neither, and zuri --help marks the command as claimed twice. zuri install refuses to create the clash in the first place.

zuri init writes a .gitignore that ignores everything under .zuri except .zuri/cmds, so a project’s commands are committed with the rest of its code while the tools that keep state in .zuri do not leave it behind in the repository.

Writing One

A command is a single .zu file, or a directory with an index.zu:

.zuri/
  cmds/
    greet.zu
    release/
      index.zu
      notes.zu
      tests/

zuri greet runs greet.zu, and zuri release runs release/index.zu. When a directory and a file share a name, the directory wins. The directory form is for a command that has grown past one file: index.zu imports its siblings relatively, as any package does, with import .notes.

The name a command answers to is its file or directory name. It is a single name, never a path: zuri tools/greet is refused as an unknown command before anything is looked up.

A name starting with _ is private, the way an identifier starting with one is. _notes.zu and _shared/index.zu are never commands: zuri --help leaves them out and zuri _notes is an unknown command. That is the place for code several commands share, which each of them imports relatively, as import .._shared.notes from release/index.zu.

Naming and Describing It

The first doc block in the file introduces the command. @command states the name it answers to, and @description is the one line zuri --help shows beside it:

/**
 * @command release
 * @description Tags a release and writes its notes from the commits
 *    since the last one.
 *
 * Run it from a clean checkout on the main branch:
 *
 *   zuri release 1.4.0
 */

import .notes

The block has to open within the first 16 KiB of the file, which in practice means at the top. A description longer than a line carries on over the indented lines beneath the tag, and ends at a blank line or the next tag. A command without a description is still listed, with its name alone. @command is what a reader of the file sees first, and it names the file’s own command; the runtime lists and runs a command by its file or directory name, so keep the two the same.

Arguments and the Exit Status

Everything after the command’s name belongs to the command, --help included, and a command reads it with the args module. Build a parser named after the command and call parse() with nothing: it reads the real command line, converts and validates each value, answers --help from the declarations, and refuses anything undeclared with a message and exit status 1. That is what gives every command, shipped or not, the same flags, the same help layout and the same errors:

/**
 * @command greet
 * @description Greets whoever is named, as often as asked.
 */

import args

def parser() {
  var p = args.Parser('greet', false)

  p.description = 'Greets whoever is named, as often as asked.'
  p.add_index('name', 'Who to greet', { required: true })
  p.add_option('times', 'How many greetings', { short_name: 't', type: args.INT, value: 1 })

  return p
}

var parsed = parser().parse()
var name = parsed.indexes[0]

iter var i = 0; i < parsed.options.times; i++ {
  echo 'Hello, ${name}!'
}

The raw list is there as well. os.args has the same shape for a command as for zuri run: the executable, the command’s own file, then what the user typed, so the command’s arguments are os.args[2,]. Reading that list by hand is how the checks and the help text drift apart, which is exactly what the parser exists to prevent, so a command should reach for args and leave os.args alone.

A command ends with status 0 when it runs to the end. An uncaught error prints its trace and ends it with status 1, and os.exit() ends it with any other status. A command that checks something, the way zuri fmt --dry-run checks formatting, should exit with a non-zero exit code when the check fails, so a shell script or a CI job can act on it.

Where It Runs

A command runs in the directory zuri was started in, so os.cwd() is the user’s directory, and for a project command that is the project’s root. __file__ is the command’s own file, which is how a command reaches something shipped beside it:

import os

var template = os.join_paths(os.dir_name(__file__), 'notes.template')

Seeing It Work

This example builds a throwaway project with one command, runs it the way a user would, asks it for its help message, and reads back the listing zuri --help gives:

import os

var project = os.create_temp_dir('zuri-commands-')
var commands = os.join_paths(project, '.zuri', 'cmds')

os.create_dir(commands, 0c755, true)

var source = file(os.join_paths(commands, 'greet.zu'), 'w')

source.write(
  '/**\n' +
  ' * @command greet\n' +
  ' * @description Greets whoever is named.\n' +
  ' */\n' +
  'import args\n' +
  'var parser = args.Parser("greet", false)\n' +
  'parser.description = "Greets whoever is named."\n' +
  'parser.add_index("name", "Who to greet", { value: "world" })\n' +
  'echo "Hello, " + parser.parse().indexes[0] + "!"\n'
)
source.close()

# Runs `zuri` in the project with `arguments`, and returns what it printed.
def zuri(arguments) {
  var child = os.spawn(os.exe_path, arguments, { cwd: project, stdin: 'null' })

  child.wait()

  return child.read_stdout().to_string()
}

echo zuri(['greet', 'Ada']).trim('\n')
echo zuri(['greet', '--help']).trim('\n')

# The part of the listing this project adds.
var listing = zuri(['--help'])
var start = listing.index_of('PROJECT COMMANDS')

echo listing[start, listing.index_of('\n\n', start)]

os.remove_dir(project, true)
Hello, Ada!
Usage: greet [OPTIONS] [name]

  Greets whoever is named.

POSITIONAL ARGUMENTS:
  [name]    Who to greet (default: world)

OPTIONS:
  -h, --help  Show this help message and exit
PROJECT COMMANDS:
  greet      Greets whoever is named.

Testing a Command

Build the parser in a function, as the greet command above does, and the argument handling can be tested by handing parse() a list, the seam Chapter 20 describes. Everything else is ordinary Zuri: keep the command’s logic in the functions and files beside index.zu, and test those.

The shipped commands keep their tests in a tests directory inside the command, and a project command can do the same. zuri test runs a directory anywhere in the project:

zuri test .zuri/cmds/release/tests