The Zuri Programming Language
by the Zuri Project
Zuri is a language for building whole applications, and it ships with everything that takes. The runtime, the web server, the templating engine, the data formats, the cryptography and the toolchain were designed together and are delivered as one binary, so starting a real program does not begin with assembling an ecosystem out of third-party parts. Zuri has first-class language-level support for package management and vendoring, so the code you bring in from outside needs no third-party tooling either.
The standard library ships with the language, and it is large. Templating,
HTTP/1.1 and HTTP/2, WebSockets, TLS, JSON, YAML, CSV, compression,
cryptography, image decoding and drawing, HTML parsing, date arithmetic
with the IANA time zone database, and OS-thread concurrency are all part
of the installation. Every one of them is reachable with a bare import.
The tools ship in the same binary. zuri fmt lays out your code,
zuri test runs your tests, and zuri lsp gives every editor
completion, navigation and diagnostics. zuri install and
zuri publish manage packages against Nyssa, a package registry any
installation can serve, and zuri bundle ships a finished program to
machines where Zuri is not installed, as a single executable if you
like.
Zuri is dynamically typed, and it declines to be vague. A declared parameter type is enforced at the call, classes are sealed, a function cannot be quietly redefined, and a name beginning with an underscore stays inside the module that declared it. Your code runs as bytecode until a piece of it gets hot, and a JIT compiles that piece to machine code while the program is still running.
This book teaches the language from the first line of code to a complete web application. It is written against the Rust implementation of Zuri. Every program printed in these pages was run against that implementation, and the output shown underneath each one is the output it produced.
Introduction
Welcome. Open a terminal, keep it next to this page, and let’s begin.
This book teaches Zuri. It starts with installing the language and printing a line of text, and it ends with a task-board web application that stores data on disk, renders HTML from templates, serves a JSON API, and runs behind a stack of middleware. Everything in between is the road from one to the other.
What Zuri Looks Like
Here is a complete program. You do not need to understand all of it yet; read it the way you would read a paragraph in a language you are learning, and see how much comes through.
class Account {
@new(owner: string, balance: number) {
self.owner = owner
self.balance = balance
}
deposit(amount: number) {
if amount <= 0 {
raise ValueError('deposit must be positive')
}
self.balance += amount
return self.balance
}
to_string() {
return '${self.owner}: ${self.balance}'
}
}
var accounts = [
Account('ada', 100),
Account('grace', 250),
]
for account in accounts {
account.deposit(50)
echo account.to_string()
}
ada: 150
grace: 300
Most of that will be familiar if you have written code before. The pieces worth pointing at now, because they come up on the very first page of Chapter 3:
vardeclares a variable, andclassdeclares a class.- A constructor is called
@new. Methods whose names begin with@hook into the language’s own syntax, and Chapter 6 covers the full set. selfrefers to the instance, and reading a field always goes through it:self.balance, never a barebalance.'${...}'interpolates an expression into a string.echoprints a value and a newline.raisesignals an error; Chapter 7 shows how to catch one.- Every block is written with braces, and every control-flow statement requires one.
What Comes With Zuri
Installing Zuri installs its whole toolchain with it. Every tool is a
subcommand of zuri, and every one has its place in this book:
| Command | What it does | Where |
|---|---|---|
zuri | the interactive REPL | Chapter 1 |
zuri run | runs a script, a package or a directory | Chapter 1 |
zuri init | starts a project | Chapter 1 |
zuri fmt | lays out Zuri source in the house style | Chapter 1 |
zuri test | runs a project’s tests | Chapter 23 |
zuri lsp | the language server behind your editor | Chapter 29 |
zuri install | adds packages, resolved into a lockfile | Chapter 27 |
zuri publish | shares a package on a registry | Chapter 27 |
zuri serve | runs Nyssa, a package registry of your own | Chapter 27 |
zuri bundle | ships a program to machines without Zuri | Chapter 27 |
zuri upgrade | moves the installation to a newer release | Chapter 27 |
A project adds commands of its own the same way, and so can a package; Appendix I shows how.
Who This Book Is For
Chapters 1 through 6 assume you can open a terminal and nothing else. If this is your first programming language, start at the beginning and go slowly. Every idea is introduced before it is used, and every example is short enough to type out.
If you already write Python, JavaScript, Ruby, Go or Java, you can move through Chapters 3 to 5 quickly. Read Appendix H first: it lists the places where Zuri does something different from what the same syntax does in the language you already know, which is where the hours get lost.
How This Book Is Organised
Getting started, Chapters 1 and 2. Install the language, run something, start a project and format it, then build a small command-line application end to end so you have seen the shape of a real program before we take one apart.
The language, Chapters 3 to 8. Variables, types, operators, control flow, strings, numbers, collections, functions, closures, type annotations, classes, inheritance, decorated methods, errors and modules. Read these in order.
The world outside your program, Chapters 9 to 13. Files, binary data and byte streams, isolates and concurrency, sockets and networking, and a tour of the standard library.
The large modules, Chapters 14 to 20. Wire for templating, HTTP for clients and servers, Imagine for images, SQL for databases, Mail for the three protocols that move messages, env for the configuration all of them read, and args for the command line they are started from. Each is big enough to need a chapter of its own, and each is a reference you will come back to.
Depth, Chapters 21 to 24. Reflection and the compiler API, how Zuri executes your code and how to make it faster, how to test it, and how to debug a program that is doing something you did not expect.
The capstone, Chapter 25. One application, built in seven steps, using almost everything the book has covered.
Interoperability, Chapter 26. Calling C and Rust libraries, and being called back by them.
Sharing and shipping, Chapter 27. Installing and publishing packages, the lockfile, packaging a program to run where Zuri is not installed, and running Nyssa, the package repository, for a team or the public.
Talking to other programs, Chapter 28. JSON-RPC, for calling methods on another program and answering its calls over any connection.
Tooling, Chapter 29. The language server, and setting up your editor to use it.
Appendices. Keywords, operators and precedence, decorated methods, built-in functions, every method on every built-in type, the standard library index, the error hierarchy, the notes for readers arriving from another language, and writing commands of your own. These are reference material; the rest of the book is prose.
How to Read It
Read it with the interpreter running. Zuri has a REPL, every short example in this book can be pasted straight into it, and the fastest way to understand a rule is to break it on purpose and read the error.
Chapters build on each other. When a chapter needs something from later in the book, it says so and links to it, and you can carry on without following the link. When a chapter introduces something you will need again, it says that too.
Conventions
Commands you type in a shell appear with a $ prompt, and the output
follows underneath:
$ zuri run main.zu
Hello, world!
Code that belongs in a file appears with the filename above it when the filename matters:
Filename: main.zu
echo 'Hello, world!'
Sessions in the interactive prompt use %> for the first line of an input
and .. for its continuations, which is exactly what the REPL itself
prints:
%> var name = 'zuri'
%> name.upper()
'ZURI'
Where a rule has an exception, the exception is stated in the same paragraph as the rule. Where a limit exists, it is stated as a limit. Nothing in this book is a guess about how the language behaves; every claim was checked by running it.
Let’s get started.
Getting Started
Three things before the language itself: putting Zuri on your machine, running a file, and using the interactive prompt.
The prompt matters more than it sounds. Zuri prints the value of any expression you type at it, which makes it the fastest way to answer “what does this actually return?” — and that question comes up constantly in the next five chapters.
Installation
Zuri installs with one command, which downloads the release for your
machine, checks it, and puts zuri on your PATH. Releases cover:
- Linux on x86-64 and ARM (
x86_64-unknown-linux-gnu,aarch64-unknown-linux-gnu) - macOS on Intel and Apple silicon (
x86_64-apple-darwin,aarch64-apple-darwin) - Windows on x86-64 (
x86_64-pc-windows-msvc)
Zuri also builds from source with Cargo, which the end of this page covers.
Linux and macOS
$ curl -fsSL https://zuri-lang.github.io/zuri/install.sh | sh
Downloading zuri 0.1.0 for x86_64-unknown-linux-gnu
Installed zuri 0.1.0 in /home/ada/.zuri/runtime
Added /home/ada/.zuri/bin to PATH in /home/ada/.bashrc
Open a new terminal, or run this one to use zuri straight away:
export PATH="$HOME/.zuri/bin:$PATH"
Get started with:
zuri --help
The installer downloads the archive for your machine and checks it
against the checksum published beside it. It then runs the new zuri
once to confirm the version it reports, and only after that moves it
into ~/.zuri/runtime, so a failed or tampered download never replaces
a working installation. zuri is linked into ~/.zuri/bin, the same
directory that globally installed tools put their launchers in, and
that one directory goes on your PATH through your shell’s startup
file: .zshrc for zsh, .bashrc for bash (.bash_profile on macOS),
config.fish for fish, and .profile for anything else.
Options go after sh -s --:
$ curl -fsSL https://zuri-lang.github.io/zuri/install.sh | sh -s -- --version 0.2.0
| Option | What it does |
|---|---|
--version <version> | installs that release rather than the newest |
--prerelease | considers pre-releases when picking the newest |
--no-modify-path | leaves your shell’s startup files alone |
Windows
In PowerShell:
> powershell -c "irm https://zuri-lang.github.io/zuri/install.ps1 | iex"
The Windows installer makes the same checks and installs into
%USERPROFILE%\.zuri\runtime. That directory goes on your user PATH,
and so does %USERPROFILE%\.zuri\bin, where globally installed tools
put their launchers. A terminal opened afterwards has both; the one the
installer ran in has them already.
The same options are parameters, given by running the installer as a script block:
> & ([scriptblock]::Create((irm https://zuri-lang.github.io/zuri/install.ps1))) -Version 0.2.0
-Version, -Prerelease and -NoModifyPath match the options above.
Checking the Install
$ zuri
Zuri 0.1.0 (running on ZuriVM 0.1.0), REPL/Interactive mode = ON
Build No. => 2026-09-10 23:19:10 UTC
Type ".exit" to quit, ".help" for help or ".credits" for more information
%>
That %> is the Zuri prompt. Type .exit to leave.
zuri --version reports the same build without opening a session, which
is the one to reach for from a script or a CI job, and zuri --help
adds everything the runtime can run from here:
$ zuri --version
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC
Where It Goes
Both installers keep everything under ZURI_HOME, which is .zuri in
your home directory unless the variable says otherwise:
~/.zuri/
runtime/ zuri, with libs and cmds beside it
bin/ the zuri link, and globally installed tools
runtime holds the executable with copies of the libs and cmds
directories beside it. That pairing matters: the standard library is
written mostly in Zuri itself, and the runtime finds it by looking for a
libs directory beside the executable. cmds is the same arrangement
for the commands the runtime ships.
Running an installer again installs over the top. zuri upgrade does
the same from inside an installation, and
Bundles and Upgrades covers it.
Removing Zuri is deleting ~/.zuri/runtime and the zuri link in
~/.zuri/bin, along with the line the installer added to your shell’s
startup file.
ZURI_RELEASES_URL points both installers, and zuri upgrade, at
another listing of releases in the shape GitHub publishes them, such as
a mirror. GITHUB_TOKEN, when set, is sent to GitHub’s API, where it
raises the rate limit.
Downloading by Hand
Every release on the
releases page carries an
archive per platform, named zuri-<version>-<target>.tar.gz, or .zip
for Windows, with a .sha256 beside it. Check the archive against it,
unpack it, and put the zuri it holds on your PATH, keeping the
libs and cmds directories beside it.
Building From Source
Building needs Rust’s package manager, Cargo. If you do not have Rust
installed, get it from rustup.rs; it takes a minute
and installs cargo for you.
$ git clone https://github.com/zuri-lang/zuri
$ cd zuri
$ cargo patch-crates
$ cargo build --release
cargo patch-crates prepares the few dependencies Zuri builds with
changes of its own. It runs once per clone, and again only when one of
those changes is updated; the build says when that is.
The first build compiles a large set of native dependencies, so it takes a few minutes. Builds after that are incremental and quick.
When it finishes you have an executable at target/release/zuri, and,
sitting right next to it, copies of the libs/ and cmds/ directories,
in the same arrangement a release has.
Putting a Build on Your PATH
The simplest arrangement is to keep the binary and its two directories together and symlink to the binary:
$ sudo ln -s "$PWD/target/release/zuri" /usr/local/bin/zuri
A symlink is resolved before the runtime looks for libs, so this works.
Copying just the binary does not.
If you want them somewhere else entirely, set ZURI_ROOT to the
directory that contains them:
$ export ZURI_ROOT=/opt/zuri
$ ls /opt/zuri
cmds libs
ZURI_ROOT wins over the executable-adjacent lookup, which makes it handy
when you are hacking on the standard library itself and want a build to
pick up your edits immediately:
$ ZURI_ROOT=$PWD ./target/debug/zuri run myscript.zu
A Debug Build, and Why You Might Want One
Plain cargo build produces target/debug/zuri. It keeps the runtime
assertions that the optimised build strips out, which makes it the build to
reach for when a program is doing something you did not expect. It runs
more slowly in exchange. Either build runs everything in this book.
Hello, World!
Make a directory, put one file in it, and run it.
$ mkdir hello
$ cd hello
Create a file called main.zu. The .zu extension is what the runtime
looks for when resolving imports, so get in the habit early.
Filename: main.zu
echo 'Hello, world!'
Run it:
$ zuri run main.zu
Hello, world!
That is the whole program. No main function, no imports, no boilerplate.
A Zuri file is a script, and its top level is code that runs.
Anatomy of One Line
echo 'Hello, world!'
echo is a keyword, not a function. You do not write echo(...), you
write echo followed by an expression. It prints the value and adds a
newline.
Strings are written between single or double quotes, and the two are identical in meaning. Most Zuri code uses single quotes and saves double quotes for strings that contain an apostrophe.
There is no semicolon. Zuri ends a statement at the end of the line. You can write a semicolon if you want two statements on one line, but almost nobody does.
echo and print
There is also a print() function, and the difference is worth learning
now because it bites people later.
echo 'first'
echo 'second'
print('third')
print('fourth\n')
first
second
thirdfourth
echo appends a newline; print() does not, and it takes any number of
arguments. Use echo for output meant for a human reading a terminal, and
print() when you are assembling output character by character.
Running a Directory
run takes a directory as readily as a file. Point it at one and it
looks for index.zu inside and runs that:
$ zuri run hello
Hello, world!
If there is no index.zu, you get told so plainly:
$ zuri run hello
(Zuri):
Launch aborted for hello
Reason: No entrypoint found in the directory
This is the same rule the module system uses for packages, which we will
get to in Chapter 8. A directory with an index.zu
is a unit you can run or import.
Leave the path off and run launches the directory you are standing in,
which is how a project is usually started:
$ cd hello
$ zuri run
Hello, world!
Everything after the path belongs to the program rather than to zuri,
so a script reads its own flags exactly as it would anywhere else:
$ zuri run main.zu --name Ada --verbose
Chapter 20 covers reading them.
Starting a Project
One file is where everybody starts, and it stops being enough about as
soon as there are two of them. zuri init writes the layout a project
grows into:
$ zuri init myproject
Run in a terminal, it asks what the project is called, what version it
starts at and who wrote it, offering an answer to each. Press enter to
take what it offers. Run anywhere there is nobody to ask — a script, a
CI job — it takes those answers itself, which --yes also says
outright.
What it leaves behind:
myproject/
project.toml what the project is
index.zu the entry point
app/
index.zu the application
tests/
app.test.zu its tests
README.md
.gitignore
.gitattributes
It runs, and its tests pass, before you have written anything:
$ cd myproject
$ zuri run
Hello, world!
$ zuri test
zuri test 1 file in tests
PASS app.test.zu 36ms 2 tests
1 file
2 passed • 2 total
The root index.zu is two things at once:
Filename: index.zu
import @.app { * }
if __root__ == __file__ {
main()
}
The import re-exports everything app declares, so another project can
import myproject and reach all of it. The if starts the application,
but only when this file is the one that was run: __root__ is the file
zuri was pointed at, and __file__ is the file the line is written
in. Importing the project leaves main() alone. That is the Zuri
spelling of a main guard, and it is worth knowing early because every
runnable package uses it.
app/index.zu is the application itself. Whatever it exports, the
project exports:
Filename: app/index.zu
def greet(name: ?string) {
return 'Hello, ${name or "world"}!'
}
def main() {
echo greet()
}
project.toml is what the project is, for anything that reads it:
[project]
name = "myproject"
version = "0.1.0"
description = ""
authors = ["Ada Lovelace <ada@example.com>"]
license = "MIT"
[dependencies]
The name comes from the directory, the author from git config, and a
git repository is created unless there is already one above. zuri init --help has the rest, and every question it asks has a flag that answers
it:
$ zuri init myproject --name web-server --license Apache-2.0 --yes
Run inside a directory that already has files, init adds what is
missing and writes over nothing. Run where a project already is, it
refuses rather than overwrite one.
Chapter 8 is where packages and re-exports are covered properly, and Chapter 25 grows this layout into a full application.
Formatting
zuri fmt lays out Zuri source in one house style: two spaces to a
level, one statement to a line, an explicit block wherever one belongs,
spacing the way a reader expects it, and lines past eighty columns broken
where the expression holds together least. Line breaks you made yourself
are kept, so a list you laid out one entry to a line stays that way.
Here is a file written in a hurry:
Filename: rooms.zu
def area(width,height){return width*height}
var rooms=[ [3,4],[5,6] ]
for room in rooms { echo area(room[0],room[1]) }
$ zuri fmt rooms.zu
formatted rooms.zu
Formatted 1 file.
Filename: rooms.zu
def area(width, height) {
return width * height
}
var rooms = [[3, 4], [5, 6]]
for room in rooms {
echo area(room[0], room[1])
}
Given a directory, zuri fmt formats every .zu file beneath it apart
from the packages a project has installed, and with no path at all it
formats the current directory. It never changes what a program does:
before it writes a file, it parses its own result and checks that the
program and every comment in it are exactly what it started with, and a
file it cannot format that way is left as it was.
--dry-run writes nothing. It names each file that would change, and
exits with status 1 when any would, which is the check a commit hook or
a CI job wants:
$ zuri fmt --dry-run
would format /home/ada/myproject/app/extra.zu
1 file would change. 3 files were already formatted.
Your editor runs the same formatter through the language server, on a whole file or a selection; Chapter 29 sets that up.
Commands
A first word that is not run names a command instead of a path:
$ zuri greet Ada
Hello, Ada!
A command is a .zu file, or a directory with an index.zu, sitting in
a cmds directory. A project keeps its own in .zuri/cmds, so either of
these answers to zuri greet:
.zuri/cmds/greet.zu # a command in one file
.zuri/cmds/greet/index.zu # a command with room to grow
The directory wins if both are there, so a command that has outgrown one
file takes over the name as soon as its index.zu lands. Until then the
directory is not a command, and the file goes on answering.
The commands the runtime itself ships live in the cmds directory beside
the executable, and those win over a project’s. Everything after the name
is forwarded to the command untouched, flags included.
A name that matches nothing is refused rather than guessed at:
$ zuri gret Ada
(Zuri):
Launch aborted for gret
Reason: Unknown command
A command says what it is in its own doc block, with two tags. zuri --help lists every command it can reach by exactly those:
Filename: .zuri/cmds/greet.zu
/**
* @command greet
* @description Say hello to somebody by name.
*/
import os
echo 'Hello, ${os.args[2]}!'
$ zuri --help
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC
Usage: zuri start the interactive REPL
zuri run [PATH] run a script, a package, or this directory
zuri <command> [ARGS] run a command
OPTIONS:
-h, --help Show this help message and exit
-v, --version Show version information and exit
COMMANDS:
format Lay out Zuri source in the project style.
init Scaffold a new Zuri project.
test Run a project's test files, each in its own process.
PROJECT COMMANDS:
greet Say hello to somebody by name.
Run "zuri <command> --help" for help on a specific command.
A description too long for one line carries on below it, indented:
/**
* @command deploy
* @description Ship the current build to staging, then wait for
* the health check to come back green.
*/
Two flags sit outside all of this and run nothing. zuri --help is the
listing above, and zuri --version reports just the build:
$ zuri --version
Zuri 0.1.0 (running on ZuriVM 0.1.0)
Build No. => 2026-09-10 23:19:10 UTC
Both take the short spelling too, -h and -v.
The REPL
Run zuri with no arguments and you get an interactive prompt:
$ zuri
Zuri 0.1.0 (running on ZuriVM 0.1.0), REPL/Interactive mode = ON
Build No. => 2026-09-10 23:19:10 UTC
Type ".exit" to quit, ".help" for help or ".credits" for more information
%>
REPL stands for read-eval-print loop, and the print is the part that
makes it useful. At the top level of the REPL, any expression whose value
is not nil is printed for you, so you never need echo just to look at
something:
%> 1 + 2
3
%> var name = 'Zuri'
%> name.upper()
ZURI
%> [1, 2, 3].map(@(n) { return n * n })
[1, 4, 9]
Notice that var name = 'Zuri' printed nothing. A declaration is not an
expression, and anything that evaluates to nil stays quiet. That single
rule is what keeps list.append(4) and echo x from doubling up on your
screen.
Auto-printing applies only to the outermost level. Type a loop and the body does not print once per iteration:
%> iter var i = 0; i < 3; i++ {
.. i * 10
.. }
%>
Use echo inside a block when you want to see something.
Multi-line Input
The prompt changes from %> to .. when what you have typed so far is
not yet a complete statement. Press Enter at the end of a line with an
unclosed brace and keep going:
%> def double(n) {
.. return n * 2
.. }
%> double(21)
42
The same applies to an unclosed bracket, an unclosed parenthesis, or a string that has not been closed. The REPL decides when you are finished; you never have to signal it.
Everything Stays Defined
Variables, functions and classes you declare at the top level of the REPL remain defined for the rest of the session, because the whole session shares one namespace:
%> class Point {
.. @new(x, y) {
.. self.x = x
.. self.y = y
.. }
.. }
%> var p = Point(3, 4)
%> (p.x ** 2 + p.y ** 2).sqrt()
5
Redeclaring a name at the REPL’s top level replaces it rather than raising an error, so you can retype a function until it is right.
Getting Out
Type .exit. That is the only thing that ends the session.
Ctrl+C and Ctrl+D do not exit. Both print an interrupt line and hand the prompt back:
%> <KeyboardInterrupt [CtrlC]>
Type '.exit' to exit the REPL session
%>
Ctrl+C is how you abandon a half-typed line: the line is discarded and you start again on a fresh prompt.
Syntax Highlighting
On a terminal, the REPL colours what you type as you type it. Keywords are
highlighted and everything else is left plain, so a misspelled retrun or
whiel stands out before you press Enter — it simply does not change
colour.
The highlighting understands string literals. A keyword inside quotes is text, not a keyword, and is left uncoloured:
%> var note = 'return it later'
return there stays plain, because it is part of the string.
Alongside it, a greyed-out suggestion appears to the right of the cursor when what you have typed so far matches something earlier in your history. Press the right arrow to accept it.
Both are terminal features. Piping input to zuri or redirecting its
output produces plain text, so a session captured to a file has no escape
codes in it.
The Other Dot Commands
.help reminds you that Tab offers completions.
.credits opens the project’s licence in a pager. Press q to come
back.
Both are completed by Tab, along with every keyword in the language.
History
Up and Down walk through previous lines, and Ctrl+R searches backwards
through them. History is written to a history.txt file sitting next to
the zuri executable, so it survives between sessions.
What the REPL Is Good At
Each line is compiled and run on its own, which makes the REPL the right tool for a specific kind of question:
- What does this method return?
'a,b,,c'.split(',') - Is this value truthy?
!!(-1) - What type am I actually holding?
typeof(x) - Does this syntax parse the way I think? Type it and find out.
It is a poor place to write a program, because you cannot go back and
edit line four. Once you are past a few lines, put them in a .zu file and
run it — which is what the next chapter does.
Programming a Bookmark Keeper
We are going to build something small and real: a command-line bookmark keeper. It asks you for a title and a URL, stores what you give it in a JSON file, and lets you list, search and delete entries later.
By the end of this chapter you will have used variables, functions, lists, dictionaries, loops, branching, string methods, file handling, JSON and error handling. We will not explain any of them thoroughly. That is what Chapters 3 to 8 are for. The goal here is to see a whole program, working, before we take it apart.
Follow along by typing the code rather than copying it. Mistakes are the fastest way to learn what the error messages mean.
Setting Up
$ mkdir bookmarks
$ cd bookmarks
Create main.zu and start with the modules we will need:
Filename: main.zu
import io
import json
var STORE = 'bookmarks.json'
echo 'Bookmark keeper.'
$ zuri run main.zu
Bookmark keeper.
Three things to notice. import brings in a module and binds it to a name
you then reach through with a dot. var declares a variable. A .zu file’s
top level is just code, running from top to bottom.
Asking a Question
The io module has readline(), which prints a prompt and waits for the
user to type a line:
var title = io.readline('Title: ')
echo 'You typed: ' + title
$ zuri run main.zu
Title: Zuri docs
You typed: Zuri docs
The value that comes back includes whatever the user typed, whitespace
included, so we almost always follow it with .trim():
var title = io.readline('Title: ').trim()
trim() is a method on strings. Every value in Zuri has methods,
numbers, booleans and nil included, and you call them with a dot.
Storing One Bookmark
A bookmark has a title, a URL and some tags. The natural shape for that is a dictionary: a set of keys with values attached.
var bookmark = {
title: 'Zuri docs',
url: 'https://zuri.dev',
tags: ['lang', 'docs'],
}
Square brackets make a list, an ordered sequence. Curly braces make a dictionary. You read a dictionary’s values with a dot, the same way you reach into a module:
echo bookmark.title
echo bookmark.tags
Zuri docs
[lang, docs]
Zuri has a shorthand you will use constantly. When the key you want and the variable holding its value have the same name, write the name once:
var title = 'Zuri docs'
var url = 'https://zuri.dev'
var bookmark = { title, url }
That is exactly the same dictionary as { title: title, url: url }, with
half the noise.
Many Bookmarks
One bookmark is a dictionary. Many bookmarks is a list of them:
var bookmarks = []
bookmarks.append({ title: 'Zuri docs', url: 'https://zuri.dev' })
bookmarks.append({ title: 'Cranelift', url: 'https://cranelift.dev' })
echo bookmarks.length()
2
Saving to Disk
The json module turns Zuri values into text and back. file() opens a
file; 'w' means open it for writing.
def save(bookmarks) {
var handle = file(STORE, 'w')
handle.write(json.encode(bookmarks, false))
handle.close()
}
def declares a function. The parameter list needs no types (though it can
have them; see Chapter 5). The body is a
block, and blocks in Zuri always use braces.
That second argument to json.encode() is compact. It defaults to true,
which gives you one dense line. Passing false asks for the indented form,
which is what you want in a file a human might open:
[
{
"title": "Cranelift",
"url": "https://cranelift.dev",
"tags": [
"compilers"
]
}
]
Loading From Disk, Carefully
Reading is where a real program has to start thinking about what can go wrong. The file might not exist yet. It might exist and contain garbage because someone edited it by hand. Neither should stop the program.
def load() {
var handle = file(STORE)
if !handle.exists() {
return []
}
catch {
return json.decode(handle.read())
} as error {
echo 'Could not read ${STORE}: ${error.message}'
return []
}
}
Two new pieces here.
The first is ${...} inside a string. That is interpolation: the
expression between the braces is evaluated and its value spliced into the
string. It works in both single- and double-quoted strings.
The second is catch. Zuri does not have try. You write catch { ... }
around the code that might fail, and as error { ... } to handle it. If
nothing goes wrong, the handler never runs. If something does, error is
an object with a message, a type and a stacktrace.
Notice there is no finally. Zuri does not have one. Code after the
catch statement runs either way, which covers most of what finally is
used for.
Adding a Bookmark
Now we can put the pieces together:
def add(bookmarks) {
var title = io.readline('Title: ').trim()
var url = io.readline('URL: ').trim()
if title.is_empty() or url.is_empty() {
echo 'Both a title and a URL are required.'
return
}
var tags = io.readline('Tags (comma separated): ').trim()
bookmarks.append({
title,
url,
tags: tags.is_empty() ? [] : tags.split(',').map(@(t) { return t.trim() }),
})
save(bookmarks)
echo 'Saved "${title}".'
}
The tags line packs in three ideas.
cond ? a : b is the conditional operator: if cond is truthy the whole
expression is a, otherwise b.
split(',') cuts a string into a list at every comma.
map() takes a function and applies it to every element, giving back a new
list. @(t) { return t.trim() } is an anonymous function: @ followed
by a parameter list and a body. You can also spell it def(t) { ... }; @
is the shorthand and it is what most Zuri code uses for a one-liner.
There is a shorter form still. When the body is a single expression that
you want returned, => replaces the braces and the return:
tags.split(',').map(@(t) => t.trim())
That is the same function written three ways. Chapter 5 covers every spelling, and when each one reads best.
Also notice or rather than ||. Zuri spells its logical operators and,
or and !.
Printing a Bookmark
def show(bookmark, index) {
echo '${index + 1}. ${bookmark.title}'
echo ' ${bookmark.url}'
if !bookmark.tags.is_empty() {
echo ' [' + ', '.join(bookmark.tags) + ']'
}
}
join lives on the string, not on the list, and it reads exactly as it
works: take this separator, and stitch that list together with it.
Listing and Searching
def list_all(bookmarks) {
if bookmarks.is_empty() {
echo 'Nothing saved yet.'
return
}
bookmarks.each(@(bookmark, index) {
show(bookmark, index)
})
}
each() calls your function once per element, handing it the value
first and the index second. That order catches people out; it is the
same for lists, dictionaries and strings.
Search is a filter plus a test:
def find(bookmarks) {
var needle = io.readline('Search: ').trim().lower()
var hits = bookmarks.filter(@(bookmark) {
if bookmark.title.lower().contains(needle) {
return true
}
return bookmark.tags.some(@(tag) { return tag.lower() == needle })
})
if hits.is_empty() {
echo 'No match for "${needle}".'
return
}
hits.each(@(bookmark, index) {
show(bookmark, index)
})
}
filter() keeps the elements your function returns true for. some()
answers “is this true of at least one element?” and stops at the first one
that matches.
Removing a Bookmark
def remove(bookmarks) {
list_all(bookmarks)
if bookmarks.is_empty() {
return
}
var answer = io.readline('Remove which number? ').trim()
var position = answer.to_number() - 1
if position < 0 or position >= bookmarks.length() {
echo 'There is no bookmark ${answer}.'
return
}
var gone = bookmarks[position]
bookmarks.remove_at(position)
save(bookmarks)
echo 'Removed "${gone.title}".'
}
to_number() converts a string; bookmarks[position] indexes a list;
remove_at() deletes by position and shifts the rest down.
The bounds check is not optional politeness. Indexing a list past its end raises an error, and an unhandled error ends the program.
The Main Loop
The last piece is a loop that reads a command and dispatches on it:
def main() {
var bookmarks = load()
echo 'Bookmark keeper. ${bookmarks.length()} saved.'
while true {
var command = io.readline('\n(a)dd (l)ist (f)ind (r)emove (q)uit > ').trim().lower()
using command {
when 'a' add(bookmarks)
when 'l' list_all(bookmarks)
when 'f' find(bookmarks)
when 'r' remove(bookmarks)
when 'q' {
echo 'Bye.'
return
}
default {
echo 'Unknown command "${command}".'
}
}
}
}
main()
using is Zuri’s multi-way branch, and when is its branch keyword. It
compares the subject against each when value and runs the first match;
default catches everything else. A when whose body is a single short
statement can stay on one line, which is what makes the dispatch table
above readable. Anything longer gets a block.
Unlike a C switch, there is no fall-through and no break to remember.
One branch runs, then the statement is over.
The Whole Program
Filename: main.zu
import io
import json
var STORE = 'bookmarks.json'
def load() {
var handle = file(STORE)
if !handle.exists() {
return []
}
catch {
return json.decode(handle.read())
} as error {
echo 'Could not read ${STORE}: ${error.message}'
return []
}
}
def save(bookmarks) {
var handle = file(STORE, 'w')
handle.write(json.encode(bookmarks, false))
handle.close()
}
def add(bookmarks) {
var title = io.readline('Title: ').trim()
var url = io.readline('URL: ').trim()
if title.is_empty() or url.is_empty() {
echo 'Both a title and a URL are required.'
return
}
var tags = io.readline('Tags (comma separated): ').trim()
bookmarks.append({
title,
url,
tags: tags.is_empty() ? [] : tags.split(',').map(@(t) { return t.trim() }),
})
save(bookmarks)
echo 'Saved "${title}".'
}
def show(bookmark, index) {
echo '${index + 1}. ${bookmark.title}'
echo ' ${bookmark.url}'
if !bookmark.tags.is_empty() {
echo ' [' + ', '.join(bookmark.tags) + ']'
}
}
def list_all(bookmarks) {
if bookmarks.is_empty() {
echo 'Nothing saved yet.'
return
}
bookmarks.each(@(bookmark, index) {
show(bookmark, index)
})
}
def find(bookmarks) {
var needle = io.readline('Search: ').trim().lower()
var hits = bookmarks.filter(@(bookmark) {
if bookmark.title.lower().contains(needle) {
return true
}
return bookmark.tags.some(@(tag) { return tag.lower() == needle })
})
if hits.is_empty() {
echo 'No match for "${needle}".'
return
}
hits.each(@(bookmark, index) {
show(bookmark, index)
})
}
def remove(bookmarks) {
list_all(bookmarks)
if bookmarks.is_empty() {
return
}
var answer = io.readline('Remove which number? ').trim()
var position = answer.to_number() - 1
if position < 0 or position >= bookmarks.length() {
echo 'There is no bookmark ${answer}.'
return
}
var gone = bookmarks[position]
bookmarks.remove_at(position)
save(bookmarks)
echo 'Removed "${gone.title}".'
}
def main() {
var bookmarks = load()
echo 'Bookmark keeper. ${bookmarks.length()} saved.'
while true {
var command = io.readline('\n(a)dd (l)ist (f)ind (r)emove (q)uit > ').trim().lower()
using command {
when 'a' add(bookmarks)
when 'l' list_all(bookmarks)
when 'f' find(bookmarks)
when 'r' remove(bookmarks)
when 'q' {
echo 'Bye.'
return
}
default {
echo 'Unknown command "${command}".'
}
}
}
}
main()
A session looks like this:
$ zuri run main.zu
Bookmark keeper. 0 saved.
(a)dd (l)ist (f)ind (r)emove (q)uit > a
Title: Zuri docs
URL: https://zuri.dev
Tags (comma separated): lang, docs
Saved "Zuri docs".
(a)dd (l)ist (f)ind (r)emove (q)uit > l
1. Zuri docs
https://zuri.dev
[lang, docs]
(a)dd (l)ist (f)ind (r)emove (q)uit > q
Bye.
What You Just Learned
You wrote a program with persistent state, user input, error recovery and a command loop, in about a hundred lines, using nothing that is not in the box.
Along the way you met: var, def, if/else, while, using/when,
catch/as, return, lists, dictionaries, dictionary shorthand, string
interpolation, the conditional operator, anonymous functions, and a handful
of built-in methods.
The next six chapters take every one of those and explain it properly.
Common Programming Concepts
Every language gives you a way to name a value, a set of types those values come in, operators to combine them, and statements to decide what runs next. This chapter is those four things in the form Zuri gives them.
If this is your first language, read the sections in order; each one uses only what came before it. If you already program, read Operators and Control Flow attentively — those are where the same syntax you already know does something different here.
The Reserved Words
Before anything else, the words you cannot use as names. Zuri reserves 31 keywords:
and as assert break catch class const
continue def default do echo else false
for if import in iter nil or
parent raise return self static true using
var when while
Every one is lowercase, and every one is explained somewhere in this book. Appendix A is the index, with a one-line summary and a pointer for each.
Three things are not in that list, and are worth noticing now:
printis not a keyword. It is an ordinary built-in function, and you could shadow it with a variable of your own if you wanted to.echois a keyword.- There is no
function,int,stringorboolkeyword. Type names are ordinary identifiers, which is why they can appear as annotations without being reserved. - There is no
try,finally,switch,case,new,this,publicorprivate. The equivalents arecatch,using,when, calling the class directly,self, and a leading underscore.
Variables and Constants
A variable is a name for a value. In Zuri you introduce one with var,
and from then on the name stands for whatever you last put in it.
var greeting = 'hello'
var count = 0
echo greeting
echo count
hello
0
That is the whole idea. The rest of this section is the detail: what a name
may be called, what happens when you leave the value out, where the name
can be seen, and what const adds.
Declaring
A var with no initializer holds nil, the value that means “nothing
here yet”:
var pending
echo pending
nil
You can declare several names in one statement by separating them with commas, and each one may or may not have a value:
var a = 1, b = 2, c
echo [a, b, c]
[1, 2, nil]
This is worth using when the names genuinely belong together — the three components of a colour, the lower and upper bound of a window — and worth avoiding when they do not, because one long line of declarations is harder to read than three short ones.
var is a declaration, not an expression. It has no value, which is why
you cannot write if var x = f() or echo var y = 1. Declare first, then
use the name.
Naming
A name is made of letters, digits and underscores, and may not start with
a digit. firstName, first_name and _first_name are all legal, and all
three are different names.
Zuri code follows one set of conventions everywhere, and the standard library is written in it:
| Kind of name | Convention | Example |
|---|---|---|
| variable, function, method | snake_case | read_line, total_count |
| class | PascalCase | Account, HttpClient |
| constant | SCREAMING_SNAKE_CASE | MAX_RETRIES |
| private anything | leading underscore | _cache, _retry() |
The leading underscore is the one convention that is not only a
convention. A name beginning with _ is private, and the compiler enforces
it: another file cannot reach module._helper, and code outside a class
cannot reach instance._field. Chapter 6 and
Chapter 8 cover what that means in each case.
The 31 keywords listed at the end of the previous section cannot be used as names at all. Trying produces a syntax error at the point of declaration.
Reassignment
Zuri is dynamically typed. A variable holds a value, not a type, and assigning a different kind of value to the same name is legal:
var value = 42
value = 'now a string'
value = [1, 2, 3]
echo value
[1, 2, 3]
Legal is not the same as advisable. A name that holds a number on one line and a list twenty lines later is a name no reader can predict. Reach for a second variable instead; they are free.
Assignment is an expression, and it evaluates to the value assigned. That lets you chain:
var x, y
x = y = 5
echo '${x} ${y}'
5 5
Compound Assignment
Every arithmetic and bitwise operator has an assignment form, which applies the operator to the variable’s current value and stores the result:
var n = 5
n += 10 # 15
n **= 2 # 225
n //= 7 # 32
n <<= 1 # 64
echo n
64
The full set is +=, -=, *=, /=, //=, **=, %=, &=, |=,
^=, ~=, <<=, >>= and >>>=. Each one means exactly what the
matching binary operator means, applied in place.
Compound assignment works on anything you can assign to, not just plain variables:
var totals = { food: 0 }
var scores = [1, 2, 3]
totals.food += 12
scores[0] *= 10
echo totals
echo scores
{food: 12}
[10, 2, 3]
Increment and Decrement
++ and -- add or subtract one in place:
var n = 5
n++
echo n
n--
echo n
6
5
Two rules to remember. First, both are postfix only. ++n is a syntax
error; the operator goes after the name.
Second, when n++ appears inside a larger expression it updates n and
then evaluates to the new value:
var j = 5
echo j++
echo j
6
6
If you have written C, Java or JavaScript, this is the opposite of what you
expect: there, j++ gives you the old value. In Zuri it gives you the new
one. The habit that avoids the question entirely is to let ++ be a
statement of its own, or to use it in the update clause of an iter loop
where the value is discarded:
iter var i = 0; i < 3; i++ {
echo i
}
0
1
2
const
const declares a name that cannot be reassigned:
def area(radius) {
const PI = 3.141592653589793
return PI * radius * radius
}
echo area(2)
12.566370614359172
Writing to it is caught when the file is compiled, before anything runs:
def f() {
const K = 1
K = 2
}
SyntaxError: cannot assign to constant 'K'
--> /path/to/main.zu:3:3
|
3 | K = 2
| ^
A const must be given a value where it is declared. const K on its own
does not compile, which is the point: a constant with no value would be a
constant nil forever.
The const keyword is enforced in local scopes — inside functions,
methods, and blocks. At the top level of a module, const declares an
ordinary module global and the reassignment check does not apply. Use it
where it does its work, which is inside the code that would otherwise be
tempted to reassign.
Constant Does Not Mean Frozen
const prevents rebinding the name. It says nothing about the value.
A constant list is still a list, and a list is still mutable:
def demo() {
const items = [1, 2]
items.append(3)
echo items
}
demo()
[1, 2, 3]
items = [4] would be an error; items.append(3) is not. If you need a
collection nothing can change, copy it at the boundary where you hand it
out rather than relying on the declaration.
Scope
A block is a scope. Braces open one, and every name declared inside it is gone when the block closes:
var x = 'outer'
{
var x = 'inner'
echo x
}
echo x
inner
outer
The inner x hides the outer one for the length of the block. This is
called shadowing, and it is allowed deliberately: a short block can use
a short name without worrying about what that name means outside it.
Function bodies, loop bodies, if branches and bare { ... } blocks all
work this way. So does the initialiser of an iter loop, whose variable
belongs to the loop:
iter var i = 0; i < 2; i++ {
echo i
}
echo i
0
1
Unhandled UndefinedError: undefined global 'i'
--> /path/to/main.zu:4
After the loop, i is not a variable at all, and reading it is an error.
That error is the good outcome: a name that does not exist is a mistake
worth hearing about, and Zuri never invents a silent nil to paper over
one.
One Declaration Per Scope
Shadowing across scopes is allowed. Declaring the same name twice in one scope is not:
def f() {
var a = 1
var a = 2
}
SyntaxError: 'a' is already declared in this scope
--> /path/to/main.zu:3:7
|
3 | var a = 2
| ^
The exception is the top level of a script or module, where a second
var a rebinds the existing global instead of erroring. That is what lets
you retype a declaration in the REPL while you are experimenting.
Functions Follow the Same Rule
A def is scoped exactly like a var. One written at the top level of a
file binds a module-level name; one written inside a function is a local of
that function, and nothing outside can reach it:
def outer() {
def helper() {
return 'from helper'
}
return helper()
}
echo outer()
echo helper()
from helper
Unhandled UndefinedError: undefined global 'helper'
--> /path/to/main.zu:9
The first call works because helper is a local of outer. The second
fails because, outside outer, no such name was ever created.
Chapter 5 covers this in full.
Data Types
Zuri has ten built-in types. typeof() names the one you are holding:
var values = [1, 1.5, 'text', true, nil, [1], {a: 1}, 1..3, 10n, bytes(2)]
for v in values {
echo typeof(v)
}
number
number
string
bool
nil
list
dict
range
bigint
bytes
On top of those there are function, class, instance, module, file
and ptr, which you get from declaring things rather than from a literal.
number
A number is an IEEE-754 double. There is no separate integer type, which
means 1 and 1.0 are the same value:
echo 1 == 1.0
true
is_int() asks whether a number’s value is integral, not whether it was
written without a decimal point:
echo is_int(1)
echo is_int(1.0)
echo is_int(1.5)
true
true
false
Doubles hold integers exactly up to 2^53. Past that, precision goes:
echo 9007199254740993
9007199254740992
Numeric literals come in five forms:
var decimal = 1_000_000
var binary = 0b1010 # 10
var octal = 0c17 # 15
var hexadecimal = 0xff # 255
var scientific = 6.02e23
echo [decimal, binary, octal, hexadecimal, scientific]
[1000000, 10, 15, 255, 602000000000000000000000]
Underscores are digit separators, and they are accepted in decimal literals
only. 1_000_000 and 1_0.5_5 are fine; 0xdead_beef and 0b1010_1010
are not. Numbers has the full rules.
The special values behave the way the standard says:
echo 1 / 0
echo -1 / 0
echo 0 / 0
inf
-inf
NaN
Every number carries methods, so mathematics reads left to right:
echo 2.sqrt()
echo 16.log2()
echo (-3).abs()
echo 3.7.round()
echo 3.14159.fixed(2)
echo 255.hex()
1.4142135623730951
4
3
4
3.14
ff
A literal takes a method directly; the parentheses around (-3) are there
because a method call binds tighter than the minus sign, so -3.abs()
would negate the result instead of the operand.
Numbers covers the rule in full.
bigint
When 2^53 is not enough, suffix the literal with n and you get an
arbitrary-precision integer:
echo 9007199254740993n + 1n
9007199254740994n
Bigints never lose precision and never overflow. They are slower than
numbers and they do not silently mix with them, so convert explicitly with
to_bigint() and to_number(). Chapter 4 covers
them properly.
bool
true and false. Every value in Zuri is either truthy or falsy,
and that determines what if, while, and, or, ! and ? : do with
it. The falsy values are:
| Value | Why |
|---|---|
false | itself |
nil | absence of a value |
0, 0.0, -0.0 | zero |
NaN | not a number at all |
0n | the bigint zero |
'' | the empty string |
bytes(0) | an empty byte buffer |
Everything else is truthy, including every negative number, [], {}
and '0'.
Two consequences deserve a second look.
A position of 0 is falsy. index_of() returns -1 when it finds
nothing, which is truthy, and 0 when it finds the item first, which is
falsy. So if list.index_of(x) { ... } gets both cases backwards. Compare
the result explicitly:
var haystack = ['a', 'b', 'c']
var position = haystack.index_of('z')
if position == -1 {
echo 'not found'
}
not found
An empty list is truthy. [] and {} are objects, and objects are
truthy. Use is_empty():
var items = []
if items.is_empty() {
echo 'nothing here'
}
nothing here
nil
nil is the absence of a value. An uninitialised var is nil, a
function that falls off the end returns nil, and a missing dictionary key
read through get() gives nil.
nil is falsy, and it still has a to_string():
echo nil.to_string()
nil
Calling any other method on nil raises a TypeError, which is usually
exactly the error you wanted.
string
This is the short version. Strings covers quoting, escapes, interpolation, concatenation, repetition and regular expressions in full.
Strings are written in single or double quotes, with no difference in meaning:
echo 'single'
echo "double"
Both kinds may span multiple lines:
echo 'multi
line'
multi
line
The escape sequences are \0, \a, \b, \f, \n, \r, \t, \v,
\\, \', \", plus \xNN for a byte, \uNNNN for a code point and
\UNNNNNNNN for one outside the basic plane:
echo "tab:\there"
echo "hex: \x41"
echo "emoji: \U0001F600"
tab: here
hex: A
emoji: 😀
A backslash followed by anything else is left alone, backslash and all.
Interpolation
${...} inside a string evaluates the expression and splices the result
in. It works in both quote styles and the expression can be anything:
var name = 'Zuri'
echo 'Hi ${name}, ${1 + 2} and ${name.upper()}'
Hi Zuri, 3 and ZURI
To produce a literal ${, build it by concatenation:
echo 'B: $' + '{x}'
B: ${x}
Strings Are Sequences
Indexing gives you a one-character string, and negative indices count from the end:
var s = 'hello world'
echo s[0]
echo s[0,5]
echo s[-5,]
h
hello
world
s[a,b] is a slice: from a up to but not including b. Either side
may be left out, so s[,3] is the first three characters and s[3,] is
everything from index three onward.
+ concatenates and * repeats:
echo 'ab' + 'cd'
echo 'ab' * 3
abcd
ababab
list
An ordered, growable sequence, written in square brackets. Elements may be of any type, including other lists:
var mixed = [1, 'two', [3], { four: 4 }]
echo mixed.length()
4
Lists index and slice exactly like strings, negative indices included. Chapter 4 covers the forty methods they carry.
dict
An insertion-ordered mapping from keys to values:
var config = { host: 'localhost', port: 8080, debug: true }
A bare word key is taken as a string, so { host: ... } and
{ 'host': ... } are the same dictionary. Keys can also be numbers, and
they can be computed:
var d = { name: 'Ada', 'age': 36, 3: 'three' }
echo d['name']
echo d.name
echo d[3]
Ada
Ada
three
Dot access and bracket access are the same operation. Use the dot when the key is a fixed name, brackets when it is computed or not a valid identifier.
When a key and the variable holding its value share a name, write it once:
var host = 'localhost'
var port = 8080
var config = { host, port }
range
a..b describes the integers from a up to but not including b:
echo 1..5
echo 1..5.to_list()
1..5
[1, 2, 3, 4]
A range is a real value, not loop syntax. You can store one, pass it around, and ask it questions:
var r = 0..10
echo r.lower()
echo r.upper()
echo r.within(7)
0
10
true
bytes
A fixed-size buffer of 8-bit values, for binary data. bytes(n) allocates
n zero bytes; bytes(list) builds one from numbers in 0..256:
var b = bytes([72, 101, 108, 108, 111])
echo b
echo b.to_string()
echo b[0]
(48 65 6c 6c 6f)
Hello
72
Bytes print as hexadecimal in parentheses, which is how you can always tell one from a list at a glance. Chapter 10 is the full treatment.
Checking Types
typeof() gives you a name. The is_* family gives you a boolean, and
there are fifteen of them:
is_bigint is_bool is_bytes is_callable is_class
is_dict is_file is_function is_instance is_int
is_iterable is_list is_number is_object is_string
echo is_string('x')
echo is_callable(print)
echo is_iterable([1, 2])
true
true
true
Use instance_of(value, SomeClass) for classes, which walks the
inheritance chain. Chapter 6 covers that.
Operators
Arithmetic
| Operator | Meaning | Example |
|---|---|---|
+ | addition | 2 + 3 is 5 |
- | subtraction | 5 - 2 is 3 |
* | multiplication | 4 * 3 is 12 |
/ | division, always floating point | 7 / 2 is 3.5 |
// | floor division | 7 // 2 is 3 |
% | remainder | 7 % 3 is 1 |
** | exponentiation | 2 ** 10 is 1024 |
/ never truncates. 7 / 2 is 3.5 even though both operands look like
integers, because there is only one numeric type. When you want the
integer, ask for it with //.
// rounds towards negative infinity, while % keeps the sign of the
left operand:
echo -7 // 2
echo 7 // -2
echo -7 % 3
echo 7 % -3
-4
-4
-1
1
% works on non-integers too: 7.5 % 2 is 1.5.
Comparison
| Operator | Meaning |
|---|---|
== | equal |
!= | not equal |
< <= > >= | ordering, numbers only |
== compares by value for numbers, strings, lists and dictionaries, and by
identity for everything else. A class changes that for its instances with
@eq.
echo 1 == 1.0
echo [1, 2] == [1, 2]
echo {a: 1} == {a: 1}
echo '1' == 1
true
true
true
false
There is no coercion in ==. A string is never equal to a number.
The ordering operators are for numbers. Comparing two strings with <
raises a TypeError; use compare(), which returns -1, 0 or 1:
echo 'abc'.compare('abd')
-1
Logic
Zuri spells these as words:
| Operator | Meaning |
|---|---|
and | true when both sides are truthy |
or | true when either side is truthy |
! | negation |
and and or short-circuit, and they return the operand, not a
boolean:
echo 1 and 2
echo nil or 'fallback'
2
fallback
That is what makes var name = given or 'anonymous' work. Remember that
0, NaN, '' and false are falsy, so this idiom is only safe when
those are not legitimate values.
The Conditional Operator
echo true ? 'yes' : 'no'
yes
cond ? a : b evaluates cond, then exactly one of the branches.
Bitwise
| Operator | Meaning |
|---|---|
& | and |
| | or |
^ | xor |
~ | not |
<< | left shift |
>> | arithmetic right shift, sign preserving |
>>> | logical right shift, zero filling |
echo ~5
echo 5 >>> 1
echo 1 << 10
-6
2
1024
Bitwise operators work on the integer value of a number, and on bigints.
Concatenation and Repetition
+ on a string concatenates, and a number on either side is converted:
echo 'n=' + 5
n=5
+ on a list concatenates; * repeats:
echo [1, 2] + [3]
echo [1, 2] * 2
[1, 2, 3]
[1, 2, 1, 2]
Dictionaries do not support +. Use extend().
Anything the language does not define raises a TypeError that names the
exact signature:
catch {
echo nil + 1
} as e {
echo e.message
}
operator '+' not defined for call signature (nil, number)
Classes can define what + and every other operator mean for their own
instances. That is Chapter 6.
Member, Index and Slice
| Syntax | Meaning |
|---|---|
x.name | member of an object, dictionary or module |
x[key] | index with a computed key |
x[a, b] | slice from a up to but not including b |
x[, b] | slice from the start |
x[a, ] | slice to the end |
a..b | a range value |
Negative indices count back from the end, for both strings and lists.
Precedence
From tightest to loosest. Everything on one row binds equally and
associates left to right, except **, which associates right to left.
| Level | Operators |
|---|---|
| 1 | literals, (...), [...], {...}, self, parent |
| 2 | .. |
| 3 | . () [] |
| 4 | ++ -- |
| 5 | ** |
| 6 | ! - ~ (unary) |
| 7 | * / // % |
| 8 | + - |
| 9 | << >> >>> |
| 10 | & |
| 11 | ^ |
| 12 | | |
| 13 | < <= > >= == != |
| 14 | and |
| 15 | or |
| 16 | ? : |
| 17 | = and every compound assignment |
** follows the mathematical convention. It binds tighter than
multiplication and tighter than a unary operator written before it, and it
groups from the right:
echo 2 * 3 ** 2
echo 2 ** 3 ** 2
echo -2 ** 2
echo 2 ** -1
18
512
-4
0.5
2 * 3 ** 2 is 2 * (3 ** 2), 2 ** 3 ** 2 is 2 ** (3 ** 2), and
-2 ** 2 is -(2 ** 2). The exponent itself may carry a sign, so
2 ** -1 needs no parentheses. Write (-2) ** 2 to raise a negative
number.
.. binds very tightly, to primaries only. 1 + 2..5 parses as
1 + (2..5). Parenthesise any range whose endpoints are expressions.
Operators Zuri Does Not Have
There is no in operator for membership; use contains(). There is no
?. optional chaining and no ?? null coalescing; or covers the common
case. There is no comma operator, and no += on a member that does not
already exist.
Comments and Doc Blocks
Line Comments
# runs to the end of the line:
# Rates are quoted per thousand, not per unit.
var rate = 0.0125 # 1.25%
Block Comments
/* ... */ spans as many lines as you like, and nests:
/* This whole section is off.
/* Including this inner comment. */
Still off. */
Nesting is worth knowing about, because it means commenting out a region
that already contains a block comment works the way you expect, and it also
means an unbalanced /* inside a comment swallows the rest of your file.
Doc Blocks
A block comment that opens with /** is a doc block. The parser keeps
doc blocks in the syntax tree rather than discarding them, which is what
lets the zuri module read a file’s documentation without a separate
parser. Put one directly above the thing it documents:
/**
* Converts a duration in seconds to a human-readable string.
*
* Rounds to the nearest whole second. Durations below one second
* render as `'0s'`.
*
* @param number seconds
* @returns string
*/
def humanize(seconds) {
# ...
}
The tag vocabulary used across the standard library is:
| Tag | Meaning |
|---|---|
@param {type} name: description | one parameter |
@returns type | what the function gives back |
@throws ErrorClass | an error it can raise |
@note | something the caller must know |
@default | the default a parameter falls back to |
Two conventions from the standard library are worth copying. State the default of every optional parameter, and state what happens at the edges: an empty input, a zero length, a value out of range. Anything a caller would otherwise have to discover by experiment belongs in the doc block.
Doc blocks are not only for readers. Because the parser keeps them, a
program can read them: zuri.parse() returns each one as a DocBlock node
sitting immediately before the declaration it documents, which is enough to
build a documentation generator in a few dozen lines.
Chapter 21 shows how.
Commenting Style
A comment earns its place by explaining why, not what. The code already says what it does:
# Bad: restates the code.
# Add one to the counter.
counter++
# Good: explains a decision the code cannot.
# Servers count from one, and the wire protocol has no zero frame.
counter++
Control Flow
Control flow is how a program decides what to run next: run this only when that is true, run this until that stops being true, run this once for every item in a collection. Zuri has seven statements for it, and this section covers all of them.
Blocks and Single Statements
Every control-flow statement takes a body. A body is either a block in braces, or a single statement:
var ready = true
if ready {
echo 'launching'
}
if ready echo 'launching again'
launching
launching again
Both forms are the language. if, else, while, do, for, and a
when branch inside using all accept either one.
iter is the exception. Its body must be a block, because the semicolons
in its header would otherwise be ambiguous with the statement that follows:
iter var i = 0; i < 3; i++ echo i
SyntaxError: Expected '{' at the start of iter body.
The standard library and the examples in this book brace every body except a short
whenbranch. Braces make a one-line body easy to extend into a three-line one later, and they keep the shape of a nested statement obvious at a glance. It is a convention, not a rule — write whichever form reads better in the code you are writing.
if / else
var score = 73
if score >= 90 {
echo 'A'
} else if score >= 70 {
echo 'B'
} else {
echo 'C'
}
B
There is no elif; else if is two keywords, and it works because the
body of an else can itself be an if statement.
The condition is any expression at all, and it is judged by truthiness
rather than by being a bool:
var name = ''
if name {
echo 'have a name'
} else {
echo 'no name'
}
no name
That is convenient and it is also where most bugs in new Zuri code come
from, because a legitimate 0 is falsy. Re-read the table in
Data Types before you write if count or
if position.
The Conditional Expression
? : is the expression form of if. It chooses between two values rather
than between two statements:
var age = 20
var status = age >= 18 ? 'adult' : 'minor'
echo status
adult
It nests, and it can span several lines. When breaking one across lines,
? and : may either end a line or begin the next, whichever reads
better; leading each branch with its operator keeps the shape of the
choice visible:
var n = 7
var size = n > 100
? 'large'
: n > 5
? 'medium'
: 'small'
echo size
medium
Use ? : when you are producing a value and if when you are performing
an action. A conditional expression whose branches are both side effects is
harder to read than the if it replaced.
while
while tests before each pass, so a body may run zero times:
var n = 0
while n < 3 {
echo 'while ${n}'
n++
}
while 0
while 1
while 2
do / while
do runs the body once and then tests, so the body always runs at least
once. Reach for it when the test depends on something the body produces:
var m = 10
do {
echo 'do ${m}'
m++
} while m < 3
do 10
The condition was false from the very start, and the body still ran.
iter
iter is the counting loop. Its header has three clauses separated by
semicolons: an initialiser, a condition, and an update.
iter var i = 0; i < 3; i++ {
echo 'iter ${i}'
}
iter 0
iter 1
iter 2
The initialiser runs once. The condition is tested before every pass. The update runs after every pass. A variable declared in the initialiser belongs to the loop and does not exist after it.
The initialiser may declare several variables and the update may hold several expressions, each list separated by commas. The updates run in the order they are written:
iter var i = 0, j = 10; i < 3; i++, j-- {
echo '${i} ${j}'
}
0 10
1 9
2 8
Every clause is optional:
var k = 0
iter ; k < 2; {
echo 'bare ${k}'
k++
}
bare 0
bare 1
iter ; ; { ... } loops forever, though while true says it more plainly.
The header may be spread over several lines, which is worth doing when the condition is long:
iter var i = 1;
i <= 3;
i++
{
echo i
}
1
2
3
for / in
for walks an iterable: a list, a dictionary, a string, a range, a byte
stream, or any class that implements the iterator protocol.
With one variable you get the value:
for ch in 'hey' {
echo ch
}
h
e
y
With two, you get the key first and the value second:
for index, value in ['a', 'b'] {
echo '${index} ${value}'
}
for key, value in { x: 1, y: 2 } {
echo '${key}=${value}'
}
0 a
1 b
x=1
y=2
For a list the key is the index, for a dictionary it is the key, and for a string it is the character position. Dictionaries iterate in insertion order, and so does everything else with an order to preserve.
Ranges are exclusive at the top:
for i in 0..3 {
echo i
}
0
1
2
The iterable expression is evaluated exactly once, before the first pass. Calling a function in that position calls it once, not once per element:
def source() {
echo 'source() called'
return [1, 2, 3]
}
for n in source() {
echo n
}
source() called
1
2
3
Chapter 6 shows how to make your own class
work with for by defining @key() and @value().
break and continue
continue skips to the next pass, and break leaves the loop entirely:
iter var i = 0; i < 5; i++ {
if i == 1 {
continue
}
if i == 3 {
break
}
echo i
}
0
2
Both apply to the innermost enclosing loop, and there are no loop labels. To leave a nested loop you either carry a flag:
var found = false
iter var i = 1; i < 4; i++ {
iter var j = 1; j < 4; j++ {
if i * j == 4 {
echo 'found ${i} x ${j}'
found = true
break
}
}
if found {
break
}
}
found 2 x 2
or, more often, put the nested loop in a function and return out of it:
def first_product(target) {
iter var i = 1; i < 4; i++ {
iter var j = 1; j < 4; j++ {
if i * j == target {
return [i, j]
}
}
}
return nil
}
echo first_product(4)
echo first_product(99)
[2, 2]
nil
The second version says what it is looking for in its name, and it has no flag to keep in sync. Prefer it.
using / when
using compares one subject against several candidate values:
var grade = 'B'
using grade {
when 'A' echo 'excellent'
when 'B', 'C' echo 'good'
default echo 'try again'
}
good
Several values on one when are alternatives, not a sequence: this when
matches 'B' or 'C'. The first branch that matches runs, and then the
statement is over — there is no fall-through, and no break to remember.
Matching uses the same equality as ==, which means using works on
numbers, strings and booleans, and compares instances by identity, or
through @eq when their class defines one.
default is optional. When nothing matches and there is no default,
nothing happens:
using 99 {
when 1 echo 'one'
}
echo 'nothing ran, and that is not an error'
nothing ran, and that is not an error
A when branch takes a block when it needs one:
var command = 'q'
using command {
when 'a' echo 'adding'
when 'q' {
echo 'saving'
echo 'goodbye'
}
default echo 'unknown command'
}
saving
goodbye
using is the right shape whenever you are dispatching on one value to
many outcomes. A chain of else if comparing the same variable over and
over is the same thing written longer.
assert
assert checks a condition and raises an AssertError when it is false:
var items = [1]
assert 1 + 1 == 2
assert items.length() == 1, 'list should have one item'
echo 'both assertions held'
both assertions held
The optional second argument is the message the error carries:
assert 1 == 2, 'arithmetic is broken'
Unhandled AssertError: arithmetic is broken
Assertions state what you believe is already true. They are for conditions
that should be impossible, not for checking input that might legitimately
be wrong — raise a ValueError for that.
Chapter 7 draws the line in detail.
return
return leaves the current function, with a value if you give it one and
nil if you do not:
def classify(n) {
if n < 0 {
return 'negative'
}
if n == 0 {
return 'zero'
}
return 'positive'
}
echo classify(-5)
echo classify(0)
echo classify(3)
negative
zero
positive
A function that reaches its closing brace without a return yields nil.
There are no implicit returns; the last expression in a body is not its
value.
return is only legal inside a function. At the top level of a script it
is a syntax error, because there is nothing to return from.
Text, Numbers and Collections
Chapter 3 introduced the built-in types and showed what the operators do to them. This chapter is the working manual for the five you will reach for every day: strings, numbers, lists, dictionaries and ranges.
One idea shapes all five. Zuri puts behaviour on the value itself.
There is no len(x) and no str.upper(x); there is x.length() and
x.upper(). That holds for numbers too, which carry the whole of the
mathematics library as methods:
echo 'zuri'.upper()
echo [3, 1, 2].sort()
echo 16.sqrt()
echo 255.hex()
ZURI
[1, 2, 3]
4
ff
Because the methods live on the value, they chain, and a chain reads left to right in the order the work happens:
var line = ' Ada, Grace , Alan '
echo line.trim().split(',').map(@(name) => name.trim()).length()
3
Each section here covers the methods worth knowing by name, the ones with a behaviour you would not guess, and the mistakes that come up most often. Appendix E is the complete list, generated from the same documentation the runtime ships.
Strings
A Zuri string is UTF-8 text. Indexing, slicing and length() work in
characters, not bytes, so a string of emoji behaves the way a reader
expects.
Text is the type you will write most, so this section starts with how you write one down — quoting, escapes and interpolation — before moving on to what you can do with it.
Writing a String
A string literal is written between single or double quotes:
echo 'single quotes'
echo "double quotes"
single quotes
double quotes
The two styles are identical. They are not the two different things they are in some other languages. Both process escape sequences, both interpolate expressions, and both may span several lines. Nothing at all changes except which quote character ends the string:
var name = 'Zuri'
echo 'escape \t and interpolate ${name}'
echo "escape \t and interpolate ${name}"
escape and interpolate Zuri
escape and interpolate Zuri
Pick whichever avoids escaping. A string containing an apostrophe is easiest in double quotes; a string containing a quotation mark is easiest in single quotes:
echo "it's clearer this way"
echo 'she said "hello"'
it's clearer this way
she said "hello"
Most Zuri code, and all of the standard library, uses single quotes by default and switches to double quotes when the content calls for it.
Multi-line Strings
A literal may simply contain newlines. There is no separate triple-quoted form and no continuation character:
var message = 'Dear reader,
Thank you for the strings.'
echo message
Dear reader,
Thank you for the strings.
Everything between the quotes is part of the string, indentation included.
When a long string is indented inside a function, the leading spaces on
each line are in the string too, which is why it is advisable that code
that builds indented text should assemble it with join() rather than
writing one long literal.
Escape Sequences
A backslash starts an escape. The full set is:
| Escape | Produces |
|---|---|
\0 | the null character, U+0000 |
\a | the alert (bell) character, U+0007 |
\b | backspace, U+0008 |
\t | horizontal tab, U+0009 |
\n | line feed, U+000A |
\v | vertical tab, U+000B |
\f | form feed, U+000C |
\r | carriage return, U+000D |
\\ | a single backslash |
\' | a single quote |
\" | a double quote |
\xNN | the byte with hexadecimal value NN |
\uNNNN | the code point U+NNNN |
\UNNNNNNNN | the code point U+NNNNNNNN |
The three numeric forms take exactly the number of hex digits shown — two, four and eight — and the digits may be upper or lower case:
echo 'hex byte: \x41'
echo 'code point: \u00e9'
echo 'astral: \U0001F600'
hex byte: A
code point: é
astral: 😀
\u covers everything up to U+FFFF. Anything above that — emoji, most
historic scripts, mathematical alphanumerics — needs \U and its eight
digits. Zero-pad to fill them.
Unknown Escapes Are Left Alone
A backslash followed by anything not in the table above is not an error, and it is not silently dropped either. The backslash and the character both survive:
echo 'a regex: \d+ and \w+'
echo 'a path: C:\temp\notes'
a regex: \d+ and \w+
a path: C: emp
otes
The first line came through untouched: \d and \w are not escapes, so
both backslashes survived, which is exactly why regular expressions can be
written as plain strings.
The second line did not. \t is an escape, so C:\temp became C:
followed by a tab and then emp; \n is an escape too, so \notes became
a newline followed by otes. One path, two silent substitutions. The rule
is easy to state and easy to forget: unknown escapes pass through, known
ones do not, and nothing in the way you write them tells you which is
which.
For Windows paths and regular expressions — the two places this bites — double the backslashes:
echo 'a path: C:\\temp\\notes'
a path: C:\temp\notes
Escaping the Quote You Chose
You only need to escape the quote character that would end the string:
echo 'it\'s escaped'
echo "she said \"hi\""
it's escaped
she said "hi"
Escaping the other quote is harmless but pointless, and because an unnecessary escape is an unknown escape, the backslash stays in the result:
echo 'this \" keeps its backslash'
this \" keeps its backslash
That is a good reason to choose the quote style that lets you write the text plainly and escape nothing at all.
Interpolation
${...} inside a string evaluates the expression between the braces and
splices its text into the result:
var name = 'Ada'
var year = 1843
echo 'In ${year}, ${name} wrote the first program.'
In 1843, Ada wrote the first program.
This works in both quote styles, and it is the way most Zuri code builds a
string. The section after next covers +, which does the same job with
more punctuation.
Any Expression Fits
The braces take a whole expression, not just a variable name. Arithmetic, method calls, indexing, comparisons and conditionals all work:
var items = ['a', 'b', 'c']
var user = { name: 'ada', admin: true }
echo 'count: ${items.length()}'
echo 'first: ${items[0].upper()}'
echo 'math: ${(2 ** 10) - 24}'
echo 'role: ${user.admin ? 'admin' : 'user'}'
echo 'name: ${user.name.capitalize()}'
count: 3
first: A
math: 1000
role: admin
name: Ada
Notice the third line. The quotes inside ${...} are the same quote
character that opened the string, and it still parses correctly, because
the expression inside the braces is lexed as its own piece of code rather
than as text. You do not need to switch quote styles or escape anything to
put a string literal inside an interpolation.
Interpolations Nest
A string literal inside ${...} is a string literal like any other, which
means it may contain interpolations of its own, to any depth:
var n = 1
echo 'outer ${ 'inner ${n + 1}' }'
echo '${ '${ '${2 ** 3}' }' }'
outer inner 2
8
This is not a curiosity. It is what makes a conditional inside a string able to produce formatted text rather than a bare word, which comes up constantly in messages meant for a person to read:
var user = { name: 'ada', unread: 3 }
echo 'Hi ${user.name.capitalize()}, you have ${user.unread > 0 ? '${user.unread} new ${user.unread == 1 ? 'message' : 'messages'}' : 'nothing new'}.'
Hi Ada, you have 3 new messages.
Three levels of interpolation are at work there: the outer message, the
branch that chooses between a count and “nothing new”, and the branch that
picks the singular or the plural. Each one is an ordinary string literal in
an ordinary expression, and each one uses single quotes, because the
expression inside ${...} is lexed as code and its quotes never collide
with the ones around it.
Nesting is equally useful inside a callback, where the inner string is building one element of a larger result:
var items = ['a', 'b']
echo 'list: ${', '.join(items.map(@(i) => '<${i}>'))}'
list: <a>, <b>
Depth costs readability quickly. When a line like the message above stops being scannable, lift the inner pieces into local variables on the lines before it; the language is happy either way.
What Each Type Looks Like
Interpolation renders a value the same way echo does:
var settings = { a: 1 }
echo 'number: ${1 / 3}'
echo 'list: ${[1, 2]}'
echo 'dict: ${settings}'
echo 'nil: ${nil}'
echo 'bool: ${true}'
number: 0.3333333333333333
list: [1, 2]
dict: {a: 1}
nil: nil
bool: true
A dictionary literal written directly inside ${...} does not parse,
because its closing brace runs into the interpolation’s closing brace:
echo 'dict: ${{ a: 1 }}'
SyntaxError: Expected '}' after dictionary
Put it in a variable first, as above. A list literal has no such problem,
since ] and } are different characters.
There is one important exception, and it catches everyone once.
Interpolating a class instance does not call its to_string() method:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
to_string() {
return '(${self.x}, ${self.y})'
}
}
var p = Point(3, 4)
echo 'implicit: ${p}'
echo 'explicit: ${p.to_string()}'
implicit: <instance of Point>
explicit: (3, 4)
Call the method yourself. This applies to 'text ' + p too. Interpolation
and + never call a method on your behalf to produce text; if you want
to_string() to run, write it. echo is the exception, through
@to_string(), which Decorated Methods
covers.
Writing a Literal ${
A backslash does not escape an interpolation. \${x} is an unknown escape,
so the backslash survives and the interpolation still does not happen —
which is rarely what anyone wants:
echo 'literal attempt: \${x}'
literal attempt: \${x}
To produce the two characters ${ followed by text, split the string so
the $ and the { are never adjacent inside one literal:
echo 'shell syntax: $' + '{HOME}'
shell syntax: ${HOME}
A lone $ is not special, so only the exact sequence ${ needs this
treatment:
echo 'price: $5.00'
echo 'total: $${12 + 8}'
price: $5.00
total: $20
Joining Strings Together
Concatenation with +
+ joins two strings end to end:
echo 'Hello, ' + 'world'
echo 'a' + 'b' + 'c'
Hello, world
abc
When one side is a string, the other side is converted to its text form first. This works in both directions, and for every built-in type:
echo 'n=' + 5
echo 5 + 'n'
echo 'pi is about ' + 3.14
echo 'flag: ' + true
echo 'missing: ' + nil
echo 'items: ' + [1, 2]
n=5
5n
pi is about 3.14
flag: true
missing: nil
items: [1, 2]
Read the second line again: 5 + 'n' gives '5n', not an error and not
arithmetic. A string on either side turns the whole expression into
concatenation. That is worth knowing when a value arrives from somewhere
you do not control:
var quantity = '2'
echo quantity + 3
echo quantity.to_number() + 3
23
5
The first line is text, the second is arithmetic, and nothing in the expression tells you which you are getting. When a number must be a number, convert it at the point it enters your program rather than hoping.
The one type that does not convert usefully is a class instance:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
to_string() {
return '(${self.x}, ${self.y})'
}
}
echo 'at ' + Point(1, 2).to_string()
at (1, 2)
Concatenating the instance directly would produce <instance of Point>,
for the same reason interpolation does: Zuri does not call to_string()
for you.
Concatenation Versus Interpolation
Both produce the same string, so choose on readability:
var host = 'localhost'
var port = 8080
echo 'connecting to ' + host + ':' + port + '/health'
echo 'connecting to ${host}:${port}/health'
connecting to localhost:8080/health
connecting to localhost:8080/health
Interpolation wins whenever there is more than one hole to fill: the
literal text stays in one piece, the quotes do not multiply, and there is
no chance of losing a separator between two + signs. Reach for + when
you are gluing exactly two pieces together, or when one of them is already
a variable holding a complete string.
Repetition with *
* repeats a string a whole number of times:
echo 'ab' * 3
echo '-' * 20
echo ' ' * 2 + 'indented'
ababab
--------------------
indented
That second line is the reason this operator earns its place. Separators, rules, indentation and padding are all one short expression instead of a loop.
The rules for the count are worth stating exactly:
- A count of
1returns the string unchanged. - A count of
0returns the empty string. - A negative count also returns the empty string, rather than raising.
- A fractional count is truncated toward zero, so
* 2.7repeats twice.
echo '[' + 'ab' * 1 + ']'
echo '[' + 'ab' * 0 + ']'
echo '[' + 'ab' * -1 + ']'
echo '[' + 'ab' * 2.7 + ']'
[ab]
[]
[]
[abab]
The negative case is the one to watch. '-' * (width - label.length())
silently produces nothing when the label is longer than the width, and no
error tells you so. When the count is computed, clamp it:
def rule(width, label) {
var padding = (width - label.length()).max(0)
return label + '-' * padding
}
echo rule(10, 'ab')
echo rule(10, 'a much longer label')
ab--------
a much longer label
Unlike +, repetition does not commute. The string has to be on the left:
echo 3 * 'ab'
Unhandled TypeError: operator '*' not defined for call signature (number, string)
Building a String from Many Pieces
For more than a handful of pieces, neither operator is the right tool.
Collect them in a list and join():
var names = ['ada', 'grace', 'alan']
echo ', '.join(names)
echo '\n'.join(names.map(@(n) => '- ' + n.capitalize()))
ada, grace, alan
- Ada
- Grace
- Alan
join() lives on the separator, not on the list, and it converts each
element to text on the way — so a list of numbers joins as readily as a
list of strings:
echo '-'.join([1, 2, 3])
1-2-3
Note also that join() puts the separator between elements, never at
either end, which is exactly the behaviour a hand-written loop usually gets
wrong on the last item.
Inspecting
var s = 'Hello, Zuri'
echo s.length()
echo s.is_empty()
echo s.index_of('Zuri')
echo s.count('l')
echo s.starts_with('Hello')
echo s.ends_with('Zuri')
echo s.contains('Zuri')
11
false
7
2
true
true
true
index_of() returns -1 when there is no match, and takes an optional
second argument to start the search from.
last_index_of() searches from the other end, which is what you want
whenever the interesting separator is the final one:
var path = 'src/vm/value.rs'
echo path.index_of('/')
echo path.last_index_of('/')
echo path[path.last_index_of('/') + 1, path.length()]
echo path.last_index_of('\\')
3
6
value.rs
-1
Both take a second argument, and in both it bounds where a match may
begin. That makes the pair split a string at one index: for any n,
index_of(str, n) finds the first match at or after n and
last_index_of(str, n) the last match at or before it.
There is a family of character-class predicates, each true when every character in the string qualifies:
echo 'abc'.is_alpha()
echo '123'.is_number()
echo 'ABC'.is_upper()
echo ' '.is_space()
true
true
true
true
The full set is is_alpha, is_alnum, is_number, is_lower,
is_upper, is_space, plus is_empty.
Changing Case
var s = 'Hello, Zuri'
echo s.upper()
echo s.lower()
echo s.capitalize()
echo s.title()
HELLO, ZURI
hello, zuri
Hello, zuri
Hello, Zuri
capitalize() uppercases the first character and lowercases the rest.
title() does that to every word.
For comparing text from users, reach for case_fold() rather than
lower(). Case folding is the Unicode operation defined for
case-insensitive matching, and it handles cases simple lowercasing gets
wrong:
echo 'Straße'.case_fold()
strasse
Trimming and Padding
echo '[' + ' pad '.trim() + ']'
echo '[' + ' pad '.ltrim() + ']'
echo '[' + ' pad '.rtrim() + ']'
echo 'xxhixx'.trim('x')
echo '-=hi=-'.trim('-=')
[pad]
[pad ]
[ pad]
hi
hi
With no argument, all three strip whitespace: spaces, tabs, newlines,
carriage returns, vertical tabs and form feeds. Given a string, they strip
every character in it, in any order, so '-=hi=-'.trim('-=') is hi.
echo 'x'.lpad(5, '.')
echo 'x'.rpad(5, '.')
....x
x....
lpad and rpad take a target width and an optional fill character that
defaults to a space. A string already at or over the width comes back
unchanged.
Splitting and Joining
echo 'Hello, Zuri'.split(', ')
echo ', '.join(['a', 'b', 'c'])
echo 'abc'.to_list()
echo 'line1\nline2'.lines()
[Hello, Zuri]
a, b, c
[a, b, c]
[line1, line2]
join lives on the separator, not on the list. to_list() splits into
single characters. lines() splits on newlines and handles both \n and
\r\n.
Replacing
echo 'Hello, Zuri'.replace('Zuri', 'World')
Hello, World
Every occurrence is replaced. The pattern may be a regular expression, and
the replacement can refer back to capture groups with $1, $2 or
${1}:
echo 'John Smith'.replace('/(\w+) (\w+)/', '$2, $1')
Smith, John
The optional third argument turns regex handling off, so a pattern that would otherwise be read as a regex is matched literally, delimiters included:
echo 'a/b/c'.replace('/b/', 'X')
echo 'a/b/c'.replace('/b/', 'X', false)
a/X/c
aXc
When the replacement depends on what was matched, use replace_with(),
which calls a function for each match:
echo 'a1b2'.replace_with('/[0-9]/', @(m) => '<' + m + '>')
a<1>b<2>
Regular Expressions
Zuri’s regular expressions are PCRE2, the same engine that powers regular expressions in PHP, R and a long list of other tools. It is the real library rather than a subset of it, so named groups, backreferences, lookahead, lookbehind, atomic groups, possessive quantifiers, conditionals and recursion all work exactly as PCRE2 documents them.
This section covers how a pattern is written in Zuri, which methods accept one, what each modifier does, and the handful of places where Zuri’s surface differs from what you may expect. It does not teach regular expression syntax itself; any PCRE or Perl reference applies unchanged.
Writing a Pattern
A pattern is an ordinary string whose first character is a non-word character, repeated to close the pattern, with any modifier letters after the closing delimiter:
'/[a-z]+/'
'/[a-z]+/i'
'#\d{3}#'
'!https?://\S+!i'
A “non-word character” is anything that is not a letter, a digit or an underscore. Forward slash is the conventional choice, and every example in this book uses it, but the delimiter is yours to pick — which is worth doing when the pattern itself is full of slashes:
echo 'see https://example.com/docs now'.match('!https?://\S+!')
{0: https://example.com/docs}
Because the pattern is a string, the backslashes in it are subject to
Zuri’s own escape rules first. This is almost never a problem, because
\d, \w, \s, \b and \S are not escape sequences and therefore pass
through untouched. The exceptions are the letters that are escapes —
\t, \n, \r, \f, \v, \a, \b, \0, \x and \u. Of those,
\t and \n mean the same thing to both, so they are harmless; \b
(word boundary in a pattern, backspace in a string) is the one that will
catch you:
echo 'one two'.match('/\btwo\b/')
echo 'one two'.match('/\\btwo\\b/')
false
{0: two}
Double the backslash whenever the pattern needs \b.
A Plain String Is a Literal
Every method in this section also accepts a plain string, and treats it as a literal substring rather than a pattern:
echo 'a.b'.replace('.', 'X')
echo 'a.b'.replace('/./', 'X')
aXb
XXX
The first call replaced the single full stop. The second compiled . as a
pattern, where it means “any character”, and replaced all three.
That distinction is decided purely by whether the string looks like a
delimited pattern. When you have text from elsewhere that might
accidentally look like one, replace() takes a third argument that turns
pattern handling off:
echo 'a/b/c'.replace('/b/', 'X')
echo 'a/b/c'.replace('/b/', 'X', false)
a/X/c
aXc
With false, the delimiters are matched as characters like anything else.
Which Methods Take a Pattern
Five, and only five:
| Method | What it does with the pattern |
|---|---|
match(pattern) | the first match and its groups, or false |
matches(pattern) | every match, grouped by capture group |
replace(pattern, with, use_regex) | replaces every match |
replace_with(pattern, callback) | replaces every match with a call |
split(pattern) | splits on every match |
Everything else — contains(), index_of(), count(), starts_with(),
ends_with(), trim() — takes a literal string only. A
regex-looking argument passed to one of those is matched as the literal
characters it contains:
echo 'a1b'.contains('/[0-9]/')
echo 'a1b'.index_of('/[0-9]/')
echo 'a1b'.count('/[0-9]/')
false
-1
0
All three answered about the six-character string /[0-9]/, which is not
in a1b. When you need a pattern there, reach for match() instead.
match() and matches()
match() finds the first match and returns it as a dictionary under
key 0. It returns false — not nil — when nothing matches:
echo 'one two'.match('/\w+/')
echo 'nope'.match('/zzz/')
echo typeof('nope'.match('/zzz/'))
{0: one}
false
bool
Both are falsy, so if s.match(p) { ... } reads correctly. An equality
test against nil does not, and this is the single most common regex
mistake in Zuri code.
The dictionary holds the pattern’s capture groups as well, each under its
number, and a named group under its name too. A group that takes no part in
the match is nil:
echo '2024-01'.match('/(?<year>\d+)-(\d+)/')
echo 'ab'.match('/([a-z]+)(\d+)?/')
{0: 2024-01, 1: 2024, 2: 01, year: 2024}
{0: ab, 1: ab, 2: nil}
matches() finds every match and groups the results by capture group.
Key 0 holds the whole matches, key 1 the first group’s, and so on —
each as a list, parallel across the keys:
echo 'a1b2c3'.matches('/[0-9]/')
echo 'a1b2'.matches('/([a-z])([0-9])/')
{0: [1, 2, 3]}
{0: [a1, b2], 1: [a, b], 2: [1, 2]}
In the second result, match zero is a1 with groups a and 1, and
match one is b2 with groups b and 2. Read down the lists, not
across.
When nothing matches, matches() returns {0: []} rather than false.
The two methods differ here, so test matches() with is_empty() on key
zero:
var found = 'nope'.matches('/zzz/')
echo found
echo found[0].is_empty()
{0: []}
true
Capture Groups in a Replacement
replace() refers to a capture group with $1, $2 and so on:
echo 'John Smith'.replace('/(\w+) (\w+)/', '$2, $1')
Smith, John
Do not write ${1}. The braced form is Zuri’s own string
interpolation, and it is consumed before replace() ever sees it. The
expression 2 evaluates to the number two, so the replacement string is
already '2, 1' by the time it arrives:
echo 'John Smith'.replace('/(\w+) (\w+)/', '${2}, ${1}')
2, 1
Use the bare $1 form throughout. This is one of the few places where
Zuri’s interpolation and another language’s syntax collide, and it fails
quietly rather than loudly.
Named Groups
Named groups are supported in the pattern, because PCRE2 supports them:
echo 'John Smith'.matches('/(?<first>\w+) (?<last>\w+)/')
{0: [John Smith], 1: [John], 2: [Smith]}
The results are keyed by position, not by name. A name in a pattern
documents the group and lets a backreference such as \k<first> refer to
it; it does not change how the result is keyed. Count the groups to find
the one you want.
Computing the Replacement
When the replacement depends on what matched, replace_with() calls a
function for each match:
echo 'a1b2'.replace_with('/[0-9]/', @(m) => '<' + m + '>')
a<1>b<2>
The callback receives the whole match first, then one argument per capture group, then the offset of the match, then the whole subject string. Take only the arguments you need, since extra arguments are dropped:
echo 'aXbXc'.replace_with('/X/', @(match, offset) => '[${offset}]')
a[1]b[3]c
Splitting on a Pattern
split() accepts a pattern, which is how you split on “one or more of
something” rather than on an exact separator:
echo 'a1b22c'.split('/[0-9]+/')
echo 'a, b,c , d'.split('/\s*,\s*/')
[a, b, c]
[a, b, c, d]
The Modifiers
Modifier letters go after the closing delimiter, in any order and any
combination. /[a-z]+/mi is both multi-line and case-insensitive.
| Modifier | Effect |
|---|---|
i | Case-insensitive matching. |
m | Multi-line: ^ and $ match at every line boundary, not only at the start and end of the subject. |
s | Dot-all: . matches a newline as well as everything else. |
x | Extended: unescaped whitespace in the pattern is ignored, and # starts a comment to end of line. |
u | Unicode properties: \d, \w, \s and friends become Unicode-aware instead of ASCII-only. |
U | Ungreedy: quantifiers become lazy by default, and ? after one makes it greedy. |
A | Anchored: a match is only accepted if it begins exactly where the search started. |
J | Allow two capture groups in one pattern to share a name. |
Each one in use:
echo 'HELLO'.match('/hello/i')
echo 'a\nb'.matches('/^./m')
echo 'a\nb'.match('/a.b/s')
echo 'abc'.match('/ a b c /x')
echo 'héllo'.matches('/\w/u')
echo 'aaa'.match('/a+?/U')
echo 'abc'.match('/abc/A')
echo 'xabc'.match('/abc/A')
{0: HELLO}
{0: [a, b]}
{0: a
b}
{0: abc}
{0: [h, é, l, l, o]}
{0: aaa}
{0: abc}
false
Read the last two together: the pattern matched when it was at the start of
the subject and failed when it was not, which is what A is for. Without
it, the second would have matched at offset one.
u is worth a second look as well. Without it, \w is ASCII-only and
é is not a word character; with it, the Unicode properties apply and it
is. Matching itself is always per-character rather than per-byte, so
indexes and offsets are correct on multi-byte text regardless of u.
The One Exception to PCRE2 Compatibility
PCRE2’s D (dollar-endonly) modifier has no effect in Zuri. $
always matches before a trailing newline as well as at the absolute end of
the subject, and there is no way to change that. The letter is accepted
rather than rejected, so a pattern carrying it compiles and runs; it simply
does not do anything.
The same is true of any other unrecognised modifier letter: it is accepted and ignored rather than raising. A typo in a modifier is therefore silent, which is worth remembering when a pattern behaves as though a flag you wrote is not set.
Everything else in the table above is supported and behaves exactly as PCRE2 documents it.
Invalid Patterns
A pattern that PCRE2 cannot compile raises, naming the pattern and the engine’s own explanation:
catch {
echo 'text'.match('/(unclosed/')
} as e {
echo e.message
}
invalid regular expression '(unclosed': PCRE2: error compiling pattern at offset 9: missing closing parenthesis
Patterns are compiled once and reused, so repeating the same pattern in a loop costs nothing after the first time through.
Conversion
echo '5'.to_number() + 1
echo 'ff'.to_number(16)
echo 'H'.ord()
echo 72.chr()
echo 'Hi'.to_bytes()
6
255
72
H
(48 69)
to_number() takes an optional base. ord() requires a single-character
string and gives its code point; chr() on a number goes the other way.
to_number() Never Fails
This is the one conversion behaviour worth memorising. to_number() does
not raise, and it does not produce NaN. It reads the first number
written in the string, and text with no number in it at all becomes
zero:
echo '42'.to_number()
echo '3.5'.to_number()
echo '-5'.to_number()
echo 'eighty'.to_number()
echo ''.to_number()
echo '12abc'.to_number()
echo ' 7 '.to_number()
echo '96.3 of 31'.to_number()
42
3.5
-5
0
0
12
7
96.3
Look at the last three. What surrounds the number is ignored, so '12abc'
is 12, the spaces around ' 7 ' are simply not part of it, and only
the first number is read however many follow.
That cuts both ways. '3 apples' giving 3 is usually what was meant;
'2026-09-16' giving 2026 usually is not. A number is taken to start at
a digit, or at a sign or decimal point directly in front of one, so '-5'
is negative five and 'a - 42', where the sign stands apart, is 42.
There is no error to catch and no sentinel to test, so a form field a user
left blank and a form field they filled with '0' produce the same number,
and so do 'eighty' and '0'. When the difference matters, check the text
before converting it:
def to_count(text) {
var trimmed = text.trim()
if !trimmed.match('/^\d+$/') {
raise ValueError('not a whole number: ${text}')
}
return trimmed.to_number()
}
echo to_count(' 12 ')
catch {
to_count('twelve')
} as e {
echo e.message
}
12
not a whole number: twelve
is_number() is a narrower test than it looks — it is true only for a
string of digits, so '3.5' and '-5' both fail it. A regular expression
is the reliable check for anything beyond unsigned integers.
Walking a String
A string is iterable, one character at a time, and the same four forms available for a list work here.
for, One Character at a Time
for ch in 'hey' {
echo ch
}
h
e
y
This is the form to reach for by default. Each ch is a one-character
string, not a code point number — use ord() when you want the number.
for With Two Variables, for the Position
for index, ch in 'hey' {
echo '${index}: ${ch}'
}
0: h
1: e
2: y
The key comes first and the value second, and for a string the key is the character position.
iter, When the Position Drives the Walk
var s = 'hey'
iter var i = s.length() - 1; i >= 0; i-- {
echo s[i]
}
y
e
h
s[i] indexes by character, not by byte, so this is correct on text that
is not ASCII. Use iter when you need to skip, step backwards, or look at
s[i + 1] from inside the body.
each(), With a Function
'Hi'.each(@(ch, index) {
echo '${index}:${ch}'
})
0:H
1:i
As everywhere else, the callback receives the value first and the index
second — the opposite order to for.
each_line() and lines(), for Text in Lines
var doc = 'first\nsecond\nthird'
doc.each_line(@(line, index) {
echo '${index}: ${line}'
})
echo doc.lines()
0: first
1: second
2: third
[first, second, third]
each_line() calls the function once per line; lines() hands you the
whole list so you can for over it, filter it, or count it. Both split on
\n and \r\n, so both read a file written on any platform.
Ordering and Comparison
== compares two strings by value, which is almost always what you want:
echo 'abc' == 'abc'
echo 'abc' == 'ABC'
true
false
The ordering operators are a different story. <, <=, > and >= are
defined for numbers only, and comparing two strings with one raises a
TypeError. Use compare(), which returns -1, 0 or 1:
echo 'abc'.compare('abd')
echo 'abc'.compare('abc')
echo 'abd'.compare('abc')
-1
0
1
That three-way result is exactly the shape a sort comparison wants, and it compares by code point, so it is stable and locale-independent.
Strings Are Immutable
Every method in this section returns a new string. Nothing mutates in place:
var original = 'hello'
var shouted = original.upper()
echo original
echo shouted
hello
HELLO
This is worth internalising, because it is the source of the single most common string mistake:
var name = ' ada '
name.trim()
echo '[${name}]'
name = name.trim()
echo '[${name}]'
[ ada ]
[ada]
The first trim() produced a trimmed string and threw it away. Assign the
result, or chain onto it.
Immutability also means you can pass a string anywhere without copying it defensively. Nothing you hand a string to can change the one you still hold.
A Worked Example
Here is a small parser that puts most of this section together. It takes a
block of key = value configuration text and produces a dictionary,
skipping blank lines and comments, and tolerating whatever spacing the
author used.
def parse_config(text) {
var config = {}
for line in text.lines() {
var trimmed = line.trim()
if trimmed.is_empty() or trimmed.starts_with('#') {
continue
}
var split_at = trimmed.index_of('=')
if split_at == -1 {
continue
}
var key = trimmed[0, split_at].trim()
var value = trimmed[split_at + 1, trimmed.length()].trim()
config[key] = value
}
return config
}
var source = '# server settings
host = localhost
port = 8080
# empty lines and comments are skipped
name = zuri app
'
var config = parse_config(source)
echo config.host
echo config.port.to_number() + 1
echo config.name
echo config.length()
localhost
8081
zuri app
3
Three things in there are worth naming. lines() handles both \n and
\r\n, so the same code reads a file written on Windows. index_of()
returning -1 is checked explicitly rather than relied on for truthiness,
because -1 is falsy and so is 0 — and 0 is a legitimate position.
And every value arrives as a string, because that is what text is;
to_number() is how you leave.
Numbers
There is one numeric type, number, and it is a 64-bit IEEE-754 double.
Integers and fractions are the same type, and the distinction you care
about is whether a particular value happens to be integral.
Writing a Number
There are five ways to write a numeric literal, and they all produce the same type:
echo 42
echo 3.5
echo 0b1011
echo 0c17
echo 0xff
echo 6.02e23
echo 1.5e-3
42
3.5
11
15
255
602000000000000000000000
0.0015
0b is binary, 0c is octal, and 0x is hexadecimal. The letters in a
hexadecimal literal may be either case, so 0xff and 0xFF are the same
number. Note that octal is 0c, not the 0o some other languages use.
The exponent form takes e or E, and the exponent may be negative.
1e3 is 1000; the mantissa needs no decimal point. Note that the
exponent is a way of writing the literal, not a property the value keeps:
6.02e23 prints as its full decimal expansion, because that is the number
it is.
Digit Separators
An underscore between digits is ignored, which makes long numbers readable:
echo 1_000_000
echo 1_0.5_5
1000000
10.55
Separators work in decimal literals only. 0xdead_beef and
0b1010_1010 are syntax errors, not clever formatting.
Two Forms That Do Not Exist
A literal needs a digit on both sides of its decimal point. Neither of these parses:
echo .5
echo 5.
Write 0.5 and 5.0. The second case matters more than it looks, because
5. is also how a method call on a literal starts — which is the subject
of the next-but-one section.
Methods, Not Functions
Everything you would reach into a math library for is a method on the number:
echo 2.sqrt()
echo 8.log2()
echo 100.log10()
echo 100.log()
echo 1.exp()
echo 2.cbrt()
1.4142135623730951
3
2
4.605170185988092
2.718281828459045
1.2599210498948732
log() is the natural logarithm. There is also log1p() and expm1()
for the precision-sensitive forms near zero.
The full trigonometric set is there: sin, cos, tan, asin, acos,
atan, atan2, and the hyperbolic sinh, cosh, tanh, asinh,
acosh, atanh.
echo 1.atan2(1)
0.7853981633974483
Calling a Method on a Literal
A numeric literal takes a method directly. No parentheses, no temporary variable, in every base and with a decimal point or without:
echo 2.sqrt()
echo 3.7.round()
echo 255.hex()
echo 0xff.bin()
echo 1e3.int()
echo 2n.bits()
1.4142135623730951
4
ff
11111111
1000
2
There is exactly one case where you need parentheses, and it is a negative literal. A method call binds tighter than unary minus, so the minus applies to the result rather than to the number:
echo -3.abs()
echo (-3).abs()
-3
3
-3.abs() is -(3.abs()), which is -3. When the receiver is negative,
parenthesise it — or put it in a variable, where the question disappears:
var n = -3
echo n.abs()
3
The same rule covers any expression you want to call a method on: wrap it, because the call would otherwise bind to the last term alone.
echo (1 / 0).is_inf()
echo (2 ** 10).hex()
true
400
Rounding
echo 2.5.ceil()
echo 2.5.floor()
echo 2.5.trunc()
echo 2.5.int()
echo (-2.5).int()
3
2
2
2
-2
trunc() and int() both drop the fractional part, rounding towards zero.
round() rounds half away from zero:
echo 2.5.round()
echo 3.5.round()
echo (-2.5).round()
3
4
-3
fixed() rounds to a number of decimal places and gives you a number back:
echo 2.567.fixed(2)
2.57
fraction() gives the digits after the decimal point as a whole number:
echo 2.567.fraction()
567
Sign, Magnitude and Comparison
echo (-5).sign()
echo (-3).abs()
echo 5.max(9)
echo 5.min(9)
-1
3
9
5
sign() is -1, 0 or 1.
Bases and Characters
echo 255.bin()
echo 255.oct()
echo 255.hex()
echo 65.chr()
11111111
377
ff
A
chr() turns a code point into a one-character string; 'A'.ord() goes
back the other way.
The Special Values
echo (0 / 0).is_nan()
echo (1 / 0).is_inf()
echo 1.is_finite()
true
true
true
NaN is falsy, like zero, so var x = a / b or fallback replaces the
NaN from a 0 / 0 with the fallback. Test with is_nan() when you need
to tell the two apart.
NaN is also not equal to itself, as the standard requires, so
x == x is a valid way to spot one.
Other Methods
factorial() for small integers, to_string(), to_bool() and
to_bigint() for conversion.
echo 5.factorial()
echo 17.to_bigint()
120
17n
The math Module
Constants live in math, because a constant is not a method on anything:
import math
echo math.PI
echo math.E
echo math.Infinity
echo math.NaN
3.141592653589793
2.718281828459045
inf
NaN
It also carries LOG_2, LOG_10, LOG_2_E, LOG_10_E, ROOT_2,
ROOT_3 and ROOT_HALF.
Bigints
When 2^53 is not enough, use a bigint. Write one with an n suffix:
var a = 2n ** 100n
echo a
echo a.bits()
1267650600228229401496703205376n
101
Bigints are arbitrary precision. They never overflow and never lose a digit.
They also never mix with numbers implicitly:
catch {
echo 5n + 3
} as e {
echo e.message
}
operator '+' not defined for call signature (bigint, number)
Convert explicitly, in whichever direction you need:
echo 5.to_bigint() * 2n
echo (2n ** 100n).to_number()
10n
1267650600228229400000000000000
Going to number is lossy once you are past 2^53, which is the whole
reason bigints exist. Going the other way is exact.
The same separation holds in type annotations. bigint is a type name
alongside number and int, and it accepts nothing else:
def scale(n: bigint, factor: number) {
return n * factor.to_bigint()
}
echo scale(2n, 50)
catch {
scale(2, 50)
} as e {
echo e.message
}
100n
scale() expects parameter 'n' (argument 1) to be a bigint, got number
The number-theory methods are the reason bigints are worth having:
echo 100n.gcd(75n)
echo 100n.lcm(75n)
echo 2n.modpow(10n, 1000n)
echo 3n.modinv(11n)
echo 144n.sqrt()
25n
300n
24n
4n
12n
modpow(exponent, modulus) is the operation every public-key algorithm is
built on, and it is computed without ever materialising the full power.
modinv gives the modular multiplicative inverse.
There is also nth_root(), cbrt(), bit(), set_bit(),
trailing_zeros(), is_zero(), is_even(), is_odd(), abs(),
sign(), max(), min(), pow(), and bin()/oct()/hex().
A bigint is the right choice when exactness past 2^53 is the point:
cryptography, currency in minor units, factorials, identifiers that must
survive a round trip. For everything else — measurements, coordinates,
counters, ratios — a number is the type you want.
Lists
A list is an ordered, growable sequence. It holds values of any type, including other lists, and it grows and shrinks as you work with it.
var items = [3, 1, 4, 1, 5]
var mixed = [1, 'two', [3], { four: 4 }]
var empty = []
echo items.length()
echo mixed[3].four
echo empty.is_empty()
5
4
true
Lists carry 39 methods, and they fall into six groups: reading, searching, adding and removing, ordering, transforming, and walking. This section takes them in that order. The one distinction to keep in mind throughout is whether a method mutates the list you called it on or returns a new one, because Zuri’s lists do both and the method names do not tell you which.
Reading
var l = [3, 1, 4, 1, 5]
echo l.length()
echo l.is_empty()
echo l[0]
echo l[-1]
echo l.first()
echo l.last()
echo l[1, 3]
5
false
3
5
3
5
[1, 4]
Indexing
An index counts from zero, and a negative index counts back from the end:
var l = ['a', 'b', 'c', 'd']
echo l[0]
echo l[3]
echo l[-1]
echo l[-4]
a
d
d
a
l[-1] is the last element, which saves writing l[l.length() - 1]
everywhere.
Indexing out of range raises rather than returning nil:
var l = [1, 2, 3]
catch {
echo l[99]
} as e {
echo e.message
}
index 99 out of bounds (length 3)
get() is the form that can fall back, but only when you give it
something to fall back to. get(i) with one argument raises exactly like
l[i] does:
var l = [1, 2, 3]
echo l.get(1)
echo l.get(99, 'missing')
catch {
echo l.get(99)
} as e {
echo e.message
}
2
missing
list index 99 out of range at get()
Slicing
l[a, b] takes the elements from a up to but not including b, and
gives you a new list:
var l = [1, 2, 3, 4, 5]
echo l[1, 3]
echo l[, 3]
echo l[3, ]
echo l[-2, ]
[2, 3]
[1, 2, 3]
[4, 5]
[4, 5]
Either bound may be left out: l[, b] starts at the beginning and l[a, ]
runs to the end. Negative bounds count from the end, just as an index does.
A slice checks its bounds, exactly as an index does. Running past the end raises rather than returning what it can:
catch {
echo [1, 2, 3][1, 99]
} as e {
echo e.message
}
slice bounds 1..99 out of range (length 3)
The one bound that is always legal is length() itself, since a slice’s
upper bound is exclusive. l[0, l.length()] is the whole list, and
l[3, 3] on a three-element list is empty rather than an error.
var l = [1, 2, 3]
echo l[0, l.length()]
echo l[3, 3]
[1, 2, 3]
[]
This same rule governs string slicing, which is why
Strings uses s[split_at + 1, s.length()] rather
than a large number to mean “the rest”.
A slice is a copy. Changing one does not touch the other:
var original = [1, 2, 3]
var part = original[0, 2]
part[0] = 99
echo original
echo part
[1, 2, 3]
[99, 2]
Nested Lists
A list holds lists, and indexes chain:
var grid = [[1, 2], [3, 4], [5, 6]]
echo grid[1]
echo grid[1][0]
echo grid.length()
echo grid[1].length()
[3, 4]
3
3
2
grid.length() counts rows, not elements. There is no two-dimensional
index; grid[1, 0] is a slice from row one to row zero, which is empty.
Combining Lists
Concatenation with +
+ joins two lists into a new one, leaving both operands alone:
var a = [1, 2]
var b = [3]
echo a + b
echo a
echo b
[1, 2, 3]
[1, 2]
[3]
That is the difference between + and extend(): a + b builds a third
list, while a.extend(b) modifies a in place and returns it. Use +
when you want to keep the originals, and extend() when you are
accumulating.
Repetition with *
* repeats a list a whole number of times:
echo [1, 2] * 3
echo [0] * 5
[1, 2, 1, 2, 1, 2]
[0, 0, 0, 0, 0]
[0] * 5 is the idiom for a fixed-size list of a starting value, and it is
worth knowing before you write the loop.
The count behaves exactly as it does for strings: zero and any negative number both produce an empty list, and a fractional count is truncated.
echo [1] * 0
echo [1] * -3
[]
[]
There is one trap. Repetition copies the elements, and when an element is itself a list, all the copies are the same list:
var grid = [[0, 0]] * 3
grid[0][0] = 9
echo grid
[[9, 0], [9, 0], [9, 0]]
One assignment changed every row, because there is only one row. Build
nested structure with a loop, or with map():
var grid = 0..3.to_list().map(@(i) => [0, 0])
grid[0][0] = 9
echo grid
[[9, 0], [0, 0], [0, 0]]
Comparing Lists
== compares lists by value, element by element, however deeply they
nest:
echo [1, 2] == [1, 2]
echo [1, [2, 3]] == [1, [2, 3]]
echo [1, 2] == [2, 1]
echo [1, 2] == [1, 2, 3]
true
true
false
false
Order matters and length matters. To compare two lists as sets, sort
copies of them first, or use the set module from
Chapter 13.
Searching
var l = [3, 1, 4, 1, 5]
echo l.contains(4)
echo l.index_of(1)
echo l.count(1)
true
1
2
index_of() gives -1 when the value is absent. last_index_of()
searches from the other end:
var l = [3, 1, 4, 1, 5]
echo l.index_of(1)
echo l.last_index_of(1)
echo l.last_index_of(1, 2)
echo l.last_index_of(9)
1
3
1
-1
Both take a second argument bounding where a match may sit, so the
pair splits the list at one index: index_of(x, n) finds the first
match at or after n, last_index_of(x, n) the last one at or before
it. Both compare by value, so a list of dictionaries can be searched
with a dictionary literal.
Adding and Removing
These mutate the list in place:
| Method | Effect |
|---|---|
append(v) | add to the end |
insert(v, at) | insert at a position |
extend(other) | append every element of another list |
pop() | remove and return the last element |
shift() | remove and return the first element |
remove(v) | remove the first element equal to v |
remove_at(i) | remove by position |
delete(from, to) | remove an inclusive range, return how many went |
clear() | empty it |
var l = [3, 1, 4]
l.insert(9, 1)
echo l
echo l.shift()
echo l
[3, 9, 1, 4]
3
[9, 1, 4]
var l = [1, 2, 3, 4, 5]
echo l.delete(1, 3)
echo l
3
[1, 5]
Note that delete() takes a from and a to, both inclusive, and
returns the number of elements removed rather than the list.
Ordering
var l = [3, 1, 2]
echo l.sort()
echo l
echo l.reverse()
echo l
[1, 2, 3]
[1, 2, 3]
[3, 2, 1]
[1, 2, 3]
Look closely at those two. sort() mutates and returns the list.
reverse() returns a new list and leaves the original alone. That
asymmetry is the single most common source of list bugs in Zuri code.
sort() takes no comparator. To sort by a computed key, decorate, sort,
and undecorate:
var people = [{ name: 'Ada', age: 36 }, { name: 'Bob', age: 24 }]
var by_age = people
.map(@(p) => [p.age, p.name])
.sort()
.map(@(pair) => pair[1])
echo by_age
[Bob, Ada]
Transforming
Every one of these returns a new list and leaves the receiver untouched:
var l = [1, 2, 3, 4]
echo l.map(@(x) => x * 2)
echo l.filter(@(x) => x > 2)
echo l.unique()
echo l.compact()
echo l.take(2)
echo l.clone()
[2, 4, 6, 8]
[3, 4]
[1, 2, 3, 4]
[1, 2, 3, 4]
[1, 2]
[1, 2, 3, 4]
compact() drops nil entries:
echo [1, nil, 2, nil].compact()
[1, 2]
reduce() folds the list down to one value. With no initial value it
starts from the first element:
echo [1, 2, 3].reduce(@(acc, x) => acc + x)
echo [1, 2, 3].reduce(@(acc, x) => acc + x, 100)
6
106
partition() splits in one pass, matches first:
echo [1, 2, 3, 4].partition(@(x) => x % 2 == 0)
[[2, 4], [1, 3]]
Finding
var l = [1, 2, 3]
echo l.find(@(x) => x > 1)
echo l.find_index(@(x) => x > 1)
echo l.find_last(@(x) => x > 1)
echo l.find_last_index(@(x) => x > 1)
echo l.find_all(@(x) => x > 1)
2
1
3
2
[2, 3]
Asking About Every Element
echo [1, 2, 3].every(@(x) => x > 0)
echo [1, 2, 3].some(@(x) => x > 2)
true
true
some() stops at the first match; every() stops at the first failure.
Walking
There are four ways to visit every element, and they are not interchangeable. Pick by what you need in the body.
for, When You Want the Values
for value in ['a', 'b', 'c'] {
echo value
}
a
b
c
This is the default. Reach for it whenever the position does not matter.
for With Two Variables, When You Want the Index Too
for index, value in ['a', 'b', 'c'] {
echo '${index}: ${value}'
}
0: a
1: b
2: c
With two variables you get the key first and the value second, and for
a list the key is the index. No call to length(), no manual counter.
iter, When You Need Control of the Position
var items = ['a', 'b', 'c']
iter var i = 0; i < items.length(); i++ {
echo '${i}: ${items[i]}'
}
0: a
1: b
2: c
iter costs more typing and earns it only when the traversal is not one
step forward per pass: walking backwards, walking in twos, comparing an
element with its neighbour, or advancing the index from inside the body.
var items = ['a', 'b', 'c', 'd']
iter var i = items.length() - 1; i >= 0; i-- {
echo items[i]
}
iter var i = 0; i < items.length(); i += 2 {
echo items[i]
}
d
c
b
a
a
c
each(), When You Have a Function Already
[10, 20].each(@(value, index) {
echo '${index}: ${value}'
})
0: 10
1: 20
each() takes a function, which makes it the one that composes: it chains
onto a map() or a filter() without a temporary variable, and it accepts
a function you already have by name.
Watch the argument order. each() hands the callback the value first and
the index second, which is the opposite of for. Every list callback in
this chapter follows the same rule — map, filter, find, some,
every all take (value, index) — so it is one thing to remember rather
than several. for is a loop over key-value pairs; each is a callback
over values that happens to tell you where it is.
break and continue work inside for and iter. They do not exist
inside an each() callback; return there ends that one call, not the
walk. When you need to stop early, use a loop, or find() / some(),
which stop on their own.
Combining
echo ['a', 'b'].zip([1, 2])
echo [1, 2, 3].zip_from([[4, 5, 6]])
[[a, 1], [b, 2]]
[[1, 4], [2, 5], [3, 6]]
zip() pairs with one other list. zip_from() takes a list of lists and
zips across all of them at once.
Lists Are References
Assigning a list does not copy it:
var a = [1, 2]
var b = a
b.append(3)
echo a
[1, 2, 3]
Use clone() when you need an independent copy. The clone is shallow:
nested lists inside it are still shared.
var original = [[1, 2], 3]
var copy = original.clone()
copy[1] = 99
copy[0].append(4)
echo original
echo copy
[[1, 2, 4], 3]
[[1, 2, 4], 99]
Replacing the top-level 3 affected only the copy. Appending to the
nested list affected both, because both lists point at the same inner
list. When you need a deep copy, copy the levels you care about yourself.
Mutating or Not: The Summary
This is the table to come back to.
| Mutates the receiver | Returns a new list |
|---|---|
append, insert, extend | map, filter, find_all |
pop, shift, remove, remove_at | unique, compact, take |
delete, clear | reverse, clone, zip, zip_from |
sort | partition |
sort() is the one that catches people out: it is in the left column, and
it also returns the list, so var sorted = items.sort() leaves items
sorted too. If you need both orders, clone first:
var items = [3, 1, 2]
var sorted = items.clone().sort()
echo items
echo sorted
[3, 1, 2]
[1, 2, 3]
A Worked Example
A tiny report generator: take a list of records, drop the incomplete ones,
group what is left, and print a summary. It uses filter, map,
reduce, sort and each together, which is how they usually show up.
var sales = [
{ region: 'north', amount: 120 },
{ region: 'south', amount: 80 },
{ region: 'north', amount: 45 },
{ region: 'east', amount: nil },
{ region: 'south', amount: 200 },
]
def totals_by_region(records) {
var totals = {}
records
.filter(@(r) => r.amount != nil)
.each(@(r) {
totals[r.region] = totals.get(r.region, 0) + r.amount
})
return totals
}
var totals = totals_by_region(sales)
var grand = totals.values().reduce(@(acc, n) => acc + n, 0)
totals
.to_list()[0]
.sort()
.each(@(region) {
echo '${region.rpad(6)} ${totals[region]}'
})
echo 'total ${grand}'
north 165
south 280
total 445
east is absent from the report, not present with a zero, because its one
record was filtered out before any accumulating happened and nothing ever
created the key. That is usually what you want from a report; when it is
not, seed the dictionary with every region first.
Two other details are worth pulling out. totals.get(r.region, 0) supplies
a starting value for a key that does not exist yet, which is what turns a
dictionary into an accumulator. And totals.to_list()[0] takes the keys —
to_list() returns keys and values as two parallel lists — which are then
sorted so the report comes out in a stable order rather than in whatever
order the records happened to arrive.
Dictionaries
A dictionary maps keys to values and remembers the order you inserted them.
var user = { name: 'Ada', age: 36 }
var empty = {}
A bare word key is a string, so { name: ... } and { 'name': ... } are
the same. Keys may also be numbers:
echo { 1: 'one', 2.5: 'two-point-five' }
{1: one, 2.5: two-point-five}
A key that is a bare identifier is always taken as that literal name. To compute a key, make it an expression the parser cannot mistake for a name, which in practice means wrapping it in parentheses:
var key = 'dynamic'
echo { (key): 1 }
echo { 'a' + 'b': 2 }
{dynamic: 1}
{ab: 2}
Assigning through brackets is usually clearer:
var d = {}
d[key] = 3
echo d
{dynamic: 3}
Shorthand Keys
When a key and the variable holding its value have the same name, write it once:
var name = 'Ada'
var age = 36
echo { name, age }
{name: Ada, age: 36}
{ name, age } means exactly { name: name, age: age }. This comes up
constantly in functions that build a result out of locals they have just
computed, and Zuri code uses the shorthand whenever the names line up.
Nesting
A value may be another dictionary, or a list, to any depth. Access chains:
var config = {
server: { host: 'localhost', port: 8080 },
tags: ['web', 'internal'],
}
echo config.server.host
echo config['server']['port']
echo config.tags[0]
localhost
8080
web
Dot and bracket access mix freely, because they are the same operation. Use the dot when the key is a fixed identifier and brackets when it is computed, contains punctuation, or is a number.
Reading
Dot and bracket access are the same operation:
var user = { name: 'Ada', age: 36 }
echo user.name
echo user['name']
echo user.length()
echo user.is_empty()
echo user.keys()
echo user.values()
Ada
Ada
2
false
[name, age]
[Ada, 36]
Reading a key that is not there raises a PropertyError. get() is the
safe form:
echo user.get('name')
echo user.get('nope', 'default')
echo user.contains('age')
Ada
default
true
With no fallback, get() on a missing key returns nil.
Writing
var user = { name: 'Ada' }
user.set('city', 'London')
user.age = 36
user['country'] = 'UK'
echo user
{name: Ada, city: London, age: 36, country: UK}
Assignment through a dot, a bracket or set() all do the same thing, and
all three create the key if it does not exist. add() is a synonym for
set().
To take a key back out:
echo user.remove('city')
echo user.contains('city')
London
false
remove() returns the value that was there.
clear() empties the dictionary.
Merging
var defaults = { host: 'localhost', port: 8080 }
var given = { port: 9000 }
defaults.extend(given)
echo defaults
{host: localhost, port: 9000}
extend() mutates the receiver and the right-hand side wins on conflicts.
There is no + on dictionaries.
Transforming
var user = { name: 'Ada', age: 36, city: nil }
echo user.compact()
echo user.filter(@(value, key) => key == 'name')
echo user.some(@(value, key) => value == 36)
echo user.every(@(value, key) => value != nil)
echo user.reduce(@(acc, value, key) => acc + 1, 0)
{name: Ada, age: 36}
{name: Ada}
true
false
3
compact() drops entries whose value is nil.
Every callback here takes value, then key, the same order lists use.
Walking
Every form below visits the pairs in insertion order, which is the order they were first added, not the order they were last written to.
for With Two Variables, the Usual Form
var user = { name: 'Ada', age: 36 }
for key, value in user {
echo '${key} = ${value}'
}
name = Ada
age = 36
Key first, value second. This is what you want almost every time.
for With One Variable Gives You the Values
var user = { name: 'Ada', age: 36 }
for value in user {
echo value
}
Ada
36
This is worth stating plainly, because the equivalent loop in several other languages hands you the keys. In Zuri, one variable is always the value, whatever you are iterating.
iter Over the Keys, When You Need an Index
var user = { name: 'Ada', age: 36 }
var keys = user.keys()
iter var i = 0; i < keys.length(); i++ {
echo '${i}. ${keys[i]} = ${user[keys[i]]}'
}
0. name = Ada
1. age = 36
A dictionary has no positional index of its own, so iter walks the key
list instead. Reach for it when you need to number the output, or to look
ahead to the next key.
each(), With a Function
var user = { name: 'Ada', age: 36 }
user.each(@(value, key) {
echo '${key} -> ${value}'
})
name -> Ada
age -> 36
The callback receives the value first and the key second — the reverse
of for, matching every other each() in the language.
Walking Just One Side
var user = { name: 'Ada', age: 36 }
for key in user.keys() {
echo key
}
echo user.values()
name
age
[Ada, 36]
keys() and values() each return a plain list, so everything from
Lists applies: sort them, filter them, reduce them.
Sorting keys() is how you get output in a stable order regardless of
insertion.
Converting
var user = { name: 'Ada', age: 36 }
echo user.to_list()
echo user.find_key('Ada')
echo user.clone()
[[name, age], [Ada, 36]]
name
{name: Ada, age: 36}
to_list() gives you two parallel lists, keys then values, not a list of
pairs. find_key() searches by value and returns the first key holding
it.
Dictionaries of Functions
A dictionary value can be a function, and that turns a dictionary into a
dispatch table: a set of named behaviours you can choose between at
runtime by looking a name up. It is the structure that replaces eval(),
and it is worth knowing well.
def plus(a, b) {
return a + b
}
var ops = {
plus: plus,
minus: @(a, b) => a - b,
times: @(a, b) { return a * b },
}
All three forms are equivalent: a named function by name, an arrow function, and an anonymous function with a body.
Calling One
This is the part that surprises people, so it comes first. You cannot call one with a dot in a single step:
echo ops.plus(1, 2)
Unhandled TypeError: object of type dict does not define method 'plus'
A dot followed by a call looks for a method on the dictionary itself —
length(), keys(), get() and the rest — not for a value stored under
that key. Dictionaries have no method called plus, so the call fails.
There are two forms that do work. Read the function out first:
def plus(a, b) {
return a + b
}
var ops = { plus: plus, minus: @(a, b) => a - b }
var chosen = ops.plus
echo chosen(1, 2)
3
Or index with brackets and call the result:
var ops = { plus: @(a, b) => a + b, minus: @(a, b) => a - b }
echo ops['minus'](5, 2)
3
The bracket form is the one to reach for when the key is computed, which is the whole point of a dispatch table:
var ops = {
plus: @(a, b) => a + b,
minus: @(a, b) => a - b,
times: @(a, b) => a * b,
}
for name in ['plus', 'minus', 'times'] {
echo '${name}: ${ops[name](6, 3)}'
}
plus: 9
minus: 3
times: 18
Beware of Names That Are Already Methods
A dictionary’s own methods take priority over its keys when you use a dot
call. Storing a function under add, keys, get, length or any other
method name produces a call that succeeds and does the wrong thing:
def plus(a, b) {
return a + b
}
var d = { add: plus }
echo d.add(1, 2)
echo d.keys()
nil
[add, 1]
d.add(1, 2) called the dictionary’s own add(), which is a synonym for
set() — so it stored the key 1 with the value 2 and returned nil.
Nothing raised. The only sign is the extra key in the output.
Two habits avoid this entirely: use bracket calls for dispatch tables, and avoid method names as keys. Appendix E lists every name a dictionary already uses.
A Dispatch Table in Practice
The pattern is a table of handlers, a lookup with a fallback, and a call:
def _unknown(args) {
return 'unknown command: ${args[0]}'
}
var commands = {
greet: @(args) => 'hello, ${args[1]}',
add: @(args) => args[1].to_number() + args[2].to_number(),
version: @(args) => '1.0.0',
}
def run(line) {
var args = line.split(' ')
var handler = commands.get(args[0], _unknown)
return handler(args)
}
echo run('greet ada')
echo run('add 2 40')
echo run('version')
echo run('explode')
hello, ada
42
1.0.0
unknown command: explode
commands.get(args[0], _unknown) is doing the work. A key that exists
gives you its handler; one that does not gives you the fallback, so there
is no separate contains() check and no branch for the error case.
The important property is that only the handlers in the table can ever
run. A user typing anything at all selects one of four functions you
wrote, or the fallback. That is the difference between a program that
handles input and one that executes it, and it is why Zuri has no
eval(). Chapter 21 covers the reasoning.
Storing Methods
zuri.reflect.bind_method() puts an instance’s method in a table with the
instance still attached:
import zuri
class Counter {
@new() {
self.n = 0
}
increment() {
self.n++
return self.n
}
}
var counter = Counter()
var actions = { up: zuri.reflect.bind_method(counter, 'increment') }
echo actions['up']()
echo actions['up']()
echo counter.n
1
2
2
The bound method keeps its receiver, so calling it through the table changes the counter it came from.
Comparing Dictionaries
== compares dictionaries by value, and order is not part of the
comparison:
echo { a: 1 } == { a: 1 }
echo { a: 1 } == { a: 2 }
echo { a: 1, b: 2 } == { b: 2, a: 1 }
true
false
true
A dictionary remembers its insertion order for iteration, but two dictionaries with the same pairs in different orders are equal. That is the opposite of lists, where order is the whole point.
Dictionaries Are References
Like lists, assigning a dictionary shares it. clone() gives you a
shallow copy.
Dictionaries and Objects
A dictionary is the right shape for data with a variable set of keys: configuration, a parsed JSON document, a set of HTTP headers. When the keys are fixed and there is behaviour attached to them, that is a class. Chapter 6 covers the difference.
A Worked Example
Dictionaries are the natural shape for counting, grouping and indexing. This example does all three over a list of log lines: it counts how often each level appears, groups the messages under their level, and builds an index from a request id back to the line that mentioned it.
var lines = [
'INFO req=a1 started',
'WARN req=a1 slow upstream',
'INFO req=b2 started',
'ERROR req=b2 upstream refused',
'INFO req=a1 finished',
]
var counts = {}
var grouped = {}
var by_request = {}
for line in lines {
var parts = line.split(' ').filter(@(p) => p != '')
var level = parts[0]
var request = parts[1].replace('req=', '', false)
var message = ' '.join(parts[2, parts.length()])
counts[level] = counts.get(level, 0) + 1
if !grouped.contains(level) {
grouped[level] = []
}
grouped[level].append(message)
if !by_request.contains(request) {
by_request[request] = []
}
by_request[request].append(level)
}
echo counts
echo grouped.ERROR
echo by_request.a1
echo by_request.get('zz', ['no such request'])
{INFO: 3, WARN: 1, ERROR: 1}
[upstream refused]
[INFO, WARN, INFO]
[no such request]
Four habits in there are worth taking away.
counts.get(level, 0) + 1 is the counting idiom. get() with a fallback
means you never have to check whether the key exists before adding to it.
if !grouped.contains(level) { grouped[level] = [] } is the grouping
idiom, and it has to be spelled out because get() would hand back a fresh
default list each time rather than one you can keep appending to.
grouped.ERROR and by_request.a1 read keys with a dot, which works
because both are valid identifiers. by_request['b2'] would be needed if
the key were computed or awkward.
And the whole thing preserves order. counts came out INFO, WARN,
ERROR — the order those levels were first seen, not alphabetical and not
arbitrary.
Ranges and Iteration
Ranges
a..b builds a range value: the integers from a up to but not including
b.
.. binds tighter than a method call, so a range takes a method directly
with no parentheses around it:
echo 1..5.to_list()
echo 1..10.step(3).get_step()
[1, 2, 3, 4]
3
Parenthesise only when the bounds are themselves expressions, since ..
takes primaries: (n * 2)..(n * 3).
A property access and a call are expressions too, so
0..self.sizeand0..items.length()are(0..self).sizeand(0..items).length(). Write0..(self.size)and0..(items.length())when the bound is the property rather than the range.This is the price of the line above it.
..has to bind tighter than.for1..10.step(3)to mean a range that steps, and once it does, there is no way for0..self.step(5)to mean the instance’s ownstepinstead. The parentheses say which one is meant, and the compiler cannot guess.
var r = 1..10
echo r
echo r.lower()
echo r.upper()
echo r.range()
1..10
1
10
9
range() is the distance between the bounds.
A range is a value, not loop syntax. Store it, pass it to a function, return it:
def page(number) {
var size = 20
return (number * size)..((number + 1) * size)
}
echo page(2)
40..60
Counting Down
Write the larger bound first and the range runs backwards:
for i in 10..1 {
echo i
}
10
9
8
7
6
5
4
3
2
The rule is the same in both directions: the first bound is included, the second is not.
Stepping
for i in 1..10.step(3) {
echo i
}
1
4
7
step() gives back a new range carrying that stride. get_step() reads
it; the default is 1.
to_list() materialises a range into a list at stride one, ignoring any
step you set:
echo 1..5.to_list()
[1, 2, 3, 4]
Membership
within() tests against the bounds inclusively on both sides, and it
normalises direction, so 10..1.within(10) and 1..10.within(10) both
answer the same:
var r = 1..10
echo r.within(5)
echo r.within(10)
true
true
That differs from iteration, which excludes the upper bound. When you want
“would the loop visit this number”, compare against lower() and upper()
yourself.
Walking a Range
for with one variable gives you the values, and with two it gives you
the position first and the value second:
for value in 3..6 {
echo value
}
for index, value in 3..6 {
echo '${index}: ${value}'
}
3
4
5
0: 3
1: 4
2: 5
The index counts from zero regardless of where the range starts, which is
what makes it useful: 3..6 yields values 3, 4, 5 at positions 0, 1, 2.
loop() is the callback form. It honours the step and the direction:
25..18.loop(@(i) { print('${i} ') })
print('\n')
25 24 23 22 21 20 19
An iter loop needs no range at all — its three clauses already say
everything a range says, and more, since the step can be any expression:
iter var i = 3; i < 6; i++ {
echo i
}
3
4
5
Use a range with for when the bounds are the interesting part, and iter
when the stepping is.
What Is Iterable
for ... in works on:
- lists, giving index and value
- dictionaries, giving key and value, in insertion order
- strings, giving character position and character
- bytes, giving index and the numeric byte
- ranges, giving position and value
- any class that defines
@key()and@value()
is_iterable() answers the question for any value:
echo is_iterable([1, 2])
echo is_iterable('text')
echo is_iterable(42)
true
true
false
The Iterator Protocol
for is not magic. It desugars into a loop over two method calls.
Given for value in thing, Zuri evaluates thing once, then repeats:
key = thing.@key(key), starting fromnil- stop if
keyisnil value = thing.@value(key)- run the body
So @key(previous) answers “what comes after this one?”, returning nil
when there is nothing left, and @value(key) answers “what is stored
here?”.
Everything built in implements this. So can your own classes, which is what Chapter 6 shows.
Because the iterable is evaluated exactly once, this is safe:
for line in read_the_whole_file() {
echo line
}
The function runs once, not once per line.
Functions
A function is a piece of behaviour with a name, a list of parameters, and a result. You have been calling them since Chapter 1; this chapter is about writing them.
Functions in Zuri are values. You can put one in a list, hand it to another function, return one from a function, store one in a dictionary, and call whatever comes back:
def double(n) {
return n * 2
}
var operations = { twice: double, thrice: @(n) => n * 3 }
var chosen = operations.twice
echo chosen(5)
echo operations['thrice'](5)
echo [1, 2, 3].map(double)
10
15
[2, 4, 6]
That one property is what makes map, filter, reduce, every callback
in the standard library, and every route handler in Chapter 15 possible.
Note the shape of those two calls. operations.twice reads the function
out of the dictionary, and then you call what you read. Writing
operations.twice(5) in one step does not work: a dot followed by a call
looks for a method on the dictionary, and a dictionary has no method
called twice. Either read it into a variable first, as above, or index
with brackets and call the result: operations['twice'](5).
This chapter covers declaring a function, the spellings an anonymous one can take, how a closure captures the variables around it, and the optional type annotations that make the runtime check arguments for you.
Defining Functions
Declaration
def greet(name) {
return 'Hello, ' + name
}
echo greet('Zuri')
Hello, Zuri
def, a name, a parameter list in parentheses, a block. There is no return
type to declare, no forward declaration, and no separate signature.
A function with no return yields nil:
def silent() {
echo 'working'
}
echo silent()
working
nil
There are no implicit returns. The last expression in a body is not its
value; if you want something back, say return.
Arity Is Not Enforced
Call a function with fewer arguments than it declares and the missing ones
arrive as nil. Call it with more and the extras are dropped:
def greet(name) {
return 'Hello, ' + name
}
echo greet()
echo greet('Zuri', 'ignored')
Hello, nil
Hello, Zuri
This is not an oversight. It is the mechanism optional parameters are built on, since there is no default-value syntax in a parameter list. You write the default in the body instead:
def greet(name, greeting) {
greeting = greeting or 'Hello'
return '${greeting}, ${name}!'
}
echo greet('Ada')
echo greet('Ada', 'Hi')
Hello, Ada!
Hi, Ada!
The or Trap in Defaults
Because or does the work, a legitimately falsy argument is replaced by
the default. In Zuri that set includes 0, NaN, '' and false:
def indent(text, width) {
width = width or 2
return ' ' * width + text
}
echo '[' + indent('a') + ']'
echo '[' + indent('a', 0) + ']'
[ a]
[ a]
The caller asked for zero indentation and got two. When a falsy value is a
legitimate input, test for nil explicitly:
def indent(text, width) {
if width == nil {
width = 2
}
return ' ' * width + text
}
echo '[' + indent('a') + ']'
echo '[' + indent('a', 0) + ']'
[ a]
[a]
Use or for a default when every falsy value is genuinely absent — a name,
a path, a list. Use == nil for numbers and booleans.
If you want the runtime to reject a missing argument rather than default it, annotate the parameter. Type Annotations covers that.
Variadic Functions
A final parameter prefixed with ... collects every remaining argument
into a list:
def count(...items) {
return items
}
echo count()
echo count(1)
echo count(1, 2, 3)
[]
[1]
[1, 2, 3]
Three facts about what it captures are worth being precise about.
It is always a list, including when nothing was passed. count() gives
[], not nil, so items.length() is safe without a check.
It captures arguments, not contents. Passing a list passes one argument, which arrives as a list nested inside the variadic list:
def count(...items) {
return items
}
echo count([1, 2])
echo count([1, 2], 3)
[[1, 2]]
[[1, 2], 3]
... marks a parameter, not an argument: count(...list) does not parse.
To call a function with a list of arguments, use
apply():
def count(...args) {
return args.length()
}
echo count.apply([1, 2, 3])
3
When a function should accept a collection, though, take the list as an ordinary parameter and skip both.
Named parameters bind first. A variadic can follow named ones, and it takes whatever is left over:
def named(a, ...rest) {
return [a, rest]
}
echo named()
echo named(1)
echo named(1, 2, 3)
[nil, []]
[1, []]
[1, [2, 3]]
The variadic must be last. def mid(a, ...rest, b) is a syntax error,
because there is no rule that could decide how many arguments rest
should keep.
A practical use is a function that formats an arbitrary number of pieces:
def log_line(level, ...parts) {
return '[${level}] ' + ' '.join(parts)
}
echo log_line('warn', 'disk', 'almost', 'full')
echo log_line('info')
[warn] disk almost full
[info]
Where a def Binds
A def scopes exactly like a var. Written at the top level of a file it
binds a module-level name. Written anywhere else it binds a local of the
block it appears in, and it goes away with that block:
def outer() {
def helper() {
return 'from helper'
}
return helper()
}
echo outer()
catch {
helper()
} as e {
echo e.message
}
from helper
undefined global 'helper'
A helper defined inside a function belongs to that function. Nothing outside can reach it, and nothing outside can be broken by it.
That extends to any block, not just a function body:
def choose(verbose) {
if verbose {
def describe(n) {
return 'the number ${n}'
}
return describe(7)
}
return '7'
}
echo choose(true)
echo choose(false)
the number 7
7
Calling One Helper From Another
Helpers declared next to each other can call each other, in either direction:
def parity(n) {
def even(k) {
if k == 0 {
return true
}
return odd(k - 1)
}
def odd(k) {
if k == 0 {
return false
}
return even(k - 1)
}
return even(n)
}
echo parity(10)
echo parity(7)
true
false
even mentions odd before odd is written, and that works because the
compiler reserves a slot for every def in a block before it compiles any
of them. What it does not do is move the declaration: until the def line
itself runs, the slot holds nil, so calling a helper from above its
declaration raises.
A helper can also call itself:
def factorial_of(n) {
def factorial(k) {
if k <= 1 {
return 1
}
return k * factorial(k - 1)
}
return factorial(n)
}
echo factorial_of(6)
720
A Helper Captures What Surrounds It
A nested def is a closure over the locals around it, and it outlives the
call that created it:
def make_counter(start) {
var count = start
def bump() {
count++
return count
}
return bump
}
var bump = make_counter(10)
echo bump()
echo bump()
11
12
An anonymous function bound to a var does the same job, and is the better
choice when the helper is a one-liner or is being passed straight into
something else:
def outer() {
var helper = @() => 'private'
return helper()
}
echo outer()
private
Pick whichever reads better. A def gives the function a real name in
stack traces and can recurse without the extra var; a lambda is shorter.
One Declaration Per Name
Declaring the same function name twice in one scope is a compile error:
def pick() {
return 'first'
}
def pick() {
return 'second'
}
SyntaxError: multiple declaration for function 'pick' found
--> /path/to/main.zu:5:5
|
5 | def pick() {
| ^
A different parameter list does not make it a different function. Zuri has no overloading: one name, one function.
def render(value) {}
def render(value, width) {}
SyntaxError: multiple declaration for function 'render' found
When you want one name to handle several shapes of input, take the extra
arguments as optional and branch in the body — which is what the
greeting parameter above is doing.
The check is per scope, exactly like var’s. A function declared
inside another is a separate declaration in a separate scope, so the
compiler allows it, and it shadows the outer one for as long as its own
scope lasts:
def render() {
return 'top level'
}
def wrapper() {
def render() {
return 'inner'
}
return render()
}
echo wrapper()
echo render()
inner
top level
wrapper’s own render is a local of wrapper. It answers every call
made inside that function and disappears when the call ends, leaving the
module-level render untouched.
Classes follow the same rule for their methods:
class Duplicate {
m() {}
m() {}
}
SyntaxError: multiple declaration for method 'm' found in class 'Duplicate'
The REPL is the one exception. Retyping a def there replaces the previous
one, because correcting something you just typed is what the prompt is for.
Declaration Order
A function must be declared before the top-level line that calls it:
echo later()
def later() {
return 'nope'
}
Unhandled UndefinedError: undefined global 'later'
There is no hoisting. Declarations execute in order, like everything else.
Inside a function body the rule does not apply, because a name is resolved when the call runs rather than when it is compiled. That is what makes mutual recursion work with no forward declaration:
def is_even(n) {
if n == 0 {
return true
}
return is_odd(n - 1)
}
def is_odd(n) {
if n == 0 {
return false
}
return is_even(n - 1)
}
echo is_even(10)
echo is_odd(7)
true
true
is_even refers to is_odd before it exists. That is fine, because the
reference is resolved when is_even(10) runs, by which time both
declarations have executed.
Helpers nested inside a function get the same freedom by a different route: their slots are reserved together, before any of their bodies compile. Either way, what matters is that every declaration has run by the time the first call is made.
Functions as Values
A function is a value, so it can be passed, stored and returned:
def apply_twice(f, x) {
return f(f(x))
}
def increment(n) {
return n + 1
}
echo apply_twice(increment, 5)
echo apply_twice(@(s) => s + '!', 'hi')
7
hi!!
What a Function Knows About Itself
Every callable carries four methods:
def named(a, b, ...c) {
return [a, b, c]
}
echo named.name()
echo named.arity()
echo named.is_variadic()
echo named.to_string()
named
3
true
<function named(3)>
arity() counts declared parameters, the variadic one included, so a
variadic function’s arity is a minimum rather than a requirement.
An anonymous function is named after the order the compiler met it:
var f = @(x) => x
echo f.name()
@anon0
A method read off an instance keeps its receiver, and counts it:
class Greeter {
hello() {
return 'hi'
}
}
var bound = Greeter().hello
echo bound()
echo bound.name()
echo bound.arity()
hi
hello
1
hello() declares no parameters, and arity() reports 1, because the
instance is the first one.
call()
call() invokes a function with the arguments you give it:
def add(a, b) {
return a + b
}
echo add.call(2, 3)
5
call() takes the arguments individually, exactly as a normal call
does. It is not a spread: add.call([2, 3]) passes one argument, a list.
Its use is calling something whose identity you only have as a value —
a handler out of a dictionary, a method from zuri.reflect — in a place
where the ordinary call syntax reads badly.
Testing Callability
def f() {}
echo is_callable(f)
echo is_callable(print)
echo is_callable(Error)
echo is_callable(42)
echo is_function(f)
echo is_function(Error)
true
true
true
false
true
false
is_callable() is true for anything you can put parentheses after,
classes included — calling a class constructs an instance.
is_function() is narrower: true for a def, an anonymous function and a
built-in, false for a class.
A Worked Example
A small pipeline builder, using most of this section: functions as values, a variadic parameter, a default, and a returned closure.
def pipeline(...stages) {
return @(input) {
var value = input
for stage in stages {
value = stage(value)
}
return value
}
}
def strip(text) {
return text.trim()
}
def collapse_spaces(text) {
return text.replace('/\s+/', ' ')
}
def truncate(text, limit) {
limit = limit or 20
if text.length() <= limit {
return text
}
return text[0, limit - 1] + '…'
}
var tidy = pipeline(strip, collapse_spaces, @(t) => truncate(t, 12))
echo '[' + tidy(' too many spaces here ') + ']'
echo '[' + tidy(' short ') + ']'
echo '[' + pipeline()('untouched') + ']'
[too many sp…]
[short]
[untouched]
Four things to take from it.
pipeline() takes its stages variadically and returns a closure over them,
so the returned function is a new function specialised to those stages.
pipeline() with no arguments still works, because a variadic parameter is
an empty list rather than nil, and a for over an empty list runs zero
times.
The third stage is wrapped in @(t) => truncate(t, 12) because
truncate takes two parameters and a stage takes one. That wrapper is the
thing a spread operator would otherwise be for.
And limit = limit or 20 is safe here precisely because 0 is not a
sensible limit. Had it been, this would need the == nil form.
Anonymous Functions and Closures
The Spellings
An anonymous function has two halves you can vary independently: how it opens, and how its body is written. That gives a small grid rather than a list to memorise.
Opening. def is the keyword; @ is its shorthand. They are the same
thing.
Body. A block in braces returns with return. An arrow returns the one
expression after it.
var block_long = def(x) { return x * x }
var block_short = @(x) { return x * x }
var arrow_long = def(x) => x * x
var arrow_short = @(x) => x * x
echo [block_long(3), block_short(3), arrow_long(3), arrow_short(3)]
[9, 9, 9, 9]
When there are no parameters, the empty parentheses may be dropped as well:
var a = @() { return 1 }
var b = @() => 2
var c = def() { return 3 }
var d = def() => 4
var e = @ => 5
var f = @{ return 6 }
var g = def { return 7 }
var h = def => 8
echo [a(), b(), c(), d(), e(), f(), g(), h()]
[1, 2, 3, 4, 5, 6, 7, 8]
All eight are the same construct. Which to use:
@(x) => exprfor a one-expression callback. This is what most Zuri code uses, and what the standard library is written in.@(x) { ... }when the body needs more than one statement.def(x) { ... }when the function is long enough that you want it to look like the declarations around it.@ => exprfor a thunk — a value computed on demand, with nothing passed in.
Compare the two common ones where they actually turn up:
echo [' a ', ' b'].map(@(t) => t.trim())
echo [1, 2, 3].filter(@(n) {
if n == 2 {
return false
}
return n > 0
})
[a, b]
[1, 3]
The arrow form has no return because the expression is the return
value. Adding one is a mistake the compiler will not catch: @(x) => return x
does not parse, but @(x) { x } parses and returns nil.
Anonymous functions take variadic parameters and type annotations exactly like named ones:
var sum = @(...numbers) => numbers.reduce(@(a, b) => a + b, 0)
var doubled = @(n: number) => n * 2
echo sum(1, 2, 3)
echo doubled(4)
catch {
doubled('four')
} as e {
echo e.message.replace('/@anon\d+/', '@anonN')
}
6
8
@anonN() expects parameter 'n' (argument 1) to be a number, got string
The real message carries a number rather than the N shown here. An
anonymous function is named @anonN in the order the compiler met it in
the file, which is worth knowing when one turns up in a stack trace — and
worth not depending on, since inserting another anonymous function above it
renumbers everything below.
Closures
A function carries the variables it referenced from its surroundings, and keeps them alive after the enclosing function has returned:
def make_counter() {
var n = 0
return @() {
n++
return n
}
}
var a = make_counter()
var b = make_counter()
echo a()
echo a()
echo b()
1
2
1
a and b each closed over their own n. Each call to
make_counter() created a fresh one.
Capture Is by Reference
A closure captures the variable, not a snapshot of its value:
def demo() {
var total = 0
var add = @(n) { total += n }
add(5)
add(10)
return total
}
echo demo()
15
This is what makes accumulator patterns work, and it is what makes loop variables surprising. If you want a per-iteration capture, declare a fresh variable inside the loop body:
var fns = []
iter var i = 0; i < 3; i++ {
var captured = i
fns.append(@() => captured)
}
echo fns.map(@(f) => f())
[0, 1, 2]
captured is a new variable on each pass, so each closure gets its own.
Closures over Parameters
Parameters are captured the same way, which is the basis of partial application:
def adder(amount) {
return @(n) => n + amount
}
var add_five = adder(5)
var add_ten = adder(10)
echo add_five(1)
echo add_ten(1)
6
11
Recursion in an Anonymous Function
An anonymous function has no name to call itself by. Bind it first:
var factorial
factorial = @(n) {
if n <= 1 {
return 1
}
return n * factorial(n - 1)
}
echo factorial(5)
120
The variable is captured by reference, so by the time the body runs,
factorial is bound.
Functions Compare by Identity
def f() {}
var g = f
echo f == g
echo f == @() {}
true
false
Two separately written functions are never equal, even with identical bodies.
A Worked Example
Closures are at their most useful when a function needs to remember something between calls without that something becoming a global. Here is a rate limiter: it hands back a function that answers “may I do this now?”, and keeps its own tally where nothing else can reach it.
def make_limiter(max_per_window: number) {
var used = 0
return {
allow: @() {
if used >= max_per_window {
return false
}
used++
return true
},
remaining: @() => max_per_window - used,
reset: @() {
used = 0
},
}
}
var limiter = make_limiter(2)
var allow = limiter.allow
var remaining = limiter.remaining
var reset = limiter.reset
echo allow()
echo allow()
echo allow()
echo remaining()
reset()
echo allow()
true
true
false
0
true
Three separate functions share one used, because all three closed over
the same variable in the same call to make_limiter. A second call to
make_limiter would produce a second, independent trio.
This is the closest Zuri gets to a private field without a class, and it is
worth knowing for exactly that reason: there is no way for a caller to read
or write used except through the three functions you gave them.
Type Annotations
Zuri is dynamically typed, and parameters may be annotated. An annotated
parameter is checked at every call, and a mismatch raises a TypeError
naming the parameter, its position and what actually arrived.
def repeat(text: string, times: number) {
return text * times
}
echo repeat('ab', 3)
catch {
repeat(42, 3)
} as e {
echo e.message
}
ababab
repeat() expects parameter 'text' (argument 1) to be a string, got number
The Type Names
| Name | Accepts |
|---|---|
any | anything, including nil |
bool | true or false |
number | any number |
int | a number with no fractional part |
bigint | an arbitrary-precision integer |
string | a string |
bytes | a byte buffer |
list | a list |
dict | a dictionary |
range | a range |
file | a file handle |
function | a function or an anonymous function |
callable | anything callable, classes included |
iterable | anything for can walk |
type | a class |
| ClassName | an instance of that class or a subclass |
Anything that is not one of those names is read as a class name.
number and bigint are separate, in annotations exactly as they are
everywhere else. A parameter typed bigint rejects 10, and one typed
number rejects 10n:
def factorial(n: bigint) {
if n <= 1n {
return 1n
}
return n * factorial(n - 1n)
}
echo factorial(25n)
catch {
factorial(25)
} as e {
echo e.message
}
15511210043330985984000000n
factorial() expects parameter 'n' (argument 1) to be a bigint, got number
Write bigint|number when a function genuinely takes either, and convert
inside it with to_bigint().
Required by Default
A plain annotation makes the argument required. Omitting it is a
TypeError, because the parameter arrives as nil:
def needs_number(x: number) {}
catch {
needs_number()
} as e {
echo e.message
}
needs_number() expects parameter 'x' (argument 1) to be a number, got nil
That is how you get argument checking out of a language that otherwise lets you call anything with anything.
Optional Parameters
Prefix the type with ? to allow nil:
def connect(host: string, port: ?number) {
port = port or 8080
echo '${host}:${port}'
}
connect('localhost')
connect('localhost', 9000)
localhost:8080
localhost:9000
?any is the same as no annotation at all, which is what an unannotated
parameter gets.
Union Types
Separate alternatives with |:
def render(value: string|number|list) {
echo value
}
render('text')
render(42)
render([1])
catch {
render({ a: 1 })
} as e {
echo e.message
}
text
42
[1]
render() expects parameter 'value' (argument 1) to be a string, a number, or a list, got dict
The ? goes at the front and applies to the whole union: ?string|number.
Classes and Subclasses
A class name accepts instances of that class and of anything that inherits from it:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
}
def distance_from_origin(p: Point) {
return (p.x * p.x + p.y * p.y).sqrt()
}
echo distance_from_origin(Point(3, 4))
5
This is what makes error handling readable, since every built-in error
subclasses Error:
def report(e: Error) {
echo '${e.type}: ${e.message}'
}
catch {
raise ValueError('bad value')
} as e {
report(e)
}
ValueError: bad value
Where Annotations Work
Annotations go on parameters: in a def, in a method, and in an anonymous
function.
class Report {
render(rows: list, title: ?string) {
# ...
}
}
var f = @(n: int) => n * 2
The same syntax goes on var declarations and on class fields:
var count: number = 0
var name: ?string
class Report {
var rows: list = []
var title: ?string
}
A variable annotation is a statement of intent. It is not checked at
runtime, so the runtime will not stop you from putting a string in
count. What it does is make the declaration self-documenting, and give
static analysis tools something to work with: an annotated declaration
tells a linter, an editor’s completion engine or a type checker exactly
what belongs there, and lets them flag the assignment your eyes would
otherwise have to catch.
Annotate declarations for the reader and the tooling. Annotate parameters for the runtime.
When to Annotate
Annotations are optional, and a program with none of them is perfectly ordinary Zuri. The question is where they earn their place.
Annotate a boundary. A function that receives data from outside your program — a request handler, a file parser, a public function in a module other people import — is where a wrong type first arrives. An annotation there turns a confusing failure deep in the call stack into a clear one at the door:
def parse_port(raw: string) {
var port = raw.to_number()
if port < 1 or port > 65535 {
raise ValueError('port out of range: ${raw}')
}
return port
}
catch {
parse_port(8080)
} as e {
echo e.message
}
echo parse_port('8080')
parse_port() expects parameter 'raw' (argument 1) to be a string, got number
8080
Note the division of labour there. The annotation handles wrong type, and
the explicit check handles wrong value. An annotation can never do the
second job, because 70000 is a perfectly good number.
Annotate to replace a manual check. Any function that opens with
if !is_string(x) { raise TypeError(...) } is spelling out by hand what an
annotation says in one word, and the annotation produces a better message:
def shout(text: string) {
return text.upper() + '!'
}
echo shout('hello')
HELLO!
Leave internal helpers alone if you prefer. A private function called from three places in the same file, all of which you can see, gains less. Annotate it if it documents something non-obvious; skip it if it does not.
One thing an annotation is not: a substitute for validation. text: string
guarantees you have a string, not that the string is a valid email address,
a well-formed date or a non-empty name. The validate module covers that
job, and Chapter 13 introduces it.
Classes and Objects
A dictionary holds data. A class holds data and the behaviour that goes with it, under a name you can check for and inherit from.
class Temperature {
@new(celsius) {
self.celsius = celsius
}
fahrenheit() {
return self.celsius * 9 / 5 + 32
}
to_string() {
return '${self.celsius}C'
}
}
var t = Temperature(100)
echo t.fahrenheit()
echo t.to_string()
212
100C
Three facts about Zuri classes shape everything in this chapter, and it is worth having them up front.
The constructor is @new. Methods whose names start with @ are
decorated methods, and each one connects the class to a piece of the
language’s own syntax: construction, arithmetic, comparison, iteration,
JSON encoding. There is a fixed set of them, and
Decorated Methods covers all of it.
Fields are reached through self. Inside a method, self is the
instance. There is no bare balance that means self.balance, and no
self parameter in the declaration either.
A class is sealed. The set of fields and methods is fixed when the class is declared. Nothing at runtime can add a new field to an instance or a new method to a class, and an attempt raises an error rather than quietly creating one. That rules out a family of bugs — a typo in a field name is an error, not a new field — and it rules out monkey-patching, which is a technique some languages rely on and Zuri does not offer.
Inheritance is single: a class has at most one parent, written with <.
There are no interfaces, no mixins and no abstract keyword. The sections
that follow show what you write instead.
Sealed does not mean unchangeable forever, though. A separate declaration
written with > instead of < — a class extension — can add methods
to a class after the fact, including one you did not write.
Class Extensions covers it.
Defining a Class
class Account {
var balance = 0
var owner
@new(owner, balance) {
self.owner = owner
self.balance = balance or 0
}
deposit(amount) {
self.balance += amount
return self
}
}
class Name { ... }declares it.PascalCaseis the convention.varinside the body declares a field, with an optional default. A field with no default starts asnil.- A bare
name(params) { ... }is a method. There is nodefkeyword on methods. @newis the constructor.selfis the current instance.
Creating an Instance
Call the class. There is no new keyword:
var a = Account('Ada', 50)
echo a.owner
echo a.balance
Ada
50
If the class has no @new, calling it with no arguments gives you an
instance with every field at its default.
The Constructor
@new is the constructor. It runs once, when the class is called, and its
job is to put the instance into a usable state — which is what
Account’s did above.
It Is Optional
A class with no @new is constructed with no arguments, and every field
takes its declared default:
class Settings {
var theme = 'dark'
var retries
}
var s = Settings()
echo s.theme
echo s.retries
dark
nil
Extra arguments to a class with no @new are ignored rather than
rejected, exactly as they are for a function.
It Declares Fields
self.x = value inside @new declares a field. This is the one place
in the language where assignment creates a member, and it exists so a
constructor does not have to repeat every field as a var line above it:
class Point {
@new(x, y) {
self.x = x
self.y = y
self.distance = (x * x + y * y).sqrt()
}
}
var p = Point(3, 4)
echo p.distance
5
distance was never declared with var, and it is a real field.
That power belongs to @new’s own body and nowhere else. A constructor
that delegates its setup to a helper must declare those fields:
class Delayed {
var ready # required: setup() cannot declare it
@new() {
self.setup()
}
setup() {
self.ready = true
}
}
echo Delayed().ready
true
Remove the var ready line and setup() raises
undefined field 'ready'.
Arguments Are Not Checked Unless You Ask
@new is a function, so the usual rules apply: missing arguments arrive as
nil, extra ones are dropped. A constructor that assumes it got something
fails later, in a confusing place:
class Ctor {
@new(x) {
self.derived = x * 2
}
}
Ctor()
Unhandled TypeError: operator '*' not defined for call signature (nil, float)
Annotate the parameter and the failure moves to the call, where it names the problem:
class Ctor {
@new(x: number) {
self.derived = x * 2
}
}
echo Ctor(5).derived
catch {
Ctor('five')
} as e {
echo e.message
}
10
@new() expects parameter 'x' (argument 1) to be a number, got string
Constructors are the highest-value place in a program to annotate, because a badly built object goes wrong somewhere else entirely.
Validating in the Constructor
A constructor may raise. Nothing is returned to the caller, so an object that cannot be valid never exists:
class Port {
@new(number: number) {
if number < 1 or number > 65535 {
raise ValueError('port out of range: ${number}')
}
self.number = number
}
}
echo Port(8080).number
catch {
Port(99999)
} as e {
echo '${e.type}: ${e.message}'
}
8080
ValueError: port out of range: 99999
This is worth doing whenever a class has an invariant. Every other method can then assume it holds, instead of re-checking.
@new Cannot Return a Value
Calling a class always produces an instance of that class. A return in
@new ends the constructor early; it does not change what the caller
gets:
class Returns {
@new() {
self.v = 1
return 'this is discarded'
}
}
echo typeof(Returns())
Returns
If you want a call that may hand back something else — a cached instance, a
subclass, nil on bad input — write a static factory method and call
that instead.
Methods
A method is a name(params) { ... } declaration inside the class body.
There is no def keyword on it:
class Account {
@new(owner, balance) {
self.owner = owner
self.balance = balance
}
deposit(amount) {
self.balance += amount
return self
}
}
var a = Account('Ada', 50)
a.deposit(25).deposit(25)
echo a.balance
100
deposit returns self, which is what makes the chain work. Returning
self from a mutator is a common Zuri idiom.
Inside a method, self is required to reach a field or another method.
There is no implicit receiver:
class Greeter {
var name = 'world'
greet() {
return 'hello ' + self.name
}
}
Writing name there would look for a local or a global, not a field.
Fields Are Declared, Not Discovered
The set of fields is fixed when the class is declared. Two things follow from that.
First, assigning to a field that was never declared is an error:
class Account {
var balance = 0
}
var a = Account()
catch {
a.nickname = 'rainy day'
} as e {
echo e.message
}
undefined field 'nickname' on instance of 'Account'
Second, self.x = value inside @new does declare a field, as a
convenience so that constructors do not have to repeat themselves:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
}
echo Point(3, 4).x
3
That convenience applies to @new’s own body and nowhere else. A
constructor that calls a helper to do its initialisation must declare those
fields with var:
class A {
var from_helper # required
@new() {
self.setup()
}
setup() {
self.from_helper = 2
}
}
Without the var line, the assignment inside setup() raises
undefined field 'from_helper'.
Static Members
static puts a field or method on the class rather than on each instance:
class Account {
static var count = 0
@new(owner) {
self.owner = owner
Account.count++
}
static open(owner) {
return Account(owner)
}
}
Account('Ada')
Account('Bob')
echo Account.count
echo Account.open('Carol').owner
2
Carol
A static method has no self:
SyntaxError: 'self' used outside of a method
Reach the class by name instead.
Static fields are the one mutable part of a class. The set of members is sealed; the values of static fields are not.
Duplicate Members Are Errors
Declaring the same method twice in one class does not silently keep the last one:
class B {
m() {}
m() {}
}
SyntaxError: multiple declaration for method 'm' found in class 'B'
The same applies to declaring the same class name twice in one module.
Instances and Dictionaries
dict | class | |
|---|---|---|
| keys | any, added at any time | fixed at declaration |
| access | d.key or d['key'] | obj.field only |
| missing member | get() returns a fallback | error |
| behaviour | none | methods |
| cost of a read | hash lookup | array index |
Use a dictionary for data whose shape you learn at runtime: parsed JSON, HTTP headers, a config file. Use a class when the shape is known and there is behaviour to attach.
Inheritance
A class extends another with <:
class Shape {
@new(name) {
self.name = name
}
area() {
raise NotImplementedError('${self.name} must define area()')
}
describe() {
return '${self.name} with area ${self.area()}'
}
}
class Circle < Shape {
@new(radius) {
parent('circle')
self.radius = radius
}
area() {
return 3.141592653589793 * self.radius ** 2
}
}
echo Circle(1).describe()
circle with area 3.141592653589793
A subclass inherits every field and method. Single inheritance only; a class has at most one parent.
Note the direction of the arrow. class Circle < Shape creates a new
class that borrows from Shape. Turning it around —
class Anything > Shape — is a different declaration entirely: it adds
methods to Shape itself and creates no class at all. See
Class Extensions.
Naming the Parent
A class is a value like any other, so the parent does not have to be a bare name in the current file. Anything an expression starting with a name can reach works, which is what lets you build on a class another module owns:
import log
class MemoryTransport < log.Transport {
@new() {
self.lines = []
}
handle(record) {
self.lines.append(record)
}
}
echo instance_of(MemoryTransport(), log.Transport)
true
Property access, indexing and calls all work the same way, so a parent
picked out of a registry (class Store < backends['redis']) or handed
back by a function (class Store < chosen_backend()) is as valid as a
name. The expression is evaluated once, where the class is declared.
The one rule is that it has to begin with a name. That is what keeps the
{ opening the body from ever being read as the start of a dictionary.
parent
parent means two related things.
parent(args) calls the parent’s constructor. That is the
parent('circle') line in Circle above, and it comes first in @new,
before the subclass sets up anything of its own. Calling it late means the
parent’s constructor overwrites what you just assigned; not calling it at
all means the parent’s fields are never initialised.
parent.method(args) calls the parent’s version of a method that this
class has overridden:
class Square < Shape {
@new(side) {
parent('square')
self.side = side
}
area() {
return self.side ** 2
}
describe() {
return 'a ' + parent.describe()
}
}
echo Square(2).describe()
a square with area 4
Note what happened there. parent.describe() ran Shape’s describe,
which calls self.area(), which dispatched back to Square’s area. A
method always dispatches on the actual object, not on the class the code
was written in.
Overriding
Redeclaring a method in a subclass replaces it. There is no keyword for it and no way to forbid it.
describe() above is the useful shape of this: a parent method written in
terms of a method the child supplies. Shape.area() raises
NotImplementedError, so a subclass that forgets to override it says so
clearly:
catch {
Shape('blob').area()
} as e {
echo e.message
}
blob must define area()
That is Zuri’s abstract method. There is no abstract keyword because
there does not need to be one.
A Subclass Without a Constructor
If a subclass declares no @new, it uses its parent’s:
class NoCtor < Shape {}
echo NoCtor('plain').name
plain
Testing Ancestry
instance_of() walks the whole chain:
echo instance_of(Circle(1), Shape)
echo instance_of(Circle(1), Square)
true
false
typeof() gives the most specific class name, as a string:
echo typeof(Circle(1))
Circle
The Depth Is Free
Inheritance chains cost nothing to walk at runtime. A method lookup on a class four levels deep is the same operation as one on a class with no parent, because every class’s method table is complete at declaration time. Write the hierarchy the design wants.
What You Write Instead of an Interface
Zuri has single inheritance and no interfaces, so the two patterns below do the jobs an interface would do elsewhere.
A base class that raises. When a base class needs every subclass to supply a method, declare it and raise:
class Shape {
@new(name: string) {
self.name = name
}
area() {
raise NotImplementedError('${self.name} must define area()')
}
describe() {
return '${self.name} has area ${self.area()}'
}
}
class Circle < Shape {
@new(radius: number) {
parent('circle')
self.radius = radius
}
area() {
return 3.141592653589793 * self.radius ** 2
}
}
class Blob < Shape {
@new() {
parent('blob')
}
}
echo Circle(2).describe()
catch {
echo Blob().describe()
} as e {
echo '${e.type}: ${e.message}'
}
circle has area 12.566370614359172
NotImplementedError: blob must define area()
describe() calls self.area(), and self is the actual instance, so the
subclass’s version runs. That is the whole of dynamic dispatch in Zuri:
there is nothing to declare and nothing to mark virtual.
A parameter typed by the base class. An annotation naming a class accepts any subclass of it, which is how you say “anything that is a Shape”:
def total_area(shapes: list) {
return shapes.reduce(@(sum, shape: Shape) => sum + shape.area(), 0)
}
Checking What Something Is
Two built-ins answer questions about an instance’s type, and they answer different questions.
instance_of(value, Class)
instance_of() asks “does this behave like a Class?”, and it walks the
whole inheritance chain to answer:
class Shape {}
class Circle < Shape {}
class Square < Shape {}
class Ellipse < Circle {}
var c = Circle()
echo instance_of(c, Circle)
echo instance_of(c, Shape)
echo instance_of(c, Square)
echo instance_of(Ellipse(), Shape)
true
true
false
true
Ellipse is two levels below Shape and still answers true. There is no
depth limit: the check follows parents until it finds a match or runs out.
A sibling — Circle against Square — is false, because siblings share
an ancestor rather than a lineage.
It never raises for the first argument. Anything that is not an
instance of the class simply answers false, including values that are not
instances at all:
class Shape {}
echo instance_of(42, Shape)
echo instance_of('circle', Shape)
echo instance_of(nil, Shape)
echo instance_of([1, 2], Shape)
false
false
false
false
That makes it safe to use as a guard on a value you know nothing about; no
is_instance() check has to come first.
A class is not an instance of itself. Passing the class rather than an
object gives false:
class Shape {}
class Circle < Shape {}
echo instance_of(Circle, Shape)
echo instance_of(Circle(), Shape)
false
true
This trips people up when a variable might hold either. Circle is a
class; Circle() is an instance; only the second is “a Shape”.
The second argument must be a class, and here it does raise:
class Shape {}
catch {
instance_of(Shape(), 'Shape')
} as e {
echo e.message
}
instance_of() expects argument 2 to be a class, got string
The class name as a string is the common version of this mistake. Pass the class itself.
A class held in a variable works, because the check is on the value rather than on the spelling:
class Base {}
class Derived < Base {}
var Alias = Base
echo instance_of(Derived(), Alias)
true
So does a class reached through a module:
import set
echo instance_of(set.set([1, 2]), set.Set)
true
It Is How the Error Hierarchy Works
Every built-in error inherits from Error, so instance_of() is what lets
one handler sort them:
echo instance_of(ValueError('bad'), Error)
echo instance_of(ValueError('bad'), TypeError)
true
false
That is the mechanism behind the “handle one kind, re-raise the rest” pattern in Error Handling, and behind mapping a domain error to an HTTP status in Chapter 25.
typeof(value)
typeof() asks a different question: “what exactly is this?”. On an
instance it names the concrete class, and it never mentions ancestors:
class Shape {}
class Circle < Shape {}
echo typeof(Circle())
echo typeof(Shape())
echo typeof(Circle)
echo typeof(42)
echo typeof('text')
Circle
Shape
class
number
string
Note the third line. typeof() on the class itself answers class, not
Circle — the class is a value of kind “class”, and its name is not what
typeof() reports.
Which to Use
| Question | Use |
|---|---|
can I treat this as a Shape? | instance_of(x, Shape) |
| which exact class is this? | typeof(x) |
| is this any instance at all? | is_instance(x) |
Reach for instance_of() by default. It is the one that respects
inheritance, and code written with typeof(x) == 'Circle' breaks the day
someone subclasses Circle — the subclass is a perfectly good Circle,
and the string comparison says otherwise.
A Type Annotation Is the Same Check
Annotating a parameter with a class name performs exactly the
instance_of() test, at every call, with a better error message:
class Shape {}
class Circle < Shape {}
def describe(shape: Shape) {
return 'a shape of type ${typeof(shape)}'
}
echo describe(Circle())
echo describe(Shape())
catch {
describe(42)
} as e {
echo e.message
}
a shape of type Circle
a shape of type Shape
describe() expects parameter 'shape' (argument 1) to be a Shape, got number
Prefer the annotation when the answer decides whether the function should
run at all, and instance_of() when the answer decides which branch to
take.
Decorated Methods
A method whose name begins with @ is a decorated method. The runtime
calls it for you when a piece of language syntax is applied to your
instance. That is how a class hooks into construction, arithmetic,
comparison and iteration without any special syntax of its own.
Decorated methods are ordinary methods. They can be inherited, overridden and called by name.
@new: Construction
Already covered. It runs when the class is called, receives the arguments,
and is the one place self.x = value may declare a new field.
Arithmetic
Define the operator you want and it works on your instances:
class Vector {
@new(x, y) {
self.x = x
self.y = y
}
@add(other) {
return Vector(self.x + other.x, self.y + other.y)
}
@sub(other) {
return Vector(self.x - other.x, self.y - other.y)
}
@mul(k) {
return Vector(self.x * k, self.y * k)
}
@neg() {
return Vector(-self.x, -self.y)
}
to_string() {
return '(${self.x}, ${self.y})'
}
}
var a = Vector(1, 2)
var b = Vector(3, 4)
echo (a + b).to_string()
echo (b - a).to_string()
echo (a * 3).to_string()
echo (-a).to_string()
(4, 6)
(2, 2)
(3, 6)
(-1, -2)
The full arithmetic set:
| Decorator | Operator |
|---|---|
@add | + |
@sub | - (binary) |
@mul | * |
@div | / |
@floordiv | // |
@mod | % |
@pow | ** |
@neg | - (unary) |
Bitwise and Logic
| Decorator | Operator |
|---|---|
@and | & |
@or | | |
@xor | ^ |
@lshift | << |
@rshift | >> |
@urshift | >>> |
@not | ~ |
class Flags {
@new(bits) {
self.bits = bits
}
@and(mask) {
return self.bits & mask
}
@not() {
return Flags(~self.bits)
}
}
echo Flags(5) & 4
echo (~Flags(5)).bits
4
-6
@not is bound to ~, the bitwise complement. ! is logical negation and
it is not overridable: an instance is always truthy, so !instance is
always false.
Comparison
| Decorator | Operator |
|---|---|
@eq | == and != |
@lt | < |
@lte | <= |
@gt | > |
@gte | >= |
class Version {
@new(major, minor) {
self.major = major
self.minor = minor
}
@lt(other) {
if self.major != other.major {
return self.major < other.major
}
return self.minor < other.minor
}
@gt(other) {
return other < self
}
}
echo Version(1, 2) < Version(1, 10)
echo Version(2, 0) > Version(1, 10)
true
true
Each Operator Is Separate
There is no derivation between them. Defining @lt does not give you
>, and defining @add does not give you += on the other side:
class Price {
@new(cents) {
self.cents = cents
}
@lt(other) {
return self.cents < other.cents
}
}
echo Price(250) < Price(500)
catch {
echo Price(500) > Price(250)
} as e {
echo e.message
}
true
operator '>' not defined for Price and Price
Define every operator you want to support. @gt is usually one line, as
the Version example above shows.
The Other Operand Is Not Checked
A decorated method is an ordinary method, and its parameter is an ordinary parameter. Nothing guarantees the other side is the same class:
class Amount {
@new(cents) {
self.cents = cents
}
@add(other) {
return Amount(self.cents + other.cents)
}
}
catch {
echo Amount(500) + 5
} as e {
echo '${e.type}: ${e.message}'
}
catch {
echo 5 + Amount(500)
} as e {
echo '${e.type}: ${e.message}'
}
TypeError: cannot read property 'cents' on a number
TypeError: operator '+' not defined for call signature (number, Amount)
Read those two together. With the instance on the left, @add ran and
failed inside your own method with a confusing message. With it on the
right, the operator was never dispatched to your class at all — a
decorated method only handles the case where its own instance is the left
operand.
Annotate the parameter to fix the first message, and accept the second as the rule:
class Sum {
@new(cents) {
self.cents = cents
}
@add(other: Sum) {
return Sum(self.cents + other.cents)
}
}
catch {
echo Sum(500) + 5
} as e {
echo e.message
}
@add() expects parameter 'other' (argument 1) to be a Sum, got number
Equality
@eq defines == for your class, and != is always its negation:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
@eq(other) {
if !instance_of(other, Point) {
return false
}
return self.x == other.x and self.y == other.y
}
}
var p = Point(1, 2)
echo p == Point(1, 2)
echo p != Point(1, 2)
echo p == Point(2, 1)
echo p == nil
true
false
false
false
@eq runs only when the value on the right is an object too: a string, a
list, another instance and so on. p == nil, p == 5 and p == true
compare the ordinary way without calling it, so a nil check stays a nil
check whatever the class defines. It must return a bool; anything else
raises a TypeError. Without @eq, two instances are equal only when they
are the same object.
using matches through @eq as well, since it compares the way == does.
@eq decides == and nothing else. Lists and dictionaries compare the
instances inside them by identity, and so do contains(), index_of() and
dictionary keys:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
@eq(other) {
if !instance_of(other, Point) {
return false
}
return self.x == other.x and self.y == other.y
}
}
var p = Point(1, 2)
echo [p] == [Point(1, 2)]
echo [p].contains(Point(1, 2))
echo [p].contains(p)
false
false
true
Iteration
@key and @value together make a class work with for ... in. They are
the largest of the decorated methods to get right, so they have a section
of their own: Making a Class Iterable.
| Decorator | Called by | Signature |
|---|---|---|
@key | for ... in | @key(previous) |
@value | for ... in | @value(key) |
@to_json
json.encode() calls @to_json() on an instance and encodes whatever it
returns:
import json
class User {
@new(name, password) {
self.name = name
self.password = password
}
@to_json() {
return { name: self.name }
}
}
echo json.encode(User('ada', 'hunter2'))
{"name":"ada"}
Without it, encoding an instance has nothing to work from. With it, you decide exactly what crosses the wire, which is the right place to leave a password behind.
@to_string
@to_string() decides what echo and print() show for an instance:
class Money {
@new(cents) {
self.cents = cents
}
@to_string() {
return '$' + (self.cents / 100)
}
}
var m = Money(500)
echo m
echo [m, Money(250)]
echo { total: m }
$5
[$5, $2.5]
{total: $5}
It applies wherever the instance sits in what is being shown, inside lists
and dictionaries, as a key or as a value. Without it, an instance shows as
<instance of Money>. It must return a string; anything else raises a
TypeError.
to_string()
to_string() has no @ because it is not a decorator; it is a real method
every value already has, and a class may override it to give its instances
a plain string form.
Nothing calls it for you. String interpolation and + render an instance
as <instance of Money>, so call it inside the interpolation. A class that
has both usually builds one from the other:
class Money {
@new(cents) {
self.cents = cents
}
to_string() {
return '$' + (self.cents / 100)
}
@to_string() {
return '<Money ${self.to_string()}>'
}
}
var m = Money(500)
echo m
echo 'cost: ${m.to_string()}'
<Money $5>
cost: $5
That is the split the standard library follows: to_string() is the value
as text, and @to_string() is how the instance looks when you print it.
Keep both cheap and free of side effects. Error messages, logging and debugging all reach for them.
Encapsulation and Class Immutability
Private Members
A field or method whose name starts with _ is private. It can be reached
through self or parent, and nowhere else:
class Box {
var _items = []
add(value) {
self._items.append(value)
return self
}
count() {
return self._items.length()
}
}
echo Box().add(1).add(2).count()
2
Reaching in from outside does not compile:
var b = Box()
echo b._items
SyntaxError: '_items' is private and can only be accessed via 'self' or 'parent'
--> /path/to/main.zu:2:8
|
2 | echo b._items
| ^
This is a compile-time check, not a runtime one, so it costs nothing and cannot be worked around by computing the name.
A subclass can reach its parent’s private members, through self for
fields and parent for methods:
class Base {
var _secret = 'base secret'
_internal() {
return 'base internal'
}
}
class Child < Base {
reveal() {
return self._secret + '/' + parent._internal()
}
}
echo Child().reveal()
base secret/base internal
The same underscore convention governs modules: a module member whose name
starts with _ cannot be imported by name. Chapter 8
covers that side of it.
Classes Are Sealed
Once a class is declared, its shape is final. You cannot add a field or a method to it, and you cannot add one to an instance:
class Account {
var balance = 0
}
var a = Account()
catch {
a.nickname = 'rainy day'
} as e {
echo e.message
}
undefined field 'nickname' on instance of 'Account'
The practical consequence is that a misspelled field name is an error at
the point you write it, rather than a new field that silently shadows the
one you meant. Every field a class has is declared in one place, and that
place is @new.
Static field values are mutable; the set of static fields is not.
Reflective Access
Four built-in functions read and write fields by name:
var b = Box()
echo hasprop(b, '_items')
echo getprop(b, '_items')
echo delprop(b, '_items')
echo getprop(b, '_items')
true
[]
true
nil
setprop(obj, name, value) writes; delprop resets the slot to nil.
Both return false when the field does not exist on the class, because
neither can create one.
These bypass the underscore rule, which is deliberate: they exist for serialisers, debuggers and test helpers, where reaching into an object is the entire point. Regular code should not use them.
The zuri module goes further, with a full reflection API over classes,
functions and modules. Chapter 21 covers it.
Designing With Sealed Classes
Two habits follow from sealing.
Declare every field on the class, even the ones the constructor fills
in. self.x = value inside @new will declare one for you, but a var
line at the top of the body is documentation the next reader gets for free,
and it is required the moment a helper method does the assigning instead.
Model optional state as a field holding nil, not as an absent field.
There is no such thing as an absent field, so a var cached_result that
starts nil is the shape you want.
A Worked Example
Privacy earns its place when a class has an invariant to protect. Here is a
bounded history buffer: it keeps the last n entries and nothing else, and
there is no way for a caller to break that from outside.
class History {
var _entries = []
var _limit = 0
@new(limit: number) {
if limit < 1 {
raise ValueError('limit must be at least 1, got ${limit}')
}
self._limit = limit
}
record(entry) {
self._entries.append(entry)
if self._entries.length() > self._limit {
self._entries.shift()
}
return self
}
# A copy, so a caller cannot append through the value we hand back.
entries() {
return self._entries.clone()
}
length() {
return self._entries.length()
}
}
var history = History(3)
history.record('a').record('b').record('c').record('d')
echo history.entries()
echo history.length()
var taken = history.entries()
taken.append('e')
echo history.entries()
catch {
History(0)
} as e {
echo '${e.type}: ${e.message}'
}
[b, c, d]
3
[b, c, d]
ValueError: limit must be at least 1, got 0
Four decisions are doing the work.
The list is private, so nothing outside can append to it and skip the
trimming. history._entries.append('x') does not compile.
entries() returns a clone. Without it, the caller would hold the
real list and could grow it past the limit — which is exactly what the
fourth output line shows not happening. Handing out a private mutable
collection is the most common way encapsulation leaks.
The invariant is established in the constructor. _limit is validated
once, so record() never has to wonder whether it is sensible.
The class is sealed, so history.limit = 999 is an error rather than a
second, ignored field sitting alongside _limit.
Making a Class Iterable
Any class can be walked with for ... in. It takes two decorated methods
and no other ceremony: no interface to declare, no iterator object to
return, and no state kept between passes.
This section covers the protocol, the two shapes it usually takes, what happens when a piece is missing, and the rules that keep a custom iterator well behaved.
The Protocol
for value in thing is not magic. Zuri evaluates thing once, then
repeats three steps:
key = thing.@key(key), starting fromnil- stop if
keyisnil value = thing.@value(key), and run the body
So the two methods answer two separate questions:
@key(previous)— “given the last position, what is the next one?” It receivesnilon the first call and must returnnilwhen there is nothing left.@value(key)— “what is stored at this position?”
The loop never stores a cursor of its own. The key is the cursor, which
is why @key gets the previous one back each time.
A Counting Example
class Countdown {
@new(from) {
self.from = from
}
@key(previous) {
if previous == nil {
return self.from
}
if previous <= 1 {
return nil
}
return previous - 1
}
@value(key) {
return key * 10
}
}
for value in Countdown(3) {
echo value
}
30
20
10
Trace it once and the protocol stops being abstract. @key(nil) returned
3, so the first value is @value(3), which is 30. @key(3) returned
2, then @key(2) returned 1, then @key(1) returned nil and the
loop ended.
Note that the key and the value are genuinely different things here: the key counts down from three, the value is ten times it. Two variables in the loop give you both:
for key, value in Countdown(3) {
echo '${key} -> ${value}'
}
3 -> 30
2 -> 20
1 -> 10
Wrapping a Collection
The more common case is a class holding a list, where the key is an index. The shape is always the same three checks:
class Stack {
@new() {
self.items = []
}
push(value) {
self.items.append(value)
return self
}
@key(previous) {
if self.items.is_empty() {
return nil
}
if previous == nil {
return 0
}
if previous >= self.items.length() - 1 {
return nil
}
return previous + 1
}
@value(key) {
return self.items[key]
}
}
var stack = Stack().push('a').push('b').push('c')
for value in stack {
echo value
}
for index, value in stack {
echo '${index}: ${value}'
}
a
b
c
0: a
1: b
2: c
Each of the three checks in @key earns its place:
The empty check comes first. Without it, @key(nil) would return 0
for an empty stack and @value(0) would index past the end.
previous == nil starts the walk, and 0 is the first index.
previous >= length() - 1 ends it. Using >= rather than == means a
collection that shrank mid-loop still terminates.
An empty collection iterates zero times rather than failing:
for value in Stack() {
echo 'never printed'
}
echo 'done'
done
What You Get for Free
Defining both methods makes is_iterable() answer true, and makes the
class usable everywhere for is:
echo is_iterable(Stack().push('a'))
echo is_iterable(Countdown(1))
true
true
Both Are Required
for calls @key and then @value. Defining only one produces an error
naming the one that is missing:
class OnlyKey {
@key(previous) {
return previous == nil ? 0 : nil
}
}
catch {
for value in OnlyKey() {
echo value
}
} as e {
echo '${e.type}: ${e.message}'
}
PropertyError: undefined property '@value' on instance of 'OnlyKey'
A class with neither fails on @key instead:
class Plain {}
catch {
for value in Plain() {
echo value
}
} as e {
echo '${e.type}: ${e.message}'
}
PropertyError: undefined property '@key' on instance of 'Plain'
Rules Worth Knowing
A nil key ends the loop, always. That means a collection whose keys
could legitimately be nil cannot use them as keys. Use indices and look
the real key up in @value.
The instance is evaluated once. for x in build_thing() calls
build_thing() a single time, before the first pass, so @key and
@value are always called on the same object.
Neither method should mutate. They are called once per pass, in a loop
you do not control, and a @key with a side effect is a loop whose
behaviour depends on how many times something asked for the next key.
Iteration order is whatever @key says. There is no requirement to
count upwards, or to visit everything; Countdown walks backwards, and a
class could just as well skip, filter or repeat.
Class Extensions
An extension adds methods to a class that already exists. It is written
with > rather than <, and it is a different thing from inheritance:
inheritance creates a new class that borrows from an old one, while an
extension modifies the old one in place.
class Account {
@new(owner) {
self.owner = owner
}
describe() {
return 'account for ${self.owner}'
}
}
class AccountExtras > Account {
static shout(account) {
return account.describe().upper()
}
static initials(account) {
return account.owner[0]
}
}
var a = Account('ada')
echo a.shout()
echo a.initials()
echo a.describe()
ACCOUNT FOR ADA
a
account for ada
Account gained two methods. There is no new class, no subclass, and no
wrapper object: a is the same Account instance it was, and it now
answers to shout() and initials() alongside the methods its own
declaration gave it.
< Versus >
The two are easy to tell apart once you read the arrow as pointing at where the methods end up.
| Inheritance | Extension | |
|---|---|---|
| Syntax | class Child < Parent | class Anything > Target |
| Produces | a new class | nothing; it modifies Target |
| Affects existing instances | no | yes |
| May declare fields | yes | no |
May declare @new | yes | no, it is not a field-holder |
| Methods are | ordinary methods | static, with the receiver as a parameter |
self inside a method | the instance | not available |
The Rules
Four rules are enforced at compile time, and the error messages say exactly what is wrong.
The Name Is a Label
An extension’s own name is never bound to anything. It exists so the declaration reads like a declaration and so the error messages have something to say:
class AccountExtras > Account {
static shout(account) { return 'hi' }
}
echo typeof(AccountExtras)
Unhandled UndefinedError: undefined global 'AccountExtras'
Name it after what it adds. AccountExtras, ShapeFormatting,
NodeDebugging all read well; the name appears in no other code.
Every Method Must Be static
class Target {}
class Bad > Target {
helper() {
return 1
}
}
SyntaxError: extension method 'helper' must be declared 'static' and receive the instance explicitly as their own first parameter if desired
static here does not mean the method ends up on the class rather than
on instances. It means the method is compiled without an implicit receiver,
so self does not exist inside it. The instance arrives as an ordinary
argument instead.
No Fields
class Target {}
class Bad > Target {
var cache = {}
}
SyntaxError: extension 'Bad' can only declare static methods but not fields
An extension adds behaviour, never state. A class’s set of fields is fixed when the class is declared, and nothing — including an extension — can add one afterwards. When your extension needs somewhere to keep something, keep it in the extending module, keyed by whatever identifies the instance.
The Target Must Exist
The target is evaluated when the extension is declared, so it has to be a name that is already bound to a class:
class Orphan > NoSuchClass {
static m(x) {
return 1
}
}
Unhandled UndefinedError: undefined global 'NoSuchClass'
The name does not have to be a bare one. Anything an expression starting with an identifier reaches works, so a class another module owns can be named straight through that module:
import set
class SetExtras > set.Set {
static summary(s) {
return 'set of ${s.length()}'
}
}
echo set.Set([1, 2, 3]).summary()
set of 3
Built-in types are not classes in scope, so string, list, number and
the rest cannot be extended. Error and every class the standard library
exports can be.
The Receiver
An extension method receives the instance as its first parameter. The name is yours to choose:
class Reading {
@new(celsius) {
self.celsius = celsius
}
}
class ReadingConversion > Reading {
static fahrenheit(reading) {
return reading.celsius * 9 / 5 + 32
}
static scaled(reading, factor) {
return reading.celsius * factor
}
}
var r = Reading(100)
echo r.fahrenheit()
echo r.scaled(3)
212
300
r.scaled(3) passes r as reading and 3 as factor. Every argument
at the call site shifts one place to the right of the receiver, exactly as
it would for an ordinary method.
Because arity is not enforced, the receiver parameter is optional. A method that does not need the instance can simply not declare one:
class Widget {
@new() {}
}
class WidgetVersion > Widget {
static api_version() {
return '1.0'
}
}
echo Widget().api_version()
1.0
self Does Not Exist Here
class Thing {
@new(n) {
self.n = n
}
}
class ThingBroken > Thing {
static value(t) {
return self.n
}
}
SyntaxError: 'self' used outside of a method
Use the receiver parameter. t.n, not self.n.
Private Members Stay Private
An extension is outside code, and the underscore rule applies to it in full:
class Safe {
var _hidden = 'secret'
@new() {}
}
class Peek > Safe {
static reveal(s) {
return s._hidden
}
}
SyntaxError: '_hidden' is private and can only be accessed via 'self' or 'parent'
This is the important limit on extensions. They can add behaviour built out of a class’s public surface, and they cannot reach inside it. An extension is not a way around encapsulation.
Replacing an Existing Method
An extension method with the same name as one the class already has replaces it:
class Greeter {
@new(name) {
self.name = name
}
greet() {
return 'hello ${self.name}'
}
}
var g = Greeter('ada')
echo g.greet()
class GreeterLoud > Greeter {
static greet(greeter) {
return 'HELLO ${greeter.name.upper()}'
}
}
echo g.greet()
hello ada
HELLO ADA
Read that carefully. g was constructed before the extension was
declared, and calling greet() on it after the declaration runs the new
one. The replacement is not a shadow or a wrapper: the method table of the
class itself changed, and every instance of it — past, present and future —
sees the change.
This holds no matter how thoroughly the original has already been used.
Build an eight-level tree, call the original count() four thousand times,
then replace it:
var t = tree_with(8)
var total = 0
iter var i = 0; i < 4000; i++ {
total = total + t.count()
}
echo total
class TreeNodeExt > TreeNode {
static count(node) {
return 999
}
}
echo t.count()
2044000
999
Four thousand calls to the old method, and the next call runs the new one. Replacement is unconditional: there is no warm-up state, cached lookup or earlier result that can keep an old body alive past the extension.
Decorated Methods
An extension can add decorated methods, which is how you give operators, iteration or a string form to a class you did not write. The receiver still comes first, and the decorator’s own parameters follow it:
class Box {
@new(n) {
self.n = n
}
}
class BoxOps > Box {
static @add(a, b) {
return Box(a.n + b.n)
}
static @key(box, previous) {
return previous == nil ? 0 : nil
}
static @value(box, key) {
return box.n
}
static to_string(box) {
return 'Box(${box.n})'
}
}
echo (Box(2) + Box(3)).n
echo Box(9).to_string()
for value in Box(7) {
echo value
}
5
Box(9)
7
@add(a, b) gets the left operand as a and the right as b.
@key(box, previous) gets the instance and then the previous key, which is
the argument @key would normally take on its own.
They Are Instance Methods, Not Statics
static in the declaration describes how the method is compiled, not
where it lands. The methods go onto instances:
class Item {
@new() {}
}
class ItemExtras > Item {
static label(item) {
return 'an item'
}
}
echo Item.label(Item())
Unhandled PropertyError: undefined static member 'label' on class 'Item'
Call it on the instance: Item().label().
Extensions and Inheritance
The two interact in one way worth knowing, and the rule is about when each declaration runs.
An extension changes the target’s method table. A subclass copies its parent’s methods when the subclass is declared. So a subclass sees an extension only if the extension came first:
class Vehicle {
@new(wheels) {
self.wheels = wheels
}
}
class VehicleExtras > Vehicle {
static doubled(vehicle) {
return vehicle.wheels * 2
}
}
class Car < Vehicle {
@new() {
parent(4)
}
}
echo Vehicle(2).doubled()
echo Car().doubled()
4
8
Car was declared after the extension, so it inherited doubled. Move the
class Car declaration above the extension and Car().doubled() raises
undefined property 'doubled' on instance of 'Car', while Vehicle keeps
it.
Declare extensions before the subclasses that should inherit them. In practice that means at the top of a module, or in a module imported at the top.
Extending a subclass never touches its parent:
class Animal {
@new() {}
}
class Dog < Animal {
@new() {}
}
class DogExtras > Dog {
static speak(dog) {
return 'woof'
}
}
echo Dog().speak()
catch {
echo Animal().speak()
} as e {
echo e.message
}
woof
undefined property 'speak' on instance of 'Animal'
Extensions Are Global
This is the property that makes extensions powerful and the one that makes them worth using sparingly. An extension is not scoped to the module that declares it. It changes the class everywhere in the program, and it takes effect the moment the declaring module runs — which, for an imported module, is the moment it is imported.
Filename: shape.zu
class Shape {
@new(name) {
self.name = name
}
}
Filename: extras.zu
import .shape { Shape }
class ShapeExtras > Shape {
static label(s) {
return 'shape: ${s.name}'
}
}
Filename: main.zu
import .shape { Shape }
var s = Shape('circle')
catch {
echo s.label()
} as e {
echo 'before import: ${e.message}'
}
import .extras
echo s.label()
echo Shape('square').label()
before import: undefined property 'label' on instance of 'Shape'
shape: circle
shape: square
main.zu never mentions label’s definition. Importing extras — for any
reason, including a reason unrelated to Shape — added a method to a class
declared in a third file, and the instance created before the import
gained it too.
Two consequences follow.
An import can change behaviour you did not ask it to change. A module that extends a class you use will alter it for your code as well, and nothing at your call site says so.
Two extensions of the same method collide silently. The last one to run wins, and the winner depends on import order:
class Slot {
@new() {}
}
class First > Slot {
static which(s) {
return 'first'
}
}
class Second > Slot {
static which(s) {
return 'second'
}
}
echo Slot().which()
second
When to Use One
Extensions answer a question inheritance cannot: how do you add behaviour to a class whose instances are created somewhere you do not control? A subclass only helps if you are the one calling the constructor.
Good reasons:
Adding a view or a format to a domain class. to_string(), a
@to_json, a summary() — presentation that does not belong in the model
itself, added from the module that cares about it.
Giving a standard library class an operator or an iterator so it works with syntax it was not written for.
Instrumenting during debugging. Replacing a method with one that logs, then deleting the extension, is a diagnostic technique that costs nothing in the original file.
Reasons to think twice:
It is invisible at the call site. account.shout() gives no hint that
shout lives in a different file from Account. A reader who greps the
class body will not find it.
It is global. Everything above.
A plain function is usually enough. shout(account) is a function, it
is obvious where it lives, and it cannot collide with anything. Reach for
an extension when the thing genuinely needs to be a method — because
syntax demands it, as with a decorated method, or because callers already
have the instance and nothing else.
Error Handling
Things go wrong. A file is not there, a number arrives as text, a network
peer stops answering. Zuri has one mechanism for all of it: an Error
object, raised with raise and intercepted with catch. There is no
try, no finally, and no checked exceptions.
This chapter covers raising, the three shapes of catch, what an error
carries, the built-in hierarchy, writing your own, and the patterns that
replace finally.
Raising
def withdraw(balance, amount) {
if amount > balance {
raise ValueError('cannot withdraw ${amount} from ${balance}')
}
return balance - amount
}
raise takes an instance of Error or any subclass, and nothing else. A
bare string or number is a TypeError in its own right:
catch {
raise 'something went wrong'
} as e {
echo e.message
}
can only raise an Error or subclass, got a string
That rule is worth the small inconvenience: every value that travels
through the error system has a type, a message and a stack trace,
because there is no way to put anything else in.
An error that nobody catches ends the program with a message, a source excerpt and a stack trace:
Unhandled ValueError: cannot withdraw 100 from 50
--> /path/to/main.zu:3
1 | def withdraw(balance, amount) {
2 | if amount > balance {
> 3 | raise ValueError('cannot withdraw ${amount} from ${balance}')
4 | }
5 | return balance - amount
Stack trace (most recent call last):
at withdraw() /path/to/main.zu:3
at @.script() /path/to/main.zu:8
Catching
catch {
risky()
} as e {
echo '${e.type}: ${e.message}'
}
The catch block runs. If anything inside it raises, execution jumps to
the handler with the error bound to e. If nothing raises, the handler
never runs.
There are three shapes, and each does something different.
catch { ... } with nothing after it swallows the error and carries
on:
catch {
raise Error('silent')
}
echo 'swallowed'
swallowed
Use this when failure genuinely does not matter, and nowhere else.
catch { ... } as e with no handler block binds the error to a
variable that survives the statement. e is nil when nothing went
wrong:
catch {
raise Error('boom')
} as e
if e {
echo e.message
}
boom
This is the shape to use when the recovery does not belong inside a handler, for example when you want to check several things in a row.
catch { ... } as e { ... } is the full form, and the one you will
write most.
What an Error Carries
catch {
raise ValueError('bad input')
} as e {
echo e.type
echo e.message
echo e.stacktrace
}
ValueError
bad input
[/path/to/main.zu:2 -> @.script()]
messageis the text.typeis the class name, as a string.stacktraceis a list of frames, innermost first, each naming a file, a line and a function.
The stack trace is captured where the error was raised, not where it was caught, so it points at the origin no matter how many frames it travelled through:
def inner() {
raise ValueError('deep')
}
def middle() {
inner()
}
catch {
middle()
} as e {
echo e.stacktrace.length()
}
3
Three frames: inner, middle, and the script’s top level. Printing them
is often the fastest way to answer “how did we get here?” in code you did
not write.
The Built-in Errors
Every one of these is a class, and every one inherits from Error:
| Class | Raised when |
|---|---|
Error | the base; a general failure |
TypeError | an operation got the wrong type |
ValueError | the type was right, the value was not |
NumericError | an arithmetic operation failed |
ArgumentError | wrong number of arguments |
NotImplementedError | a method that must be overridden was not |
RangeError | an index or bound was out of range |
AccessError | a permission or access check failed |
AssertError | an assert failed |
PropertyError | a member that does not exist was read |
UndefinedError | an undefined name was read |
ModuleNotFoundError | an import could not be resolved |
Because they all inherit from Error, catching Error catches everything,
and a parameter annotated Error accepts any of them.
Custom Errors
Subclass Error and carry whatever the caller needs:
class HttpError < Error {
@new(message, status) {
parent(message)
self.type = 'HttpError'
self.status = status
}
}
catch {
raise HttpError('not found', 404)
} as e {
echo '${e.type} ${e.status}: ${e.message}'
echo instance_of(e, Error)
}
HttpError 404: not found
true
Two things make this work well. Call parent(message) so the base
constructor sets message and captures the stack trace. Set self.type so
the class name shows up in logs and in the uncaught-error banner.
Deciding Which Error to Catch
catch catches everything inside its block. To handle one kind and let the
others through, test and re-raise:
catch {
load_config()
} as e {
if !instance_of(e, ModuleNotFoundError) {
raise e
}
echo 'no config, using defaults'
}
instance_of() walks the inheritance chain, so a test against Error
matches everything and a test against HttpError matches only that branch.
There Is No finally
Code after the catch statement runs whether the block raised or not,
because the handler either recovers or re-raises:
def cleanup_demo() {
catch {
raise Error('mid')
} as e {
echo 'caught'
}
echo 'always runs'
}
cleanup_demo()
caught
always runs
That covers the common case. When the handler re-raises, or when the block
contains a return, the trailing code is skipped, so a resource that must
be released either way goes in the handler as well:
var handle = file(path, 'w')
catch {
write_everything(handle)
} as e {
handle.close()
raise e
}
handle.close()
The Bare as e Form Does It Once
The version above repeats handle.close(), once in the handler and once
after the statement, because those are two different paths out. The third
shape of catch — as e with no handler block — collapses them into
one:
var handle = file(path, 'w')
catch {
write_everything(handle)
} as e
handle.close()
if e {
raise e
}
The difference is the missing { ... } after as e, and it changes the
control flow rather than just the layout. With no handler there is nothing
to jump into, so a raise inside the block is recorded in e and
execution simply continues on the next line. Both paths — the one that
raised and the one that did not — now run the same trailing code:
fallthrough: closing
fallthrough: re-raising
caught: disk full
That gives you close() written once, running unconditionally, followed by
an explicit decision about whether to re-raise. It is the closest thing
Zuri has to finally, and it is assembled out of the ordinary pieces
rather than being a separate construct:
| Line | Job |
|---|---|
catch { ... } as e | run it, record any failure, do not jump |
handle.close() | the cleanup, on every path |
if e { raise e } | pass the failure on, unchanged |
Two things to be deliberate about.
e is nil when nothing went wrong, which is what makes
if e { raise e } the whole of the decision. On the success path the
cleanup runs and the function carries on normally.
Re-raising e itself keeps the original stack trace, pointing at the
line that actually failed rather than at the raise you just wrote. Wrap
it in a new error only when you have something to add, as
Nesting and Re-raising shows.
Use the handler form when the failure needs handling. Use this form when it needs only cleaning up after.
return Inside catch
A return from inside a catch block or its handler leaves the enclosing
function, exactly as it would anywhere else:
def find(items, needle) {
catch {
var i = items.index_of(needle)
if i == -1 {
raise ValueError('${needle} not in list')
}
return items[i]
} as e {
echo 'handled: ' + e.message
return nil
}
}
echo find([1, 2], 2)
echo find([1, 2], 9)
2
handled: 9 not in list
nil
Nesting
Handlers can raise, and an outer catch will see it:
catch {
catch {
raise Error('inner')
} as inner {
raise Error('outer: ' + inner.message)
}
} as outer {
echo outer.message
}
outer: inner
assert Versus raise
assert items.length() > 0, 'caller must pass a non-empty list'
assert raises an AssertError when its condition is falsy. The message
is optional.
Draw the line like this. raise is for things that happen: a file that
is not there, a number that does not parse, a network that times out. The
caller is expected to handle it. assert is for things that cannot
happen: an invariant the code itself is responsible for maintaining. A
failed assertion means the program has a bug, not that the world was
uncooperative.
Style
Keep the catch block small. The block should contain the operation that
can fail and nothing else, so the handler is not accidentally catching a
mistake somewhere further down:
Too wide — a failure inside render() is reported as a parse failure:
def show_wide(text) {
catch {
var data = json.decode(text)
render(data)
} as e {
echo 'bad json'
}
}
Right — only the call that can fail is inside the block:
def show_narrow(text) {
var data
catch {
data = json.decode(text)
} as e {
echo 'bad json: ' + e.message
return
}
render(data)
}
Note the var data outside the block. A catch block is a scope like any
other, so a variable declared inside it is gone by the time the next
statement runs.
Write messages that name the value:
raise ValueError('port must be between 1 and 65535, got ${port}')
The person reading that message is trying to work out what went wrong from one line of a log file. Give them the number.
Catching Inside a Loop
A catch inside a loop body handles one iteration and lets the rest carry
on. This is the shape for processing a batch where individual items are
allowed to fail:
def parse_positive(text) {
if !text.match('/^\d+$/') {
raise ValueError('not a number: ${text}')
}
var n = text.to_number()
if n <= 0 {
raise ValueError('must be positive: ${text}')
}
return n
}
var inputs = ['12', 'not a number', '30']
var total = 0
var rejected = []
for raw in inputs {
catch {
total += parse_positive(raw)
} as e {
rejected.append(raw)
}
}
echo total
echo rejected
42
[not a number]
Put the catch outside the loop instead, and the first failure ends
the whole loop — which is the right choice when one bad item makes the
rest meaningless, and the wrong one when it does not. The placement of the
block is the decision; there is no flag to set.
Nesting and Re-raising
A catch inside a handler works like any other, which is how you translate
a low-level failure into one your caller understands:
import json
class ConfigError < Error {
@new(message) {
parent(message)
self.type = 'ConfigError'
}
}
def load(text) {
catch {
return json.decode(text)
} as e {
raise ConfigError('config is not valid JSON: ${e.message}')
}
}
catch {
load('{ broken')
} as e {
echo '${e.type}: ${e.message}'
}
The caller now gets an error in its own vocabulary. Include the original message, as above, so the detail is not lost on the way up.
Re-raising the same error, rather than a new one, keeps the original stack trace pointing at the original line:
def only_handle_missing(work) {
catch {
return work()
} as e {
if !instance_of(e, ModuleNotFoundError) {
raise e
}
return 'defaulted'
}
}
echo only_handle_missing(@() => 'fine')
catch {
only_handle_missing(@() { raise ValueError('not mine') })
} as e {
echo '${e.type}: ${e.message}'
}
fine
ValueError: not mine
That is the pattern for “handle one kind and let everything else through”, and it is worth reaching for whenever a handler would otherwise swallow a bug along with the failure it meant to catch.
A Worked Example
Here is the whole chapter in one function: a loader that validates its input, distinguishes the failures a caller can act on from the ones it cannot, and closes what it opened on every path.
import json
class StoreError < Error {
@new(message) {
parent(message)
self.type = 'StoreError'
}
}
def read_records(path) {
var handle = file(path)
if !handle.exists() {
raise StoreError('no store at ${path}')
}
var records
catch {
records = json.decode(handle.read())
} as e {
handle.close()
raise StoreError('${path} is corrupt: ${e.message}')
}
handle.close()
if !is_list(records) {
raise StoreError('${path} should hold a list, found ${typeof(records)}')
}
return records
}
file('records.json', 'w').write('[{"id": 1}]')
echo read_records('records.json').length()
file('records.json', 'w').write('{ not json')
catch {
read_records('records.json')
} as e {
echo e.type
}
file('records.json').delete()
catch {
read_records('records.json')
} as e {
echo '${e.type}: ${e.message}'
}
1
StoreError
StoreError: no store at records.json
Four things in there are worth naming.
Every failure the caller might reasonably handle arrives as one error type,
StoreError, so catch on the calling side needs one branch rather than
three.
The message always names the path. A log line saying “file is corrupt” with no filename costs someone an hour.
handle.close() appears on both paths — once in the handler before the
re-raise, once after the block. There is no finally to do it for you, and
forgetting the one in the handler is the most common resource leak in Zuri
code.
And the shape check at the end is a raise, not an assert. A file on
disk containing the wrong thing is the world being uncooperative, not a bug
in this function.
The Module System
A module is a .zu file. It runs at most once per program, it gets its own
global namespace, and nothing it declares is visible anywhere else until
someone imports it.
That is the whole model. The rest of this chapter is the syntax and the resolution rules.
Importing
import math
import os
echo math.PI
echo os.cwd()
import name binds the module under that name. Reach into it with a dot.
Picking Out Members
import math { PI, E }
echo PI
The named members are bound directly, and the module name itself is not.
{ * } brings in everything public:
import math { * }
echo PI
echo ROOT_2
Renaming
import http.websocket as ws
Use this when the natural name is long, or when it would collide.
as renames the module, not a member. There is no way to rename an
individual name in a member list — import math { PI as pi } does not
parse. When you want a different name for one imported thing, bind it
yourself:
import math { PI }
var pi = PI
echo pi
3.141592653589793
Relative Imports
A path starting with . or .. is resolved against the directory of the
file doing the importing, and is never searched for anywhere else.
import .helpers # next to this file
import .models.user # models/ next to this file, then user
import ..shared.config # up one directory, then shared/, then config
import .. ..shared.config # up two directories, then shared/, then config
Each .. climbs one directory, so .. .. climbs two and .. .. .. three.
The dots may also be written together, ....shared.config, and it is the
same import: a run of dots counts two for each directory up. zuri fmt
leaves whichever way the path is written as it is.
Each segment resolves the same way a bare name does: name.zu is tried
first, then name/index.zu. So import .helpers finds either
helpers.zu or helpers/index.zu, and import .models.user finds
models/user.zu or models/user/index.zu, with models itself being a
directory either way.
Which one it finds is invisible at the import site, and that is what lets a
module grow into a package: split helpers.zu into helpers/index.zu plus
some siblings, and every import .helpers keeps working untouched.
Inside a package, a sibling is always import .sibling, never the full
path from the project root. Writing import myapp.models.user from inside
myapp/models/ sends the resolver out to the library search path and it
will not find anything.
Exporting
Imports are local by default. If a.zu imports b.zu, a third file
that imports a does not see b’s contents through it.
Prefix the path with @ to re-export:
import @.util { * }
Now everything util.zu exposed is part of this module’s public surface
too. All three forms take the prefix:
import @.module # the module itself is re-exported
import @.module { item } # just that item
import @.module { * } # everything
This is how a package’s index.zu assembles a public API out of several
private files:
Filename: pkg/index.zu
import @.util { * }
import .sub.deep { deep_slug }
def hello() {
return 'hello from pkg ' + VERSION
}
util’s members are re-exported. deep_slug is imported for this file’s
own use and stays private.
Privacy
A member whose name starts with _ is private to its module and cannot be
imported by name:
import .pkg.util { _secret }
SyntaxError: Cannot import private items from module
--> /path/to/bad.zu:1:20
|
1 | import .pkg.util { _secret }
| ^
{ * } skips private members too. Prefix anything that is an
implementation detail and the module system will keep it that way.
Packages
A directory with an index.zu is a package, and importing the directory
runs its index.zu:
pkg/
index.zu
util.zu
sub/
deep.zu
import .pkg # runs pkg/index.zu
import .pkg.util # runs pkg/util.zu
import .pkg.sub.deep # runs pkg/sub/deep.zu
The same rule applies to zuri run pkg on the command line, which is what
makes a package runnable as well as importable.
How a Bare Name Is Resolved
import http, with no leading dot, is searched for in this order:
.zuri/libs/httpin the project$ZURI_ROOT/libs/http, orlibs/httpbeside thezuriexecutable: the standard library- a built-in native module named
http $ZURI_HOME/libs/http, the packages installed for your user, which is~/.zuri/libsunlessZURI_HOMEmoves it
At each step, http.zu is tried first and then http/index.zu.
The project is the nearest directory holding a project.toml, looking
upwards from the script zuri run was given, or from the working
directory for a command or the REPL. A program run from anywhere inside
a project imports that project’s packages, and a program in no project
uses .zuri/libs in the working directory. The project is found once,
when the program starts, so changing directory part way through does
not change where imports come from.
zuri install fills .zuri/libs, and
Packages and Nyssa covers it in full. Step one is
also what makes vendoring work. Dropping a file into .zuri/libs/
shadows a standard library module of the same name:
$ cat .zuri/libs/mylib.zu
def hi() { return 'from user libs' }
$ cat uses.zu
import mylib
echo mylib.hi()
$ zuri run uses.zu
from user libs
Modules Run Once
A module’s top level executes the first time it is imported, and never again. Every later import of the same file gets the same module object:
import .once
import .once as again
import .once
side effect ran
One line of output, three imports. Identity is by canonical filesystem path, so two different relative paths to the same file are the same module.
This makes a module’s top level the natural place for setup that must happen exactly once: opening a connection pool, reading a config file, registering handlers.
Circular Imports
A circular import works. A module is registered before its body runs, so
when b.zu imports a.zu while a.zu is still loading, it gets the
partially built module rather than looping forever:
Filename: a.zu
import .b
def from_a() {
return 'a'
}
echo 'a loaded'
Filename: b.zu
import .a
echo 'b loaded'
b loaded
a loaded
a
The catch is visible in that output: b finished loading before a did,
so anything b reads from a at its top level is not there yet.
Reading it from inside a function is fine, because by then a has
finished.
If a module fails while loading, it is dropped from the cache rather than left behind half built, so a later import genuinely retries.
Module Variables
Every module gets two names for free:
echo __file__
echo __root__
__file__ is this module’s own canonical path. __root__ is the entry
file the program was started from, and it is the same in every module. Use
them to locate files relative to your source rather than relative to
whatever directory the user happened to run from:
import os
var templates = os.join_paths(os.dir_name(__file__), 'templates')
In the REPL both are defined, with placeholder values standing in for the file that does not exist:
%> __file__
@.repl
%> __root__
@.repl.root
They differ from each other there, so the __root__ == __file__ check
below is false at the prompt — a REPL session is never the entry point of
a program.
Running as a Program, Importing as a Module
Because __root__ is the entry file and __file__ is this one, comparing
them tells a module whether it is the program being run or something being
imported:
Filename: tool.zu
def add(a, b) {
return a + b
}
if __root__ == __file__ {
echo 'running as a program: ' + add(2, 3)
}
$ zuri run tool.zu
running as a program: 5
import .tool
echo tool.add(10, 20)
30
The same file is a clean, side-effect-free library when imported and a runnable command-line program when launched directly. Put the argument parsing and the entry point behind that check, and everything else above it.
These two names describe which file this is, so they are not exports. A
wildcard import copies every public name out of the module it names, and
__file__ and __root__ are deliberately excluded from that:
import math { * }
echo __file__ == __root__
true
That guarantee is what makes the __root__ == __file__ check above
reliable in every file, including one that wildcard-imports a sibling.
Structuring a Project
A small program is one file. Past that, the shape that works is a package
per area of responsibility, each with an index.zu that re-exports what is
public:
myapp/
index.zu # import @.routes, import @.models, then start
config.zu
models/
index.zu # import @.user { * }, import @.task { * }
user.zu
task.zu
routes/
index.zu
api.zu
pages.zu
storage/
index.zu
_json_store.zu # private: the leading underscore says so
Run it with zuri run myapp. The capstone in Chapter 25
is laid out exactly this way.
A Package, End to End
Here is the smallest complete version of that shape. Three files, one package, one public entry point.
Filename: greet/english.zu
def hello(name) {
return 'Hello, ${name}'
}
def _shout(text) {
return text.upper()
}
Filename: greet/french.zu
def hello(name) {
return 'Bonjour, ${name}'
}
Filename: greet/index.zu
import @.english
import @.french
def greet(name, language) {
return language == 'fr' ? french.hello(name) : english.hello(name)
}
Filename: index.zu
import .greet
echo greet.greet('Ada', 'en')
echo greet.greet('Ada', 'fr')
echo greet.english.hello('Grace')
$ zuri run
Hello, Ada
Bonjour, Ada
Hello, Grace
Four things to take from it.
greet/index.zu is what import .greet loads. A directory with an
index.zu is a package, and importing the directory runs that file.
The @ on import @.english is what re-exports it. Without it,
greet.english would not be reachable from outside greet/index.zu, even
though greet() itself would still work — which is often exactly what you
want.
Both submodules define hello, and they do not collide. Each lives in
its own namespace, reached through its own module name. That is the whole
reason to use import @.english rather than import @.english { * } here.
_shout is unreachable from outside. Writing greet.english._shout(x)
anywhere else is a compile error, not a runtime one:
SyntaxError: '_shout' is private and can only be accessed via 'self' or 'parent'
The leading underscore is the only declaration of privacy there is, and it is checked before the program runs.
Files and the Filesystem
Two things do the work here. The built-in file() function gives you a
handle to one file. The os module covers everything else: directories,
paths, globs, temporary files and the environment.
Opening a File
var handle = file('notes.txt')
var writer = file('notes.txt', 'w')
file(path) defaults to read mode. The second argument is the mode:
| Mode | Meaning |
|---|---|
r | read; the file must exist |
w | write; creates the file, truncates an existing one |
a | append; writes always go to the end, creates the file |
r+ | read and update; the file must exist |
w+ | read and update; creates the file, does not truncate |
a+ | read and append; creates the file |
x | write; creates the file, fails if the path already exists |
x+ | read and write; creates the file, fails if the path already exists |
Append b to any of them for binary mode: 'rb', 'wb', 'ab'.
x is the one to reach for when two programs might create the same file.
The existence check and the creation happen in a single step, so exactly
one of them succeeds and every other one fails, which is what a lock file
needs.
Creating a handle does not touch the disk. Nothing happens until you read,
write or call open().
Reading a Whole File
file('notes.txt', 'w').write('It works!')
echo file('notes.txt').read()
It works!
read() with no argument opens the file, reads all of it, and closes it
again. That is the one-liner for “give me this file’s contents”, and it is
what you want most of the time.
In text mode you get a string, decoded strictly as UTF-8. In binary mode
you get bytes.
Reading in Chunks
read(length) reads at most that many bytes and leaves the handle open, so
you can call it again. This is how you process a file too large to hold in
memory at once:
file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')
var handle = file('notes.txt')
handle.open()
var chunks = 0
while true {
var chunk = handle.read(6)
if chunk.is_empty() {
break
}
chunks++
}
handle.close()
echo chunks
3
Six bytes at a time is a demonstration; in real code the chunk is tens of kilobytes. The shape is what matters: open once, read until you get an empty result, close once.
Note that chunks fall wherever the byte count lands, not on line boundaries. A chunked reader that needs whole lines has to keep the tail of each chunk and join it to the front of the next.
Reading Lines
For text you want line by line, the simplest form reads the file and splits it:
file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')
for line in file('notes.txt').read().lines() {
echo '[${line}]'
}
[alpha]
[beta]
[gamma]
lines() handles both \n and \r\n, and drops the trailing empty
piece a final newline would otherwise produce.
Writing
file('notes.txt', 'w').write('It works!')
Like read(), write() opens the handle if it is closed, writes, flushes,
and closes it again. One call, one complete file.
That auto-close has a consequence worth being precise about. Two
consecutive write() calls on a closed handle in w mode each truncate
the file, so only the last one survives. When you are writing more than
once, open the handle yourself:
var handle = file('notes.txt', 'w')
handle.open()
handle.write('first line\n')
handle.write('second line\n')
handle.close()
echo file('notes.txt').read()
first line
second line
The explicit open() is what keeps the handle open across both writes.
Without it, each write() would open, truncate, write and close, and only
second line would survive.
An already-open handle is written to where it stands, so the sequence above does exactly what it reads like.
puts() writes without ever opening or closing. It requires an open
handle, and it is the method to reach for inside a loop.
Reading and Writing at a Position
var handle = file('notes.txt')
handle.open()
handle.seek(6, 0)
echo handle.tell()
echo handle.read(4)
handle.close()
seek(offset, whence) takes 0 for the start of the file, 1 for the
current position and 2 for the end. The io module names them:
import io
handle.seek(0, io.SEEK_SET)
handle.seek(-10, io.SEEK_END)
tell() reports the current offset.
Asking About a File
var handle = file('notes.txt')
echo handle.exists()
echo handle.path()
echo handle.abs_path()
echo handle.name()
echo handle.mode()
echo handle.is_open()
echo handle.is_closed()
stats() returns a dictionary of metadata about the file on disk:
file('notes.txt', 'w').write('alpha\nbeta\n')
var info = file('notes.txt').stats()
echo info.size
echo info.is_readable
echo info.keys()
11
true
[is_readable, is_writable, is_executable, is_symbolic, size, mode, dev, ino, nlink, uid, gid, mtime, atime, ctime, blocks, blksize]
size is in bytes. mtime, atime and ctime are epoch seconds, ready
to hand to the date module. mode is the raw permission-and-type word,
and the stat module is what turns it into an answer:
import stat
file('notes.txt').chmod(0c644)
var info = file('notes.txt').stats()
echo stat.S_ISREG(info.mode)
echo stat.S_ISDIR(info.mode)
echo stat.file_mode(info.mode)
true
false
-rw-r--r--
S_ISREG, S_ISDIR, S_ISLNK and the rest of the family each answer one
question about the kind of entry. file_mode() renders the permission bits
the way ls -l does.
Managing Files
file('notes.txt').copy('backup.txt')
file('backup.txt').rename('archive.txt')
file('archive.txt').delete()
There is also truncate(length), chmod(mode), set_times(access, modify) and symlink(target).
Always Close What You Opened
read() and write() clean up after themselves. Anything you opened with
open() is yours to close, and the cleanest way to guarantee it is a
catch that closes on the way out:
var handle = file(path, 'w')
handle.open()
catch {
write_everything(handle)
} as e
handle.close()
if e {
raise e
}
For something that has to be cleaned up however the program ends, rather
than however one block ends, register it with os.at_exit():
import os
var scratch = file('scratch.txt', 'w')
scratch.open()
os.at_exit(@{
scratch.close()
scratch.delete()
})
scratch.write('working notes')
echo scratch.is_open()
true
Handlers run last registered first, and they run whether the program
reached the end of its script, called os.exit(), or stopped on an
uncaught error. A handler that raises is reported on standard error and
the rest still run, so one failed cleanup cannot cancel the others.
Directories
import os
os.create_dir('sub/deep', nil, true)
The three arguments are the path, the permission bits, and whether to
create intermediate directories. It returns false when the directory
already existed.
echo os.dir_exists('sub')
echo os.is_dir('sub')
echo os.remove_dir('sub', true)
remove_dir’s second argument makes it recursive.
Listing and Globbing
Everything in this section works on a real tree, so build one first:
import os
os.create_dir('tree/sub', nil, true)
file('tree/top.txt', 'w').write('a')
file('tree/sub/nested.txt', 'w').write('b')
file('tree/sub/notes.md', 'w').write('c')
read_dir() lists one directory:
echo os.read_dir('tree')
echo os.read_dir('tree', true)
[., .., sub, top.txt]
[., .., sub, sub/nested.txt, sub/notes.md, top.txt]
Three things to notice. . and .. are included, so a loop over the result
almost always wants to skip them. The recursive form returns nested entries
as paths relative to the directory you asked about, not as bare names,
which is what makes them usable directly. And entries come back sorted by
name, with a directory’s contents following immediately after it, so the
listing reads the same on every machine and filesystem.
glob() is usually what you actually want:
echo os.glob('*.txt', 'tree')
echo os.glob('**/*.txt', 'tree')
[top.txt]
[sub/nested.txt]
* matches within one path segment; ** matches across segments. The
second argument is the base directory, and results come back relative to
it.
Read those two results together, because the distinction catches people
out. *.txt found the file at the top and not the nested one. **/*.txt
found the nested one and not the top-level one, because **/ means
“in a subdirectory”. Neither pattern finds both.
To match at every depth, glob ** and filter:
echo os.glob('**', 'tree')
echo os.glob('**', 'tree').filter(@(p) => p.ends_with('.txt'))
[sub, sub/nested.txt, sub/notes.md, top.txt]
[sub/nested.txt, top.txt]
** on its own matches every entry at every depth, directories included,
which is why the filter is doing real work in the second line. glob()
walks with read_dir() underneath, so matches arrive in that same sorted
order.
Clean up when you are done:
echo os.remove_dir('tree', true)
true
Paths
Every path function is pure string manipulation except where noted:
echo os.join_paths('sub', 'a.txt')
echo os.base_name('sub/a.txt')
echo os.dir_name('sub/a.txt')
echo os.real_path('sub')
echo os.relative_path(os.cwd(), '/full/path/to/sub')
sub/a.txt
a.txt
sub
/home/you/project/sub
sub
real_path() resolves symlinks and requires the path to exist.
abs_path() does not. expand_user() turns a leading ~ into the home
directory. path_contains(base, candidate) answers whether one path is
inside another, which is the check you need before serving a file a user
named.
echo os.cwd()
echo os.home_dir()
os.change_dir('/some/where')
Temporary Files
echo os.temp_dir()
var path = os.create_temp_file('report-', '.csv')
var dir = os.create_temp_dir('build-')
Both create the thing and hand you its path. Clean them up yourself when you are done.
Environment Variables
echo os.get_env('HOME')
echo os.get_env('NOPE', 'fallback')
os.set_env('ZURI_BOOK', '1')
os.unset_env('ZURI_BOOK')
echo os.environ()
echo os.expand_vars('$HOME/projects')
get_env() takes a fallback. environ() gives the whole set as a
dictionary.
Locating Files Relative to Your Code
The current working directory is wherever the user ran zuri from, which
is not where your source lives. Use __file__:
import os
var HERE = os.dir_name(__file__)
var templates = os.join_paths(HERE, 'templates')
Doing this in every module that reads a file next to itself is the difference between a program that works and one that works only from the project root.
A Worked Example
Counting words across every text file in a directory tree, start to finish:
import os
def text_files(root) {
return os.glob('**', root).filter(@(p) => p.ends_with('.txt'))
}
def word_count(root) {
var counts = {}
for path in text_files(root) {
var text = file(os.join_paths(root, path)).read()
for word in text.lower().split('/\W+/') {
if word.is_empty() {
continue
}
counts.set(word, counts.get(word, 0) + 1)
}
}
return counts
}
os.create_dir('wc/sub', nil, true)
file('wc/a.txt', 'w').write('Hello world, hello!')
file('wc/sub/b.txt', 'w').write('World of Zuri.')
file('wc/sub/skip.md', 'w').write('not counted')
echo word_count('wc')
os.remove_dir('wc', true)
{hello: 2, world: 2, of: 1, zuri: 1}
Four things in that function are worth pointing at.
glob('**', root) returns paths relative to the root, so they have to
be joined back onto it before they can be opened. Forgetting that join is
the most common mistake in code that globs.
The .md file is absent from the result because text_files() filtered it
out, which is the filter doing the job ** alone cannot.
split('/\W+/') is a regular expression, which is why punctuation does not
end up in the keys — world, and hello! became world and hello. A
plain split(' ') would have kept both.
And counts.get(word, 0) + 1 supplies the starting value for a key that
does not exist yet. Without the fallback, the first sighting of every word
would raise.
Where to Go Next
Binary files, byte streams and the io module’s in-memory files are
Chapter 10. Reading a file over the network is
Chapter 12. The complete list of methods a file
handle carries is Appendix E.
Binary Data and Streams
Text is convenient. Protocols, file formats, images and checksums are not
text, and this chapter is about the four tools Zuri gives you for them:
the bytes type, the struct module, io.BytesIO, and compress.
The bytes Type
bytes is a mutable buffer of 8-bit values.
var b = bytes(3)
var c = bytes([1, 2, 3, 4, 5])
var d = 'Hi'.to_bytes()
echo b
echo c
echo d
(00 00 00)
(01 02 03 04 05)
(48 69)
bytes(n) allocates n zero bytes. bytes(list) builds one from numbers,
each of which must be in 0..256. to_bytes() on a string gives you its
UTF-8 encoding.
The hexadecimal-in-parentheses rendering is how bytes always prints,
which makes it impossible to confuse one with a list at a glance.
Reading
var b = bytes([1, 2, 3, 4, 5])
echo b.length()
echo b[0]
echo b[1, 3]
echo b.first()
echo b.last()
echo b.get(2)
echo b.index_of(3)
echo b.last_index_of(3)
echo b.to_list()
echo 'Hi'.to_bytes().to_string()
5
1
(02 03)
1
5
3
2
2
[1, 2, 3, 4, 5]
Hi
Indexing gives a number. Slicing gives bytes. to_string() decodes
as UTF-8, to_list() gives numbers.
index_of() and last_index_of() both return -1 when the byte is not
there, and both take a second argument bounding where a match may sit,
so they search the two halves either side of one index.
Slicing
b[a, b] takes the bytes from a up to but not including b, and
returns a new byte stream:
var b = bytes([10, 20, 30, 40, 50])
echo b[1, 3]
echo b[, 3]
echo b[3, ]
echo b[-2, ]
(14 1e)
(0a 14 1e)
(28 32)
(28 32)
The rules are exactly the list’s. Either bound may be omitted: b[, n]
starts at the beginning and b[n, ] runs to the end. Negative bounds count
back from the end, so b[-2, ] is the last two bytes.
Note the difference between an index and a slice, because for bytes the
two return different types:
var b = bytes([10, 20, 30])
echo b[0]
echo typeof(b[0])
echo b[0, 1]
echo typeof(b[0, 1])
10
number
(0a)
bytes
One index gives you the numeric value of that byte. A slice of length one
gives you a byte stream containing it. Reaching for b[0] when you meant
b[0, 1] is the most common slip here, and it shows up as a number where
a bytes was expected rather than as an error at the slicing site.
Bounds Are Checked
A slice that runs past the end raises rather than returning what it can:
var b = bytes([10, 20, 30, 40, 50])
catch {
echo b[1, 99]
} as e {
echo '${e.type}: ${e.message}'
}
RangeError: slice bounds 1..99 out of range (length 5)
length() itself is always a legal upper bound, because the bound is
exclusive, and an empty slice is legal rather than an error:
var b = bytes([10, 20, 30, 40, 50])
echo b[0, b.length()]
echo b[2, 2]
echo b[2, 2].is_empty()
(0a 14 1e 28 32)
()
true
An empty byte stream prints as ().
A Slice Is a Copy
Slicing allocates a new stream, so writing through one does not disturb the original:
var original = bytes([10, 20, 30])
var part = original[0, 2]
part[0] = 99
echo original
echo part
(0a 14 1e)
(63 14)
That matters when you are parsing a buffer. Pulling a header out with
frame[0, 4] gives you something you can modify freely, and the frame you
are still reading from is untouched. It also means slicing in a loop copies
every time, so a parser that walks a large buffer should carry an offset
and slice once per field rather than re-slicing the remainder each step.
Slice, Then Decode
The common shape when a buffer holds text with a known extent:
var b = bytes([72, 101, 108, 108, 111, 33])
echo b[0, 5].to_string()
echo b.to_string()
Hello
Hello!
to_string() decodes the whole stream it is called on, so the slice is
what limits the extent. Slicing on a byte boundary in the middle of a
multi-byte character produces a stream that is not valid UTF-8; decode
whole units, or keep the tail for the next read.
There Is No Slice Assignment
A slice can be read but not written to:
var b = bytes([10, 20, 30])
b[0, 2] = bytes([1, 1])
SyntaxError: invalid assignment target
Assign to one index at a time, or rebuild the stream with extend(). A
single index does accept assignment, and the value has to be a real byte:
var b = bytes([72, 101, 108, 108, 111, 33])
b[0] = 74
echo b.to_string()
catch {
b[1] = 300
} as e {
echo '${e.type}: ${e.message}'
}
Jello!
NumericError: bytes element must be an integer in 0..=255, got 300
Note the difference from bytes([300]), which wraps rather than raising.
Construction is lenient; assignment is not.
Writing
var b = bytes([1, 2, 3])
b.append(4)
b.extend(bytes([9]))
echo b
echo b.pop()
echo b.reverse()
(01 02 03 04 09)
9
(04 03 02 01)
Unlike a list, bytes.reverse() mutates in place. So do append,
extend, pop and remove.
dispose() releases the buffer’s memory immediately rather than waiting
for the collector, which matters when you have just finished with something
very large.
Splitting
split() takes a bytes delimiter, not a number:
echo bytes([1, 2, 3, 4, 5]).split(bytes([3]))
[(01 02), (04 05)]
Walking
A byte stream is iterable, and every form that works on a list works here. What you get out is always a number between 0 and 255, never a one-character string.
var b = bytes([72, 105])
for value in b {
echo value
}
72
105
Two variables give you the index first and the value second:
var b = bytes([72, 105])
for index, value in b {
echo '${index}: ${value}'
}
0: 72
1: 105
iter is the form to use when you are decoding a structure and the
position drives the walk — reading a two-byte length, then skipping that
many bytes, then reading the next field:
var b = bytes([72, 105, 33])
iter var i = 0; i < b.length(); i += 2 {
echo '${i}: ${b[i]}'
}
0: 72
2: 33
And each() takes a function, with the value first and the index
second as everywhere else:
bytes([72, 105]).each(@(value, index) {
echo '${index}=${value}'
})
0=72
1=105
When you want the characters rather than the numbers, convert first:
b.to_string() decodes the whole stream as UTF-8, and b.to_list() gives
you the numbers as an ordinary list.
Binary Files
Append b to the file mode and reads give you bytes instead of a string,
and writes accept bytes:
var out = file('data.bin', 'wb')
out.write(bytes([1, 2, 3]))
out.close()
echo file('data.bin', 'rb').read()
(01 02 03)
Reading a non-UTF-8 file without the b fails rather than silently
producing replacement characters, which is the behaviour you want: a
decoding failure is information.
struct: Packing and Unpacking
struct converts between Zuri values and a fixed binary layout. A format
string is a sequence of /-separated fields, each CODE[COUNT][:NAME].
import struct
var header = struct.pack('N:magic/n:version/Z8:name/C:flags',
0x5A555249, 1, 'task', 3)
echo header
echo header.length()
echo struct.unpack('N:magic/n:version/Z8:name/C:flags', header)
(5a 55 52 49 00 01 74 61 73 6b 00 00 00 00 03)
15
{magic: 1515541065, version: 1, name: task, flags: 3}
unpack() hands back a dictionary keyed by the :NAME you gave each
field. Fields with no name get numbered:
echo struct.unpack('N/N', struct.pack('N2', 7, 8))
{1: 7, 2: 8}
A count greater than one under a single name numbers the keys:
echo struct.unpack('n3:vals', struct.pack('n3', 1, 2, 3))
{vals1: 1, vals2: 2, vals3: 3}
The Format Codes
Strings:
| Code | Size | Meaning |
|---|---|---|
a | count | NUL-padded string |
A | count | space-padded string |
Z | count | NUL-padded and NUL-terminated, like C |
h / H | ceil(count/2) | hex string, low or high nibble first |
Integers. The letter tells you the width and the byte order:
| Code | Size | Meaning |
|---|---|---|
c / C | 1 | signed / unsigned 8-bit |
? | 1 | boolean |
s / S | 2 | signed / unsigned 16-bit, native order |
n / v | 2 | unsigned 16-bit, big- / little-endian |
i l / I L | 4 | signed / unsigned 32-bit, native order |
N / V | 4 | unsigned 32-bit, big- / little-endian |
q / Q | 8 | signed / unsigned 64-bit, native order |
J / P | 8 | unsigned 64-bit, big- / little-endian |
u / U | 16 | signed / unsigned 128-bit, little-endian |
Floats:
| Code | Size | Meaning |
|---|---|---|
f / g / G | 4 | float: native / little / big endian |
d / e / E | 8 | double: native / little / big endian |
w / W | 2 | half-precision: little / big endian |
Padding:
| Code | Meaning |
|---|---|
x | write count NUL bytes, consumes no argument |
X | back up count bytes |
@ | seek to absolute position count |
n, N and J are the network-order codes, which is what you want for
almost every wire protocol.
A count of * means “the rest”.
Precision
Zuri numbers are doubles, so they hold integers exactly only up to 2^53.
Every 64-bit and 128-bit code (q, Q, J, P, u, U) automatically
produces a bigint when the unpacked value falls outside that range,
rather than quietly losing digits. Packing accepts either.
The Rest of the Module
calcsize(format) gives the fixed byte size of a format with no * in it.
pack_into(buffer, offset, format, ...) and unpack_from(format, buffer, offset) work on an existing buffer at a position, which is how you build
one large frame without allocating and concatenating per field.
io.BytesIO
BytesIO is a file-shaped object backed by memory. It implements the same
interface file() handles do, which means anything that takes a file takes
a BytesIO too:
import io
var buffer = io.BytesIO(bytes(0), 'w')
buffer.write('hello ')
buffer.write('world')
echo buffer.source.to_string()
hello world
The constructor takes a bytes source and a mode. It has read, gets,
write, puts, seek, tell, flush, close and stats, so a
function written against files needs no changes to work in memory.
That is the useful part: test a function that writes a file without touching the disk, and parse an in-memory buffer with code written for a stream.
Compression
The compress module has a submodule per format, each with compress()
and decompress():
import compress
var raw = ('the quick brown fox ' * 20).to_bytes()
echo raw.length()
echo compress.gzip.compress(raw).length()
echo compress.deflate.compress(raw).length()
echo compress.zstd.compress(raw).length()
echo compress.brotli.compress(raw).length()
echo compress.bzip2.compress(raw).length()
echo compress.lz4.compress(raw).length()
400
47
29
41
31
75
37
The formats are deflate, zlib, gzip, zstd, lz4, bzip2 and
brotli. zlib is re-exported at the top level, so compress.compress()
and compress.decompress() are the zlib pair:
echo compress.decompress(compress.compress(raw)).to_string() == raw.to_string()
true
Which to reach for: gzip when something else has to read it, zstd when
you want the best ratio-to-speed trade, lz4 when speed is the only thing
that matters, brotli for text you will serve over HTTP.
Archives
compress.tar and compress.zip read and write archives, and
compress.checksum has crc32() and adler32():
import compress
echo compress.checksum.crc32('hello'.to_bytes())
907060870
Hashing and Encoding
base64 moves binary through text channels:
import base64
var encoded = base64.encode('hello'.to_bytes())
echo encoded
echo base64.decode(encoded).to_string()
aGVsbG8=
hello
hash covers the digest algorithms:
import hash
echo hash.sha256('hello')
echo hash.hmac_sha256('key', 'hello')
2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
9307b3b915efb5171ff14d8cb55fbcc798c6c0ef1456d66ded1a6aa723a58b7b
Every function takes an optional final as_bytes argument; pass true to
get raw bytes instead of a hex string.
The algorithms are md2, md4, md5, sha1, sha224, sha256, sha384, sha512, the
sha3 family, shake128/256, blake2b512, blake2s256, ripemd160, whirlpool and
gost, each with an hmac_ counterpart, plus pbkdf2() for key derivation.
hash.hash(algorithm, data) selects by name at runtime.
For password storage, use bcrypt rather than any of these. A fast hash is
the wrong tool for a password, and Chapter 13
covers the right one.
A Worked Example: A Length-Prefixed Frame
Most binary protocols are a header followed by a payload. Here is the whole round trip:
import struct
import compress
var HEADER = 'N:length/C:compressed'
def encode_frame(payload) {
var body = payload.to_bytes()
var compressed = 0
if body.length() > 128 {
body = compress.gzip.compress(body)
compressed = 1
}
var frame = struct.pack(HEADER, body.length(), compressed)
frame.extend(body)
return frame
}
def decode_frame(frame) {
var size = struct.calcsize('N/C')
var header = struct.unpack(HEADER, frame[0, size])
var body = frame[size, size + header.length]
if header.compressed == 1 {
body = compress.gzip.decompress(body)
}
return body.to_string()
}
var frame = encode_frame('ping')
echo frame
echo decode_frame(frame)
echo decode_frame(encode_frame('the quick brown fox ' * 20)).length()
(00 00 00 04 00 70 69 6e 67)
ping
400
calcsize() tells you where the header ends, slicing gives you the two
halves, and the compressed flag is one byte because a protocol that has
to guess is a protocol that breaks.
Concurrency with Isolates
Zuri’s concurrency model has one rule: nothing is shared.
An isolate is a real operating-system thread running a complete, private VM. It has its own heap, its own garbage collector, its own register stack and its own module namespace. Two isolates never hold a pointer to the same object, which means there are no data races to reason about, no locks to take, and no stop-the-world pause in one thread caused by allocation in another.
Values cross between isolates by being copied. The copy is a faithful one: cycles are preserved, shared identity inside one payload is preserved, and class instances arrive as instances of the same class.
Spawning
import isolate
def square(n) {
return n * n
}
var task = isolate.spawn(square, 12)
echo task.join()
144
spawn(fn, ...args) queues the call and returns an Isolate handle
immediately. join() waits for it and gives you the result. Calling
join() again returns the same value; it is not a one-shot.
What Can Be Spawned
Any function: a top-level def, a helper declared inside another function,
a lambda, a bound method. What it captures travels with it, including the
modules it uses:
import isolate
import json
def encode_all(items) {
return items.map(@(item) => json.encode(item)).length()
}
echo isolate.spawn(encode_all, [{ a: 1 }, { b: 2 }]).join()
2
A module is not copied the way a list is. The isolate loads the same module for itself and uses its own copy, which is why this works at all: module top levels are declarations, they run once, and the result is cached. A module built out of live per-process state is the case to think twice about, since the isolate gets a fresh one rather than yours.
Two things still cannot be spawned. A native function or a class as the
spawn target itself; wrap it in a def. And a resource handle such as
an open socket or file is moved rather than copied, so the side that
handed it over no longer has it.
Putting the worker in its own module stays the better structure once it grows past a few lines:
Filename: work.zu
import net
def serve(channel) {
var listener = net.TcpStream()
# ...
}
Filename: main.zu
import isolate
import .work
isolate.spawn(work.serve, channel)
That is the shape the rest of this book uses. A worker reached by name is resolved in the isolate’s own namespace, so nothing about it has to be reconstructed.
An import written inside the function body works too, and runs in the
isolate:
import isolate
def make_id() {
import uuid
return uuid.v4().to_string().length()
}
echo isolate.spawn(make_id).join()
36
The arguments are copied, not shared. An isolate that mutates a list it was handed is mutating its own copy:
import isolate
def append_to(items) {
items.append('from the isolate')
return items.length()
}
var original = ['a']
echo isolate.spawn(append_to, original).join()
echo original
2
[a]
The isolate saw two elements. The caller still has one. That is the whole concurrency model in one example: there is no way to write a data race, because there is nothing shared to race over.
Running Many at Once
Spawn a list, join a list:
import isolate
def square(n) {
return n * n
}
var tasks = [1, 2, 3, 4].map(@(n) => isolate.spawn(square, n))
echo tasks.map(@(task) => task.join())
[1, 4, 9, 16]
isolate.map() does exactly that in one call:
import isolate
def square(n) {
return n * n
}
echo isolate.map(square, [1, 2, 3, 4, 5])
[1, 4, 9, 16, 25]
wait_all(tasks, timeout) waits for every task; wait_any(tasks, timeout)
returns as soon as one finishes.
The Lifecycle of a Task
A spawned task moves through exactly two states, and four methods let you ask about it without blocking.
import isolate
import .work
var task = isolate.spawn(work.slow, 3)
echo task.try_join()
echo task.is_done()
echo task.status()
echo task.join()
echo task.status()
nil
false
pending
finished
done
| Method | Answers | Blocks? |
|---|---|---|
join(timeout) | the result | yes |
try_join() | the result, or nil if not ready | no |
is_done() | whether it has finished | no |
status() | 'pending' or 'done' | no |
name() | the name it was given, or nil | no |
try_join() returning nil is ambiguous when the task’s own result could
be nil; pair it with is_done() when that matters.
join() may be called more than once. It is not a one-shot: the second
call returns the same value immediately.
Naming a Task
spawn() leaves a task anonymous. spawn_named() gives it a name that
shows up in diagnostics:
var named = isolate.spawn_named('importer', work.quick, 5)
var plain = isolate.spawn(work.quick, 5)
echo named.name()
echo plain.name()
importer
nil
Name anything long-running. A stuck program is far easier to diagnose when the task can say what it is.
Timeouts Are in Seconds
Every timeout in the isolate module is measured in seconds, and it
takes a fraction:
import isolate
var empty = isolate.channel(1)
var start = time()
catch {
empty.recv(0.3)
} as e {
echo '${e.type} after about ${((time() - start) * 10).round() * 100}ms'
}
IsolateTimeoutError after about 300ms
This applies to join(), send(), recv(), select(), wait_any(),
wait_all() and shutdown() alike.
It is worth stating loudly because the net module uses milliseconds
for its own timeouts. socket.set_read_timeout(5000) is five seconds;
channel.recv(5000) is an hour and twenty minutes. The two modules are
easy to use in one program, and the mistake is silent — a timeout that
never fires simply looks like a hang.
Omitting the timeout means “wait forever”, which is the right default when the other side is code you control and the wrong one when it is not.
Cancellation Is Cooperative
cancel() requests that a task stop. It does not kill anything.
What happens next depends on what the task is doing:
Blocked in a channel operation, a join(), a wait_* or a select() —
the call is interrupted within roughly 50ms and raises
IsolateCancelledError.
Running ordinary code — nothing is interrupted. The task’s own function must notice and return:
import isolate
def slow(seconds) {
var start = time()
while time() - start < seconds {
if isolate.is_cancelled() {
return 'stopped early'
}
}
return 'ran to completion'
}
isolate.is_cancelled() is the module-level function a worker calls about
itself. task.is_cancelled() is the method the spawner calls to ask
whether it requested cancellation — and it answers true from the moment
cancel() was called, whether or not the task noticed:
var task = isolate.spawn(work.slow, 3)
task.cancel()
echo task.join()
echo task.is_cancelled()
ran to completion
true
That output is not a contradiction. The worker in this example does not
poll, so it finished normally; is_cancelled() reports the request, not
the outcome. A loop with no cancellation check is a loop that cannot be
stopped.
Put the check where the loop turns over, and make it cheap. Checking once per iteration of an outer loop is usually enough; checking inside the innermost arithmetic is not worth it.
What Can Cross, and What It Costs
Arguments in, results out, and channel traffic all cross the same way: everything is copied. There is no sharing and no reference that survives the boundary.
The Copy Is Faithful
A copy is not a shallow snapshot. Cycles survive, shared identity inside one payload survives, and a class instance arrives as an instance of the same class with its methods intact:
import isolate
class Point {
@new(x) {
self.x = x
}
doubled() {
return self.x * 2
}
}
def identity(value) {
return value
}
# A list that contains itself.
var cyclic = [1]
cyclic.append(cyclic)
# Two slots holding one list.
var shared = [1]
var pair = [shared, shared]
echo isolate.spawn(identity, Point(3)).join().doubled()
echo typeof(isolate.spawn(identity, cyclic).join())
var back = isolate.spawn(identity, pair).join()
back[0].append(2)
echo back[1].length()
6
list
2
The last line is the one to notice. pair held the same list twice, and on
the other side it still does: appending through back[0] is visible
through back[1]. The copy preserved the shape of the sharing, not just
the values.
What Cannot Cross
A handful of things are tied to the isolate that made them, and sending one raises rather than silently producing something broken:
import isolate
def identity(value) {
return value
}
catch {
isolate.spawn(identity, file('notes.txt'))
} as e {
echo e.message.lines()[0]
}
cannot send a file across isolates; only nil, bool, number, string, bytes, bigint, range, list, dict, instance, class, bound method, function, module, and native-pointer values can cross
The message lists what can cross, which is the more useful half. An open file is the one you will meet: it is a handle onto a position in a descriptor this process owns, and there is no copy of that to hand over. Send the path instead and let the worker open it.
The Cost Model
Spawning is cheap. A task is a small descriptor holding the callee and a snapshot of its arguments, pushed onto a queue the pool drains, so queueing a hundred thousand of them is reasonable.
What is not free is the copy. Handing an isolate a large list copies the whole list, once per spawn. When a worker needs a lot of data, give it a description of the work — a path, a range of indices, a query — and let it do its own reading:
# Copies the file's contents into every worker.
isolate.map(work.process, files.map(@(p) => file(p).read()))
# Copies a short path into every worker instead.
isolate.map(work.read_and_process, files)
Errors Cross the Boundary
An uncaught error inside an isolate surfaces at the join() as an
IsolateError, carrying the original error’s message and the stack trace
from inside the worker:
import isolate
def boom() {
raise ValueError('worker failed')
}
catch {
isolate.spawn(boom).join()
} as e {
echo e.type
echo e.message.lines()[0]
}
IsolateError
ValueError: worker failed
Nothing is lost. You get the failure where you can do something about it,
with enough information to find it. Note the shape of the message: the
IsolateError wraps the worker’s own error rather than replacing it, so
e.message still names ValueError and the line it came from.
An error in a task nobody joins has nowhere to surface. That is what
scope(), further down, exists to prevent.
Channels
A channel is a bounded queue that crosses isolate boundaries. One side sends, the other receives, and the values are copied on the way across like everything else:
import isolate
def producer(channel, count) {
iter var i = 0; i < count; i++ {
channel.send(i)
}
channel.close()
}
var channel = isolate.channel(4)
var task = isolate.spawn(producer, channel, 3)
while true {
var value = channel.recv()
if value == nil and channel.is_closed() {
break
}
echo value
}
task.join()
0
1
2
Read the loop condition carefully, because it is the part that is easy to
get wrong. recv() returns nil both for “a nil was sent” and for “the
channel is closed and drained”, so the test for the end of the stream is
nil and is_closed(). Checking only for nil would stop early on a
legitimate nil; checking only is_closed() would stop before the queue
had drained.
channel(capacity) bounds the queue, which gives you backpressure for
free: a producer that outruns its consumer blocks on send() rather than
growing the queue without limit.
| Method | Behaviour |
|---|---|
send(value, timeout) | blocks while the channel is full |
recv(timeout) | blocks while the channel is empty |
try_recv() | returns immediately, nil when empty |
close() | no more sends; pending receives drain, then return nil |
is_closed() | whether it has been closed |
length() | how many values are waiting |
Every one of those behaviours is observable from a single isolate, which makes a channel easy to reason about before you introduce a second one:
import isolate
var c = isolate.channel(2)
c.send('a')
c.send('b')
echo c.length()
echo c.try_recv()
c.close()
echo c.is_closed()
echo c.recv()
echo c.recv()
catch {
c.send('z')
} as e {
echo '${e.type}: ${e.message}'
}
2
a
true
b
nil
IsolateError: cannot send on a closed channel
Read the last four lines together. After close(), the queue still
drains: recv() returned the buffered b before it started returning
nil. A closed, drained channel returns nil forever rather than raising,
and only send() raises.
Closing twice is harmless. try_recv() on an empty channel returns nil
immediately and never blocks.
Selecting Across Channels
select() waits on several channels at once and tells you which one
produced a value:
import isolate
var c1 = isolate.channel(1)
var c2 = isolate.channel(1)
c2.send('from c2')
var picked = isolate.select([c1, c2], 1000)
echo picked[1]
from c2
It returns [channel, value], so you can tell where the value came from by
comparing the first element against the channels you passed in. The second
argument is a timeout in seconds, like every other timeout in this
module — see Timeouts Are in Seconds, because
it is the opposite of what the net module does.
This is the shape for a consumer fed by more than one producer — work on one channel, shutdown signals on another — without polling either.
Broadcast
A channel delivers each value to exactly one receiver. A broadcast delivers each value to every subscriber:
import isolate
var bus = isolate.broadcast(4)
var a = bus.subscribe()
var b = bus.subscribe()
bus.send('tick')
echo a.recv()
echo b.recv()
echo bus.subscriber_count()
bus.close()
tick
tick
2
subscribe() hands back an ordinary Channel, so everything in the
channel table applies to it: recv(), try_recv(), length(), and the
same backpressure.
A Subscriber Only Sees What Comes After It
There is no replay. A subscriber that arrives late has missed everything sent before it subscribed:
import isolate
var bus = isolate.broadcast(4)
var early = bus.subscribe()
bus.send('first')
var late = bus.subscribe()
bus.send('second')
echo 'early: ${early.recv()}, ${early.recv()}'
echo 'late: ${late.try_recv()}'
early: first, second
late: second
This matters when subscribers are set up concurrently with the producer. Subscribe everything before anything is sent, or accept that a consumer starting later begins mid-stream.
Unsubscribing, and Sending Into the Void
unsubscribe(channel) drops one subscriber. Sending with none at all is
not an error — the value is simply discarded:
import isolate
var bus = isolate.broadcast(4)
var sub = bus.subscribe()
bus.unsubscribe(sub)
echo bus.subscriber_count()
echo bus.send('nobody is listening')
echo sub.try_recv()
0
nil
nil
A dropped subscriber stops receiving new values immediately, and its channel keeps whatever was already queued in it:
import isolate
var bus = isolate.broadcast(4)
var sub = bus.subscribe()
bus.send('queued before unsubscribe')
bus.unsubscribe(sub)
echo sub.try_recv()
queued before unsubscribe
So unsubscribing is not a way to discard what a consumer has not read yet; it only stops the flow.
Closing
close() closes the bus and every channel it handed out. Subscribers drain
what they have and then receive nil; send() raises:
import isolate
var bus = isolate.broadcast(4)
var sub = bus.subscribe()
bus.close()
echo bus.is_closed()
echo sub.try_recv()
catch {
bus.send('too late')
} as e {
echo e.type
}
true
nil
IsolateError
Backpressure Applies Per Subscriber
The capacity you give broadcast() is the capacity of each subscriber’s
channel. A subscriber that stops reading fills its own buffer, and once it
is full the bus blocks on send() — one slow consumer holds up the
producer and therefore everyone.
When a slow consumer must not be allowed to do that, give the bus a larger capacity, or have that consumer read into its own queue and fall behind on its own time.
Channel or Broadcast?
Use a channel when the work should be done once, by whichever worker is free. Use a broadcast when every consumer needs to see every event. Shutdown signals, configuration changes and progress events are broadcasts; jobs are channels.
Scopes
A scope owns its children and does not return until all of them have finished:
import isolate
def square(n) {
return n * n
}
var result = isolate.scope(@(s) {
s.spawn(square, 3)
s.spawn(square, 4)
return s.children().map(@(child) => child.join())
})
echo result
[9, 16]
scope(body) calls body with a scope object, waits for everything
spawned through it, and returns whatever the body returned.
It Waits Whether or Not You Join
Joining inside the body is how you collect results. It is not how the waiting happens — that is the scope’s job either way:
import isolate
def square(n) {
return n * n
}
echo isolate.scope(@(s) {
s.spawn(square, 3)
s.spawn_named('four', square, 4)
return s.children().length()
})
2
Neither child was joined, and the scope still did not return until both had
finished. children() gives you the Isolate objects in spawn order, and
s.spawn_named() works exactly like the module-level one.
A Failure Stops the Group
This is the real reason to use a scope. A bare spawn() whose result is
never joined swallows its own error in silence; a scope raises:
import isolate
def ok(n) {
return n
}
def boom() {
raise ValueError('worker failed')
}
catch {
isolate.scope(@(s) {
s.spawn(ok, 1)
s.spawn(boom)
return 'never returned'
})
} as e {
echo '${e.type}: ${e.message.lines()[0]}'
}
IsolateError: ValueError: worker failed
The body’s return value is discarded when a child failed, and the failure
comes out of scope() itself. Reach for a scope whenever a piece of work
fans out and must be complete — and correct — before the next step begins.
Waiting on Several Tasks
Three helpers cover the shapes that come up, and they return different things:
import isolate
def square(n) {
return n * n
}
var tasks = [1, 2, 3].map(@(n) => isolate.spawn(square, n))
echo isolate.wait_all(tasks)
echo typeof(isolate.wait_any([isolate.spawn(square, 9)]))
echo isolate.map(square, [4, 5])
[1, 4, 9]
Isolate
[16, 25]
wait_all(tasks, timeout) returns the results, in the order the
tasks were given, not the order they finished.
wait_any(tasks, timeout) returns the Isolate that finished
first, not its result — you still call join() on it. That is what lets
you tell which one won.
map(fn, items, timeout) spawns one task per item and collects the
results, which is wait_all with the spawning done for you.
All three raise IsolateTimeoutError if the timeout passes, and all three
propagate a worker’s failure:
import isolate
def halve(n) {
if n == 0 {
raise ValueError('cannot halve zero')
}
return n / 2
}
echo isolate.map(halve, [2, 4, 6])
catch {
isolate.map(halve, [2, 0, 6])
} as e {
echo '${e.type}: ${e.message.lines()[0]}'
}
[1, 2, 3]
IsolateError: ValueError: cannot halve zero
The whole call fails on the first failing element. When individual failures
are acceptable, spawn and join yourself with a catch around each join(),
as the worked example below does.
Isolates Can Spawn Isolates
There is no restriction on nesting. A worker may spawn its own tasks and join them, and they run on the same pool:
Filename: work.zu
import isolate
def double(n) {
return n * 2
}
def nested(n) {
return isolate.spawn(double, n).join()
}
$ zuri run main.zu
6
This is worth knowing mostly as a warning, which the next section covers: nested tasks that block on each other are the fastest way to exhaust the pool.
The Pool
Isolates do not each get a thread of their own. They run on a fixed pool of OS threads, sized to the machine’s core count by default:
import isolate
echo isolate.cpu_count() > 0
echo isolate.pool_size() == isolate.cpu_count()
true
true
Sizing It
configure(threads) changes the size, and it only works before the pool
exists — which is to say, before the first spawn. It reports whether it
did anything:
import isolate
echo isolate.configure(7)
echo isolate.pool_size()
true
7
Call it after a spawn and it returns false and changes nothing. It does
not raise, so a configure() buried below some initialisation that already
spawned will silently do nothing — put it at the very top of the entry
file.
The size is a ceiling, not a head count. A thread is started only when an
isolate is waiting and every thread already started is busy, so sizing the
pool generously for a burst costs nothing while the burst is not happening.
started_count() says how many threads the pool has started so far:
import isolate
isolate.configure(8)
echo isolate.started_count()
echo isolate.spawn(@() => 21 * 2).join()
echo isolate.started_count()
0
42
1
Why the Size Matters More Than It Looks
A pool sized to your core count can deadlock, and the failure looks like a hang rather than an error.
The mechanism: a task that blocks — on join(), on recv(), on a socket
read — occupies its thread while it waits. If every thread is occupied by a
task waiting for work that has no free thread to run on, nothing can ever
progress.
That happens most easily with three patterns:
- a server and a client that talk to each other in one process;
- nested spawns where the parent joins the child;
- a pipeline with more stages than threads.
Size the pool for the number of tasks that will be blocked at once, not for the number of cores. Threads that are blocked are not competing for CPU, so over-provisioning costs little; under-provisioning costs everything.
Watching It
import isolate
echo isolate.active_count() >= 0
echo isolate.queued_count() >= 0
echo isolate.is_shutdown()
true
true
false
active_count() is how many tasks are running; queued_count() is how
many are waiting for a thread. A queued count that only grows is the
signature of a starved pool.
shutdown(timeout) drains the pool and stops it. Most programs never call
it; it is there for a long-lived process that wants to release its threads
without exiting.
Memory While Waiting
A pool thread keeps its isolate’s heap from one task to the next, and the
collector runs as a program allocates. A thread that waits allocates
nothing, so a waiting thread collects on its own: once it has waited a
second, for its next task or inside recv(), select(), join(),
wait_any() or wait_all(), it collects its garbage and the memory that
garbage held goes back to the system. The main thread does the same in
those calls. A task that built and dropped a large structure leaves the
process at its working size rather than its peak.
A wait of a second or less returns before that point, and a heap that has grown by less than 4 MB since its last such collection is left as it is, so a loop of short waits pays nothing for it.
The Errors
Three error types come out of this module, and they mean different things:
| Error | Raised when |
|---|---|
IsolateError | a worker raised, or an operation is invalid — sending on a closed channel |
IsolateTimeoutError | a timeout passed before the operation completed |
IsolateCancelledError | a blocking call was interrupted by cancel() |
All three inherit from Error, so catch on its own catches every one and
instance_of() sorts them.
The distinction matters because they call for different responses. A
timeout usually means retry or give up. A cancellation means shut down
quietly. An IsolateError carrying a worker’s failure means something in
your own code went wrong, and the message names it.
A Worked Example
Fan out a computation, collect the results, and report a failure without losing the rest:
import isolate
def classify(n) {
if n < 0 {
raise ValueError('negative input: ${n}')
}
if n < 2 {
return 'small'
}
var divisor = 2
while divisor * divisor <= n {
if n % divisor == 0 {
return 'composite'
}
divisor++
}
return 'prime'
}
def classify_all(numbers) {
var tasks = numbers.map(@(n) => [n, isolate.spawn(classify, n)])
var results = {}
for pair in tasks {
catch {
results[pair[0]] = pair[1].join()
} as e {
results[pair[0]] = 'failed'
}
}
return results
}
echo classify_all([1, 7, 9, -3, 97])
{1: small, 7: prime, 9: composite, -3: failed, 97: prime}
Three things are doing the work there.
Every task is spawned before any is joined. Spawning in one pass and joining in another is what makes the work overlap; joining inside the first loop would run them one after another and buy nothing.
The number is carried alongside its task, because the results come back in whatever order the pool produces them and a bare list of results would lose which input each belonged to.
And the catch is around one join(), so the one bad input becomes one
failed entry rather than ending the batch. Move it outside the loop and
the first failure takes the other four with it.
Choosing a Shape
Fan out, collect results. isolate.map(), or scope() when a failure
must stop the group.
A pipeline. Channels between stages, each stage an isolate, each channel bounded so backpressure propagates all the way back to the source.
An event fan-out. A broadcast, one subscriber per consumer.
A server. http.serve() builds the whole pattern for you: one isolate
per worker, a bounded backlog, connections handed out as they arrive.
Chapter 15 covers it.
Nothing at all. Concurrency costs a copy at every boundary and a great deal of care at every join. A single-threaded loop that finishes in time is the better program.
Networking
The net module is the socket layer: TCP, UDP, unix domain sockets, TLS,
DTLS, address parsing and polling. The http module sits on top of it and gives you a client and
a server that speak HTTP/1.1 and HTTP/2.
Every example in this chapter runs a server in an isolate and a client in the main program, which is how you test network code without two terminals. Chapter 11 is the background.
TCP
A TcpStream is both ends. bind() and accept() make it a listener;
connect() makes it a client. accept() blocks until someone connects,
unless the listener was put in non-blocking mode, in which case it answers
nil when nobody is waiting.
Filename: server.zu
import net
def echo_server(port_channel) {
var listener = net.TcpStream()
listener.bind('127.0.0.1:0')
port_channel.send(listener.local_address())
var client = listener.accept()
var received = client.read_as_string()
client.write_all('echo: ' + received)
client.close()
listener.close()
}
Filename: main.zu
import net
import isolate
import .server
var channel = isolate.channel(1)
var task = isolate.spawn(server.echo_server, channel)
var address = channel.recv()
var client = net.TcpStream()
client.connect(address)
client.write_all('hello')
client.shutdown(net.Shutdown.WRITE)
echo client.read_as_string()
client.close()
task.join()
echo: hello
Three things in that pair of files are doing real work.
The server lives in its own file. It has to: echo_server refers to
net, and a function that references an imported module cannot be spawned
from the file that did the importing. In server.zu, net is resolved
inside the isolate instead of being captured from the caller.
Chapter 11 covers the rule and the alternative.
Binding to port 0 asks the operating system for a free port, and
local_address() reports which one it gave. That is the right way to write
a test, and the right way to run several servers in one process. The port
travels back to the main program through the channel, because the main
program cannot know it in advance.
shutdown(net.Shutdown.WRITE) is not optional here.
read_as_string() reads until the peer stops writing, so without a
half-close from the client the server would still be waiting for more
request while the client waits for a response — a deadlock that looks
exactly like a hang. Shutdown has READ, WRITE and BOTH.
Reading
| Method | Behaviour |
|---|---|
read(length) | up to length bytes, whatever has arrived |
read_exact(length) | exactly length bytes, waiting for them |
read_all() | everything until the peer closes, as bytes |
read_as_string() | the same, decoded as UTF-8 |
peek(length) | look without consuming |
read() returning fewer bytes than you asked for is normal, not an error.
That is what read_exact() is for.
Writing
write(data) writes what it can and tells you how much. write_all(data)
loops until everything is out, and is what you want almost always.
Options
socket.set_read_timeout(5000)
socket.set_write_timeout(5000)
socket.set_nodelay(true)
socket.set_non_blocking(true)
socket.set_ttl(64)
Timeouts here are milliseconds, which is the opposite of the isolate
module’s seconds — easy to mix up in a program that uses both.
Set a read timeout on anything that talks to the network. A socket with no timeout and a peer that never answers is a thread that never comes back.
set_nodelay(true) disables Nagle’s algorithm, which is what you want for
a request/response protocol where latency matters more than packet count.
UDP
import net
var a = net.UdpSocket()
a.bind('127.0.0.1:0')
var b = net.UdpSocket()
b.bind('127.0.0.1:0')
b.set_read_timeout(2000)
a.send_to('ping', b.local_address())
var pending = b.peek_from(1024)
echo pending.data.to_string()
echo pending.address.to_string()
echo b.receive_from(1024).to_string()
b.send_to('pong', pending.address)
a.set_read_timeout(2000)
echo a.receive_from(1024).to_string()
ping
127.0.0.1:<port>
ping
pong
The second line is written with a placeholder because the real one is not
predictable: bind('127.0.0.1:0') asks the operating system for any free
port, and it picks a different one every run.
receive_from(length) gives you the datagram’s bytes. To learn who sent
it, peek_from(length) returns { data, address } without consuming the
datagram, so the pattern for a responder is peek, read, reply.
A datagram larger than the buffer you give is truncated, and the rest is discarded. Size the buffer for your protocol’s largest message.
There is also connect() to fix a peer, set_broadcast(), and
join_multicast_v4() / join_multicast_v6().
Unix Sockets
A unix domain socket is a path on the filesystem rather than an address on the network. It is how two programs on one machine usually talk when one of them is a server: databases, message brokers and system daemons nearly all listen on one.
There is no port to collide with and nothing to reach it from another host. And because the socket is a file, the filesystem decides who may connect — a socket in a directory only one user can enter is reachable by only that user, which is a stronger answer than binding to loopback and trusting everyone on the machine.
pair() gives two sockets already connected to each other, with no
path and nothing left on disk:
import net
var ends = net.pair()
ends[0].write_all('ping')
echo ends[1].read_exact(4)
ends[1].write_all('pong')
echo ends[0].read_exact(4).to_string()
ends[0].close()
ends[1].close()
(70 69 6e 67)
pong
A server binds a path and accepts on it, exactly as a TcpStream
binds an address:
import net
var server = net.UnixStream()
server.bind('/run/app.sock')
while true {
var client = server.accept()
client.write_all('hello\n')
client.close()
}
Two things differ from TCP and both come from the socket being a file.
bind() refuses a path that already exists, including one a crashed
process left behind, so a server that expects to be restarted deletes
a stale path first. And closing frees the descriptor without removing
the file, because by then another process may have bound the same path
and removing it would break them.
net.is_supported() is false on Windows, where these are not
available. Every call raises there rather than the module being
missing, so a program that can fall back to TCP tests it rather than
catching an error.
Addresses
import net
var address = net.SocketAddr.parse('127.0.0.1:8080')
echo address.to_string()
127.0.0.1:8080
resolve() turns a name into addresses:
echo net.resolve('localhost:80').map(@(a) => a.to_string())
[[::1]:80, 127.0.0.1:80]
It returns a list because a name can have several addresses, and what comes
back depends on the machine: a host with IPv6 configured answers for both
families, one without gives only [127.0.0.1:80]. Never assume a position
in that list, and never assume a family.
net.ip has IpAddress, Ipv4Address and Ipv6Address for parsing,
comparing and classifying addresses, which is what you want before you
trust an X-Forwarded-For header.
TLS
net.tls wraps an established TCP stream:
import net
import net.tls
var config = tls.TlsConfig()
config.set_cert_chain(certificate_pem, private_key_pem)
var listener = net.TcpStream()
listener.bind('127.0.0.1:0')
var client = listener.accept()
var secure = tls.TlsStream.accept(client, config)
echo secure.read_as_string()
A TlsStream has the same read and write interface a TcpStream does, so
code written against one works against the other.
On the client side, tls.TlsStream.connect(socket, config, hostname)
performs the handshake and verifies the certificate. config.add_ca_pem()
adds a trust anchor, and config.require_client_cert(true) turns on mutual
TLS. peer_certificate() gives you the other end’s certificate once the
handshake is done.
net.dtls is the same thing over UDP.
Polling
net.poll waits on many sockets at once without a thread each. It is the
right tool when you have thousands of mostly-idle connections and the wrong
tool when you have a handful of busy ones, where an isolate per connection
is simpler and faster.
Accepting without parking the thread
accept() on a blocking listener waits inside the runtime. Nothing else on
that thread gets a turn while it waits: a signal trapped with
os.on_signal() is not delivered, and a flag telling the server to stop is
not read, until a connection happens to arrive. On an idle server that is
never, which is why Ctrl+C on one can appear to do nothing at all.
net.Acceptor is the accept loop without that problem. It puts the listener
into non-blocking mode and waits on a poller instead. next() hands over a
connection when there is one and returns nil when its interval passes with
nothing arriving, which is the loop’s chance to look at whatever else it has
to look at.
import net
var listener = net.TcpStream()
listener.bind('127.0.0.1:0')
var acceptor = net.Acceptor(listener, 50)
# Nobody has connected yet, so the round comes back empty.
echo acceptor.next()
var client = net.TcpStream()
client.connect(listener.local_address())
var served = acceptor.wait_for_one()
echo served.peer_address().ip().to_string()
served.close()
client.close()
listener.close()
nil
127.0.0.1
The interval decides how soon a stopped server notices, not how soon a
connection is served: one that arrives wakes the wait at once. Nothing is
added while connections are arriving either, because next() tries
accept() first and reaches for the poller only when the queue is empty.
That turns a server into something Ctrl+C can stop:
import net
import os
var listener = net.TcpStream()
listener.bind('127.0.0.1:8080')
var running = true
os.on_signal('INT', @() {
running = false
return true
})
var acceptor = net.Acceptor(listener)
while running {
var client = acceptor.next()
if client == nil {
continue
}
serve(client)
}
listener.close()
The handler returns true to say it has taken responsibility for the
signal. A handler that returns anything falsy declines, and the process
then dies of the signal as it would have with nothing registered at all.
http accepts this way in both HttpServer.listen() and http.serve(), and
so do the SMTP and IMAP servers in mail. A server built on any of them is
already stoppable.
Above the Socket Layer
Everything so far has been bytes on a socket. Most programs want a protocol on top of that, and the standard library brings two of them.
http is a complete HTTP/1.1 and HTTP/2 client and server: routing,
middleware, cookies, multipart uploads, static files, server-sent events
and WebSockets. It is large enough to have its own chapter, and
Chapter 15 is it. The one-line version:
import http
var response = http.get('https://example.com')
echo response.status
echo response.as_text()
net.tls wraps a TcpStream in TLS, as shown above, for protocols
that are not HTTP — a mail client, a database driver, a custom binary
protocol.
A rough guide to which layer you want:
| You are writing | Reach for |
|---|---|
| a web API, or a client for one | http |
| a browser-facing server | http |
| a client for an existing non-HTTP protocol | net.tcp, plus net.tls if it is encrypted |
| a protocol of your own design | net.tcp and struct |
| discovery, telemetry, games | net.udp |
| anything waiting on many sockets at once | net.poll |
| a server that has to stop on Ctrl+C | net.Acceptor |
A Worked Example
A length-prefixed request/response protocol of the kind net is for: each
message is a four-byte big-endian length followed by that many bytes of
JSON. This is the pattern behind most binary protocols, and it is worth
writing once by hand.
Filename: framed.zu
import struct
# Reads one frame, or returns nil once the peer has hung up.
#
# read_exact() raises rather than returning short when the stream ends
# mid-read, and a clean disconnect between frames looks exactly like
# that, so the end of the conversation arrives here as an error.
def read_frame(stream) {
var header
catch {
header = stream.read_exact(4)
} as e {
return nil
}
var length = struct.unpack('N:size', header).size
return stream.read_exact(length).to_string()
}
def write_frame(stream, payload) {
stream.write_all(struct.pack('N', payload.length()))
stream.write_all(payload)
}
Filename: server.zu
import net
import json
import .framed
def serve(port_channel) {
var listener = net.TcpStream()
listener.bind('127.0.0.1:0')
port_channel.send(listener.local_address())
var client = listener.accept()
while true {
var request = framed.read_frame(client)
if request == nil {
break
}
var parsed = json.decode(request)
framed.write_frame(client, json.encode({ reply: parsed.message }))
}
client.close()
listener.close()
}
Filename: main.zu
import net
import json
import isolate
import .framed
import .server
var channel = isolate.channel(1)
var task = isolate.spawn(server.serve, channel)
var client = net.TcpStream()
client.connect(channel.recv())
client.set_read_timeout(2000)
framed.write_frame(client, json.encode({ message: 'first' }))
echo framed.read_frame(client)
framed.write_frame(client, json.encode({ message: 'second' }))
echo framed.read_frame(client)
client.close()
task.join()
{"reply":"first"}
{"reply":"second"}
The framing is the whole point. TCP is a stream of bytes with no message
boundaries in it: one write_all() may arrive as three reads, and three
writes may arrive as one. read_exact(4) followed by read_exact(length)
is what puts the boundaries back, and read() alone would not — it returns
whatever has arrived, which is why the table above distinguishes the two.
The end of the conversation is the other thing worth studying.
read_exact() raises when the stream ends before it has the bytes it
was promised, and a peer that hangs up cleanly between frames produces
exactly that. Catching it and returning nil is what turns “the connection
closed” from a crash into the loop’s normal exit.
Note also that both ends share framed.zu. A protocol implemented twice,
once per end, is a protocol that will eventually disagree with itself.
One last detail, easily missed: the reply key is reply, not echo.
echo is a keyword, so it cannot be a bare dictionary key — { echo: x }
is a syntax error. Quote it as { 'echo': x } if you need that exact
name.
The Rest of the Module
net.poll answers “which of these sockets can I read right now?” without a
thread per socket, which is how you serve many connections from one
isolate. It takes a UnixStream alongside a TcpStream or a TlsStream. net.addr and net.ip parse, format and classify addresses —
is_private(), is_loopback(), is_multicast() and the rest — which is
what you want before trusting an address a client sent you. net.dtls is
TLS over UDP.
Appendix F lists every submodule.
A Tour of the Standard Library
The standard library is part of the installation. There is nothing to add
to a manifest, no package manager to run, and no dependency to resolve;
import json works in a file you created ten seconds ago.
This chapter walks through what is there, grouped by the job you would
reach for it to do, with a working example of each. It is a tour rather
than a reference: Appendix F is the index,
every module carries doc blocks in libs/, and the three largest modules
have chapters of their own — Wire,
HTTP and Imagine.
Read it once to learn what exists. The value of a tour like this is not remembering the details; it is recognising, six months from now, that the thing you are about to write by hand is already here.
Data Formats
json
import json
var data = { name: 'Ada', langs: ['zuri', 'rust'], active: true }
echo json.encode(data)
echo json.decode('{"a":1}').a
echo json.encode({ name: 'Ada' }, false)
{"name":"Ada","langs":["zuri","rust"],"active":true}
1
{
"name": "Ada"
}
encode(value, compact, max_depth) defaults to compact. parse(path)
reads and decodes a file; dump(value, file) writes one. A class that
defines @to_json() controls its own encoding.
yaml
import yaml
echo yaml.parse('name: zuri\ntags:\n - fast\n - small')
{name: zuri, tags: [fast, small]}
Anchors, aliases, tags, multi-document streams and block scalars are all supported.
toml
import toml
echo toml.parse('[package]\nname = "zuri"\nversion = "1.0"\n')
{package: {name: zuri, version: 1.0}}
parse() and dump() treat a document as data. edit() treats it as a
file somebody wrote: the Document it returns renders back byte for byte
until it is changed, and a change disturbs only the line it lands on.
import toml
var doc = toml.edit('# what we ship\n[package]\nversion = "0.9.0" # bump me\n')
doc.set('package.version', '1.0.0')
echo doc.to_string()
# what we ship
[package]
version = "1.0.0" # bump me
That is what a program editing somebody else’s configuration file needs: the comment stayed, and so did the spacing on the line that changed.
csv
import csv
echo csv.parse('a,b\n1,2')
[[a, b], [1, 2]]
Reader and Writer stream large files, Dialect configures separators
and quoting, and sniff_dialect() guesses from a sample.
struct
Binary layouts. Covered in Chapter 10.
base64, convert
import convert
echo convert.bytes_to_hex(bytes([255, 0]))
echo convert.to_base(255, 16)
echo convert.from_base('ff', 16)
ff00
ff
255
convert handles every base-to-base conversion you would otherwise write
by hand, plus hex, binary, octal and unicode helpers.
Text and Markup
html
A WHATWG-conformant parser, a real DOM, and CSS selectors:
import html
var doc = html.parse('<ul><li class="a">one</li><li>two</li></ul>')
echo doc.query_selector('li.a').text_content()
echo doc.query_selector_all('li').length()
one
2
The DOM supports traversal, mutation and serialisation, which makes it a scraper, a templating backend and a sanitiser in one module.
wire
Templating, with directives expressed as HTML attributes rather than a second syntax layered over your markup:
import wire
echo wire.render_string('<p x-text="msg"></p>', { msg: 'hi' })
<p>hi</p>
Everything is escaped by default, and escaped correctly for where it sits:
a value in an attribute, in a URL and in a <script> block are three
different escapes, and Wire knows which is which because it parses your
template as structure rather than text.
Chapter 14 is the full treatment.
url
import url
var parsed = url.parse('https://user@example.com:8443/a/b?q=1#top')
echo parsed.host
echo parsed.port
echo parsed.get_param('q')
example.com
8443
1
encode(), decode() and parse_query() handle percent-encoding.
mime
import mime
echo mime.detect_from_name('a.png')
image/png
detect(file) sniffs content rather than trusting the extension, which is
the check you want on an upload.
colors
ANSI colour for terminal output, and conversion between every colour space you are likely to have a value in:
import colors
echo colors.hex_to_rgb('#ff8800')
echo colors.rgb_to_hex(255, 136, 0)
echo colors.rgb_to_ansi256(255, 136, 0)
[255, 136, 0, 1]
ff8800
214
colors.text(value, color, background) wraps a string in the escape codes:
import colors
echo colors.text('warning', colors.text_color.yellow)
The conversions are what make the same code work on a terminal that cannot
do what you asked: a true-colour value becomes the nearest of 256, and 256
becomes the nearest of 16. hex(), rgb(), hsl(), hsv(), hwb(),
cmyk() and xyz() each take a colour in that space and produce the
escape sequence for it.
Time
date
import date
var d = date.date(2026, 9, 11, 8, 30, 0)
echo d.format('Y-m-d H:i:s')
echo d.format('l, jS F Y')
echo date.parse('2026-09-11').format('Y-m-d')
2026-09-11 08:30:00
Friday, 11th September 2026
2026-09-11
The format codes are single letters: Y four-digit year, m zero-padded
month, d zero-padded day, H 24-hour, i minutes, s seconds, l
weekday name, F month name, jS day with an ordinal suffix.
localtime() and gmtime() give the current time, from_time(seconds)
converts a Unix timestamp, and the module carries a real IANA time zone
database, so from_timezone('Europe/London', ...) does the right thing
across a daylight-saving boundary.
Cryptography and Identity
hash
Digests and HMACs. Covered in Chapter 10.
bcrypt
Password hashing, which is a different problem from digesting:
import bcrypt
var stored = bcrypt.hash('secret')
echo bcrypt.compare('secret', stored)
echo bcrypt.get_rounds(stored)
true
10
Use this for passwords and hash for everything else. needs_rehash()
tells you when a stored hash was made with a lower cost than you now
require.
crypto
RSA signing and HKDF key derivation.
uuid
import uuid
echo uuid.v4().length()
echo uuid.is_valid(uuid.v7())
36
true
Versions 1, 3, 4, 5, 6, 7 and 8 are all there. v4 is the random one you
usually want; v7 is time-ordered, which makes it a better database key.
jwt
Signing, verifying and decoding JSON Web Tokens, with a JWKS client for rotating keys.
Validation and Structure
validate
A fluent schema builder:
import validate
var schema = validate.schema({
name: validate.required().string().max_length(10),
age: validate.required().integer().min(0),
})
echo schema.check({ name: 'Ada', age: 36 })
echo schema.check({ name: 'a name that is far too long', age: 200.5 })
{valid: true, errors: []}
{valid: false, errors: [{field: name, message: The name field must not exceed 10 characters.}, {field: age, message: The age field must be an integer.}]}
check_or_raise() raises instead of returning. extend(), only() and
except() build one schema from another, which is how a create schema and
an update schema stay in sync.
types
Type predicates, one per type, as an alternative to the is_* built-ins:
import types
echo types.of(42)
echo types.int(42)
echo types.int(4.2)
echo types.digit('7')
echo types.alpha('a')
echo types.iterable([1])
echo types.instance(ValueError('x'), Error)
number
true
false
true
true
true
true
Each one answers a question and returns a boolean; none of them
converts anything. types.int(4.2) is false because 4.2 is not an
integer, not because it failed to become one.
types.of() is typeof(). digit(), alpha() and char() are the
string-shape tests the built-ins do not cover, and instance(value, Class)
walks the inheritance chain.
When you want conversion rather than a question, the methods on the value
do it: to_number(), to_string(), to_bigint(), to_list(),
to_bytes().
set
import set
var s = set.set([1, 2, 2, 3])
echo s.length()
3
Union, intersection, difference and subset tests, with insertion order preserved.
enum
import enum
var Color = enum.enum(['RED', 'GREEN'])
echo Color.RED
0
Pass a dictionary instead of a list to choose the values yourself.
array
Typed, fixed-width numeric arrays: Int8Array through Uint64Array, plus
FloatArray and DoubleArray. They store values in their declared width
rather than as doubles, which matters for memory and for talking to binary
formats.
import array
var ints = array.Int32Array([1, 2, 3])
echo ints.length()
3
The System
os
Processes, the filesystem, paths and the environment. Covered in Chapter 9.
os.exec(command) runs a shell command and gives you its output.
os.spawn(command, args, options) starts a process you can talk to.
os.on_signal(name, handler) installs a signal handler.
os.at_exit(handler) registers cleanup that runs however the program
ends, and os.set_exit_code(code) decides the status it ends with
without ending it there and then.
env
Configuration: a .env file read into the process environment, and
values read back out already converted. Covered in
Chapter 19.
import env
env.load()
var port = env.int('PORT', 8080)
var debug = env.bool('DEBUG', false)
var secret = env.require('SESSION_SECRET')
Names already set in the environment are left alone, so the file is a set of defaults and the deployment is what overrides them.
io
Standard streams, terminal control and in-memory files:
import io
var name = io.readline('Your name: ')
var secret = io.readline('Password: ', true)
echo io.stdout.is_tty()
io.TTY puts the terminal into raw mode, reads single keypresses and moves
the cursor, which is what an interactive program needs. io.BytesIO is the
in-memory file from Chapter 10.
io.capture(body) collects everything body writes to standard output
instead of printing it, echo, print() and io.stdout alike:
import io
def greet(name) {
echo 'Hello, ${name}!'
}
var out = io.capture(@{ greet('Ada') })
echo 'captured ${out.length()} characters'
captured 12 characters
Captures nest, and capture_begin()/capture_end() are the manual pair
for when the body might raise and you want its output anyway. This is
what lets Chapter 23 assert on what a function
prints, and keep a passing test’s output out of the report.
test
Suites, matchers, mocks, snapshots and reports. Covered in Chapter 23.
import test { * }
describe('slug', @{
it('lowercases and joins', @{
expect(slug('Hello World')).to_be('hello-world')
})
})
run()
stat
The S_IS* predicates over the mode word from file().stats(), plus a
renderer for it:
import stat
file('notes.txt', 'w').write('x')
file('notes.txt').chmod(0c644)
var info = file('notes.txt').stats()
echo stat.S_ISREG(info.mode)
echo stat.S_ISDIR(info.mode)
echo stat.file_mode(info.mode)
true
false
-rw-r--r--
S_ISREG, S_ISDIR, S_ISLNK, S_ISCHR, S_ISBLK, S_ISFIFO and
S_ISSOCK each answer one question about the kind of entry.
S_IMODE(mode) strips the type bits and leaves the permissions;
file_mode(mode) renders the whole thing the way ls -l does.
args
A command-line parser with subcommands, typed options, automatic
--help and wrapped terminal output. Covered in
Chapter 20.
import args
var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
parser.add_command('history', 'Show past greetings')
var parsed = parser.parse()
The help text is written from the same declarations that read the arguments, so the two cannot drift apart.
log
import log
log.info('server started')
log.error('connection refused')
A Logger binds structured fields, a child() logger inherits them, and
transports send records to the console, a file or somewhere you write
yourself.
isolate
Concurrency. Covered in Chapter 11.
ffi
Calling C and Rust. ffi loads shared libraries, reads their C headers
or Rust sources as declarations, calls their functions with checked
conversions, and turns Zuri functions into callbacks; it links static
libraries into loadable ones too.
import ffi
var c = ffi.open(ffi.LIBC).declare('size_t strlen(const char *s);')
echo c.strlen('interop')
7
Covered in Chapter 26.
The Network
net
TCP, UDP, unix domain sockets, TLS, DTLS, addresses and polling. Covered in Chapter 12.
http
Client and server, HTTP/1.1 and HTTP/2, with routing, middleware, WebSockets, server-sent events, multipart uploads, static files and a reverse proxy. Introduced in Chapter 12, covered fully in Chapter 15, and used throughout Chapter 25.
rpc
JSON-RPC 2.0, the protocol for calling methods in another program, on
both ends. A service answers calls with handlers, behind an http
route or on a connection; a client calls a service over HTTP; and an
endpoint on a WebSocket, a socket, a child process’s standard streams
or a pipe between isolates both answers and calls, so either side can
call the other at any time. rpc.serve() gives every connection to a
listening socket an endpoint of its own, over TCP, TLS or a Unix
domain socket.
import rpc
import isolate
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('add', @(params) => params[0] + params[1])
.serve()
}, ends[1])
var client = rpc.endpoint(ends[0])
echo client.request('add', [2, 3])
client.close()
5
Covered in Chapter 28.
Databases
sql
One way to talk to a relational database, whichever one it is. sql
defines what an adapter has to provide and supplies everything that is
the same across engines: parameters, transactions and savepoints,
cursors, pooling, introspection, and one error hierarchy. SQLite,
PostgreSQL and MySQL adapters ship with it, and changing between them
means changing the connection string.
import sql
var db = sql.open('sqlite://./app.db')
var id = db.insert('posts', { title: 'Hello' })
for post in db.query('select * from posts where id = ?', [id]) {
echo post.title
}
Covered in Chapter 17.
mail
Messages and the three protocols that move them, written in Zuri from
the socket up. mail.message() builds a message out of text, HTML and
files and works out the MIME tree from what went in; mail.parse()
reads one back. mail.smtp sends, mail.imap reads mail where it is
kept, and mail.pop3 takes it away. Both server ends are here too: an
SMTP server that decides what to accept through handlers of your own,
and an IMAP server that answers out of a mail store, of which one keeps
mail on disk in Maildir format and one keeps it in the process.
mail.dkim signs outgoing mail and checks incoming mail.
import mail
mail.send('smtp://mail.example.com', mail.message({
from: 'reports@example.com',
to: 'ann@example.com',
subject: 'Quarterly report',
text: 'The numbers are in.',
}), { username: 'reports', password: secret })
Covered in Chapter 18.
Compression
compress
deflate, zlib, gzip, zstd, lz4, bzip2 and brotli, plus tar
and zip archives and checksum for CRC32 and Adler-32. Covered in
Chapter 10.
Graphics
imagine
Image creation and manipulation on an RGBA buffer: drawing primitives, text with a built-in stroke font, filters, colour-space conversion, and reading and writing the common formats.
import imagine { Image }
Image(400, 200, '#0f172a')
.fill_circle(200, 100, 70, '#38bdf8')
.circle(200, 100, 70, 'white', { thickness: 3 })
.save('badge.png')
Almost every method returns an image, so operations chain.
Image.open(path) decodes an existing file, with the format taken from the
contents rather than the extension:
Image.open('photo.jpg')
.thumbnail(400, 400)
.save('thumb.webp')
Decoding and encoding are native; everything between them — filters,
drawing, colour conversion — is ordinary Zuri you can read in libs/imagine
and extend. Chapter 16 is the full treatment.
The Language Itself
zuri
Lexing, parsing, compiling and reflection, all reachable from Zuri code:
import zuri
echo zuri.tokenize('var x = 1').length() > 0
echo zuri.reflect.kind([1])
true
list
Chapter 21 is the full treatment.
math
The constants. Everything else is a method on number; see
Chapter 4.
Finding the Rest
Every module’s source is in libs/, and every public function in it
carries a doc block with its parameters, its defaults and its edge cases.
Reading libs/set.zu is a faster way to learn set than any summary, and
the standard library is written to be read.
Wire
Wire is Zuri’s built-in templating engine. Unlike most templating languages, Wire does not invent its own syntax on top of your markup. Every Wire feature is powered by attributes on ordinary HTML elements, which means anything a designer already knows about HTML transfers directly, and every valid HTML5 document is already a valid Wire template.
Under the hood, a Wire template is compiled once — not re-interpreted
on every render — into a small instruction tree, using the exact same
WHATWG-conformant parser that backs the html module.
That has a consequence worth knowing up front: Wire understands your
markup as structure, not as text. It knows that one interpolation
sits inside a paragraph, another inside an href, and a third inside a
<script> tag, and it escapes each one correctly for where it actually
is. A value handed to a template can never turn into new markup by
accident — that has to be requested explicitly, and Wire makes you say
so out loud.
- Introduction
- Rendering Templates
- Displaying Data
- The Expression Language
- Filters
- Conditionals
- Loops
- Content and Attributes
- Comments
- Including Templates
- Template Inheritance
- Security
- Extending Wire
- Configuration
- Compiling and Caching
- Error Handling
- Full Directive Reference
- Cheat Sheet
Introduction
Wire and Blade
If you have used Laravel’s Blade, PHP’s own answer to templating, a lot of Wire will feel familiar in spirit even though the syntax is different. Both engines compile templates rather than interpret them on every request, both let you extend a base layout and override named sections of it, and both escape everything by default so that printing a user’s name can never become a way for that user to run script in someone else’s browser.
Where Wire departs from Blade is in how it expresses control flow.
Blade adds its own directive syntax on top of plain text
(@if, @foreach, {{ }}) that a template author has to learn as a
second language layered over HTML. Wire instead expresses everything
as attributes on the HTML you were already going to write:
{{-- Blade --}}
@if ($user->isAdmin())
<p>Welcome back, administrator.</p>
@endif
{{-- Wire --}}
<p x-if="user.is_admin">Welcome back, administrator.</p>
The practical benefit is that a Wire template can be handed to a
designer who has never seen Zuri and they can still read it: it is
HTML, with some attributes they can look up. It also means every Wire
template validates as HTML5, can be opened directly in a browser to
check its structure, and can be run through the html
module’s own tools (html.format(), a linter, a
selector query) without anything special-casing Wire’s own syntax.
Your First Template
Here is the smallest possible Wire template, rendered from a string:
import wire
echo wire.render_string('<p>Hello {{ name }}</p>', { name: 'Ada' })
# <p>Hello Ada</p>
{{ name }} is an interpolation: it evaluates the expression name
against the variables you supplied and writes the result into the
page, escaped for wherever it landed. Everything else in this guide is
built out of that one idea, plus a handful of x- prefixed attributes
that control whether and how many times an element is rendered.
Rendering Templates
Templates From Files
For anything beyond a one-off snippet, templates live in files. Build
a Wire instance, point it at a directory, and render by path:
import wire
var view = wire.wire()
view.set_root('./views')
echo view.render('pages/home', { user, posts })
render()’s first argument is a path relative to the root. The
.html extension is added automatically when the path as written
names no file, so 'pages/home' finds pages/home.html; a path that
already carries an extension ('pages/home.wire') is tried exactly as
given first. set_extension() changes what gets tried when none is
given.
Note A
Wireinstance is meant to be built once, configured, and reused for the life of your program — typically at startup, alongside however you already configure the rest of your application. Rendering does not mutate it, so it is safe to render from several places at once.
Templates From Strings
render_string() renders a template given directly as a string,
without touching the filesystem. It behaves identically to render()
in every other respect — the same directives, the same escaping, the
same filters — and any x-include or x-extend inside the string
still resolves against the configured root.
echo view.render_string('<p>{{ greeting }}</p>', { greeting: 'Hi there' })
The third argument names the source for error messages (it defaults
to <source>), which is worth passing when the string came from
somewhere with its own identity — a database row, a file you already
had open for another reason:
view.render_string(row.body, { user }, 'cms:page:${row.id}')
render_string() is not cached, since there is no file path to key a
cache entry on. Reach for render() for anything rendered more than
once.
The Template Root
Every path — the one passed to render(), and the one written in
every x-include and x-extend in every template — resolves inside
the configured root directory, and nowhere else. This is not a
convention; it is enforced by the loader on every single resolution,
and it matters for security, not
just organization.
view.set_root('./views')
view.root()
# '/home/you/project/views' — always the absolute path
The root does not have to exist yet. create_root() makes it, and
reports whether it had to:
if view.create_root() {
echo 'Created a fresh views/ directory.'
}
Wire never creates this directory on its own initiative — a typo in a root path should read as “template not found,” not silently produce an empty folder somewhere unexpected.
Displaying Data
You have already seen the basic form. {{ expression }} evaluates
whatever is between the braces and writes the result into the page:
<h1>{{ post.title }}</h1>
<p>By {{ post.author.name }}</p>
expression is not limited to a bare variable name — it is a full
expression, covered in its own section below — so all of the following
are valid:
<p>{{ post.views > 1000 ? 'Popular' : 'New' }}</p>
<p>{{ post.tags|join(', ') }}</p>
<p>{{ user.nickname ?? user.name }}</p>
Escaping Data
Every interpolation is escaped for the specific place it lands, and
this cannot be turned off from inside a template. If name holds
<b>Ada</b>:
<p>Hello {{ name }}</p>
renders as:
<p>Hello <b>Ada</b></p>
which is exactly what you want when name came from a form field, a
database column, or anywhere else outside your own control. Wire is
not guessing at this — because it compiles a real parsed document
rather than gluing strings together, it always knows precisely which
kind of place an interpolation sits in, and it applies the escaping
that place needs:
| Where the interpolation is | What happens |
|---|---|
| Text between tags | &, <, > become entities |
An ordinary attribute (title, alt, data-*, …) | & and " become entities |
href, src, action, and other URL attributes | the value’s URL scheme is checked too |
<script>, or an on* event handler attribute | the value is encoded as JSON |
<style> | anything outside a CSS-safe set is dropped |
That is a materially stronger guarantee than “the special characters
got replaced”: a value dropped into an href cannot smuggle in a
javascript: scheme, and a value dropped into a <script> block
cannot break out of its string literal, close the tag early, or open a
comment — even though none of those things are <, >, or &.
Rendering Raw Markup
Escaping everything by default is right, but sometimes you genuinely
have markup — the output of another render, a snippet you built and
trust — and you want it written out as markup. The raw filter is how
you say so:
<div class="article-body">{{ post.rendered_html|raw }}</div>
You can build the same value on the Zuri side and hand it to a
template already marked as safe, with wire.safe():
import wire
view.render_string('<div>{{ body }}</div>', {
body: wire.safe('<em>already-trusted markup</em>'),
})
Warning
rawandwire.safe()are a promise that the value is safe to write unescaped at the place it lands. Applying either to anything a user submitted — a comment, a bio, a search query — reopens exactly the cross-site scripting hole the rest of Wire exists to close. Only reach for them on markup your own code produced.
If you need to render markup as text on purpose — showing someone a
literal <script> tag in a code sample, say — that is what plain
interpolation already does; there is nothing extra to opt into.
Wire and JavaScript Frameworks
Wire uses {{ }} for interpolation, the same delimiter many
JavaScript templating libraries (Vue, Angular, Handlebars, Mustache)
use for their own. If a Wire template also contains inline script that
a browser-side framework is meant to interpret, escape the braces with
a leading % so Wire leaves them alone:
<div id="app">
<p>{{ user.name }}</p> {{-- rendered by Wire, server-side --}}
<p>%{{ message }}</p> {{-- left as literal {{ message }}, for Vue --}}
</div>
%{{ renders as the literal text {{, and %{! does the same for
template function calls. If you find yourself
escaping braces constantly because most of a template belongs to a
client-side framework, it may be worth keeping that section in its own
file and serving it untouched rather than through Wire at all.
The Expression Language
Everything between {{ and }} — and everything given to x-if,
x-for, x-attr, and the rest of the directives below — is written in
Wire’s own small expression language. It is deliberately not a full
embedded copy of Zuri: there is no assignment, no way to declare
anything, and no way to reach a global variable. A template describes
a page; the logic that decides what the page contains belongs in the
code that calls render(), not in the template itself.
Literals
{{ 42 }} {{-- a number --}}
{{ 3.14 }} {{-- a decimal --}}
{{ 'a string' }} {{-- single quotes --}}
{{ "a string" }} {{-- or double, interchangeably --}}
{{ true }} {{-- true / false / nil --}}
{{ [1, 2, 3] }} {{-- a list literal --}}
{{ { id: 1, name } }} {{-- a dict literal; { name } is short for { name: name } --}}
{{ 0..pages }} {{-- a range, exactly as Zuri writes one --}}
Looking Up Values
{{ user.name }} {{-- a dotted lookup --}}
{{ user.address.city }} {{-- chains as deep as you like --}}
{{ items.0 }} {{-- a numeric key reads a list position --}}
{{ items[index] }} {{-- a computed index --}}
{{ items[-1] }} {{-- negative counts from the end --}}
{{ items.length }} {{-- a collection answers `length` by name --}}
Reading a variable that was never supplied gives nil rather than
raising, and reading a key off nil gives nil too. That is what
makes an optional value safe to reach through without a guard in front
of it:
{{ user.profile.avatar_url }}
renders as nothing at all if user has no profile, instead of
failing three levels down. Reach for x-if when the
difference between “empty” and “genuinely missing” matters to what you
show.
Note A name starting with an underscore can never be read from a template, the same way Zuri treats
_fieldas private. If you find yourself wanting to read one, expose a public accessor from Zuri instead.
Operators
{{ price * quantity }}
{{ subtotal + tax }}
{{ stock - reserved }}
{{ total / count }}
{{ index % 2 }}
{{ a == b }} {{ a != b }}
{{ a < b }} {{ a <= b }}
{{ a > b }} {{ a >= b }}
{{ 'admin' in user.roles }}
{{ 'admin' not in user.roles }}
{{ is_admin and is_active }}
{{ is_admin && is_active }} {{-- && is the same as `and` --}}
{{ is_guest or is_banned }}
{{ is_guest || is_banned }} {{-- || is the same as `or` --}}
{{ !is_active }}
{{ not is_active }} {{-- ! is the same as `not` --}}
{{ stock > 0 ? 'In stock' : 'Sold out' }}
{{ nickname ?? name }}
A string on either side of + concatenates rather than raising:
{{ 'Hello, ' + user.name }}
?? and or look similar but answer different questions, and mixing
them up is the single most common Wire mistake:
??asks “was this ever supplied?” and only falls back onnil.orasks “is this worth showing?” using Wire’s own truthiness, and falls back on anything falsy —nil,false, an empty string, an empty collection, or the number0.
{{ discount ?? 0 }} {{-- a missing discount becomes 0 --}}
{{ discount or 0 }} {{-- a discount that IS 0 also becomes 0, harmlessly here --}}
{{ stock_count ?? 'unknown' }} {{-- 0 in stock still shows as 0 --}}
{{ stock_count or 'unknown' }} {{-- 0 in stock is treated the same as never supplied --}}
Use ?? whenever zero is a legitimate value you want to keep, and
or whenever you only want to show something when there is genuinely
something to show.
Truthiness
x-if, x-not, and/or, !/not, and ? : all use Wire’s own
notion of truthy and falsy, which differs from Zuri’s own rules in one
place, deliberately, for template authoring:
| Value | Wire |
|---|---|
nil, false | falsy |
0, NaN | falsy |
| any other number, including negative ones | truthy |
| an empty string | falsy |
| a non-empty string | truthy |
| an empty list or dict | falsy |
| a non-empty list or dict | truthy |
| anything else | truthy |
The difference is the empty collection. Zuri treats [] and {} as
truthy; Wire treats them as falsy, so x-if="results" correctly hides a
section for a search that came back with nothing.
Calling Functions
A value registered as a global can be called directly:
<a href="{{ route('user.profile', user.id) }}">{{ user.name }}</a>
There is also an older, standalone spelling for calling a function that takes no arguments, kept from Wire’s very first version because it reads well on its own:
{! current_year !}
is exactly the same as writing:
{{ current_year() }}
Prefer {{ fn() }} for anything new; {! !} exists for templates that
already use it and for the handful of cases — a page’s build stamp, a
feature flag — where a function taking nothing at all reads a little
cleaner without the parentheses.
Filters
A filter transforms the value on the left of a |. It is the same
idea as a Unix pipe:
{{ name|upper }}
{{ price|round(2) }}
{{ post.body|truncate(150) }}
The value being filtered is always the filter’s first argument;
anything written in parentheses follows it. truncate(150) above
calls the truncate filter as truncate(post.body, 150).
Chaining Filters
Filters read left to right, each one’s output feeding the next:
{{ name|trim|title }}
{{ comment.body|strip_tags|truncate(200) }}
The = Argument Shorthand
For a filter that takes exactly one argument, name=value is
shorthand for name(value):
{{ status|is='active' }}
is the same as:
{{ status|is('active') }}
This spelling exists for symmetry with Wire’s very first version and reads naturally for a short comparison; the parenthesised form is equally valid everywhere and is the only option once a filter needs more than one argument.
Available Filters
Escaping
| Filter | What it does |
|---|---|
raw | Marks the value as markup, skipping escaping entirely. |
escape / e | Escapes for a context other than the one the value is being written into. Takes 'text' (the default), 'attribute', 'url', 'script', or 'style'. |
Text
| Filter | What it does |
|---|---|
upper | Converts to upper case. |
lower | Converts to lower case. |
title | Title Cases Every Word. |
capitalize | Capitalizes only the first letter, leaving the rest alone. |
trim | Removes leading and trailing whitespace of every kind (space, tab, newline). |
truncate(length, suffix?) | Cuts to length characters, appending suffix (default '…') only if anything was actually cut. |
replace(search, replacement?) | Replaces every literal occurrence of search. Never a regular expression. replacement defaults to an empty string. |
lpad(width, fill?) | Pads on the left to width characters, with fill (default a space). |
rpad(width, fill?) | Pads on the right. |
repeat(count) | Repeats the value count times. |
nl2br | Turns line breaks into <br>, escaping the text first. Returns markup. |
line_breaks | An alias for nl2br. |
strip_tags | Removes every HTML tag, keeping only the text — a real parse, not a pattern match. |
slug | Lower-cased, hyphen-separated, safe for a URL segment. |
url_encode | Percent-encodes for use inside a URL. |
json | Encodes as JSON text. |
json_script(id?) | Wraps the value’s JSON encoding in a <script type="application/json">, optionally with an id, ready to be read back by a script on the page. Returns markup. |
Numbers
| Filter | What it does |
|---|---|
abs | Removes the sign. |
round(places?) | Rounds to places decimal places (default 0), halves rounding away from zero. |
floor | Rounds down to the nearest whole number. |
ceil | Rounds up. |
number_format(places?, point?, separator?) | Groups thousands and fixes the decimal places, like 1,234,567.89. Pass point/separator to use another convention, e.g. 1.234.567,89. |
filesize(binary?) | A byte count written the way a person reads it — 1.4 MB by default, or 1.3 MiB with binary set to true. |
Collections
| Filter | What it does |
|---|---|
length | How many entries — works on a string, list, dict, or bytes. nil has length 0. |
first | The first entry, or nil if there is none. |
last | The last entry. |
join(glue?) | Joins entries into one string with glue (default '') between them. |
sort(key?) | Sorted ascending; key sorts a list of dicts or instances by one field. |
reverse | The entries backwards, or a string reversed. |
unique | Duplicates removed, keeping the first of each. |
keys | A dict’s keys, in insertion order. |
values | A dict’s values, in insertion order. |
slice(start, end?) | The entries from start up to but not including end. Negative positions count from the end. |
sum(key?) | Adds the entries together; key sums one field of a list of dicts or instances. |
split(separator?) | Splits a string into a list. separator defaults to any run of whitespace. |
Choice
| Filter | What it does |
|---|---|
default(fallback) / alt | fallback when the value is falsy, otherwise the value. |
empty | Whether the value has nothing in it. Unlike falsiness, a number is never empty — not even 0. |
is(expected) | Whether the value equals expected. |
not(expected) | Whether it differs. |
Dates
| Filter | What it does |
|---|---|
date(format?) | Formats a date.Date, a Unix timestamp, or a parseable date string, using the same format directives as Date.format(). Defaults to 'Y-m-d H:i:s'. |
<time datetime="{{ post.published_at|date('Y-m-d') }}">
{{ post.published_at|date('jS F Y') }}
</time>
Writing Your Own Filter
See Custom Filters below.
Conditionals
x-if, x-elif, and x-else
x-if renders an element, and everything inside it, only when its
expression is truthy:
<p x-if="user.is_admin">You have administrator access.</p>
If user.is_admin is falsy, the whole <p> — tag and contents — is
left out of the page entirely. There is no empty element left behind.
Chain further conditions with x-elif, and close the chain with a
plain x-else:
<p x-if="user.role == 'admin'">Administrator</p>
<p x-elif="user.role == 'staff'">Staff member</p>
<p x-elif="user.role == 'contributor'">Contributor</p>
<p x-else>Member</p>
Exactly one of these renders. Only whitespace and HTML comments are
allowed between the elements in a chain — any real content in between
ends it, and an x-elif or x-else with no x-if in front of it is
rejected when the template compiles, not silently ignored:
{{-- this chain is broken by the text between the two elements --}}
<p x-if="a">A</p>
some text
<p x-else>not A</p> {{-- error: x-else has no x-if in front of it --}}
x-not
x-not is the plain inverse of x-if — it renders when its expression
is falsy — and does not take part in a chain:
<div x-not="user.has_verified_email">
<p>Please verify your email address.</p>
</div>
x-if="!condition" and x-not="condition" mean the same thing;
x-not exists because it often reads more naturally for a guard
clause.
Loops
x-for repeats an element once per entry of whatever its expression
evaluates to — a list, a dict, a string, or a range:
<ul>
<li x-for="posts" x-value="post">{{ post.title }}</li>
</ul>
Notice that the element itself repeats, not just its contents — the
example above produces one whole <li> per post, not one <li>
wrapping every post. x-value names the variable each iteration binds
its current entry to; leave it out entirely if you do not need to
refer to the entry by name.
An optional x-key binds the position (for a list) or the key (for a
dict):
<tr x-for="users" x-key="id" x-value="user">
<td>{{ id }}</td>
<td>{{ user.name }}</td>
</tr>
Every kind of collection iterates naturally:
<li x-for="tags" x-value="tag">{{ tag }}</li> {{-- a list --}}
<li x-for="scores" x-key="name" x-value="score"> {{-- a dict --}}
{{ name }}: {{ score }}
</li>
<span x-for="word" x-value="letter">{{ letter }}</span> {{-- a string, by character --}}
<option x-for="1..5" x-value="n">{{ n }}</option> {{-- a range --}}
A missing or empty collection simply renders nothing — there is no
need to guard a loop with an x-if first:
<li x-for="comments" x-value="comment">{{ comment.body }}</li>
{{-- renders nothing at all if `comments` is empty or was never supplied --}}
The loop Variable
Every iteration publishes a loop variable with the following fields:
| Field | Value |
|---|---|
loop.index | The position, counting from 1. |
loop.index0 | The position, counting from 0. |
loop.first | true on the first pass. |
loop.last | true on the last pass. |
loop.length | How many entries there are in total. |
loop.even | true on the 2nd, 4th, 6th, … pass — even-numbered by loop.index. |
loop.odd | true on the 1st, 3rd, 5th, … pass. |
loop.key | The current key or index, whether or not x-key binds it too. |
loop.value | The current value, whether or not x-value binds it too. |
loop.parent | The enclosing loop’s own loop, for a nested x-for. |
<tr x-for="rows" x-value="row" x-attr="{ 'class': loop.odd ? 'zebra' : nil }">
<td>{{ loop.index }}</td>
<td>{{ row.name }}</td>
</tr>
Nesting Loops
x-loop renames the metadata variable a specific x-for publishes,
which is what lets an inner loop’s own loop and an outer loop’s
loop both be reached at once:
<table x-for="rows" x-loop="row" x-value="cells">
<tr>
<td x-for="cells" x-value="cell">
{{ row.index }}, {{ loop.index }}: {{ cell }}
</td>
</tr>
</table>
Without x-loop, the inner loop’s own loop would simply shadow the
outer one for the scope of the inner loop — reach for loop.parent
instead if you would rather not rename anything:
<td x-for="cells" x-value="cell">
outer position {{ loop.parent.index }}, inner position {{ loop.index }}
</td>
Looping Without a Wrapper Element
Sometimes you want to repeat several elements together without a real
element wrapping them, or without introducing any element at all. Put
the directive on a <template> instead of on the element you want
repeated:
<select>
<template x-for="countries" x-value="country">
<option value="{{ country.code }}">{{ country.name }}</option>
</template>
</select>
<template> is the one HTML element the parser lets through
completely untouched, wherever it appears — inside a <head>, inside
a <table>, inside a <select> — and Wire removes it from the output
entirely, leaving only what was inside it, once per pass. This is also
the way to apply x-if to a group of elements without
picking one of them to carry the attribute:
<template x-if="user.is_admin">
<a href="/admin">Dashboard</a>
<a href="/admin/users">Users</a>
</template>
Content and Attributes
x-text
x-text replaces an element’s children with its expression, escaped
as plain text — useful when the element already has other attributes
and you would rather not write the value twice:
<p x-text="post.summary"></p>
is the same as:
<p>{{ post.summary }}</p>
Anything written inside the element in the source is discarded; it exists only to describe what would go there without JavaScript.
x-html
x-html is x-text’s unescaped counterpart — it replaces the
element’s children with its expression’s value, written as markup
rather than text:
<div x-html="post.rendered_body|raw"></div>
Note the |raw: x-html still expects a value marked safe, the same
as an ordinary interpolation would. x-html only changes where the
markup goes (replacing the whole element’s contents, rather than
sitting inline in a text run) — it does not, on its own, turn escaping
off. The same warning about untrusted input
applies here just as much as it does to the raw filter.
x-attr
x-attr spreads a dictionary of names and values onto the element as
attributes:
<input x-attr="{ type: 'text', name: field.name, required: field.is_required }">
Within that dictionary:
trueproduces a valueless attribute (required), exactly as writingrequiredby hand would.falseandnilleave the attribute off entirely, rather than writing it with an empty or literal"false"value.- Anything else is written as the attribute’s value, escaped for whatever that attribute means (a URL attribute is checked as a URL, same as always).
This is what turns a boolean into a real HTML boolean attribute without a ternary in every place one is needed:
<button x-attr="{ disabled: !form.is_valid }">Submit</button>
A computed attribute takes over from one written directly on the same element, so you can set a sensible default and only override it when there is something to override:
<div class="card" x-attr="{ 'class': featured ? 'card card-featured' : nil }">
Comments
HTML’s own comment syntax is a server-side note in a Wire template. It never reaches the rendered page, and nothing inside one is evaluated — which is exactly what makes it safe to leave notes for other people maintaining the template, without those notes shipping to a browser:
<!-- TODO: replace this hard-coded banner once marketing sends the real copy -->
<div class="banner">Coming soon</div>
<!-- this variable is not evaluated: {{ some.internal.detail }} -->
If you genuinely want a comment to reach the browser — a conditional
comment, a build stamp a deploy script checks for — turn comments back
on for that Wire instance with set_comments(true).
Including Templates
<include path="..." /> renders another template in its place. It
also has a directive spelling, <template x-include="...">, which
means exactly the same thing — the pseudo-element forms in this guide
(<include>, <extend>, <declare>, <define>) are sugar, rewritten
into the directive form before the template is even parsed, which is
what lets them work correctly inside a <head> or a <table> where an
element the parser does not recognise would otherwise be moved or
dropped.
<include path="partials/nav.html" />
is the same as:
<template x-include="partials/nav.html"></template>
The path is resolved inside the template root,
the same as render()’s own path argument.
By default, an include sees every variable the page around it can see — it is not a separate scope:
{{-- page.html --}}
<include path="partials/greeting.html" />
{{-- partials/greeting.html --}}
<p>Hi, {{ user.name }}!</p>
renders correctly without user ever being passed explicitly to the
partial.
Passing Data to an Include
x-with adds variables for the included template, alongside whatever
it already inherits:
<include path="partials/badge.html" x-with="{ label: 'New', tone: 'green' }" />
only withholds the surrounding scope entirely, leaving the partial
with only what x-with gave it — turning a plain partial into
something closer to a real component with a defined interface:
<include path="components/price-tag.html" x-with="{ amount: item.price }" only />
Includes as Components
Anything written inside an <include> tag is handed to the included
template as a named region called content, declared with
<declare name="content">:
{{-- components/card.html --}}
<section class="card">
<h3>{{ title }}</h3>
<declare name="content"></declare>
</section>
{{-- the page --}}
<include path="components/card.html" x-with="{ title: 'Recent Orders' }">
<p>{{ orders|length }} orders this week</p>
</include>
renders as:
<section class="card">
<h3>Recent Orders</h3>
<p>3 orders this week</p>
</section>
The content you write inside the <include> tag keeps the scope it
was written in — the page’s, not the partial’s — which is exactly what
lets {{ orders|length }} above read a variable from the page even
though x-only was not used and the card component never mentions
orders at all.
Computed Include Paths
A path is read as literal text, with {{ }} interpolation allowed
inside it — not as an expression in its own right, so
x-include="header" names a file called header, rather than reading
a variable of that name:
<include path="themes/{{ current_theme }}/header.html" />
resolves a different file per render depending on current_theme,
while everything after path="themes/" up to the next {{ is taken
literally.
Template Inheritance
Includes are for small, reusable fragments — a navbar, a footer, a badge. For whole-page structure — the parts of a layout every page on your site shares — template inheritance is the better fit. Where an include composes fragments together, inheritance lets a base layout define the shape of a page once, and lets each page fill in only the parts that differ.
Defining a Layout
A base template marks the regions a page is allowed to fill in with
<declare name="...">:
{{-- layouts/app.html --}}
<!DOCTYPE html>
<html>
<head>
<title>{{ title }}</title>
<declare name="head"></declare>
</head>
<body>
<nav><!-- shared navigation --></nav>
<main>
<declare name="content"></declare>
</main>
<footer>
<declare name="footer">
<p>© {{ year }} My Company</p>
</declare>
</footer>
</body>
</html>
Notice that footer has content already inside its <declare> tag.
That becomes the default — what renders when
a page does not define that region at all.
Extending a Layout
A page declares which layout it extends with <extend base="...">,
and fills in regions with <define name="...">:
{{-- pages/dashboard.html --}}
<extend base="layouts/app.html">
<define name="head">
<meta name="description" content="Your account dashboard.">
</define>
<define name="content">
<h1>Welcome back, {{ user.name }}</h1>
<p>You have {{ notifications|length }} new notifications.</p>
</define>
</extend>
Rendering pages/dashboard.html produces the layout’s full structure,
with head and content replaced by what the page defined, and
footer left exactly as the layout’s own default. Both the layout and
the page are rendered against the same variables render() was given
— there is no separate scope to pass anything through.
Everything inside an <extend> must be inside a <define>. The base
template owns the page’s structure; a page’s job is only to fill in
the regions it was offered, not to add markup of its own outside them.
A template can extend exactly one base.
Default Slot Content
A region a page does not define keeps whatever the layout put inside
its own <declare> tag:
{{-- pages/minimal.html --}}
<extend base="layouts/app.html">
<define name="content">
<p>Just this page's content — head and footer both use the layout's defaults.</p>
</define>
</extend>
A <declare> tag with nothing inside it — like head in the layout
above — simply renders empty when nothing defines it.
x-super
<super /> — or its directive spelling, <template x-super></template>
— renders whatever the region it sits inside would have shown before
this definition replaced it. That lets a page add to a section
instead of fully restating it:
<extend base="layouts/app.html">
<define name="footer">
<super />
<p><a href="/privacy">Privacy Policy</a></p>
</define>
</extend>
renders the layout’s own copyright line, followed by the extra privacy link — without the page having to know or repeat what the layout’s default footer actually said.
Multi-Level Inheritance
A template that extends a base can itself declare regions of its own, letting a further template extend it. This is how a site with, say, a general layout and several page-type-specific layouts (a blog post, a product page) is usually structured:
{{-- layouts/app.html --}}
<html><body>
<declare name="body"><p>default</p></declare>
</body></html>
{{-- layouts/article.html --}}
<extend base="layouts/app.html">
<define name="body">
<article>
<declare name="article-content"></declare>
</article>
</define>
</extend>
{{-- pages/post.html --}}
<extend base="layouts/article.html">
<define name="article-content">
<h1>{{ post.title }}</h1>
{{ post.body|raw }}
</define>
</extend>
Rendering pages/post.html walks the whole chain: it extends
layouts/article.html, which extends layouts/app.html, and the
final page is assembled from all three.
Overriding a Definition
Defining the same region twice in one template is almost always an
accident — two people editing the same file, a copy-paste that was
never cleaned up — so Wire refuses it at compile time unless you say
the replacement is deliberate with override:
<extend base="layouts/app.html">
<define name="content"><p>First draft</p></define>
<define name="content" override><p>Final version</p></define>
</extend>
Without override on the second one, this template fails to compile
with a clear message naming the region and both locations, rather than
silently keeping whichever definition happened to come last.
Security
Why Everything Is Escaped by Default
A templating engine’s job includes keeping a value that came from outside your program — a form submission, a query string, another user’s profile — from becoming markup, a script, or a link, unless you explicitly say it is safe to. Wire treats this as the default rather than an opt-in setting, for the same reason a seatbelt only works if putting it on is the ordinary thing to do rather than the thing you remember to do under pressure: the moment escaping is something you have to remember to turn on, the templates that skip it are the ones that get exploited.
Every interpolation is escaped for the exact place it lands — see the
table earlier in this guide — and there is no
template-level setting to disable that. The only way past it is the
explicit, visible raw filter or x-html
directive, which is exactly the friction you
want between “displaying a value” and “trusting a value with
unescaped markup.”
URLs Are Checked, Not Just Escaped
An attribute a browser reads as a URL — href, src, action, and
several others — gets more than entity escaping. Its scheme is checked
against an allowlist, and a disallowed scheme is replaced with
about:blank rather than written through:
<a href="{{ profile_link }}">{{ user.name }}</a>
If profile_link were javascript:alert(document.cookie), the
rendered href is about:blank, not the script URL. A URL with no
scheme at all — every relative link, every absolute path, every
protocol-relative URL — is always allowed, since none of those can
name a handler in the first place.
The default allowlist covers http, https, mailto, tel, and a
handful of others; it deliberately leaves out javascript,
vbscript, data, and file. set_url_schemes()
replaces the list for an application that genuinely needs another
scheme — a custom app: handler, say.
srcset and ping, which each hold a list of URLs, have every entry
in the list checked the same way, with descriptors (2x, 640w) and
separators preserved exactly as written.
Scripts and Stylesheets
A value interpolated inside a <script> element, or inside an event
handler attribute like onclick, is encoded as JSON rather than
merely escaped — automatically, with no filter needed:
<script>
var currentUser = {{ user }};
</script>
renders the whole user value as a JSON literal, with its own quotes
included. That is worth internalising, because it changes how you
write the surrounding script: do not wrap the interpolation in
your own quotes.
{{-- correct: renders var name = "Ada"; --}}
<script>var name = {{ user.name }};</script>
{{-- wrong: renders var name = '"Ada"'; which is not what you meant --}}
<script>var name = '{{ user.name }}';</script>
A value inside a <style> block has anything outside a conservative,
CSS-safe character set dropped rather than escaped — there is no
escape sequence that would stop a < from being read by the HTML
tokenizer looking for </style>, so the only sound answer is for the
character not to be there at all.
The Template Root Is a Sandbox
Every path — whether it came from render()’s own argument, or from
an x-include/x-extend written inside a template, even one built
from an interpolated value — is resolved inside the configured root
and refused if it resolves anywhere else, .. segments and symbolic
links both included:
<include path="{{ theme }}/header.html" />
If theme ever came from something a request controls, an unbounded
loader would turn this into a way to read arbitrary files the process
can reach. Wire’s loader treats the root as a hard boundary instead:
the worst a hostile value can do here is fail to find a template.
Extending Wire
Custom Filters
register_filter() adds a filter, or replaces one that already
exists by that name. The value being filtered is always the first
argument; anything the template passes after it follows:
view.register_filter('excerpt', @(value, words) {
var count = words ?? 25
return ' '.join(value.split('/\\s+/').take(count)) + '…'
})
<p>{{ post.body|excerpt(40) }}</p>
An argument the template leaves out arrives as nil, which is why
words ?? 25 above supplies a sensible default rather than the
filter insisting the template spell it out every time.
A filter returning a plain value has that value escaped like any
other — the same as an ordinary interpolation. A filter that genuinely
produces markup returns wire.safe(), and
from that point on is responsible for what is inside it, exactly like
raw and x-html are.
Globals
register_global() makes a value readable from every template
without it being passed to render() explicitly:
view.register_global('site_name', 'Example Inc.')
view.register_global('route', @(name) {
return '/' + name
})
<title>{{ page_title }} — {{ site_name }}</title>
<a href="{{ route('user.profile', user.id) }}">{{ user.name }}</a>
A variable of the same name passed to render() wins over a global —
a global is a default, not a hard override, so a page can still shadow
site_name for itself if it ever needs to.
Custom Elements
For the rare case the directives genuinely cannot express,
register_element() claims an HTML tag name outright and hands every
element of that name to a Zuri function instead of writing it out:
view.register_element('icon', @(view, element) {
var name = element.attributes.get('name', 'dot')
return wire.safe('<svg class="icon"><use href="#${name}"></use></svg>')
})
<icon name="star" />
The function receives the Wire instance and a dictionary describing
the element — its tag name, its attributes (already rendered and
escaped), its already-rendered children, and the scope it was written
in. Return wire.safe(markup) to write markup, any other value to
write it as escaped text, or nil to write nothing at all. Directives
still work on a claimed element exactly as they read
(<icon x-if="..." /> behaves the way it looks), and the whole
mechanism exists for genuinely dynamic, code-driven markup that a
partial and x-include cannot express — reach for an include first;
it is easier for the next person to read and does not require them to
go find the Zuri function behind it.
Configuration
Every setting can be passed to the Wire constructor at once, or set
individually with its own method afterward:
var view = wire.wire({
root: './views',
extension: '.html',
compact: true,
comments: false,
auto_reload: true,
url_schemes: ['https', 'mailto'],
})
| Option / method | Default | What it controls |
|---|---|---|
root / set_root(path) | ./templates | The directory every template path resolves inside. |
extension / set_extension(ext) | .html | Tried when a path names no file as written. Must start with .. |
compact / set_compact(bool) | false | Drops whitespace-only text between tags. |
comments / set_comments(bool) | false | Whether HTML comments reach the rendered output. |
auto_reload / set_auto_reload(bool) | true | Whether a cached template is checked against its file, and re-read if it changed, before each render. |
url_schemes / set_url_schemes(list) | see above | Which URL schemes an href/src/etc. is allowed to use. |
root() reports the currently configured root as an absolute path,
and create_root() makes the directory if it does not exist yet
(returning whether it had to).
Compiling and Caching
render() compiles a template into its instruction tree the first
time it is used, and keeps the compiled result for next time — parsing
and directive resolution happen once per template, not once per
request. Every subsequent render just walks that tree.
With auto_reload on (the default), each render checks the file’s
modification time and size against what was cached, and recompiles if
either changed. That is one filesystem stat per template per
render — cheap, and exactly what you want while actively editing
templates. Turn it off with set_auto_reload(false) once you are
running under load and are not editing templates live; clear_cache()
is then how a long-running process picks up a new deployment.
compile(path) compiles a template without rendering it, and returns
the compiled result — or raises, if the template has a mistake in it.
That makes it a good fit for a startup-time check across a whole
directory of templates, so a broken one is caught before the first
request that would have hit it:
for name in os.read_dir('./views', true) {
view.compile(name)
}
Error Handling
Everything Wire raises is a wire.WireError, and every one of them
carries the template path and the line/column of the tag responsible,
so catching the base class is enough to handle any template failure
the same way — turning it into a 500 page, logging it with context,
whatever your application needs:
catch {
echo view.render('pages/dashboard', { user })
} as e {
if instance_of(e, wire.WireError) {
log.error('template failed: ${e.reason} at ${e.location()}')
echo view.render('errors/500')
}
}
Three more specific errors, each a WireError, tell you what kind
of thing went wrong:
| Error | Raised when |
|---|---|
wire.TemplateSyntaxError | A template does not compile: an unknown directive, a malformed expression, a broken inheritance chain. Always caught the first time a template is used, not on a later render. |
wire.TemplateNotFoundError | A path names no file, or resolves outside the template root. |
wire.RenderError | Something that depends on the values a template was given: iterating something that cannot be iterated, a filter rejecting its input, an include chain that never bottoms out. |
A variable that was simply never supplied is not one of these — that renders as empty and tests as falsy, by design, so an optional section can be written without a guard around every single field it touches.
Full Directive Reference
Every directive Wire understands. An x- prefixed attribute outside
this list fails to compile — Wire owns that whole prefix, so a typo
like x-fi is caught the moment the template is compiled rather than
silently doing nothing.
| Directive | Takes | Meaning |
|---|---|---|
x-if | expression | Renders only if truthy. Starts a chain. |
x-elif | expression | Continues an x-if chain. |
x-else | (none) | Closes an x-if chain. |
x-not | expression | Renders only if falsy. Does not chain. |
x-for | expression | Repeats the element once per entry. |
x-value | name | Binds each iteration’s value. |
x-key | name | Binds each iteration’s key/index. |
x-loop | name | Renames the loop metadata variable. |
x-text | expression | Replaces children with escaped text. |
x-html | expression | Replaces children with unescaped markup. |
x-attr | expression (dict) | Spreads attributes onto the element. |
x-include | path | Renders another template in this element’s place. |
x-with | expression (dict) | Adds variables for an include. |
x-only | (none) | Withholds the surrounding scope from an include. |
x-extend | path | This template extends the named base. |
x-slot | name | Declares a region an extending template may replace. |
x-define | name | Replaces a base template’s region. |
x-override | (none) | Permits redefining a region already defined once. |
x-super | (none) | Renders the definition this one replaces. |
Five pseudo-elements are sugar for the directives above, rewritten before the template is parsed:
| Pseudo-element | Equivalent to |
|---|---|
<include path="p" /> | <template x-include="p"></template> |
<extend base="p">…</extend> | <template x-extend="p">…</template> |
<declare name="n">…</declare> | <template x-slot="n">…</template> |
<define name="n">…</define> | <template x-define="n">…</template> |
<super /> | <template x-super></template> |
Cheat Sheet
{{-- Variables --}}
{{ name }}
{{ user.address.city }}
{{ items.0 }}
{{ items[index] }}
{{-- Filters --}}
{{ name|upper }}
{{ price|round(2) }}
{{ post.body|truncate(150) }}
{{ status|is='active' }}
{{-- Raw markup --}}
{{ trusted_html|raw }}
{{-- Conditionals --}}
<p x-if="condition">…</p>
<p x-elif="other">…</p>
<p x-else>…</p>
<p x-not="condition">…</p>
{{-- Loops --}}
<li x-for="items" x-value="item" x-key="i">{{ i }}: {{ item }}</li>
{{ loop.index }} {{ loop.first }} {{ loop.last }} {{ loop.length }}
{{-- Wrapper-free groups --}}
<template x-for="items" x-value="item">…</template>
<template x-if="condition">…</template>
{{-- Content and attributes --}}
<p x-text="value"></p>
<div x-html="trusted_html|raw"></div>
<input x-attr="{ type: 'text', disabled: !enabled }">
{{-- Comments (never rendered) --}}
<!-- a note for the next person -->
{{-- Includes --}}
<include path="partials/nav.html" />
<include path="components/card.html" x-with="{ title: 'X' }" only>
content for the component's declared region
</include>
{{-- Inheritance --}}
<extend base="layouts/app.html">
<define name="content">…</define>
<define name="footer"><super />…</define>
</extend>
Further reading: the html module that Wire’s parser
and serializer are built on, and the date module
for the format directives the date filter accepts.
HTTP
The http module is Zuri’s HTTP stack: a client, a server, and the
pieces both are built from. It speaks HTTP/1.1 and HTTP/2, over
cleartext or TLS, and it is meant to face the internet directly.
That last part is a design decision rather than a boast. Most language runtimes ship an HTTP server that is fine for development and then expect a reverse proxy in front of it in production — something else terminates TLS, serves the static files, compresses the responses, enforces the request limits, and speaks HTTP/2 to the browser. Everything on that list is in this module, done properly, because a standard library that leaves them out is not really shipping a server.
- Introduction
- Making Requests
- Serving Requests
- Built-in Middleware
- Sessions
- TLS
- HTTP/2
- WebSockets
- Server-Sent Events
- Reverse Proxying
- Running in Production
- What the Module Refuses
- Module Reference
Introduction
A First Request
import http
echo http.get('https://example.com').as_text()
get(), post(), put(), patch(), delete(), head(),
options() and trace() all exist at module level and all go through
one shared client, which keeps its connections open between calls.
Request Methods covers each of them.
A First Server
import http
var server = http.server(3000)
server.get('/', @(request, response) {
response.html('<h1>Hello</h1>')
})
server.get('/users/:id', @(request, response) {
response.json({ id: request.param('id') })
})
server.listen()
Two things are already true of that server that are worth noticing:
GET /users/42 matches the second route and hands the handler
'42', and a HEAD / request is answered by the first route with the
headers a GET would have produced and no body, because that is what
HEAD means.
Blocks in this chapter that list several calls together — three ways to save, four filters, a family of methods — are reference listings, not programs. They show the shape of each call rather than a sequence you could run, and several would conflict if pasted into one file. Anything presented as a complete program on this page runs as written.
A Round Trip You Can Run
Most examples in this chapter call a host that does not exist, or start a server that never returns — neither of which you can paste into a terminal and watch. This one you can. It runs a real server on a real port and makes real requests against it, in two files.
The server goes in its own file, because a function that references an imported module cannot be spawned from the file that imported it. See Concurrency with Isolates for the rule.
Filename: service.zu
import http
def serve(port_channel) {
var server = http.server(0, '127.0.0.1')
server.get('/health', @(request, response) {
response.json({ status: 'ok' })
})
server.post('/echo', @(request, response) {
response.json({ heard: request.json_body().message }, 201)
})
# Bind first so the port is known, hand it back, then serve.
server.bind()
port_channel.send(server.socket.local_address().port())
server.listen()
}
Filename: main.zu
import http
import isolate
import .service
var channel = isolate.channel(1)
var task = isolate.spawn(service.serve, channel)
var port = channel.recv()
var api = http.client('http://127.0.0.1:${port}')
echo api.get('/health').as_dict()
echo api.post('/echo', { message: 'hello' }).status
task.cancel()
$ zuri run main.zu
{status: ok}
201
Four details in there are worth carrying into your own code.
Port 0 asks the operating system for a free port. listen() would
bind for you, but then nobody could ask which port it got, so the example
calls bind() explicitly, reads the bound address, and only then starts
accepting. That is the pattern for any test or any process running more
than one server.
The port travels over a channel. The client cannot know it in advance, and the two halves are in different isolates with separate heaps, so a shared variable is not an option.
request.json_body() parses the request body, and
response.json(value, status) writes the response. The names are not
symmetrical because they are doing different jobs: one decodes what
arrived, the other encodes and sets a status.
task.cancel() ends it. listen() runs until the server is closed, so
without that the program would never exit.
Making Requests
The Shared Client
The module-level functions use a single HttpClient that lives for
the life of the program. It pools connections, so a second request to
a host it has already talked to skips the TCP handshake and, over
HTTPS, the TLS handshake as well.
import http
var page = http.get('https://example.com')
echo page.status # 200
echo page.headers.get('content-type')
echo page.as_text()
Reach that client with http.shared_client() when you want a setting
to apply to every casual call in a program, and build your own for
anything more specific.
Building a Client
var api = http.client('https://api.example.com', {
headers: { 'Authorization': 'Bearer ' + token },
read_timeout: 5000,
})
var me = api.get('/users/me').raise_for_status().as_dict()
A client built with a base URL prefixes it onto any target that is not
already absolute, so the rest of the program addresses the service by
path. Everything in the options dictionary is a field of HttpClient,
plus headers, and every one of them can also be set afterwards:
api.user_agent = 'my-service/2.1'
api.max_redirects = 3
api.set_header('X-Client-Version', '2.1')
Note Build a client once and keep it. A fresh client per request throws away the connection pool, the cookie jar and the TLS configuration — which is to say, it throws away most of what a client is for.
Request Methods
Every method has a function of its own. They exist twice over, with
identical signatures: on the module, where they go through the shared
client, and on any HttpClient you build yourself.
The three methods that carry a body take it as their second argument; the rest take the options dictionary in that position.
| Call | Sends a body |
|---|---|
get(url, options) | no |
post(url, data, options) | yes |
put(url, data, options) | yes |
patch(url, data, options) | yes |
delete(url, options) | no |
head(url, options) | no |
options(url, options) | no |
trace(url, options) | no |
request(method, url, options) | if options says so |
Every argument after the first is optional, so the short forms all work:
import http
# Reading
http.get('https://example.com/items')
http.get('https://example.com/items', { query: { page: 2 } })
# Creating and updating
http.post('https://example.com/items', { name: 'Widget' })
http.put('https://example.com/items/1', { name: 'Widget', price: 9 })
http.patch('https://example.com/items/1', { price: 11 })
# Removing
http.delete('https://example.com/items/1')
# Asking about a resource without fetching it
var probe = http.head('https://example.com/large.iso')
echo probe.headers.get('content-length')
# Asking what a resource supports
echo http.options('https://example.com/items').headers.get('allow')
The same calls on a client of your own, which is what you want for anything that runs more than once:
var api = http.client('https://api.example.com')
api.get('/items')
api.post('/items', { name: 'Widget' })
api.delete('/items/1')
request() takes the method as an argument, which is how you send one
that has no function of its own — a WebDAV PROPFIND, or anything
else a service has invented:
api.request('PROPFIND', '/files/', {
headers: { 'Depth': '1' },
body: '<propfind xmlns="DAV:"><allprop/></propfind>',
content_type: 'application/xml',
})
Method names are case-sensitive on the wire, and every registered one is uppercase; a lowercase name is uppercased for you.
Note
head()does not follow redirects by default, anddelete(),options()andtrace()send no body. ADELETEthat genuinely needs one — some APIs want a reason in the body — can userequest('DELETE', url, { json: ... }).
Request Bodies
The second argument to post(), put() and patch() is the body,
and what it is decides how it is sent:
api.post('/items', { name: 'Widget', price: 9 }) # JSON
api.post('/items', ['a', 'b']) # JSON
api.post('/items', 'raw text') # sent as-is
api.post('/items', file('photo.jpg', 'rb').read()) # sent as-is
api.post('/items', form_builder) # multipart
A dictionary or list becomes JSON with a matching Content-Type; a
string or bytes is sent exactly as given, with no content type unless
you name one.
For anything else, or to be explicit, name it in the options dictionary:
| Option | Sends | Content-Type |
|---|---|---|
body | a string or bytes, exactly as given | none, unless content_type is set |
json | any value, JSON-encoded | application/json |
form | a dictionary | application/x-www-form-urlencoded |
multipart | a MultipartBuilder | multipart/form-data, boundary included |
content_type | — | overrides whichever of the above applied |
# A login form
api.post('/login', nil, {
form: { username: 'ada', password: secret },
})
# An explicit content type over a raw body
api.put('/documents/1', nil, {
body: markdown_source,
content_type: 'text/markdown; charset=utf-8',
})
# A form whose field repeats
api.post('/search', nil, {
form: { tag: ['new', 'featured'], q: 'zuri' },
})
Content-Length is set for you from the body, and cannot be
overridden — two parties disagreeing about how long a body is, is a
request-smuggling bug rather than a formatting choice.
Uploading Files
A file upload is a multipart/form-data body, which
MultipartBuilder assembles:
import http
var form = http.MultipartBuilder()
form.add_field('title', 'Holiday')
form.add_field('album', '2026')
form.add_file('photo', 'beach.jpg', file('beach.jpg', 'rb').read(), 'image/jpeg')
var response = http.post('https://example.com/photos', form)
Passing the builder as the body is enough — the Content-Type, the
boundary parameter and the Content-Length all follow from it.
| Method | Does |
|---|---|
add_field(name, value) | adds a plain form field; numbers and booleans are stringified |
add_file(name, filename, content, content_type) | adds a file part; content is bytes or a string, content_type defaults to application/octet-stream |
content_type() | the Content-Type header value, boundary included |
boundary() | the boundary string |
build() | the assembled body, as bytes |
length() | how many parts have been added |
Several files under one field name is just several calls — that is
what an <input type="file" multiple> sends:
for path in paths {
form.add_file('attachments', os.base_name(path), file(path, 'rb').read())
}
A filename that is not plain ASCII is sent twice: once as a
transliterated filename, and once as an RFC 5987 filename*, which
is what lets the real name survive a recipient that only understands
one of the two.
If you need the body separately — to sign it, to log its size, to send it somewhere this client is not going — build it yourself:
var body = form.build()
api.post('/photos', nil, {
body,
content_type: form.content_type(),
})
Note
build()assembles the whole body in memory. For an upload large enough that this matters, send the file as a raw body with its own content type instead —multipart/form-dataonly earns its overhead when there are fields alongside the file.
The boundary is generated from the platform’s secure random source, not from a timestamp or a counter — a boundary a peer can predict is a boundary a peer can write into a field value to forge extra parts.
Reading a Response
var response = api.get('/items/1')
response.status # 200
response.reason_phrase() # 'OK'
response.version # '1.1' or '2'
response.headers.get('etag')
response.as_text() # the body, decoded as UTF-8
response.as_dict() # the body, parsed as JSON
response.as_bytes() # the body, untouched
response.is_ok() # 2xx
response.is_redirect() # 3xx
response.is_error() # 4xx or 5xx
raise_for_status() turns a 4xx or 5xx into a StatusError and
returns the response otherwise, so it can be used inline. The response
stays reachable on the raised error, which matters because that is
where an API usually explains what went wrong:
catch {
var data = api.get('/items/1').raise_for_status().as_dict()
} as error {
echo error.response.status
echo error.response.as_text()
}
A compressed response body is decompressed automatically — gzip,
deflate, brotli and zstd — and the Content-Encoding header is
removed once it has been, so nothing downstream tries to decode it a
second time.
Query Parameters and Headers
api.get('/search', {
query: { q: 'zuri', page: 2, tag: ['new', 'featured'] },
headers: { 'Accept-Language': 'en-GB' },
})
A list value repeats the parameter, which is how a query string carries more than one value under one name. Per-request headers are merged over the client’s own.
Authentication
api.get('/me', { auth: ['bearer', token] })
api.get('/me', { auth: ['basic', 'ada', 'lovelace'] })
Credentials passed this way are scoped to the origin you addressed. If
the response is a redirect to a different scheme, host or port, they
are dropped rather than followed — a redirect to a host of someone
else’s choosing is otherwise a way to collect whatever Authorization
header was in flight.
Redirects
Redirects are followed by default, up to max_redirects (10), and the
number followed is on the response:
var final = http.get('https://example.com/old')
final.redirects # how many were followed
final.responder # the URL that finally answered
A 303, and in practice a 301 or 302, becomes a GET with no
body when followed — that is what every browser and every other client
does, and a server that meant otherwise should have sent 307 or
308, both of which preserve the method and body here.
Turn following off per request or per client:
http.get(url, { follow_redirects: false })
head() does not follow redirects by default, since the point of a
HEAD is usually to inspect the very response a redirect would hide.
Cookies
A client with a cookie jar carries cookies between requests, so a login and the requests after it behave the way a browser would:
var session = http.client('https://example.com')
session.enable_cookies()
session.post('/login', nil, { form: { user: 'ada', password: secret } })
session.get('/dashboard') # sends the session cookie
The jar applies RFC 6265’s matching rules: a cookie without a Domain
is host-only, one with a Domain reaches subdomains, a path prefix
matches below itself, and a Secure cookie is never sent over
cleartext. A server trying to set a cookie for a domain it does not
control is refused.
Timeouts and Retries
var api = http.client('https://api.example.com', {
connect_timeout: 5000,
read_timeout: 10000,
write_timeout: 10000,
})
api.get('/slow', { timeout: 30000 }) # this one request
All timeouts are milliseconds. max_retries retries a failed request,
but only when the method is idempotent — replaying a POST can mean
charging a card twice, so GET, HEAD, PUT, DELETE, OPTIONS
and TRACE are retried and nothing else is.
A pooled connection the far end closed while it was idle fails on the
next write and looks exactly like a network failure. That specific
case is retried once for an idempotent request even at the default
max_retries of zero, because otherwise every connection a server
reaps would surface as a spurious error.
Streaming a Response
A large response does not have to be held in memory:
var response = api.get('/exports/large.csv', { stream: true })
var out = file('large.csv', 'wb')
while true {
var chunk = response.body_reader.read(65536)
if chunk.length() == 0 {
break
}
out.write(chunk)
}
out.close()
api.finish(response)
finish() hands the connection back to the pool once you are done
with the body. Until then the connection is yours, since there is no
way for the client to know when you have finished reading.
TLS and Certificates
Certificates are verified against the platform’s trust store by default. To talk to a service with an internal or self-signed certificate, trust its authority:
api.add_ca(file('/etc/ssl/internal-ca.pem').read())
There is also verify = false, which turns verification off
completely. It exists for local development, and it makes the
connection encrypted but unauthenticated — which is to say, trivially
interceptable. add_ca() is the right answer everywhere else.
Proxies
A client goes through the proxy the environment names, the way command
line tools do: HTTPS_PROXY for https addresses, HTTP_PROXY for
http ones, and ALL_PROXY for either, each in upper or lower case.
NO_PROXY lists the hosts reached directly, separated by commas: a
name covers every name under it, an entry with a port covers only that
port, and * covers everything. This machine’s own loopback addresses
never go through a proxy.
A client of your own can name its proxy instead, or ignore the environment altogether:
var api = http.client('https://api.example.com', {
proxy: 'http://ada:secret@proxy.internal:3128',
})
var direct = http.client('https://api.example.com', { trust_env: false })
An https request is tunnelled through the proxy with CONNECT, so the
proxy carries the encrypted bytes and never sees what they say, and the
certificate checked is the destination’s own. An http request is
handed to the proxy whole. Credentials in the proxy address are sent to
the proxy alone, as Proxy-Authorization. The proxy itself is reached
over plain HTTP; an https:// proxy address is refused.
Serving Requests
Routing
A route pattern is made of literal segments, :name parameters, and
an optional trailing catch-all:
| Pattern | Matches | Captures |
|---|---|---|
/users | /users | |
/users/:id | /users/42 | id = '42' |
/users/:id/posts/:post | /users/4/posts/7 | id, post |
/files/ + *path | /files/css/site.css | path = 'css/site.css' |
server.get('/users/:id', @(request, response) {
response.json({ id: request.param('id') })
})
A literal segment always beats a parameter, and a parameter always
beats a catch-all, so /users/new and /users/:id can both exist and
the specific one wins — regardless of which was registered first.
Matching runs over a trie, so a router with a thousand routes costs
the same per request as one with ten.
get(), post(), put(), patch(), delete(), head() and
options() register for one method. any() registers GET, POST,
PUT, PATCH, DELETE and OPTIONS at once. handle() takes the
method as an argument, which is how a method with no function of its
own — PROPFIND, or anything a service has invented — gets a route:
server.handle('PROPFIND', '/files/' + '*path', list_properties)
Three things are answered without a handler:
HEADfalls back to theGETroute for the same path, and the body is dropped on the way out.OPTIONSanswers204with anAllowheader listing what the path actually accepts.- A path that exists for other methods answers
405, again withAllow— rather than a404, which would be a lie.
Name a route to build URLs from it later:
server.get('/users/:id', show_user, 'user.show')
server.routes().url_for('user.show', { id: 42 }) # '/users/42'
The Request Object
server.post('/items', @(request, response) {
request.method # 'POST'
request.path # decoded and normalised
request.target # exactly as it arrived on the wire
request.version # '1.1' or '2'
request.secure # whether it came over TLS
request.param('id') # a route parameter
request.query_param('page', '1') # a query parameter, with a default
request.header('accept') # a header field
request.cookie('session') # a cookie
request.bearer_token() # the token from an Authorization header
request.text() # the body as text
request.json_body() # the body parsed as JSON
request.form() # urlencoded or multipart fields
request.file('avatar') # an uploaded file
})
path is percent-decoded and then normalised, in that order — which
is the only order that works, since %2e%2e%2f is ../ written to
survive a naive check. Route on path; target is the raw form and
matching on it is how directory traversal gets through.
An uploaded file is an UploadedFile:
var upload = request.file('avatar')
upload.filename # what the client claimed
upload.safe_name() # that, reduced to one path-safe segment
upload.content_type # what the client declared
upload.size()
upload.save_to('/var/uploads/' + upload.safe_name())
filename and content_type are both attacker-controlled and say
nothing about what the bytes are. safe_name() drops any directory
component — including a Windows one — and reduces the rest to letters,
digits, ., - and _.
Forms and File Uploads
A submitted form arrives as a request body, in one of two encodings,
and both are read through the same methods. Which one a browser sends
is decided by the enctype on the <form>: the default
application/x-www-form-urlencoded for a form of plain fields, and
multipart/form-data for one that carries a file.
server.post('/signup', @(request, response) {
var email = request.form_field('email', '')
var password = request.form_field('password', '')
if email.is_empty() {
response.json({ error: 'email is required' }, 422)
return
}
create_account(email, password)
response.redirect('/welcome', 303)
})
| Method | Returns |
|---|---|
form() | every field as name -> value, keeping the first of a repeated name |
form_all() | every field as name -> [value, ...] |
form_field(name, fallback) | one field’s first value |
files() | uploaded files as name -> [UploadedFile, ...] |
file(name) | the first file uploaded under name, or nil |
The body is parsed on first use and cached, so reading form() in a
middleware and again in the handler costs one parse. A body that is
neither encoding gives an empty dictionary rather than raising — use
request.json_body() for a JSON body and request.text() for
anything else.
A field a form can repeat — a set of checkboxes, a multi-select —
needs form_all(), since form() keeps only the first value:
var tags = request.form_all().get('tags', [])
Note A redirect after a successful
POSTshould be303 See Other, as above. It is what stops a browser re-submitting the form when the user reloads the resulting page.
Receiving an upload
<form method="post" action="/avatar" enctype="multipart/form-data">
<input type="text" name="caption">
<input type="file" name="avatar">
<button>Upload</button>
</form>
import os
server.post('/avatar', @(request, response) {
var upload = request.file('avatar')
if upload == nil {
response.json({ error: 'no file was submitted' }, 400)
return
}
if upload.size() > 2 * 1024 * 1024 {
response.json({ error: 'the image must be under 2 MB' }, 413)
return
}
var destination = os.join_paths('./uploads', upload.safe_name())
upload.save_to(destination)
response.json({
saved: upload.safe_name(),
caption: request.form_field('caption', ''),
bytes: upload.size(),
})
})
An UploadedFile carries:
| Member | Is |
|---|---|
filename | the name the client claimed, exactly as sent |
safe_name() | that name reduced to one path-safe segment |
content_type | the type the client declared, or application/octet-stream |
content | the bytes |
size() | how many of them |
name | the form field it arrived under |
headers | every header on that part of the body |
save_to(path) | writes the content, and returns the byte count |
to_text() | the content decoded as UTF-8 |
An <input type="file" multiple> sends several parts under one name,
which is what files() returns a list for:
for upload in request.files().get('attachments', []) {
upload.save_to(os.join_paths('./uploads', upload.safe_name()))
}
Two things a handler must not trust
The filename. It is attacker-controlled, and it routinely contains
a full local path from a Windows client, .. segments from a hostile
one, or a NUL byte meant to truncate a later check. Never join it to a
path directly:
os.join_paths('./uploads', upload.filename) # no
os.join_paths('./uploads', upload.safe_name()) # yes
safe_name() drops every directory component — both separators — and
reduces the rest to letters, digits, ., - and _, returning
'unnamed' when nothing usable is left. Better still, name the file
yourself and keep the client’s name as a label:
var stored = '${uuid.v4()}.jpg'
upload.save_to(os.join_paths('./uploads', stored))
record_upload(stored, upload.filename)
The content type. content_type is whatever the client wrote in
the part header and says nothing about what the bytes are. A file
claiming image/png may be anything at all. When the difference
matters, look at the bytes — mime.detect_from_header() reads a real
file’s leading bytes, so sniffing an upload means writing it somewhere
first:
import mime
import os
var staged = os.join_paths(os.temp_dir(), uuid.v4())
upload.save_to(staged)
var actual = mime.detect_from_header(file(staged, 'rb'))
if actual != 'image/png' and actual != 'image/jpeg' {
file(staged).delete()
response.json({ error: 'that is not an image' }, 415)
return
}
os.rename(staged, os.join_paths('./uploads', stored_name))
Size limits
The body is read into memory, bounded by the server’s
max_body_size — 10 MiB by default, which is deliberately small:
server.max_body_size = 50 * 1024 * 1024 # accept uploads up to 50 MB
A request that announces a body over the limit is refused with 413 Content Too Large before a byte of it is read, which is the whole
point of Content-Length. One that lies about its length is cut off
at the limit and also refused. Either way the connection then closes,
since a body that was only partly read leaves nothing safe to parse
after it.
A client asking permission first with Expect: 100-continue gets its
answer before it sends anything — the server checks the announced
length against the limit and either says 100 Continue or refuses
outright, so a rejected upload costs one round trip rather than a
whole transfer.
For an upload too large to want in memory at all, take it as a raw
body and stream it rather than as a form field. multipart/form-data
earns its overhead only when there are fields alongside the file.
Request Validation
A request validates itself against a validate schema:
import http
import validate
var create_user = validate.schema({
name: validate.required().string().max_length(100),
email: validate.required().string().email(),
age: validate.required().integer().gte(18).lte(120),
})
server.post('/users', @(request, response) {
catch {
var data = request.validate(create_user)
response.json(create_account(data), 201)
} as error {
response.json({ errors: create_user.group_errors(error.errors) }, 422)
}
})
validate() returns the input it checked, so the happy path is one
line and the data you go on to use is the data that was validated.
Failure raises the schema’s own validate.ValidationError, carrying
an errors list of { field, message }; group_errors() turns that
into a dictionary keyed by field, which is the shape most front ends
want:
{
"errors": {
"email": ["The email field must be a valid email address."],
"age": ["The age field must be greater than or equal to 18."]
}
}
To branch rather than catch, validate the input yourself — there is no separate API for it:
var result = create_user.check(request.input())
if !result.valid {
response.json({ errors: result.errors }, 422)
return
}
The rules themselves — and there are around eighty of them, including
cross-field ones like confirmed() and required_if() — belong to
the validate module rather than to this one.
What gets validated
request.input() is the dictionary validate() checks. Three sources
are merged, each overriding the one before it:
- route parameters, from the pattern that matched
- query string parameters
- the body — a JSON object’s keys, or the submitted form fields
So one schema covers POST /users with a JSON body, GET /users?…
with a query string, and /users/:id with a route parameter, without
the handler caring which arrived.
Take one source on its own by naming it:
request.validate(schema, 'body') # only the body
request.validate(schema, 'query') # only the query string
request.validate(schema, 'params') # only the route parameters
Two things are deliberately left out of the merge:
- Uploaded files. Nothing a schema can say about a file is
expressible as a rule over its bytes; reach them with
request.file()and check them as Forms and File Uploads describes. - A JSON body that is not an object. An array or a bare string has
no names to merge, so it contributes nothing; read it with
request.json_body().
A body that fails to parse as JSON also contributes nothing rather
than raising, which leaves the schema’s own required rules to report
what is missing. That is a better answer to a client than a parser
message.
Values from the wire are strings
A query string and a urlencoded form carry text and nothing else.
?age=36 is the string '36', not the number 36, and that changes
which rules hold:
| Rule | On '36' | Because |
|---|---|---|
integer(), numeric(), gt(), gte(), lt(), lte() | reads it as 36 | these coerce |
size(), min(), max(), between() | reads it as 2 | these measure size, which for a string is its character count |
That is validate’s documented behaviour, not an accident of this
module: min(8) on a password means eight characters. It only
surprises when a schema written against a JSON body is later pointed
at a query string.
Write a schema that has to serve both with the value rules:
age: validate.required().integer().gte(18).lte(120) # both
age: validate.required().integer().between(18, 120) # JSON bodies only
input() does not coerce anything on your behalf. It would have to
guess, and a postcode of '01234' or a version of '1.0' silently
becoming a number is worse than the rule you have to pick deliberately.
Repeated fields
A name may legally repeat in a query string or a form, so those
sources arrive as name -> [values]. input() flattens a name
carrying exactly one value to that value, and leaves a name carrying
several as a list:
?tag=a -> { tag: 'a' }
?tag=a&tag=b -> { tag: ['a', 'b'] }
That is what lets a scalar rule see a scalar. A field that must
always be a list, however many values arrived, is better read
through form_all() or request.query directly and validated with
validate’s .* wildcard.
Note The
httpmodule does not importvalidate.validate()takes any object with acheck_or_raise()method and calls it, so a server that validates nothing never pays to load a schema engine — and a schema of your own, or from somewhere else, works just as well.
The Response Object
response.text('plain') # text/plain
response.html('<h1>hi</h1>') # text/html
response.json({ ok: true }) # application/json
response.xml('<doc/>') # application/xml
response.write('more') # append, no content type
response.file('/var/www/report.pdf') # streamed from disk
response.download('/var/www/report.pdf') # ...as an attachment
response.render('pages/home', { user }) # a Wire template
response.redirect('/elsewhere') # 302 by default
response.redirect('/elsewhere', 308) # method-preserving
response.status = 201
response.header('X-Thing', 'value')
response.content_type('text/csv')
response.cache_for(3600)
response.no_cache()
Every writer sets a sensible Content-Type alongside the body, and
each returns the response, so they chain.
file() streams from disk rather than reading the file into memory,
which is what makes serving something larger than you would like to
hold in memory a one-liner.
Middleware
A middleware takes (request, response, next) and decides whether the
rest of the chain runs:
server.use(@(request, response, next) {
var started = time()
next()
echo '${request.method} ${request.path} ${response.status} ' +
'${(time() - started) * 1000}ms'
})
Not calling next() is how a middleware short-circuits, which is
exactly what an authentication or rate-limiting layer wants:
server.use(@(request, response, next) {
if request.header('x-api-key') != expected {
response.json({ error: 'unauthorized' }, 401)
return
}
next()
})
Middleware run in the order they were added, outermost first, so one added first wraps everything added after it.
Code after next() runs only when the rest of the chain returned. A
handler that raises never comes back there, and the 500 the visitor
is sent is decided later, by the server’s error handling. A middleware
that reports on the response the client actually receives registers
with response.on_finish() instead, which runs once the response is
final, failures included:
server.use(@(request, response, next) {
var started = time()
response.on_finish(@{
echo '${request.method} ${request.path} ${response.status} ' +
'${(time() - started) * 1000}ms'
})
next()
})
Callbacks run once each, in the order they were registered, and one
that raises is skipped rather than costing the client its response.
middleware.logger() is built this way.
Anything a middleware wants to hand to the handler goes on
request.context, which is a plain dictionary that exists for exactly
that:
server.use(@(request, response, next) {
request.context['started_at'] = time()
next()
})
A middleware that wants to act on the response calls next() first
and then works on what came back:
server.use(@(request, response, next) {
next()
response.header('X-Served-By', hostname)
})
The ones most services need are already written — see Built-in Middleware.
Errors
A handler that raises becomes a 500, and the exception message never
reaches the client — an exception message routinely carries a file
path, a query, or a fragment of the data being processed, and none of
that belongs in a reply to whoever triggered it.
server.on_error(@(error, connection) {
log.error('${error.message}')
})
server.error_handler(@(request, response, error) {
response.status = 500
response.render('errors/500')
})
server.not_found(@(request, response) {
response.status = 404
response.render('errors/404')
})
on_error() listeners see everything, including connection-level
failures that never reached a handler. error_handler() produces the
response body for a failed request; without one, the server sends a
bare 500.
Static Files
server.serve_files('/static', './public', {
cache_age: 86400,
precompressed: true,
})
That single call brings with it the things a static file needs to actually behave well in front of a browser or a CDN:
ETagandLast-Modified, and304 Not Modifiedfor a conditional request that still holdsRangerequests, answered with206and aContent-Range, and416for a range that falls outside the fileIf-Match,If-Unmodified-Since,If-None-Match,If-Modified-SinceandIf-Range, evaluated in the order RFC 9110 lays downContent-Typefrom the file extensionindex.htmlfor a request that names a directory- with
precompressed, a.bror.gzsibling served in place of the original when the client accepts that coding
Every request path is percent-decoded, normalised, joined to the root,
and then checked to still be inside the root — all three, because
doing any two of them is not enough. Dotfiles are not served at all:
.env, .git and .htpasswd all end up in a web root at some point
in a project’s life, and none of them should ever go out.
For a single-page application, fallback serves the shell for any
path that names no file, which is what makes client-side routes work
on reload:
server.serve_files('/', './dist', { fallback: 'index.html' })
Compression
Responses are compressed on the way out, and it is on by default. A response is compressed when all of the following hold:
server.compressionis on (it is),- the body is at least
server.compression_min_sizebytes (1024), - its media type is one that benefits — anything
text/, plus JSON, JavaScript, XML, SVG, WebAssembly and the+json/+xmltypes, - the client’s
Accept-Encodingacceptsbrorgzip, - nothing already set a
Content-Encoding, - and the compressed body actually came out smaller.
Brotli is preferred over gzip when the client accepts both. Vary: Accept-Encoding is added either way, so a shared cache in front
cannot serve a compressed body to a client that asked for none.
Turning it off, or moving the threshold:
server.compression = false # off entirely
server.compression_min_size = 4096 # only bodies over 4 KiB
The list is an allowlist rather than a denylist on purpose. A JPEG, an MP4 or a zip is already compressed; running it through brotli spends CPU to make it slightly larger.
Note Responses built with
response.file()— which includes everythingserve_files()serves — are not compressed on the fly. They are streamed from disk rather than held in memory, and compressing them per request would mean reading them into memory to do it. Compress those ahead of time and let the static handler pick the compressed file up:server.serve_files('/static', './public', { precompressed: true })With that on, a request for
site.cssfrom a client that accepts brotli is answered withsite.css.brif it exists, and withsite.css.gzif that exists and gzip is accepted — at no CPU cost per request, and with the correctContent-EncodingandVary.
On the client side there is nothing to configure: Accept-Encoding: gzip, br, deflate, zstd is sent by default and a compressed response
body is decoded before you see it.
Streaming a Response Body
A response body can come from a callback instead of memory:
server.get('/export.csv', @(request, response) {
response.content_type('text/csv')
response.stream(@(writer) {
writer.write('id,name\n')
for row in rows {
writer.write('${row.id},${row.name}\n')
writer.flush()
}
})
})
Unless a Content-Length was set beforehand, the body is framed with
chunked transfer encoding on HTTP/1.1 and as an ordinary DATA stream
on HTTP/2 — the handler does not have to know which.
Setting Cookies
response.set_cookie('theme', 'dark', {
max_age: 86400,
secure: true,
same_site: 'Strict',
})
response.clear_cookie('theme')
http_only defaults to true and same_site to 'Lax', which are
what a cookie carrying anything sensitive should have; pass them
explicitly to opt out. A cookie whose name carries the __Secure- or
__Host- prefix has that prefix’s rules applied for it, rather than
being sent in a form the browser will silently refuse to store.
For state that belongs to a visitor rather than to the browser, use a session, which keeps the state on the server and puts only an identifier in the cookie. Sessions covers it.
Content Negotiation
import http.negotiate
var type = negotiate.best_match(
request.header('accept'),
['application/json', 'text/html']
)
best_match() picks what the client would most like out of what you
can produce, honouring quality weights, with your own order deciding
ties. request.accepts('text/html') and request.wants_json() cover
the common cases.
Built-in Middleware
http.middleware holds the middleware nearly every public service
ends up wanting. Each function returns a middleware, so it is called
at the point you register it:
import http
import http.middleware
var server = http.server(3000)
server.use(middleware.request_id())
server.use(middleware.logger())
server.use(middleware.security_headers())
They are ordinary middleware with no special standing — read any of them as a worked example of writing your own.
None of them raises when it is built. That matters for the three that
take a verifier — basic_auth(),
bearer_auth() and jwt_auth() — where
no verifier means the middleware is disabled: it passes every
request straight through rather than refusing one, and rather than
failing at registration.
# Authentication, off. Everything else in the chain is untouched.
server.use(middleware.jwt_auth(nil))
That is the right shape for a middleware, and it is deliberately useful: blanking a verifier switches its middleware off without unpicking the chain around it, which is what you want when bisecting a request that is failing somewhere in a stack of them.
Disabling is opt-in and never a fallback: a verifier that is present and refuses a request still refuses it.
logger()
Writes one line per request, after the response is finished.
server.use(middleware.logger())
127.0.0.1 - [Wed, 09 Sep 2026 20:09:17 GMT] "GET /static/big.txt HTTP/1.1" 200 320 3.07ms
That is the Common Log Format with the response time added, which
every log analyser already understands. The size comes from
Content-Length when the body is file-backed or streamed, so a static
file logs its real size rather than zero.
| Option | Default | Meaning |
|---|---|---|
sink | echo | a function taking the finished line |
format | the format above | a function (request, response, milliseconds) returning the line |
trust_proxy | false | log the forwarded client address rather than the peer |
server.use(middleware.logger({
sink: @(line) { log_file.puts(line + '\n') },
trust_proxy: true,
}))
Structured logging is a format that returns JSON:
server.use(middleware.logger({
format: @(request, response, elapsed) {
return json.encode({
method: request.method,
path: request.path,
status: response.status,
ms: elapsed,
id: request.context.get('request_id', nil),
})
},
}))
A request that raises is logged too, with the status the client was
sent once the server had answered the failure: the 500 of the default
handling, or whatever the error handler chose.
request_id()
Gives every request an identifier, so one request can be followed across a log file and across services.
server.use(middleware.request_id())
The identifier lands on request.context['request_id'] and in the
response’s X-Request-Id. One supplied by the client is echoed rather
than replaced, which is what makes a trace continue across a hop.
The single optional argument is the header name:
server.use(middleware.request_id('X-Correlation-Id'))
An identifier that came from outside is written into your logs, so it is capped at 64 characters and stripped of anything that is not plainly printable before it is trusted that far.
security_headers()
Adds the response headers a browser acts on to harden a page.
server.use(middleware.security_headers())
| Option | Default | Header |
|---|---|---|
content_type_options | 'nosniff' | X-Content-Type-Options — stops the browser second-guessing a declared content type |
frame_options | 'DENY' | X-Frame-Options — refuses to be framed, which is what clickjacking needs |
referrer_policy | 'strict-origin-when-cross-origin' | Referrer-Policy — keeps paths and queries out of outbound referrers |
hsts_max_age | 31536000 | Strict-Transport-Security |
hsts_subdomains | false | adds includeSubDomains |
content_security_policy | not set | Content-Security-Policy |
server.use(middleware.security_headers({
content_security_policy: "default-src 'self'; img-src 'self' data:",
hsts_subdomains: true,
frame_options: 'SAMEORIGIN',
}))
Pass nil for any of them to leave that header off entirely.
Content-Security-Policy has no default because a wrong one breaks
the page and a permissive one is theatre; it depends on what your
application actually loads.
Strict-Transport-Security is only sent over HTTPS. Over cleartext it
is at best ignored, and at worst a way to lock a client out of a host
that has no certificate.
Every header is set with set_default, so a handler that set its own
keeps it.
cors()
Answers CORS preflights and adds the headers a browser needs before it will let script read a response from another origin.
server.use(middleware.cors({
origins: ['https://app.example.com'],
credentials: true,
}))
| Option | Default | Meaning |
|---|---|---|
origins | ['*'] | origins allowed, matched exactly; '*' allows any |
methods | the seven usual ones | what a preflight is told is allowed |
headers | reflect what was asked | request headers script may set |
expose | none | response headers script may read |
credentials | false | whether cookies and Authorization may be sent |
max_age | 86400 | how long a preflight may be cached, in seconds |
A request with no Origin is not a cross-origin request and passes
straight through. An Origin that is not on the list gets no CORS
headers at all — the browser then refuses to hand the response to
script, which is the correct outcome; answering with an error instead
would leak whether the resource exists.
Note A wildcard and
credentialscannot be combined — every browser rejects that pairing. Whencredentialsis set, the requesting origin is echoed back instead, which makes theoriginslist the only thing standing between an attacker’s page and an authenticated response. Name the origins explicitly whenever credentials are in play.
Vary: Origin is added whenever the answer depends on the request’s
origin, so a shared cache cannot serve one origin’s response to
another.
basic_auth()
Requires HTTP Basic authentication.
server.use(middleware.basic_auth(@(user, password) {
return user == 'admin' and http.util.secure_equals(password, secret)
}, 'Admin area'))
The first argument is called with the username and password and
returns whether they are acceptable. The second is the realm named in
the challenge, and defaults to 'Restricted'.
On success the username is put on request.context['user']. On
failure the response is 401 with a WWW-Authenticate header, which
is what makes a browser show its credentials prompt.
Passing nil in place of the verifier disables the middleware
entirely — see the note above.
Compare secrets with http.util.secure_equals() rather than ==: it
compares in constant time, so a wrong guess and a nearly-right one
take the same time and the comparison does not hand over the secret
one character at a time.
To guard part of a site rather than all of it, wrap it:
var guard = middleware.basic_auth(check_credentials)
server.use(@(request, response, next) {
if request.path.starts_with('/admin') {
guard(request, response, next)
return
}
next()
})
Note that the next() a guard calls on success continues the whole
remaining chain, not just the handler. Two authentication middleware
registered globally therefore both apply to every request; scope each
of them by prefix, as above, when they are meant for different parts
of the site.
bearer_auth()
Requires a bearer token — the usual shape for an API.
server.use(middleware.bearer_auth(@(token) {
var claims = verify_jwt(token)
return claims == nil ? false : claims.subject
}))
The verifier is called with the token and returns either something
falsy to refuse it, or a value to attach — that value lands on
request.context['user'], so returning the subject, the claims, or a
whole user record all work.
Failure is 401 with a WWW-Authenticate: Bearer challenge and a
JSON body. The realm is the optional second argument and defaults to
'api'. A nil verifier disables the middleware, as with
basic_auth().
middleware.parse_bearer(header) and middleware.parse_basic(header)
are exported separately for code that needs to read an Authorization
header without installing a middleware at all, and
request.bearer_token() reads the token straight off a request.
For a JSON Web Token specifically, jwt_auth() does the
verification, the claim handling and the RFC 6750 challenges rather
than leaving them to the verifier you write here.
jwt_auth()
Requires a valid JSON Web Token, verified by the jwt module.
import http.middleware
import jwt
server.use(middleware.jwt_auth(
jwt.Verifier(secret, { algorithms: ['HS256'], audience: 'api' })
))
The first argument is either a jwt.Verifier or any function taking
the token and returning its claims. The function form is what covers a
key set resolved by kid, or anything else the jwt module can do
that a fixed verifier cannot:
server.use(middleware.jwt_auth(@(token) {
return jwt.verify_with_jwks(token, keys, { audience: 'api' })
}))
| Option | Default | Meaning |
|---|---|---|
realm | 'api' | the realm named in the challenge |
optional | false | attach the claims when a valid token is present, but do not refuse a request without one |
scopes | none | a list of scopes every token must carry |
On success the claims land on request.context['claims'], and the
sub claim — the usual place an issuer puts the account a token
speaks for — on request.context['user']:
server.get('/me', @(request, response) {
response.json({
account: request.context.get('user', nil),
issued_at: request.context.get('claims', {}).get('iat', nil),
})
})
Failures follow RFC 6750 §3:
| Situation | Answer |
|---|---|
| no token at all | 401, WWW-Authenticate: Bearer realm="api" |
| the token does not verify | 401, with error="invalid_token" |
| valid, but missing a required scope | 403, with error="insufficient_scope" and the scope that was needed |
The distinction in that last row is the point of scopes: a token
that failed to authenticate and a token that authenticated but is not
allowed to do this are different problems, and answering both with
401 tells a client to go and get a new token when a new token will
not help.
Note A challenge never says which check a token failed. An expired token and a forged one get exactly the same
error_description, because which one it was is useful to your logs and useful to an attacker, and to nobody else. If you need the reason, log it from the verifier you passed in.
optional is for a route that behaves differently when it knows who
is asking without requiring it — a public page that shows an edit
button to its author:
server.use(middleware.jwt_auth(verifier, { optional: true }))
server.get('/posts/:id', @(request, response) {
var viewer = request.context.get('user', nil)
response.json(render_post(request.param('id'), viewer))
})
A token that is present but invalid is still refused under
optional. Ignoring a bad token would let a client tamper with one
and get the anonymous view rather than an error, which hides exactly
the problem worth surfacing.
A nil verifier disables the middleware: every request passes
through unauthenticated, and request.context['claims'] is simply
never set. Nothing here raises when it is built, so registering it
never needs a catch around it.
var verifier = production ? jwt.Verifier(secret, options) : nil
server.use(middleware.jwt_auth(verifier))
Like HttpRequest.validate(), this does not import the jwt module —
the verifier is built by the caller, which keeps the token format the
application’s business and means a server that authenticates nothing
never pays to load it.
rate_limit()
Limits how many requests one client may make in a window of time.
server.use(middleware.rate_limit({ limit: 100, window: 60 }))
| Option | Default | Meaning |
|---|---|---|
limit | 60 | requests allowed per window |
window | 60 | the window, in seconds |
key | the client address | a function of the request returning the bucket key |
trust_proxy | false | derive the address from forwarding headers |
Every response carries the current state, so a well-behaved client can back off before it is refused:
RateLimit-Limit: 100
RateLimit-Remaining: 87
RateLimit-Reset: 41
Over the limit is 429 Too Many Requests with a Retry-After.
Rate-limit per API key rather than per address by supplying a key:
server.use(middleware.rate_limit({
limit: 1000,
window: 3600,
key: @(request) {
return request.header('x-api-key', nil) or request.client_ip() or 'anonymous'
},
}))
Note The counter lives in memory, so it is per worker: running with
workers: 4makes the effective limit four timeslimit. That is a deliberate trade — a shared counter would need shared state, and this is meant to blunt a runaway client rather than to meter billing. Dividelimitby the worker count if the exact number matters, or keep the count somewhere both workers can see.
etag()
Computes a weak ETag over a finished response body and answers 304 Not Modified when the client already has that version.
server.use(middleware.etag())
The single optional argument is the smallest body worth tagging, in
bytes; it defaults to 128, below which the validator costs more than
the body it would save.
A handler that already set its own ETag is left alone — it knows
something about the resource that hashing the bytes does not — and so
are file-backed responses, which serve_files() already tags from the
file’s size and modification time.
Because it works on the finished body, it pairs with anything: a rendered template, a JSON document, a generated report.
force_https()
Redirects every request that arrived over cleartext to the same URL over HTTPS.
server.use(middleware.force_https())
| Option | Default | Meaning |
|---|---|---|
status | 308 | the redirect status; 308 preserves the method and body |
port | the default | the HTTPS port, when it is not 443 |
This belongs on the cleartext listener, which usually exists only to perform this redirect:
# port 80: redirect and nothing else
var redirector = http.server(80, '0.0.0.0')
redirector.use(middleware.force_https())
Pair it with security_headers()’s Strict-Transport-Security on the
TLS listener, so that after the first visit the browser stops making
the cleartext request at all.
Ordering
Middleware run outermost first, so the order they are registered in is the order they wrap the request. A workable default:
server.use(middleware.request_id()) # so everything after can log it
server.use(middleware.logger()) # so it sees the final status
server.use(middleware.force_https()) # before any work is done
server.use(middleware.security_headers())
server.use(middleware.cors({ origins: allowed }))
server.use(middleware.rate_limit({ limit: 100 }))
server.use(middleware.jwt_auth(verifier)) # after the cheap refusals
server.use(middleware.etag()) # innermost: it needs the finished body
The reasoning behind each position: identifiers before anything that might log, logging outside everything so it records what actually happened, cheap refusals before expensive ones, and anything that inspects the response body innermost, where the body exists.
Sessions
A session is state that belongs to one visitor, kept on the server and found again by a cookie the browser sends back. Only the identifier travels, so a visitor can neither read what the session holds nor change it, and the cookie is worth nothing to anyone who cannot present the exact value that was issued.
import http
var server = http.server(3000)
server.use(http.session.session())
server.get('/', @(request, response) {
var seen = request.session().get('seen', 0) + 1
request.session().set('seen', seen)
response.text('visit ${seen}')
})
server.listen()
session() is middleware. Register it once, above anything that reads
a session, and every handler below it reaches its own through
request.session().
Nothing Happens Until Something Uses It
A request that never touches its session costs nothing: no read from
the store, no write, and no Set-Cookie. The record is created the
first time something is written to the session, which is what keeps a
crawler working through a public site from filling the store with
empty sessions.
A request that only reads an existing session writes nothing back
either, beyond moving the idle expiry along at most once every
touch_interval seconds.
Signing In
server.post('/login', @(request, response) {
var account = authenticate(request.form())
if account == nil {
response.status = 401
response.html(render_login('Those details do not match.'))
return
}
request.session().regenerate()
request.session().set('account', account.id)
response.redirect('/')
})
regenerate() is the line to get right. It gives the session a new
identifier and destroys the record the old one named, keeping
everything the session holds.
Without it, an attacker who can set a cookie in the victim’s browser beforehand — through a stray subdomain, an open redirect, a shared machine — knows the identifier the victim will be signed in under, and can simply use it afterwards. That is session fixation, and a new identifier is the whole of the defence. Call it whenever what the session means changes: signing in, elevating to administrator, a step-up authentication.
Signing out is destroy(), which removes the record and has the
response expire the cookie:
server.post('/logout', @(request, response) {
request.session().destroy()
response.redirect('/')
})
clear() is the other one: it empties the session without ending it,
keeping the identifier and the cookie.
Flash Messages
A handler that does the work and redirects cannot render the message saying what happened. A flash carries it to the page that can:
server.post('/posts', @(request, response) {
create_post(request.form())
request.session().flash('notice', 'Your post is up.')
response.redirect('/posts')
})
server.get('/posts', @(request, response) {
response.html(render(posts(), request.session().take_flash('notice')))
})
A flash is spent by the next request that touches the session at all, whether or not that request asks for this one. A page that looked at the session and did not read the message does not leave it for the page after; a request that never touched its session — a static file, an image — leaves it waiting.
Where Sessions Are Kept
Four things can hold a session, and swapping between them changes one line.
FileStore, one file per session. This is the default, because it
needs no setup and is shared between the workers http.serve()
starts. Given no directory it uses a private subdirectory of the
platform’s temporary directory, created 0700:
server.use(http.session.session())
That directory is cleared on whatever schedule the platform keeps, and on most of them at every reboot, so name your own for anything that has to outlive the host:
server.use(http.session.session({
store: http.session.FileStore('/var/lib/app/sessions'),
}))
A session file is a bearer credential in the same way the cookie is.
The directory is created 0700 and each file 0600, and a directory
that every user on the machine can reach is refused rather than used —
which is why the temporary directory itself is never the default.
Pass strict_permissions: false to accept one anyway.
SqlStore, one row per session. It lives in its own import, so a
program using the default store never loads the sql module:
import http.session.sql { SqlStore }
import sql
var store = SqlStore(sql.pool('postgres://localhost/app'))
store.migrate()
server.use(http.session.session({ store }))
migrate() creates the table and its index if they are not there, and
is safe to call on every start. It takes a sql.Connection or a
sql.Pool; a server wants the pool.
MemoryStore, for a test. Nothing survives a restart, and nothing
is shared between workers, so a browser whose next request lands on a
different worker arrives with a session that worker has never heard
of. It is the right store for a test and the wrong one for traffic.
Something of your own. A store is five methods, none of which sees an identifier or understands a payload:
class RedisStore < http.session.SessionStore {
@new(client) {
self._client = client
}
read(key) {
return self._client.get('session:' + key)
}
write(key, payload, expires_at) {
self._client.set_with_ttl('session:' + key, payload, (expires_at - time()).ceil())
}
destroy(key) {
self._client.remove('session:' + key)
}
gc(now) {
# Redis expires keys itself.
return 0
}
}
touch() has a working default built on read() and write();
override it where the backend can move an expiry on its own.
The key a store is handed is the SHA-256 of the identifier, not the identifier. Someone who reads the directory, the table, or a backup of either learns what is in the sessions but cannot resume one, because the value the browser presents is the preimage.
When a Session Ends
Two clocks, and a session ends at whichever runs out first:
| Option | Default | |
|---|---|---|
idle_timeout | 7200 | seconds of inactivity; rolls forward while the visitor is active |
lifetime | 86400 | seconds the session may live however active it is |
Either may be nil to remove that limit, but not both.
The absolute lifetime is what asks a tab left open overnight to sign in again.
touch_interval is how often a request that only read the session
bothers to move the idle expiry. It is the difference between a store
write on every request and a store write once a minute. It defaults to
sixty seconds, or half the idle timeout where that is shorter, and one
set by hand has to stay under idle_timeout — otherwise a session in
constant use still expires, because nothing ever moves it.
Nothing schedules a sweep of expired sessions, so one rides along with
ordinary traffic: gc_probability (default 0.01) is the chance that
a write also sweeps. Set it to 0 where a cron job calls store.gc()
instead.
The Cookie
| Option | Default | |
|---|---|---|
name | 'zuri_session' | |
path | '/' | |
domain | nil | nil scopes the cookie to the exact host, which is the narrower choice |
secure | nil | nil follows the request’s own scheme |
http_only | true | |
same_site | 'Lax' | 'Strict', 'Lax' or 'None' |
persistent | false | whether the cookie outlives the browser |
secure following the request is what lets development over cleartext
work while production over TLS gets Secure without being told. A
deployment behind a proxy that terminates TLS sets secure: true
itself.
persistent: false sends a cookie that ends with the browser session,
which is what a sign-in should normally do. persistent: true sends
Max-Age instead, and moves it forward on every request.
Signing the Cookie
An identifier is 32 bytes from the platform’s cryptographic generator, which is far beyond guessing. Signing adds nothing against that, and everything against volume: with a secret set, a cookie this server did not issue is thrown out after one HMAC, rather than after a read from disk or a query to the database.
import env
server.use(http.session.session({ secret: env.require('SESSION_SECRET') }))
Set it on anything facing the open internet, and give every worker the same secret — a cookie issued by one is otherwise refused by the next. Turning signing on refuses the cookies issued before it, so it signs everyone out once.
What a Session May Hold
Whatever JSON holds: strings, numbers, booleans, nil, lists and
dictionaries of those. A class instance is not JSON, and storing one
raises when the session is written.
Sessions are for identity and small state — who is signed in, which
steps of a form are done, what to say on the next page. A payload over
max_size (default 65536 bytes) raises SessionError; the answer to
that is a row in a database with the session holding its key.
Across Workers
http.serve() runs each worker in its own isolate, so the store is
built inside setup rather than handed in from outside:
# app.zu
import http
import http.session
def setup(server) {
server.use(http.session.session({
store: http.session.FileStore('/var/lib/app/sessions'),
secret: os.get_env('SESSION_SECRET'),
}))
server.get('/', @(request, response) {
response.text(request.session().get('account', 'nobody'))
})
}
A sql connection belongs to the isolate that opened it, so a worker
using SqlStore opens its own pool in setup too.
What Sessions Do Not Do
Two requests writing the same session at the same moment — parallel requests from one browser tab, usually — both succeed, and the one that finishes last is the one that survives. The file store writes through a rename, so a reader never sees half a session, and the SQL store writes in one statement; neither takes a lock. Do not use a session as a counter that several requests increment at once.
TLS
var server = http.server(443, '0.0.0.0')
server.load_certs('/etc/certs/site.crt', '/etc/certs/site.key')
server.listen()
Or from strings, with more control:
server.use_tls(cert_chain_pem, private_key_pem, {
min_version: '1.2',
client_ca: internal_ca_pem,
require_client_cert: true,
})
cert_chain must be the server certificate followed by any
intermediates. Leaving the intermediates out is the single most common
TLS misconfiguration, and it fails only for the clients that do not
happen to have them cached — which is to say, it fails in a way that
looks fine from your own browser.
Note The certificate and key are parsed when the first handshake runs, not when they are set, so a malformed or mismatched pair surfaces as a failed connection rather than as an error from
use_tls(). Make a request against the server as part of starting it if you want to find out early.
Turning on TLS also advertises h2 through ALPN, so a browser gets
HTTP/2 without anything else being configured.
HTTP/2
HTTP/2 needs nothing turned on. HttpServer speaks it when ALPN
negotiates h2 over TLS, and when a cleartext client opens with the
HTTP/2 connection preface. HttpClient uses it whenever a server
negotiates h2 for an HTTPS connection. Routes, middleware and
handlers are the same code either way; request.version is '2' when
it applies.
What that gets you is one connection carrying every request, HPACK header compression, and flow control — rather than the six-connection scramble HTTP/1.1 forces a browser into.
Set server.http2 = false to turn it off, which also stops h2 being
advertised in ALPN.
http.h2 exposes the machinery underneath — the frame codec, HPACK
and its Huffman coding — for anything that needs to speak the protocol
directly.
Note Streams are multiplexed on the wire but handled in the order they complete: one connection is served by one isolate, so a request that arrives while another is being handled is read, buffered, and answered next rather than in parallel. That is a throughput property, not a correctness one, and
http.serve()is how you get more than one connection served at a time.
WebSockets
import http.websocket as ws
server.get('/ws', @(request, response) {
var socket = ws.accept(request, response, { protocols: ['chat'] })
while socket.is_open() {
var message = socket.receive()
if message == nil or message.is_close() {
break
}
socket.send('you said: ' + message.text())
}
})
accept() completes the RFC 6455 handshake and takes the connection
over; the server writes nothing more on it. Fragmented messages are
reassembled, ping frames are answered, and a close frame from the peer
is answered and then handed back so you can see the code and reason.
The client side is ws.connect():
var socket = ws.connect('wss://example.com/ws')
socket.send('hello')
echo socket.receive().text()
socket.close()
Client frames are masked with a key from the platform’s secure random source, as the RFC requires — the mask is what stops a hostile page from steering a proxy into caching a forged response.
Server-Sent Events
Where a WebSocket gives you a duplex connection, server-sent events give you a long-lived response body and nothing else — which is all a live feed of updates needs, and it survives proxies that would refuse an upgrade.
import http.sse
server.get('/events', @(request, response) {
sse.stream(response, @(events) {
for update in updates {
events.send(update, 'update', update.id)
}
})
})
A browser’s EventSource reconnects on its own and sends back the
last id it saw; sse.last_event_id(request) reads it, which is how a
stream resumes where it left off instead of replaying from the start.
Reverse Proxying
var api = http.ReverseProxy('http://127.0.0.1:9000', {
strip_prefix: '/api',
})
server.any('/api/' + '*path', @(request, response) {
api.handle(request, response)
})
Hop-by-hop headers are stripped in both directions — including any
field the message’s own Connection header names, which is how an
endpoint declares an extra one. Forwarding a client-supplied
Transfer-Encoding to an upstream that frames it differently is the
classic request-smuggling setup, so this is not optional.
X-Forwarded-For, X-Forwarded-Proto, X-Forwarded-Host and RFC
7239’s Forwarded are added, and the response body is streamed rather
than buffered.
Several upstreams get a LoadBalancer, which is round-robin with an
upstream that fails taken out of rotation for a while rather than
retried on every request:
var pool = http.LoadBalancer([
'http://10.0.0.1:9000',
'http://10.0.0.2:9000',
], { recovery_time: 30 })
Running in Production
Using More Than One Core
listen() serves connections on the calling isolate, one at a time.
http.serve() runs the same pipeline across a pool of isolates: one
accept loop hands each connection to whichever worker takes it next.
# app.zu
import http
def setup(server) {
server.get('/', @(request, response) {
response.text('hello')
})
server.serve_files('/static', './public', { cache_age: 86400 })
}
# main.zu
import http
import .app
http.serve(app.setup, {
host: '0.0.0.0',
port: 3000,
workers: 8,
})
setup is called once inside each worker with that worker’s own
HttpServer. It can live in the main script or in a module, and it can
use whatever it imports; a module it reaches for is loaded again inside
the worker rather than shared with it. Putting it in a module of its own,
as above, is still the better shape once it registers more than a couple
of routes.
Isolates share no memory, so anything a worker needs — a cache, a connection pool, a counter — is per worker. That is the trade the model makes: no locks and no shared-heap garbage collection pauses, in exchange for state that has to be either per worker or in something outside the process.
backlog bounds how many accepted connections may queue for a free
worker. Bounding it is deliberate: an unbounded queue under load means
accepting connections faster than they can be served and then
answering all of them late, rather than letting the kernel’s own
listen backlog apply back-pressure.
Limits
Every part of a request whose size a peer controls has a ceiling:
| Setting | Default | Bounds |
|---|---|---|
max_line_size | 8 KiB | the request line, and each header line |
max_header_size | 64 KiB | the header section |
max_header_count | 100 | how many header fields |
max_body_size | 10 MiB | the request body |
header_timeout | 10s | how long the whole head may take to arrive |
keep_alive_timeout | 5s | how long an idle connection is held |
max_keep_alive_requests | 1000 | how many requests one connection may serve |
header_timeout is the one that is easy to leave out and matters
most: without it, a client sending one byte every few seconds holds a
connection open indefinitely while never tripping any single read’s
timeout. That is what slowloris is.
Raise max_body_size deliberately for a service that takes uploads,
rather than discovering it by accident:
server.max_body_size = 100 * 1024 * 1024
A request that announces a body over the limit is refused with 413
before a byte of it is read, which is the whole point of the header.
Graceful Shutdown
import os
http.serve(app.setup, {
port: 3000,
on_ready: @(address, stop) {
echo 'listening on ${address}'
os.on_signal('INT', @() {
stop()
return true
})
os.on_signal('TERM', @() {
stop()
return true
})
},
})
stop() closes the listener, which is what breaks the blocking
accept; serve() then waits for every worker to finish the connection
it is on before returning. For a single-isolate server, close() does
the same thing.
Behind Another Proxy
If something else really is in front — a CDN, a load balancer you control — tell the server, and tell it which addresses to believe:
server.trust_proxy = true
server.trusted_proxies = ['10.0.0.1', '10.0.0.2']
request.client_ip(true, server.trusted_proxies) then walks the
forwarding chain in from the proxy end and returns the rightmost entry
that is not itself a trusted proxy. Taking the leftmost entry — the
common shortcut — hands an attacker whatever client address they care
to claim, which matters the moment an address is used for rate
limiting, allowlisting or an audit trail.
Left off, forwarding headers are ignored entirely and the peer is the client. That is the right default: those headers are request headers, which is to say anyone can write anything in them.
What the Module Refuses
Some of what this module does is refuse things, and it is worth being explicit about which, since each refusal is a request that some other implementation would have accepted:
- A header field name with whitespace before its colon.
Content-Length : 5is read as a length by some parsers and as an unknown field by others, and that disagreement is a smuggled request. - Two
Content-Lengthfields that disagree, for the same reason. - A
Content-Lengthalongside aTransfer-Encoding. - A
Transfer-Encodingwhose last coding is notchunked. - An obsolete folded header line — RFC 9112 deprecated it, and intermediaries unfold it differently.
- An HTTP/1.1 request with no
Host, or with two. - A header value containing CR, LF or NUL, at the point it is set: a value that can inject a newline is a response-splitting bug wherever it eventually lands.
- An uppercase header field name over HTTP/2, and any of the connection-specific fields there.
- A static file path that resolves outside the root, before or after percent-decoding, and any dotfile.
Module Reference
| Module | What it holds |
|---|---|
http | the facade: get(), post(), server(), client(), serve() |
http.status | status codes, reason phrases, and predicates |
http.headers | Headers, field validation, canonical names |
http.cookies | Cookie, CookieJar, and both cookie header formats |
http.session | Session, SessionStore, FileStore, MemoryStore |
http.session.sql | SqlStore, for sessions kept in a database |
http.request | HttpRequest, query string encoding and decoding |
http.response | HttpResponse |
http.router | Router, Route, RouteMatch |
http.middleware | CORS, logging, security headers, auth, rate limiting |
http.files | StaticFiles, byte ranges, validators |
http.multipart | MultipartBuilder, UploadedFile, the parser |
http.negotiate | Accept parsing and matching |
http.body | message framing, chunked encoding, content codings |
http.h1 | the HTTP/1.1 wire codec |
http.h2 | HTTP/2: frames, HPACK, connections |
http.websocket | RFC 6455, client and server |
http.sse | server-sent events |
http.proxy | ReverseProxy, LoadBalancer |
http.stream | the buffered connection both protocols read and write through |
http.util | dates, header parameters, percent coding, path normalisation |
http.errors | the error hierarchy |
Two other standard library modules meet this one without it depending
on either: HttpRequest.validate() takes a schema from validate,
and middleware.jwt_auth() takes a verifier from jwt. Both are
duck-typed, so neither module is loaded by a server that does not use
it.
Every error this module raises descends from HttpError:
ProtocolError for a malformed message, ConnectionError for a
connection that failed, TimeoutError, TooLargeError,
TooManyRedirectsError, StatusError, UnsupportedProtocolError and
SessionError.
Imagine
The imagine module is Zuri’s image library: decoding, drawing,
transforming and encoding raster images.
It covers the whole path an image takes through a program. A photograph arrives as an upload, gets checked, oriented, resized and sharpened, has a watermark composited onto it, and goes back out as a WebP. A chart is built from nothing but shapes and text. An avatar is cropped to a square, rounded off, and cached as a data URL. None of that needs anything outside the standard library.
Every image on this page is the output of the code beside it, produced by
docs/book/src/imagine/figures.zu. Re-run that script and the figures follow whatever the module actually does.
- Introduction
- Colours
- Loading and Saving
- Resizing and Reshaping
- Filters
- Drawing
- Text
- Compositing
- Inspecting an Image
- Animation
- Working With Pixels Directly
- Errors
- Performance and Memory
- Recipes
Blocks in this chapter that list several calls together — three ways to save, four filters, a family of methods — are reference listings, not programs. They show the shape of each call rather than a sequence you could run, and several would conflict if pasted into one file. Anything presented as a complete program on this page runs as written.
Following Along
Most examples below operate on a file called photo.jpg. Use your own, or
make one — the module can draw its own test subject:
import imagine { Image }
Image(320, 240, '#1e3a5f')
.fill_circle(220, 70, 40, '#ffd166')
.fill_rect(0, 170, 320, 70, '#2a9d8f')
.fill_polygon([[40, 170], [110, 80], [180, 170]], '#264653')
.save('photo.jpg')
echo Image.open('photo.jpg').size()
{width: 320, height: 240}
Everything from here on assumes that file exists in the working directory.
Introduction
A First Image
import imagine { Image }
Image.open('photo.jpg')
.thumbnail(400, 400)
.save('thumb.webp')
Three lines: decode, shrink, encode. The format on the way in comes from the file’s contents, and the format on the way out comes from the extension you saved it as.
Building an image from scratch works the same way.
import imagine { Image, Color }
Image(400, 200, '#0f172a')
.fill_circle(200, 100, 70, '#38bdf8')
.circle(200, 100, 70, 'white', { thickness: 3 })
.save('badge.png')
Almost every method returns an image, so operations chain.
Copying and Mutation
There is one rule in this module that is worth learning before anything else, because everything else follows from it:
Operations that change the image’s size return a new image. Everything else changes the image in place.
resize(), crop(), rotate(), flip(), transpose() and pad()
all leave the original untouched and hand back a new one. Drawing,
filters and compositing all modify the image you called them on and
return it so the chain continues.
var original = Image.open('photo.jpg')
var small = original.thumbnail(200, 200) # original is untouched
small.grayscale() # small is now grey
echo original.size() # still the full size, still colour
That means a mixed chain reads correctly, and it means clone() is what
you reach for when you want to keep an image before filtering it.
var greyed = photo.clone().grayscale() # photo keeps its colour
Colours
Writing a Colour
Anywhere a colour is expected, five spellings are accepted. These are all the same red:
image.fill(Color(255, 0, 0))
image.fill('#ff0000')
image.fill('red')
image.fill(0xFF0000FF)
image.fill([255, 0, 0])
Hexadecimal strings take all four CSS lengths, with or without the
leading #: #f00, #f00c, #ff0000, #ff0000cc. Names are the full
CSS Color Level 4 list, matched ignoring case, spaces and hyphens, so
'Dark Sea Green' and 'darkseagreen' are the same colour.
Alpha runs 0 (transparent) to 255 (opaque), as in CSS and PNG.
The Color Class
Color is an immutable value type. Every method that would change one
returns a new colour, so a colour held in a variable is safe to pass
around.
import imagine { Color }
var brand = Color.hex('#4f46e5')
brand.r # 79
brand.to_hex() # '#4f46e5'
brand.to_packed() # 0x4F46E5FF
brand.with_alpha(128) # half transparent
brand.fade(0.5) # halves whatever alpha it already had
brand.lighten(20) # 20 percentage points of HSL lightness
brand.darken(20)
brand.saturate(15)
brand.desaturate(15)
brand.rotate_hue(180)
brand.mix('white', 0.25) # a quarter of the way to white
brand.invert()
brand.to_grayscale()
over() flattens a translucent colour against a background, which is
what happens when it is drawn onto something solid:
Color(255, 0, 0, 128).over('white') # '#ff7f7f'
Colour Spaces
Conversions come from the colors module, so the conventions are
the same ones it and CSS use: hue in degrees, everything else in
percentage points from 0 to 100.
Color.hsl(240, 100, 50) # '#0000ff'
Color.hsv(120, 100, 100) # '#00ff00'
Color.hwb(0, 0, 0) # '#ff0000'
Color.cmyk(0, 100, 100, 0) # '#ff0000'
Color.lab(40.7, 50.6, -79.1) # perceptually uniform
Color.xyz(0.2, 0.15, 0.7)
brand.to_hsl() # { h: 243.4, s: 75.4, l: 58.6, a: 255 }
brand.to_hsv()
brand.to_hwb()
brand.to_cmyk()
brand.to_lab()
brand.to_xyz()
Lab is the right space for interpolating between two colours when the intermediate steps need to look evenly spaced to the eye. CMYK here is the naive conversion with no output profile, so it is right for generating colours and wrong for predicting a printing press.
Contrast and Accessibility
var background = Color.hex('#1e293b')
background.luminance() # 0 to 1, WCAG relative luminance
background.contrast_ratio('#ffffff') # 1 to 21
background.is_dark() # true
background.best_contrast() # white, since the background is dark
WCAG asks for a ratio of at least 4.5 for normal text and 3 for large
text, so best_contrast() is the quick way to pick a legible foreground
for a colour you did not choose:
var label = swatch.best_contrast('#ffffff', '#111111')
card.text(20, 20, name, font, label)
Loading and Saving
Opening an Image
var photo = Image.open('photo.jpg') # from a path
var photo = Image.open(file('photo.jpg')) # from an open file
var photo = Image.decode(upload) # from bytes in memory
The format is detected from the contents rather than the name, so a
mislabelled file still opens. When the contents cannot be identified —
TGA carries no signature of its own — Image.open() falls back to the
file extension.
Image.decode() takes options:
Image.decode(upload, { format: 'png' }) # fail unless it really is a PNG
Image.decode(raw, { orient: false }) # skip EXIF auto-rotation
Naming the format explicitly is worth doing when you already know it
from a Content-Type or an extension: a file claiming to be a PNG that
is really something else then fails loudly instead of being decoded as
whatever it actually is.
Checking Before Decoding
An image costs four bytes per pixel once decoded, so a 6000x4000 photograph occupies 96 MB in memory however small its file was. For anything arriving from outside the program, read the header first:
import imagine
import imagine { Image }
var header = imagine.probe(upload)
if header == nil {
raise Exception('that is not an image')
}
if header.width * header.height > 40000000 {
raise Exception('image is too large')
}
var photo = imagine.decode(upload)
probe() returns {format, width, height} or nil. It reads only the
header, which costs microseconds where decoding the same file costs tens
of milliseconds. imagine.detect() is the same check when you only want
the format name.
Saving
photo.save('out.png')
photo.save('out.jpg', { quality: 90 })
photo.save('out.dat', { format: 'webp' }) # extension overridden
The format comes from the extension unless format says otherwise. To
get the bytes instead of a file:
var data = photo.encode('webp')
var data = photo.to_png()
var data = photo.to_jpeg(90)
encode() with no format uses whatever the image was decoded from,
falling back to PNG for an image built in memory.
Format Reference
| Format | Read | Write | Alpha | Notes |
|---|---|---|---|---|
| PNG | yes | yes | yes | Lossless. The safe default. |
| JPEG | yes | yes | no | Lossy. Photographs only. |
| WebP | yes | yes | yes | Written lossless, so smaller than PNG but larger than a lossy WebP. |
| AVIF | no | yes | yes | Smallest files, slowest to encode. |
| GIF | yes | yes | 1-bit | 256 colours. The only animated format that can be written. |
| BMP | yes | yes | yes | Uncompressed and enormous. |
| TIFF | yes | yes | yes | Common in printing and scanning. |
| TGA | yes | yes | yes | No signature; needs an extension or an explicit format. |
| QOI | yes | yes | yes | Lossless, several times faster than PNG. |
| ICO | yes | yes | yes | Icon container. |
| PNM | yes | yes | no | Trivially simple, trivially large. |
| WBMP | yes | yes | no | One bit per pixel. |
Three entries need explaining.
AVIF is write-only: opening one raises DecodeError.
TGA has no magic number, so it can only be identified by its file extension or by naming the format outright.
WebP is written losslessly. Reading handles both lossy and lossless
WebP, but the encoder here only writes lossless, so quality has no
effect on it and a photograph saved as WebP will be larger than one
saved by a tool that can write lossy WebP. For a photograph where size
matters, JPEG or AVIF is the better target; WebP here is a
smaller-than-PNG lossless format with alpha.
Never assume; ask:
import imagine
var can = imagine.capabilities()
echo can.decode.contains('png')
echo can.encode.contains('webp')
echo can.animated.contains('gif')
true
true
true
decode, encode and animated are each a list of format names, so
contains() answers the question you actually have.
Encoding Options
| Option | Formats | Meaning |
|---|---|---|
quality | JPEG, AVIF | 1 to 100. Defaults to 85 for JPEG, 80 for AVIF. |
compression | PNG | 'fast', 'default' or 'best'. All lossless. |
speed | AVIF, GIF | AVIF 1-10, lower is slower and smaller. GIF 1-30, lower is slower and picks better colours; 15 by default. |
background | JPEG | What transparent pixels are flattened against. Defaults to white. |
threshold | WBMP | The brightness cut, 0 to 255. |
JPEG has no alpha channel, so transparency has to go somewhere on the
way out. It is flattened against background rather than silently
dropped, because dropping it turns transparent pixels black:
logo.save('logo.jpg', { background: '#ffffff' })
To see the result first, or to choose the colour once and keep it, do it explicitly:
logo.flatten('#ffffff').save('logo.jpg')
Data URLs
var url = icon.to_data_url('png')
# 'data:image/png;base64,iVBORw0...'
Base64 costs a third more bytes than the raw image, so this suits icons and small graphics rather than photographs.
Resizing and Reshaping
Choosing a Resize
There are six, and picking the right one is most of the work:
| Method | Aspect ratio | Enlarges? | Result |
|---|---|---|---|
resize(w, h) | not kept | yes | exactly w x h, possibly distorted |
scale(factor) | kept | yes | the same image, scaled |
thumbnail(w, h) | kept | no | fits inside the box |
fit(w, h) | kept | yes | fits inside the box |
cover(w, h) | kept | yes | fills the box, cropping the overflow |
contain(w, h) | kept | yes | fits the box, padding the remainder |
thumbnail() is the one you usually want for thumbnails, and the
refusal to enlarge is the reason: scaling a small image up to fill a
thumbnail box only makes it blurry. Use fit() when you do want it
scaled up.
cover() and contain() both produce exactly the size asked for. They
differ in what they sacrifice — cover() loses part of the image,
contain() adds bars:
photo.cover(300, 300) # crops to a square
photo.cover(300, 300, { anchor: TOP }) # keeps the top, crops the bottom
photo.contain(300, 300, { background: 'white' }) # letterboxes instead
The same 320x180 image asked for a 120x120 result four ways:
thumbnail() keeps the proportions and does not fill the box.
cover() fills it and loses the sides. contain() fills it and adds
bars. resize() fills it by distorting the picture.
The anchors are TOP_LEFT, TOP, TOP_RIGHT, LEFT, CENTER,
RIGHT, BOTTOM_LEFT, BOTTOM and BOTTOM_RIGHT. CENTER is the
default, and TOP is what you want for photographs of people, where the
face is rarely in the bottom third.
Resampling Filters
| Filter | Speed | Use |
|---|---|---|
NEAREST | fastest | Pixel art, QR codes, anything whose edges must stay hard. |
BILINEAR | fast | When speed matters more than sharpness. |
BICUBIC | moderate | A good default for photographs. |
GAUSSIAN | moderate | Deliberately soft; for noisy input, or before sharpening. |
LANCZOS | slowest | Thumbnails and any large reduction. The default. |
photo.thumbnail(200, 200, LANCZOS)
sprite.scale(4, NEAREST) # keeps pixel art crisp
A 16x16 sprite enlarged seven times over, so the differences are visible at all:
This is the case where NEAREST is right and everything else is
wrong. Shrinking a photograph inverts that judgement entirely.
LANCZOS is the default because most resizing is shrinking, and that is
where it earns its cost. It can produce faint ringing next to very
high-contrast edges, which is the price of its sharpness; BICUBIC is
the fallback when that shows.
Cropping, Padding and Trimming
photo.crop(100, 50, 400, 300) # x, y, width, height
photo.pad(20) # 20 pixels on every side
photo.pad(10, 20, 10, 20, 'white') # top, right, bottom, left, colour
photo.trim() # remove a uniform border
photo.trim({ tolerance: 8 }) # allow for JPEG noise in the border
crop() requires its rectangle to fit inside the image and raises
BoundsError otherwise. Clamping a crop that runs off the edge would
hand back different dimensions than were asked for, which is a worse
surprise than an error. Pad first when the region really is meant to
extend past the edge.
trim() takes the border colour from the top-left pixel unless you name
one. Scanned documents and screenshots almost always want a tolerance,
because a “white” border from a lossy format is not exactly white.
Rotating and Mirroring
photo.rotate_90() # lossless
photo.rotate_180()
photo.rotate_270()
photo.rotate(37, 'white') # resamples, grows the canvas
photo.flip_horizontal()
photo.flip_vertical()
photo.flip(FLIP_BOTH)
photo.transpose() # reflect across the main diagonal
Quarter turns are exact: every pixel is moved, none is resampled. Any
other angle interpolates, and grows the canvas to hold the rotated
corners with background filling the gaps.
EXIF Orientation
Phone cameras usually store the sensor’s orientation in the file rather
than rotating the pixels, so a photograph whose bytes are sideways is
meant to be displayed upright. Image.open() and Image.decode()
handle this for you.
Image.open('photo.jpg') # upright
Image.open('photo.jpg', { orient: false }) # exactly as stored
All eight orientation values are handled, including the four mirrored ones that scanners and some front-facing cameras produce.
Filters
Filters change the image in place and return it, so they chain.
photo.grayscale().contrast(15).sharpen(0.5)
Tone and Exposure
photo.brightness(20) # -255 to 255, added to each channel
photo.contrast(15) # -100 to 100, around mid-grey
photo.gamma(1.2) # above 1 lifts midtones, below 1 lowers them
photo.levels(20, 235) # stretch this input range to full scale
photo.levels(20, 235, 1.1) # ...with a midtone curve
photo.threshold(128) # every pixel to black or white
photo.posterize(6) # six evenly spaced steps per channel
photo.invert() # a photographic negative
photo.opacity(0.5) # scale the alpha channel
levels() is the single most useful correction for a flat or
washed-out photograph. Everything at or below the black point becomes
black, everything at or above the white point becomes white, and the
range between is stretched to fill the scale.
Note the difference between posterize() and quantize():
posterize() spaces its levels evenly, quantize() picks the colours
to suit the image, so a photograph survives far fewer of them.
photo.quantize(32) # 32 well-chosen colours
photo.quantize(16, true) # ...with dithering
auto_levels() does the same job as levels() without being told
where the endpoints are; it is shown alongside the detail filters
below.
Colour
photo.grayscale()
photo.sepia()
photo.saturate(1.4) # 0 removes colour, 1 is unchanged
photo.hue_rotate(45)
photo.tint('#ff8800', 0.3) # blend a flat colour in
photo.colorize('#4f46e5') # monochrome in one colour
photo.duotone('#1e1b4b', '#fbbf24')
photo.flatten('white') # composite over a colour, drop alpha
grayscale() weights the channels for perceived brightness, so a bright
yellow comes out light and a deep blue comes out dark. That is different
from saturate(0), which keeps HSL lightness and makes both mid-grey.
Blur, Sharpen and Detail
photo.blur(3) # Gaussian; the argument is its sigma
photo.sharpen(0.8)
photo.smooth() # a cheap, harsher 3x3 average
photo.emboss()
photo.edges()
photo.mean_removal()
photo.pixelate(12)
blur() is a true separable Gaussian, so a large radius costs linearly
rather than quadratically. Its argument is the standard deviation: about
two thirds of each pixel’s contribution falls within that distance, and
the visible spread is roughly three times it.
Alpha is premultiplied for the duration of a blur, so blurring a shape on a transparent background does not drag a dark halo into its edge.
Writing Your Own Filter
Nearly every filter above is one of three primitives with different numbers in it, and all three are available directly.
A lookup table covers any per-channel tone curve. The table is built once whatever the image’s size, then applied at memory speed.
import imagine { filters }
# A gentle S-curve: more contrast, but the highlights survive.
var curve = filters.build_lut(@(value) {
var t = value / 255
return 255 * t * t * (3 - 2 * t)
})
photo.apply_lut(curve, curve, curve, nil) # nil leaves alpha alone
A colour matrix covers anything that mixes channels — saturation, hue rotation, channel swaps, tinting. It is the 4x5 matrix SVG and CSS filters use, written row by row, with the last entry of each row a constant in 0-255 units.
# Swap the red and blue channels.
photo.apply_matrix([
0, 0, 1, 0, 0,
0, 1, 0, 0, 0,
1, 0, 0, 0, 0,
0, 0, 0, 1, 0,
])
Applying several matrices in a row is both slower and less accurate than combining them and applying the result once:
var both = filters.combine_matrices(
filters.grayscale_matrix(),
filters.saturation_matrix(1.2)
)
photo.apply_matrix(both)
A convolution kernel covers anything that reads a pixel’s neighbours.
photo.convolve([
0, -1, 0,
-1, 5, -1,
0, -1, 0,
])
The kernel must be square with an odd side. The default divisor of 0 means “divide by the kernel’s own sum”, which is what nearly every published kernel expects.
One trap is worth knowing. Alpha goes through the kernel along with the colour channels, which is right for a blur and wrong for anything whose weights do not sum to one. A Laplacian over a uniformly opaque image sums to zero, which would make the whole result invisible:
photo.convolve(filters.edge_kernel(), { divisor: 1, keep_alpha: true })
The built-in edges(), emboss(), sharpen() and mean_removal()
already do this.
The S-curve and the channel-swap matrix from this section, run against the same picture:
edge controls what happens off the image’s border: EDGE_CLAMP (the
default) repeats the nearest edge pixel, EDGE_TRANSPARENT treats the
outside as empty, and EDGE_WRAP wraps to the opposite side for images
meant to tile.
Drawing
Shapes
Coordinates start at the top-left corner. Integer coordinates fall on pixel corners rather than centres, so a rectangle from (0, 0) to (10, 10) covers exactly the first ten pixels in each direction.
image.pixel(x, y, color) # one pixel, blended
image.line(x1, y1, x2, y2, color)
image.rect(x, y, w, h, color) # outline
image.fill_rect(x, y, w, h, color)
image.rounded_rect(x, y, w, h, radius, color)
image.fill_rounded_rect(x, y, w, h, radius, color)
image.circle(cx, cy, r, color)
image.fill_circle(cx, cy, r, color)
image.ellipse(cx, cy, rx, ry, color)
image.fill_ellipse(cx, cy, rx, ry, color)
image.arc(cx, cy, rx, ry, start, end, color)
image.pie(cx, cy, rx, ry, start, end, color) # closed through the centre
image.chord(cx, cy, rx, ry, start, end, color) # closed along the chord
image.bezier(x1, y1, cx1, cy1, cx2, cy2, x2, y2, color)
Angles are in degrees, measured clockwise from three o’clock, matching the direction the y axis runs.
Drawing outside the image is never an error; anything that falls outside is clipped away. That is what makes it safe to draw a shape that only partly overlaps.
Strokes and Anti-aliasing
image.thickness(4) # applies to every subsequent stroke
image.antialias(false) # hard edges
Both are settings on the image rather than per-call arguments, though a single call can override the thickness:
image.line(0, 0, 100, 100, 'black', { thickness: 8 })
Strokes are centred on the path, so a thickness of 4 puts 2 pixels on each side. Joints and the ends of an open path are rounded.
Anti-aliasing is on by default. Turn it off for output that has to be pixel-exact — barcodes, QR codes, anything that will be thresholded afterwards.
Stroke widths of 1, 3, 6 and 12:
Line Caps
How an open stroke finishes at its two ends is a separate choice from its width:
image.cap(CAP_ROUND) # a half-disc. The default.
image.cap(CAP_SQUARE) # a square, reaching the same distance
image.cap(CAP_BUTT) # stops dead at the endpoint
The red marks are the exact coordinates the line was given. A round or square cap reaches half the stroke’s width past them; a butt cap does not.
CAP_BUTT is the one to reach for whenever the coordinates have to
mean exactly what they say — segments meeting end to end, a scale bar
of a known length, the pieces of a dashed line. CAP_SQUARE gives the
same reach as round with a blunt finish.
A single call can override the surface’s setting:
image.line(40, 200, 40, 40, '#334155', { thickness: 6, cap: CAP_BUTT })
Joins between a path’s segments are always round, and closed outlines have no ends, so neither is affected by this.
Paths and Polygons
Every filled shape in this module becomes a polygon and goes through one scanline rasterizer, and every outline becomes the polygon around its stroke and goes through the same one. A shape not listed above can be drawn by supplying the points:
image.fill_polygon([[10, 10], [90, 30], [50, 80]], '#4f46e5')
image.polygon([[10, 10], [90, 30], [50, 80]], 'black') # outline, closed
image.polyline([[10, 10], [90, 30], [50, 80]], 'black') # not closed
Points may be a list of [x, y] pairs or a flat list of alternating
values; both read naturally depending on where the points came from.
The outline is closed for you, so the last point does not need to repeat
the first.
Self-intersecting outlines are filled by the non-zero winding rule,
which fills a five-pointed star solid. Pass { even_odd: true } for the
other convention, which leaves its middle empty.
Filling Areas
image.fill('#0f172a') # every pixel, ignoring the clip
image.clear() # every pixel to transparent
image.flood_fill(x, y, color)
image.flood_fill(x, y, color, 12) # with a tolerance
Gradients get their own section below.
fill() replaces rather than blends, so filling with a transparent
colour empties the image instead of leaving it unchanged.
Flood fill spreads four-connected — up, down, left and right, but not
diagonally — through pixels within tolerance of the colour at the
starting point. A tolerance of 0 spreads only through exactly equal
pixels; on a photograph or anything anti-aliased you will want more.
Gradients
image.linear_gradient(x1, y1, x2, y2, stops)
image.radial_gradient(cx, cy, radius, stops)
A linear gradient runs along the vector from the first point to the second. Everything before that vector takes the first stop’s colour and everything past it takes the last stop’s, so a short vector across a large area gives a hard transition with flat bands on either side.
# top to bottom
image.linear_gradient(0, 0, 0, image.height(), ['#0f172a', '#334155'])
# diagonal, three colours
image.linear_gradient(0, 0, 400, 200, ['#4f46e5', '#f472b6', '#fbbf24'])
Stops are either bare colours, spaced evenly, or [offset, colour]
pairs with the offset running 0 to 1:
image.linear_gradient(0, 0, 200, 0, [
[0, 'black'],
[0.25, 'red'],
[1, 'white'],
])
By default a gradient replaces what is there. Pass { blend: true } to
composite it instead, which is what makes overlays and vignettes work:
# a vignette over an existing photograph
photo.radial_gradient(
photo.width() / 2,
photo.height() / 2,
photo.width() * 0.7,
[[0, Color(0, 0, 0, 0)], [1, Color(0, 0, 0, 180)]],
{ blend: true }
)
# a caption scrim along the bottom
photo.linear_gradient(0, photo.height() - 120, 0, photo.height(), [
Color(0, 0, 0, 0),
Color(0, 0, 0, 200),
], { blend: true })
{ rect: {x, y, width, height} } confines the fill to a rectangle
instead of covering the whole surface.
Both of those are the snippets above, run against the picture on the left.
Clipping
image.clip(20, 20, 100, 100)
image.fill_rect(0, 0, 500, 500, 'red') # only the clip is painted
image.clear_clip()
A clip confines every drawing operation until it is cleared. It does not
affect reading: get_pixel() sees the whole image either way.
Text
The Built-in Font
imagine ships no font file, so everything in the next section
depends on what is installed on the machine and can fail. One font
always works:
import imagine { Image, StrokeFont }
Image(300, 70, 'white')
.text(16, 20, 'Always available', StrokeFont(28), '#0f172a')
.save('label.png')
Font.builtin(size) is the same thing under a name you will find from
Font.
It is not a font file. Every glyph is defined as geometry — centre-line
strokes rather than filled outlines — in
libs/imagine/strokefont.zu, and
drawn through the same anti-aliased rasterizer as everything else. So
it scales cleanly to any size, and its weight is a parameter rather
than part of the design:
StrokeFont(24) # regular
StrokeFont(24).weight(0.13) # bold
StrokeFont(24).weight(0.05) # light
Here is the whole thing:
What it is for. Labels, chart axes, watermarks, diagrams, placeholder text, and any output that has to work on a machine with no fonts installed. The look is geometric and single-weight, closer to a technical drawing than to a typeface, because that is what centre-line strokes give you honestly.
What it is not for. Body text, headlines, or anything where the shapes themselves matter. Load a real font for those.
Coverage is printable ASCII, from space through ~. Anything else
draws the empty box a font uses for a glyph it does not have, so text
in another script comes out visibly missing rather than silently blank.
A StrokeFont has the same interface as a Font — size(),
metrics(), measure(), render(), wrap() — so anywhere a font is
accepted, either works.
Loading a Font
import imagine { Font }
var font = Font.load('assets/Inter.ttf', 24)
var font = Font.from_bytes(embedded_font, 24)
var font = Font.system('DejaVu Sans', 24)
var font = Font.sans(24)
Font.load() reads a TrueType or OpenType file and caches the parsed
face by path, so loading the same file twice in one process parses it
once. Font.system() searches the platform’s font directories by family
name, matching ignoring case, spaces and hyphens.
Font.sans() finds whichever common sans-serif font is installed. It is
a convenience for scripts and tests, not something to rely on for output
that must look the same everywhere: which font it lands on depends on
the machine, and a container image with no fonts installed has none to
find. It raises FontError in that case, pointing at
Font.builtin(), which never depends on what is installed.
A Font is immutable and cheap to copy. size() returns the same face
at another size, sharing the parsed data:
var title = font.size(32)
var body = font.size(14)
Drawing Text
image.text(20, 20, 'Hello', font, '#111111')
x and y are the top-left corner of the text’s box, not its baseline,
because the corner is what you know when placing text in a layout. Pass
{ baseline: true } when you want y to mean the first line’s
baseline instead.
A \n starts a new line.
card.text(24, 24, 'Quarterly report\n2026', title, '#111111', {
align: ALIGN_CENTER,
line_height: 1.4,
tracking: 0.5,
})
align positions the lines against each other, not against the image.
line_height is a multiplier on the font’s own recommended spacing, and
tracking adds pixels between characters.
Measuring and Positioning
var box = image.text_size('Hello', font)
# { width: 58, height: 28, baseline: 22.3, lines: 1 }
Measuring costs a fraction of drawing, so it is the right way to lay text out before committing to it — centring, wrapping, or sizing a background to fit.
var box = card.text_size(label, font)
card.fill_rounded_rect(16, 16, box.width + 24, box.height + 16, 8, '#1e293b')
card.text(28, 24, label, font, 'white')
To position by anchor instead of coordinates:
card.place_text('SOLD OUT', font, 'white', CENTER)
card.place_text('v2.1', font, '#94a3b8', BOTTOM_RIGHT, { margin: 12 })
Wrapping
var body = Font.load('assets/Inter.ttf', 16)
page.text(40, 120, article, body, '#334155', { width: 520 })
The width option wraps the text to that many pixels before drawing
it. text_size() takes the same option, so measuring wrapped text
gives the box it will actually occupy.
To get the broken text itself — to store it, or to draw it in pieces —
call wrap() on the font:
var lines = body.wrap(article, 520).split('\n')
echo '${lines.length()} lines'
Newlines already in the text are kept as paragraph breaks, and runs of spaces are collapsed. Words are kept whole where they can be; a single word too long for the width is broken between characters rather than allowed to overflow, since text spilling out of an image cannot be scrolled to.
What Text Layout Does Not Do
Glyphs are positioned by advance width with kerning applied. That covers Latin, Greek, Cyrillic and anything else written left to right without contextual shaping.
It does not do complex shaping. Arabic letters will not join, Indic clusters will not reorder, and ligatures are not substituted. Those need a shaping engine, and a standard library that pretended to do them would be worse than one that says plainly it does not.
Compositing
Drawing One Image Onto Another
photo.draw_image(logo, 20, 20)
photo.draw_image(logo, 20, 20, { opacity: 0.6 })
photo.place(logo, BOTTOM_RIGHT, { margin: 16, opacity: 0.5 })
The source is clipped to the destination, and its position may be negative, so a sprite hanging off the top-left corner draws correctly. Compositing an image onto itself works; it is copied first.
Blend Modes
photo.draw_image(texture, 0, 0, { blend: BLEND_MULTIPLY })
| Mode | Effect |
|---|---|
BLEND_NORMAL | Ordinary alpha compositing. The default. |
BLEND_MULTIPLY | Never lighter than either input. Shadows. |
BLEND_SCREEN | Never darker than either input. Glows. |
BLEND_OVERLAY | Multiply in the shadows, screen in the highlights. |
BLEND_DARKEN / BLEND_LIGHTEN | Keep the darker or lighter channel. |
BLEND_COLOR_DODGE / BLEND_COLOR_BURN | Strong brighten or darken. |
BLEND_HARD_LIGHT / BLEND_SOFT_LIGHT | Overlay with the roles swapped; and a gentler version. |
BLEND_DIFFERENCE / BLEND_EXCLUSION | Absolute difference; and a lower-contrast variant. |
BLEND_ADD / BLEND_SUBTRACT | Add or subtract, clamped. |
The formulas are the ones in the CSS compositing specification, which is also what every image editor implements.
BLEND_DIFFERENCE makes a quick visual diff: identical images blended
this way come out black.
A red circle drawn onto a gradient under eight of the modes:
var diff = before.clone().draw_image(after, 0, 0, { blend: BLEND_DIFFERENCE })
Drawing operations — lines, shapes, text — always use ordinary alpha compositing. To draw a shape under a blend mode, draw it on a layer and composite the layer.
Masks
var stencil = photo.layer()
stencil.fill_circle(200, 200, 150, 'white')
photo.mask(stencil) # everything outside the circle becomes transparent
Each pixel keeps its colour and takes its transparency from the mask: opaque white shows the image through fully, black or transparent hides it, and greys give partial transparency. The mask’s own alpha counts, so a shape drawn on a transparent background works as a mask without being filled in first.
The mask must be the same size as the image.
Layers
layer() gives a transparent image of the same size. Building a
composite out of layers is how you get effects that a single pass
cannot, and it is the answer whenever you want a blend mode or an
opacity applied to a group of operations rather than one:
var glow = photo.layer()
glow.fill_circle(200, 150, 80, '#fbbf24')
glow.blur(30)
photo.draw_image(glow, 0, 0, { blend: BLEND_SCREEN, opacity: 0.7 })
Inspecting an Image
photo.width() # dimensions
photo.height()
photo.size() # { width, height }
photo.bounds() # { x: 0, y: 0, width, height }
photo.format() # what it was decoded from, or nil
photo.is_opaque() # walks the alpha channel
photo.info() # all of the above, plus a pixel count
Beyond the shape, four methods read what is actually in the image.
histogram() counts how many pixels hold each value, as four
256-entry lists: red, green, blue and luma. Fully transparent
pixels are skipped, since their colour is not visible. A histogram is
what tells you an image is underexposed (everything bunched at the low
end), flat (bunched in the middle), or clipped (a spike at 0 or 255).
auto_levels() is the automatic form of levels(): it finds the
black and white points from the brightness histogram and stretches the
range between them, ignoring a small fraction at each end so a handful
of stray pixels cannot decide the result.
photo.auto_levels()
photo.auto_levels({ clip: 0.02, gamma: 1.1 })
average_color() returns the mean colour, weighting each pixel by
its alpha so a mostly transparent image reports the colour of the part
you can see.
dominant_colors() returns the colours occupying the most of the
image, most common first. The image is shrunk and reduced to a small
palette first, so it costs about the same whatever the original size.
var accent = photo.dominant_colors(1)[0]
page.fill(accent.darken(40)) # a placeholder while the photo loads
difference() scores how far two images are apart, from 0
(identical) to 1, which makes it usable as a rendering assertion:
if rendered.difference(expected) > 0.01 {
raise Exception('the rendering changed')
}
Animation
import imagine { Animation }
var animation = Animation.open('loading.gif')
animation.length() # frame count
animation.duration() # milliseconds for one pass
animation.frame(0) # one frame, as an Image
animation.frames() # the live list
Frames come back already composited against whatever preceded them, so frame 5 can be used on its own without replaying the first four. The disposal and transparency rules an animated GIF is built from never surface.
Building and transforming:
var frames = []
iter var i = 0; i < 30; i++ {
var frame = Image(200, 200, 'black')
frame.fill_circle(100, 100, i * 3, '#38bdf8')
frames.append(frame)
}
Animation(frames, 40, 0).save('pulse.gif') # 40ms per frame, loops forever
animation
.map(@(frame) {
return frame.grayscale().blur(1)
})
.repeat(3)
.save('out.gif')
map() copies each frame before the function sees it, so a filter chain
works directly as the body and the original animation is left alone.
That copy is why map() costs as much memory again as the animation
itself; to filter in place, walk frames() and change each image
directly.
Animated GIF and animated WebP can both be read, and only GIF can be written. A still image decodes as a one-frame animation rather than an error, so code handling both does not need to branch.
Every frame must be the same size, since animated formats have one canvas that each frame paints into.
Working With Pixels Directly
pixels() hands back the live buffer. It is 8-bit RGBA with straight
(not premultiplied) alpha, laid out row by row with no padding, so pixel
(x, y) begins at byte (y * width + x) * 4.
var buffer = image.pixels()
var total = buffer.length()
iter var at = 0; at < total; at += 4 {
buffer[at] = 255 - buffer[at] # invert red only
}
This is deliberate, and it is much faster than a method call per pixel, so an operation this module does not provide can still be written efficiently in Zuri. Annotating a function’s parameters pays off here:
def darken_edges(pixels: bytes, width: number, height: number) {
iter var y = 0; y < height; y++ {
iter var x = 0; x < width; x++ {
var at = (y * width + x) * 4
# ...
}
}
}
set_pixels() takes a buffer back, which is how you save a copy before
a destructive filter and restore it afterwards:
var saved = image.pixels().clone()
image.blur(8)
image.set_pixels(saved) # back where it started
Image.from_pixels() builds an image around a buffer that came from
somewhere else entirely.
Errors
Failures fall into two groups. An argument of the wrong type raises
TypeError, from the parameter’s own type declaration. Everything else
descends from ImageError, so one catch covers it:
import imagine { Image, ImageError, DecodeError }
catch {
var photo = Image.open(path)
} as e {
if instance_of(e, DecodeError) {
echo 'not a readable image'
} else {
echo 'something else went wrong: ${e.message}'
}
}
| Error | Raised when |
|---|---|
ImageError | The base class. Also raised directly for bad arguments. |
DecodeError | The data is not an image, is a format this build cannot read, or is truncated. |
EncodeError | The image cannot be written in the requested format. |
FormatError | A format name or file extension is not one this module knows. |
BoundsError | A rectangle, crop or resize falls outside the image, or a dimension is below 1. |
FontError | A font cannot be parsed, found, or laid out with. |
Two kinds of failure sit outside that hierarchy on purpose.
Wrong argument type raises TypeError. Parameters declare their
types, so the check happens at the boundary and the message names the
parameter:
image.rotate('sideways')
# TypeError: rotate() expects parameter 'degrees' (argument 1)
# to be a number, got string
A malformed colour raises ValueError, because that is what
[[colors]] reports for it and relabelling would lose the distinction
between “not a colour” and “not a string”:
Color.hex('nonsense') # ValueError, from colors
Color.named('chartroose') # ValueError, from colors
Color.hex(42) # TypeError, from the type declaration
A value of the right type but the wrong range is still an
ImageError subclass, since that is a judgement this module makes
rather than a type the runtime can check:
Image(0, 100) # BoundsError, not TypeError
filters.gamma_lut(-1) # ImageError
DecodeError is the one to always be ready for, since anything arriving
from outside the program can raise it.
Drawing operations do not raise BoundsError. A line running off the
edge of the canvas is clipped, which is what every drawing API does; it
is only the operations that must return an image of an exact size that
have no sensible way to continue.
Performance and Memory
Four bytes per pixel, always. A 6000x4000 photograph is 96 MB decoded
however small its file was, and a 100-frame 500x500 animation is 100 MB.
probe() before decoding anything whose size you do not control.
Some rough guidance on what costs what:
- Decoding and encoding dominate almost every pipeline. PNG at
'best'compression is several times slower than at'default'for a few percent of size; AVIF is slower still. - Resizing costs roughly in proportion to the output size, so
shrinking is cheap and enlarging is not.
LANCZOScosts a few timesNEAREST. - Blur is linear in its radius, not quadratic, because it is
separable. A 3x3
convolve()is cheaper thanblur(1), but byblur(5)the Gaussian has won by a wide margin. - Lookup tables and colour matrices are memory-bound and about as
fast as touching every pixel can be. Chain as many as you like, but
combine matrices with
combine_matrices()rather than applying them one at a time — that is both faster and more accurate, since the intermediate result is never rounded back to 8 bits. - Drawing costs in proportion to the area covered, and anti-aliasing samples each pixel row four times over. Turning it off is a real saving on very large fills.
Resize before filtering whenever the result is going to be smaller anyway. Filtering a 24-megapixel photograph and then shrinking it to a thumbnail does the same visual work at forty times the cost.
Recipes
A thumbnail pipeline for uploads
import imagine
import imagine { Image, DecodeError, LANCZOS, TOP }
def make_thumbnail(upload) {
var header = imagine.probe(upload)
if header == nil {
raise DecodeError('not an image')
}
if header.width * header.height > 50000000 {
raise DecodeError('image too large')
}
return Image.decode(upload)
.cover(400, 400, { anchor: TOP, filter: LANCZOS })
.sharpen(0.4)
.encode('webp')
}
A rounded avatar with a border
import imagine { Image }
def avatar(source, size) {
var photo = Image.decode(source).cover(size, size)
var stencil = photo.layer()
stencil.fill_circle(size / 2, size / 2, size / 2 - 2, 'white')
photo.mask(stencil)
photo.circle(size / 2, size / 2, size / 2 - 2, '#e2e8f0', { thickness: 3 })
return photo
}
A social card
import imagine { Image, Font, Color, ALIGN_LEFT }
def card(title, subtitle) {
var image = Image(1200, 630, '#0f172a')
var heading = Font.load('assets/Inter-Bold.ttf', 64)
var body = heading.size(30)
image.fill_rect(0, 0, 1200, 8, '#38bdf8')
var box = image.text_size(title, heading, { align: ALIGN_LEFT })
image.text(80, 200, title, heading, 'white', { align: ALIGN_LEFT })
image.text(80, 200 + box.height + 24, subtitle, body, '#94a3b8')
return image.to_png()
}
Serving a generated image over HTTP
import http
import imagine { Image, Font, CENTER }
import imagine.formats
var server = http.server(8000)
var font = Font.load('assets/Inter.ttf', 20)
server.get('/badge/{label}', @(request, response) {
var image = Image(220, 60, '#1e293b')
image.fill_rounded_rect(0, 0, 220, 60, 8, '#334155')
image.place_text(request.param('label'), font, 'white', CENTER)
response.content_type(formats.mime_for('png'))
response.cache_for(86400)
response.write(image.to_png())
})
server.listen()
Comparing two images
import imagine { BLEND_DIFFERENCE }
def differs(a, b) {
if a.size() != b.size() {
return true
}
var diff = a.clone().draw_image(b, 0, 0, { blend: BLEND_DIFFERENCE })
var pixels = diff.pixels()
var total = pixels.length()
iter var at = 0; at < total; at += 4 {
if pixels[at] > 8 or pixels[at + 1] > 8 or pixels[at + 2] > 8 {
return true
}
}
return false
}
Databases
The sql module is how a Zuri program talks to a relational database.
It is one module rather than one per engine, and that is the whole
design: sql defines what a database adapter has to provide, supplies
everything that is the same whichever engine answers, and picks the
adapter from the connection string.
Three adapters ship with it to support four major database engines — SQLite, PostgreSQL, MySQL and MariaDB. SQLite is a file-based database with no server to run and nothing to configure, which makes it the right choice for an application that ships with its data and for a test suite that wants a real database per run. PostgreSQL, MySQL, and MariaDB are servers, for everything that outgrows a file, and the MySQL adapter is the same adapter that drives MariaDB as well.
Changing from one to another means changing the connection string, and whatever SQL they genuinely spell differently.
- Following Along
- Introduction
- Connecting
- Querying and Fetching
- Parameters
- Rows, Columns and Types
- The CRUD Helpers
- Transactions
- Streaming Large Results
- Prepared Statements
- Connection Pools
- Errors
- Schema Introspection
- Switching Databases
- SQLite Specifics
- PostgreSQL Specifics
- MySQL Specifics
- Writing Your Own Adapter
- What the Module Refuses
- Module Reference
Blocks on this page that list several calls together are reference listings, not programs: they show the shape of each call rather than a sequence to run. Anything presented as a complete program runs as written. The PostgreSQL and MySQL examples need a server, so they are shown rather than run.
Following Along
Most examples below use a small database of posts and their authors. This builds it:
import sql
var db = sql.open('sqlite://guide.db')
db.exec_script("
drop table if exists posts;
drop table if exists authors;
create table authors (
id integer primary key,
name text not null unique
);
create table posts (
id integer primary key,
author_id integer not null references authors(id),
title text not null,
views integer not null default 0,
published boolean not null default 0
);
")
var ada = db.insert('authors', { name: 'Ada Lovelace' })
var grace = db.insert('authors', { name: 'Grace Hopper' })
db.insert_many('posts', [
{ author_id: ada, title: 'On Engines', views: 412, published: true },
{ author_id: ada, title: 'On Looms', views: 87, published: false },
{ author_id: grace, title: 'On Bugs', views: 1290, published: true },
{ author_id: grace, title: 'On Compilers', views: 640, published: true },
])
echo db.count('posts')
db.close()
4
Everything from here on assumes that file exists in the working directory.
Introduction
A database library usually ties a program to one engine. The calls are named after that engine, the placeholders are spelled its way, and the errors are its own numbers, so moving to another means rewriting every call site.
sql puts the parts that are the same in one place. A statement is
written once:
import sql
var db = sql.open('sqlite://guide.db')
for post in db.query('select title, views from posts where published = ?', [true]) {
echo '${post.title}: ${post.views}'
}
db.close()
On Engines: 412
On Bugs: 1290
On Compilers: 640
Point the first line at postgres://localhost/app or
mysql://localhost/app and the rest runs unchanged. The ? becomes
$1 on PostgreSQL because that is what PostgreSQL wants, and stays ?
on MySQL because that is what MySQL wants. An insert asking for its new
id gets a RETURNING clause on PostgreSQL, because PostgreSQL has no
last insert id and the other two do. A duplicate key arrives as
UniqueViolation from all three.
What it does not do is pretend the engines are the same. They disagree
about auto-incrementing keys, about which functions exist, about a
great deal of SQL. sql translates what can be translated and is
explicit about the rest.
Connecting
open() takes a connection string and returns a Connection.
sql.open('sqlite://./app.db') # a file
sql.open(':memory:') # a private database in memory
sql.open('./app.db') # a bare path is SQLite
sql.open('postgres://localhost/app')
sql.open('postgres://alice:secret@db.internal:5432/shop?sslmode=require')
sql.open('host=localhost dbname=app user=alice')
sql.open('mysql://localhost/app')
sql.open('mysql://alice:secret@db.internal:3306/shop?sslmode=verify')
sql.open('mariadb://localhost/app')
The scheme picks the adapter. A string with no scheme at all is taken as a SQLite path, since nothing else it could be.
Options can also be given as a dictionary, which is how anything an adapter accepts beyond the connection string is passed:
sql.open({
driver: 'sqlite',
path: './app.db',
journal_mode: 'wal',
busy_timeout: 10000,
})
or alongside a string, where they are merged over whatever it carried:
sql.open('sqlite://./app.db', { journal_mode: 'wal' })
A connection should be closed when it is finished with, and a catch
block makes that certain even when the work between raises:
import sql
var db = sql.open(':memory:')
catch {
db.exec('create table t (n integer)')
db.exec('insert into t values (?)', [1])
echo db.fetch_value('select n from t')
} as error {
echo 'failed: ${error.message}'
}
db.close()
1
A connection belongs to the isolate that opened it and cannot be handed to another. An isolate that needs the database opens its own connection or its own pool.
Querying and Fetching
query() runs a statement and reads its whole result:
import sql
var db = sql.open('sqlite://guide.db')
var result = db.query('select title, views from posts order by views desc')
echo result.length()
echo result.first().title
echo result.column('title')
db.close()
4
On Bugs
[On Bugs, On Compilers, On Engines, On Looms]
A result is iterable, so the common case needs nothing else:
for post in db.query('select * from posts') {
echo post.title
}
For the shapes that come up constantly there are shorter forms:
db.fetch_one(sql, params) # the first row, or nil
db.fetch_all(sql, params) # every row, as dictionaries
db.fetch_value(sql, params, fallback) # the first column of the first row
db.fetch_column(sql, params, column) # one column's values, as a list
import sql
var db = sql.open('sqlite://guide.db')
echo db.fetch_value('select count(*) from posts')
echo db.fetch_one('select title from posts where views > ?', [1000]).title
echo db.fetch_column('select title from posts order by title limit 2')
echo db.fetch_value('select title from posts where views > ?', [99999], 'none')
db.close()
4
On Bugs
[On Bugs, On Compilers]
none
exec() is for a statement run for its effect rather than its rows:
import sql
var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, n integer)')
var result = db.exec('insert into t (n) values (?)', [7])
echo result.rows_affected
echo result.last_insert_id
db.close()
1
1
And exec_script() runs several statements at once, which is what a
schema file is:
db.exec_script(file('schema.sql').read())
Parameters
Values are bound by the engine, never pasted into the statement. That is what makes a value a value rather than a piece of SQL, and it is the whole of the defence against injection.
Write ? for positional parameters:
import sql
var db = sql.open('sqlite://guide.db')
echo db.fetch_column(
'select title from posts where author_id = ? and published = ?',
[1, true]
)
db.close()
[On Engines]
or :name for named ones, where order stops mattering and a name may
be used more than once:
import sql
var db = sql.open('sqlite://guide.db')
echo db.fetch_column(
'select title from posts where views > :floor and views < :ceiling',
{ floor: 100, ceiling: 1000 }
)
db.close()
[On Engines, On Compilers]
sql rewrites these into whatever the adapter wants: $1 and $2 for
PostgreSQL, ?1 and ?2 for SQLite, and ? unchanged for MySQL, which
already spells them that way. The scanner knows when a ? or a : is
not a placeholder, so none of these are touched:
"a ? inside a string" -- a string
"a ? in an identifier" -- a quoted identifier
-- a ? in a line comment
/* a ? in a block comment */
$tag$ a ? in a dollar-quoted body $tag$
value::text -- a PostgreSQL cast
arr[1:3] -- an array slice
PostgreSQL uses ? as a JSON operator, so a literal one is written
??:
db.query('select * from docs where data ?? ?', ['key'])
Never build a statement by joining strings around a value, even when the value looks safe. The two forms above cover every case where a value varies. Where an identifier varies, quote it through the driver rather than interpolating it:
var column = db.driver().quote_identifier(name) db.query('select ${column} from posts')
Rows, Columns and Types
A row is a dictionary keyed by column name, which is what lets
post.title read the way it does.
import sql
var db = sql.open('sqlite://guide.db')
var result = db.query('select id, title, views from posts order by id limit 1')
echo result.first()
echo result.columns
echo result.tuples()
db.close()
{id: 1, title: On Engines, views: 412}
[{name: id, type: integer}, {name: title, type: text}, {name: views, type: integer}]
[[1, On Engines, 412]]
Two columns of the same name collapse in a dictionary, which is worth
knowing before selecting id from both sides of a join. Aliasing one
of them is the fix; tuples() is the escape hatch when it cannot be.
Values correspond like this:
| Zuri | Database |
|---|---|
nil | NULL |
bool | a boolean where the engine has one, otherwise 0 and 1 |
number | an integer where the value is whole, otherwise a float |
bigint | a 64 bit integer, for values past what a double holds |
string | text |
bytes | a blob |
date.Date | a timestamp |
sql.Time | a TIME interval, on an engine that has one |
list, dict | JSON |
Not every engine has every one of those. PostgreSQL has a real boolean
and MySQL does not, so a true written to MySQL comes back as 1.
db.supports() answers that sort of question without guessing, and
each engine’s own section below says where it differs.
An integer too large for a double comes back as a bigint rather than
silently rounded:
import sql
var db = sql.open(':memory:')
db.exec('create table t (n integer)')
db.exec('insert into t values (9223372036854775807)')
var n = db.fetch_value('select n from t')
echo typeof(n)
echo n
db.close()
bigint
9223372036854775807n
SQLite stores five things and remembers nothing about intent, so a boolean goes in as 1 and would come back as the number 1. What it does keep is the type each column was declared with, and that is what the adapter reads it back by:
import sql
var db = sql.open(':memory:')
db.exec('create table t (ok boolean, at datetime, doc json)')
db.exec('insert into t values (?, ?, ?)', [
true, '2026-09-16T14:30:00.000000+00:00', '{"a":1}',
])
var row = db.fetch_one('select * from t')
echo typeof(row.ok)
echo typeof(row.at)
echo row.doc
db.close()
bool
Date
{a: 1}
A column with no declared type is an expression, and comes back exactly as it was stored.
Exact decimals
A number is a double, which holds 0.1 only approximately. For money
that is not good enough, so sql.Decimal holds a value exactly:
import sql { Decimal }
var price = Decimal('19.99')
var tax = price.multiply(Decimal('0.20'))
echo price.to_string()
echo tax.to_string()
echo price.add(tax).to_string()
echo Decimal('0.1').add(Decimal('0.2')).to_string()
19.99
3.9980
23.9880
0.3
PostgreSQL’s numeric and MySQL’s DECIMAL columns read and write as
Decimal automatically. SQLite has no exact decimal type, so store one
as text or as an integer count of the smallest unit.
The CRUD Helpers
Four calls on a connection write no SQL at all. insert(), update(),
delete() and find() take a table name and dictionaries, build the
statement for whichever engine is on the other end, and bind every
value as a parameter.
They are here because the statements they replace are the ones least worth writing by hand. They are mechanical, they differ between dialects in small ways that only show up in production, and a string built by concatenation is where an injection gets in. Anything harder than a flat list of conditions is still written as SQL, and the two mix freely on the same connection.
Inserting and ids
insert() takes a table and a dictionary, and returns the new row’s
id:
import sql
var db = sql.open(':memory:')
db.exec('create table posts (id integer primary key, title text)')
echo db.insert('posts', { title: 'Hello' })
echo db.insert('posts', { title: 'World' })
db.close()
1
2
This is the call that hides the largest difference between the engines.
SQLite and MySQL report the id of the row just inserted; PostgreSQL does
not, and an insert that wants one has to ask with a RETURNING clause
naming the primary key. MariaDB has both and uses RETURNING.
insert() does whichever applies, and finds the key by asking the
schema rather than assuming it is called id.
Where the key is not the column to return, name it:
db.insert('events', { name: 'started' }, { returning: 'uuid' })
Several rows at once
Several rows go in one statement, split so that no single statement binds more parameters than the engine allows:
import sql
var db = sql.open(':memory:')
db.exec('create table points (x integer, y integer)')
echo db.insert_many('points', [
{ x: 1, y: 2 },
{ x: 3, y: 4 },
])
db.close()
2
Every row has to name the same columns. A row naming a different set raises rather than being padded with nulls, because a missing column and a null column mean different things.
Finding rows
find() answers with a ResultSet, count() with a number, and
find_one() with the first row as a dictionary:
import sql
var db = sql.open('sqlite://guide.db')
echo db.count('posts')
echo db.count('posts', { published: true })
echo db.count('posts', { author_id: [1, 2] })
echo db.count('posts', { author_id: [] })
echo db.find('posts', { published: true }, {
columns: ['title'],
order: ['views desc'],
limit: 2,
}).column('title')
echo db.find_one('posts', { title: 'On Bugs' }).views
db.close()
4
3
4
0
[On Bugs, On Compilers]
1290
A find_one() that matches nothing answers nil. That is the answer,
not a failure, so nothing is raised for it:
import sql
var db = sql.open('sqlite://guide.db')
echo db.find_one('posts', { title: 'On Bugs' }).views
echo db.find_one('posts', { title: 'Never Written' })
db.close()
1290
nil
What a filter can say
A condition’s value decides what it means. A plain value is equality, a
list is IN, an empty list matches nothing, and nil is IS NULL
rather than = NULL, which no row ever satisfies.
Several conditions in one dictionary are joined with AND. There is no
OR, which is the first thing to write as SQL:
import sql
var db = sql.open('sqlite://guide.db')
echo db.count('posts', { published: true, author_id: 1 })
echo db.count('posts', { published: true, views: [87, 1290] })
db.close()
1
1
Passing nil as the whole filter matches every row. That has to be
asked for rather than happening because a dictionary came out empty,
which is what stands between a filter built from user input and an
UPDATE with no WHERE.
Ordering, limits and pages
order is a list of column names, each optionally followed by asc
or desc, applied in the order given:
import sql
var db = sql.open('sqlite://guide.db')
echo db.find('posts', nil, { order: ['views desc'] }).column('title')
echo db.find('posts', nil, { order: ['author_id asc', 'views desc'] }).column('title')
db.close()
[On Bugs, On Compilers, On Engines, On Looms]
[On Engines, On Looms, On Bugs, On Compilers]
limit and offset are whole numbers of rows, and together they are a
page. A page past the end is empty rather than an error:
import sql
var db = sql.open('sqlite://guide.db')
def page(number, size) {
return db.find('posts', nil, {
columns: ['title'],
order: ['views desc'],
limit: size,
offset: (number - 1) * size,
}).column('title')
}
echo page(1, 2)
echo page(2, 2)
echo page(3, 2)
db.close()
[On Bugs, On Compilers]
[On Engines, On Looms]
[]
A limit and an offset go into the statement text rather than being bound, because not every engine allows a parameter in either place. That leaves them as the one part of a built statement that is not a parameter, so each is checked before it goes in:
import sql
var db = sql.open('sqlite://guide.db')
catch {
db.find('posts', nil, { limit: 2.5 })
} as error {
echo error.message
}
db.close()
a limit is a whole number of rows, not 2.5
Updating and deleting
update() and delete() return how many rows changed:
import sql
var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, n integer)')
db.insert_many('t', [{ n: 1 }, { n: 2 }, { n: 3 }])
echo db.update('t', { n: 0 }, { n: [1, 2] })
echo db.delete('t', { n: 0 })
echo db.count('t')
db.close()
2
2
1
They read a filter exactly as find() does, nil included: passing it
changes or deletes every row.
Values that are not values
Where a value is not a value, sql.raw() marks a fragment to be used
as written:
import sql
var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, views integer)')
db.insert('t', { views: 10 })
db.update('t', { views: sql.raw('views + 1') }, { id: 1 })
echo db.fetch_value('select views from t')
db.close()
11
A Raw is accepted everywhere a column or a value is, which is what
makes an aggregate or a subquery reachable without leaving the
helpers:
import sql
var db = sql.open('sqlite://guide.db')
echo db.find('posts', nil, {
columns: [sql.raw('count(*) as n'), sql.raw('sum(views) as total')],
}).first()
echo db.count('posts', {
author_id: sql.raw("(select id from authors where name = 'Ada Lovelace')"),
})
db.close()
{n: 4, total: 2429}
2
raw()is exactly as dangerous as it sounds. A fragment built from anything a user supplied is a SQL injection. Build them from literals, and keep values in the parameters where they belong.
Inside a transaction
A Transaction carries all of them, and they mean the same thing
there. A read inside one sees what that transaction has written and
nobody else has committed yet:
import sql
var db = sql.open(':memory:')
db.exec('create table t (id integer primary key, n integer)')
db.transaction(@(tx) {
tx.insert_many('t', [{ n: 1 }, { n: 2 }, { n: 3 }])
echo tx.count('t')
echo tx.find('t', nil, { order: ['n desc'] }).column('n')
tx.update('t', { n: 9 }, { n: 1 })
tx.delete('t', { n: 3 })
})
echo db.find('t', nil, { order: ['n asc'] }).column('n')
db.close()
3
[3, 2, 1]
[2, 9]
A rollback takes those rows with it, the reads included, and a transaction that has already committed refuses them the way it refuses everything else.
Where they stop
These calls cover a flat list of equality conditions and stop there.
There is no join, no OR and no expression tree, because SQL is a
better language for those than any chain of method calls. A query that
wants one is a query() or a fetch_all() on the same connection,
next to the helpers rather than instead of them.
Transactions
The closure form commits when the body returns and rolls back when it raises, and there is no path through it that leaves a transaction open:
import sql
var db = sql.open(':memory:')
db.exec('create table accounts (id integer primary key, balance integer)')
db.insert_many('accounts', [{ balance: 500 }, { balance: 0 }])
db.transaction(@(tx) {
tx.update('accounts', { balance: sql.raw('balance - 100') }, { id: 1 })
tx.update('accounts', { balance: sql.raw('balance + 100') }, { id: 2 })
})
echo db.fetch_column('select balance from accounts order by id')
db.close()
[400, 100]
A failure undoes the whole thing:
import sql
var db = sql.open(':memory:')
db.exec('create table t (n integer)')
catch {
db.transaction(@(tx) {
tx.exec('insert into t values (1)')
raise Error('something went wrong')
})
} as error {
echo 'rolled back: ${error.message}'
}
echo db.count('t')
db.close()
rolled back: something went wrong
0
A transaction() called inside another becomes a savepoint, so a
helper that opens one works the same whether it was called on its own
or from inside a larger piece of work. Only the outermost commits, and
a failure inside undoes just that inner piece:
import sql
var db = sql.open(':memory:')
db.exec('create table t (name text)')
db.transaction(@(outer) {
outer.insert('t', { name: 'outer' })
catch {
db.transaction(@(inner) {
inner.insert('t', { name: 'inner' })
raise Error('inner failed')
})
} as _error {
echo 'inner undone'
}
})
echo db.fetch_column('select name from t')
db.close()
inner undone
[outer]
Schema changes are part of the transaction on SQLite and PostgreSQL,
so a set of create table statements that fails half way leaves
nothing behind. MySQL and MariaDB commit the open transaction at every
create, alter and drop, and a transaction() whose body changes
the schema there raises TransactionError when it comes to commit.
db.supports('transactional_ddl') tells the two apart, so a program
that migrates its own schema can run each change inside a transaction
where that holds and on its own where it does not.
An isolation level can be asked for. Where an engine cannot provide one it says so rather than quietly giving something weaker:
db.transaction(@(tx) { ... }, sql.SERIALIZABLE)
sql.READ_UNCOMMITTED | MySQL honours it; PostgreSQL accepts it and gives read committed |
sql.READ_COMMITTED | everywhere |
sql.REPEATABLE_READ | everywhere, and MySQL’s default |
sql.SERIALIZABLE | everywhere |
begin() opens one to be committed or rolled back by hand, for a
transaction whose lifetime is not a block. The closure form is safer
and should be preferred.
Streaming Large Results
query() builds the whole result in memory, which is right for the
hundreds of rows most queries return and wrong for the millions some
do. stream() reads the same result a row at a time:
import sql
var db = sql.open('sqlite://guide.db')
var cursor = db.stream('select title from posts order by title')
for post in cursor {
echo post.title
}
db.close()
On Bugs
On Compilers
On Engines
On Looms
Running to the end closes the cursor. A loop that stops early does not,
so anything that might break has to close it:
var cursor = db.stream('select * from events')
for event in cursor {
if done(event) {
break
}
}
cursor.close()
take(n) reads a batch at a time, for work that batches naturally:
import sql
var db = sql.open('sqlite://guide.db')
var cursor = db.stream('select title from posts order by title')
echo cursor.take(2).map(@(post) => post.title)
echo cursor.take(2).map(@(post) => post.title)
echo cursor.take(2)
db.close()
[On Bugs, On Compilers]
[On Engines, On Looms]
[]
On PostgreSQL this is a server-side portal and on MySQL a server-side
cursor, both fetched in batches that { batch: n } sizes. On SQLite it
is the engine’s own behaviour: rows are computed as they are asked for,
so nothing is held but the current one.
A server-side cursor holds its connection while it is open, since the server is part way through answering. Running another statement on that connection before the cursor is read to the end or closed raises rather than letting the exchange fall out of step.
Prepared Statements
Compiling a statement is the expensive half of running one, and a statement run in a loop should pay for it once:
import sql
var db = sql.open(':memory:')
db.exec('create table points (x integer, y integer)')
var insert = db.prepare('insert into points (x, y) values (?, ?)')
for i in 0..(1000) {
insert.exec([i, i * 2])
}
insert.close()
echo db.count('points')
db.close()
1000
The placeholders are translated once, when the statement is prepared, so the loop does no string work at all. A prepared statement carries the same query and fetch methods a connection does.
Connection Pools
Opening a connection is expensive: for PostgreSQL and MySQL it is a TCP connection, a TLS handshake and an authentication exchange before the first statement runs. A server that opened one per request would spend most of its time connecting.
import sql
var db = sql.pool('sqlite://guide.db', { max: 4 })
echo db.fetch_value('select count(*) from posts')
echo db.stats()
db.close()
4
{size: 1, idle: 1, in_use: 0, max: 4}
Used that way the pool takes a connection, runs the statement and gives
it back. Where several statements have to run on the same connection,
which a transaction requires, with_connection() holds one and
releases it whatever happens:
db.with_connection(@(connection) {
connection.transaction(@(tx) {
tx.exec('...')
})
})
transaction() on the pool does the same in one call.
| Setting | |
|---|---|
max | most connections to open. Ten by default |
min | how many to open up front. None by default |
idle_timeout | how long an idle connection is kept |
max_lifetime | how long any connection is kept before being replaced |
validate_on_acquire | check a connection is alive before lending it |
on_connect | run something on each new connection |
A pool wants to be as small as the work allows. Connections are not free at the other end either, and a pool larger than the database can usefully serve turns a queue in the application into a queue in the database, where it is harder to see.
A pool belongs to the isolate that made it, which has two consequences. An isolate that needs database access makes its own. And there is no waiting when a pool is empty: nothing else can release a connection while the call is running, so running out is reported rather than waited on.
Errors
Every error is a sql.SqlError. Catch that to catch anything a
database can do; the subclasses separate what is worth handling
differently.
Error
└── SqlError
├── ConnectionError
│ ├── AuthenticationError
│ ├── TimeoutError
│ └── ProtocolError
├── QueryError
├── IntegrityError
│ ├── UniqueViolation
│ ├── ForeignKeyViolation
│ ├── NotNullViolation
│ └── CheckViolation
├── TransactionError
│ ├── SerializationError
│ └── DeadlockError
├── PoolError
│ └── PoolExhaustedError
├── NotSupportedError
└── ClosedError
The classes are the same on every engine, which is the point. A unique
constraint is SQLSTATE 23505 on PostgreSQL, error number 1062 on
MySQL, and extended result code 2067 on SQLite; no one of those means
anything to the others, and code matching on any of them would stop
working the moment the adapter changed.
import sql
var db = sql.open(':memory:')
db.exec('create table users (id integer primary key, email text unique)')
db.insert('users', { email: 'ada@example.com' })
catch {
db.insert('users', { email: 'ada@example.com' })
} as error {
echo instance_of(error, sql.UniqueViolation)
echo instance_of(error, sql.IntegrityError)
echo instance_of(error, sql.SqlError)
echo error.driver
}
db.close()
true
true
true
sqlite
The engine’s own account is kept rather than thrown away. code holds
what the engine said, sqlstate the five character code where there is
one, and query the statement that failed:
import sql
var db = sql.open(':memory:')
db.exec('create table users (id integer primary key, email text unique)')
db.insert('users', { email: 'ada@example.com' })
catch {
db.insert('users', { email: 'ada@example.com' })
} as error {
echo error.code
echo error.sqlstate
echo error.type
}
db.close()
2067
nil
UniqueViolation
SerializationError and DeadlockError are the two worth retrying:
both mean the engine gave up on a transaction to keep its promises, and
running the whole transaction again usually succeeds.
Schema Introspection
db.schema answers questions about what is in the database. The
queries underneath are entirely different per engine, and the answers
are the same shape.
Every method takes an optional schema to look in. PostgreSQL has
schemas inside a database and resolves the default through
search_path; MySQL calls a database a schema and has no layer above
it, so naming one there names a database; SQLite has neither and
ignores the argument.
import sql
var db = sql.open('sqlite://guide.db')
echo db.schema.tables()
echo db.schema.column_names('posts')
echo db.schema.primary_key('posts')
echo db.schema.has_column('posts', 'views')
db.close()
[authors, posts]
[id, author_id, title, views, published]
id
true
Each column is described the same way whichever engine answered:
import sql
var db = sql.open('sqlite://guide.db')
for column in db.schema.columns('posts') {
echo '${column.name} ${column.type} ${column.nullable ? "null" : "not null"}'
}
db.close()
id integer null
author_id integer not null
title text not null
views integer not null
published boolean not null
indexes() and foreign_keys() describe the rest:
import sql
var db = sql.open('sqlite://guide.db')
echo db.schema.foreign_keys('posts')
db.close()
[{columns: [author_id], references_table: authors, references_columns: [id], on_update: NO ACTION, on_delete: NO ACTION}]
This is also what makes insert() portable: on an engine with no last
insert id the insert needs a RETURNING clause, and the column to
return is the table’s primary key, which is asked for here.
One caution when moving a schema between engines: MySQL parses a
column level references clause and then ignores it, so a foreign key
declared that way exists on the other two and not there. Declared at
table level it exists on all three.
Switching Databases
The claim this module makes is that a program moves between engines by changing its connection string. Here is that claim, as a program:
import sql
def report(db) {
db.exec_script('drop table if exists tally')
db.exec_script(
'create table tally (id ' + key_type(db) + ' primary key, word text not null)'
)
db.insert_many('tally', [{ word: 'alpha' }, { word: 'beta' }])
var id = db.insert('tally', { word: 'gamma' })
var found = db.fetch_column('select word from tally order by word')
db.exec_script('drop table tally')
return '${db.driver_name()}: ${found} (last id ${id})'
}
def key_type(db) {
using db.driver_name() {
when 'postgres' return 'serial'
when 'mysql' return 'integer auto_increment'
when 'mariadb' return 'integer auto_increment'
}
return 'integer'
}
var sqlite = sql.open(':memory:')
echo report(sqlite)
sqlite.close()
sqlite: [alpha, beta, gamma] (last id 3)
The same report() runs against PostgreSQL or MySQL by opening a
different connection, and prints the same list.
Three things did have to be written twice, and they are the three that are genuinely different:
- The connection string, which is the point.
- An auto-incrementing key, which is not standard SQL. PostgreSQL
spells it
serial, MySQLauto_increment, SQLiteinteger primary key. - Any SQL only one of them has.
sqldoes not parse or rewrite statements beyond their placeholders, so a PostgreSQL array operator, a MySQLon duplicate key updateor a SQLitejson_extractstays what it is.
db.supports() is how a program asks rather than assumes:
import sql
var db = sql.open(':memory:')
echo db.driver_name()
echo db.supports('returning')
echo db.supports('arrays')
echo db.supports('concurrent_writers')
echo db.capabilities().placeholder_style
db.close()
sqlite
true
false
false
indexed
SQLite Specifics
db.native() reaches the adapter underneath, where the engine’s own
features live.
A SQLite database is a file, and :memory: is one that is not. An
in-memory database belongs to the connection that opened it, so a
second connection to :memory: is a second, empty database rather than
another handle on the same data.
Foreign keys are enforced. SQLite leaves them unenforced by default,
per connection, for compatibility with databases written before it had
them; a database that declares them almost certainly means them, so the
adapter turns them on. Pass foreign_keys: false to turn them back
off.
Functions written in Zuri
A function registered on a connection runs inside the engine, once per row, and can be used anywhere an expression can:
import sql
var db = sql.open('sqlite://guide.db')
db.native().create_function('initials', 1, @(name) {
return ''.join(name.split(' ').map(@(part) => part[0, 1]))
})
echo db.fetch_column('select initials(name) from authors order by name')
db.close()
[AL, GH]
An aggregate is written as a fold: step is called once per row with
the accumulator and the row’s arguments and returns the next
accumulator, and finish turns the last accumulator into the group’s
value.
import sql
var db = sql.open('sqlite://guide.db')
db.native().create_aggregate('longest', 1, @(longest, title) {
if longest == nil or title.length() > longest.length() {
return title
}
return longest
}, @(longest) => longest)
echo db.fetch_value('select longest(title) from posts')
db.close()
On Compilers
A collation is an ordering for text, reached with collate:
import sql
var db = sql.open('sqlite://guide.db')
db.native().create_collation('bylength', @(a, b) => a.length() - b.length())
echo db.fetch_column('select title from posts order by title collate bylength')
db.close()
[On Bugs, On Looms, On Engines, On Compilers]
A function must not touch the connection it was registered on. SQLite is in the middle of a statement when it calls, and reentering would deadlock. An error raised inside one is carried out and re-raised with its own class intact once the statement has finished.
Blobs, backups and hooks
A large value read with a select arrives all at once. A blob handle
opens one cell and reads windows of it:
var blob = db.native().blob('main', 'files', 'content', id, false)
var at = 0
while at < blob.length() {
out.write(blob.read(at, 65536.min(blob.length() - at)))
at += 65536
}
blob.close()
The cell has to exist and already be the right size: SQLite cannot grow
a blob this way, which is what inserting zeroblob(n) is for.
Copying the file with the filesystem is only safe when nothing is writing. SQLite’s own backup copies page by page and notices when the source changes underneath:
db.native().backup_to('./snapshot.db')
And the hooks report what the engine is doing:
db.native().on_change(@(operation, database, table, rowid) { ... })
db.native().on_commit(@() => true)
db.native().on_rollback(@() { ... })
db.native().set_authorizer(@(action, first, second, database, trigger) { ... })
db.native().on_progress(1000, @() => keep_going())
PostgreSQL Specifics
The adapter speaks version 3 of the wire protocol, over TCP or over a unix domain socket, optionally under TLS.
A host beginning with a slash is a socket directory, which is how
libpq spells it, and the socket inside is named after the port:
/var/run/postgresql with port 5432 means
/var/run/postgresql/.s.PGSQL.5432. A path that already names the
socket is taken as given. TLS is neither offered nor wanted over a
socket, since nothing sits in between, so sslmode is ignored there.
sql.open('postgres:///app?host=/var/run/postgresql')
sql.open({ driver: 'postgres', host: '/var/run/postgresql', database: 'app' })
sslmode chooses how TLS is used: disable never, prefer when the
server offers it, and require always, failing when the server
refuses. Only require verifies who is on the other end.
Statements go through the extended protocol, which compiles them on the server and binds values rather than substituting them. Values travel in the server’s own binary form wherever the adapter has a codec for the type, and as text otherwise, so a type it has never heard of still arrives readable rather than as bytes.
numeric columns arrive as sql.Decimal, arrays as lists, jsonb as
dictionaries, and timestamps as date.Date.
Listening for notifications
LISTEN and NOTIFY are PostgreSQL’s own publish and subscribe. A
connection that has run LISTEN channel receives a message whenever
anything anywhere runs NOTIFY channel.
var listener = db.native().listener()
listener.listen('jobs')
while true {
for message in listener.wait(nil) {
handle(message.channel, message.payload)
}
}
Notifications arrive between other messages, so a connection busy
running statements collects them as it goes and poll() hands over
whatever has accumulated.
Teaching it a type
A database that defines its own types can teach the adapter about them:
import sql.postgres { DEFAULT_REGISTRY }
DEFAULT_REGISTRY.register(oid, decoder, encoder)
Until it is taught, such a column arrives as text, which is correct if unexciting.
MySQL Specifics
The adapter speaks the client/server protocol directly, over TCP or over a unix domain socket, optionally under TLS, with no client library underneath it. MariaDB speaks the same protocol and the same adapter drives it.
A socket option names a socket, and so does a host beginning with a
slash. Unlike PostgreSQL the path names the socket itself rather than
the directory holding it:
sql.open({ driver: 'mysql', socket: '/var/run/mysqld/mysqld.sock', user: 'app' })
sql.open('socket=/var/run/mysqld/mysqld.sock user=app database=shop')
sql.open('mysql://app@localhost/shop?socket=/var/run/mysqld/mysqld.sock')
TLS is neither offered nor wanted there, since nothing sits in between.
The connection does count as private, which is what lets
caching_sha2_password and sha256_password send the password itself
rather than encrypting it to the server’s public key.
A statement with no values is sent as text, which is one round trip. A statement with values is prepared, so the values travel in the server’s binary encoding instead of being written into the SQL, and the question of quoting never arises. Compiled statements are kept and reused, so a query run in a loop is compiled once.
import sql
var mysql = sql.driver('mysql')
echo mysql.capabilities().placeholder_style
echo mysql.capabilities().last_insert_id
echo mysql.capabilities().returning
echo mysql.quote_identifier('order by')
echo sql.driver('mariadb').capabilities().returning
question
true
false
`order by`
true
MySQL keeps ? as it is written, and reports the id of an inserted row
rather than returning it, so insert() reads the reported id instead of
adding a RETURNING clause. MariaDB differs in two ways that change
what the layer above generates, which is why mariadb:// is a scheme
of its own: it has RETURNING, and it has no JSON type.
What arrives from where
DECIMAL columns arrive as sql.Decimal, JSON as lists and
dictionaries, DATE and DATETIME and TIMESTAMP as date.Date, and
TIME as sql.Time.
MySQL has no boolean. BOOLEAN is another name for TINYINT(1), so a
value written as true comes back as 1, and db.supports('booleans')
is false to say so.
BLOB and TEXT share a type on the wire and are told apart only by
the column’s collation, which the adapter reads: a binary column
arrives as bytes and a text one as a string. BIT and GEOMETRY
arrive as bytes, having no shape Zuri could represent without
inventing one. SET arrives as the comma separated text the server
stores rather than as a list, because a list written back would be
encoded as JSON and the column would quietly stop matching.
A date MySQL considers absent is written as zeros, and 0000-00-00 is
not a date any calendar has. It arrives as nil.
A TIME is a span, not a clock
A TIME column runs from -838:59:59.999999 to 838:59:59.999999. It
is a duration rather than a point in the day, so it can be negative and
can exceed twenty four hours, and neither of those would survive being
read into a date.Date. sql.Time holds it:
import sql
var shift = sql.time(9, 30, 0)
var late = sql.time(0, 0, 30, 0, true)
echo shift.to_string()
echo shift.total_seconds()
echo late.to_string()
echo late.total_seconds()
# Two spans of the same length are equal however each was built.
echo sql.time(1, 30).equals(sql.time(0, 90))
# And the parts are kept as they were written.
echo sql.time(0, 90).minutes
echo sql.parse_time('-838:59:59.999999').to_string()
echo sql.time_from_seconds(-5400).to_string()
09:30:00
34200
-00:00:30
-30
true
90
-838:59:59.999999
-01:30:00
Time zones
A DATETIME carries no zone, and a TIMESTAMP is converted to and
from whatever zone the session is in. So the session’s zone decides
what a timestamp means, and leaving it to the server’s configuration
would make the same database read differently from two machines.
The adapter sets the session to UTC when it connects. Every timestamp
then arrives as UTC, and a date.Date written back is converted from
whatever offset it carries. To leave the server’s own setting alone,
pass time_zone as nil, or name a zone to use instead:
sql.open('mysql://localhost/app', { time_zone: nil })
sql.open('mysql://localhost/app', { time_zone: '+01:00' })
Authentication
caching_sha2_password, which is the default from MySQL 8.0 onwards,
mysql_native_password, sha256_password, mysql_clear_password, and
MariaDB’s client_ed25519.
Two of those need the password itself rather than a proof of it, the
first time an account authenticates. Over TLS the password is sent as
it is. Over a plain connection it is encrypted to a public key the
server hands over first, so it is never readable in transit.
mysql_clear_password has no such fallback and is refused outright on
a connection that is not private, since it sends the password with
nothing protecting it.
TLS
sslmode chooses how TLS is used:
| Mode | Meaning |
|---|---|
disable | Never. |
prefer | When the server offers it, without checking who the server is. The default. |
require | Always, without checking who the server is. |
verify | Always, checking the server’s certificate and host name. |
MySQL’s own spellings are accepted too. DISABLED, PREFERRED and
REQUIRED mean what they say, and both VERIFY_CA and
VERIFY_IDENTITY become verify. That makes VERIFY_CA stricter here
than on the command line, where it checks the certificate but not the
name: the difference can only cause a connection to be refused, never
one to be wrongly trusted.
ssl_ca names a certificate authority to trust, and ssl_cert with
ssl_key present a client certificate.
Compression
The protocol can compress everything after the handshake:
sql.open('mysql://localhost/app', { compression: 'zlib' })
sql.open('mysql://localhost/app', { compression: 'zstd' })
It is worth it for large results over a slow link and costs more than it saves on a local socket, so it is off unless asked for. A packet too small to benefit is sent uncompressed regardless, which the protocol allows for.
Sending a file
LOAD DATA LOCAL INFILE has the server name a file and the client send
it. The file is read from the machine the program runs on, chosen by
the server, so a hostile or compromised server could ask for anything
the process can read. It is refused unless a reader is supplied:
import os
var db = sql.open('mysql://localhost/app', {
local_infile: @(name) {
return os.file(name).read()
},
})
The reader decides what may be sent, which is where that decision belongs. Refusing still answers the server, so the connection stays usable afterwards rather than waiting for a file that will never come.
Statements that answer more than once
A stored procedure sends one result set per select inside it, then an
OK packet. query() hands back the first and reads the rest, because
leaving them unread would not lose them: it would give them to the next
statement, which would then answer with this one’s rows.
var first = db.query('call two_results()')
A routine’s body has semicolons in it, and exec_script() splits a
script on semicolons. So a create procedure goes through exec() as
a single statement rather than through exec_script().
Reaching the adapter
db.native().flavor() # 'mysql' or 'mariadb'
db.native().connection_id() # what KILL names
db.native().parameters() # what the server said about itself
db.native().reset() # back to a fresh session
Writing Your Own Adapter
An adapter for another engine implements four things: a Driver that
reads a connection string and opens a connection, a DriverConnection
that runs statements, a DriverStatement, and a DriverCursor.
import sql
class DuckDbDriver < sql.Driver {
name() { return 'duckdb' }
schemes() { return ['duckdb'] }
capabilities() {
var caps = sql.default_capabilities()
caps.set('placeholder_style', sql.QUESTION)
caps.set('returning', true)
return caps
}
parse_dsn(dsn) { ... }
connect(options) { ... }
}
sql.register(DuckDbDriver())
var db = sql.open('duckdb://./app.duckdb')
Everything above the adapter comes for free: pooling, transactions and savepoints, the CRUD helpers, placeholder translation, cursors, and the error hierarchy. What the adapter supplies is the engine.
The base classes raise NotImplementedError for anything left out, so
an unfinished adapter fails where the gap is. An engine that genuinely
cannot do something raises NotSupportedError instead, which reads
differently on purpose: the first is an unfinished adapter, the second
is an honest limit.
What the Module Refuses
It is not an ORM. There is no model class, no lazy loading and no identity map. Rows are dictionaries.
It does not write your joins. The CRUD helpers cover a flat list of equality conditions. Anything past that is SQL.
It does not migrate schemas. exec_script() runs a schema file and
db.schema reports what is there; deciding what to change and in what
order is a separate problem.
It does not translate SQL. Placeholders are rewritten. Statements are not parsed, and a function only one engine has stays a function only one engine has.
It does not hide a difference by guessing. Where an engine cannot
do what was asked, the answer is NotSupportedError naming the adapter
and the feature, rather than an approximation nobody asked for.
Module Reference
The standard library reference documents every class and method. The shape of the module:
sql.open(dsn, options) | opens a connection |
sql.pool(dsn, options) | opens a pool of them |
sql.register(driver) | adds an adapter |
sql.drivers() | what is registered |
sql.raw(fragment) | marks SQL to be used as written |
sql.Decimal(text) | an exact decimal |
sql.time(h, m, s) | a signed span, as a TIME column holds one |
sql.parse_time(text) | the same, read from text |
On a Connection:
query, exec, exec_script | run a statement |
fetch_one, fetch_all, fetch_value, fetch_column | read a result |
stream, prepare | a cursor, and a compiled statement |
insert, insert_many, update, delete, find, find_one, count | the four statements that are always the same |
transaction, begin, in_transaction | transactions |
schema | introspection |
driver, driver_name, capabilities, supports | what this engine is |
native | the adapter underneath |
ping, close, is_closed | the connection itself |
On a Transaction:
query, exec | run a statement |
fetch_one, fetch_all, fetch_value | read a result |
insert, insert_many, update, delete, find, find_one, count | the same helpers a connection has |
commit, rollback, is_finished | ending it |
connection, nested | what it is running on, and whether it is a savepoint |
The mail module is Zuri’s mail stack: the message format, the three
protocols that move messages around, and the servers at the far end of
two of them. It is written in Zuri from the socket up, and it speaks to
real mail servers.
Mail is older than almost everything it runs on, and it shows. A message is text with a header block, defined in 1982 and extended ever since to carry anything that is not ASCII, which by now is most of what people send. Sending one is SMTP. Reading one where it is kept is IMAP. Taking one away is POP3. Three protocols, one format, and a great deal of accumulated history, almost all of which a program should not have to know about.
That is the design. A message is described by what is in it and the
module works out the rest: which parts it needs, how they nest, which
encoding each header and each body wants, and what has to be escaped so
that a line of a single dot in the body does not end the message early.
A program says who a message is from and what it says; it does not say
Content-Transfer-Encoding.
- Following Along
- Introduction
- Building a Message
- Reading a Message
- Sending
- Reading Mail with IMAP
- Collecting Mail with POP3
- Running an SMTP Server
- Running an IMAP Server
- Serving More Than One Connection
- Proving Where Mail Came From
- Authentication Mechanisms
- Errors
- What the Module Refuses
- Module Reference
Following Along
Most examples below build on one message. This is it:
import mail
var note = mail.message()
.set_from('Ada Lovelace <ada@example.com>')
.add_to('Grace Hopper <grace@example.com>')
.set_subject('On Engines')
.set_text('The analytical engine weaves algebraic patterns.')
echo note.subject()
echo note.sender().name
echo note.to()[0].address
On Engines
Ada Lovelace
grace@example.com
Nothing there touches a network. Building and reading messages is entirely separate from moving them, and a program that only needs to produce or parse one never opens a socket.
Introduction
There are three protocols and they do genuinely different things.
SMTP moves a message from wherever it was written to a server that will take responsibility for it. That is all it does. It has no notion of a folder, a read message, or a search; a message goes in and either is accepted or is not.
IMAP is for reading mail where it is kept. The mail stays on the server, the client works with it in place, and two devices looking at the same mailbox see the same thing. Almost everything a mail client does is IMAP.
POP3 is for taking mail away. There is a numbered list and there is removing things from it. It is the wrong protocol for reading mail on more than one device and the right one for a program whose job is to drain a mailbox into somewhere else.
The module covers the client end of all three and the server end of the
two that have one worth having. There is no POP3 server here and there
is not meant to be: a POP3 server is an IMAP server with almost
everything taken away, and mail.imap is the one to run.
Building a Message
mail.message() starts one. What goes in it is said one call at a
time, and each call hands the message back, so they read as one:
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_subject('On Engines')
.set_text('The analytical engine weaves algebraic patterns.')
echo note.content_type().mime_type()
text/plain
A message starts with a Date, because the time it was written is the
time it was written, and with MIME-Version. It gets its Message-ID
when it is sent, because that is when the domain it belongs to is
settled.
Addresses
Every address setter takes text, an Address, or a list of either:
import mail
var note = mail.message()
.set_from('Ada Lovelace <ada@example.com>')
.add_to(['grace@example.com', 'Charles Babbage <charles@example.com>'])
.add_cc('archive@example.com')
echo note.to().map(@(person) => person.address)
echo note.to()[1].name
echo note.cc().length()
[grace@example.com, charles@example.com]
Charles Babbage
1
add_to() keeps whoever is already there, so it can be called in a
loop. set_to() replaces them. The same pair exists for Cc and
Bcc.
Bcc is removed from the message before it is handed to a server,
after the recipients have been taken out of it. That is the whole point
of a blind copy, and it is easy to get wrong by hand.
A display name with characters a header cannot carry is encoded, and one containing a comma is quoted, without being asked:
import mail
echo mail.Address('gruss@example.com', 'Grüße').to_string()
echo mail.Address('ada@example.com', 'Lovelace, Ada').to_string()
=?utf-8?B?R3LDvMOfZQ==?= <gruss@example.com>
"Lovelace, Ada" <ada@example.com>
Reading them back gives the name, not the encoding:
import mail
echo mail.parse_address('=?utf-8?B?R3LDvMOfZQ==?= <gruss@example.com>').name
echo mail.parse_address('"Lovelace, Ada" <ada@example.com>').name
Grüße
Lovelace, Ada
mail.parse_address_list() reads a whole header, including the
comments and groups that RFC 5322 allows and nobody remembers:
import mail
var people = mail.parse_address_list(
'Ada <ada@example.com> (the author), Engineers: grace@example.com, charles@example.com;'
)
echo people.map(@(person) => person.address)
[ada@example.com, grace@example.com, charles@example.com]
The group’s members come back alongside the rest, since a program
sending mail cares who the recipients are and not how they were
gathered. mail.parse_address_groups() keeps the grouping when that
matters.
Text, HTML, or Both
Text alone is a plain message. HTML alone is an HTML message. Both
together become a multipart/alternative, with the plain text first,
because a reader that understands both is meant to take the last one it
can display:
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_subject('On Engines')
.set_text('The engine weaves algebraic patterns.')
.set_html('<p>The engine weaves <em>algebraic patterns</em>.</p>')
echo note.content_type().mime_type()
echo note.parts().map(@(part) => part.content_type().mime_type())
multipart/alternative
[text/plain, text/html]
Setting the text again replaces it wherever in the tree it is, rather than adding a second copy. The shape is a consequence of the content, so it stays correct as the content changes.
Outgoing text is written as UTF-8, or as US-ASCII when that is all it needs. Naming any other character set raises, because the module cannot encode into one and a header claiming otherwise would be a lie.
Attachments
An attachment is a thing in its own right, built the way a message is
and handed to attach():
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_subject('The figures')
.set_text('Attached.')
.attach(mail.attachment('name,total\nengines,41\n')
.set_filename('figures.csv'))
echo note.content_type().mime_type()
echo note.attachments().map(@(part) => part.filename())
echo note.attachments()[0].content_type().mime_type()
multipart/mixed
[figures.csv]
text/csv
The media type came from the filename. The message became a
multipart/mixed with whatever it held before as the first part; that
happens once however many files are attached.
set_filename(name) | the name to offer the file under |
set_content_type(type) | the media type, when the filename is not enough |
set_disposition(kind) | attachment by default, or inline |
set_cid(id) | the identifier HTML refers to it by |
set_encoding(name) | the transfer encoding, chosen from the data when absent |
set_description(text) | what some clients show beside the file |
A file on disk is one call:
note.attach(mail.Attachment.from_file('reports/q3.pdf'))
The name and the media type both come from the path unless something says otherwise. A filename with characters a header cannot carry is encoded the way RFC 2231 says, split across continuations if it is long, and read back whole at the far end.
Images Inside the HTML
An image the message shows rather than offers is attached like any other file and referred to by an identifier:
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_subject('The logo')
var source = note.embed(
mail.attachment(bytes([137, 80, 78, 71])).set_filename('logo.png')
)
note.set_html('<p>Our mark: <img src="${source}"></p>')
echo source.starts_with('cid:')
echo note.content_type().mime_type()
echo note.parts().map(@(part) => part.content_type().mime_type())
true
multipart/related
[text/html, image/png]
embed() returns the reference to point an <img> at and rearranges
the message into the multipart/related that tells a reader the two
belong together. The HTML ends up as the first part, which is what
marks it as the one to display.
Headers of Your Own
Anything the setters do not cover goes through set_header():
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_header('X-Priority', '1')
.set_header('List-Unsubscribe', '<https://example.com/unsubscribe>')
echo note.headers.get('X-Priority', nil)
1
The value is written exactly as given. A header that needs encoding needs it applied first; the setters that cover addresses and the subject do that for you, because those are the headers where getting it wrong is common.
add_header() keeps any header already there rather than replacing it,
which is what Received needs.
Replies and Threads
A reply is an ordinary message with two more headers, and getting them right is what puts it in the same conversation in a mail client:
import mail
var original = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_subject('On Engines')
.set_message_id('<first@example.com>')
var reply = mail.message()
.set_from('grace@example.com')
.add_to(original.sender())
.set_subject('Re: ${original.subject()}')
.set_in_reply_to(original.message_id())
.set_references(original.references() + [original.message_id()])
.set_text('Quite so.')
echo reply.headers.get('In-Reply-To', nil)
echo reply.references()
<first@example.com>
[<first@example.com>]
references() on the original gives whatever chain it already belonged
to, so appending its own identifier extends the thread rather than
starting a new one.
Reading a Message
mail.parse() takes the bytes and gives back the tree:
import mail
var note = mail.parse(
'From: Ada Lovelace <ada@example.com>\r\n'
+ 'To: grace@example.com\r\n'
+ 'Subject: =?utf-8?Q?On_Engines_and_Gr=C3=BC=C3=9Fe?=\r\n'
+ 'Content-Type: text/plain; charset=utf-8\r\n'
+ '\r\n'
+ 'The engine weaves algebraic patterns.'
)
echo note.sender().name
echo note.subject()
echo note.text()
Ada Lovelace
On Engines and Grüße
The engine weaves algebraic patterns.
Nothing is decoded until something asks for it. Reading the subject of a message with a twenty megabyte attachment in it costs no more than reading the headers, which is what makes it reasonable to parse everything in a mailbox and look at only some of it.
The Tree
A part is a message too: it has headers and a body, and its body may be more parts. The same class covers both, so walking a message and reading a standalone one are the same code.
import mail
var note = mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_text('plain')
.set_html('<p>html</p>')
.attach(mail.attachment('some,data\n').set_filename('figures.csv'))
for part in note.walk() {
echo part.content_type().mime_type()
}
multipart/mixed
multipart/alternative
text/plain
text/html
text/csv
walk() is the whole tree, outermost first. parts() is one level.
find_part() is the first part of a given type anywhere in it.
Finding the Body
Most programs want the text, wherever it happens to be:
import mail
var note = mail.parse(mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_text('the plain version')
.set_html('<p>the html version</p>')
.to_bytes())
echo note.text_body()
echo note.html_body()
the plain version
<p>the html version</p>
Both return nil when the message has no such part, which is not the
same as an empty one. An attachment that happens to be text/plain is
not mistaken for the body: a part that says it is an attachment, or
that carries a filename without saying either way, is left out.
Attachments Coming In
import mail
var note = mail.parse(mail.message()
.set_from('ada@example.com')
.add_to('grace@example.com')
.set_text('Attached.')
.attach(mail.attachment('name,total\nengines,41')
.set_filename('figures.csv'))
.to_bytes())
for part in note.attachments() {
echo '${part.filename()} (${part.content_type().mime_type()})'
echo part.body_bytes().to_string()
}
figures.csv (text/csv)
name,total
engines,41
body_bytes() decodes the transfer encoding, so base64 and
quoted-printable both come back as the bytes that went in. text()
goes one step further and applies the character set.
Character Sets
A header carrying anything outside ASCII carries it encoded, and there are two encodings it might have used. Reading a header through the module decodes both:
import mail.encoding
echo encoding.decode_words('=?utf-8?B?SGVsbG8=?= =?utf-8?B?IHdvcmxk?=')
echo encoding.decode_words('=?iso-8859-1?Q?caf=E9?=')
echo encoding.decode_words('this is not =?encoded? at all')
Hello world
café
this is not =?encoded? at all
Whitespace between two adjacent encoded words is dropped, which is what
lets a long subject be split across several of them without a space
appearing where none was written. Anything that is not a well-formed
encoded word is left exactly as it is, including text that merely
begins with =?.
Bodies say their character set in the Content-Type. The sets a
message is likely to name are understood: UTF-8, US-ASCII, the ISO 8859
Latin sets 1 and 15, windows-1252, and UTF-16 in either byte order. One
this does not know is read as UTF-8, which leaves whatever ASCII is in
it intact rather than discarding the part that would have been
readable.
import mail.encoding
echo encoding.decode_text(bytes([0x63, 0x61, 0x66, 0xe9]), 'iso-8859-1')
echo encoding.decode_text(bytes([0x93, 0x41, 0x94]), 'windows-1252')
café
“A”
Sending
mail.send() opens a connection, sends one message and closes it
again:
import mail
mail.send('smtp://mail.example.com', mail.message()
.set_from('reports@example.com')
.add_to('ada@example.com')
.set_subject('Quarterly report')
.set_text('The numbers are in.'),
{ username: 'reports', password: secret })
That is the whole of sending mail for a program that sends one at a time.
A Client of Your Own
A program sending many wants a connection kept open across all of them, because opening one costs a handshake and an authentication exchange:
import mail.smtp { SmtpClient }
var server = SmtpClient.connect('smtp://mail.example.com', {
username: 'reports',
password: secret,
})
for note in queue {
server.send(note, nil)
}
server.quit()
quit() says goodbye and closes. close() just closes, which is what
to do when something has gone wrong and the conversation is no longer
in a state the server would recognise.
| scheme | port | what it means |
|---|---|---|
smtp:// | 587 | submission, with TLS negotiated over it |
smtps:// | 465 | TLS from the first byte |
Port 25 is for one server relaying to another, not for a program submitting mail, and it is not a default here. Give it explicitly when that is genuinely what you are doing.
The Envelope
What the server is told and what the message says are two different
things. The envelope comes from the message unless the options say
otherwise: the sender from From, the recipients from To, Cc and
Bcc together, with duplicates removed.
server.send(note, { from: 'bounces@example.com' })
server.send(note, { to: ['someone@example.com'] })
The first is what a mailing list does: the message says who wrote it, the envelope says where a bounce should go. The second delivers to somewhere the headers do not mention at all, which is how a blind copy actually works underneath.
An empty sender is the null path, which is what a bounce is sent from so that it cannot be bounced in turn:
server.send(bounce, { from: '' })
TLS
A client connects, reads what the server can do, negotiates TLS, and only then sends anything worth protecting. That is the default:
tls | what happens |
|---|---|
require | negotiate TLS, and refuse to go on without it. The default. |
prefer | negotiate it when the server offers it |
disable | do not ask |
smtps://, imaps:// and pop3s:// handshake before the first byte
instead, and then tls has nothing left to decide.
Whatever tls says, a mechanism that puts the password on the wire is
never used on a connection that is not encrypted. A client offered
nothing else raises rather than sending it. That is not configurable,
and it is the one place the module refuses to do what it is told.
To trust a certificate the system does not, hand it a configuration:
import net.tls
var config = tls.TlsConfig()
config.add_ca_pem(file('internal-ca.pem').read())
var server = SmtpClient.connect('smtp://mail.internal', {
username: 'reports',
password: secret,
tls_config: config,
})
Authenticating
Credentials go in the options, and the client picks the strongest mechanism both ends know:
SmtpClient.connect('smtp://mail.example.com', {
username: 'reports',
password: secret,
})
SmtpClient.connect('smtp://smtp.gmail.com', {
username: 'reports@example.com',
token: access_token,
})
A token picks a token mechanism. To force one rather than choosing,
pass mechanisms; the choice is described under Authentication
Mechanisms.
A username and password in the connection string work too, which is convenient for a string that came from configuration:
mail.send('smtp://reports:secret@mail.example.com', note, nil)
What the Server Can Do
echo server.capabilities().keys()
echo server.max_size()
A message larger than the server said it would take is refused before it is sent rather than after uploading it, which matters when the message is the reason the connection is slow.
Where the server offers CHUNKING, send() can hand the message over
in pieces with nothing escaped:
server.send(note, { chunking: true })
Where it offers SIZE, 8BITMIME or DSN, those are used without
being asked for. Where it does not, the message still goes.
When Sending Fails
A refusal partway through a transaction leaves the server holding half
of one. send() abandons it before raising, so the connection is still
usable for the next message rather than answering 503 to everything
afterwards. That is worth knowing because doing it by hand is easy to
forget:
for note in queue {
catch {
server.send(note, nil)
} as error {
failures.append([note, error])
}
}
server.quit()
Every message after a failure still goes.
Reading Mail with IMAP
IMAP leaves the mail on the server. A client opens a mailbox, searches it, and fetches what it needs:
import mail.imap { ImapClient }
var inbox = ImapClient.connect('imaps://mail.example.com', {
username: 'ada',
password: secret,
})
inbox.select('INBOX')
for uid in inbox.search('UNSEEN', true) {
var note = inbox.fetch_message(uid, true, false)
echo '${note.sender().address}: ${note.subject()}'
}
inbox.logout()
Opening a Mailbox
var box = inbox.select('INBOX')
echo box.exists
echo box.uidvalidity
echo box.permanent_flags
select() opens a mailbox for reading and writing. examine() opens
it read-only, which also means that reading a message does not mark it
read. close_mailbox() closes it and removes anything marked deleted
on the way out; unselect() closes it without doing that.
uidvalidity is how a server says its numbering has been reset. A
client that remembers identifiers between sessions has to check it: if
it has changed, every identifier it remembers means something else now.
To see what is there:
for box in inbox.list('', '*') {
echo '${box.name} ${box.is_selectable() ? '' : '(container only)'}'
}
* matches anything including the separator between levels; %
matches anything except it, which is what lists one level. status()
asks what is in a mailbox without opening it, which is how a client
shows unread counts for a dozen folders without selecting each one:
echo inbox.status('Archive', ['MESSAGES', 'UNSEEN'])
Searching
The criteria are IMAP’s own, and they read close to English:
inbox.search('UNSEEN', true)
inbox.search('FROM ada@example.com SINCE 1-Jan-2026', true)
inbox.search('SUBJECT engines LARGER 10000', true)
inbox.search('FLAGGED UNDELETED', true)
The second argument asks for unique identifiers rather than positions.
A position is only good until something is removed from the mailbox and
everything after it renumbers; an identifier outlives the connection
and is what a program that runs twice should remember. Prefer true
unless the numbers are being used immediately.
Fetching Less Than Everything
Fetching every message to show a list of them is the mistake IMAP exists to prevent:
for info in inbox.fetch('1:50', 'ENVELOPE FLAGS RFC822.SIZE', false) {
echo '${info.envelope.subject} (${info.size} bytes)'
}
The envelope is the addresses and the date, parsed by the server. No body crossed the network at all.
| what to ask for | what comes back |
|---|---|
ENVELOPE | the addresses, subject and date |
FLAGS | what has been done to the message |
RFC822.SIZE | how large it is |
INTERNALDATE | when the server took it |
BODYSTRUCTURE | the shape of the message, part by part |
BODY.PEEK[HEADER] | the header block |
BODY.PEEK[] | the whole message |
BODY.PEEK[2] | one part of it |
BODYSTRUCTURE is the one worth knowing about. It reports what the
message is made of without sending any of it, so a client can decide to
fetch the text and leave a large attachment on the server:
var structure = inbox.fetch(uid, 'BODYSTRUCTURE', true)[0].structure
for part in structure.walk() {
echo '${part.section} ${part.mime_type()} ${part.size}'
}
section is the number to ask for that part on its own.
fetch_headers() is the common case of asking for the header block,
and fetch_message() the common case of asking for all of it:
inbox.fetch_message(uid, true, false) # leaves it unread
inbox.fetch_message(uid, true, true) # marks it read
Reading a message does not mark it read unless you say so. A program going through a mailbox should not change what a person sees when they next open it.
Flags
A flag is what has been done to a message:
inbox.add_flags(uid, ['\\Seen'], true)
inbox.remove_flags(uid, ['\\Flagged'], true)
inbox.mark_seen(uid, true)
inbox.mark_deleted(uid, true)
\Seen, \Answered, \Flagged, \Deleted and \Draft are the ones
with a defined meaning. A server may allow others, and says which in
the mailbox’s permanent_flags.
mark_deleted() only marks. Nothing goes until expunge():
inbox.mark_deleted(uid, true)
echo inbox.expunge()
expunge() returns the positions that went, highest first, because
each removal renumbers everything after it and that is the order they
have to be applied in.
Moving, Copying and Removing
inbox.copy(uid, 'Archive', true)
inbox.move(uid, 'Archive', true)
move() uses the server’s own MOVE where there is one, and otherwise
does what MOVE was invented to replace: copy, mark deleted, expunge.
Either way the message ends up in one place.
create(), delete(), rename(), subscribe() and unsubscribe()
do what they say.
Putting a Message Back
append() adds a message to a mailbox without sending it anywhere,
which is how a sent message gets into the Sent folder:
inbox.append('Sent', note, ['\\Seen'], nil)
The flags are the ones to file it under, and \Seen is the usual one
for something the account itself wrote. The last argument is the
internal date; the time of arrival is used when it is not given.
Waiting for New Mail
idle() waits for the server to say something rather than asking over
and over:
inbox.on_event(@(response) {
echo 'the server said ${response.name()}'
})
while true {
inbox.idle(1500000)
}
A server may drop a connection that idles for too long, which is why the wait is bounded and the loop comes back around. Twenty-five minutes is the default and is what RFC 2177 recommends.
noop() does the same thing without waiting: it gives the server a
chance to report anything that has changed, and keeps the connection
from going idle at all.
Collecting Mail with POP3
POP3 hands the mail over and forgets it. There is no searching and no folders; there is a numbered list, and there is taking things off it:
import mail.pop3 { Pop3Client }
var mailbox = Pop3Client.connect('pop3s://mail.example.com', {
username: 'ada',
password: secret,
})
for entry in mailbox.list() {
archive(mailbox.retrieve(entry.number))
mailbox.delete(entry.number)
}
mailbox.quit()
list() gives every message with its size and, where the server offers
one, an identifier that outlives the session. The number does not: it
is only good until something is removed and the rest renumber.
Nothing is actually removed until quit(). delete() only marks and
reset() unmarks everything. A connection that drops halfway leaves
the mailbox exactly as it was, which is the protocol protecting you
from a program that fails in the middle. It also means that closing
without quit() is how to abandon a run.
top() fetches the headers and the first few lines of the body, which
is enough to decide whether the rest is worth fetching:
var preview = mailbox.top(entry.number, 5)
Where the server’s greeting offers it, APOP is used in preference to
sending the password, and where it offers SASL those mechanisms are
preferred again. USER/PASS is the last resort and is only used over
TLS.
Running an SMTP Server
SmtpServer knows the protocol and nothing about policy. What to
accept is decided by handlers:
import mail.smtp { SmtpServer }
var server = SmtpServer({ port: 2525, hostname: 'mail.example.com' })
server.on_rcpt(@(session, recipient) {
if !accounts.contains(recipient.local) {
return { code: 550, message: 'no such user here', enhanced: '5.1.1' }
}
})
server.on_data(@(session, raw) {
for recipient in session.recipients {
store.append(recipient.local, 'INBOX', raw, nil, nil)
}
})
server.listen()
The Handlers
| handler | called with | when |
|---|---|---|
on_connect | the session | a connection opens, before the greeting |
on_auth | the session and the credentials | a client authenticates |
on_mail | the session and the sender | a transaction starts |
on_rcpt | the session and a recipient | for each recipient |
on_data | the session and the message | the whole message has arrived |
on_close | the session | the connection closes, however it closed |
on_error | the error and the session | something inside the server failed |
The session carries what is known so far: who connected, whether the
connection is encrypted, who authenticated, the sender, and the
recipients accepted up to now. session.state is an empty dictionary
your handlers can put anything in, and it lives as long as the
connection.
Refusing Properly
A handler that returns nothing accepts. One that returns a code and a message refuses with those:
server.on_mail(@(session, sender) {
if blocklist.contains(sender.domain) {
return { code: 550, message: 'not accepted from there', enhanced: '5.7.1' }
}
})
A handler that raises is a failure on the server’s side, not a bad
message, and the sender is told 451 and to try again later. That
distinction matters: a database that is down should not turn into a
bounce.
Refusing at on_rcpt is the useful one. It tells the sender
immediately which address is wrong, while it is still connected and can
do something about it, rather than accepting the message and generating
a bounce to an address that may not exist either.
The null sender, which is what a bounce comes from, arrives at
on_mail as an empty string rather than an address. Refusing to accept
mail from it is how a server ends up unable to receive bounces, so
handle it deliberately.
TLS and Authentication
server.use_tls(file('cert.pem').read(), file('key.pem').read())
That is what lets the server offer STARTTLS. Everything a client said
before the handshake is discarded afterwards, including who it claimed
to be, because none of it was protected.
Authentication needs one handler and, for the mechanisms that prove a password without sending it, a second:
server.on_auth(@(session, credentials) {
if !accounts.verify(credentials.username, credentials.password) {
return { code: 535, message: 'no' }
}
})
server.on_password(@(username) {
return accounts.password_of(username)
})
on_password is what makes CRAM-MD5 possible: checking that proof
means working out the same one, which means knowing the password. A
server that cannot produce one does not advertise the mechanism.
Without on_auth the server advertises no authentication at all.
PLAIN and LOGIN are advertised only once the connection is
encrypted. require_auth refuses mail from a client that has not
authenticated, and require_tls refuses it on a connection that is not
encrypted.
Limits
| option | default | what it does |
|---|---|---|
max_size | 35 MB | the largest message to take |
max_recipients | 100 | recipients one message may have |
timeout | 300000 | milliseconds a client may go quiet |
A client that declares a size past the limit is refused at MAIL,
before it uploads anything. One that does not declare it is refused
when it goes past. Ten malformed commands in a row and the connection
is dropped.
Running an IMAP Server
ImapServer runs the session state machine and answers every command
out of a MailStore. What mail there is and where it lives is the
store’s business:
import mail.imap { ImapServer, MaildirStore }
var store = MaildirStore('/var/mail')
store.add_account('ada', secret)
ImapServer({ port: 143 }, store).listen()
The server implements IMAP4rev1 along with UNSELECT, MOVE, IDLE,
LITERAL+, SASL-IR and ID. It advertises exactly what it
implements, so a client that reads the capability list and trusts it
will not go wrong.
Where the Mail Lives
Two stores ship. MemoryStore keeps everything in the process and
forgets it on exit, which is what a test wants and what a server
embedded in something larger wants when the mail it holds is not the
point:
import mail
import mail.imap { MemoryStore }
var store = MemoryStore(['INBOX', 'Archive'])
store.add_account('ada', 'secret')
store.append('ada', 'INBOX', mail.message()
.set_from('grace@example.com')
.add_to('ada@example.com')
.set_subject('On Bugs')
.set_text('Found one.')
.to_bytes(), nil, nil)
echo store.mailboxes('ada')
echo store.messages('ada', 'INBOX').length()
echo store.messages('ada', 'INBOX')[0].uid
[Archive, INBOX]
1
1
MaildirStore is the same contract on disk.
Maildir
Each account is a directory, holding its INBOX directly and every other mailbox beside it under a leading dot, which is the Maildir++ arrangement every other mail tool understands. A mailbox written by this server can be read by anything else, and mail delivered by anything else turns up here.
Every mailbox is the three directories Maildir defines. A message is
written into tmp, where nothing reads from, and only moved into place
once it is whole, so a reader never sees half of one. new is where a
delivery agent leaves mail nobody has looked at yet, and cur is where a
message lives once a client has seen the mailbox, with its flags recorded
in the filename after :2,. Windows does not allow a colon in a filename,
so there the flags follow ;2, instead, as they do for mbsync. So an
account on disk looks like this:
ada/
cur/ new/ tmp/ zuri-uidlist
.Archive/
cur/ new/ tmp/ zuri-uidlist
That interoperability is the reason to choose it. A delivery agent can drop a message in and the server finds it on the next look, with no shared database and no protocol between them.
IMAP needs a message number that survives a rename and Maildir has no such thing, so each mailbox keeps a small index file beside its directories recording which file is which number. A message whose file has gone keeps its number retired rather than reused, so a client that remembered one is told the message is missing rather than handed a different one.
Mail is on disk; accounts are not:
store.add_account('ada', secret)
store.set_authenticator(@(username, password) {
return accounts.verify(username, password)
})
Where an application keeps its passwords is the application’s business and not a mail library’s. A store with an authenticator can no longer produce a password, so the server stops advertising the mechanisms that need one.
A Store of Your Own
A store only has to answer the same calls. Nothing in the contract requires a file:
authenticate(username, password) | are these the right credentials |
password_of(username) | the password, for the mechanisms that need it |
has_passwords() | whether it can produce one at all |
mailboxes(username) | what mailboxes there are |
exists, create, remove, rename | the mailboxes themselves |
messages(username, mailbox) | everything in one, oldest first |
append(username, mailbox, raw, flags, received) | add a message |
set_flags(username, mailbox, uid, flags) | change one’s flags |
expunge(username, mailbox) | remove what is marked deleted |
counters(username, mailbox) | uidnext and uidvalidity |
Subclass MailStore and implement them, and an ImapServer will serve
whatever is behind it.
Serving More Than One Connection
Both servers serve one connection to the end before taking the next.
That is the right shape for the protocols and the wrong shape for more
than one client at a time. mail.pool puts a pool of isolates behind
one listening socket:
import mail.pool
import .my_server
pool.serve(my_server.build, { host: '0.0.0.0', port: 143, workers: 8 })
build runs inside each worker to construct that worker’s own server,
since isolates share nothing and each needs its own. It can be any
function; keeping it in a module of its own is just a tidy place for a
worker’s setup to live:
# my_server.zu
import mail.imap { ImapServer, MaildirStore }
def build() {
var store = MaildirStore('/var/mail')
store.set_authenticator(accounts.verify)
return ImapServer({}, store)
}
An IMAP connection can be open for hours, so the pool is a ceiling on how many clients can be served at once rather than on how fast they are served. Size it for the number of clients, not the rate of requests.
pool.start() does everything except run the accept loop, which is
what to use when the address has to be known before the first
connection, as it does in a test.
Proving Where Mail Came From
A receiving server has no reason to believe a From header. DKIM is
the domain signing the message on the way out, and the receiver
checking the signature against a key published in that domain’s own
DNS.
Signing
import mail.dkim { Signer }
var signer = Signer('example.com', 'default', private_key)
signer.sign(note)
The signature covers the body and the headers worth covering, and the
header it adds goes at the top. It goes on last, once everything else
about the message is settled: changing a signed header afterwards
breaks it, and mail.send() sets the Message-ID and Date if they
are absent, so let it or set them yourself before signing.
| option | default | what it does |
|---|---|---|
algorithm | rsa-sha256 | or ed25519-sha256 |
canonicalisation | relaxed/relaxed | headers then body |
headers | a sensible set | which headers to cover |
expires_in | none | seconds until the signature stops counting |
relaxed forgives the whitespace and folding changes a mail server may
make in passing. simple covers the bytes exactly, which means any
change at all on the way breaks the signature. Both are implemented;
relaxed/relaxed is what almost everything uses and is the default for
that reason.
The public half goes in DNS under the selector:
default._domainkey.example.com TXT "v=DKIM1; k=rsa; p=MIIBIjANBg..."
Checking
import mail.dkim
for result in dkim.verify(incoming) {
if result.valid {
echo '${result.domain} takes responsibility for this'
} else {
echo '${result.domain}: ${result.reason}'
}
}
One result per signature, in the order they appear. A message with no signatures gives an empty list, which is not a failure: it is a message nobody signed.
The key is looked up through net.resolver unless verify() is handed
something else to look it up with, which is what a program with its own
resolver or its own cache wants:
dkim.verify(incoming, @(name) {
return my_dns.text_records(name)
})
A message that was parsed is checked against the bytes it was parsed from, which is the only thing a signature can be checked against. Changing a parsed message and checking it again reports on the message that arrived, not the one now in hand.
What a Signature Does Not Say
A valid signature says the domain vouches for the message. It does not
say the message is wanted, that the From header matches the signing
domain, or that the sender is who they claim to be to a human reader.
Deciding what a signature from a given domain is worth is a separate
question with a separate answer, and that answer is usually DMARC.
DNSSEC is not validated. Signatures are carried through when a server
sends them and the dnssec option asks for them, but nothing here
checks one. A program that needs a validated answer wants a validating
resolver and a trusted path to it, which is what net.resolver’s tls
option gives.
Authentication Mechanisms
All three protocols carry the same mechanisms, and mail.sasl holds
them once rather than three times. A client is given credentials and
picks the strongest thing both ends know:
SCRAM-SHA-256, SCRAM-SHA-1 | proves the password without sending it, and proves the server knew it too |
CRAM-MD5 | proves it without sending it, and proves nothing about the server |
XOAUTH2, OAUTHBEARER | a bearer token, which is what the large providers want |
PLAIN, LOGIN | the password itself, so never without TLS |
EXTERNAL | nothing: the client certificate already said who this is |
The order is the order of that table. SCRAM is the one to use where a server offers it: the server stores something derived from the password rather than the password, the client never sends it, and the exchange ends with the server proving it knew it too, which is what stops a server that simply says yes to everything.
Channel binding, the -PLUS form of SCRAM, is not offered. It needs a
value out of the TLS session that net.tls does not expose.
To force a mechanism rather than choosing:
SmtpClient.connect(url, {
username: 'reports',
password: secret,
mechanisms: ['SCRAM-SHA-256'],
})
Errors
Every error is a MailError. Catching that catches everything the
module raises on its own.
MessageError | the bytes are not a message, or are one that contradicts itself |
ProtocolError | the server said something the protocol does not allow |
ConnectionClosed | the connection went away mid-conversation |
AuthenticationError | the credentials were refused, or nothing usable was offered |
StateError | a command that makes no sense where it was issued |
MailboxError | a mailbox that does not exist, or cannot be created |
SmtpError | the SMTP server refused, with a code |
ImapError | the IMAP server answered NO or BAD |
Pop3Error | the POP3 server answered -ERR |
The one distinction worth building a mail program around is SMTP’s:
catch {
mail.send(url, note, credentials)
} as error {
if instance_of(error, mail.SmtpTransientError) {
queue.retry(note)
} else if instance_of(error, mail.SmtpPermanentError) {
queue.bounce(note, error.code)
} else {
raise error
}
}
A 4xx reply means the server could not take the message now and the sender should try again later. A 5xx means it will not take the message and trying again changes nothing. Queue the first and bounce the second; treating them alike is how a mail queue either loses mail or sends it forty times.
SmtpError carries code, the three digit reply, and enhanced, the
finer grained code from RFC 3463 when the server sends one. Neither is
meant to be matched on beyond its first digit.
ImapError carries status, which is NO when the server understood
the command and refused it, and BAD when it did not understand it at
all. BAD points at the client rather than at the request.
What the Module Refuses
It will not send a password over an unencrypted connection. Not
with an option, not with a flag. A client offered only mechanisms that
would do that raises instead. tls: 'disable' turns off negotiating
TLS; it does not turn off this.
It will not write a character set it cannot encode. Asking for a body in ISO 8859-1 raises rather than writing UTF-8 bytes under a header that claims otherwise.
It does not validate DNSSEC. See What a Signature Does Not Say.
It does not implement a POP3 server. A POP3 server is an IMAP
server with almost everything removed, and running mail.imap is the
better answer.
It does not decide whether mail is wanted. There is no spam filtering, no reputation, and no policy. DKIM tells you who signed something; what to do about that is yours.
Module Reference
The standard library reference documents every class and method. The shape of the module:
mail.message() | starts a message |
mail.parse(data) | reads one |
mail.send(url, note, options) | sends one, connection and all |
mail.smtp(url, options) | a connection to a server that sends mail |
mail.imap(url, options) | a connection to a server that stores it |
mail.pop3(url, options) | a connection to a server that hands it over |
mail.parse_address(text) | one address |
mail.parse_address_list(text) | every address in a header |
mail.format_addresses(people) | the other direction |
mail.attachment(data) | starts an attachment |
mail.Attachment.from_file(path) | one read from disk |
On an Attachment:
set_filename, set_content_type, set_encoding | what it is |
set_disposition, set_cid, set_description | how it is presented |
filename, content_type, disposition, cid, size, data | reading them back |
to_part | the message part it becomes |
On a Message:
set_from, set_sender, set_reply_to | who it is from |
add_to, add_cc, add_bcc, set_to, set_cc, set_bcc | who it is for |
set_subject, set_text, set_html | what it says |
attach, embed | what it carries, as an Attachment |
set_header, add_header, set_date, set_message_id | anything else |
set_in_reply_to, set_references | which conversation it belongs to |
sender, to, cc, bcc, reply_to, recipients | reading them back |
subject, text, text_body, html_body | reading what it says |
parts, walk, find_part, attachments | the tree |
to_bytes, to_string | writing it out |
The protocol modules:
mail.smtp | SmtpClient, SmtpServer, SmtpSession, Reply |
mail.imap | ImapClient, ImapServer, Mailbox, Envelope, BodyPart, MessageInfo |
mail.imap | MailStore, MaildirStore, MemoryStore, StoredMessage |
mail.pop3 | Pop3Client, Entry |
mail.dkim | Signer, Signature, Result, verify, is_signed |
mail.sasl | of, choose, and a class per mechanism |
mail.pool | serve, start, Cluster |
And the pieces underneath, for a program that needs them directly:
mail.address | parse, parse_list, parse_groups, format_list |
mail.headers | Headers, fold, unfold |
mail.encoding | the transfer encodings, encoded words and parameters |
mail.content | ContentType, ContentDisposition |
mail.stream | LineStream, connect, start_tls, endpoint |
mail.imap.parser | the IMAP grammar, for reading a response by hand |
Configuration and the Environment
Every program that talks to anything else needs to be told where it is. A database address, an API key, a port to bind, a flag that turns verbose logging on for one afternoon. None of that belongs in the source: it changes between your laptop and the server, and some of it must never reach a repository at all.
The answer the industry settled on is the process environment. Orchestrators set it, CI runners set it, shells set it, and every language can read it. What none of them solve is the first machine in the chain: yours, where there is no orchestrator and nobody wants to prefix every run with eight assignments.
The env module closes that gap. It reads a .env file into the
process environment at startup, and it reads values back out already
converted to the number, boolean or list your program is going to use.
The file is a development convenience; the environment is the
interface. A program written against env does not know or care which
one supplied a value, which is exactly what lets it run unchanged in
both places.
- The First Line
- Writing a
.envFile - What Is Already Set Wins
- References Between Values
- Reading Values Back
- Building a Load
- Writing a File Back Out
- When It Goes Wrong
- Keeping Secrets Out of Git
- Module Reference
The First Line
One call, as early in the program as you can put it:
import env
file('.env.demo', 'w').write('PORT=8080\nGREETING=hello\n')
env.load('.env.demo')
echo env.get('GREETING')
echo env.int('PORT')
hello
8080
env.load() with no argument reads .env in the current working
directory, which is what almost every program wants. The examples in
this chapter write their file first and name it, so that you can run
each one as it stands.
The file is optional. If it is not there, load() leaves the
environment exactly as it found it and the program carries on with
whatever the shell already provided. That is deliberate: the same
binary runs on a laptop with a .env file and on a server without one.
Writing a .env File
One name per line, a =, and a value:
HOST=127.0.0.1
PORT=8080
DATABASE_URL=postgres://app@localhost/app
A name is a letter or underscore followed by letters, digits and
underscores, which is the same rule a shell applies. Anything else
stops the load with a ParseError rather than being silently dropped,
because a name no shell could ever export is a typo every time.
A leading export is allowed and ignored, so the same file can be fed
to source in a terminal when you want the variables in your shell
too.
Quoting
An unquoted value is the text up to the end of the line, with the whitespace trimmed off both ends. Wrap it in quotes when that is not what you want:
import env
var values = env.parse(
'BARE= plain text \n' +
'COMMENTED=value # not part of it\n' +
'PASSWORD=hunter#2\n' +
"RAW='no \\n escape, no expansion'\n" +
'COOKED="caf\\u00e9"\n'
)
for name, value in values {
echo '${name} -> [${value}]'
}
BARE -> [plain text]
COMMENTED -> [value]
PASSWORD -> [hunter#2]
RAW -> [no \n escape, no expansion]
COOKED -> [café]
The three quoting forms differ in exactly one way each:
| Written | Means |
|---|---|
KEY=value | trimmed, ends at a comment or the line |
KEY='value' | every character as written, nothing resolved |
KEY=`value` | the same, for values containing both other quotes |
KEY="value" | backslash escapes resolved, $ references expanded |
Inside double quotes, \n, \r, \t, \f, \v, \b, \a, \e
and \0 mean what they do in Zuri, and \xHH, \uHHHH and \u{H...}
name a codepoint. A backslash before anything else keeps both
characters, so "C:\Users\me" is the path you meant and not a lesson
in escaping.
Reach for single quotes whenever a value is a secret. A generated
password containing $ or \ goes through untouched, and nothing in
it can accidentally name another variable.
Comments
A # starts a comment, either on its own line or after a value. Inside
an unquoted value it only does so when it begins the value or follows a
space, which is why PASSWORD=hunter#2 above kept its #. Where that
rule is too subtle to rely on, quote the value and the question does
not arise.
Values That Span Lines
A quoted value runs until its closing quote, newlines included. This is how a private key goes in a file:
SIGNING_KEY="-----BEGIN PRIVATE KEY-----
MIIEvQIBADANBgkqhkiG9w0BAQEFAASCBKcwggSjAgEAAoIBAQC7...
-----END PRIVATE KEY-----"
Single quotes work the same way and resolve nothing, which is usually the better choice for key material.
What Is Already Set Wins
A name that is already set in the process environment is left alone. The file fills in the rest:
import env
import os
os.set_env('DATABASE_URL', 'postgres://production/app', true)
file('.env.demo', 'w').write(
'DATABASE_URL=postgres://localhost/app\n' +
'CACHE_TTL=300\n'
)
var result = env.load('.env.demo')
echo os.get_env('DATABASE_URL')
echo result.applied
echo result.skipped
postgres://production/app
[CACHE_TTL]
[DATABASE_URL]
This is the whole point of the default. The file carries what a developer needs to run the program at all; production sets the real values through the shell, the orchestrator or the CI runner, and the file never has to know that it did.
load() returns a Result that accounts for every name and every
file, so a program never has to guess whether its configuration
arrived. applied is what the load set, skipped is what was already
set, values is everything the file defined either way, and files
and missing say which paths were read and which were not there.
Overriding
Where the file genuinely should win, say so:
import env
import os
os.set_env('LOG_LEVEL', 'warn', true)
file('.env.demo', 'w').write('LOG_LEVEL=debug\n')
env.loader().path('.env.demo').override().load()
echo os.get_env('LOG_LEVEL')
debug
Empty Means Unconfigured
Throughout the module, a variable set to the empty string counts as
unset. FLAG= in a file, an empty shell variable, and a name nobody
ever set all mean the same thing, because that is the one thing anybody
means by any of them. A load will fill in an empty name, has() is
false for it, require() raises on it, and every typed reader returns
its default.
os.get_env() is the unfiltered view for the rare case that needs to
tell an empty value apart from an absent one.
References Between Values
A value may refer to another value, or to the environment around it, with the syntax a shell uses:
HOST=localhost
PUBLIC_URL=http://${HOST}:${PORT:-8080}
import env
file('.env.demo', 'w').write(
'HOST=localhost\n' +
'PUBLIC_URL=http://$' + '{HOST}:$' + '{PORT:-8080}\n'
)
var loaded = env.load('.env.demo')
echo loaded.values.PUBLIC_URL
http://localhost:8080
The split string in that example is a Zuri detail, not a .env one: a
${ written inside a Zuri literal is an interpolation, so building
.env text in source means keeping the $ and the { apart. A file
you type in an editor has no such problem, as the dosini block above
shows.
A reference resolves to the value the name will hold once the load has
finished. That single rule is what makes PORT=${PORT:-8080} read the
way it looks: whatever the environment already says, and 8080 when it
says nothing. It holds no matter which line of the file the reference
sits on, and it changes with override() exactly as precedence does.
A definition may also reach past itself to the value it is replacing, which is how a path gets extended rather than clobbered:
PATH=${PATH}:/opt/app/bin
Fallbacks and Demands
The three modifiers are POSIX’s, and the leading colon on each widens the test from “is it set” to “is it set to something”:
| Written | Means |
|---|---|
${NAME:-fallback} | the value, or fallback when it is unset or empty |
${NAME-fallback} | the value, or fallback when it is unset |
${NAME:+instead} | instead when the value is set and not empty, otherwise nothing |
${NAME+instead} | instead when the value is set, otherwise nothing |
${NAME:?reason} | the value, or a MissingVariable carrying reason |
${NAME?reason} | the same, counting an empty value as set |
The text after a modifier is itself expanded, so a fallback may have a fallback:
CACHE_DIR=${XDG_CACHE_HOME:-${HOME}/.cache}/app
:? is worth knowing. It turns the file itself into the place where a
deployment’s requirements are written down, and the failure arrives at
startup with the reason attached:
SESSION_SECRET=${SESSION_SECRET:?the deploy must supply a session secret}
Literal Dollar Signs
\$ is a dollar sign and nothing else. A $ that does not name
anything, as in PRICE=$5.00, stands for itself already. A value in
single quotes or backticks is never expanded at all, so a secret
containing ${ needs no thought:
API_KEY='sk_live_${not_a_reference}'
Reading Values Back
An environment variable is always text, and almost nothing in a program wants text. The typed readers convert it and refuse anything that is not what it claims to be:
import env
file('.env.demo', 'w').write(
'PORT=8080\n' +
'DEBUG=yes\n' +
'REQUEST_TIMEOUT=2.5\n' +
'CORS_ORIGINS= https://a.example , https://b.example \n' +
'SENTRY_DSN=\n'
)
env.load('.env.demo')
echo env.int('PORT', 3000)
echo env.bool('DEBUG', false)
echo env.float('REQUEST_TIMEOUT', 1)
echo env.list('CORS_ORIGINS', ',', [])
echo env.get('SENTRY_DSN', 'not configured')
echo env.has('SENTRY_DSN')
8080
true
2.5
[https://a.example, https://b.example]
not configured
false
get(name, default) | the text, or the default |
require(name, reason) | the text, or a MissingVariable |
has(name) | whether it is configured at all |
int(name, default) | a decimal integer, optionally signed |
float(name, default) | a number, exponents included |
bool(name, default) | 1, true, yes, y, on and their opposites |
list(name, separator, default) | split, trimmed, empty items dropped |
A value that is set but not convertible is a ValueError naming the
variable, not a silent fall back to the default. PORT=eighty is a
mistake in the configuration, and the moment to hear about it is
startup rather than the first request that needed a port.
These read the process environment, not the file. They answer the same
whether a value came from .env or from the shell, which is what lets
one program run in both places without a branch anywhere in it.
require() deserves its own line in most programs. A setting the code
cannot invent a default for should stop the program at the top, with
its own name in the message:
var secret = env.require('SESSION_SECRET', 'set SESSION_SECRET in .env')
Building a Load
env.loader() returns a Loader when the one-line form is not enough.
Every setting returns the loader, so a whole configuration is one
expression, and nothing about the loader changes when it runs, so the
same one can be kept and used again.
Layering Files
Sources are read in the order they are added, and the last one to define a name is the one that defines it:
import env
file('.env.demo', 'w').write('HOST=localhost\nPORT=8080\n')
file('.env.local.demo', 'w').write('HOST=127.0.0.1\n')
var result = env.loader()
.path('.env.demo')
.path('.env.local.demo')
.path('.env.missing.demo')
.read()
echo result.values
echo result.files
echo result.missing
{HOST: 127.0.0.1, PORT: 8080}
[.env.demo, .env.local.demo]
[.env.missing.demo]
That pairing is the useful one: .env holds what the team shares and
is worth committing as .env.example, and .env.local holds what one
machine does differently and is not. A path may begin with ~, which
expands to the home directory.
Text Instead of a File
source() adds text rather than a path, so configuration that arrived
over the network or out of a secret store goes through exactly the same
parsing, expansion and precedence as a file on disk:
env.loader()
.path('.env')
.source(vault.fetch('app/production'))
.load()
Demanding the File
A file that is not there is not an error by default, and its path is
recorded in Result.missing. Where the file is genuinely part of the
deployment, say so and let it fail loudly:
env.loader().path('/etc/app/env').required().load()
Looking Without Loading
read() does everything load() does except the last step. Nothing is
written to the process environment, and $ references still resolve
the way they would have, so what comes back is exactly what load()
would have set. It is how a tool inspects a file, and how a test checks
one without changing the process it is running in.
expand(false) turns expansion off entirely, for a file of opaque
secrets where every value should be taken exactly as written.
Writing a File Back Out
stringify() is parse() backwards, for the program that generates a
.env file rather than reading one:
import env
print(env.stringify({
HOST: '127.0.0.1',
PORT: 8080,
GREETING: 'hello there',
DEBUG: false,
}))
HOST=127.0.0.1
PORT=8080
GREETING="hello there"
DEBUG=false
Values are written bare where that is unambiguous and double-quoted
where it is not, and a $ inside a value is escaped so that loading
the file back does not expand it. The output round-trips: feeding it to
parse() returns the same names and the same values.
Numbers, bigints, booleans and bytes are converted to their text.
Anything else is a TypeError, because guessing what a list should
look like in an environment file is how configuration goes wrong
quietly.
When It Goes Wrong
Every error the module raises is an EnvError, and each subclass
carries the detail a message alone cannot:
import env
catch {
env.parse('12FACTOR=yes')
} as error {
echo '${error.type}: ${error.message}'
echo 'line ${error.line}, column ${error.column}'
}
catch {
env.loader().path('.env.nowhere').required().load()
} as error {
echo '${error.type}: ${error.message}'
}
catch {
env.require('NOTHING_SET_THIS')
} as error {
echo '${error.type}: ${error.message}'
}
ParseError: expected a variable name
line 1, column 1
MissingFile: no environment file at .env.nowhere
MissingVariable: NOTHING_SET_THIS is not set
| Class | Raised when | Carries |
|---|---|---|
ParseError | a source is not a valid environment file | line, column |
MissingFile | a required file is not there | path |
MissingVariable | a demanded variable is not configured | name |
Catching EnvError catches all three, and a ValueError from a typed
reader is the ordinary one, so it goes wherever the rest of your
validation failures go.
Keeping Secrets Out of Git
A .env file holds the values that differ between one deployment and
the next, which is to say it holds the secrets. Three lines of
housekeeping and the subject never comes up again:
$ echo '.env' >> .gitignore
$ echo '.env.local' >> .gitignore
$ cp .env .env.example # then replace every real value
Commit .env.example with every name in it and no real value. It is
the only documentation of what the program needs that cannot go stale,
because the day someone adds a setting without adding it there is the
day a new checkout stops working.
Resist the pull towards a .env.production checked in beside a
.env.staging. Configuration varies per deploy, not per named
environment, and the moment there are two files somebody edits the
wrong one. Use one file per machine and let required() and :? say
out loud what that machine still owes you.
Module Reference
The whole surface:
load(path) | read a file into the process environment |
loader() | a Loader, for a load that needs more |
parse(source) | text to names and values, nothing else |
stringify(values) | names and values back to text |
is_name(name) | whether a name is a legal variable name |
Reading values back:
get, require, has | text, or the absence of it |
int, float, bool, list | text converted, or a ValueError |
The classes:
Loader | path, paths, source, override, expand, required, read, load |
Result | values, applied, skipped, files, missing |
EnvError | ParseError, MissingFile, MissingVariable |
And the piece underneath, for a program that needs it directly:
env.expand | expand(value, resolve), the $ reference syntax on its own |
Parsing the Command Line
A program started from a shell is handed a list of strings and nothing
else. Everything a user meant by -v, --output report.csv or
commit -m "fix the thing" has to be recovered from that list, and the
recovering is where command-line programs quietly rot. The first
version reads os.args and checks a few positions. The second adds a
flag, and the checks become a chain of conditions. By the fourth nobody
can say what -vo does without running it, the help text has drifted
from the code that reads the flags, and a typo silently means something
else instead of saying so.
The args module takes the declaration instead. You say which options
exist, what type each one carries and which are required; it works out
what the user typed, converts it, validates it, writes the help text
from the same declarations, and refuses anything that does not fit.
- The First Parser
- What
parse()Returns - Options
- Positional Arguments
- Sub-commands
- The Shapes a Command Line Can Take
- Help
- When It Does Not Fit
- Testing a Command Line
- Module Reference
The First Parser
Three lines of declaration and a call:
import args
var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
var parsed = parser.parse(['--name', 'Ada'])
echo parsed.options.name
Ada
parse() with no argument reads the real command line, which is what a
program does. Everywhere in this chapter it is given an explicit list
instead, because that is also how you test one, and because a book
example cannot rely on how you invoked it.
What parse() Returns
One dictionary with three keys, always:
import args
var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_index('source', 'The file to read')
parser.add_command('check', 'Check it over')
echo parser.parse(['-v', 'check', 'notes.txt'])
{options: {verbose: true}, command: {name: check, value: nil}, indexes: [notes.txt]}
options holds every option that was supplied or has a default.
command is nil or the sub-command that was named. indexes holds
the positional arguments in the order they were declared. The three are
independent: a program can use one of them and ignore the rest.
A declared positional that nobody filled and that has no default keeps
its place in indexes as a nil, so the argument after it is still
found at the index it was declared at rather than sliding down one.
Options
An option is declared by its long name. The short name is optional, and so is everything else:
import args
var parser = args.Parser('tool')
parser.add_option('output', 'Where to write', { short_name: 'o', type: args.STRING })
parser.add_option('force', 'Overwrite without asking', { short_name: 'f' })
echo parser.parse(['-o', 'out.csv', '-f']).options
{output: out.csv, force: true}
--output and -o are the same option.
Attaching a Value
A long option takes its value either as the next argument or attached
with =. The two spellings mean the same thing:
import args
def parser() {
var p = args.Parser('tool')
p.add_option('output', 'Where to write', { short_name: 'o', type: args.STRING })
return p
}
echo parser().parse(['--output', 'out.csv']).options.output
echo parser().parse(['--output=out.csv']).options.output
out.csv
out.csv
The split is on the first =, so --output=a=b writes to a=b rather
than losing half the name, and --output= is an explicit empty string
rather than a missing value. The attached form is also the only way to
pass a value that begins with a dash, because in --output -5 the
parser reads -5 as an option:
import args
var parser = args.Parser('seek')
parser.add_option('offset', 'Where to start', { type: args.INT })
echo parser.parse(['--offset=-5']).options.offset
-5
A short option takes its value as the next argument only: -o out.csv,
never -oout.csv and never -o=out.csv. A short token is a bundle of
single-character flags, and neither = nor the letters of a value are
flags.
Types
The type decides what the string becomes and whether a value is taken at all.
| Constant | The option takes | Becomes |
|---|---|---|
args.NONE | nothing | true when present |
args.STRING | a value | the string itself |
args.INT | a value | a whole number |
args.NUMBER | a value | a number, fractions allowed |
args.BOOL | a value | a boolean |
args.LIST | a value, repeatable | a list of strings |
args.CHOICE | a value from a fixed set | the string, or what it maps to |
args.OPTIONAL | a value, if one is there | the string, or true |
NONE is the default and is the plain flag. BOOL is the one that
takes an explicit answer, and it accepts every spelling a user is
likely to reach for:
import args
def parser() {
var p = args.Parser('tool')
p.add_option('colour', 'Use colour', { short_name: 'c', type: args.BOOL })
return p
}
echo parser().parse(['-c', 'yes']).options.colour
echo parser().parse(['-c', 'off']).options.colour
echo parser().parse(['-c', '0']).options.colour
true
false
false
1, true, yes, y and on all mean true; 0, false, no, n
and off all mean false.
Defaults and Absence
value is what the option is worth when nobody supplies it:
import args
var parser = args.Parser('tool')
parser.add_option('count', 'How many', { short_name: 'c', type: args.INT, value: 1 })
parser.add_option('verbose', 'Say more', { short_name: 'v' })
echo parser.parse([]).options
echo parser.parse(['-c', '5', '-v']).options
{count: 1}
{count: 5, verbose: true}
An option with no default and no value on the command line is not in
the dictionary at all. That is the difference between a flag that was
left off and one that was set to false, and it is why verbose is
missing from the first line rather than sitting there as false. Use
options.contains('verbose') to ask, or give the option a default and
stop having to.
Requiring an Option
import args
var parser = args.Parser('tool')
parser.add_option('output', 'Where to write', {
short_name: 'o',
type: args.STRING,
required: true,
})
echo parser.parse(['-o', 'out.csv']).options.output
out.csv
Leave it off and the program stops with error: required option --output is missing, the usage line, and exit status 1. Nothing else
runs.
Restricting the Value
CHOICE with a list accepts only what is in the list:
import args
var parser = args.Parser('tool')
parser.add_option('level', 'How loud', {
type: args.CHOICE,
choices: ['quiet', 'normal', 'loud'],
})
echo parser.parse(['--level', 'loud']).options.level
loud
Anything else stops the program with error: --level expects one of {'quiet', 'normal', 'loud'}, got "shouty", which names the offender
and lists the alternatives without you writing either.
Give it a dictionary instead and the user types the key while the program receives the value. This is how a short spelling on the command line becomes a meaningful value inside:
import args
var parser = args.Parser('tool')
parser.add_option('mode', 'What to do', {
short_name: 'm',
type: args.CHOICE,
choices: { r: 'read', w: 'write', rw: 'read-write' },
})
echo parser.parse(['-m', 'rw']).options.mode
read-write
Collecting Repeats
LIST accumulates every time the option appears:
import args
var parser = args.Parser('tool')
parser.add_option('tag', 'Tag to apply', { short_name: 't', type: args.LIST })
echo parser.parse(['-t', 'urgent', '-t', 'docs', '--tag', 'review']).options.tag
[urgent, docs, review]
Retiring an Option
An option you no longer want but cannot remove yet keeps working and says so on stderr:
import args
var parser = args.Parser('tool')
parser.add_option('old-name', 'Use --name instead', {
type: args.STRING,
deprecated: true,
})
echo parser.parse(['--old-name', 'Ada']).options['old-name']
Ada
The warning goes to stderr, so it reaches the person running the program without contaminating output that something else is reading.
Positional Arguments
A positional argument is declared with add_index, and they fill in
the order they were declared:
import args
var parser = args.Parser('copy')
parser.add_index('source', 'The file to read', { required: true })
parser.add_index('dest', 'Where to write it', { value: 'out.txt' })
echo parser.parse(['in.csv']).indexes
echo parser.parse(['in.csv', 'report.csv']).indexes
[in.csv, out.txt]
[in.csv, report.csv]
They take type, choices, value, required and metavar, the
same as options do. A positional that is neither declared nor expected
is an error rather than something quietly ignored, which is what makes
a mistyped flag fail loudly instead of being swallowed as a filename.
Sub-commands
A program with several jobs gives each one a name and its own options:
import args
var parser = args.Parser('notes')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_command('list', 'Show every note')
parser.add_command('remove', 'Delete a note').
add_option('force', 'Do not ask first', { short_name: 'f' })
echo parser.parse(['-v', 'list'])
echo parser.parse(['remove', '-f'])
{options: {verbose: true}, command: {name: list, value: nil}, indexes: []}
{options: {force: true}, command: {name: remove, value: nil}, indexes: []}
add_command returns the command, so its options chain straight off
it. A command’s own options land in the same options dictionary as
the global ones; command.name is what tells you which job was asked
for.
Order matters, and it is the order every tool of this shape uses: the
parser’s own options come before the command name, and the command’s
options come after it. notes -v list works and notes list -v does
not, because after list the parser is reading list’s options and
-v is not one of them.
A Value of Its Own
A command can take one value directly, the way git commit -m takes a
message:
import args
var parser = args.Parser('notes')
parser.add_command('add', 'Write a note', {
type: args.STRING,
metavar: 'text',
}).add_option('pin', 'Keep it at the top', { short_name: 'p' })
echo parser.parse(['add', 'buy milk', '-p'])
{options: {pin: true}, command: {name: add, value: buy milk}, indexes: []}
The value comes immediately after the command name, before the
command’s own options. metavar is the word the help text shows in
place of the value, so the usage line reads notes add <text> rather
than notes add <value>.
Running Something Directly
A command can carry the function that implements it, called after parsing with the options and the command’s value:
import args
var parser = args.Parser('notes')
parser.add_command('add', 'Write a note', {
type: args.STRING,
metavar: 'text',
action: @(options, value) {
var prefix = options.contains('pin') ? '[pinned] ' : ''
echo prefix + value
},
}).add_option('pin', 'Keep it at the top', { short_name: 'p' })
parser.parse(['add', 'buy milk', '-p'])
[pinned] buy milk
This is worth reaching for once there are more than two or three
commands, because the alternative is a chain of comparisons on
command.name that has to be kept in step with the declarations by
hand.
The Shapes a Command Line Can Take
Users type things the parser was not asked about directly. These are the conventions it honours.
Bundling
Several flags behind one dash:
import args
var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_option('quiet', 'Say less', { short_name: 'q' })
echo parser.parse(['-vq']).options
{verbose: true, quiet: true}
Every character in a bundle has to name an option. -vqz is an error
rather than -vq with the z quietly dropped, which is the difference
between a typo you find now and one you find after the run.
Abbreviation
A long option may be shortened to any prefix that is still unambiguous:
import args
var parser = args.Parser('tool')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
echo parser.parse(['--verb']).options
{verbose: true}
Set parser.allow_abbrev = false to turn this off, which is worth
doing for a program whose option set is still growing: a prefix that is
unambiguous today stops being so the day somebody adds --version, and
a script written against the short form breaks.
End of Options
A bare -- ends option parsing. Everything after it is positional,
even when it starts with a dash:
import args
var parser = args.Parser('run')
parser.add_option('verbose', 'Say more', { short_name: 'v' })
parser.add_index('command', 'What to run')
parser.add_index('argument', 'What to pass it')
echo parser.parse(['-v', '--', 'grep', '-i']).indexes
[grep, -i]
This is how a program passes arguments through to another one without having to know what they mean.
Arguments From a File
A token beginning with @ is replaced by the contents of that file,
one argument per line:
import args
file('greet.args', 'w').write('--name\nAda\n--count\n3\n')
var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { type: args.STRING })
parser.add_option('count', 'How many times', { type: args.INT, value: 1 })
echo parser.parse(['@greet.args']).options
echo parser.parse(['@greet.args', '--name', 'Grace']).options
{name: Ada, count: 3}
{name: Grace, count: 3}
The file is expanded in place, so anything after it still wins. This is
the escape hatch for a command line that has outgrown what a shell will
comfortably hold, and for arguments a program would rather not put
where ps can read them. Set parser.allow_atfile = false to turn it
off for a program that should treat a leading @ as ordinary text.
Help
The help text is written from the declarations, so it cannot drift away from what the program accepts:
import args
var parser = args.Parser('greet')
parser.set_terminal_width(68)
parser.description = 'Greet somebody, once or several times.'
parser.epilog = 'Set NO_COLOR to turn off colour.'
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
parser.add_option('count', 'How many times', { short_name: 'c', type: args.INT, value: 1 })
parser.add_index('output', 'Where to write the greeting')
parser.add_command('history', 'Show past greetings')
parser.help()
Usage: greet [OPTIONS] [COMMAND] [output]
Greet somebody, once or several times.
POSITIONAL ARGUMENTS:
[output] Where to write the greeting
OPTIONS:
-h, --help Show this help message and exit
-n, --name <name> Who to greet
-c, --count <count> How many times (default: 1)
COMMANDS:
history Show past greetings
Set NO_COLOR to turn off colour.
Run "greet --help [COMMAND]" for help on a specific command.
-h and --help are declared for you and handled wherever they
appear, including after a command, where they describe that command
instead of the whole program. help() is the same thing called
directly; both print and exit 0.
Colour is used when stdout is a terminal and NO_COLOR is unset, so a
run whose output is piped or redirected gets plain text rather than
escape codes. set_terminal_width() fixes the wrapping width, which is
what the example above does so the output is the same on every
terminal; left alone, the parser asks the terminal and falls back to
COLUMNS, then to 80.
A parser whose command line matched nothing at all prints its help and
carries on. An option with a default counts as a match, so a parser
that defaults anything never does this; pass false as the second
argument to args.Parser to turn the behaviour off outright.
When It Does Not Fit
Every failure follows the same shape: a line on stderr beginning
error:, the usage text, and exit status 1.
| What happened | What it says |
|---|---|
| an option nobody declared | unknown option: --colour |
| a value-taking option at the end | option --name expects <name> |
| a required option left off | required option --output is missing |
a value outside choices | --level expects one of {'quiet', 'loud'}, got "shouty" |
| a required positional left off | required positional argument <source> is missing |
| an argument that fits nothing | unexpected argument: report.csv |
None of these raise. A command-line program that has been handed
something it cannot use has nothing useful left to do, and unwinding a
stack trace into a user’s terminal tells them less than one line does.
The errors that do raise are the ones in the declarations — a
duplicate option name, a short name already taken, a choices that is
neither a list nor a dictionary — because those are bugs in the program
rather than mistakes by its user, and they raise ArgsError or a
TypeError at the add_option that caused them.
Testing a Command Line
parse() takes an explicit list, and that is the seam:
import args
def build() {
var parser = args.Parser('greet')
parser.add_option('name', 'Who to greet', { short_name: 'n', type: args.STRING })
parser.add_option('count', 'How many times', { short_name: 'c', type: args.INT, value: 1 })
return parser
}
var parsed = build().parse(['-n', 'Ada', '-c', '3'])
assert parsed.options.name == 'Ada', 'the name is read'
assert parsed.options.count == 3, 'the count is coerced to a number'
assert build().parse([]).options.count == 1, 'the default applies'
echo 'all good'
all good
Build the parser in a function so each test gets a clean one; a parser
carries the results of the last parse, and sharing one between tests
makes them depend on their order. The paths that end in os.exit() —
help, and every error above — need a real subprocess to observe, which
is what os.exec() is for; Chapter 23 covers the
rest of testing.
Module Reference
Building a parser:
Parser(name, default_help) | a new parser |
add_option(name, help, opts) | an option, global or on a command |
add_command(name, help, opts) | a sub-command, returned for chaining |
add_index(name, help, opts) | a positional argument |
parse(custom_args) | read the command line, or a list |
help() | print the help text and exit 0 |
set_terminal_width(width) | fix the wrapping width |
Properties you can set after construction:
description, epilog | prose above and below the options |
allow_abbrev | unambiguous long-option prefixes, default true |
allow_atfile | @file expansion, default true |
terminal_width | the wrapping width directly |
Keys opts understands:
short_name | the single-character form |
type | one of the type constants |
value | what it is worth when absent |
choices | a list of allowed values, or a map from key to value |
required | refuse to run without it |
metavar | the word help shows in place of the value |
deprecated | warn on stderr when it is used |
action | on a command, the function to run after parsing |
The type constants:
NONE, STRING, INT, NUMBER | a flag, and the three plain values |
BOOL, LIST, CHOICE, OPTIONAL | an answer, a repeat, a fixed set, a maybe |
Metaprogramming and Reflection
Zuri’s own compiler is available to Zuri programs. The zuri module hands
you the lexer, the parser and the bytecode compiler as ordinary functions,
alongside a reflection API over live functions, classes, modules and
instances.
That is an unusual amount of the language to expose, so it is worth being clear about what it is for. This is the machinery behind documentation generators, linters, plugin loaders, serialisers, debuggers and editor tooling — programs whose subject is other programs. It is not machinery for ordinary code, and the chapter ends with the one thing it deliberately cannot do.
Two Halves That Cannot Answer for Each Other
The module divides cleanly, and confusing the halves is the most common mistake:
| Reflection | The compiler API | |
|---|---|---|
| Subject | objects that exist right now | source text |
| Entry points | zuri.reflect.* | tokenize(), parse(), compile() |
| Can tell you | a function’s arity, a class’s fields, an instance’s values | a doc block, a comment, a line number, what the compiler emitted |
| Cannot tell you | anything about the source it came from | anything about a running value |
A function’s arity is a runtime fact: the function object carries it, and reflection reads it. That same function’s doc block is a source fact: it exists only in the file, and only the parser can find it. There is no call that crosses the gap, which is why a documentation generator parses files rather than importing them.
One Shape for Everything
tokenize(), parse() and compile() all return the same shape of
result: a flat list of small, uniformly tagged records.
| Function | Returns | Tag | Payload |
|---|---|---|---|
tokenize(source) | Token list | kind | value, text |
parse(source) | Node list | kind | fields |
compile(source) | Instr list | op | fields |
There is one Token class, one Node class and one Instr class — not
one class per token kind, grammar rule or opcode. The grammar has around
fifty shapes and the instruction set around sixty opcodes; a class for each
would be a hundred-odd near-identical classes to keep in step with the
compiler forever.
The cost of that choice is real: nothing tells you which fields a given
kind carries. You cannot autocomplete your way to a Binary node’s
left, op and right. So the module leans on a different habit — run
it and look:
import zuri
echo zuri.parse('var total = a + 1')[0].dump()
Stmt@1:1
statement:
Var@1:1
name: 'total'
name_span: 1:5-1:10
value:
Binary@1:13
left:
Identifier@1:13
name: 'a'
op: 'Plus'
right:
Integer@1:17
value: 1
type_hint: nil
is_constant: false
Everything about that node is in front of you: it is a Stmt wrapping a
Var, the initialiser is a Binary whose op is the string 'Plus', and
the literal 1 is an Integer node rather than a generic literal. Each
header names where that node starts, and name_span says exactly where the
name total is. Nothing had to be looked up.
dump() is the fastest way to learn any node’s shape, and it is the first
thing to reach for whenever you are unsure. zuri.dump_file(path) does the
same for a whole file.
Runtime Reflection
Reflection answers questions about a value you are holding. Every entry
point is under zuri.reflect.
What Kind of Thing Is This?
import zuri
def add(a, b) {
return a + b
}
class Marker {}
echo zuri.reflect.kind(add)
echo zuri.reflect.kind(Marker)
echo zuri.reflect.kind(Marker())
echo zuri.reflect.kind(42)
echo zuri.reflect.kind(print)
function
class
instance
number
function
kind() is typeof()’s more literal cousin. Note the last line: a
built-in native function reports 'function', as does a closure and a
bound method, because from the language’s point of view all three are
interchangeably callable.
Functions
import zuri
def add(a, b) {
return a + b
}
var info = zuri.reflect.function_info(add)
echo info.name
echo info.arity
echo info.variadic
echo info.is_method
echo info.owning_class_name
add
2
false
false
nil
function_info() returns { name, arity, variadic, is_method, owning_class_name, source_path }. The last one is the only field that
reaches back toward the source, and it is a path, not content — to read
the function’s doc block you still have to parse that file.
Classes
import zuri
class Shape {
static var sides = 0
@new(name) {
self.name = name
}
area() {
return 0
}
_secret() {
return 'hidden'
}
@to_json() {
return { name: self.name }
}
}
var info = zuri.reflect.class_info(Shape)
echo info.name
echo info.superclass_name
echo info.fields
echo info.statics
echo info.methods.keys()
Shape
nil
[name]
[sides]
[@new, area, @to_json, _secret]
Four things in that output are worth pausing on.
fields lists name, which was never declared with var — it was
created by self.name = name inside @new. Reflection reports the class’s
real shape, not just what the var lines said.
statics is separate from fields, because a static belongs to the
class and a field belongs to each instance.
Decorated methods appear under their @ names. @new and @to_json
are in methods alongside ordinary ones.
_secret is listed. Reflection sees private members; more on that
below.
Each entry in methods is a full function_info, so
info.methods.area.arity works without another call.
superclass_name is a string or nil, not a class:
import zuri
class Base {}
class Derived < Base {}
echo zuri.reflect.class_info(Derived).superclass_name
echo zuri.reflect.class_info(Base).superclass_name
Base
nil
Instances
import zuri
class Point {
@new(x, y) {
self.x = x
self.y = y
}
distance() {
return (self.x ** 2 + self.y ** 2).sqrt()
}
}
var p = Point(3, 4)
echo zuri.reflect.get_props(p)
echo zuri.reflect.has_prop(p, 'x')
echo zuri.reflect.get_prop(p, 'x')
zuri.reflect.set_prop(p, 'x', 10)
echo p.x
[x, y]
true
3
10
get_props() lists the field names an instance actually has.
get_prop(), set_prop(), has_prop() and del_prop() read and write
them by computed name — the reflective equivalents of the getprop family
of built-ins from Chapter 6.
They cannot create a field. A class is sealed, so set_prop() on a name
the class never declared returns false and changes nothing.
Methods by Name
Three calls turn a method name into something callable:
import zuri
class Greeter {
@new(name) {
self.name = name
}
greet() {
return 'hello ${self.name}'
}
@to_json() {
return { name: self.name }
}
}
var g = Greeter('ada')
echo zuri.reflect.has_method(g, 'greet')
echo zuri.reflect.has_decorator(g, 'to_json')
echo typeof(zuri.reflect.get_method(g, 'greet'))
echo zuri.reflect.bind_method(g, 'greet')()
true
true
function
hello ada
The distinction between the last two matters. get_method() returns the
unbound function; calling it needs the instance supplied yourself.
bind_method() returns it already attached to the instance, so it can
be stored, passed around and called with nothing extra — which is what
makes a dispatch table of methods possible:
import zuri
class Counter {
@new() {
self.n = 0
}
up() {
self.n++
return self.n
}
down() {
self.n--
return self.n
}
}
var counter = Counter()
var actions = {
up: zuri.reflect.bind_method(counter, 'up'),
down: zuri.reflect.bind_method(counter, 'down'),
}
echo actions['up']()
echo actions['up']()
echo actions['down']()
echo counter.n
1
2
1
1
has_decorator() and get_decorator() do the same for @-prefixed
methods, taking the name without the @.
Reflection Sees Private Members
Everything above ignores the leading-underscore rule:
import zuri
class Vault {
var _combination = '1234'
@new() {}
}
echo zuri.reflect.get_props(Vault())
echo zuri.reflect.get_prop(Vault(), '_combination')
[_combination]
1234
This is deliberate, and it is the same decision getprop() makes. A
serialiser has to see every field or it writes an incomplete record; a
debugger has to see every field or it shows a lie. Hiding them would defeat
the only purpose these functions have.
The rule to take from it: reflection is infrastructure, not a way around
encapsulation. Code that reaches for get_prop() to read a private field
it could not otherwise read has not found a loophole; it has written
something the next reader will not expect.
Modules
import zuri
import math
var info = zuri.reflect.module_info(math)
echo info.name
echo info.members.contains('PI')
math
true
module_info() reports a module’s name, its path and the members it
exports. That is enough for a plugin loader: import a directory of modules,
ask each whether it has a known function, and call the ones that do.
The Collector
The garbage collector runs on its own as a program allocates. gc() runs
a full collection on the spot:
import zuri
var scratch = [1, 2, 3]
scratch = nil
zuri.reflect.gc()
echo 'collected'
collected
Every object nothing reaches is freed before gc() returns, and anything
that releases a resource when collected releases it then: an ffi pointer
taken over with own() runs its destructor. The memory it frees goes back
to the system before it returns too. No program needs gc() to
stay correct. It is for the moments timing matters, such as releasing
native resources at a known point, or a test checking what collection
does. A full collection visits every live object, so calling it in a loop
is slow.
Tokens
tokenize() is the first stage: text in, a flat list of lexical tokens
out.
import zuri
for token in zuri.tokenize('var x = 1 # note') {
echo '${token.kind} at ${token.line}:${token.column} -> ${token.text}'
}
Var at 1:1 -> var
Identifier at 1:5 -> x
Equal at 1:7 -> =
Integer at 1:9 -> 1
Comment at 1:11 -> # note
Eof at 1:17 ->
Nothing is filtered. Comments, doc blocks and newlines are all real
tokens, and the list always ends with one Eof. The parser’s grammar skips
trivia when it runs; tokenize() reports everything the lexer saw, which
is exactly what a formatter or a syntax highlighter needs.
Each Token carries:
| Field | What it is |
|---|---|
kind | the lexer’s own variant name — 'Identifier', 'Plus', 'Comment' |
line, column | 1-indexed position; column counts characters, not bytes |
start, end | character offsets spanning the token’s exact text |
text | the exact source text, delimiters included |
value | the payload, where there is one — a literal’s value, an identifier’s name, a comment’s content |
start and end are what make edits possible: they let you recover or
replace a token’s original text without re-lexing, which is how a rename
tool or an automatic formatter works.
is_trivia() is true for comments and doc blocks, so filtering them is one
call:
import zuri
var tokens = zuri.tokenize('var x = 1 # note')
echo tokens.filter(@(t) => t.is_trivia()).map(@(t) => t.kind)
echo tokens.filter(@(t) => !t.is_trivia()).length()
[Comment]
5
Lexing Never Raises
A malformed input does not throw. It produces a token of kind 'Error'
carrying what went wrong:
import zuri
var tokens = zuri.tokenize("var s = 'unterminated")
echo tokens.map(@(t) => t.kind)
echo tokens.find(@(t) => t.kind == 'Error') != nil
[Var, Identifier, Equal, Error, Eof]
true
tokenize() is therefore total over any input at all, which is what
makes it safe to point at a file you did not write, or at a half-typed
buffer in an editor. Neither parse() nor compile() has that property —
both raise on bad input, because neither can produce a meaningful result
from it.
The Syntax Tree
parse() is the second stage: tokens become a tree.
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
echo tree[0].kind
echo tree[0].fields.keys()
echo tree[0].fields.name
echo tree[0].line
Function
[name, name_span, parameters, body, is_variadic]
double
1
A Node has a kind, a position and a fields dictionary. The fields hold
more nodes, plain lists, plain dictionaries or scalars and nothing else, so
walking the tree is uniform no matter which node you are looking at.
Where Each Node Is
A node’s position covers the whole of the source it was read from.
start and end are character offsets, the first character and the one
just past the last, so slicing the source with them gives back the node’s
own text. line and col are where it starts and end_line and end_col
where it ends, all counted from 1:
import zuri
var source = 'def double(n) {\n return n * 2\n}'
var function = zuri.parse(source)[0]
var product = function.fields.body.fields.statements[0].fields.value
echo source[product.start, product.end]
echo '${function.line}:${function.col} to ${function.end_line}:${function.end_col}'
echo source[function.fields.name_span.start, function.fields.name_span.end]
n * 2
1:1 to 3:2
double
A declaration starts at its keyword, so the function starts at def, and
its name has a name_span of its own. That is everything a tool needs to
underline a mistake, jump to a definition or select a whole statement.
Walking It
walk_nodes(tree, visitor) visits every node, depth first:
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
var kinds = []
zuri.walk_nodes(tree, @(node) {
kinds.append(node.kind)
})
echo kinds
[Function, Argument, TypeHint, Any, Block, Return, Binary, Identifier, Integer]
That output repays a second look, because it shows how much the parser
makes explicit. The single unannotated parameter n still produced an
Argument wrapping a TypeHint wrapping an Any — the absence of an
annotation is represented in the tree, not omitted from it. A tool walking
this never has to special-case “no type was written”.
The visitor may also be a dictionary from kind to handler, in which case only matching nodes are called:
import zuri
var tree = zuri.parse('def a() {}
def b() {}
var c = 1')
var names = []
zuri.walk_nodes(tree, {
Function: @(node) {
names.append(node.fields.name)
},
})
echo names
[a, b]
find_nodes(tree, kind) is the shortcut when you want one kind and nothing
else:
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
echo zuri.find_nodes(tree, 'Return').length()
echo zuri.find_nodes(tree, 'Binary')[0].fields.op
1
Multiply
Comments Survive
This is the property that makes documentation tooling possible. Comments and doc blocks are kept in the tree, in their original position, as their own nodes:
import zuri
var source = "# a note
def add(a, b) {
return a + b
}
"
echo zuri.parse(source).map(@(n) => n.kind)
[Comment, Function]
A doc block sits as a sibling immediately before whatever it documents, so pairing them is one pass with one variable of state — which is exactly what the worked example below does.
Most parsers throw comments away. Keeping them is what separates a parser you can build a formatter or a documentation generator on from one you can only build an interpreter on.
Bytecode
compile() is the third stage: the tree becomes VM instructions.
import zuri
for instr in zuri.compile('var a = 1 + 2') {
echo '${instr.op} from line ${instr.line}'
}
LoadConst from line 0
AddImm from line 1
SetGlobal from line 1
LoadNil from line 1
Return from line 1
An Instr has an op, a line and a fields dictionary — the same shape
a Node has — and walk_instrs() and find_instrs() mirror the AST
walkers exactly.
Notice AddImm. The compiler folded the constant 2 into the add
instruction rather than loading it separately. That is the sort of question
only the bytecode can answer, and reading it is the most direct way to find
out what the compiler actually made of something:
import zuri
echo zuri.compile('var a = 2 * 3 ** 2').map(@(i) => i.op)
[LoadConst, LoadConst, LoadConst, Pow, Mul, SetGlobal, LoadNil, Return]
The power comes before the multiply, which is ** binding tighter than
*. One line of bytecode settles an argument that reading the expression
does not.
Resolved Operands
Several opcodes carry only a raw index into the compiler’s constant pool, which on its own tells you nothing. Rather than make every caller fetch and correlate a constants table, each such field arrives with its resolved value alongside the index:
import zuri
var load = zuri.find_instrs(zuri.compile('var a = 42'), 'LoadConst')[0]
echo load.fields.keys()
echo load.fields.value
[dst, const_idx, value]
42
const_idx is the raw index, value is what it points at, and dst is
the register the result lands in. The same
pairing appears on GetGlobal, GetField, Invoke and the rest.
A Closure instruction’s resolved value is a nested function prototype
({ name, arity, variadic, instructions }) whose own instructions are
wrapped the same way, recursively — so compiling a file with functions in
it exposes their bodies too, not only the code around them.
Line Numbers Are Statement-Grained
Every instruction carries the source line it came from, but never a
column. The compiler’s own tracking is line-only. A Token and a Node
both have real column information from the lexer; an Instr does not.
Checking Source Without Running It
parse() and compile() raise on the first sign of bad source, which is
right for a tool that needs the whole program. A tool working on source as
it is being written, half a statement at a time, needs the opposite:
everything that can be read, and an exact account of what is wrong.
parse_partial() reads as much as it can. Where a statement fails, the
parser reports it, skips to where the next statement starts and carries
on, so one mistake costs one statement and is reported once:
import zuri
var result = zuri.parse_partial('var a = 1
var = 2
var c = 3')
echo result.is_clean()
echo result.errors
echo result.nodes.map(@(n) => n.fields.statement.fields.name)
false
[2:5: Variable name expected.]
[a, , c]
Each error is a Diagnostic: a message and the same six-part position a
node has, so it can be underlined exactly. The statement that failed is
still in the tree, its missing name an empty string sitting where the name
belongs.
check() goes one stage further. Source that parses goes on to the
compiler, which catches what the grammar cannot see, such as a break
with no loop around it or a name declared twice in one scope:
import zuri
var source = 'def total(items) {
var sum = 0
for item in items {
sum += item
}
break
}'
for problem in zuri.check(source) {
echo problem
echo source[problem.start, problem.end]
}
6:3: 'break' used outside of a loop
break
An empty list from check() means the source compiles. Neither function
runs anything or raises on bad source, so both can be pointed at whatever
an editor’s buffer holds at the time.
Reading a File
Each of these has a _file counterpart that reads the path first:
| Source string | File |
|---|---|
tokenize(source) | tokenize_file(path) |
parse(source) | parse_file(path) |
compile(source) | compile_file(path) |
parse_partial(source) | parse_partial_file(path) |
check(source) | check_file(path) |
There is a fourth with no string equivalent: dump_file(path) reads,
parses and returns the dump() of every top-level node. It is usually the
fastest possible answer to “what is actually in this file”.
A Worked Example: A Documentation Extractor
Here is the pattern the book’s own reference appendices are built on. It takes source text and returns every documented function in it, pairing each declaration with the doc block above it.
The key fact is that zuri.parse() keeps comments in the tree. A doc block
comes back as its own DocBlock node, sitting as a sibling immediately
before whatever it documents. zuri.doc.attach() walks a list of siblings
and pairs each one with the doc block right above it, and zuri.doc reads
the inside of each block into its prose and its @tags:
import zuri
def documented_functions(text) {
var found = []
for attached in zuri.doc.attach(zuri.parse(text)) {
var node = attached.node
if node.kind == 'Function' and attached.documented {
var params = node.fields.parameters.map(@(p) => p.fields.name)
found.append({ name: node.fields.name, params, doc: attached.block })
}
}
return found
}
var source = "/**\n * Adds two numbers.\n *\n * @param number a\n * @param number b\n */\n" +
"def add(a: number, b: number) {\n return a + b\n}\n\n" +
"def undocumented(x) {\n return x\n}\n"
for fn in documented_functions(source) {
var takes = fn.doc.all('param').map(@(tag) => '${tag.name}: ${tag.type}')
echo '${fn.name}(${', '.join(fn.params)})'
echo ' ' + fn.doc.summary()
echo ' takes ' + ', '.join(takes)
}
add(a, b)
Adds two numbers.
takes a: number, b: number
undocumented is absent from the output: attach() hands it back with
documented false, because no doc block sits above it.
Three things make this approach worth preferring over scanning the text yourself.
The parser knows what a declaration is. A function named def_handler,
a def inside a string literal, a doc block inside a block comment — all of
them fool a text scan and none of them fool the parser.
Parameter names and types come from the real nodes. A parameter’s
type_hint node carries its types and whether it is nullable, so a
generated signature matches what the runtime will actually enforce rather
than what the source happened to look like.
It cannot drift. When the grammar gains something, the parser gains it too, and a tool built this way keeps working.
The tags are read for you. zuri.doc joins a tag’s wrapped lines,
reads its type and the name it documents, and treats @return as
@returns. It is the same reading the book’s own reference is generated
with, so a block reads the same here as it does there.
A Second Example: A Linter
The other half of the module’s use is checking rather than generating. Here is a rule of the kind a team accumulates: flag every function whose parameter list is longer than some limit.
import zuri
def long_signatures(source, limit) {
var offenders = []
zuri.walk_nodes(zuri.parse(source), {
Function: @(node) {
var count = node.fields.parameters.length()
if count > limit {
offenders.append({ name: node.fields.name, line: node.line, count })
}
},
})
return offenders
}
var source = "def small(a, b) {}\n" +
"def large(a, b, c, d, e) {}\n" +
"def also_large(a, b, c, d) {}\n"
for problem in long_signatures(source, 3) {
echo 'line ${problem.line}: ${problem.name}() takes ${problem.count}'
}
line 2: large() takes 5
line 3: also_large() takes 4
Three properties make this worth doing with the parser rather than with a regular expression over the text.
It cannot be fooled by text that looks like code. A function named
def_handler, the word def inside a string literal, a commented-out
declaration — none of them are Function nodes, so none of them are
counted.
It reports the real line. node.line comes from the lexer, so the
message points at the declaration whatever the formatting around it.
It keeps working. When the grammar gains something, the parser gains it too, and a rule written this way does not quietly stop matching.
The dictionary form of the visitor is doing real work here: only Function
nodes reach the handler, so there is no if node.kind == ... and no chance
of matching a kind you did not mean to.
Serialising Without a @to_json
Reflection covers the case where you would otherwise write the same method on every class:
import zuri
import json
class User {
@new(name, age) {
self.name = name
self.age = age
}
}
class Product {
@new(title, price) {
self.title = title
self.price = price
}
}
def to_record(instance) {
var record = {}
for field in zuri.reflect.get_props(instance) {
record[field] = zuri.reflect.get_prop(instance, field)
}
return record
}
echo json.encode(to_record(User('ada', 36)))
echo json.encode(to_record(Product('desk', 120)))
{"name":"ada","age":36}
{"title":"desk","price":120}
One function, every class, no per-class method to keep in step with the
fields. Note that this reads private fields too, which is right for a
debugging dump and wrong for an API response — for the latter, filter on
the leading underscore, or write a real @to_json and let the class decide
what it exposes.
What Else This Is Good For
Plugin systems. Load modules from a directory, ask each whether it has
the function your host expects, and call the ones that do. module_info()
answers the question without importing blindly and hoping.
Editor tooling. tokenize(), parse_partial() and check() never
raise, so they can run against a half-typed buffer. Every token and every
node carries start and end, so a rename or a reformat can rewrite exact
spans, and every problem check() finds can be underlined where it is.
Understanding the compiler. zuri.compile() on a snippet is faster
than reading the compiler’s source, and it can never go out of date with
the compiler you are actually running.
Answering questions about this book. Appendices D and E are generated
from libs/ with exactly the techniques above, and the audit that checks
those stubs is written the same way.
What This Is Not Good For
There is no eval(). You can compile source to bytecode and inspect it;
you cannot execute a string as code.
Zuri leaves it out for security. eval() erases the line between data and
code, and every program that has one eventually runs a string it did not
mean to: a form field, a query parameter, a config value, a webhook
payload. The moment any of those reaches an eval(), whoever supplied it
is running code with your program’s full permissions. It reads your files,
opens your sockets and sends your secrets anywhere it likes. This is the
single most damaging vulnerability class in dynamic languages, and it keeps
happening because the dangerous call looks harmless in review: one
function, three characters of input, and nothing in the language warns you.
So Zuri does not provide one, and every job eval() is usually reached for
has a better answer here:
| Instead of | Use |
|---|---|
| evaluating a user-supplied expression | a dictionary of handlers keyed by the input |
| calling a method whose name you computed | zuri.reflect.bind_method() |
| varying behaviour at runtime | pass a function in |
| turning text into data | json.decode() |
Each of those does the job with a fixed, auditable set of things that can happen. That is the difference between a program that handles input and a program that obeys it.
Classes are sealed, so reflection cannot add a method to one at runtime. Patterns from other languages that rely on monkey-patching do not translate, and the alternative is the one the language wants: express the variation as a subclass, or as a function you pass in.
Performance and the JIT
Zuri runs your bytecode in an interpreter until a piece of it gets hot, then compiles that piece to machine code and runs it there instead. Two compilers share the work, one quick and one thorough, and a hot function passes through both. The tiering happens on its own, on background threads, with no annotations and no flags.
Here is what that is worth on a plain recursive Fibonacci, measured on one idle machine:
def fib(n) {
if n < 2 {
return n
}
return fib(n - 1) + fib(n - 2)
}
var start = time()
echo fib(27)
echo 'took ${((time() - start) * 1000).round()}ms'
$ zuri run fib.zu
196418
took 14ms
$ ZURI_JIT=0 zuri run fib.zu
196418
took 47ms
Over three times, for a function that does nothing but add and compare, and you write nothing to get it. The rest of this chapter explains how the compilers work and covers the handful of cases where what you write decides whether you get it.
The Two Compilers
Kebbi compiles first. It turns a function’s bytecode into machine code
one instruction at a time, keeps values in machine registers, and proves
what it can about types before it compiles: a parameter annotated list is
a list, and a counter that starts at 0 and steps by 1 is a whole number.
Kebbi compiles quickly, and its code runs several times faster than the
interpreter.
Bayelsa compiles second. It builds the whole function as a graph of typed values, then optimizes that graph before it generates any machine code. It carries a fact proven once to every later use, moves checks that cannot change out of loops, keeps loop counters as plain integers, builds small functions into the functions that call them, and bets on what a function has done so far wherever nothing proves it. Bayelsa takes longer to compile, and its code is the fastest Zuri produces.
Both compile with Cranelift, and both run on background threads. A program never waits for a compile: it carries on in the interpreter, or in the code it already has, until the new code is ready.
How a Function Moves Between Them
- The interpreter. Every function starts here. It counts its calls and its loop turns, and records the kind of value each operation meets.
- Kebbi, profiling. Once a function is warm, Kebbi compiles it with the counting and recording built in. The function is fast from here on, and still watching itself.
- Bayelsa. Once the profiling code has done enough work, Bayelsa compiles the function from what it recorded. A loop running at that moment moves into the new code at its next turn.
A function every operation of which has already run in the interpreter skips the profiling step and goes straight to Bayelsa, since profiling would only record what the interpreter already knows.
When a Bet Fails
Bayelsa compiles what it has seen. A loop that only ever added whole
numbers does integer arithmetic; a module constant that held 10 is
compared as the integer 10; an index that only ever met bytes reads a
byte. Each bet is checked where it is made, and a failed check hands the
frame back to the interpreter at that exact instruction, with every
variable as the compiled code left it. This is a deoptimisation.
The place that failed is remembered. The next compile of the function makes no bet there, so a function deoptimises at a given place once, not every time round.
What Stays in Kebbi
Bayelsa leaves a function to Kebbi in two cases:
- The function contains a
catch. Kebbi compiles it, handlers included. - The function’s hot operations are ones Kebbi handles inline and Bayelsa would hand to the runtime: arithmetic on values that were not always numbers, method calls on strings, and indexing lists and strings whose kind nothing settles. Kebbi’s code is the faster of the two there.
A function left in Kebbi is rebuilt without its profiling and runs at Kebbi’s full speed.
How Tiering Works
Every function counts its own calls, and every loop counts its own back-edges. When a counter crosses that function’s threshold, a compile job goes to a background thread; when the machine code comes back, later calls use it.
The threshold is not a fixed number. It scales with the function’s size:
threshold = K / sqrt(instruction_count)
A large function gets a low threshold, because it is doing more work per call and the compilation pays for itself sooner. A three-instruction getter gets a high one, because compiling it is speculative and most tiny functions never run enough to earn it back.
Loops get their own, lower threshold through on-stack replacement. A loop inside a function that is only ever called once still compiles, and execution jumps from the interpreter into the middle of the compiled version without unwinding anything. That is what makes a one-shot batch script fast.
You can watch it happen:
$ ZURI_JIT_LOG=1 zuri run fib.zu
[jit] 'fib' goes straight to Bayelsa
[jit] compiled 'fib' in Bayelsa (13 bytecode ops, 0 osr point(s))
196418
fib goes straight to Bayelsa because every one of its operations ran in
the interpreter before it warmed up. With Bayelsa off, Kebbi compiles it
instead, and the line lists what Kebbi proved:
$ ZURI_JIT_BAYELSA=0 ZURI_JIT_LOG=1 zuri run fib.zu
[jit] compiled 'fib' in Kebbi (13 bytecode ops, 0 osr point(s), speculative_params=0x1, speculative_regs=0x0, speculative_lists=0x0, speculative_ints=0x1)
196418
A function Bayelsa leaves to Kebbi gets a line saying why:
[jit] 'report' stays in Kebbi: catch at ip 5
And you can see what did and did not make it:
$ ZURI_JIT_COVERAGE=1 zuri run fib.zu
=== jit coverage ===
status calls threshold ops osr name
compiled 62009 169 11 0 fib
cold 0 56 100 0 @.script
...
calls against threshold tells you how close something came.
Keep Errors for Errors
A function with a catch compiles like any other. While no error comes
through, a catch costs next to nothing: the handler is registered and
dropped around the body, and the body runs as compiled code.
What costs is an error that is actually raised. A raise hands the frame
back to the interpreter at that instruction, and a caught error resumes at
its handler in the interpreter too. The function returns to compiled code
the next time its loop comes round, or the next time it is called. For a
genuine failure that is the right trade. For an outcome a loop meets every
few iterations, it is not:
def parse_or_raise(i) {
if i % 10 == 0 {
raise ValueError('not a digit')
}
return i % 10
}
def parse_or_nil(i) {
if i % 10 == 0 {
return nil
}
return i % 10
}
def with_raise(n) {
var total = 0
iter var i = 0; i < n; i++ {
catch {
total += parse_or_raise(i)
} as e {
total += 0
}
}
return total
}
def with_nil(n) {
var total = 0
iter var i = 0; i < n; i++ {
var digit = parse_or_nil(i)
if digit != nil {
total += digit
}
}
return total
}
def timed(label, work) {
var start = time()
work()
echo '${label}: ${((time() - start) * 1000).round()}ms'
}
timed('raise and catch', @() {
iter var r = 0; r < 200; r++ {
with_raise(1000)
}
})
timed('return nil ', @() {
iter var r = 0; r < 200; r++ {
with_nil(1000)
}
})
Two hundred rounds of a thousand iterations, one in ten of them failing:
| Variant | Time |
|---|---|
raise, caught in the loop | 73ms |
nil returned and checked | 11ms |
When failing is an expected outcome, input that might not parse or a key
that might be missing, return a value that says so and test it. Keep
raise for what the caller cannot reasonably carry on from. A raise that
never fires, a guard clause at the top of a function, costs the compiled
path nothing.
Type Annotations Are a Performance Feature
The JIT speculates on the types it has seen. When it has been told instead, it does not need to guess, and the guard disappears.
def sum_untyped(values) {
var total = 0
iter var i = 0; i < values.length(); i++ {
total += values[i]
}
return total
}
def sum_typed(values: list) {
var total = 0
iter var i = 0; i < values.length(); i++ {
total += values[i]
}
return total
}
Twenty thousand rounds over a thousand-element list:
| Variant | Time |
|---|---|
values: list | 85ms |
| no annotation | 150ms |
Running them in both orders gives the same answer, which is the check that tells you it is the annotation and not the warm-up.
Annotate the parameters of any function on a hot path. It costs one word per parameter and it documents the function at the same time.
Keep Types Stable
Speculation works by assuming the future looks like the past. A variable that holds a number on ten thousand iterations and a string on the ten thousand and first causes a deoptimisation: the compiled code bails to the interpreter, and the function may be recompiled with a weaker assumption.
One deopt is cheap. A loop that deopts every iteration is slower than never compiling at all.
In practice this means:
- One variable, one kind of value.
- A list of numbers, not a list of numbers-and-sometimes-strings.
- A field that starts
niland later holds a number is two types. Start it at0.
That last one is the common case. var count on a class field is a nil
until the constructor runs, and every read of it in compiled code has to
handle both. var count = 0 does not.
Where to Put a Hot Loop
A per-element loop belongs in a typed free function, not in a method reading a field:
Slower, because every iteration reads self.pixels back through a field
guard:
class Canvas {
var pixels = bytes(0)
brighten(amount) {
iter var i = 0; i < self.pixels.length(); i++ {
self.pixels[i] = self.pixels[i] + amount
}
}
}
Faster, because the loop sees plain locals whose types are declared:
def _brighten(pixels: bytes, amount: number) {
iter var i = 0; i < pixels.length(); i++ {
pixels[i] = pixels[i] + amount
}
}
class Canvas {
var pixels = bytes(0)
brighten(amount) {
_brighten(self.pixels, amount)
}
}
The standard library’s imagine module is written this way throughout, and
the difference on a per-pixel loop is measured in multiples, not percent.
Reading the Bytecode
When you want to know what the compiler actually did, ask it:
import zuri
echo zuri.compile('var a = 1 + 2').map(@(i) => i.op)
[LoadConst, AddImm, SetGlobal, LoadNil, Return]
AddImm rather than a separate load tells you the constant was folded into
the instruction. This is the fastest way to check whether a rewrite did
what you hoped, and it never goes stale.
Measuring Honestly
Every number in this chapter is one measurement on one machine. Treat them as ratios rather than as figures to reproduce: yours will differ with your processor, your build and what else the machine is doing. Five rules make the difference between a measurement and a guess.
Warm up before you time. The first few hundred iterations run interpreted, and if your benchmark is short, that is all you measured.
Run each variant on its own. Two variants in one process share warm-up state and a heap. Order them both ways; if the answer changes, you measured the order.
Watch the machine. A laptop under sustained load throttles, and a second run is not comparable to the first. Check the load before the timed run, not before the build that precedes it.
Change one thing. A “fix” that touches three sites has three possible explanations for its effect, and at least one of them is usually a regression hiding behind the other two.
Measure the real program. A microbenchmark of one function tells you
about that function in isolation, with a warm cache and no competing
allocation. ZURI_JIT_COVERAGE=1 on the actual workload tells you which
function to look at in the first place, which is nearly always the more
valuable answer.
Timing Something Yourself
time() returns epoch seconds with microsecond resolution, so a timing
harness is four lines:
def timed(label, work) {
var start = time()
var result = work()
echo '${label}: ${((time() - start) * 1000).round()}ms'
return result
}
var squares = timed('build a list', @() {
var out = []
iter var i = 0; i < 200000; i++ {
out.append(i * i)
}
return out
})
echo squares.length()
The elapsed line will differ every run; the length will not. Print both, so a change in the second tells you the harness broke rather than the code getting faster.
Choosing a Mode
Both compilers run by default. ZURI_JIT_BAYELSA=0 turns Bayelsa off, and
every hot function then stays in Kebbi.
Keep the default for anything that runs long enough for its speed to matter: servers, batch jobs, numeric work, anything whose hot loops run for more than a moment. That is where Bayelsa’s code repays its compile many times over.
Turn Bayelsa off when:
- The program is short. A script that finishes in a fraction of a second spends a real share of its life compiling a second time, and the faster code arrives too late to pay for itself.
- Cores are scarce. Bayelsa’s compiles take more processor time than Kebbi’s. On a machine or container held to one core, or with every core busy with isolates, that time comes out of the program’s own.
- Start-up is the workload. A command-line tool run over and over, a few milliseconds each time, gains nothing from code that is faster on its thousandth iteration.
- Timings have to hold from the first run. Kebbi reaches its speed sooner and stays there. Under Bayelsa a function speeds up once more, part way through a run.
Measure both. The switch is one variable, and running the real workload each way settles the question for that workload.
The Environment Variables
These exist for measurement and debugging. Ordinary programs need none of them.
| Variable | Effect |
|---|---|
ZURI_JIT=0 | disable the JIT entirely |
ZURI_JIT_BAYELSA=0 | compile with Kebbi alone |
ZURI_JIT_LOG=1 | one line per compilation attempt |
ZURI_JIT_LOG_IR=1 | dump the Cranelift IR, and Bayelsa’s own before it |
ZURI_JIT_COVERAGE=1 | a table of what compiled, at exit |
ZURI_JIT_NO_SPECIALIZATION=1 | compile, but do not speculate on types |
ZURI_JIT_THREADS=n | background compiler threads |
ZURI_JIT_CALL_K, ZURI_JIT_OSR_K | the warm-up curve constants |
ZURI_JIT_TIERUP_K | how much profiling work comes before Bayelsa |
ZURI_GC_LOG=1 | garbage collector activity |
ZURI_GC_NURSERY_MB=n | the young generation’s starting and smallest size, in megabytes |
ZURI_OPCODE_PROFILE=1 | interpreter opcode histogram |
MIMALLOC_PURGE_DELAY=n | milliseconds freed memory waits before going back to the system; 10 unless set |
ZURI_JIT=0 is the most useful one. Running a benchmark with and without it
tells you immediately whether your hot path is being compiled at all, which
is the first question to ask when something is slower than it should be.
When to Stop
The interpreter is fast and the JIT is automatic. Most Zuri code needs no performance work at all, and the code that does usually needs exactly one of the three things in this chapter: return a value instead of raising for an expected outcome, annotate a parameter, or stop putting two types in one variable.
Reach for anything more exotic only after ZURI_JIT_COVERAGE has told you
which function is actually the problem.
Testing
A test is a program that runs your program and says whether it did the
right thing. The test module gives you the pieces: a way to name and
group tests, a way to state what you expected, and a report that tells
you which expectation failed and where.
Nothing needs installing. import test and you have it.
Your First Test
import test { * }
def subtotal(items) {
return items.reduce(@(total, item) {
return total + item.price * item.quantity
}, 0)
}
describe('subtotal', @{
it('is zero for an empty cart', @{
expect(subtotal([])).to_be(0)
})
it('multiplies price by quantity', @{
expect(subtotal([{ price: 250, quantity: 3 }])).to_be(750)
})
it('adds every line together', @{
expect(subtotal([
{ price: 250, quantity: 3 },
{ price: 100, quantity: 1 },
])).to_be(850)
})
})
Run it the way you run anything else:
$ zuri run cart.zu
subtotal
✓ is zero for an empty cart
✓ multiplies price by quantity
✓ adds every line together
PASS
3 passed • 3 total
suites 1 time 1ms
Three things are worth noticing straight away.
import test { * }. A test file wants describe, it, expect
and the rest in scope, not behind a module name. import test on its
own works too, and gives you test.describe, test.expect and so on;
use that from code that is not itself a test file.
Nothing says “now run”. describe and it do not run anything;
they build a tree, and the tree is run once the file has finished
declaring it, through os.at_exit(). That separation is what lets the
framework count the tests before the first one starts, focus on one of
them, filter by name, and run them in a random order. run() exists
for when you want to pass options or read the result, and
Controlling a Run comes back to it.
@{ ... } is a function. it('name', @{ ... }) hands the
framework a body to call later, which is why it can decide not to.
When It Fails
Change 850 to 900 and run it again:
subtotal
✓ is zero for an empty cart
✓ multiplies price by quantity
✗ adds every line together
Failures
1) subtotal › adds every line together
Expected 850 to be 900
+ expected 900
- received 850
at cart.zu:23 in @anon3
21 | { price: 250, quantity: 3 },
22 | { price: 100, quantity: 1 },
> 23 | ])).to_be(900)
24 | })
25 |
FAIL
1 failed • 2 passed • 3 total
suites 1 time 4ms
The message, the two values, the line, and the code around it. The
process exits with status 1, which is what a CI system reads.
Grouping
describe nests as deeply as the thing you are describing:
describe('Cart', @{
describe('subtotal()', @{
it('is zero when empty', @{ ... })
})
describe('add()', @{
it('appends a line', @{ ... })
it('merges a duplicate SKU', @{ ... })
})
})
The nesting shows up in the report and in the name a failure is filed
under: Cart > add() > merges a duplicate SKU. Choose names that read
as a sentence when joined like that, because that is how you will read
them.
Expectations
expect(value) gives you an object with every matcher on it.
expect(total).to_be(850)
expect(cart).to_have_length(2)
expect(user).to_match_object({ name: 'Ada' })
expect(items).not.to_be_empty()
.not negates any matcher. There is one implementation behind both
directions, so to_contain and not.to_contain can never disagree
about what containing means.
Matchers chain, because each returns the object it was called on:
expect(port).to_be_int().to_be_between(1024, 65535)
A second argument names the value, which earns its keep the moment the number alone would not say which number it was:
expect(response.status, 'status').to_be(200)
Expected status (404) to be 200
Choosing Between to_be and to_equal
to_be is Zuri’s ==. That already compares lists, dictionaries and
bytes by their contents, so it is the right matcher for most things:
expect([1, 2]).to_be([1, 2]) # passes
expect({ a: 1 }).to_be({ a: 1 }) # passes
It compares two instances by identity, though, unless their class
defines @eq, because that is what == does:
expect(Point(1, 2)).to_be(Point(1, 2)) # fails without @eq: two objects
expect(Point(1, 2)).to_equal(Point(1, 2)) # passes: same contents
to_equal is the structural one. It walks an instance field by field,
treats NaN as equal to itself, and survives a value that contains
itself. Reach for it whenever the objects are built fresh on both sides
of the comparison.
The Matchers
| Group | Matchers |
|---|---|
| Equality | to_be, to_equal, to_be_same_as, to_be_close_to, to_be_within, to_match_object |
| Truthiness | to_be_true, to_be_false, to_be_truthy, to_be_falsy, to_be_nil, to_be_defined |
| Types | to_be_a, to_be_string, to_be_number, to_be_int, to_be_float, to_be_bigint, to_be_bool, to_be_list, to_be_dict, to_be_bytes, to_be_function, to_be_callable, to_be_iterable, to_be_class, to_be_instance, to_be_instance_of |
| Numbers | to_be_greater_than, to_be_greater_than_or_equal, to_be_less_than, to_be_less_than_or_equal, to_be_between, to_be_positive, to_be_negative, to_be_zero, to_be_divisible_by, to_be_even, to_be_odd, to_be_nan, to_be_finite, to_be_infinite |
| Text | to_contain, to_contain_ignoring_case, to_equal_ignoring_case, to_start_with, to_end_with, to_match, to_be_blank |
| Collections | to_have_length, to_be_empty, to_contain_equal, to_contain_all, to_contain_any, to_contain_none, to_contain_exactly, to_have_key, to_have_keys, to_have_value, to_have_property, to_be_sorted, to_have_unique_items |
| Errors | to_raise, to_raise_instance_of, to_raise_with_message, to_not_raise |
| Mocks | to_have_been_called, to_have_been_called_times, to_have_been_called_with, to_have_been_last_called_with, to_have_been_nth_called_with, to_have_returned, to_have_returned_with, to_have_raised |
| Output | to_print, to_print_exactly, to_print_nothing |
| Snapshots | to_match_snapshot |
| Anything else | to_satisfy |
Each one’s exact behaviour, and what it refuses, is in its doc block.
A few are worth calling out.
to_be_close_to is how you compare anything that has been through
floating-point arithmetic. expect(0.1 + 0.2).to_be(0.3) fails; the
sum is 0.30000000000000004.
expect(0.1 + 0.2).to_be_close_to(0.3)
to_raise* takes a function, not a value, because the framework has
to be the one to call it:
expect(@{ parse('') }).to_raise_instance_of(ValueError)
expect(@{ parse('{}') }).to_not_raise()
to_have_property walks a dotted path through dictionaries,
instances and list indices alike, which is what a decoded response
usually is all the way down:
expect(payload).to_have_property('data.items.0.id', 7)
to_satisfy is the escape hatch when nothing else fits:
expect(port).to_satisfy(@(n) { return n % 2 == 0 }, 'to be an even port')
Failing on Purpose
using response.status {
when 200 handle_ok()
when 404 handle_missing()
default fail('unexpected status ${response.status}')
}
And when the assertions live inside a callback that might never run, say how many you expect:
it('reports every error', @{
assertions(2)
validate(bad_input, @(error) {
expect(error.field).to_be_string()
})
})
If the callback ran once instead of twice, the test fails with
Expected 2 assertions but 1 ran rather than passing on a technicality.
has_assertions() is the looser form: at least one.
Setup and Teardown
import test { * }
describe('Session', @{
var store = nil
before_all(@{
echo 'connecting'
})
after_all(@{
echo 'disconnecting'
})
before_each(@{
store = { rows: [] }
})
it('starts empty', @{
expect(store.rows).to_be_empty()
})
it('is a fresh store every time', @{
store.rows.append('a')
expect(store.rows).to_have_length(1)
})
})
$ zuri run session.zu
connecting
Session
✓ starts empty
✓ is a fresh store every time
disconnecting
PASS
2 passed • 2 total
suites 1 time 2ms
The rules:
before_allruns once, immediately before the first test in its suite that is actually going to run. A suite everything was filtered out of never connects to anything.after_allruns after the last one, and only ifbefore_allran. It runs whenbefore_allraised as well, since a setup that failed part way may already hold a connection or a server, so write it to cope with whatever the setup got as far as.before_eachruns outermost suite first,after_eachinnermost first, so teardown undoes setup in the order it was done.after_eachruns even when the test failed, which is exactly when you need it to.
A hook that raises is reported as a hook failure against the test it was
preparing, rather than taking the run down. A before_all that raises
fails every test in its suite, because none of them got the setup they
were written against.
Focusing, Skipping and Todo
While you are working on one thing:
it_only('the one I am fixing', @{ ... })
describe_only('the area I am in', @{ ... })
As soon as anything is marked only, everything else is skipped and
still listed, so you cannot forget it is on.
it_skip('known broken, see #412', @{ ... })
describe_skip('the old API', @{ ... })
it_todo('handle an empty payload')
A todo is a test with no body. it('name') with nothing after it means
the same thing. It counts towards nothing and shows up in the report as
a reminder that survives being committed.
Two options change what a failure means:
it('reproduces issue 412', @{ ... }, { failing: true })
it('talks to a flaky endpoint', @{ ... }, { retries: 2 })
failing: true passes when the body fails, and fails when the body
passes, so the day someone fixes the bug the test tells you. retries
runs the body again on failure, and a test that passes on a later
attempt is reported as flaky rather than quietly green.
One Test, Many Inputs
import test { * }
def slug(title) {
return title.lower().replace('/[^a-z0-9]+/', '-').trim('-')
}
it_each([
['Hello World', 'hello-world'],
[' Spaced Out ', 'spaced-out'],
['Zuri 1.0!', 'zuri-1-0'],
], 'turns $0 into $1', @(title, expected) {
expect(slug(title)).to_be(expected)
})
$ zuri run slug.zu
✓ turns 'Hello World' into 'hello-world'
✓ turns ' Spaced Out ' into 'spaced-out'
✓ turns 'Zuri 1.0!' into 'zuri-1-0'
PASS
3 passed • 3 total
suites 0 time 3ms
Each row becomes a separate test with its own name and its own place in
the report, so one bad row does not hide the others. $0, $1 and so
on stand for the row’s values and $# for the row number.
describe_each does the same for whole suites, which is how you run one
set of tests against several implementations of the same interface.
Test Doubles
mock() gives you a function that records how it was called and does
whatever you tell it to.
import test { * }
def retry(operation, attempts) {
iter var attempt = 1; attempt <= attempts; attempt++ {
catch {
return operation()
} as error {
if attempt == attempts {
raise error
}
}
}
}
describe('retry', @{
it('returns the first success', @{
var operation = mock()
operation.returns('ok')
expect(retry(operation.fn, 3)).to_be('ok')
expect(operation).to_have_been_called_times(1)
})
it('tries again after a failure', @{
var operation = mock()
operation.raises_once(Error('connection reset'))
operation.returns('ok')
expect(retry(operation.fn, 3)).to_be('ok')
expect(operation).to_have_been_called_times(2)
})
it('gives up eventually', @{
var operation = mock()
operation.raises(Error('connection reset'))
expect(@{ retry(operation.fn, 3) }).to_raise_with_message('connection reset')
expect(operation).to_have_been_called_times(3)
})
})
mock() returns a Mock, and mock.fn is the plain function you hand
to the code under test. The matchers accept either, so
expect(operation) and expect(operation.fn) mean the same thing.
The behaviour is scripted in front-to-back order: everything queued with
a _once suffix runs first, one call each, and then the standing
behaviour takes over.
| Method | Effect |
|---|---|
returns(v) / returns_once(v) | return v |
raises(e) / raises_once(e) | raise e |
implements(fn) / implements_once(fn) | run fn with the real arguments |
reset() | forget the calls |
clear_behaviour() | forget the behaviour |
Spying on Something That Already Exists
spy_on swaps a function out in place and gives you a Mock that both
records and stands in for it:
import test { * }
def checkout(cart, gateway) {
var total = cart.reduce(@(sum, line) { return sum + line }, 0)
return gateway['charge'](total)
}
it('charges the cart total once', @{
var gateway = { charge: @(amount) { return 'live-charge' } }
var charge = spy_on(gateway, 'charge')
charge.returns('receipt-1')
expect(checkout([250, 100], gateway)).to_be('receipt-1')
expect(charge).to_have_been_called_times(1)
expect(charge).to_have_been_called_with(350)
})
A spy calls through to the original unless you tell it otherwise, and is restored automatically after the test that installed it, even one that failed halfway through. Nothing has to be undone by hand.
It works on a dictionary entry and on an instance property that holds a function. It does not work on a class method or a module function: Zuri classes are immutable once declared, and a module’s members cannot be assigned from outside it. Code you want to substitute takes its collaborators as arguments or holds them in properties, which is the shape worth designing for anyway.
Snapshots
For a value too big to write out by hand, record it once and compare against the recording from then on:
def invoice(customer, lines) {
return {
customer,
lines,
total: lines.reduce(@(sum, line) { return sum + line.amount }, 0),
}
}
it('builds the document', @{
expect(invoice('Ada', [{ label: 'Design', amount: 4200 }])).to_match_snapshot()
})
The first run writes the file and passes:
$ zuri run invoice.zu
invoice
✓ builds the document
PASS
1 passed • 1 total
suites 1 time 2ms
snapshots 1 written
It lands beside the test file, in __snapshots__:
# Zuri snapshot file v1
=== invoice > builds the document 1 ===
{
customer: 'Ada',
lines: [
{
amount: 4200,
label: 'Design'
}
],
total: 4200
}
Commit it. Reviewing the change to that file in a pull request is the entire value of the technique: an unexplained diff there is exactly the thing worth noticing.
When the value changes, you get the diff:
1) invoice › builds the document
Snapshot 'invoice > builds the document 1' no longer matches
+ expected - received
{
customer: 'Ada',
lines: [
{
- amount: 5200,
+ amount: 4200,
label: 'Design'
}
],
- total: 5200
+ total: 4200
}
If the new value is right, rewrite the snapshots:
$ ZURI_UPDATE_SNAPSHOTS=1 zuri run invoice.zu
That also deletes entries nothing asks for any more. And on CI, where
CI is set in the environment, writing a brand new snapshot is a
failure rather than a silent pass, because a snapshot nobody has looked
at asserts nothing.
Testing What Something Prints
expect(@{ greet('Ada') }).to_print('Hello, Ada')
expect(@{ quiet_mode() }).to_print_nothing()
Or take the output and assert on it yourself:
var printed = capture_output(@{ report(rows) })
expect(printed.lines()).to_have_length(4)
This works because the runtime can redirect everything Zuri writes to
standard output. The same mechanism is why a passing test’s output does
not clutter the report: it is captured and shown only when the test
fails. run({ verbose: true }) shows it either way, and
run({ capture: false }) turns it off entirely.
More Than One File
A real project has a directory of them:
project/
tests/
cart.zu
pricing.zu
zuri test runs the lot:
$ zuri test
zuri test 2 files in tests
PASS cart.zu 188ms 2 tests
FAIL pricing.zu 281ms 1 test
✗ pricing > applies the discount
Expected 90 to be 100
at tests/pricing.zu:4
2 files • 1 with failures
1 failed • 2 passed • 3 total
time 476ms
There is nothing to write for this. The command finds the tests
directory you ran it from, and ends the process with 1 when anything
failed, so a CI job needs nothing added to it either.
Each file runs in a process of its own. That is not an implementation detail you can ignore, because it is what you are buying:
- A file that loops forever is killed, and the rest still run.
--timeout 30ssays how long is too long. - A file that crashes, or calls
os.exit()halfway through, is reported as a file that never reported rather than taking the run with it. - Global state, a module loaded for its side effect, a changed working directory: none of it leaks from one file into the next.
Files are reported in the order they were discovered whatever order they
finish in, so --jobs 4 makes a big suite faster without making the
report move around.
A test file needs nothing special to be conducted. It declares its tests, and is equally runnable on its own.
Running One File
Name it, with or without its .zu:
$ zuri test pricing
$ zuri test pricing.zu
$ zuri test tests/pricing.zu
All three run tests/pricing.zu. A name is looked for under tests
first, so a test keeps its own name even when something else in the
project shares it, and a name that is nowhere to be found under that
directory is looked for once more by filename alone anywhere beneath it,
so a file in a subdirectory answers to its own name.
A directory works too, wherever it sits:
$ zuri test tests/api
$ zuri test packages/store/tests
The Flags
| Flag | What it does |
|---|---|
-j, --jobs <count> | How many files to run at once. auto is one per CPU. Default 1. |
-t, --timeout <duration> | How long to give a single file before killing it. 500ms, 30s, 2m, 1h, or a bare number of milliseconds. Default: no limit. |
-b, --bail [count] | Stop after this many failing files. On its own, stop at the first. |
-m, --match <pattern...> | Filename patterns to run, instead of *.zu. |
-i, --ignore <pattern...> | Filename patterns to skip, instead of _*, .* and index.zu. |
--no-recursive | Only the files directly in the directory. |
-e, --env <assignment...> | Extra environment for every test process, as KEY=VALUE. |
-l, --list | Print the files that would run, one to a line, and stop. |
$ zuri test --jobs auto --timeout 30s
$ zuri test --bail
$ zuri test --match '*_test.zu' '*_spec.zu'
--bail and the two pattern flags take as many words as follow them, so
put the file you are naming ahead of them, or close them with --:
$ zuri test pricing --bail
$ zuri test --match '*_test.zu' -- pricing
A count for --bail has to be attached, --bail=3, for the same
reason: a count written as a separate word is indistinguishable from the
file.
Conducting a Suite Yourself
conduct() is the function zuri test is built on, and a project that
wants the run under its own control can call it directly. A directory
handed to zuri run runs its index.zu, so one file makes the suite
zuri run tests:
Filename: tests/index.zu
import os
import test
test.conduct(os.dir_name(__file__), { jobs: 4, timeout: 30000 })
It takes the same choices the flags do, as a dictionary, and returns the
run rather than only reporting it. conduct leaves index.zu out of
discovery, and never runs the script that called it either, so the index
cannot end up running itself.
Controlling a Run
Call run() yourself when you want to change how the run behaves, or
to read the result. It takes an options dictionary, and calling it
takes over from the automatic run, so nothing happens twice.
The options you will reach for:
| Option | What it does |
|---|---|
filter | run only tests whose full name contains this, or matches it as a regular expression |
bail | stop after this many failures |
shuffle / seed | run in a random order, and reproduce that order later |
reporter | spec, dot, tap, junit, json, ndjson, silent |
exit | exit the process with the run’s status when it finishes |
run({ filter: 'subtotal', bail: 1 })
filter and reporter can also come from ZURI_TEST_FILTER and
ZURI_TEST_REPORTER, which is what lets one command be pointed at one
test without editing the file.
Shuffling is the one worth turning on deliberately. Tests that only pass because an earlier test left something behind are a real and common problem, and running them in a different order every time is how you find out:
run({ shuffle: true })
The seed is printed with the summary, and passing it back reproduces that exact order.
On CI
Nothing to add. A test file exits 0 when everything passed and 1
when it did not, whether or not it called run().
For a CI system that wants a machine-readable report,
{ reporter: 'junit' } writes JUnit XML to standard output and
{ reporter: 'tap' } writes TAP version 14.
Writing Your Own Reporter
The runner knows nothing about output. It walks the tree and calls methods on a reporter, and every method has a do-nothing default:
import test { * }
import test.reporter { Reporter }
class Quiet < Reporter {
test_finished(one) {
if one.status == 'failed' {
echo one.full_name()
}
}
}
describe('a suite', @{
it('passes quietly', @{ expect(1).to_be(1) })
it('fails loudly', @{ expect(1).to_be(2) })
})
run({ reporter: Quiet(), exit: false })
The objects handed to it are the same ones the built-in reporters see:
Case, Suite, Failure and Summary, documented in
test.result.
Two Things to Know
A timeout on a test is measured after the body returns. Zuri runs
synchronously, so a test runs to completion and is then timed. That
catches one that got too slow; it does not rescue one that hangs. A hang
needs a process boundary, which is exactly what conduct puts around
each file, and why its timeout is the enforcing one.
Everything in one file shares one interpreter. Reset what you change
in before_each, and reach for shuffle to find out whether you missed
anything.
Where to Go Next
The module’s own documentation carries the full matcher list with each
one’s edge cases, every option run() and conduct() accept, and the
snapshot file format. Appendix F lists
the submodules. Chapter 24 is what to do once a
test has told you something is wrong.
Debugging
Something is not doing what you expected. This chapter is about closing that gap: reading what the runtime tells you, recognising the handful of messages that account for most confusion, and the techniques that turn a vague “it’s broken” into a line number.
Reading an Error
An uncaught error prints three things: what went wrong, where, and how you got there.
Unhandled ValueError: bottomed out
--> /path/to/main.zu:3
1 | def recurse(n) {
2 | if n <= 0 {
> 3 | raise ValueError('bottomed out')
4 | }
5 | recurse(n - 1)
Stack trace (most recent call last):
at recurse() /path/to/main.zu:3
at recurse() /path/to/main.zu:5
... 19 more frames ...
at recurse() /path/to/main.zu:5
The type and message come first. Then the source around the failure, with the offending line marked. Then the call stack, innermost first.
A deep stack is truncated in the middle, because the top and the bottom are the parts that tell you anything: the top is where it broke, the bottom is where you started, and two hundred identical recursive frames in between are noise.
The process exits with status 1.
The Messages You Will Actually See
Every message below is real output. Run this and you get all of them:
class Config {
@new() {
self.name = 'default'
}
}
def needs_string(s: string) {
return s
}
def show(label, work) {
catch {
work()
} as e {
echo '${label}: ${e.type} — ${e.message}'
}
}
show('name never declared', @() => undeclared_name)
show('key not in dict', @() { var d = { a: 1 }; return d['b'] })
show('field not on class', @() { var c = Config(); c.nope = 1 })
show('property not on class', @() { var c = Config(); return c.missing })
show('method on nil', @() { var x = nil; return x.f() })
show('operator on nil', @() => nil + 1)
show('index past the end', @() => [1][9])
show('wrong argument type', @() => needs_string(5))
name never declared: UndefinedError — undefined global 'undeclared_name'
key not in dict: PropertyError — undefined key 'b' in dict
field not on class: PropertyError — undefined field 'nope' on instance of 'Config'
property not on class: PropertyError — undefined property 'missing' on instance of 'Config'
method on nil: TypeError — object of type nil does not define method 'f'
operator on nil: TypeError — operator '+' not defined for call signature (nil, number)
index past the end: RangeError — index 9 out of bounds (length 1)
wrong argument type: TypeError — needs_string() expects parameter 's' (argument 1) to be a string, got number
What each one usually means in practice:
| Message | What to look for |
|---|---|
undefined global 'x' | a typo, or a def that appears below the top-level line calling it |
undefined key 'x' in dict | a key that is genuinely absent; get(key, fallback) is the fix when absence is legal |
undefined field 'x' on instance of 'C' | a typo in a field name — classes are sealed, so this cannot create one |
undefined property 'x' on instance of 'C' | reading a field or method the class never declared |
undefined member 'x' on module m | the module did not export it, or the export needs an @ |
object of type nil does not define method 'f' | something upstream returned nil |
operator '+' not defined for call signature (nil, number) | the same, one step earlier |
'x' is private and can only be accessed via 'self' or 'parent' | a leading underscore, reached from outside |
'x' is already declared in this scope | two vars of one name in one block |
cannot assign to constant 'x' | writing to a local const |
module 'x' could not be found | usually a missing leading . on a sibling import |
The two nil messages are worth internalising. Zuri never tells you
where a nil came from, because by the time it causes trouble the value
has already been passed along. When you see one, stop looking at the line
that failed and look at whatever produced the value.
Techniques
Echo the Value and Its Type
Dynamic typing means the surprise is almost always that something is not what you assumed it was:
var value = '42'
echo typeof(value)
echo value
echo value + 1
echo value.to_number() + 1
string
42
421
43
One line of typeof() would have saved the third line’s confusion.
Give a Class @to_string()
echo shows an instance through its class’s @to_string(), and as
<instance of Point> when the class has none:
class Point {
@new(x, y) {
self.x = x
self.y = y
}
@to_string() {
return '(${self.x}, ${self.y})'
}
}
var p = Point(1, 2)
echo p
echo [p, Point(3, 4)]
(1, 2)
[(1, 2), (3, 4)]
Define @to_string() on any class you expect to look at while debugging.
The five minutes it costs are repaid the first time you print a list of
them.
Encode Nested Data Instead of Echoing It
A deep dictionary printed by echo is one unreadable line. json.encode()
with compact off indents it:
import json
var request = {
method: 'POST',
headers: { accept: 'application/json' },
body: { title: 'write chapter 19', tags: ['docs', 'zuri'] },
}
echo json.encode(request, false)
{
"method": "POST",
"headers": {
"accept": "application/json"
},
"body": {
"title": "write chapter 19",
"tags": [
"docs",
"zuri"
]
}
}
Add Context on the Way Up
An error raised deep in a call chain says what failed, not what you were doing at the time. Catch it, say what you were doing, and re-raise:
def parse_port(raw) {
if !raw.match('/^\d+$/') {
raise ValueError('not a number')
}
return raw.to_number()
}
def load_settings(source, raw) {
catch {
return { port: parse_port(raw) }
} as e {
raise ValueError('while reading ${source}: ${e.message}')
}
}
catch {
load_settings('config.json', 'eighty')
} as e {
echo e.message
}
while reading config.json: not a number
“not a number” is a fact. “while reading config.json: not a number” is a fact you can act on.
Print the Stack Trace Yourself
The trace is a list on the error, so a handler can log it without letting the program die:
def inner() {
raise ValueError('deep')
}
def outer() {
inner()
}
catch {
outer()
} as e {
for frame in e.stacktrace {
echo frame
}
}
/path/to/main.zu:2 -> inner()
/path/to/main.zu:6 -> outer()
/path/to/main.zu:10 -> @.script()
Ask the Compiler What It Made of Your Code
When an expression does not behave the way you read it, the bytecode settles the argument:
import zuri
echo zuri.compile('var a = -2 ** 2').map(@(i) => i.op)
[LoadConst, LoadConst, Pow, Neg, SetGlobal, LoadNil, Return]
Read the order: the power happens before the negation. ** binds
tighter than unary minus, so -2 ** 2 is -(2 ** 2), which is -4, and
not the 4 a reader expecting (-2) ** 2 would get. The bytecode settles
it in one line. Chapter 21
covers zuri.compile() and zuri.parse() properly.
Narrow It With assert
An assert is a claim you can leave in the code:
def average(numbers) {
assert !numbers.is_empty(), 'average() needs at least one number'
return numbers.reduce(@(a, b) => a + b, 0) / numbers.length()
}
echo average([2, 4, 6])
catch {
average([])
} as e {
echo '${e.type}: ${e.message}'
}
4
AssertError: average() needs at least one number
Without it, average([]) would have returned NaN and the problem would
have surfaced somewhere else entirely, in a value that looks like a
number.
Watch the Collector
ZURI_GC_LOG=1 reports garbage collector activity, which is the one to
reach for when memory rather than logic is the question:
$ ZURI_GC_LOG=1 zuri run main.zu
Chapter 22 lists the rest of the runtime’s diagnostic switches alongside what each one measures.
The Traps Worth Knowing by Heart
These are the behaviours that produce a wrong answer rather than an error, which makes them far more expensive to find.
Zero is falsy. var n = count or 10 turns a real 0 into 10, and
if index is false for the first position and true for -1, “not
found”. Compare explicitly.
to_number() returns 0 for text it cannot parse. 'eighty', ''
and '12abc' all become 0, and so does ' 7 ' with its spaces. There is
no NaN and no error to catch, so a bad input silently becomes a valid
zero. Validate the text before converting it.
[] and {} are truthy. if items is true for an empty list. Use
is_empty().
sort() mutates; reverse() does not. var s = items.sort() leaves
items sorted as well. Clone first when you need both orders.
A method call on a string result was discarded. name.trim() does
nothing on its own; strings are immutable, so you must assign the result.
A nested def is local. A function declared inside another is scoped
to it, like a var, and nothing outside that scope can call it.
x++ evaluates to the new value. Unlike C and JavaScript.
Structuring Code You Can Reason About
Three habits pay for themselves the first time something breaks.
Separate the decision from the effect. A function that reads a file, parses it and decides something is three functions. Split them, and the parsing and the decision can both be exercised without a filesystem in the way.
Pass dependencies in. A function that calls time() behaves
differently every second. One that takes a timestamp behaves identically
every time you call it with the same number, which means you can reproduce
a failure instead of waiting for it.
Return values instead of printing them. echo inside a function is
invisible to its caller and useless to anything that wants to check the
result. Return the string and let the caller decide.
A Full-Stack Task Board
We are going to build a real web application: a shared task board with three columns, a browser interface, and a JSON API over the same data.
It is about four hundred lines of Zuri, and it uses almost everything this
book has covered. Classes and inheritance for the domain model. Custom
errors for validation. The module system for structure. Files and JSON for
persistence. The HTTP server for routing, middleware and static files. Wire
for server-rendered HTML. os for configuration and signals. log for
output.
No dependencies. Nothing to install. zuri run taskboard and it runs.
What It Does
- Three columns:
todo,doing,done. - A page at
/showing the board, with forms to add, move and delete. - A JSON API under
/apidoing the same things, for anything that is not a browser. - A
board.jsonfile holding the state, written atomically. - One error handler turning domain errors into the right status code, for both the HTML and the JSON side.
How the Chapter Is Organised
Each section builds one layer, from the inside out:
- Laying Out the Project: the directory structure and why it is shaped this way.
- The Storage Layer: the JSON file and the board that sits on top of it.
- Validation and the Domain Model: the
Taskclass, which owns every rule about what a task is. - The JSON API: six routes over the board.
- Server-Rendered Pages with Wire: the templates, the forms and the redirect-after-post pattern.
- Middleware, Logging and Errors: the two pieces of cross-cutting behaviour every request goes through.
- Running It for Real: configuration, signals and what changes when you want more than one core.
- Testing the Board: a suite over all three layers, and what the layering bought.
Read it in order. Each section assumes the previous one exists.
Laying Out the Project
taskboard/
index.zu entry point
app.zu wiring: board + templates + server
config.zu every setting, with defaults
models/
index.zu
task.zu the Task class and its rules
storage/
index.zu the Board
_json_store.zu private: reading and writing the file
routes/
index.zu
api.zu the JSON routes
pages.zu the HTML routes
templates/
layout.html
board.html
static/
app.css
tests/
task.zu one file per layer
board.zu
api.zu
Four decisions are worth explaining, because they are the ones that make the rest of the code short.
Before them, the shape itself. Every arrow in this application points one way:
index.zu -> app.zu -> routes/ -> storage/ -> models/
routes knows about storage; storage knows about models; models
knows about nothing but config. Nothing points back the other way, which
is what makes each layer readable on its own and testable without the ones
above it. Testing the Board is where that second
half is cashed in.
config.zu sits outside that chain — everything may read it, and it reads
nothing. That is the one module allowed to be depended on from anywhere,
and it earns the exemption by containing no behaviour at all.
One Package per Responsibility
models, storage and routes are packages: directories with an
index.zu. The index.zu decides what is public:
Filename: models/index.zu
import @.task { * }
That one line re-exports everything task.zu declares, so the rest of the
application writes import .models { Task } and never needs to know there
is a task.zu at all. Splitting task.zu into two files later changes
that one line and nothing else.
The @ is what makes it a re-export rather than a private import. Without
it, models/index.zu could use Task itself but nothing outside could
reach it through models — the import would be local, and
import .models { Task } elsewhere would fail. That distinction is the
subject of Exporting, and this is the
single most common place to get it wrong.
The index.zu is therefore a deliberate, editable list of what a package
offers, rather than an accident of which files happen to exist.
The Underscore Is Load-Bearing
storage/_json_store.zu starts with an underscore, so no module outside
storage can import it. The compiler enforces that:
SyntaxError: Cannot import private items from module
The point is not secrecy. It is that swapping the JSON file for a database
means rewriting one file, and the compiler guarantees nothing else
reached past the Board to touch it. Not “we checked and nothing does” —
nothing can, and an attempt does not compile.
That guarantee is what turns a convention into a boundary. A comment saying “internal, do not use” is advice; a leading underscore is enforced.
app.zu and index.zu Are Separate
app.zu builds a fully configured server and returns it. index.zu
starts one:
Filename: index.zu
import log
import os
import @.app
import .config
if __root__ == __file__ {
var server = app.build()
os.on_signal('INT', @() {
log.info('shutting down')
os.exit(0)
})
log.info('task board on http://${config.HOST}:${config.PORT}')
server.listen()
}
Three things are happening in that short file.
if __root__ == __file__ is the whole trick. zuri run taskboard runs the
directory, finds index.zu, and the two are equal, so the body runs and
the server starts. import .taskboard from another program leaves them
different — __root__ is that program’s entry file — so nothing starts,
and the importer gets app.build() to use however it likes. One file, both
jobs, and Chapter 8 covers the idiom.
The signal handler is installed before serving. listen() does not
return, so anything that needs to happen on the way out has to be arranged
first. os.on_signal('INT', ...) is what turns Ctrl+C into an orderly exit
rather than a killed process.
The log line comes before listen(), so the address is printed the
moment the process is ready rather than after it stops. It is one line, and
it is the difference between “did it start?” and knowing.
Note what index.zu does not do. It builds nothing, configures nothing
and knows nothing about boards, templates or routes. Every one of those
decisions is in app.zu, which is why the whole of
Middleware, Logging and Errors can walk through
one function and cover the entire assembly.
Configuration Has Defaults
Filename: config.zu
import os
var HERE = os.dir_name(__file__)
var PORT = os.get_env('PORT', '8000').to_number()
var HOST = os.get_env('HOST', '127.0.0.1')
var DATA_DIR = os.get_env('DATA_DIR', os.join_paths(HERE, 'data'))
var TEMPLATE_DIR = os.join_paths(HERE, 'templates')
var STATIC_DIR = os.join_paths(HERE, 'static')
var COLUMNS = ['todo', 'doing', 'done']
Four things to notice.
Every environment lookup has a fallback, so the application runs with
no configuration at all. git clone, zuri run taskboard, and it works. A
program that requires six environment variables before it will start is a
program nobody tries.
PORT is converted with to_number(). Environment variables are
always strings, and '8000' is not a port a socket will accept. The
conversion is here, once, rather than at the place the port is used.
HERE is derived from __file__, not from os.cwd(). The templates
and the stylesheet live next to the source, so the person running the
program is not required to be standing in the right directory. This is the
rule from Chapter 9, and a web application is where
ignoring it hurts most: the program starts, serves a page, and fails to
find a template only once someone requests it.
COLUMNS is here because it is the one piece of knowledge the model,
the storage layer and the templates all share. _clean_column() validates
against it, Board.columns() and Board.summary() iterate it, and the
template renders one section per entry. Adding a fourth column is a
one-line change in this file, and every one of those follows.
That is the test for whether something belongs in config.zu: not “is it a
setting”, but “would changing it otherwise mean editing several files
consistently?”
The Storage Layer
Two files. One knows about the disk and nothing about tasks. The other knows about tasks and nothing about the disk.
The File
Filename: storage/_json_store.zu
import json
import os
/**
* Reads every stored task dictionary.
*
* A missing file is an empty board, not an error: a fresh install has
* nothing saved yet and should start cleanly. A file that exists but
* cannot be parsed IS an error, because silently discarding someone's
* data is worse than refusing to start.
*
* @param string path
* @returns list[dict]
* @throws Error if the file exists and does not contain a JSON list.
*/
def read_all(path: string) {
var handle = file(path)
if !handle.exists() {
return []
}
var decoded
catch {
decoded = json.decode(handle.read())
} as error {
raise Error('${path} is not readable as JSON: ' + error.message)
}
if !is_list(decoded) {
raise Error('${path} should contain a JSON list, found ' + typeof(decoded))
}
return decoded
}
/**
* Replaces the stored board with `records`.
*
* Writes to a temporary file beside the target and renames it into
* place, so a write interrupted halfway leaves the previous board
* intact rather than a truncated file.
*
* @param string path
* @param list records
*/
def write_all(path: string, records: list) {
var directory = os.dir_name(path)
if !os.dir_exists(directory) {
os.create_dir(directory, nil, true)
}
var temporary = path + '.tmp'
var handle = file(temporary, 'w')
handle.open()
handle.write(json.encode(records, false))
handle.close()
file(temporary).rename(path)
}
Four decisions in fifty lines.
A missing file is not an error. A first run has nothing saved, and the program should start. A file that exists but cannot be parsed is an error, because the alternative is overwriting whatever was in there with an empty board.
The error message names the path. The person reading it is looking at one line of a log.
The write is atomic. Writing into a temporary file and renaming it into place means a process killed mid-write leaves the previous board intact, because a rename within a directory either happens or does not.
The handle is opened explicitly. write() on a closed handle opens,
writes and closes again, so two write() calls in w mode would each
truncate the file. open() first, then write, then close(). That is the
rule from Chapter 9, and this is exactly where it
bites.
The Board
storage/index.zu is the layer above. It holds Task objects, answers
questions about them, and persists through _json_store — and it is the
only thing in the application that knows a file is involved at all.
Rather than read it as one hundred lines, here it is a piece at a time.
The Error It Raises
Filename: storage/index.zu
import os
import ..config
import ..models { Task, from_dict }
import ._json_store
/**
* Raised when a task id does not name a task on the board.
*/
class NotFoundError < Error {
@new(message) {
parent(message)
self.type = 'NotFoundError'
}
}
A custom error class, four lines, and it earns them. Every layer above can
say instance_of(error, NotFoundError) instead of matching on a message
string, and the middleware in
Middleware, Logging and Errors turns exactly this
class into a 404. Had get() raised a plain Error('not found'), that
mapping would be a substring search.
Note the two lines inside @new. parent(message) lets the base
constructor set message and capture the stack trace; self.type is what
makes the class name show up in logs and in the uncaught-error banner.
Both are the pattern from Chapter 7.
The Fields and the Constructor
/**
* A board backed by a JSON file.
*/
class Board {
/** Every task, newest last. */
var tasks = []
/** The file this board reads from and writes to. */
var path
/**
* @param ?string directory: defaults to `config.DATA_DIR`
*/
@new(directory) {
self.path = os.join_paths(directory or config.DATA_DIR, 'board.json')
self.tasks = _json_store.read_all(self.path).map(@(record) => from_dict(record))
}
Two fields, declared with var even though @new assigns both. The
constructor could have declared them implicitly; writing them out means the
class’s shape is visible at the top without reading the constructor, and it
is required the moment any other method assigns to them.
The constructor does two things and no more. It works out where the file is, and it loads what is in it.
directory or config.DATA_DIR is the optional-parameter idiom from
Chapter 5. It is safe here precisely
because a directory is never legitimately '' or 0 — the falsy-default
trap does not apply to paths.
The map on the second line is the boundary between two worlds.
read_all() returns plain dictionaries, because that is what JSON is;
from_dict() turns each into a real Task. Everything above this line
deals in objects, everything below it deals in dictionaries, and this is
the single place they meet.
Reading the Board
all() {
return self.tasks
}
in_column(column: string) {
return self.tasks.filter(@(task) => task.column == column)
}
all() is a one-liner, and it hands back the real list rather than a copy
— deliberate, because every caller in this application only reads it. If
that changed, this is the line that would need clone().
in_column() is filter() and nothing else. It is a method rather than a
loop at each call site so that “which column is this task in” is decided in
one place.
Shaping It for the Page
/**
* The board grouped for display: one entry per configured column,
* in board order, each with its own tasks.
*
* @returns list[dict]: `{ name, tasks }`
*/
columns() {
return config.COLUMNS.map(@(name) {
var tasks = self.in_column(name).map(@(task) => task.to_view())
return { name, tasks }
})
}
This is the method the HTML template consumes, and three decisions inside it are worth naming.
It iterates config.COLUMNS, not the tasks. That means a column with
no tasks still appears, as an empty column — which is what a board should
look like. Grouping by walking the tasks instead would silently drop empty
columns and produce them in whatever order the data happened to be in.
It returns to_view() results, not Task objects. A view carries the
formatted date and the is_done flag already computed, so the template
never has to. Server-Rendered Pages explains why that
matters for templates specifically.
{ name, tasks } uses the shorthand. Both keys match the variables
holding them, so writing { name: name, tasks: tasks } would be noise.
Finding One
/**
* @throws NotFoundError if no task has that id.
*/
get(id: string) {
var found = self.tasks.find(@(task) => task.id == id)
if found == nil {
raise NotFoundError('no task with id ${id}')
}
return found
}
The important decision is that get() raises rather than returning
nil.
That choice propagates. Every caller either has a real task or has an
error, so no route handler contains if task == nil. The alternative —
returning nil — would put that check in five places and guarantee one of
them was eventually forgotten, producing a nil field access somewhere far
away from the cause.
The message includes the id, because the person reading the log has the request and nothing else.
Changing It
add(title, options) {
var task = Task(title, options)
self.tasks.append(task)
self._save()
return task
}
update(id: string, changes: dict) {
var task = self.get(id).update(changes)
self._save()
return task
}
remove(id: string) {
var task = self.get(id)
self.tasks.remove_at(self.tasks.index_of(task))
self._save()
return task
}
All three follow one shape: change the in-memory list, then save, then return what changed.
update() and remove() both start by calling get(), which means a bad
id raises NotFoundError before anything is modified. There is no path
that half-applies a change.
Each returns the affected task rather than nothing, which is what lets the JSON API answer with the created or updated record without a second lookup.
remove_at(index_of(task)) rather than remove(task) is deliberate: list
remove() compares by value, and two tasks with identical fields would
make that ambiguous. Removing by position removes exactly the object
get() found.
Counting, and Saving
summary() {
var counts = {}
for name in config.COLUMNS {
counts.set(name, self.in_column(name).length())
}
return counts
}
_save() {
_json_store.write_all(self.path, self.tasks.map(@(task) => task.to_dict()))
}
}
summary() walks the configured columns for the same reason columns()
does: an empty column should report 0, not be missing.
_save() is the other side of the constructor’s map. Objects go out as
dictionaries via to_dict(), exactly as they came in through
from_dict(). The two are a matched pair, and a field added to Task
needs to appear in both or it will not survive a restart.
_save() is private, and that is the class’s most important property.
There is no public method that changes the board without persisting,
because add, update and remove each call it and nothing else can. A
caller cannot forget to save, because a caller is never given the choice.
The Trade It Makes
The whole board is held in memory and rewritten in full after every change.
For a board a team can read on one screen, that is the right trade: the code is simple enough to hold in your head, and a full rewrite of a few kilobytes is immaterial. It is also the first assumption to revisit if this ever has to hold a hundred thousand tasks, at which point the answer is a real database rather than a cleverer file format.
Concurrency
Board holds the whole file in memory. Two processes writing the same
board.json would each overwrite the other’s changes, and the atomic
rename does not fix that: it guarantees the file is never half-written, not
that two writers agree.
This application runs one server process, so the question does not arise. Running It for Real is where it does.
Validation and the Domain Model
Every rule about what a task is lives in one class. Nothing above it re-checks a title, and nothing below it stores a task that broke a rule.
Filename: models/task.zu
import date
import uuid
import ..config
/**
* Raised when a task cannot be built from the data given.
*/
class TaskError < Error {
@new(message, field) {
parent(message)
self.type = 'TaskError'
self.field = field
}
}
TaskError carries a field as well as a message. That one extra value is
what lets the API answer {"error": "...", "field": "title"}, which is
what lets a form highlight the input that is wrong. A custom error class
exists precisely so it can carry more than a string.
The Class
/**
* A single task on the board.
*/
class Task {
/** The task's stable identifier, a UUID v7 so ids sort by age. */
var id
/** One line describing the work. Required, at most 120 characters. */
var title
/** Free-form detail. Optional, defaults to an empty string. */
var notes = ''
/** Which column the task sits in. One of `config.COLUMNS`. */
var column = 'todo'
/** Unix timestamp, in seconds, of when the task was created. */
var created_at
/**
* @param string title
* @param ?dict options: `notes`, `column`, `id`, `created_at`
* @throws TaskError if the title is empty, too long, or the column
* is not one this board has.
*/
@new(title, options) {
options = options or {}
self.id = options.get('id', nil) or uuid.v7()
self.title = _clean_title(title)
self.notes = options.get('notes', nil) or ''
self.column = _clean_column(options.get('column', nil) or 'todo')
self.created_at = options.get('created_at', nil) or time()
}
Every field is declared with var and documented, even the ones the
constructor fills in. Classes are sealed, so the list is the whole truth
about what a task holds, and writing it out means the next reader gets that
truth without reading the constructor.
The constructor takes a required title and an options dictionary for
everything else. That shape is worth copying: a positional argument for the
thing that is always there, and named options for the rest, so a call site
never reads Task('x', nil, nil, 'todo', nil).
uuid.v7() rather than v4(), because v7 embeds a timestamp, so ids sort
by age. When your identifier is going to end up as a key in something
ordered, that is free value.
Mutation
/**
* Moves the task to another column.
*
* @param string column
* @returns Task: this task, so calls chain.
* @throws TaskError if the column is not one this board has.
*/
move_to(column) {
self.column = _clean_column(column)
return self
}
/**
* Applies a partial update. Only the keys present in `changes` are
* touched; anything absent keeps its current value.
*
* @param dict changes: any of `title`, `notes`, `column`
* @returns Task: this task, so calls chain.
* @throws TaskError on an invalid title or column.
*/
update(changes) {
if changes.contains('title') {
self.title = _clean_title(changes.title)
}
if changes.contains('notes') {
self.notes = changes.notes or ''
}
if changes.contains('column') {
self.column = _clean_column(changes.column)
}
return self
}
update() uses contains() rather than truthiness. That is the whole
difference between a partial update that works and one that does not: a
PATCH body of {"notes": ""} means “clear the notes”, and
if changes.notes would read that as “no change requested”.
Both return self, so board.get(id).move_to('done') reads as one thought.
Conversion
/**
* The task as a plain dictionary, which is what the store writes
* and what `from_dict()` reads back.
*/
to_dict() {
return {
id: self.id,
title: self.title,
notes: self.notes,
column: self.column,
created_at: self.created_at,
}
}
/**
* The same dictionary, plus the derived fields a template or an API
* client wants and should not have to compute.
*/
to_view() {
var view = self.to_dict()
view.set('created_on', date.from_time(self.created_at).format('M j, Y'))
view.set('is_done', self.column == 'done')
return view
}
/**
* What `json.encode()` uses, so a Task can be handed straight to a
* JSON response with no conversion at the call site.
*/
@to_json() {
return self.to_dict()
}
@to_string() {
return 'Task(${self.id}, ${self.column}, ${self.title})'
}
}
Two representations, on purpose. to_dict() is what gets stored, and it
holds exactly what from_dict() needs to rebuild the task. to_view() is
what gets displayed, and it adds things that are derived rather than
stored: a formatted date, a boolean the template can branch on.
Keeping them apart means a change to the display format never changes the file format.
@to_json() means a Task can be passed straight to response.json().
@to_string() means echo shows something useful when you look at a
task while debugging.
Rebuilding From Storage
/**
* Rebuilds a task from the dictionary `to_dict()` produced.
*
* @param dict data
* @returns Task
* @throws TaskError if the stored data is not a valid task.
*/
def from_dict(data: dict) {
return Task(data.get('title', nil), {
id: data.get('id', nil),
notes: data.get('notes', nil),
column: data.get('column', nil),
created_at: data.get('created_at', nil),
})
}
This is four lines with one important property: it goes through the same constructor as everything else.
It would have been shorter to assign the fields directly. Doing it this way
means a hand-edited board.json containing an unknown column, or a task
with no title, is rejected when the board loads rather than becoming a task
that nothing can render and nothing can fix.
data.get(key, nil) rather than data.key throughout, because a stored
record written by an older version of the program may be missing a key
entirely. get() with a fallback turns that into nil, which the
constructor’s own defaults then handle; data.column would raise
undefined key.
It is also the exact inverse of to_dict(). The two are a matched pair,
and a field added to Task has to appear in both or it will not survive a
restart — the kind of bug that only shows up the second time you run the
program.
Where the Rules Live
def _clean_title(title) {
if !is_string(title) {
raise TaskError('a task needs a title', 'title')
}
var cleaned = title.trim()
if cleaned.is_empty() {
raise TaskError('a task needs a title', 'title')
}
if cleaned.length() > 120 {
raise TaskError('a title must be 120 characters or fewer', 'title')
}
return cleaned
}
Read the order of those three checks, because it is the whole function.
The type check comes first. A PATCH body of {"title": 42} reaches
this function as a number, and 42.trim() would be a TypeError about
methods rather than a TaskError about titles. Checking first means the
caller gets an error in the application’s own vocabulary.
The trim happens before the emptiness check, so ' ' is rejected. A
title of three spaces is not a title, and checking title.is_empty() on
the raw input would have accepted it.
The length check happens after the trim, so trailing whitespace does not count against the limit.
And the function returns the cleaned value. It is not a validator that
answers yes or no; it is a normaliser that either produces a good value or
raises. That is why the constructor can write self.title = _clean_title(title) with nothing around it.
def _clean_column(column) {
if !config.COLUMNS.contains(column) {
raise TaskError(
'unknown column "${column}", expected one of ' + ', '.join(config.COLUMNS),
'column'
)
}
return column
}
The message names both what was wrong and what would have been right.
unknown column "backlog", expected one of todo, doing, done tells the
caller how to fix it; invalid column does not.
Note that the valid set comes from config.COLUMNS rather than being
written out here. Adding a column to the board is a one-line change in one
file, and this check, columns(), and summary() all follow it.
Two Call Sites Each
Both helpers are module-private — the leading underscore means no other
file can reach them — and each is called from exactly two places: the
constructor and update().
That is the property the whole chapter is built on. There is no path into
a Task that skips them. Not from_dict(), which goes through the
constructor. Not move_to(), which calls _clean_column() itself. Not a
route handler, which cannot reach the private function at all.
Every layer above can therefore assume a Task it is holding is valid,
which is why no route handler in
The JSON API re-checks a title.
Why Not the validate Module?
Zuri has a schema validator, and for a form with fifteen fields it is exactly right:
import validate
var schema = validate.schema({
title: validate.required().string().max_length(120),
column: validate.required().string(),
})
Here the rules are three lines of Zuri that also normalise (the trim()),
produce a domain error with a field on it, and live next to the data they
constrain. A schema would be a second place to look.
Use validate when the shape of the input is the problem. Use methods on
the class when the rules are part of what the thing is.
The JSON API
Six routes. Every one of them reads or writes through the board and returns JSON, and not one of them handles an error.
The Routes, One at a Time
The Shape of the File
Filename: routes/api.zu
import ..models { TaskError }
import ..storage { NotFoundError }
/**
* Registers every `/api` route on `server`, reading and writing
* through `board`.
*
* @param HttpServer server
* @param Board board
*/
def register(server, board) {
register(server, board) takes the server and the board rather than
importing them. That is what lets the same routes run against a different
board — a temporary one in a test, a seeded one in a demo — and it keeps
this file free of any decision about where the data lives.
Listing, With an Optional Filter
server.get('/api/tasks', @(request, response) {
var column = request.query_param('column', nil)
var tasks = column == nil ? board.all() : board.in_column(column)
response.json({
tasks: tasks.map(@(task) => task.to_view()),
summary: board.summary(),
})
})
One route serving two questions: every task, or every task in one column.
query_param('column', nil) supplies the fallback, and the test is
column == nil rather than if column. That distinction matters: a query
string of ?column= produces an empty string, which is falsy but present
— and an empty column name should be rejected by in_column(), not
silently treated as “no filter”.
The response carries summary alongside tasks because a client
rendering a board wants both, and making it ask twice would be two requests
for one screen.
Fetching One
server.get('/api/tasks/:id', @(request, response) {
response.json(board.get(request.param('id')).to_view())
})
One line, and it can fail. board.get() raises NotFoundError for an
unknown id, and this handler does nothing about it — which is the subject
of the section below.
Creating
server.post('/api/tasks', @(request, response) {
var body = request.json_body() or {}
var task = board.add(body.get('title', nil), {
notes: body.get('notes', nil),
column: body.get('column', nil),
})
response.json(task.to_view(), 201)
})
request.json_body() returns nil for a request with no body, or a body
that is not JSON at all. or {} turns that into an empty dictionary, so
body.get('title', nil) works either way and a missing title becomes the
model’s problem — which is where the message about it already lives.
The 201 is the second argument to response.json(). A created resource
is not a 200, and saying so is one character of effort.
task.to_view() rather than task, for the reason below.
Updating and Deleting
server.patch('/api/tasks/:id', @(request, response) {
var body = request.json_body() or {}
response.json(board.update(request.param('id'), body).to_view())
})
server.delete('/api/tasks/:id', @(request, response) {
board.remove(request.param('id'))
response.json({ deleted: true })
})
server.get('/api/summary', @(request, response) {
response.json(board.summary())
})
}
The PATCH handler passes the decoded body straight to board.update()
with no filtering. That is safe because update() only looks at the three
keys it knows — title, notes, column — using contains(), so a body
containing {"id": "hacked"} changes nothing. The whitelist lives in the
model, once, rather than in every route that accepts a body.
DELETE returns a body rather than a bare 204, because a client that
parses every response as JSON should not have to special-case one route.
Handlers Do Not Handle Errors
This is the chapter’s real point, and it is easiest to see by counting: six
handlers, zero catch blocks.
board.get() raises NotFoundError for an unknown id. board.add() and
board.update() raise TaskError for a bad title or an unknown column.
Not one handler catches either, and every one of them is one to five lines
as a direct result.
The middleware in the next section catches both and turns them into a 404 and a 422. There is exactly one place in this application that knows which domain error means which status code, and it is not in a route.
Consider the alternative for a moment. Six handlers each wrapping their
board call in a catch, each deciding on a status, each formatting an
error body — thirty-odd lines of duplication, and a seventh route added
next month that gets one of them subtly wrong.
Note also the two imports at the top of the file. TaskError and
NotFoundError are imported and never mentioned again in this file; they
are there because the module that catches them needs them re-exported
through routes. That is the module system doing its job, and
Chapter 8 covers why the @ matters there.
to_view() at the Boundary
Every response sends to_view(), never the Task itself.
That single habit means the API’s shape is a decision Task makes in one
method, rather than something that leaks out of however the object happens
to be laid out. Add a private field to Task tomorrow and the API returns
exactly what it returned yesterday.
It is also why the responses carry created_on and is_done, which are
not stored anywhere — to_view() computes them, so every client gets a
formatted date and a boolean without doing the work itself.
The Routes in Practice
$ curl -s localhost:8000/api/tasks -H 'content-type: application/json' \
-d '{"title":"write the capstone","notes":"chapter 17"}'
{"id":"01a08e15-e758-741d-b368-d5a988cb90cd","title":"write the capstone",
"notes":"chapter 17","column":"todo","created_at":1789090195.9,
"created_on":"Sep 11, 2026","is_done":false}
$ curl -s localhost:8000/api/tasks
{"tasks":[...],"summary":{"todo":1,"doing":0,"done":0}}
$ curl -s -X PATCH localhost:8000/api/tasks/01a08e15-... \
-H 'content-type: application/json' -d '{"column":"doing"}'
{"id":"01a08e15-...","column":"doing",...}
$ curl -s localhost:8000/api/tasks?column=doing
{"tasks":[...],"summary":{"todo":0,"doing":1,"done":0}}
$ curl -s localhost:8000/api/tasks -H 'content-type: application/json' -d '{"title":""}'
{"error":"a task needs a title","field":"title"}
$ curl -s localhost:8000/api/tasks/does-not-exist
{"error":"no task with id does-not-exist"}
The last two are the interesting ones. A 422 with a field, and a 404 with
a message, from handlers that said nothing at all about status codes.
Server-Rendered Pages with Wire
The browser side is two templates and four routes. Wire’s full reference is its own chapter; this section uses the parts a real page needs.
The Layout
Filename: templates/layout.html
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title x-text="title">Task Board</title>
<link rel="stylesheet" href="/static/app.css">
</head>
<body>
<header class="masthead">
<h1>Task Board</h1>
<p class="tagline" x-text="tagline"></p>
</header>
<main>
<template x-slot="content">
<p>Nothing here yet.</p>
</template>
</main>
<footer>
<p>Served by Zuri.</p>
</footer>
</body>
</html>
This is a valid HTML5 document. Open it in a browser and you get a page with a heading and a paragraph. That is the whole point of Wire’s design: the directives are attributes, so the template is still the thing it renders.
x-slot="content" declares a region an extending template may replace. The
content inside it is the default, used when nothing replaces it.
The Page
Filename: templates/board.html
<extend base="layout.html">
<define name="content">
<form class="new-task" method="post" action="/tasks">
<input name="title" placeholder="What needs doing?" maxlength="120" required>
<select name="column">
<option x-for="columns" x-value="column"
x-attr="{ value: column.name }" x-text="column.name"></option>
</select>
<button type="submit">Add</button>
</form>
<p class="error" x-if="error" x-text="error"></p>
<div class="board">
<section class="column" x-for="columns" x-value="column">
<h2>
<span x-text="column.name"></span>
<span class="count">{{ column.tasks|length }}</span>
</h2>
<p class="empty" x-not="column.tasks">Nothing here.</p>
<article class="task" x-for="column.tasks" x-value="task">
<h3 x-text="task.title"></h3>
<p class="notes" x-if="task.notes" x-text="task.notes"></p>
<p class="meta">Added <span x-text="task.created_on"></span></p>
<form method="post" x-attr="{ action: '/tasks/' + task.id + '/move' }">
<select name="column">
<option x-for="columns" x-value="target"
x-attr="{ value: target.name }" x-text="target.name"></option>
</select>
<button type="submit">Move</button>
</form>
<form method="post" x-attr="{ action: '/tasks/' + task.id + '/delete' }">
<button type="submit" class="danger">Delete</button>
</form>
</article>
</section>
</div>
</define>
</extend>
Six directives carry the whole page.
x-for with x-value repeats an element once per entry, binding each
one to a name. The nested x-for="column.tasks" inside
x-for="columns" is an ordinary nested loop, and the inner one can still
see columns from the outer scope, which is how the move dropdown lists
every column.
x-text replaces an element’s children with escaped text. Nothing a
task’s title contains can become markup.
x-attr takes a dictionary and spreads it onto the element, which is
how a form’s action gets built from a task id.
x-if / x-not render an element conditionally. x-not="column.tasks"
shows the “Nothing here.” line when the list is empty, because in Wire an
empty list is falsy. That is not Zuri’s rule, where [] is truthy; Wire
uses its own, friendlier one for template authoring.
{{ column.tasks|length }} is an interpolation through a filter.
length is one of Wire’s built-in filters, and it works on strings, lists,
dictionaries and bytes.
The Routes
routes/pages.zu registers four routes: one that renders the board, and
three that handle the forms on it. Here it is a piece at a time.
The Shape of the File
Filename: routes/pages.zu
import ..config
/**
* Registers the page and form routes on `server`.
*
* @param HttpServer server
* @param Board board
* @param Wire view: the template engine to render through
*/
def register(server, board, view) {
The whole file is one register() function that takes its dependencies as
arguments. Nothing here reaches for a global board or a global template
engine, which is what makes the routes testable: hand register() a board
backed by a temporary directory and the same code runs against it.
This is the “pass dependencies in” habit from Debugging, applied at the layer where it costs nothing and buys the most.
Rendering the Board
server.get('/', @(request, response) {
response.html(view.render('board.html', {
title: 'Task Board',
tagline: _tagline(board),
columns: board.columns(),
error: request.query_param('error', nil),
}))
})
One route, one call, no logic. Everything it hands the template is either a constant or a method call on the board — there is no loop, no formatting and no branching in this handler, because all of that already happened somewhere better suited to it.
board.columns() does the grouping. _tagline() does the counting.
to_view(), back in the domain model, did the date formatting. By the time
the template runs, every value it needs is sitting in front of it.
request.query_param('error', nil) is the other half of the redirect
pattern below: a failed form redirects with ?error=..., and this is where
that message comes back in to be rendered.
Handling a Form
server.post('/tasks', @(request, response) {
var form = request.form()
catch {
board.add(form.get('title', nil), { column: form.get('column', nil) })
} as error {
response.redirect('/?error=' + _escape(error.message))
return
}
response.redirect('/')
})
All three form handlers follow this exact shape, so it is worth reading once carefully.
request.form() parses a URL-encoded body into a dictionary.
form.get('title', nil) rather than form.title, because a browser can
post a body with any fields at all — or none — and form.title would raise
undefined key on a request that simply omitted it. With get(), a
missing field arrives as nil, and _clean_title() turns that into a
proper TaskError.
The catch wraps only the board call. response.redirect('/') on the
success path sits outside it, so a mistake in the redirect is not reported
as a validation failure. That is the “keep the catch block small” rule from
Chapter 7.
The return inside the handler is what stops execution falling through to
the success redirect. Without it, a failed add would issue two redirects.
The Other Two
server.post('/tasks/:id/move', @(request, response) {
catch {
board.update(request.param('id'), {
column: request.form().get('column', nil),
})
} as error {
response.redirect('/?error=' + _escape(error.message))
return
}
response.redirect('/')
})
server.post('/tasks/:id/delete', @(request, response) {
catch {
board.remove(request.param('id'))
} as error {
response.redirect('/?error=' + _escape(error.message))
return
}
response.redirect('/')
})
}
:id in the path is a route parameter, and request.param('id') reads it.
Notice what is not in these handlers. Neither checks that the id names
a real task — board.get() raises NotFoundError and the catch picks it
up. Neither validates the column — _clean_column() does. Neither checks
that the task exists before deleting it — remove() calls get() first.
That is the payoff for putting the rules in the domain model. A route handler is four lines because there is nothing left for it to do.
The Two Helpers
def _tagline(board) {
var counts = board.summary()
return '${counts.todo} to do, ${counts.doing} in progress, ${counts.done} done'
}
def _escape(message) {
return message.replace('/[^a-zA-Z0-9 .,-]/', '').replace(' ', '+', false)
}
_tagline() turns the board’s counts into the line under the heading. It
lives here rather than in Board because it is a presentation decision:
another front end would word it differently, and the board should not have
an opinion.
_escape() is doing something more careful than it looks. The message is
about to be put into a URL, so it strips everything that is not a letter,
digit, space or basic punctuation, then turns spaces into +.
Two details in that one line. The first replace() uses a regular
expression; the second passes false as the third argument to turn
pattern handling off, so the single space is matched literally rather
than as a pattern. And stripping rather than percent-encoding is the
deliberate choice: this is a message we generated, not user input echoed
back, so a conservative allow-list is simpler than encoding and cannot
produce a malformed URL.
Redirect After Post
Every form handler ends in a redirect rather than rendering a page. That is
the post/redirect/get pattern, and it exists because a browser that
rendered a page in response to a POST will re-submit that POST when the
user presses refresh. Adding a task twice because someone hit F5 is not a
bug you want to explain.
The failure path redirects too, carrying the message as a query parameter
that the next GET renders into the error banner. So both outcomes leave
the browser sitting on a plain GET /, which is refreshable, bookmarkable
and safe to go back to.
Why These Handlers Catch and the API’s Do Not
The API handlers in The JSON API let errors propagate, because the middleware turns them into status codes. These catch, because a browser submitting a form does not want a 422 page — it wants the board back with a message on it.
Same domain errors, two presentations, and the choice is made at the layer
that knows which kind of client it is talking to. Neither the Board nor
the Task has to know that a browser is involved.
The Stylesheet
server.serve_files('/static', config.STATIC_DIR) mounts the directory.
The static-file handler deals with content types, ETags, conditional
requests and range requests on its own, so app.css is a plain file with
nothing around it:
$ curl -s -o /dev/null -w '%{http_code} %{content_type}\n' localhost:8000/static/app.css
200 text/css
What It Renders
<!DOCTYPE html><html lang="en"><head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Task Board</title>
<link rel="stylesheet" href="/static/app.css">
</head>
<body>
<header class="masthead">
<h1>Task Board</h1>
<p class="tagline">1 to do, 0 in progress, 0 done</p>
</header>
...
<div class="board">
<section class="column">
<h2>
<span>todo</span>
<span class="count">1</span>
</h2>
<article class="task">
<h3>write the capstone</h3>
<p class="notes">chapter 17</p>
<p class="meta">Added <span>Sep 11, 2026</span></p>
<form method="post" action="/tasks/01a08e15-e758-741d-b368-d5a988cb90cd/move">
...
Server-rendered HTML, no JavaScript, and every value escaped by the engine rather than by the person who wrote the template.
Middleware, Logging and Errors
Two pieces of behaviour belong to every request and to no handler: logging what happened, and turning a domain error into a status code. Both are middleware.
Filename: app.zu
import http
import log
import wire
import .config
import .models { TaskError }
import .routes { api, pages }
import .storage { Board, NotFoundError }
/**
* Builds the application.
*
* @param ?number port: defaults to `config.PORT`; pass `0` to let the
* operating system choose a free one.
* @param ?string data_dir: defaults to `config.DATA_DIR`
* @returns HttpServer: bound to nothing yet; the caller decides how to
* serve it.
*/
def build(port, data_dir) {
var board = Board(data_dir)
var view = wire.wire({
root: config.TEMPLATE_DIR,
auto_reload: true,
})
var server = http.HttpServer(port == nil ? config.PORT : port, config.HOST)
server.max_body_size = 64 * 1024
_install_middleware(server)
api.register(server, board)
pages.register(server, board, view)
server.get('/health', @(request, response) {
response.json({ ok: true, tasks: board.all().length() })
})
server.serve_files('/static', config.STATIC_DIR)
return server
}
That one function is the whole assembly, and the order of its lines is the order of the application.
It builds the board first, because everything else needs it. Passing
data_dir through rather than letting Board find it means a test, or a
second instance, can point at a different directory without touching
config.
auto_reload: true on the template engine re-reads a template when the
file changes, which is what you want while writing one. It is also the
first line to reconsider before a deployment, since it costs a filesystem
check per render.
port == nil ? config.PORT : port is not port or config.PORT. That
distinction is load-bearing here: port 0 is a legitimate value meaning
“let the operating system pick a free one”, and 0 is falsy, so the or
form would silently turn it into the configured port. This is the
negative-and-zero trap from Chapter 3 in a place
where it would really bite.
max_body_size is set explicitly. A server that accepts a request body
of any size is a server anyone can exhaust with a single request. 64 KiB is
generous for a form that carries a title and a column.
Middleware is installed before the routes. Middleware runs outermost
first, so anything registered here wraps every route added afterwards —
including /health and the static files below.
The routes are registered by handing each module what it needs.
api.register(server, board) and pages.register(server, board, view) are
the same shape: a function, given its dependencies, that attaches things to
the server. The page routes get the template engine; the API routes do not,
because they never render one.
/health is defined here rather than in a route module, because it is
about the process rather than about tasks. It reports a task count so that
a monitoring check proves the board actually loaded, not merely that the
socket answers.
serve_files() comes last, mapping /static onto a directory. It is
last because more specific routes should be registered before a catch-all
prefix.
And build() returns the server without binding a socket. Nothing here
listens. That separation is what lets the entry point decide how to serve
it — one process, a pool of isolates, or a test harness that never listens
at all — and it is what makes the next section possible.
The Middleware
/**
* Request logging, and one error handler that turns every exception
* the routes can raise into the right status code. Handlers below it
* are then free to raise and say nothing about HTTP.
*/
def _install_middleware(server) {
server.use(@(request, response, next) {
var started = time()
next()
var elapsed = ((time() - started) * 1000).round()
log.info('${request.method} ${request.path} -> ${response.status} (${elapsed}ms)')
})
server.use(@(request, response, next) {
catch {
next()
} as error {
_render_error(request, response, error)
}
})
}
A middleware takes three arguments: the request, the response, and a
next function. next() is where everything below it runs — the other
middleware, and eventually the route handler itself.
That one fact explains the shape of both functions here. Code written
before the next() call happens on the way in; code written after it
happens on the way out, once the handler has finished and the response is
populated. A middleware is therefore a pair of moments, not a single step,
and the call in the middle is the seam between them.
Reading the Logger
server.use(@(request, response, next) {
var started = time()
next()
var elapsed = ((time() - started) * 1000).round()
log.info('${request.method} ${request.path} -> ${response.status} (${elapsed}ms)')
})
The timestamp is taken on the way in, before anything else runs. The log
line is written on the way out, which is the only point at which
response.status is known — on the way in, nothing has decided it yet.
That is the whole reason timing middleware works: the same function body runs at both ends of the request, with a local variable surviving in between.
Reading the Error Handler
server.use(@(request, response, next) {
catch {
next()
} as error {
_render_error(request, response, error)
}
})
This one has nothing before next() and nothing after it. All its work is
in the handler, and what it catches is everything that raised anywhere
below it — a route, the board, the model, the store.
That is the mechanism the whole application leans on. A catch around
next() catches the routes, because next() is the routes.
Why the Logger Comes First
Middleware run outermost first, in the order they were added. So the logger wraps the error handler, which wraps the routes:
logger starts the clock
errors catches whatever escapes
routes handles the request
errors turns an error into a status
logger writes the line, with that status
Read the order bottom to top on the way out and the reason becomes obvious: a request that failed still gets logged, and it gets logged with the status the error handler chose rather than with whatever the response held when the exception was raised.
Swap the two and a failing request would be logged before the error handler had set a status — or not at all, if the exception escaped the logger first.
The Module Already Has One
http.middleware.logger() is a ready-made request logger, and in a real
application it is what you would reach for:
import http.middleware
server.use(middleware.logger())
It writes the Common Log Format extended with the response time, which
every log analyser already parses, and it takes a sink to send lines
somewhere other than standard output, a format to build the line
yourself, and trust_proxy to log the forwarded client address instead of
the peer.
The nine lines above are written out by hand here because a middleware you
have read the whole of teaches more than one you called. Once you can see
what next() does, swap in the real one.
One Place That Knows About Status Codes
def _render_error(request, response, error) {
var status = 500
var body = { error: 'internal error' }
if instance_of(error, NotFoundError) {
status = 404
body = { error: error.message }
} else if instance_of(error, TaskError) {
status = 422
body = { error: error.message, field: error.field }
} else {
log.error('unhandled: ' + error.message)
}
if request.path.starts_with('/api') or request.wants_json() {
response.json(body, status)
return
}
response.html('<h1>' + status + '</h1><p>' + body.error + '</p>', status)
}
This is the payoff for defining NotFoundError and TaskError as real
classes instead of raising Error('not found') everywhere.
instance_of() maps a domain error to a status code, once.
Three details matter.
Unknown errors are a 500 with a generic body, and the real message goes to the log. An error message can contain a path, a query, or a fragment of a file. It goes where operators can read it, not where users can.
The response shape follows the client. A request under /api, or one
whose Accept header asks for JSON, gets JSON. Everything else gets HTML.
request.wants_json() reads the header for you.
A handler that has already committed a response is not overwritten,
because the error handler only ever runs when next() raised.
What It Looks Like
2026-09-11T02:29:55+01:00 INFO [taskboard]: task board on http://127.0.0.1:8000
2026-09-11T02:29:55+01:00 INFO [taskboard]: POST /api/tasks -> 201 (15ms)
2026-09-11T02:29:55+01:00 INFO [taskboard]: GET / -> 200 (409ms)
2026-09-11T02:30:02+01:00 INFO [taskboard]: POST /api/tasks -> 422 (3ms)
2026-09-11T02:30:02+01:00 INFO [taskboard]: GET /api/tasks/nope -> 404 (0ms)
2026-09-11T02:30:03+01:00 INFO [taskboard]: GET /static/app.css -> 200 (1ms)
log puts the timestamp, the level and the module name on every line and
colours it when the output is a terminal. The module name comes from where
the call was made, so a message from routes/api.zu says so without being
told.
The 409ms on that first GET / is Wire compiling the two templates. Every
render after it walks the cached instruction tree instead, which is the
1-3ms the later lines show.
Middleware Worth Adding
http.middleware has ready-made pieces for the things every application
eventually wants: CORS, compression, rate limiting, and authentication.
Each is a function you pass to server.use(), in the position you want it
in the chain.
Running It for Real
$ zuri run taskboard
2026-09-11T02:29:55+01:00 INFO [taskboard]: task board on http://127.0.0.1:8000
Open http://127.0.0.1:8000 and the board is there. Add a task, move it,
delete it. Stop the server with Ctrl+C and start it again, and everything
is still there.
The Entry Point
Filename: index.zu
import log
import os
import @.app
import .config
if __root__ == __file__ {
var server = app.build()
os.on_signal('INT', @() {
log.info('shutting down')
os.exit(0)
})
log.info('task board on http://${config.HOST}:${config.PORT}')
server.listen()
}
import @.app re-exports app, so another program can
import .taskboard { app } and build a server of its own. The
__root__ == __file__ check means importing this package starts nothing.
os.on_signal('INT', ...) catches Ctrl+C. This application saves after
every change, so there is nothing to flush, and the handler is here because
the place to put “finish what you were doing” is obvious once the hook
exists.
Configuration
Every setting comes from the environment with a default:
$ PORT=9000 HOST=0.0.0.0 DATA_DIR=/var/lib/taskboard zuri run taskboard
Running it with no environment at all works too, which is what makes the first run painless.
Exercising It
$ curl -s localhost:8000/health
{"ok":true,"tasks":0}
$ curl -s -X POST -H 'content-type: application/json' \
-d '{"title":"write the capstone","notes":"chapter 17"}' \
localhost:8000/api/tasks
{"id":"01a08e15-...","title":"write the capstone","notes":"chapter 17",
"column":"todo","created_at":1789090195.9,"created_on":"Sep 11, 2026",
"is_done":false}
$ curl -s -X POST -d 'title=review+it&column=doing' localhost:8000/tasks -o /dev/null -w '%{http_code} -> %{redirect_url}\n'
302 -> http://localhost:8000/
$ curl -s -X POST -d 'title=&column=todo' localhost:8000/tasks -o /dev/null -w '%{redirect_url}\n'
http://localhost:8000/?error=a+task+needs+a+title
$ curl -s localhost:8000/api/summary
{"todo":0,"doing":1,"done":1}
The HTML form and the JSON API reach the same board through the same validation, and each reports failure in the form its own client understands.
What Is Stored
$ cat taskboard/data/board.json
[
{
"id": "01a08e15-e758-741d-b368-d5a988cb90cd",
"title": "write the capstone",
"notes": "chapter 17",
"column": "todo",
"created_at": 1789090195.9
}
]
Readable, editable, and greppable. An invalid column typed in by hand is
rejected when the board next loads, because from_dict() goes through the
same constructor everything else does.
The Single-Process Assumption
server.listen() serves one connection at a time on one isolate. For a
team board that is more than enough: the slowest thing in a request is
Wire’s first compile, and everything after it is single-digit milliseconds.
Zuri can serve across cores, and http.serve() is how. It takes a setup
function, calls it once inside each worker with that worker’s own server,
and runs the pool:
import http
import .app
# In app.zu, alongside build():
#
# def setup(server) {
# ...register the same routes and middleware here...
# }
http.serve(app.setup, { port: 8000, workers: 4 })
Doing it here would change one thing that matters: each isolate has its
own heap, so each worker would have its own Board, its own copy of
every task, and its own idea of what board.json should contain. Two
workers saving at once would each write a complete file, and one of them
would win.
The board would need to stop being in-memory state. The choices are the usual ones:
- A database. Move
_json_store.zuto a real store with transactions. Nothing above it changes; that is why it is one private file. - One owner. Keep the board in a single isolate and have the workers talk to it over a channel. Every mutation becomes a message, and one isolate serialises them.
- A lock. Guard the file with an advisory lock and re-read before every write. Simplest to add, and it makes every request pay for the file.
Which one is right depends on how many people are using the board, and none of them is worth doing before the answer is “more than this can handle.” The design already leaves the door open, and that is the part to get right early.
What This Application Used
Almost all of it.
| Chapter | What the application uses it for |
|---|---|
| 3, 4 | every line |
| 5 | anonymous handlers, closures over board, typed parameters |
| 6 | Task, Board, custom errors, @to_json, to_string |
| 7 | TaskError, NotFoundError, catch at the form boundary |
| 8 | packages, index.zu, @ re-export, _ privacy, __file__ |
| 9 | file(), atomic rename, os.join_paths, os.get_env |
| 12 | the HTTP server, routing, middleware, static files |
| 13 | json, uuid, date, log |
| 14 | Wire: layout, slots, loops, filters, escaping |
| 16 | the request log, and the error shapes that make it readable |
| 19 | the suite in Testing the Board |
What it did not use is as informative. There are no isolates, because one
process is enough. There is no validate, because the rules belong to the
model. There is no binary handling, no compression, no reflection. Those
are all available, and reaching for them here would have made the
application longer without making it better.
Where to Take It
The natural next steps, roughly in order of how much they teach:
Authentication. bcrypt for password hashing, a signed cookie for the
session, a middleware that rejects anything unauthenticated. The middleware
slot is already there.
Live updates. http.sse pushes an event to the browser when the board
changes, so two people looking at it see each other’s edits.
Per-user boards. One board.json per user, a Board cache keyed by
user id, and path becoming a parameter rather than a default.
A database. Replace _json_store.zu, change nothing else, and watch
the underscore earn its keep.
Search and filtering. filter() over the board, exposed as query
parameters, with the same code serving the HTML page and the API.
Every one of those is a change to one layer. That is what the layout was for, and Testing the Board is how you make any of them without holding your breath.
Testing the Board
The application works. This section is about keeping it working, and it is the last thing the layering from Laying Out the Project pays for.
Recall the shape:
index.zu -> app.zu -> routes/ -> storage/ -> models/
That is also the order to test in. models depends on nothing but
config, so its tests need nothing. storage depends on models and a
file, so its tests need a directory. routes depends on storage and a
server, and takes both as arguments, so its tests need neither a socket nor
a real board.
Each layer is tested against the layer below it as it really is, and against the layer above it not at all.
Where the Tests Live
taskboard/
index.zu
app.zu
config.zu
models/
storage/
routes/
tests/
task.zu
board.zu
api.zu
$ zuri test
That is the whole arrangement. The application is zuri run taskboard
and its tests are zuri test, and neither needs a file name
remembering. The command exits 1 when anything failed, which is all CI
needs.
Each test file is an ordinary script: it declares its tests and stops.
Each one runs in a process of its own, which matters here for one
concrete reason: every one of these files is going to create a Board,
and a Board is a file on disk. One process per file means one file’s
leftovers can never reach another’s.
The Model
models/task.zu is pure. No file, no clock you care about, no network. Its
tests are the fastest and the ones worth writing first, because every rule
about what a task is lives there and nothing above re-checks any of them.
Filename: tests/task.zu
import test { * }
import ..models { Task, TaskError, from_dict }
describe('Task', @{
describe('construction', @{
it('needs a title', @{
expect(@{ Task(nil) }).to_raise_instance_of(TaskError)
expect(@{ Task('') }).to_raise_instance_of(TaskError)
expect(@{ Task(' ') }).to_raise_instance_of(TaskError)
})
it('says which field was wrong', @{
catch {
Task(nil)
} as error {
expect(error.field).to_be('title')
}
})
it('refuses a title over 120 characters', @{
expect(@{ Task('x' * 121) }).to_raise_instance_of(TaskError)
expect(@{ Task('x' * 120) }).to_not_raise()
})
it('defaults everything but the title', @{
var task = Task('Write the tests')
expect(task.notes).to_be('')
expect(task.column).to_be('todo')
expect(task.id).to_be_string()
expect(task.created_at).to_be_number()
})
it('sorts by id, because v7 embeds the time', @{
var first = Task('one')
var second = Task('two')
expect([first.id, second.id]).to_be_sorted()
})
})
describe('update()', @{
it('touches only the keys it was given', @{
var task = Task('Write the tests', { notes: 'keep me' })
task.update({ column: 'doing' })
expect(task.column).to_be('doing')
expect(task.notes).to_be('keep me')
expect(task.title).to_be('Write the tests')
})
it('clears a value when the key is present and empty', @{
var task = Task('Write the tests', { notes: 'remove me' })
task.update({ notes: '' })
expect(task.notes).to_be('')
})
it('refuses a column the board does not have', @{
var task = Task('Write the tests')
expect(@{ task.update({ column: 'sideways' }) }).to_raise_instance_of(TaskError)
expect(task.column).to_be('todo')
})
})
describe('round-tripping', @{
it('rebuilds a task from what it stored', @{
var original = Task('Write the tests', { notes: 'a note', column: 'doing' })
expect(from_dict(original.to_dict())).to_equal(original)
})
it('rejects a stored record that broke a rule', @{
expect(@{ from_dict({ id: 'x', title: '', column: 'todo' }) })
.to_raise_instance_of(TaskError)
})
it('survives a record written before a field existed', @{
var task = from_dict({ title: 'Write the tests' })
expect(task.column).to_be('todo')
expect(task.notes).to_be('')
})
})
})
Four of those are worth pointing at.
expect(@{ task.update(...) }).to_raise_instance_of(TaskError) followed
by expect(task.column).to_be('todo'). Two assertions, because there are
two claims: that it refused, and that it refused before changing
anything. A validator that raises after assigning is a bug you only find
by checking the second one.
to_equal, not to_be, for the round trip. from_dict() builds a new
Task, so identity is never going to match. The whole question is whether
the contents survived.
expect([first.id, second.id]).to_be_sorted() is the cheapest possible
way to state what uuid.v7() was chosen for. If someone ever changes that
line to v4(), this is the test that says so.
The record with no column and no notes. That is the
older-version-of-the-program case that data.get(key, nil) exists to
handle. Writing the test is what stops the next person from simplifying
get() into .column and finding out on someone’s real board.json.
The Storage Layer
A Board is a file, so its tests need a directory of their own. One per
test, created and removed by the hooks:
Filename: tests/board.zu
import test { * }
import os
import ..models { Task }
import ..storage { Board, NotFoundError }
describe('Board', @{
var directory = nil
var board = nil
before_each(@{
directory = os.create_temp_dir('taskboard-test')
board = Board(directory)
})
after_each(@{
os.remove_dir(directory, true)
})
it('starts empty', @{
expect(board.all()).to_be_empty()
expect(board.summary()).to_match_object({ todo: 0, doing: 0, done: 0 })
})
it('keeps what it was given', @{
board.add('Write the tests')
expect(board.all()).to_have_length(1)
expect(board.all()[0].title).to_be('Write the tests')
})
it('survives a restart', @{
var created = board.add('Write the tests', { notes: 'a note' })
expect(Board(directory).get(created.id)).to_equal(created)
})
it('raises for an id it does not have', @{
expect(@{ board.get('no-such-id') }).to_raise_instance_of(NotFoundError)
})
it('filters by column', @{
board.add('one')
var moved = board.add('two')
board.update(moved.id, { column: 'doing' })
expect(board.in_column('todo')).to_have_length(1)
expect(board.in_column('doing')).to_have_length(1)
expect(board.in_column('done')).to_be_empty()
})
it('counts what it holds', @{
board.add('one')
board.add('two')
expect(board.summary()).to_match_object({ todo: 2 })
})
it('removes', @{
var created = board.add('Write the tests')
board.remove(created.id)
expect(board.all()).to_be_empty()
expect(@{ board.get(created.id) }).to_raise_instance_of(NotFoundError)
})
})
The one that earns its place is survives a restart. It constructs a
second Board over the same directory and asks it for the task the first
one created. That is the only test here that exercises _json_store,
to_dict() and from_dict() together, and it is the test that fails the
day someone adds a field to Task and forgets one half of the pair.
after_each removes the directory whether the test passed or not, so a
failing test leaves nothing behind for the next one to trip over. That is
the property to rely on: teardown that only runs on success is teardown you
cannot trust.
Note that none of these tests read board.json themselves. They ask the
Board what it holds. The file format is storage’s business, and a test
that parsed it would fail the day the format changed for a reason that has
nothing to do with what the test was checking.
The Routes
register(server, board) takes its two collaborators as arguments. That
was presented in The JSON API as being about seeding a
demo board; here is the other half of what it buys.
A handler needs three things: something to register on, a request, and a response. Stand in for all three:
Filename: tests/api.zu
import test { * }
import os
import ..routes { register_api }
import ..storage { Board }
class FakeServer {
var routes = {}
@new() {
self.routes = {}
}
get(path, handler) {
self.routes.set('GET ${path}', handler)
}
post(path, handler) {
self.routes.set('POST ${path}', handler)
}
patch(path, handler) {
self.routes.set('PATCH ${path}', handler)
}
delete(path, handler) {
self.routes.set('DELETE ${path}', handler)
}
}
class FakeRequest {
var params = {}
var query = {}
var body = nil
@new(options) {
options = options or {}
self.params = options.get('params', {})
self.query = options.get('query', {})
self.body = options.get('body', nil)
}
param(name) {
return self.params.get(name, nil)
}
query_param(name, fallback) {
return self.query.get(name, fallback)
}
json_body() {
return self.body
}
}
class FakeResponse {
var body = nil
var status = 200
json(body, status) {
self.body = body
self.status = status or 200
}
}
Three small classes and the routes become ordinary functions:
describe('the JSON API', @{
var directory = nil
var server = nil
var board = nil
before_each(@{
directory = os.create_temp_dir('taskboard-test')
board = Board(directory)
server = FakeServer()
register_api(server, board)
})
after_each(@{
os.remove_dir(directory, true)
})
def call(route, options) {
var response = FakeResponse()
server.routes[route](FakeRequest(options), response)
return response
}
it('registers every route', @{
expect(server.routes).to_have_keys([
'GET /api/tasks',
'GET /api/tasks/:id',
'POST /api/tasks',
'PATCH /api/tasks/:id',
'DELETE /api/tasks/:id',
'GET /api/summary',
])
})
it('answers 201 when it creates a task', @{
var response = call('POST /api/tasks', { body: { title: 'Write the tests' } })
expect(response.status).to_be(201)
expect(response.body).to_match_object({ title: 'Write the tests', column: 'todo' })
expect(board.all()).to_have_length(1)
})
it('lets the model reject a bad body', @{
expect(@{ call('POST /api/tasks', { body: {} }) }).to_raise_with_message('title')
})
it('treats a missing body as an empty one', @{
expect(@{ call('POST /api/tasks', {}) }).to_raise_with_message('title')
})
it('lists everything, with a summary', @{
board.add('one')
var response = call('GET /api/tasks', {})
expect(response.body).to_have_keys(['tasks', 'summary'])
expect(response.body.tasks).to_have_length(1)
})
it('filters by column when asked', @{
board.add('one')
expect(call('GET /api/tasks', { query: { column: 'done' } }).body.tasks).to_be_empty()
})
it('sends the view, not the stored task', @{
var created = board.add('Write the tests')
var response = call('GET /api/tasks/:id', { params: { id: created.id } })
expect(response.body).to_have_keys(['created_on', 'is_done'])
})
it('ignores a key the model does not own', @{
var created = board.add('Write the tests')
call('PATCH /api/tasks/:id', {
params: { id: created.id },
body: { id: 'hacked', column: 'doing' },
})
expect(board.get(created.id).column).to_be('doing')
expect(board.get(created.id).id).to_be(created.id)
})
})
Several things worth saying about that.
The board is real. Only the HTTP machinery is faked. A fake board would
have meant writing down what board.add() returns, and the test would then
pass forever afterwards regardless of what board.add() actually did.
Faking is for the things that are slow, remote or awkward, and a temporary
file is none of those.
lets the model reject a bad body. From
Handlers Do Not Handle Errors:
the handler does not catch anything, so an invalid title comes back out of
the handler as a TaskError. That is the behaviour the middleware relies
on, and this asserts it directly rather than through the middleware.
ignores a key the model does not own is the whitelist claim from
update(), tested where an attacker would aim it. One line of test for a
property the code gets for free from contains(), and the day someone
“simplifies” that method it fails.
Use a class, not a dictionary, for a fake. A dictionary already has
get, add, set and keys of its own, so fake.get('id') calls the
dictionary’s method rather than yours. A small class gives you the names
you meant.
What to Test Next
Two layers are deliberately not tested above.
The pages. routes/pages.zu renders templates. The same fake server
works, and the assertion becomes expect(response.body).to_contain(...)
against the rendered HTML, or to_match_snapshot(), which records the
whole page once and watches it from then on:
it('renders the board', @{
expect(call('GET /', {}).body).to_match_snapshot()
})
That records the page in tests/__snapshots__/pages.zu.snap. Reviewing
the diff when it changes is the point; a template edit that alters more of
the page than you intended shows up there and nowhere else.
The whole thing. Starting the real server on port 0, making real
requests with the http client, and shutting it down in after_all gives
you one test that covers routing, middleware, templates and storage
together. Write a handful of those, not a hundred: they are the slowest
tests you own and the ones that break for reasons unrelated to what they
were checking.
What This Bought
Run it:
$ zuri test
zuri test 3 files in tests
PASS api.zu 312ms 8 tests
PASS board.zu 198ms 7 tests
PASS task.zu 164ms 10 tests
3 files
25 passed • 25 total
time 674ms
Twenty-five tests, under a second, no server and no network. That is a direct consequence of the layering: every arrow points one way, and every layer takes its collaborators as arguments rather than importing them.
The layering was justified in Laying Out the Project on the grounds that it makes each layer readable on its own. This is the other half of the claim, and it is the half you feel every day.
Foreign Functions: C and Rust
The ffi module calls code written in C and in Rust. It loads a shared
library, describes the functions in it, and calls them with ordinary Zuri
values; it turns Zuri functions into C function pointers so a library can
call back; and it reads the declarations a header or a crate already
contains, so the description is written once, by the people who wrote the
library.
Every value that crosses is converted and checked against its C type. An
integer that does not fit is a RangeError at the call, not a truncated
argument inside the library. Memory the module allocates knows its size,
and a read past its end is refused before it happens. Where a check cannot
be made, because C handed back an address with no extent attached, the
module says so plainly rather than guessing.
A C library, a Rust cdylib, and a static library of either are all
reachable, and the layout of every record, packed, over-aligned or full of
bitfields, is the one the platform’s own compiler would give it.
- Following Along
- Introduction
- Loading a Library
- Calling a Function
- Types
- Numbers at the Boundary
- Declaring from C
- Declaring from Rust
- Pointers and Memory
- Strings and Text
- Structs, Unions and Arrays
- Enums and Constants
- Callbacks
- Callbacks and Threads
- Variadic Functions
- Errors and errno
- Ownership and Lifetimes
- Static Libraries
- Isolates
- Rust in Depth
- Platform Differences
- What Is Checked
- What the Module Refuses
- Module Reference
Blocks on this page that list several calls together are reference listings, not programs: they show the shape of each call rather than a sequence to run. Anything presented as a complete program runs as written. The programs that call the C runtime run on Linux as shown; the ones that need a library of your own are shown rather than run.
Following Along
The C runtime is on every machine, so most examples below call it.
ffi.LIBC names it for the platform the program is running on, and
ffi.LIBM names the maths library:
import ffi
var libc = ffi.open(ffi.LIBC)
var strlen = libc.function('strlen', ffi.size_t, [ffi.string])
echo strlen('Hello, C')
8
Some sections use a small library of your own. Save this as
geometry.c:
#include <math.h>
#include <stdlib.h>
typedef struct { double x, y; } point;
double distance(point a, point b) {
return hypot(a.x - b.x, a.y - b.y);
}
point midpoint(const point *a, const point *b) {
point m = { (a->x + b->x) / 2, (a->y + b->y) / 2 };
return m;
}
void scale_all(point *points, size_t n, double k) {
for (size_t i = 0; i < n; i++) {
points[i].x *= k;
points[i].y *= k;
}
}
and build it as a shared library with the C compiler:
cc -shared -fPIC geometry.c -o libgeometry.so -lm # Linux
cc -dynamiclib geometry.c -o libgeometry.dylib # macOS
cl /LD geometry.c # Windows
The Rust sections use a crate built with crate-type = ["cdylib"], shown
where it is used.
Introduction
Every language that grows up eventually needs code it did not write: a
database engine, a codec, a cryptography library, a scientific kernel, a
system call the standard library has no wrapper for. The code exists, it
is fast and it is tested, and it speaks the C calling convention, which
is the one convention every language on every platform agrees on. Rust
libraries speak it too, through extern "C".
Calling across that boundary is a matter of describing, exactly, what the other side expects: how wide each integer is, where each member of a struct sits, which register a value travels in. The machine does not check any of it. A description that is off by one byte corrupts memory silently, and the bug shows up somewhere unrelated, later.
ffi makes that description a value the program can inspect, builds it
from the source the library’s authors already wrote, lays out every type
the way the platform’s compiler does, and checks every value against it
at the moment it crosses:
import ffi
var c = ffi.open(ffi.LIBC).declare('
int abs(int n);
typedef struct { int quot; int rem; } div_t;
div_t div(int numerator, int denominator);
')
echo c.abs(-7)
echo c.div(17, 5)
7
{quot: 3, rem: 2}
The declarations are the ones <stdlib.h> contains. The struct came back
as a dictionary. And the conversion rules held on the way in:
c.abs(3000000000)
RangeError: abs() argument 1 'n': 3000000000 is out of range for 'int', which holds -2147483648 to 2147483647
Loading a Library
ffi.open() loads a shared library and returns a Library. It takes a
path, a file name, or a bare name the platform’s conventions complete:
var sqlite = ffi.open('sqlite3') # libsqlite3.so, libsqlite3.dylib, sqlite3.dll
var local = ffi.open('./build/libgeometry.so') # a path, as given
var vendored = ffi.open('geometry', { paths: ['vendor/lib'] })
A bare name is tried in the directories given as paths first, then in
the platform’s own search order. On Linux, a library’s unversioned
libname.so is usually only installed with its development package, so
ffi.open() also asks the dynamic loader’s cache for the versioned file
the runtime package ships, such as libsqlite3.so.0. On macOS the name is
tried as a .dylib and as a framework.
ffi.find() answers the same question without loading anything:
import ffi
echo ffi.find('zuri_has_no_such_library')
nil
With no name at all, ffi.open() returns the running process, whose
symbols include the C runtime and everything already loaded globally.
Two options change how a library is loaded. lazy: true resolves its
symbols as they are first used instead of all at once. global: true
makes its symbols visible to libraries loaded after it and to the process
handle, which a plugin that expects its host’s symbols needs. Both are
Unix concepts and have no effect on Windows, where every loaded module is
searched when the process is asked for a symbol.
A library that cannot be found or loaded raises LoadError, carrying the
loader’s own explanation: a missing dependency, a library for another
architecture, a file that is not a library.
Symbols
A Library answers whether it exports a name, and where:
import ffi
var libc = ffi.open(ffi.LIBC)
echo libc.has('strlen')
echo libc.has('zuri_has_no_such_symbol')
echo libc.symbol('zuri_has_no_such_symbol')
true
false
nil
symbol() returns a Pointer to the symbol. On glibc, name@VERSION
asks for a particular version of a versioned symbol, for the rare program
that must pin one.
Closing
close() stops a library being used: every function bound from it
raises LoadError from then on. The code itself is unloaded once nothing
refers to the library any longer, functions bound from it included, so a
closed library is never unloaded out from under a call in progress.
Calling a Function
A function is bound by name, return type and parameter types, and comes back as an ordinary Zuri function:
import ffi
var libm = ffi.open(ffi.LIBM)
var pow = libm.function('pow', ffi.double, [ffi.double, ffi.double])
echo pow(2, 10)
echo pow.name()
echo pow.arity()
1024
pow
2
It can be stored, passed to map(), spawned onto an isolate, and called
like any other function. The number of arguments is checked like any
function’s:
pow(2)
ArgumentError: 'pow' expects 2 arguments, got 1
ffi.describe() reports what a foreign function is, and
ffi.is_foreign() tells a foreign function from a Zuri one:
import ffi
var abs = ffi.open(ffi.LIBC).function('abs', ffi.int, [ffi.int])
var about = ffi.describe(abs)
echo about.name
echo about.type.name()
echo ffi.is_foreign(abs)
echo ffi.is_foreign(@(x) => x)
abs
int (*)(int)
true
false
Binding one function at a time suits a handful. For more, declarations read a header, as the next sections show.
Types
Every C type is a Type. The module exports the built-in ones:
| Zuri | C | Size |
|---|---|---|
ffi.void | void | |
ffi.bool | bool | 1 |
ffi.char, ffi.schar, ffi.uchar | char, signed char, unsigned char | 1 |
ffi.short, ffi.ushort | short, unsigned short | 2 |
ffi.int, ffi.uint | int, unsigned int | 4 |
ffi.long, ffi.ulong | long, unsigned long | 8, or 4 on Windows |
ffi.longlong, ffi.ulonglong | long long, unsigned long long | 8 |
ffi.int8 … ffi.uint64 | int8_t … uint64_t | 1 to 8 |
ffi.int128, ffi.uint128 | __int128, unsigned __int128 | 16 |
ffi.size_t, ffi.ssize_t, ffi.ptrdiff_t | the same | 8 |
ffi.intptr_t, ffi.uintptr_t | the same | 8 |
ffi.wchar_t, ffi.char16_t, ffi.char32_t | the same | 4, 2 on Windows; 2; 4 |
ffi.float, ffi.double | the same | 4, 8 |
ffi.longdouble | long double | 16, or 8 on Windows and Apple Arm |
ffi.complex_float, ffi.complex_double | float _Complex, double _Complex | 8, 16 |
ffi.ptr | void * | 8 |
ffi.string | const char *, as text | 8 |
ffi.wstring | const wchar_t *, as text | 8 |
and Rust’s names for the same types: ffi.i8 through ffi.u128,
ffi.isize, ffi.usize, ffi.f32, ffi.f64, and ffi.rust_char for
Rust’s char. A Rust name and its C counterpart are the same type:
import ffi
echo ffi.i32.equals(ffi.int32)
echo ffi.int32.equals(ffi.int)
echo ffi.usize.equals(ffi.size_t)
true
true
true
A type knows its size, alignment and kind on this platform:
import ffi
echo ffi.int.size()
echo ffi.double.align()
echo ffi.uint16.kind()
echo ffi.string.kind()
4
8
int
string
Building types
Pointers, arrays, const and function types are built from other types:
import ffi
var names = ffi.pointer(ffi.char).array(4)
var compare = ffi.function_type(ffi.int, [ffi.ptr, ffi.ptr])
echo names.name()
echo names.size()
echo compare.pointer().name()
echo ffi.char.as_const().pointer().name()
char *[4]
32
int (*)(void *, void *)
const char *
ffi.type() reads a C spelling of a built-in type, which is often the
shortest way to write one:
import ffi
echo ffi.type('unsigned long long').equals(ffi.ulonglong)
echo ffi.type('void (*)(int, const char *)').kind()
true
function pointer
Structs, unions and enums are built member by member; Structs, Unions and Arrays and Enums and Constants cover them.
Numbers at the Boundary
A Zuri number is a double, which holds every integer up to 2^53 exactly
and no further. C integer types reach 2^64 and, with __int128, 2^128.
So integers cross the boundary this way:
- Going in, a number must be a whole number that fits the type, and
a bigint may be given for any integer type. A fraction, a string, or a
bool for an integer is a
TypeError; a value outside the type’s range is aRangeError. Nothing is ever truncated or wrapped. - Coming out, an integer is a number when it lies within 2^53 of zero, and a bigint beyond that, so no value is ever rounded.
import ffi
var strtoull = ffi.open(ffi.LIBC).function('strtoull', ffi.uint64,
[ffi.string, ffi.ptr, ffi.int])
echo strtoull('42', nil, 10)
echo strtoull('18446744073709551615', nil, 10)
echo typeof(strtoull('18446744073709551615', nil, 10))
42
18446744073709551615n
bigint
A type that can be either is the program’s to handle:
is_bigint(value), or comparing against a bigint, which compares
correctly with a number too.
Floating-point types take any number. long double is wider than a
double on some platforms, 80-bit extended precision on x86-64 Linux and
macOS and 128-bit quad precision on Arm Linux; going in is exact, and
coming out rounds to the nearest double, as C’s own conversion does.
bool takes and gives a Zuri bool, and a C character type takes a number
or a one-character string:
import ffi
var toupper = ffi.open(ffi.LIBC).function('toupper', ffi.int, [ffi.int])
echo toupper('q'.ord())
echo toupper(97).chr()
81
A
A complex number is a list of two, [real, imaginary], and passing one
by value is available wherever the C compiler has _Complex, which is
everywhere but Windows.
Declaring from C
ffi.declare() reads C declarations and returns a Declarations: a set
of types, functions, variables and constants that is not tied to any
library. Binding it to one produces a namespace:
import ffi
var api = ffi.declare('
typedef struct { long quot; long rem; } ldiv_t;
ldiv_t ldiv(long numerator, long denominator);
long labs(long n);
')
var c = api.bind(ffi.open(ffi.LIBC))
echo c.labs(-12)
echo c.ldiv(100, 7)
echo c.ldiv_t.size()
12
{quot: 14, rem: 2}
16
Library.declare() does both steps at once, and is what most programs
use.
The namespace holds a function for each declared function, a Pointer
for each declared variable, each constant, and each type that has a name
of its own. It is read the way a module is. A tagged type with no typedef,
such as struct stat, is reached through the declarations instead:
api.type('struct stat').
A declared function that the library does not export is a SymbolError
naming every one that is missing. allow_missing: true binds the rest
and leaves those out, for a header describing several versions of a
library.
What is read
Everything a header says about an API:
typedef,struct,unionandenum, including forward declarations, self-referencing records, nested records, anonymous members, bitfields and flexible array members;- function prototypes, variadic ones included, and function pointer types however deeply they nest;
externvariables;_Static_assert, which is evaluated, so a header that checks its own assumptions checks them here too;__attribute__((packed)),aligned(n),ms_abiandsysv_abi,__declspec(align(n)),_Alignas, andasmlabels that give a function a different symbol name; every other attribute is read past;extern "C" { ... }, as headers shared with C++ write it.
A function with a body, as an inline helper in a header has one, is read and not bound, because the library need not export it.
import ffi
var api = ffi.declare('
struct node;
typedef struct node node;
struct node { int value; node *next; };
typedef struct {
struct { double x, y; } origin;
union { int id; char tag[8]; };
} shape;
_Static_assert(sizeof(shape) == 24, "shape is 24 bytes");
')
echo api.type('node').size()
echo api.type('shape').offset_of('tag')
echo api.type('shape').has_field('id')
16
16
true
The preprocessor
Real header text is full of preprocessor lines, so enough of the preprocessor runs for it to read as written:
#defineof a value is expanded wherever the name appears and becomes a constant: numbers in any base with any suffix, floating-point numbers, character constants, strings, adjacent strings joined, and expressions over other constants. A value cast to a pointer type, the way a header spells a sentinel such as((void *) -1), is aPointerof that type at that address, andnilwhen the address is zero. A macro can use a type declared after it, as C allows, since a macro means something only where it is used.#if,#ifdef,#ifndef,#elif,#elseand#endifare evaluated, withdefined(), against the macros the target platform’s compiler predefines:__linux__,__APPLE__,_WIN32,__x86_64__,__aarch64__,__LP64__and the rest.__has_include()and the other__has_queries answer no.#pragma packapplies to the records declared after it, exactly where it is written.#undefremoves a macro, and#errorstops with its message.- Including a C standard header is accepted, because everything it
declares is already known;
<stdio.h>also declaresFILE.
import ffi
var api = ffi.declare('
#include <stdint.h>
#define VERSION_MAJOR 2
#define VERSION_MINOR 7
#define VERSION ((VERSION_MAJOR << 8) | VERSION_MINOR)
#define NAME "geo" "metry"
#if VERSION >= 0x200
typedef struct { uint32_t flags; double scale; } options;
#else
typedef struct { uint32_t flags; } options;
#endif
#pragma pack(push, 1)
typedef struct { uint8_t kind; uint32_t length; } header;
#pragma pack(pop)
')
echo api.constant('VERSION')
echo api.constant('NAME')
echo api.type('options').size()
echo api.type('header').size()
519
geometry
16
5
Two things are refused, each with the line and column of the problem. Including any other file, because declarations are read as given, never fetched: paste in the ones the program needs. And using a function-like macro, which would need the preprocessor’s full expansion rules; define a function-like macro and it is recorded, and only using one is an error.
Text returns
C cannot say who owns a returned pointer, but its convention is clear
enough to follow: a function returning const char * hands back text it
keeps, and one returning char * usually hands over memory. So a
declared function returning const char * returns a string, nil for a
null pointer, and one returning char * returns a Pointer, which the
program reads and releases. The same holds for const wchar_t *.
import ffi
var c = ffi.open(ffi.LIBC).declare('
int setenv(const char *name, const char *value, int overwrite);
const char *getenv(const char *name);
')
c.setenv('ZURI_FFI_DEMO', 'hello', 1)
echo c.getenv('ZURI_FFI_DEMO')
echo c.getenv('ZURI_FFI_NOT_SET')
hello
nil
getenv is declared char *getenv(const char *) in the real header;
writing it as const char * here is the program saying it will not free
the result, which is true.
Growing a set, and sharing one
Sources are added to a set one after another, and each sees what came before. One set can include another, whose types it can then use without owning them, which is how two libraries share one set of common types:
import ffi
var common = ffi.declare('typedef struct { double x, y; } point;')
var shapes = ffi.declarations()
.include(common)
.declare('typedef struct { point from, to; } segment;')
.declare('double length(segment s);')
echo shapes.type('segment').size()
echo shapes.functions().length.signature().params[0].name()
32
segment
types(), functions(), variables() and constants() list what a set
holds, and type() resolves any C spelling against it:
shapes.type('segment *[2]').
Declaring from Rust
A Rust library is reached through the C ABI it exports: extern "C"
functions, and the #[repr(C)] types they take. ffi.declare_rust()
reads that surface as a crate writes it, so its source can be handed over
as it stands. Given this lib.rs, built with crate-type = ["cdylib"]:
#![allow(unused)]
fn main() {
use std::ffi::{CStr, CString, c_char};
#[repr(C)]
pub struct Point {
pub x: f64,
pub y: f64,
}
#[repr(C)]
pub enum Shape {
Circle { radius: f64 },
Rect { w: f64, h: f64 },
}
#[unsafe(no_mangle)]
pub extern "C" fn area(shape: Shape) -> f64 {
match shape {
Shape::Circle { radius } => std::f64::consts::PI * radius * radius,
Shape::Rect { w, h } => w * h,
}
}
#[unsafe(no_mangle)]
pub extern "C" fn midpoint(a: &Point, b: &Point) -> Point {
Point { x: (a.x + b.x) / 2.0, y: (a.y + b.y) / 2.0 }
}
#[unsafe(no_mangle)]
pub unsafe extern "C" fn greet(name: *const c_char) -> *mut c_char {
let name = unsafe { CStr::from_ptr(name) }.to_string_lossy();
CString::new(format!("hello, {name}")).unwrap().into_raw()
}
#[unsafe(no_mangle)]
pub unsafe extern "C" fn free_greeting(text: *mut c_char) {
drop(unsafe { CString::from_raw(text) });
}
}
the Zuri side reads the same file:
import ffi
var source = file('shapes/src/lib.rs').read()
var shapes = ffi.open('shapes', { paths: ['shapes/target/release'] }).declare_rust(source)
echo shapes.area({ variant: 'Rect', w: 2, h: 3 })
echo shapes.midpoint({ x: 0, y: 0 }, { x: 4, y: 2 })
var greeting = shapes.greet('zuri').own(shapes.free_greeting)
echo greeting.read_string()
6
{x: 2, y: 1}
hello, zuri
What is read, and what is skipped:
- functions in
extern "C"blocks (andunsafe extern "C"ones), with#[link_name]for a different symbol; extern "C" fndefinitions marked#[no_mangle]or#[export_name], with their bodies skipped; one with neither is refused, because it is exported under a mangled name nothing can look up;#[repr(C)],#[repr(C, packed)],#[repr(C, align(N))]and#[repr(transparent)]structs, tuple structs and unions;- enums with
#[repr(C)]or an integer#[repr], with and without fields; typealiases,constitems and statics;#[cfg(...)], evaluated for the platform the program is running on;- everything else, from
uselines toimplblocks, macros and private functions, is read past.
A type may be used before it is declared, as Rust allows, and a struct
without #[repr(C)] can only be used through a pointer, because Rust’s
own layout is unspecified. Rust in Depth covers how
each Rust type crosses.
C and Rust declarations go in the same set, and each sees the other’s types, which suits a crate that ships a C header beside its source.
Pointers and Memory
A Pointer is an address with, optionally, the type of what it points
at. Memory comes from ffi.alloc(), zeroed and typed:
import ffi
var numbers = ffi.alloc(ffi.int, 4)
numbers.set(0, 10)
numbers.set(3, 40)
echo numbers.get(0)
echo numbers.to_list(4)
echo numbers.add(3).get()
echo numbers.size()
10
[10, 0, 0, 40]
40
16
get() and set() read and write elements of the pointer’s type, as C’s
ptr[i] does, and add() steps by elements, as C’s ptr + i does.
read() and write() take a type of their own and a byte offset, and work
on any pointer:
import ffi
var header = ffi.alloc_bytes(16)
header.write(ffi.uint32, 3405691582)
header.write(ffi.double, 2.5, 8)
echo header.read(ffi.uint32)
echo header.read(ffi.double, 8)
echo header.read(ffi.uint8)
echo header.cast(ffi.uint16).get(1)
3405691582
2.5
190
51966
Memory is little-endian on every platform the module runs on, which the
third line shows. cast() gives a pointer a type, or takes it away.
Checked access
Memory the module allocated knows its extent, and every access through a pointer into it is checked against that extent before it happens:
import ffi
var numbers = ffi.alloc(ffi.int, 4)
catch {
numbers.get(4)
} as e {
echo e.type
echo e.message
}
PointerError
4 bytes at offset 16 fall outside the 16-byte block this pointer belongs to
A write that fails to convert writes nothing, so a record is never left half updated. Freed memory refuses every access, and freeing it twice is an error rather than a double free.
A pointer C handed back carries no extent, because C did not say how much memory is behind it. Only a null pointer is caught through one; how much is there is whatever the library’s documentation says.
Pointers as arguments
A pointer parameter accepts more than a Pointer, each for the duration
of the call:
| Passed | Becomes |
|---|---|
nil | a null pointer |
a Pointer | its address |
| a string | a NUL-terminated copy, for a pointer to characters |
| bytes | the bytes’ own storage, with nothing copied |
| a list | a temporary array of the target type, written back after the call |
| a dictionary | a temporary record, written back after the call |
| a Zuri function | a callback, for a function pointer |
a Callback | its code |
| a foreign function | its address |
A list or dictionary is written back only when the pointer is not to a
const type, so it works as an out-parameter:
import ffi
var frexp = ffi.open(ffi.LIBM).function('frexp', ffi.double,
[ffi.double, ffi.pointer(ffi.int)])
var exponent = [0]
echo frexp(48, exponent)
echo exponent[0]
0.75
6
Bytes pass their own storage, so a C function that fills a buffer fills the bytes directly:
import ffi
var memset = ffi.open(ffi.LIBC).function('memset', ffi.ptr,
[ffi.ptr, ffi.int, ffi.size_t])
var buffer = bytes(6)
memset(buffer, 65, 4)
echo buffer
(41 41 41 41 00 00)
A bytes value passed to a call must not be resized by a callback while the call is running; its storage is only guaranteed to stay where it is for as long as nothing changes its length.
Other allocations
ffi.alloc_bytes(size, align) allocates untyped memory. ffi.malloc()
allocates from the C allocator, for memory that will be handed to C code
which frees it with free(); it is never freed when the pointer is
collected. ffi.at(address, type) makes a pointer from a known address.
copy_from(), fill() and compare() are memmove, memset and
memcmp over checked memory.
Strings and Text
A string passed where a pointer to characters is expected becomes a
NUL-terminated copy for the duration of the call, in the encoding the
character type implies: UTF-8 for char, UTF-16 for char16_t, UTF-32
for char32_t, and wchar_t’s own, which is UTF-16 on Windows and
UTF-32 elsewhere.
ffi.string and ffi.wstring are pointer types that also convert back:
a function declared to return one returns a Zuri string, and nil for a
null pointer.
import ffi
var libc = ffi.open(ffi.LIBC)
var strstr = libc.function('strstr', ffi.string, [ffi.string, ffi.string])
echo strstr('needle in a haystack', 'hay')
echo strstr('needle in a haystack', 'pin')
haystack
nil
Text that has to outlive a call, because it is stored in a struct or kept
by the library, goes in memory of its own with ffi.alloc_string():
import ffi
var text = ffi.alloc_string('naïve')
echo text.read_string()
echo text.read_bytes(6)
echo ffi.alloc_string('naïve', 'utf-16').read_string(nil, 'utf-16')
naïve
(6e 61 c3 af 76 65)
naïve
read_string() reads up to the terminator, or exactly a given number of
code units, in any of the four encodings, and raises ValueError for
text that is not valid in its encoding. write_string() writes text and
its terminator. When a read has no terminator inside a block of known
size, it is refused rather than run past the end.
Structs, Unions and Arrays
A record is built member by member, and laid out the way the platform’s C compiler lays it out:
import ffi
var Header = ffi.struct('Header')
.add_field('tag', ffi.uint8)
.add_field('length', ffi.uint32)
.add_field('checksum', ffi.uint64)
echo Header.size()
echo Header.offset_of('length')
echo Header.offset_of('checksum')
16
4
8
It is sealed the first time anything needs its size, including passing a
value of it, and cannot change after that. set_packed() packs it as
#pragma pack does, set_align() raises its alignment, and add_field()
takes a minimum alignment for one member, as _Alignas gives:
import ffi
var Packed = ffi.struct('Packed').add_field('tag', ffi.uint8).add_field('length', ffi.uint32).set_packed()
var Wide = ffi.struct('Wide').add_field('tag', ffi.uint8).set_align(64)
echo Packed.size()
echo Wide.size()
5
64
Values
A record value is a dictionary of its members, going in and coming out. A
member left out of a dictionary going in is zero. Nested records are
nested dictionaries, and array members are lists, except that an array of
char reads as the string in it and an array of other bytes reads as
bytes:
import ffi
var Item = ffi.struct('Item')
.add_field('name', ffi.char.array(8))
.add_field('counts', ffi.int.array(3))
var item = ffi.alloc(Item)
item.set(0, { name: 'widget', counts: [1, 2] })
echo item.get()
echo item.get_field('counts')
{name: widget, counts: [1, 2, 0]}
[1, 2, 0]
Through a pointer, get_field() and set_field() read and write one
member, as ptr->name does, and field() points at one, as &ptr->name
does. fields() lists the members with their offsets.
Bitfields
add_bitfield() adds a member of a given width. Bitfields follow the
platform’s rules for packing them, GCC’s and Clang’s on Linux and macOS
and MSVC’s on Windows, including a zero-width bitfield ending a storage
unit:
import ffi
var Flags = ffi.struct('Flags')
.add_bitfield('ready', ffi.uint, 1)
.add_bitfield('mode', ffi.uint, 3)
.add_bitfield('delta', ffi.int, 4)
var flags = ffi.alloc(Flags)
flags.set_field('mode', 5)
flags.set_field('delta', -3)
echo flags.get()
echo Flags.fields()[2].bit_offset
catch {
flags.set_field('mode', 8)
} as e {
echo e.message
}
{ready: 0, mode: 5, delta: -3}
4
8 does not fit in the 3-bit field 'mode', which holds 0 to 7
Unions
A union value going in is a dictionary holding one member, the one to store. Coming out, it holds every member, each read from the same bytes, because only the program knows which one is meaningful:
import ffi
var Bits = ffi.union('Bits').add_field('f', ffi.float).add_field('u', ffi.uint32)
var value = ffi.alloc(Bits)
value.set(0, { f: 1 })
echo value.get()
{f: 1, u: 1065353216}
A union read back and passed in again is accepted as it is, because its
members agree; members that disagree are a TypeError.
By value
Records pass to and return from functions by value, whatever their shape. Which registers a record travels in depends on its size and on the types of its members, differently on every platform, and those rules are applied to the record’s real layout, packed and over-aligned records and bitfields included:
var geometry = ffi.open('geometry').declare('
typedef struct { double x, y; } point;
double distance(point a, point b);
point midpoint(const point *a, const point *b);
void scale_all(point *points, size_t n, double k);
')
echo geometry.distance({ x: 0, y: 0 }, { x: 3, y: 4 })
echo geometry.midpoint({ x: 0, y: 0 }, { x: 4, y: 2 })
var points = [{ x: 1, y: 1 }, { x: 2, y: 3 }]
geometry.scale_all(points, 2, 10)
echo points
5
{x: 2, y: 1}
[{x: 10, y: 10}, {x: 20, y: 30}]
midpoint takes pointers to records, and the dictionaries went in as
temporary records. scale_all takes a pointer to an array of them, and a
list of dictionaries went in as a temporary array and was written back.
Enums and Constants
An enum is an integer type with named constants. A value of one is a number, and anywhere one goes in, a constant’s name may go instead:
import ffi
var Level = ffi.enum('Level')
.add_constant('LOW', 1)
.add_constant('HIGH', 10)
echo Level.constants()
echo Level.value('HIGH')
echo Level.size()
var level = ffi.alloc(Level)
level.set(0, 'HIGH')
echo level.get()
{LOW: 1, HIGH: 10}
10
4
10
Declared enums come with their constants, and a declared enum is stored
as int unless its values need more, or it says otherwise
(enum x : uint8_t), or GCC’s packed attribute asks for the smallest
type. Every enum constant, #define constant and Rust const is a
constant of the declarations, and a member of the namespace they bind to.
Callbacks
A Zuri function passed where a function pointer is expected becomes a C function pointer for the length of that call:
import ffi
var c = ffi.open(ffi.LIBC).declare('
void qsort(void *base, size_t count, size_t size,
int (*compare)(const void *, const void *));
')
var numbers = ffi.alloc(ffi.int, 5)
for i, n in [42, 7, 19, 3, 25] {
numbers.set(i, n)
}
c.qsort(numbers, 5, 4, @(a, b) {
return a.cast(ffi.int).get() - b.cast(ffi.int).get()
})
echo numbers.to_list(5)
[3, 7, 19, 25, 42]
The arguments arrive converted from their C types, pointers as
Pointers, records as dictionaries, and the return value is converted to
the callback’s return type, a small integer widened to a full register the
way a C compiler returns one.
Lasting callbacks
A library that keeps a function pointer and calls it later, a signal
handler, an event callback, a logging hook, needs a callback that
outlives the call that handed it over. ffi.callback() makes one:
var on_event = ffi.callback(@(code) {
echo 'event ${code}'
}, ffi.function_type(ffi.void, [ffi.int]))
library.set_handler(on_event)
A lasting callback lives until release(). Nothing else frees it,
because nothing can know when the library has finished with the pointer;
a callback that is never released lives as long as the program. After
release() it can no longer be passed, and C must not call it again.
Errors inside a callback
A Zuri error cannot unwind through C frames; the C code in between
expects to finish. So an error raised inside a callback is trapped, C gets
a zero back (or the callback’s error_value), and the error is raised
again, as it was, the moment the C function returns:
import ffi
var c = ffi.open(ffi.LIBC).declare('
void qsort(void *base, size_t count, size_t size,
int (*compare)(const void *, const void *));
')
var numbers = ffi.alloc(ffi.int, 3)
catch {
c.qsort(numbers, 3, 4, @(a, b) {
raise ValueError('comparison refused')
})
} as e {
echo '${e.type}: ${e.message}'
}
ValueError: comparison refused
The first error is the one raised; later calls into the callback during the same C call see the error value.
Callbacks and Threads
C libraries run threads of their own, and those threads call callbacks.
A Zuri isolate runs on one thread, so a call from any other thread is
posted to the isolate that made the callback, and the calling thread
waits until the isolate has run it. The isolate answers at its next
safepoint, the same points where it checks for signals, or at once when
it is inside a foreign call or in ffi.serve().
That covers a library whose threads call back while the program does
something else. It does not cover a function that blocks the isolate
until its own threads have finished calling back: the isolate would be
waiting for the function, the function for its threads, and the threads
for the isolate. ffi.threaded() gives such a function a variant that
runs on a helper thread while the isolate keeps answering:
import ffi
var c = ffi.open(ffi.LIBC).declare('
typedef unsigned long pthread_t;
int pthread_create(pthread_t *thread, const void *attributes,
void *(*start)(void *), void *argument);
int pthread_join(pthread_t thread, void **result);
')
var seen = []
var start = ffi.callback(@(argument) {
seen.append(argument.address())
return nil
}, ffi.function_type(ffi.ptr, [ffi.ptr]))
var thread = ffi.alloc(ffi.ulong)
c.pthread_create(thread, nil, start, ffi.at(42))
var join = ffi.threaded(c.pthread_join)
join(thread.get(), nil)
echo seen
start.release()
[42]
The callback ran on the isolate’s own thread, in the middle of the
threaded call. A program with nothing else to do while it waits for calls
from another thread can wait in ffi.serve(timeout), which answers
whatever is posted and returns how many it answered.
Variadic Functions
A variadic function, printf and its family, takes its fixed parameters
by type and any number after them. Each argument past the fixed ones
travels as the type its value suggests:
| Value | Travels as |
|---|---|
a whole number that fits an int | int |
| a larger whole number | long long |
| any other number | double |
| a bigint | long long, or unsigned long long when it needs to be |
| a bool | int |
| a string | const char * |
a pointer, bytes or nil | void * |
Anything else is given its type with Type.of(), and C’s promotions still
apply on top: a float travels as a double, and a type narrower than
int as an int.
import ffi
var snprintf = ffi.open(ffi.LIBC).function('snprintf', ffi.int,
[ffi.pointer(ffi.char), ffi.size_t, ffi.string], { variadic: true })
var buffer = ffi.alloc(ffi.char, 64)
snprintf(buffer, 64, '%s has %d items at %.2f each', 'cart', 3, 4.5)
echo buffer.read_string()
snprintf(buffer, 64, '%.1f, %ld, %c', ffi.double.of(2), ffi.long.of(-7), ffi.char.of('z'))
echo buffer.read_string()
cart has 3 items at 4.50 each
2.0, -7, z
The second call shows why of() exists: %.1f with a bare 2 would pass
an int, and printf would read a double that was never there.
Errors and errno
The module’s own errors all descend from FfiError:
| Class | Raised when |
|---|---|
FfiError | a type cannot be used as asked: a record with no size passed by value, a signature libffi cannot call |
LoadError | a library cannot be found or loaded, or a function is called after its library was closed |
SymbolError | a library does not export a symbol |
DeclarationError | C or Rust source cannot be read; line and column point at the problem |
PointerError | memory would be accessed out of bounds, through null, after being freed, or freed twice |
CallbackError | a released callback is used |
LinkError | a static library cannot be linked |
A value that does not convert to its C type raises the prelude’s
TypeError or RangeError, the same errors any function raises for a bad
argument, and the message names the function, the argument and the type.
errno
C reports failure through errno, which anything that runs afterwards may
change. So errno is cleared immediately before every foreign call and
read immediately after it, and ffi.errno() reports what the most recent
call on this isolate left there:
import ffi
var strtol = ffi.open(ffi.LIBC).function('strtol', ffi.long,
[ffi.string, ffi.ptr, ffi.int])
echo strtol('123', nil, 10)
echo ffi.errno()
strtol('99999999999999999999999', nil, 10)
echo ffi.errno()
123
0
34
34 is ERANGE. On Windows, ffi.last_error() reports GetLastError() the
same way; it is always zero elsewhere. ffi.set_errno() sets errno for
the rare function that reads it.
Ownership and Lifetimes
Three kinds of memory cross the boundary, and each has one owner.
Memory from ffi.alloc() belongs to the program. It is freed when the
last pointer into it is collected, or earlier with free(). Pointers made
from it with add(), offset(), cast() and field() share it and keep
it alive. The one rule is that a pointer must stay reachable for as long
as C holds the address; memory C was given and Zuri forgot is memory freed
under C’s feet.
Memory from ffi.malloc() belongs to whoever frees it. It is never
freed on collection, because the usual reason to allocate it is to hand
it to C code that frees it with free(). free() releases it otherwise.
Memory C allocated belongs to C until the program takes it over with
own(), naming the function that releases it, or with no function for
memory the C allocator’s free() releases:
import ffi
var c = ffi.open(ffi.LIBC).declare('
char *strdup(const char *text);
void free(void *pointer);
')
var copy = c.strdup('owned by Zuri now').own(c.free)
echo copy.read_string()
echo copy.is_owned()
copy.free()
echo copy.is_freed()
owned by Zuri now
true
true
An owned pointer is released when it is collected, or with free(); every
pointer derived from it refuses access afterwards. A destructor is any
foreign function of one pointer: sqlite3_close, png_destroy,
CString::from_raw wrapped in an extern "C" function. Only a pointer to
the start of an allocation can free it.
The collector runs a destructor in the middle of its own work, where no
Zuri code can run, so a destructor that calls a callback during a
collection gets zero back, or the callback’s error_value, every time. free() runs the
destructor from the program, where callbacks work as they do anywhere
else.
Static Libraries
A static library is object code waiting for a linker; nothing can load it
at run time as it stands. ffi.link() hands it to the platform’s linker,
which links every object in it into a shared library, and loads that:
var geometry = ffi.link('build/libgeometry.a', { libraries: ['m'] })
var rust = ffi.link('shapes/target/release/libshapes.a')
The linker is the C compiler, cc or whatever $CC names, on Linux and
macOS, and MSVC’s link.exe on Windows, found through the Visual Studio
installation. The result is cached under a name derived from the archives’
contents and the options, in ffi.default_link_cache() or the cache
option’s directory, so the linker runs once for a given input and every
later program start loads the cached library.
A Rust staticlib carries the Rust standard library with it, which in
turn needs a handful of system libraries. An archive holding Rust code is
recognised and linked against them without being asked.
On Windows, a static library’s functions are not marked for export, so
the linked DLL exports every symbol the archives define with an
unmangled name, or exactly the ones listed in the exports option.
libraries, search_paths and flags pass further libraries, their
directories and raw arguments to the linker, and linker replaces it.
A failure raises LinkError carrying the linker’s own output.
Isolates
Types, pointers, libraries, foreign functions, callbacks and declarations all cross to another isolate, and none of them is moved: both sides keep a working handle on the same thing. A type is a description, a library a handle the loader shares across threads, and a pointer an address, so memory one isolate writes, another reads:
import ffi
import isolate
def fill(memory, value) {
memory.set(0, value)
return memory.get(0)
}
var shared = ffi.alloc(ffi.int, 1)
echo isolate.spawn(fill, shared, 99).join()
echo shared.get(0)
99
99
Whatever the memory holds is shared without any synchronisation, as it is between C threads. A callback belongs to the isolate that made it: passed elsewhere, its calls still run on that isolate.
Rust in Depth
Every Rust type that has a defined C ABI crosses, and the rest are refused with the reason.
| Rust | Crosses as |
|---|---|
i8 … u128, isize, usize | integers, 128-bit ones by value included |
f32, f64 | numbers |
bool | a bool |
char | a one-character string; checked to be a Unicode scalar value |
*const T, *mut T | a Pointer, nil for null |
&T, &mut T, NonNull<T>, Box<T> | a Pointer; nil is refused going in |
Option<&T>, Option<NonNull<T>>, Option<Box<T>> | a Pointer or nil |
extern "C" fn(...) | a function pointer: a Zuri function or a Callback in, a callable out |
Option<extern "C" fn(...)> | the same, or nil |
NonZeroU32 and the other NonZero types | a number; zero is refused |
Option<NonZeroU32> | a number, or nil for None |
[T; N] | a list |
#[repr(C)] struct | a dictionary |
#[repr(transparent)] struct | its field |
#[repr(C)] or #[repr(u8)] enum without fields | a number, or a variant name going in |
#[repr(C)], #[repr(C, u8)] or #[repr(u8)] enum with fields | a dictionary with a variant key |
MaybeUninit<T>, ManuallyDrop<T>, Cell<T> | as T |
PhantomData<T> | nothing; it takes no space |
An enum with fields is laid out by the rules those representations
define. A value names its variant, and its fields are the rest of the
dictionary; a tuple variant’s fields are '0', '1' and so on, and a
variant without fields may be passed as just its name:
shapes.area({ variant: 'Circle', radius: 2 })
shapes.area({ variant: 'Rect', w: 2, h: 3 })
tokens.value({ variant: 'Number', '0': 42 })
tokens.value('Plus')
A slice crosses as the pointer and the length Rust’s own FFI convention
pairs it into; ffi.slice(type) builds that #[repr(C)] struct:
var Slice = ffi.slice(ffi.i32)
var values = ffi.alloc(ffi.i32, 3)
lib.sum_slice({ ptr: values, len: 3 })
Refused, each with the reason: &str, &[T], String, Vec, and every
other type without a stable layout; trait objects; tuples; generic types
and functions; Option of a type without a niche; a function without
#[no_mangle] or #[export_name]; and the extern "Rust" ABI.
A Rust function declared extern "C" aborts the process if it panics,
which is Rust’s own rule for that ABI, and a panic never reaches Zuri.
extern "C-unwind" functions are called the same way; a panic unwinding
out of one likewise ends the process, because nothing can unwind safely
through a Zuri frame.
Platform Differences
The module runs on 64-bit little-endian platforms, x86-64 and Arm, on Linux, macOS and Windows, and follows each one’s C compiler:
| Linux x86-64 | Linux Arm | macOS x86-64 | macOS Arm | Windows x86-64 | |
|---|---|---|---|---|---|
long | 8 | 8 | 8 | 8 | 4 |
char | signed | unsigned | signed | signed | signed |
wchar_t | 4, signed | 4, unsigned | 4, signed | 4, signed | 2, unsigned |
long double | 80-bit | 128-bit | 80-bit | 64-bit | 64-bit |
_Complex by value | yes | yes | yes | yes | no |
enum past int | wider type | wider type | wider type | wider type | int, truncated |
| bitfield rules | GCC | GCC | Clang | Clang | MSVC |
ffi.platform() reports these for the running platform. A program that
passes records between platforms through files or sockets gets the same
layouts the C compiler on each one produces, which is what the C code
there expects.
Calling conventions follow the platform too: System V on Linux and macOS
x86-64, AAPCS64 on Arm, with Apple’s variations on it, and the Microsoft
x64 convention on Windows. On x86-64, a function type or Library.function()
can ask for abi: 'win64', and on Unix abi: 'sysv64', for a function
compiled for the other convention; declarations read ms_abi and
sysv_abi attributes and Rust’s extern "win64" and extern "sysv64".
128-bit integers cross by value under every one of them, in both
directions. The Microsoft x64 convention passes one by reference and
returns it in a vector register, as rustc, Clang and GCC compile it, so a
callback returning i128 works there as it does everywhere else.
What Is Checked
A foreign call runs code the module cannot see into, so the guarantees stop where C begins. Inside them:
- every value is converted to its declared type, with integers checked against their range and every other kind against its type;
- memory the module allocated is bounds-checked and refuses use after
free(); freeing twice is an error; - a null pointer is caught before it is read or written through;
- a failed write changes nothing;
- a Zuri error inside a callback never unwinds through C, and a Rust panic inside the module never reaches C;
- a callback called from any thread runs on its own isolate.
Beyond them, what C does is C’s. A description that disagrees with the library, a pointer read past what the library says is there, a callback C calls after it was released, bytes resized by a callback while C writes to them: each of these is undefined behaviour in C, and it stays so here. Reading declarations from the library’s own header or crate, rather than writing them by hand, is the surest way to keep the first of these away.
What the Module Refuses
It does not run the C preprocessor in full. Macros that define values and conditional blocks are evaluated; including other files and expanding function-like macros are not, and are refused where they appear.
It does not compile C. Static libraries are linked by the platform’s own linker, which needs a C toolchain on the machine that links them.
It does not guess ownership. A pointer C returns is not freed until
the program says how, with own().
It does not free a callback on its own. A callback lives until
release(), because only the program knows when the library is done
with it.
It does not unwind through foreign frames. Errors are trapped and raised again once C has returned; a panic crossing the boundary ends the process, as it does in Rust.
Module Reference
The standard library reference documents every class and method. The shape of the module:
ffi.open(name, options) | loads a library, or the process |
ffi.find(name, paths) | where a library would load from |
ffi.link(archives, options) | links static libraries and loads the result |
ffi.declare(source), ffi.declare_rust(source), ffi.declarations() | sets of declarations |
ffi.struct(name), ffi.union(name), ffi.enum(name, type) | records and enums, built by hand |
ffi.pointer(type, options), ffi.array(type, length), ffi.function_type(returns, params, options), ffi.slice(type), ffi.type(spelling) | other types |
ffi.alloc(type, count), ffi.alloc_bytes(size), ffi.malloc(size), ffi.alloc_string(text), ffi.at(address, type) | memory |
ffi.function(pointer, type), ffi.threaded(function), ffi.describe(function), ffi.is_foreign(value) | foreign functions |
ffi.callback(function, type, options), ffi.serve(timeout) | callbacks |
ffi.errno(), ffi.set_errno(n), ffi.last_error() | error codes |
ffi.platform(), ffi.default_link_cache() | facts about the platform |
ffi.LIBC, ffi.LIBM | the C runtime and maths libraries |
On a Library: function, variable, symbol, has, declare,
declare_rust, bind, path, close, is_closed.
On a Pointer: get, set, read, write, get_field, set_field,
field, read_string, write_string, read_bytes, write_bytes,
to_list, add, offset, cast, copy_from, fill, compare,
own, free, is_owned, is_freed, address, is_null, type,
size, equals.
On a Type: name, kind, size, align, is_const, target,
length, signature, pointer, array, as_const, of, equals; on
a StructType or UnionType also add_field, add_bitfield,
set_packed, set_align, fields, offset_of, has_field, variants;
on an EnumType, add_constant, constants, value.
On a Declarations: declare, declare_rust, include, type,
constant, constants, types, functions, variables, bind.
On a Callback: pointer, type, release, is_released.
Packages and Nyssa
A package is a Zuri project other projects can use. Zuri installs, publishes and serves them itself: the commands ship with the runtime, and so does Nyssa, the repository they talk to. There is nothing else to install.
$ zuri install http-extra
Resolving dependencies
Installing http-extra 1.5.0
Installing json-schema 1.2.0
+ http-extra 1.5.0
+ json-schema 1.2.0
Installed 2 packages.
import http_extra
This chapter covers the whole of it:
- Projects and Versions: what
project.tomlsays, how packages are named, and how version ranges read. - Installing Packages: adding, updating and removing dependencies, the lockfile, and where packages can come from.
- Publishing Packages: accounts, tokens, and putting a version on a registry.
- Commands From Packages: packages that add
zuricommands, and installing tools for your user. - Bundles and Upgrades: shipping a program to machines without Zuri, and keeping Zuri itself up to date.
- Running Nyssa: hosting a repository for a team, a company, or the public.
The Commands
| Command | What it does |
|---|---|
zuri init | starts a project |
zuri install | adds packages, or installs everything a project declares |
zuri uninstall | removes packages and whatever only they needed |
zuri update | moves packages to newer versions |
zuri restore | installs exactly what the lockfile names |
zuri info | describes the project’s packages, or one on a registry |
zuri search | finds packages on a registry |
zuri account | signs in, and manages tokens |
zuri publish | publishes a version |
zuri yank | stops a version from being chosen |
zuri owner | manages who may publish a package |
zuri clean | frees the space downloads and installs take |
zuri bundle | packages a program with a runtime |
zuri upgrade | replaces this Zuri with a newer release |
zuri serve | runs a Nyssa repository |
Every one of them answers --help, and every one that changes a
project answers --dry-run with exactly what it would do.
What Holds It Together
- A project is a directory with a
project.toml. Every command works on the project around the directory it runs in, found by looking upwards, the same way imports find it. - Packages install into the project. They land in
.zuri/libs, whichimportsearches before the standard library. Two projects on one machine never share, or fight over, an installed package. - One version of each package. Every requirement in the project is satisfied at once or the install stops and says which requirements clash. Nothing is installed twice at two versions.
- The lockfile is the record.
project.lockpins every package to an exact version and the checksum of what was downloaded, andzuri restorereproduces it anywhere. - Nothing half done. An install is staged beside
.zuri/libsand swapped in whole, so a failure or an interrupted command leaves the project as it was.
Projects and Versions
Starting a Project
$ zuri init weather
$ cd weather
$ zuri run
Hello, world!
zuri init writes a project that runs and a test that passes, with
project.toml describing it. zuri init --help lists what it asks and
what it can be told up front.
project.toml
[project]
name = "weather"
version = "0.1.0"
description = "Shows the weather."
authors = ["Ada Lovelace <ada@example.com>"]
license = "MIT"
readme = "README.md"
homepage = "https://example.com/weather"
repository = "https://github.com/example/weather"
keywords = ["weather", "cli"]
zuri = ">=0.1"
[dependencies]
http-extra = "^1.4"
[dev-dependencies]
fixtures = "^0.3"
| Key | Meaning |
|---|---|
name | what the package is called, and what it is imported as |
version | the version publishing sends, a semantic version |
description | one sentence, shown in search results |
authors, license, readme, homepage, repository, keywords | shown on the package’s page; license is an SPDX expression |
zuri | the versions of Zuri the package works with, as a range |
include, exclude | which files publishing sends, as globs |
[dependencies] is what the project needs to run. [dev-dependencies]
is what working on it needs, such as test fixtures, and is never
required of a project that depends on this one. Both are managed by the
package commands, which keep every comment and blank line the file
already has.
Four more sections come up later in the chapter: [registries] in
Installing Packages,
[install] and [hooks] in
Install Scripts, and
[bundle] in Bundles and Upgrades.
Names
A package name is lowercase letters, digits, hyphens and underscores, starting with a letter and ending with a letter or digit, at most 64 characters. It is imported with every hyphen made an underscore:
| Package | Import |
|---|---|
http-extra | import http_extra |
json-schema | import json_schema |
orm_lite | import orm_lite |
Because of that, two names that differ only in hyphens and underscores
are the same package, and a registry refuses the second. A name the
standard library already uses, such as json or http, cannot be a
package, since the package could never be imported past the standard
library module.
Versions
A version is major.minor.patch, as
Semantic Versioning lays it out:
- patch for fixes that change nothing anyone relies on,
- minor for additions that break nothing,
- major for anything that can break code written against the version before.
A pre-release, such as 2.0.0-rc.1, sorts below its release, and
build metadata after a + is ignored when comparing.
Ranges
A dependency says which versions it accepts:
| Range | Accepts |
|---|---|
1.4.2 or ^1.4.2 | >=1.4.2, <2.0.0: anything compatible |
^0.4.2 | >=0.4.2, <0.5.0: below 1.0, a minor change may break |
^0.0.4 | exactly 0.0.4 |
~1.4.2 | >=1.4.2, <1.5.0: patches only |
=1.4.2 | exactly 1.4.2 |
1.4 | >=1.4.0, <2.0.0, a caret range like any bare version |
1.4.*, 1.4.x | >=1.4.0, <1.5.0 |
* | any release |
>=1.2, <1.8 | both at once; a comma or a space joins comparators |
^1.2 || ^2.1 | either |
A bare version is a caret range, so the common case takes compatible
updates without saying so. = pins.
A pre-release is only chosen for a range that names a pre-release of
the same major.minor.patch itself: ^2.0.0-rc.1 accepts
2.0.0-rc.3, and ^1.4 accepts no pre-release at all. Nobody who
asked for releases is handed an unfinished version.
Installing Packages
Adding a Dependency
$ zuri install http-extra
Resolving dependencies
Installing http-extra 1.5.0
Installing json-schema 1.2.0
+ http-extra 1.5.0
+ json-schema 1.2.0
Installed 2 packages.
The newest version is taken and recorded in project.toml as a caret
range, http-extra = "^1.5.0". Name a range to choose otherwise, pass
--exact to record =1.5.0, or --dev to add a development
dependency:
zuri install http-extra@^1.4
zuri install http-extra --exact
zuri install fixtures --dev
Everything the package needs is installed with it, and everything
lands in .zuri/libs under its import name:
weather/
├── project.toml
├── project.lock
└── .zuri/
├── installed.toml
└── libs/
├── http_extra/
└── json_schema/
zuri init ignores .zuri in git, apart from .zuri/cmds, so
installed packages are never committed. project.toml and
project.lock are, and together they reproduce the directory.
The Lockfile
project.lock records the exact version of every package the project
was resolved to, direct or not, where it came from, and the checksum of
what was downloaded:
version = 1
[[package]]
name = "http-extra"
version = "1.5.0"
source = "registry+https://pub.zurilang.org"
checksum = "sha256:c5b2e20f517379696623b2bb9af8b9f5459a81f62f1aaf3273d9eebbbca3a8af"
dependencies = ["json-schema"]
[[package]]
name = "json-schema"
version = "1.2.0"
source = "registry+https://pub.zurilang.org"
checksum = "sha256:c92464109e86305502e8acaca519d1f7cb2f515a3ab12f3aa26bd1ec67e66b55"
dependencies = []
zuri restore installs exactly that, and is what a fresh checkout, a
teammate or a build server runs:
zuri restore
zuri restore --frozen
zuri restore --production
A package whose download does not match its recorded checksum is
refused. --frozen refuses to go on when the lockfile no longer
matches project.toml instead of resolving again, which is what a
build that must install exactly what was reviewed wants.
--production leaves out what only development dependencies need.
zuri install with no package named does the same as zuri restore
after a change to project.toml: it resolves only what changed and
keeps every other package at its locked version.
When Requirements Clash
Every package gets exactly one version, chosen so that every requirement in the project holds at once. When no choice does, the install stops, changes nothing, and explains the clash step by step:
$ zuri install report-kit
Resolving dependencies
install: no set of versions satisfies every requirement:
(1) Because http-extra 1.5.0 depends on json-schema ^1.2 and http-extra 1.4.2 depends on json-schema ^1.2, http-extra requires json-schema ^1.2.
(2) Because report-kit 1.0.0 depends on json-schema ^2.0 and http-extra requires json-schema ^1.2 (1), report-kit 1.0.0 cannot be used with http-extra.
(3) Because report-kit 1.0.0 cannot be used with http-extra (2) and weather depends on http-extra 1.4, report-kit 1.0.0 cannot be used.
Because report-kit 1.0.0 cannot be used (3) and weather depends on report-kit *, version solving failed.
Read from the bottom: report-kit needs json-schema 2, http-extra
needs json-schema 1, and a project cannot have both. The way out is
a newer http-extra that accepts json-schema 2, when there is one,
or doing without one of the two.
Updating
$ zuri update --dry-run
Resolving dependencies
package current wanted latest
json-schema 1.2.0 1.2.0 2.0.0 (breaking)
current is what is installed, wanted the newest the declared range
allows, and latest the newest there is, marked when moving to it can
break code written against the current one.
zuri update # everything, within its range
zuri update http-extra # one package; the rest stay locked
zuri update json-schema --latest # move the range itself
zuri update --check # exit 1 when anything is out of date
--latest rewrites the range in project.toml to the newest release,
keeping an exact range exact, and says so for every move to a new major
version.
Removing
$ zuri uninstall http-extra
Resolving dependencies
- http-extra 1.5.0
Removed 1 package.
Whatever was installed only because the package needed it goes too. A
package something else still needs stays, and zuri uninstall says
what needs it.
Looking Around
$ zuri info
weather 0.1.0
Shows the weather
package declared locked installed
http-extra 1.4 1.5.0 1.5.0
json-schema - 1.2.0 1.2.0
$ zuri info --tree
weather
└── http-extra 1.5.0
└── json-schema 1.2.0
A package whose declared range, locked version and installed version do
not agree is marked: missing, drifted, unlocked, extraneous,
or, with --verify, modified when its files changed since it was
installed.
The registry answers questions too:
$ zuri search json
json-schema 2.0.0 Validates data against JSON Schema drafts 4 to 2020-12.
report-kit 1.0.0 Builds reports from JSON data.
2 packages on https://pub.zurilang.org, page 1 of 1
$ zuri info http-extra
http-extra 1.5.0
Helpers for building HTTP services: routing, sessions and rate limits.
license MIT
owners ada
downloads 0
depends on json-schema ^1.2
versions: 1.5.0, 1.4.2, 1.0.0 (yanked)
Git and Local Packages
A package does not have to be on a registry:
zuri install --tag v1.2.0 --git https://example.com/tools.git
zuri install --branch main --git git@example.com:acme/tools.git
zuri install --path ../shared
[dependencies]
tools = { git = "https://example.com/tools.git", tag = "v1.2.0" }
shared = { path = "../shared" }
The package’s name comes from its own project.toml. A git dependency
is locked to the commit its tag, branch or revision pointed at and to
the checksum of what that commit packs to, so zuri restore installs
the same files long after the branch has moved. zuri update follows a
branch to its newest commit. A path dependency is copied afresh each
time the project is installed, which suits a package being worked on
beside the project that uses it.
A published package can only depend on registry packages, since a git or path dependency means nothing on the machine of whoever installs it.
Other Registries
Packages come from the default registry, https://pub.zurilang.org,
unless something says otherwise:
zuri install internal-tools --registry https://packages.example.com
The dependency records the registry it came from, so everyone who installs the project gets it from the same place. A project that uses a registry often names it once:
[registries]
company = "https://packages.example.com"
[dependencies]
internal-tools = { version = "^2", registry = "company" }
Aliases can also live in $ZURI_HOME/config.toml for your own use in
every project. default there, or ZURI_REGISTRY in the environment,
changes the registry used when nothing names one; a project that names
its own default keeps it.
A package never falls back from one registry to another. A package asked for from one registry is only ever installed from that registry, so a package with the same name elsewhere can never stand in for it.
Install Scripts
A package may name scripts to run around its installation:
[hooks]
post-install = "scripts/setup.zu"
pre-uninstall = "scripts/teardown.zu"
A script can do anything the person running zuri install can, so a
dependency’s scripts run only when the project allows that package by
name:
[install]
allow-hooks = ["native-sqlite"]
The project’s own scripts always run. --allow-hooks allows more for
one run, and --no-hooks runs none. A script runs with zuri run,
from the package’s directory, with no shell in between, and must finish
within ten minutes with status 0; its output goes to a log under
$ZURI_HOME/logs, which a failure names. ZURI_PACKAGE_NAME,
ZURI_PACKAGE_VERSION, ZURI_PACKAGE_DIR and ZURI_PROJECT_DIR tell
it where it is. A failing script undoes the whole install.
Offline, Caches and Space
Downloads are kept in a cache, checked by checksum, and shared by every
project on the machine. --offline installs from the cache alone, and
fails naming what is missing rather than reaching the network.
zuri clean # the cache and the script logs
zuri clean --libs # this project's installed packages
zuri clean --older-than 30d
Everything zuri clean removes comes back by itself, downloaded again
or restored from the lockfile, the next time it is needed.
The Environment
| Variable | What it does |
|---|---|
ZURI_HOME | where your packages, settings and tokens live; ~/.zuri by default |
ZURI_CACHE | where downloads are kept; $ZURI_HOME/cache by default |
ZURI_REGISTRY | the default registry, by alias or address |
ZURI_TOKEN, ZURI_TOKEN_<ALIAS> | a registry token, before the saved one |
HTTPS_PROXY, HTTP_PROXY, NO_PROXY | the proxy every download goes through |
NO_COLOR | plain output |
Publishing Packages
An Account
Publishing needs an account on the registry. Installing needs none.
zuri account create
That asks for a username, an email address and a password, creates the account, and signs the command line in. It also prints a recovery key, once: the only way back into the account if the password is lost. An account can be made on the registry’s website just as well, and signed in to afterwards:
$ zuri account login
$ zuri account whoami
ada on https://pub.zurilang.org
Signing in exchanges the password for a token, and the token is what
every later command sends. The password is never stored. Tokens are
kept one per registry in $ZURI_HOME/credentials.toml, readable by you
alone.
Tokens
$ zuri account token ci --scopes publish --days 90
$ zuri account tokens
id name scopes expires
74fed0927f53cc20 ci publish 2026-10-28
544af66ae6171536 zuri account create publish, yank, owners, account 2027-09-28
$ zuri account revoke 74fed0927f53cc20
A token carries scopes that bound what it can do:
| Scope | Allows |
|---|---|
publish | publishing new versions |
yank | yanking and restoring versions |
owners | adding and removing a package’s owners |
account | issuing and revoking tokens |
A token is shown once, when it is issued, and the registry keeps only a
digest of it. It lasts 365 days unless --days says otherwise, and
zuri account logout revokes the saved one and forgets it.
A build server signs in with a token rather than a password, read from standard input so it never lands in the shell history or the process list, or taken from the environment:
echo "$PUBLISH_TOKEN" | zuri account login --token-stdin
ZURI_TOKEN="$PUBLISH_TOKEN" zuri publish
Give it a token with the publish scope alone, and nothing it leaks
can yank, change owners or issue more tokens.
What Goes In
A package is the project’s files, packed. Inside a git repository that
is what git tracks or would track; outside one, every file below the
project. .git and .zuri never go. include and exclude in
[project] narrow or widen that with globs, * within one directory
and ** across any number:
[project]
exclude = ["docs/drafts/**", "*.log"]
Files that look like they hold secrets, such as .env, private keys
and credential files, are left out, and the publish stops and names
them. A file that is meant to go is named in include.
See exactly what would be sent first. This project leaves its tests
out with exclude = ["tests"]:
$ zuri publish --dry-run
Would publish weather 0.1.0
53 B .gitattributes
217 B .gitignore
461 B README.md
541 B app/index.zu
376 B index.zu
206 B project.toml
6 files, 1.8 KiB unpacked, 1.6 KiB packed
sha256:0a6b340e81682bde7e4364e2209333ba32c1f3852ff10558bf513550e2758b30
The archive is built the same way every time: sorted entries, fixed times and owners, normalised permissions. The same files always make the same bytes and the same checksum, on any machine.
Publishing a Version
zuri publish
Before anything is sent, the project must have a valid name and
version, git must have no uncommitted changes (--allow-dirty goes
ahead anyway), and the package must hold at most 20,000 files that
unpack to at most 256 MiB. A license that does not read as an SPDX
expression is a warning. The registry then checks the archive again
for itself, and refuses anything unsafe to unpack.
A published version never changes. Publishing the same files again is reported as already published; publishing different files under a version that exists is refused. To fix a version, publish the next one.
The first account to publish a name owns it.
Yanking
zuri yank http-extra@1.4.1
zuri yank http-extra@1.4.1 --undo
A yanked version stays on the registry, and every project whose lockfile already names it keeps installing it, so yanking never breaks a build. It is only left out when versions are chosen anew. Yank a version with a serious bug; do not yank it to hide that it existed.
Owners
zuri owner list http-extra
zuri owner add http-extra grace
zuri owner remove http-extra ada
Every owner may publish, yank and change the owners. A package always keeps at least one.
Another Registry
publish, yank, owner and account work with the default registry
unless --registry names another, by address or by alias:
zuri account login --registry company
zuri publish --registry company
A dependency from another registry must say which in project.toml,
so that whoever installs the package finds it. A dependency naming no
registry comes from the registry the package itself is published on.
Commands From Packages
A package can add commands to zuri, the way the runtime’s own are
written: a cmds directory in the package, one command per .zu file
or directory, as Appendix I describes.
lint-tools/
├── project.toml
├── index.zu
└── cmds/
└── lint/
└── index.zu
Installed into a project, the package’s commands work from anywhere
inside it and show up in zuri --help, marked with the package they
come from:
$ zuri install lint-tools --dev
$ zuri lint
$ zuri --help
...
PACKAGE COMMANDS:
lint Check the project for the mistakes CI rejects. (from lint-tools)
The runtime’s commands come first and a project’s own .zuri/cmds
next, so a package can never replace either. Two installed packages
providing the same command is refused when the second is installed.
Tools for Your User
--global installs into $ZURI_HOME instead of a project, which is
how a tool you use everywhere is installed:
zuri install lint-tools --global
zuri uninstall lint-tools --global
zuri update --global
A globally installed package’s commands work from any directory. Each
one also gets a launcher in $ZURI_HOME/bin, so with that directory on
your PATH, the command runs on its own:
export PATH="$HOME/.zuri/bin:$PATH"
lint
zuri install --global says when the directory is missing from PATH.
Globally installed packages are importable from any script too, after
the project’s packages and the standard library.
Bundles and Upgrades
Bundling a Program
zuri bundle packages a project with a runtime into something that
runs on a machine where Zuri is not installed:
$ zuri bundle --format exe
$ ./dist/weather-0.1.0-x86_64-unknown-linux-gnu
Hello, world!
A bundle holds the runtime renamed after the program, the standard
library it was built with, and the project with the packages it needs
in production. Running it starts the project’s index.zu, whatever
directory it is started from, and its arguments reach the program in
os.args. It uses nothing installed on the machine it runs on, neither
a ZURI_ROOT nor anything in a ZURI_HOME, so it behaves the same
everywhere.
| Format | What it makes |
|---|---|
archive | a .tar.gz of the bundle directory, or a .zip for Windows; the default |
dir | the bundle directory itself |
exe | one file, which unpacks itself into the user’s cache the first time it runs and starts from there afterwards |
app | a macOS application |
Bundles go in dist in the project, named
<name>-<version>-<platform>, unless --output and --name say
otherwise.
What Goes In
The lockfile has to match project.toml, and every package it names,
apart from development dependencies, has to be installed at its locked
version, so a bundle never ships packages nobody resolved. zuri restore puts either right. A project with no dependencies needs no
lockfile at all. The project’s files are chosen the same way
publishing chooses them, so include and exclude apply here too.
Other Platforms
zuri bundle --target aarch64-apple-darwin --target x86_64-pc-windows-msvc
zuri bundle --target all
A bundle for another platform is built with the release of this same
Zuri version for that platform, downloaded once, checked against its
published checksum, and kept in $ZURI_HOME/runtimes. --runtime
names an unpacked runtime to use instead. The platforms are:
x86_64-unknown-linux-gnuaarch64-unknown-linux-gnux86_64-apple-darwinaarch64-apple-darwinx86_64-pc-windows-msvc
A single-file bundle for macOS is built on any machine. Its payload goes inside the executable’s image rather than after it, and the result is signed ad hoc, which is all an Apple silicon Mac needs to run it.
Signing for Distribution
A program downloaded onto a Mac runs only when it is signed with a
Developer ID and notarized, and notarization requires the hardened
runtime. Under the hardened runtime, Zuri’s JIT needs the
com.apple.security.cs.allow-jit entitlement to create executable
memory, so sign with an entitlements file that grants it:
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>com.apple.security.cs.allow-jit</key>
<true/>
</dict>
</plist>
$ codesign --force --options runtime --entitlements entitlements.plist \
--sign "Developer ID Application: Example Ltd" dist/weather-0.1.0-aarch64-apple-darwin
--force replaces the ad hoc signature. Sign an app bundle the same
way, naming the .app directory in place of the executable.
macOS Applications
[bundle]
name = "Weather"
identifier = "com.example.weather"
icon = "assets/weather.icns"
An app bundle needs identifier, in reverse domain form. name is
what Finder shows, and icon an .icns file in the project.
Upgrading Zuri
$ zuri upgrade --check
$ zuri upgrade
--check says whether a newer release exists, and exits 1 when one
does. zuri upgrade downloads the release for this platform, checks it
against its published checksum, unpacks it beside the installation, and
runs it to confirm the version it reports. Only then are the
executable, the standard library and the shipped commands swapped in
together, and if any part of that fails, all of it is put back.
A release with no published checksum is refused, and so is an
installation you cannot write to, before anything is downloaded.
--version picks a release, --prerelease considers pre-releases, and
moving to an older release needs --allow-downgrade. Releases come
from the project’s GitHub releases; ZURI_RELEASES_URL points at
another listing in the same shape, such as a mirror, and GITHUB_TOKEN
is sent when it is set.
Running Nyssa
Nyssa is the package repository, and every Zuri installation can run one: the registry every package command talks to, and a website for finding packages, reading their documentation, and managing an account and its tokens.
$ zuri serve
Nyssa is serving http://127.0.0.1:3000
listening on 127.0.0.1:3000, storage in /home/ada/.zuri/nyssa
That is a working repository, with nothing to set up first. Point the package commands at it:
zuri account create --registry http://127.0.0.1:3000
zuri publish --registry http://127.0.0.1:3000
Settings
Every setting is a flag, a NYSSA_ environment variable, or a line in
nyssa.toml in the storage directory, in that order of precedence:
| Flag | Variable | Default |
|---|---|---|
--host | NYSSA_HOST | 127.0.0.1 |
--port | NYSSA_PORT | 3000 |
--storage | NYSSA_STORAGE | $ZURI_HOME/nyssa |
--config | NYSSA_CONFIG | nyssa.toml in the storage directory |
--public-url | NYSSA_PUBLIC_URL | http://<host>:<port> |
--workers | NYSSA_WORKERS | one per CPU core, up to 8 |
--database | NYSSA_DATABASE | SQLite in the storage directory |
--signup | NYSSA_SIGNUP | open |
--max-archive-size | NYSSA_MAX_ARCHIVE_SIZE | 10485760 bytes |
--trust-proxy | NYSSA_TRUST_PROXY | off |
--read-only | NYSSA_READ_ONLY | off |
--tls-cert, --tls-key | NYSSA_TLS_CERT, NYSSA_TLS_KEY | none |
NYSSA_NAME | Nyssa, the name the site shows | |
NYSSA_MAIL_URL, NYSSA_MAIL_USERNAME, NYSSA_MAIL_PASSWORD, NYSSA_MAIL_FROM | none |
[server]
host = "0.0.0.0"
port = 8080
public_url = "https://packages.example.com"
workers = 4
trust_proxy = true
[database]
url = "postgres://nyssa@db.internal/nyssa"
[registry]
name = "Example Packages"
signup = "closed"
max_archive_size = 20971520
[mail]
url = "smtp://mail.example.com:587"
username = "nyssa"
from = "Example Packages <packages@example.com>"
public_url is the address people reach the repository at, which links
and emails use and which package commands name it by. Passwords belong
in the environment, never in a flag, where they would show in the
process list: NYSSA_DATABASE for a connection string that carries
one, and NYSSA_MAIL_PASSWORD.
The Database
SQLite in the storage directory needs nothing set up, and serves a
repository for a team or a company comfortably. PostgreSQL, MySQL and
MariaDB serve the same site through the same connection strings the
sql module opens:
NYSSA_DATABASE=postgres://nyssa:secret@db.internal/nyssa zuri serve
NYSSA_DATABASE=mysql://nyssa:secret@db.internal:3306/nyssa zuri serve
The schema is kept up to date by numbered migrations, applied when the
repository starts, so starting a newer Zuri against an older database
brings it up to date. zuri serve migrate does the same and exits, for
preparing a database ahead of a deployment. On PostgreSQL and SQLite
each migration runs in a transaction; MySQL and MariaDB commit at every
schema change, so there each migration runs on its own and is recorded
once it has finished.
Published archives are stored under the storage directory by checksum, so the database and that directory are what a backup has to hold.
Serving It Safely
Passwords and tokens cross the network whenever anyone signs in or publishes, so a repository reachable beyond one machine is served over HTTPS, one of two ways:
zuri serve --host 0.0.0.0 --tls-cert fullchain.pem --tls-key privkey.pem
zuri serve --trust-proxy
The second is for a repository behind a proxy that terminates TLS, such as nginx or a load balancer, and makes it believe the client addresses the proxy forwards. Listening beyond this machine with neither prints a warning.
Accounts
With signup = "open", anyone may create an account. A private
repository closes sign-up and creates accounts itself:
$ zuri serve admin create grace
Email: grace@example.com
Password:
Password again:
Created grace. Their recovery key, shown only this once:
| Command | What it does |
|---|---|
zuri serve admin create <username> | creates an account, whether or not sign-up is open |
zuri serve admin promote <username> | makes the account an administrator |
zuri serve admin demote <username> | makes it an ordinary publisher again |
zuri serve admin suspend <username> | suspends it and revokes every token it holds |
zuri serve admin restore <username> | lifts a suspension |
An administrator may yank any version and change the owners of any package, which is how a malicious release is stopped or an abandoned package handed on. Publishing a new version stays with the package’s owners. Every change anyone makes is recorded in the audit log, with who made it and from where.
With mail configured, a new account confirms its email address before it can publish, and a lost password can be reset by email. Without mail, the recovery key an account is given is the way back in.
Looking After It
zuri serve check # every stored archive against its checksum
zuri serve backup nyssa.db # a SQLite database, while it is in use
zuri serve --read-only # browse and install, nothing published
A PostgreSQL or MySQL database is backed up with its own tools, such as
pg_dump or mysqldump. Read-only mode keeps a repository serving
during maintenance or a migration elsewhere.
Running It as a Service
On Linux, systemd keeps it running:
[Unit]
Description=Nyssa package repository
After=network.target
[Service]
User=nyssa
Environment=NYSSA_STORAGE=/var/lib/nyssa
Environment=NYSSA_PUBLIC_URL=https://packages.example.com
EnvironmentFile=/etc/nyssa/secrets.env
ExecStart=/usr/local/bin/zuri serve --host 127.0.0.1 --port 3000 --trust-proxy
Restart=on-failure
[Install]
WantedBy=multi-user.target
secrets.env holds NYSSA_DATABASE and NYSSA_MAIL_PASSWORD,
readable by the service’s user alone. Stopping the service lets every
request already running finish first.
How It Protects Itself
- Passwords are hashed with Argon2id. Tokens and recovery keys are kept only as digests, and every token carries scopes and an expiry.
- Five failed sign-ins lock an account for fifteen minutes, counted in the database so every worker sees them, and each worker allows an address 30 attempts a minute to sign in, sign up or recover.
- Every form carries a token only the site’s own pages have, pages are sent with a content security policy that runs no scripts, and README HTML is sanitised before it is shown.
- An archive is checked before it is stored: its checksum, its size,
every path in it, and that its
project.tomlnames the package and version it was published as. Anything that could not be unpacked safely is refused. - A name that differs from a published one only by case, hyphens or underscores is refused, so no package can pass for another.
- A failure answers the visitor with a plain error page, and the whole of it, with its stack, goes to standard error for whoever runs the repository.
The API
The package commands speak version 1 of a JSON API, which any client
can use. Every failure is { "error": { "code", "message" } } with a
status that describes it, and the code is stable.
| Request | What it does |
|---|---|
GET /api/v1/config | what the repository is and accepts |
GET /api/v1/index/:name | every version of a package, with checksums and dependencies |
GET /api/v1/packages/:name | a package as its page shows it |
GET /api/v1/packages/:name/:version/archive | one archive |
GET /api/v1/search?q= | packages matching a query, with sort, page and per_page |
PUT /api/v1/packages | publishes a version, as a multipart upload |
POST, DELETE /api/v1/packages/:name/:version/yank | yanks a version, and restores it |
GET, PUT /api/v1/packages/:name/owners | the owners, and adding one |
DELETE /api/v1/packages/:name/owners/:username | removes an owner |
POST /api/v1/accounts | creates an account |
POST, GET /api/v1/tokens | signs in or issues a token, and lists tokens |
DELETE /api/v1/tokens/:id | revokes a token, current being the one sent |
GET /api/v1/me | the account behind the token |
GET /healthz | answers 200 while the repository is up |
A request that changes anything sends its token as
Authorization: Bearer nys_....
JSON-RPC
The rpc module is JSON-RPC 2.0, on both ends and over everything it
travels on: as the body of an HTTP request, over a WebSocket, a TCP or
TLS socket, a Unix domain socket, a pipe to a child process, or any
other stream of bytes. It is written in Zuri on top of json, http,
isolate and net, and it holds to the specification exactly.
JSON-RPC is a remote procedure call protocol, and a small one. One program sends the name of a method and its parameters; the other runs it and sends back the result or an error. That is very nearly all of it, and the smallness is the point. The protocol says nothing about what the methods are, how the two programs are connected, or which of them is in charge, so it fits wherever two programs need to call each other. Services call each other with it over HTTP. Blockchain nodes publish their whole API through it, over HTTP for calls and over WebSockets for what they push to subscribers. Mining pools, wallets, trading systems, editors and their tools, and plugin hosts speak it over sockets and pipes.
The module is built in the same spirit. A service answers messages and nothing else, so the same service answers whatever carried the message there; an HTTP route, a WebSocket and a socket server are each a few lines around it.
- Following Along
- Introduction
- Messages
- Services
- JSON-RPC Over HTTP
- Framing
- Transports
- Endpoints
- Reading the Connection
- Serving Many Connections
- Watching the Conversation
- Errors
- What the Module Refuses
- Module Reference
Following Along
A service says which methods it answers, and answers one message at a time:
import rpc
var calculator = rpc.Service()
.on_request('add', @(params) => params[0] + params[1])
echo calculator.answer('{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [2, 3]}')
{"jsonrpc":"2.0","id":1,"result":5}
The message went in as JSON text and the answer came out the same way. Nothing touched a network. Everything else in this chapter is about how the message gets to a service and how its answer gets back: as the body of an HTTP request, over a WebSocket or a socket, or between two isolates of one program.
Introduction
JSON-RPC has three kinds of message, each a small JSON object whose
jsonrpc member is "2.0".
A request names a method, carries its params, and carries an
id. The other side must answer it, and the answer carries the same
id.
A notification is a request without an id. Nothing answers it,
not even when it fails. It is for things the sender wants done but has
no need to hear back about: a log line, a progress report, a change
the other side should know of.
A response carries the id of the request it answers, and exactly
one of result and error.
Parameters are either a list, matched to the method’s parameters by
position, or a dictionary, matched by name. Nothing else is allowed: a
call to square(4) sends [4], never 4.
Messages can also travel together as a batch, a JSON array of
them. The requests in a batch are answered together, as an array of
responses, each matched to its request by id.
The protocol has no transport of its own, and two kinds carry it. In an exchange, such as an HTTP request and its response, one message or batch goes each way and that is the end of it. On a connection, such as a WebSocket or a socket, messages flow both ways for as long as it stays open, and neither side is the client: either may send requests and notifications at any time, including while it is in the middle of answering one. On a stream of bytes, both ends must also agree on where one message stops and the next begins. That agreement is the framing.
The module has a layer for each of these:
| Layer | What it holds |
|---|---|
| messages | Request, Notification, Response, RpcError, and encode(), decode() and read() |
| services | Service, and the Context its handlers are given |
| HTTP | http_handler() to serve a service, and HttpClient to call one |
| framing | HeaderFraming, LineFraming and MessageFraming |
| transports | StdioTransport, SocketTransport, ProcessTransport, WebSocketTransport, ChannelTransport and pipe() |
| endpoints | Endpoint, a service on a connection |
| servers | serve(), an endpoint for every connection to a listening socket |
import rpc reaches all of it.
Messages
Requests and Notifications
A message is built from what it says, and rpc.encode() writes it as
JSON-RPC sends it:
import rpc
var call = rpc.Request('add', [2, 3], 1)
var note = rpc.Notification('log', { level: 'info', text: 'started' })
echo call
echo rpc.encode(call)
echo rpc.encode(note)
Request(add #1)
{"jsonrpc":"2.0","id":1,"method":"add","params":[2,3]}
{"jsonrpc":"2.0","method":"log","params":{"level":"info","text":"started"}}
An id is a string or a number, and is whatever the sender chooses to
match the answer by. An endpoint numbers its own requests from 1.
params of nil leaves the member out of the message, which the
specification allows for a method that takes nothing. Anything other
than a list, a dictionary or nil is refused when the message is
built, before it can be sent:
import rpc
catch {
rpc.Request('square', 4, 1)
} as error {
echo error.message
}
params must be a list, a dictionary or nil, not number
Responses
Response.success() answers a request with its result, and
Response.failure() with an error:
import rpc
var done = rpc.Response.success(1, 5)
var missing = rpc.RpcError(rpc.METHOD_NOT_FOUND, 'Method not found: mul')
var failed = rpc.Response.failure(2, missing)
echo rpc.encode(done)
echo rpc.encode(failed)
echo failed.is_error()
{"jsonrpc":"2.0","id":1,"result":5}
{"jsonrpc":"2.0","id":2,"error":{"code":-32601,"message":"Method not found: mul"}}
true
A response’s id is nil only when it answers a message whose id
could not be read, such as one that was not JSON at all. A result of
nil is a result like any other, and is sent as null.
Errors and Their Codes
An error is an RpcError: a code, a short message, and optional
data carrying whatever else the other side needs to know.
import rpc
var error = rpc.RpcError(rpc.INVALID_PARAMS, 'a name is required', { field: 'name' })
echo error.code
echo error.message
echo error.to_dict()
-32602
a name is required
{code: -32602, message: a name is required, data: {field: name}}
The specification reserves the codes from -32768 to -32000, and
defines these:
| Constant | Code | Meaning |
|---|---|---|
PARSE_ERROR | -32700 | the message is not valid JSON |
INVALID_REQUEST | -32600 | the JSON is not a valid message |
METHOD_NOT_FOUND | -32601 | the method does not exist |
INVALID_PARAMS | -32602 | the method cannot take these parameters |
INTERNAL_ERROR | -32603 | the method failed while it ran |
SERVER_ERROR_MIN to SERVER_ERROR_MAX | -32099 to -32000 | errors an implementation defines for itself |
An application’s own errors take any code outside the reserved range. Pick them once, write them down, and keep them stable: the code is what the other side’s program matches on, and the message is for the person reading the log.
RpcError is an Error, so it is raised and caught like any other.
Raised from a request handler, it is the answer.
Reading and Writing JSON
rpc.decode() reads the JSON text of a message, and checks it against
every rule of the specification:
import rpc
var text = '{"jsonrpc": "2.0", "id": "a7", "method": "subtract", ' +
'"params": {"minuend": 42, "subtrahend": 23}}'
var call = rpc.decode(text)
echo call
echo call.params.minuend - call.params.subtrahend
Request(subtract #a7)
19
What it reads is a Request, a Notification or a Response, told
apart the way the specification tells them apart: a message with a
method and an id is a request, one with a method and no id is a
notification, and one with a result or an error is a response.
Anything that breaks a rule is refused with an RpcError carrying the
code it would be answered with:
import rpc
var attempts = [
'{"jsonrpc": "2.0", "method": 1}',
'{"jsonrpc": "1.0", "id": 1, "method": "ping"}',
'{"jsonrpc": "2.0", "id": 1, "method": "ping", "params": 1}',
'{"jsonrpc": "2.0", "id": 1, "result": 1, "error": null}',
'{"jsonrpc": "2.0", "id": 1',
]
for text in attempts {
catch {
rpc.decode(text)
} as error {
echo '${error.code} ${error.message}'
}
}
-32600 Invalid Request: method must be a string
-32600 Invalid Request: the jsonrpc member must be '2.0'
-32600 Invalid Request: params must be an array or an object
-32600 Invalid Request: a response needs exactly one of result and error
-32700 Parse error: json.decode(): expected ',' or '}' in object
rpc.read() does the same for a value already decoded from JSON, for
a program that got the JSON from somewhere that decodes it already.
rpc.encode() is the other direction, for any message or list of
them.
Batches
A list of messages is encoded as a batch:
import rpc
echo rpc.encode([
rpc.Request('add', [1, 2], 1),
rpc.Notification('log', ['adding']),
])
[{"jsonrpc":"2.0","id":1,"method":"add","params":[1,2]},{"jsonrpc":"2.0","method":"log","params":["adding"]}]
A batch decodes to a list. Each entry in it is read on its own, so one
bad entry does not spoil the rest: in its place is the RpcError
saying what is wrong with it, ready to be answered.
import rpc
var batch = rpc.decode('[' +
'{"jsonrpc": "2.0", "id": 1, "method": "sum", "params": [1, 2]},' +
'{"jsonrpc": "2.0", "method": "notify_hello"},' +
'{"foo": "boo"}' +
']')
for entry in batch {
if instance_of(entry, rpc.RpcError) {
echo 'refused: ${entry.message}'
} else {
echo entry
}
}
Request(sum #1)
Notification(notify_hello)
refused: Invalid Request: the jsonrpc member must be '2.0'
An empty batch, [], is not a message at all, and decoding one raises
INVALID_REQUEST.
Services
A Service holds the handler for every method a program answers.
answer() takes one message, or a batch, as JSON text or as its
bytes, and returns the JSON text of the answer, or nil when there is
nothing to answer.
Answering Requests
on_request() gives a method its handler. The handler is called with
the request’s params and a Context, and what it returns is the
result sent back:
import rpc
var service = rpc.Service()
.on_request('add', @(params) => params[0] + params[1])
.on_request('whoami', @(params, context) => '${context.method} #${context.id}')
echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [2, 3]}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "whoami"}')
echo service.answer('{"jsonrpc": "2.0", "id": 3, "method": "mul", "params": [2, 3]}')
{"jsonrpc":"2.0","id":1,"result":5}
{"jsonrpc":"2.0","id":2,"result":"whoami #2"}
{"jsonrpc":"2.0","id":3,"error":{"code":-32601,"message":"Method not found: mul"}}
The Context names the request’s id and method, the endpoint
handling it, and the request it arrived in, which for a message from
HTTP is the HTTP request. A handler that has no use for the context
leaves it out, as add does. A request for a method with no handler is
answered with METHOD_NOT_FOUND, as mul is.
A handler is replaced by calling on_request() again with the same
method. Method names that start with rpc. are reserved by the
specification, and on_request() refuses them.
A message that is not valid UTF-8, or not valid JSON, is answered with
PARSE_ERROR and an id of nil, since its id cannot be read.
Failing Properly
A handler fails a request by raising. An RpcError is sent back as it
is, code, message, data and all. Any other error is sent back as an
INTERNAL_ERROR carrying the error’s message:
import rpc
var service = rpc.Service()
.on_request('divide', @(params) {
if params.length() != 2 or params[1] == 0 {
raise rpc.RpcError(rpc.INVALID_PARAMS, 'divide takes a number and a non-zero divisor')
}
return params[0] / params[1]
})
.on_request('save', @(params) {
raise Error('the disk is full')
})
echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "divide", "params": [1, 0]}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "save", "params": ["notes"]}')
{"jsonrpc":"2.0","id":1,"error":{"code":-32602,"message":"divide takes a number and a non-zero divisor"}}
{"jsonrpc":"2.0","id":2,"error":{"code":-32603,"message":"the disk is full"}}
An error that is not an RpcError also goes to the service’s
on_error() handler, since it is a
failure in the program rather than a refusal the program meant to
make. A handler that must not let its failures’ messages reach the
other side catches them and raises an RpcError of its own instead.
On the calling side, an error answer raises from request() as the
RpcError the other side sent.
Notifications
on_notification() gives a notification its handler. It is called
the same way, with the params and a Context, and whatever it
returns is ignored, since nothing is sent back:
import rpc
var seen = []
var service = rpc.Service()
.on_notification('log', @(params) {
seen.append(params.text)
})
echo service.answer('{"jsonrpc": "2.0", "method": "log", "params": {"text": "started"}}')
echo seen
nil
[started]
A notification for a method with no handler is dropped, and a
notification handler that raises sends nothing back either: the error
goes to on_error(). That is the specification’s rule, and it is
deliberate. The sender asked not to be answered.
Methods Without a Handler
on_unhandled() sets a fallback for every request and notification
whose method has no handler of its own. It is called like any other
handler, and context.method says which method it is standing in for:
import rpc
var service = rpc.Service()
.on_unhandled(@(params, context) {
if context.method.starts_with('legacy.') {
return 'retired: ${context.method}'
}
raise rpc.RpcError(rpc.METHOD_NOT_FOUND, 'Method not found: ${context.method}')
})
echo service.answer('{"jsonrpc": "2.0", "id": 1, "method": "legacy.export"}')
echo service.answer('{"jsonrpc": "2.0", "id": 2, "method": "export"}')
{"jsonrpc":"2.0","id":1,"result":"retired: legacy.export"}
{"jsonrpc":"2.0","id":2,"error":{"code":-32601,"message":"Method not found: export"}}
It is the place for a family of methods answered the same way, for a
proxy that passes calls on, and for logging what a client asks for
that nothing answers. Raising METHOD_NOT_FOUND from it refuses a
method exactly as a service without a fallback would. Methods whose
names start with rpc. never reach it.
Batches and Their Limit
A batch is answered as a batch. Each request in it gets its answer, in place, and each notification gets none; a batch of nothing but notifications has no answer at all:
import rpc
var service = rpc.Service()
.on_request('add', @(params) => params[0] + params[1])
echo service.answer('[' +
'{"jsonrpc": "2.0", "id": 1, "method": "add", "params": [1, 2]},' +
'{"jsonrpc": "2.0", "method": "add", "params": [3, 4]},' +
'{"jsonrpc": "2.0", "id": 3, "method": "sub"}' +
']')
[{"jsonrpc":"2.0","id":1,"result":3},{"jsonrpc":"2.0","id":3,"error":{"code":-32601,"message":"Method not found: sub"}}]
One message holding a million requests is one message, and a service
that took it whole would do a million calls’ work for it. A batch of
more than batch_limit() messages, 1000 unless set otherwise, is
refused whole, before any of it runs:
import rpc
var service = rpc.Service()
.set_batch_limit(2)
.on_request('ping', @() => 'pong')
var call = '{"jsonrpc": "2.0", "id": 1, "method": "ping"}'
echo service.answer('[${call}, ${call}, ${call}]')
{"jsonrpc":"2.0","id":null,"error":{"code":-32600,"message":"Invalid Request: a batch of 3 messages is over the limit of 2"}}
set_batch_limit(nil) takes a batch of any size, for a service that
only ever hears from programs it trusts.
JSON-RPC Over HTTP
Over HTTP, each message, or batch, is the body of a POST, and its
answer is the body of the response. Every exchange stands on its own:
nothing is framed, and no connection is kept between calls beyond what
HTTP keeps for itself. It is how services call each other, and how
most of the JSON-RPC in the world is spoken.
Serving a Service
rpc.http_handler() turns a service into a route handler for an
http server. Here the server runs on an isolate of its own and the
client calls it:
import rpc
import isolate
def serve_calculator(ready) {
import http
import rpc
var calculator = rpc.Service()
.on_request('add', @(params) => params[0] + params[1])
var server = http.server(0, '127.0.0.1')
server.post('/rpc', rpc.http_handler(calculator))
server.bind()
ready.send(server.socket.local_address().port())
server.listen()
}
var ready = isolate.channel(1)
var server = isolate.spawn(serve_calculator, ready)
var calculator = rpc.HttpClient('http://127.0.0.1:${ready.recv()}/rpc')
echo calculator.request('add', [2, 3])
echo calculator.request_batch([['add', [1, 2]], ['add', [3, 4]]])
server.cancel()
5
[3, 7]
In production the service is built in the setup each http worker
runs, and the server spreads its requests across a pool of isolates:
import http
import rpc
def setup(server) {
var api = rpc.Service()
.on_request('add', @(params) => params[0] + params[1])
server.post('/rpc', rpc.http_handler(api))
}
http.serve(setup, { host: '0.0.0.0', port: 8545 })
Everything http offers a route applies to this one: middleware,
TLS, HTTP/2, compression, body limits, and the rest of
Chapter 15.
A request that arrives over HTTP is answered in the response to that
HTTP request, so its handler answers by returning: context.defer()
raises there, since there is no connection to send a later answer on.
Calling a Service
rpc.HttpClient calls a service at a URL. Each call is one POST:
import rpc
var node = rpc.HttpClient('https://node.example.com', {
headers: { Authorization: 'Bearer ${token}' },
timeout: 10,
})
echo node.request('eth_blockNumber', [])
node.notify('log', ['checked the block number'])
var answers = node.request_batch([
['eth_blockNumber', []],
['eth_gasPrice', []],
])
request() returns the result or raises the RpcError the service
answered with. notify() expects nothing back. request_batch()
sends every call in one POST and returns the answers in the order of
the calls, whatever order the server sent them in, with each failed
call’s RpcError in its place.
The options are headers, sent with every call, timeout, the most
seconds to wait for an answer, and client, the http.HttpClient to
send through, for its proxy, TLS and connection settings. Each call
can also take a timeout of its own.
Statuses
The statuses on the wire follow common practice:
| Status | When |
|---|---|
200 | an answer, whether it is a result or a JSON-RPC error |
204 | nothing to answer: a notification, or a batch of them |
405 | a method other than POST, with an Allow: POST header |
415 | a body that is not application/json in UTF-8 |
A JSON-RPC error never changes the status: METHOD_NOT_FOUND is a
200 whose body says so. The http server’s own limits apply before
the service sees anything, so a body over its max_body_size is a
413. Register the handler with server.any() and every method other
than POST gets its 405; with server.post(), they get the server’s
404.
On the calling side, any other status raises an RpcHttpError
carrying the status and the body. Some servers send a JSON-RPC
error with a status other than 200, as older JSON-RPC over HTTP
drafts asked for; when the body of such a response is a JSON-RPC
answer, HttpClient reads it as one, and the error raises as the
RpcError it is.
Who Is Calling
Every handler gets the HTTP request it came in as context.request,
so it can read headers, the client’s address, or whatever a middleware
left in the request’s context. Authentication belongs in a
middleware, which refuses a request before the service sees it:
def setup(server) {
var api = rpc.Service()
.on_request('balance', @(params, context) {
return accounts.balance(context.request.context.user)
})
server.use(@(request, response, next) {
var user = tokens.user_for(request.bearer_token())
if user == nil {
response.text('who are you?\n', 401)
return
}
request.context.user = user
next()
})
server.post('/rpc', rpc.http_handler(api))
}
Framing
A stream of bytes has no edges. A program reading a socket gets bytes in whatever pieces the network delivers them, and two messages may arrive in one piece, or one message in ten. A framing says where each message ends, so the reader can put them back together.
Both framings work the same way. feed() takes the next bytes, in
whatever pieces they come, and returns every message they complete.
frame() goes the other way, turning a message’s text into the bytes
to send.
Content-Length Headers
HeaderFraming puts a short header block in front of each message,
giving its length in bytes:
import rpc
var framed = rpc.HeaderFraming().frame('{"jsonrpc":"2.0","method":"ping"}')
echo framed.to_string().split('\r\n')
[Content-Length: 33, , {"jsonrpc":"2.0","method":"ping"}]
The header block is ASCII: lines of Name: value, each ended by
\r\n, then a blank line, then exactly as many bytes as
Content-Length says. The length counts bytes rather than characters,
so a message in any language frames correctly.
Reading it back, the pieces can be split anywhere at all:
import rpc
var framing = rpc.HeaderFraming()
var stream = framing.frame('{"a":1}') + framing.frame('{"b":2}')
var first = framing.feed(stream[0, 30])
var second = framing.feed(stream[30, stream.length()])
echo first.map(@(m) => m.to_string())
echo second.map(@(m) => m.to_string())
echo framing.pending()
[{"a":1}]
[{"b":2}]
0
The first piece held all of one message and the start of the next.
feed() handed back the whole one and kept the rest, and pending()
says how many bytes it is keeping. The second piece finished the
second message.
Header names are read without regard to case. Content-Length is
required. Content-Type may be given, and if it names a charset, the
charset must be UTF-8. Any other header is read past and ignored.
One Message to a Line
LineFraming ends each message with a line break instead, the framing
known as newline-delimited JSON:
import rpc
var framing = rpc.LineFraming()
var lines = framing.feed('{"a":1}\n{"b":'.to_bytes())
echo lines.length()
echo framing.feed('2}\r\n'.to_bytes())[0].to_string()
echo framing.frame('{"c":3}').to_string().trim()
1
{"b":2}
{"c":3}
A line may end in \n or \r\n, and an empty line is skipped.
json.encode() never writes a raw line break, since a line break
inside a string is escaped, so a line is always exactly one message.
A peer that pretty-prints its JSON over several lines cannot use this
framing.
One Message to Each Read
MessageFraming frames nothing. It is for a transport that keeps
messages apart itself, a WebSocket above all, whose every read returns
exactly one message: feed() takes each piece as one whole message,
and frame() hands the text back as it is.
import rpc
var framing = rpc.MessageFraming()
echo framing.feed('{"a":1}'.to_bytes())[0].to_string()
echo framing.frame('{"b":2}').to_string()
{"a":1}
{"b":2}
A transport that wants it says so with a framing() method of its
own, and an endpoint over it uses that framing unless told otherwise.
rpc.websocket() does exactly that.
Choosing a Framing
Both ends must use the same framing, so the choice is usually made by
whatever is on the other end. An endpoint uses its transport’s own
framing when it has one, and HeaderFraming otherwise, unless told
something else:
var endpoint = rpc.endpoint(transport).set_framing(rpc.LineFraming())
Where the choice is yours, HeaderFraming is the sturdier of the two:
a reader knows how much is coming before it arrives, and refuses a
message that is too large without reading it.
Limits
Every framing refuses a message larger than max_size(), 64 MiB by
default, and HeaderFraming refuses a header block longer than 8 KiB.
Without them, a peer could make a reader hold any amount of memory by
never finishing a message.
import rpc
var framing = rpc.HeaderFraming().set_max_size(1024)
catch {
framing.feed('Content-Length: 4096\r\n\r\n'.to_bytes())
} as error {
echo error.message
}
a message of 4096 bytes is over the 1024-byte limit
The message is refused as soon as its header arrives, before any of it is read.
A framing that cannot read its stream raises RpcFramingError. Every
reason is the same in one respect: nothing after it in the stream can
be found, because the reader no longer knows where the next message
starts. The connection is over at that point.
Transports
A transport carries the bytes. It is any object with three methods:
read(max)returns the next bytes that arrive, at mostmaxof them, waiting until at least one does. Once the stream has ended, it returns empty bytes.write(data)sends all ofdata.close()ends the stream in the direction it writes.
A transport may also have can_listen(), true when another isolate
can read it while the one that made it writes to it. That is what
listen() needs; without the
method, the answer is false.
Standard Streams
rpc.stdio() talks over the program’s own standard input and output.
It is the transport of a program that another program starts and
talks to: the parent writes to the child’s stdin and reads its stdout.
# calculator.zu
import io
import rpc
rpc.endpoint(rpc.stdio())
.on_request('add', @(params) => params[0] + params[1])
.on_notification('log', @(params) {
io.stderr.write('${params[0]}\n')
})
.serve()
Stdout belongs to the protocol. Anything else the program prints there lands in the middle of the stream and breaks it for the other side, so a program serving over stdio sends everything else to stderr, as this one does with its log. Every write the transport makes is flushed at once.
Sockets
rpc.socket() talks over a connected TcpStream, UnixStream or
TlsStream from net:
import net
import rpc
import isolate
var ready = isolate.channel()
var server = isolate.spawn(@(ready) {
import net
import rpc
var listener = net.TcpStream()
listener.bind('127.0.0.1:0')
ready.send(listener.local_address().to_string())
var connection = listener.accept()
rpc.endpoint(rpc.socket(connection))
.on_request('add', @(params) => params[0] + params[1])
.serve()
connection.close()
listener.close()
}, ready)
var stream = net.TcpStream()
stream.connect(ready.recv())
var client = rpc.endpoint(rpc.socket(stream))
echo client.request('add', [20, 22])
client.close()
server.join()
42
The server binds to port 0, so the system picks a free one, and
sends the address back over a channel for the client to connect to. It
serves the one connection it accepts, and its serve() returns when
the client closes its end.
A socket is held by one isolate at a time, so a socket endpoint reads
with serve() and request() rather than listen().
This server serves one connection and ends. A server for many clients
at once is rpc.serve(), which gives
every connection an endpoint of its own.
Child Processes
rpc.process() talks to a child process over its standard streams.
Spawn it with stdin and stdout both 'pipe':
import os
import rpc
var child = os.spawn('zuri', ['run', 'calculator.zu'], {
stdin: 'pipe',
stdout: 'pipe',
})
var calculator = rpc.endpoint(rpc.process(child))
echo calculator.request('add', [2, 3]) # 5
calculator.notify('log', ['added two numbers'])
calculator.close()
child.wait()
Closing the endpoint closes the child’s stdin. The calculator above
reads that as the end of the connection, its serve() returns, and
the program ends, which is what wait() waits for. A child’s streams
are held by the isolate that spawned it, so this endpoint reads with
serve() and request() too.
WebSockets
rpc.websocket() talks over a WebSocket from http.websocket, either
one a route accepted or one websocket.connect() opened. Each
JSON-RPC message is one WebSocket message, so an endpoint over it uses
a MessageFraming without being told to.
A WebSocket is a connection in both directions, so either side calls the other whenever it likes. That is what makes it the transport of choice for a server that pushes to its clients, such as a node streaming new blocks to whoever subscribed:
import http.websocket
import rpc
import isolate
def serve_shop(ready) {
import http
import http.websocket
import rpc
var server = http.server(0, '127.0.0.1')
server.get('/ws', @(request, response) {
var socket = websocket.accept(request, response)
rpc.endpoint(rpc.websocket(socket))
.on_request('total', @(params, context) {
var price = context.endpoint.request('price_of', [params.item])
return price * params.count
})
.serve()
})
server.bind()
ready.send(server.socket.local_address().port())
server.listen()
}
var ready = isolate.channel(1)
var shop = isolate.spawn(serve_shop, ready)
var socket = websocket.connect('ws://127.0.0.1:${ready.recv()}/ws')
var client = rpc.endpoint(rpc.websocket(socket))
.on_request('price_of', @(params) => params[0] == 'tea' ? 3 : 5)
echo client.request('total', { item: 'tea', count: 4 })
client.close()
shop.cancel()
12
The route accepts the WebSocket and serves an endpoint on it for as
long as the client stays connected. Asked for a total, the server asks
the client for a price before it answers. A WebSocket is held by one
isolate at a time, so an endpoint over one reads with serve() and
request().
Between Isolates
rpc.pipe() returns two transports joined to each other: what one
writes, the other reads. Each end is a ChannelTransport over a pair
of isolate channels, and channels cross isolates, so either end can be
handed to another isolate and used from there.
import rpc
var ends = rpc.pipe()
ends[0].write('ping'.to_bytes())
echo ends[1].read(4096).to_string()
ping
A pipe is the natural way to give one part of a program a JSON-RPC interface to another, and the easiest way to test an endpoint: the server and its test talk exactly as they would over a socket, with nothing listening on the network.
A Transport of Your Own
Anything with the three methods is a transport. This one wraps another and counts what it sends:
import rpc
import isolate
class Counted {
@new(inner) {
self.inner = inner
self.sent = 0
}
read(max) {
return self.inner.read(max)
}
write(data) {
self.sent += data.length()
self.inner.write(data)
}
close() {
self.inner.close()
}
}
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('add', @(params) => params[0] + params[1])
.serve()
}, ends[1])
var counted = Counted(ends[0])
var client = rpc.endpoint(counted)
client.request('add', [1, 2])
client.close()
echo '${counted.sent} bytes sent'
76 bytes sent
Endpoints
An endpoint is a service on a connection. rpc.endpoint(transport)
makes one, and everything in Services holds for it: its
handlers, its fallback, its batch limit and its errors. The connection
adds the other direction. An endpoint calls the other side, and a
handler on one can answer later than it returns.
Calling the Other Side
request() sends a request and waits for its answer:
var sum = client.request('add', [2, 3])
It returns the result, raises the RpcError the other side answered
with, and raises RpcClosedError if the connection ends first. While
it waits, it handles everything else that arrives, exactly as
serve() would: other answers, notifications, and requests from the
other side.
notify() sends a notification, and returns as soon as it is sent:
client.notify('log', { level: 'info', text: 'started' })
send_request() sends a request and returns at once, with the
request’s id. The answer goes to a callback when it arrives:
import rpc
import isolate
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('add', @(params) => params[0] + params[1])
.on_request('fail', @(params) {
raise rpc.RpcError(-32001, 'no')
})
.serve()
}, ends[1])
var client = rpc.endpoint(ends[0])
client.send_request('add', [1, 2], @(result, error) {
echo 'add: ${result}'
})
client.send_request('fail', nil, @(result, error) {
echo 'fail: ${error.message}'
})
echo 'waiting'
echo client.request('add', [3, 4])
client.close()
waiting
add: 3
fail: no
7
A callback is called with the result and nil, or with nil and the
error. Callbacks run on the isolate that owns the endpoint, while it
is reading: here, while request() waited for its own answer, the two
earlier answers arrived first and their callbacks ran. When the
connection ends with a callback still waiting, it is called with an
RpcClosedError.
Sending a Batch
request_batch() sends several requests as one batch and waits for
every answer. Each call is a pair of a method and its params, and the
answers come back in the order of the calls, whatever order the other
side sent them in:
import rpc
import isolate
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('add', @(params) => params[0] + params[1])
.serve()
}, ends[1])
var client = rpc.endpoint(ends[0])
var answers = client.request_batch([
['add', [1, 2]],
['multiply', [3, 4]],
['add', [5, 6]],
])
for answer in answers {
if instance_of(answer, rpc.RpcError) {
echo 'failed: ${answer.message}'
} else {
echo answer
}
}
client.close()
3
failed: Method not found: multiply
11
A batch can partly succeed, so a failed call does not raise: its
place in the list holds the RpcError it was answered with. Only an
ending connection or a timeout raises, since then no answer can be
trusted to come.
Answering Later
A handler normally answers by returning. Sometimes the answer comes
from work still to be done: a job on another isolate, a reply from a
third program, an event that has not happened yet. The handler then
calls context.defer(), keeps the context, and answers through it
with reply() or fail() when the answer is ready:
import rpc
var ends = rpc.pipe()
var waiting = []
var server = rpc.endpoint(ends[0])
.set_framing(rpc.LineFraming())
.on_request('next_job', @(params, context) {
waiting.append(context.defer())
})
.on_notification('add_job', @(params) {
for context in waiting {
context.reply(params)
}
waiting = []
})
server.handle({ jsonrpc: '2.0', id: 1, method: 'next_job' })
server.handle({ jsonrpc: '2.0', id: 2, method: 'next_job' })
server.handle({ jsonrpc: '2.0', method: 'add_job', params: { file: 'cat.png' } })
echo ends[1].read(4096).to_string().trim()
echo ends[1].read(4096).to_string().trim()
{"jsonrpc":"2.0","id":1,"result":{"file":"cat.png"}}
{"jsonrpc":"2.0","id":2,"result":{"file":"cat.png"}}
Once a request is deferred, what its handler returns is ignored. It is
answered exactly once: answering it a second time, or answering a
request that was never deferred, raises ValueError. A handler that
raises after deferring fails the request with that error, unless it
was already answered. A request deferred from inside a batch is
answered on its own, after the batch’s other answers.
Calls in Both Directions
Either side can call the other at any time, including from inside a handler. Here the server, asked for an order’s total, asks the client for a price first:
import rpc
import isolate
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('total', @(params, context) {
var price = context.endpoint.request('price_of', [params.item])
return price * params.count
})
.serve()
}, ends[1])
var prices = { tea: 3, cake: 5 }
var client = rpc.endpoint(ends[0])
.on_request('price_of', @(params) => prices[params[0]])
echo client.request('total', { item: 'tea', count: 4 })
client.close()
12
The client’s request('total') is waiting when the server’s request
for price_of arrives, and it answers that while it waits, then goes
on waiting for its own answer. Nothing has to be arranged for this:
every wait handles what arrives.
Reading the Connection
An endpoint reads its transport in one of two ways, and every handler runs on the isolate that owns the endpoint either way.
Serving
serve() reads and answers messages until the connection ends,
stop() is called, or the endpoint is closed. It is the whole program
for a server that does nothing but answer, as every server so far in
this chapter has been.
A message that is not valid UTF-8, or not valid JSON, is answered with
PARSE_ERROR and an id of nil, since its id cannot be read. A
stream whose framing cannot be read raises RpcFramingError out of
serve().
Listening While Doing Other Work
A program that waits on other things as well cannot sit inside
serve(). listen() reads the transport on an isolate of its own,
decodes each message there, and returns a Channel of what arrives.
The program takes items from it when it is ready, alongside its other
channels, and passes each to dispatch():
import rpc
import isolate
var ends = rpc.pipe()
var caller = isolate.spawn(@(transport) {
import rpc
var client = rpc.endpoint(transport)
var squares = [3, 4, 5].map(@(n) => client.request('square', [n]))
client.close()
return squares
}, ends[0])
# The work runs on a worker of its own, and its results come back on
# `done` whenever they are ready.
var jobs = isolate.channel()
var done = isolate.channel()
isolate.spawn(@(jobs, done) {
var job = jobs.recv()
while job != nil {
done.send({ id: job.id, result: job.value * job.value })
job = jobs.recv()
}
}, jobs, done)
var waiting = {}
var server = rpc.endpoint(ends[1])
.on_request('square', @(params, context) {
waiting[context.id] = context.defer()
jobs.send({ id: context.id, value: params[0] })
})
var inbox = server.listen()
while !server.is_closed() {
var ready = isolate.select([inbox, done])
if ready[0] == inbox {
server.dispatch(ready[1])
} else {
waiting[ready[1].id].reply(ready[1].result)
waiting.remove(ready[1].id)
}
}
jobs.close()
echo caller.join()
[9, 16, 25]
The server defers each request and hands the work to a worker. Its loop waits on both the connection and the worker’s results, and whichever is ready first is dealt with first, so the endpoint never stops answering while work is under way.
The reader isolate does the framing and the JSON decoding, so a large message costs the owning isolate nothing until it is dispatched. Everything is still sent from the owning isolate, and every handler still runs there.
Only a transport another isolate can read can be listened to. That is
stdio and a pipe; a socket, a WebSocket and a child process are held
by one isolate at a time, and listen() on one raises ValueError. Set the
framing before calling listen(), since the reader isolate takes the
framing with it.
serve() and request() work on a listening endpoint too, taking
what arrives from the same channel.
Reading It Yourself
A program that reads the transport itself, such as one polling many
connections at once, hands each read to feed(). It handles every
message the bytes complete, exactly as serve() would have, and empty
bytes end the connection. rpc.serve() drives every connection it
holds this way.
Timeouts
request() and request_batch() take a timeout in seconds:
catch {
var report = client.request('build_report', nil, 30)
} as error {
if instance_of(error, rpc.RpcTimeoutError) {
echo 'gave up on the report'
}
}
A timeout needs the endpoint to be listening. An endpoint reading its
transport itself is inside that transport’s read while it waits, and
waits as long as the read does. Over a socket, the socket’s own
set_read_timeout() is the limit, and a read that runs out raises the
socket’s error.
An answer that arrives after its request timed out has nothing waiting
for it, and goes to on_error().
Stopping and Closing
stop() makes serve() return once the message it is handling is
done, which lets a handler end the conversation:
server.on_request('shutdown', @(params, context) {
context.endpoint.stop()
return 'bye'
})
close() closes the transport. Nothing more can be sent, and sending
raises RpcClosedError; every send_request() callback still waiting
is called with an RpcClosedError. is_closed() is true once the
endpoint is closed or the connection has ended from the other side.
Serving Many Connections
rpc.serve() binds a listening socket and gives every connection
accepted on it an endpoint of its own, spread across a pool of worker
isolates. Each worker polls the connections it holds rather than
blocking on one, so a quiet client costs a descriptor and not a
thread, and one worker serves hundreds of long-lived connections side
by side.
import isolate
import net
import rpc
isolate.configure(6)
def setup(endpoint, peer) {
endpoint.on_request('add', @(params) => params[0] + params[1])
}
def run(ready) {
import rpc
rpc.serve(setup, {
port: 0,
workers: 2,
framing: 'line',
max_connections: 1,
on_ready: @(address, stop) {
ready.send(address.to_string())
},
})
}
var ready = isolate.channel(1)
var server = isolate.spawn(run, ready)
var stream = net.TcpStream()
stream.connect(ready.recv())
var client = rpc.endpoint(rpc.socket(stream)).set_framing(rpc.LineFraming())
echo client.request('add', [20, 22])
client.close()
server.join()
42
setup runs inside a worker for every connection, with the
connection’s endpoint and the client’s address, and gives the endpoint
its handlers. A worker shares nothing with the others, so setup is a
function of a module, or one that uses nothing but its own imports,
and builds whatever a connection needs. Each endpoint is a full peer:
its handlers can call the client back, and it can notify the client
whenever it has something to say.
| Option | Meaning | Default |
|---|---|---|
host | the address to listen on | '127.0.0.1' |
port | the port to listen on; 0 picks a free one | 8000 |
path | a Unix domain socket to listen on, in place of host and port | nil |
workers | how many worker isolates serve connections | the number of CPUs |
backlog | how many accepted connections may wait for a worker | workers * 4 |
framing | 'header' for Content-Length headers, 'line' for one message to a line | 'header' |
max_message_size | the largest message accepted, in bytes | 64 MiB |
max_connections_per_worker | the most connections one worker holds | 256 |
idle_timeout | seconds a connection may stay silent before it is closed | nil, never |
read_timeout | seconds a read may wait once a message has begun | 30 |
write_timeout | seconds a write may wait | 30 |
cert_chain, private_key | PEM strings that put every connection behind TLS | nil |
on_ready | called once bound, with the address and a stop function | nil |
max_connections | stop accepting after this many | nil, never |
on_ready is called on the isolate that called serve(), once the
socket is bound. The stop it is given stops accepting connections;
every worker then finishes the message it is handling, closes its
connections and ends, and serve() returns. That is what a signal
handler calls to shut a server down cleanly. Reaching
max_connections also stops accepting, but serves the connections
already accepted until each closes, which is what the example above
relies on.
Each worker holds a thread of the isolate pool for as long as the
server runs. When serve() is the first thing in a program to use an
isolate it sizes the pool itself; otherwise, call isolate.configure()
first thing, as the example does, with room for the workers and
whatever else the program runs.
With cert_chain and private_key, every connection is TLS, and a
client that fails the handshake is closed without reaching a worker’s
handlers. With path, the server listens on a Unix domain socket,
the way local daemons offer an API to the programs on their machine;
the socket file is removed when the server stops, and Unix domain
sockets need a platform that has them, which net.unix.is_supported()
answers.
A handler that calls its client with request() holds its worker
until the answer arrives, and every other connection on that worker
waits with it. For a call that can take a while, send_request() with
a callback keeps the worker free.
Watching the Conversation
on_trace() sees the text of every message, as it is sent and as it
arrives:
import rpc
import isolate
var ends = rpc.pipe()
isolate.spawn(@(transport) {
import rpc
rpc.endpoint(transport)
.on_request('add', @(params) => params[0] + params[1])
.serve()
}, ends[1])
var client = rpc.endpoint(ends[0])
.on_trace(@(direction, text) {
echo '${direction}: ${text}'
})
client.request('add', [1, 2])
client.close()
out: {"jsonrpc":"2.0","id":1,"method":"add","params":[1,2]}
in: {"jsonrpc":"2.0","id":1,"result":3}
on_error() sees every problem the other side is not told about: an
error a notification handler raised, an error other than an RpcError
a request handler raised, an error a callback raised, a response
nothing was waiting for, and a connection that failed. It is called
with the error and the message it concerns, as a dictionary, or nil
when no one message is to blame:
server.on_error(@(error, message) {
io.stderr.write('${error.type}: ${error.message}\n')
})
Without an on_error() handler, these are dropped. A server that
leaves it unset has no way to learn its own handlers are failing, so
set it.
Errors
Every error the module raises is one of these:
RpcError | a JSON-RPC error, with a code: the other side’s answer, or a message that broke the rules |
RpcHttpError | an HTTP status with no JSON-RPC answer, with the status and the body |
RpcFramingError | the stream’s framing cannot be read, and nothing after it can be either |
RpcClosedError | the connection ended before an answer, or the endpoint has been closed |
RpcTimeoutError | a request’s timeout ran out before its answer |
ValueError | a mistake in the calling program: params that are not structured, a reserved method name, a request answered twice |
The first is the one a program handles as part of its work. The next four are about the transport, and the last is a bug to fix.
What the Module Refuses
It will not send params that are not a list or a dictionary. The specification allows nothing else, and a message that breaks it is refused when it is built, not when the other side receives it.
It will not accept a message that breaks the specification. A
missing or wrong jsonrpc member, a method that is not a string, an
id that is not a string, a number or null, a response with both a
result and an error: each is refused with the code the specification
gives it, never guessed at.
It will not read text that is not UTF-8. JSON is UTF-8, and a
message that is not is answered with PARSE_ERROR rather than read
with its bad bytes replaced. A header block must be ASCII, and an HTTP
body that declares another charset is refused with 415.
It will not answer a response. A malformed response goes to
on_error(), never back to the peer, so two endpoints can never fall
into answering each other’s errors forever.
It will not let a handler claim a reserved name. Method names that
start with rpc. belong to the specification, and neither
on_request(), on_notification() nor on_unhandled() will answer
them.
It will not do unbounded work for one message. A batch over the service’s limit is refused whole, and a framing refuses an oversized message from its header, before reading a byte of it.
It will not defer what has nowhere to be answered. A request that
arrived over HTTP is answered in its HTTP response, and defer() on
it raises rather than leaving the client waiting for an answer that
cannot come.
It will not listen to a transport only one isolate can hold. A
socket, a WebSocket or a child process is read where it is held, and
listen() on one raises rather than reading it from somewhere it
cannot be.
Module Reference
The standard library reference documents every class and method. The shape of the module:
rpc.Service() | a service: handlers, and answer() |
rpc.http_handler(service) | an http route handler answering with a service |
rpc.HttpClient(url, options) | a client calling a service over HTTP |
rpc.endpoint(transport) | an Endpoint over a transport |
rpc.serve(setup, options) | an endpoint for every connection to a listening socket |
rpc.stdio() | the program’s own standard streams |
rpc.socket(stream) | a connected net stream |
rpc.process(child) | a child process’s standard streams |
rpc.websocket(socket) | a WebSocket from http.websocket |
rpc.pipe() | two transports joined to each other |
rpc.encode(message) | a message, or a list of them, as JSON text |
rpc.decode(text) | JSON text as a message, or a batch of them |
rpc.read(value) | a decoded value as a message, or a batch of them |
The messages:
Request(method, params, id) | a call that expects an answer |
Notification(method, params) | a call that expects none |
Response(id, result, error) | an answer, also made by Response.success() and Response.failure() |
RpcError(code, message, data) | an error, with to_dict() |
PARSE_ERROR, INVALID_REQUEST, METHOD_NOT_FOUND | the codes the specification defines |
INVALID_PARAMS, INTERNAL_ERROR | the rest of them |
SERVER_ERROR_MIN, SERVER_ERROR_MAX | the range set aside for implementations |
On a Service, and so on an Endpoint:
on_request, on_notification, on_unhandled | what it answers |
on_error, on_trace | what it reports |
set_batch_limit, batch_limit | the most messages a batch may hold, DEFAULT_BATCH_LIMIT by default |
answer | one message or batch in, its answer out |
On an Endpoint as well:
set_framing | how its messages are told apart |
request, request_batch, send_request, notify | calling the other side |
send, handle | sending and handling messages as they are |
serve, stop | reading on the isolate that calls it |
listen, dispatch | reading on an isolate of its own |
feed | handling what the program read itself |
close, is_closed | ending the connection |
On an HttpClient: request, request_batch, notify and send.
On a Context:
endpoint, id, method, request | what is being handled, where, and what it arrived in |
is_notification | whether anything is answered |
defer, is_deferred | taking over answering |
reply, fail, is_answered | answering a deferred request |
The framings, HeaderFraming, LineFraming and MessageFraming:
feed(data) | the messages the next bytes complete |
frame(message) | the bytes that send a message |
set_max_size(size), max_size() | the largest message accepted |
pending() | how many bytes are held for an unfinished message |
The transports, StdioTransport, SocketTransport, ProcessTransport,
WebSocketTransport and ChannelTransport, each have read(max),
write(data), close() and can_listen(), and WebSocketTransport
has framing().
Editor Support
Zuri ships a language server. zuri lsp speaks the Language Server
Protocol, so any editor that speaks it too gets completion, navigation,
diagnostics as you type, refactoring, formatting and test running for
Zuri, from the same server that ships with the runtime. The server is
written in Zuri, on the zuri module’s own parser and compiler, so what
it reports is what the runtime would do with the same code.
Visual Studio Code has an extension that sets everything up. Neovim, Helix, Sublime Text and Emacs each need a few lines of configuration, shown at the end of this chapter.
- What the Server Does
- How the Server Reads a Project
- Settings
- Starting the Server
- Visual Studio Code
- Neovim
- Helix
- Sublime Text
- Emacs
- When Something Goes Wrong
What the Server Does
Writing Code
Completion offers what can be written at the cursor:
- after a
., the members of the receiver’s type: an instance’s fields and methods, inherited ones and the ones extensions add, a class’s statics, a module’s exports, a built-in type’s methods, and the keys of a dictionary written out where its variable was declared; - elsewhere, the names in scope, the built-in globals, and the keywords that can start what is being written;
- after
import, module paths, and between an import’s braces, what the module exports; - at the start of a class member,
@newand the operator decorators the class does not have yet, as method skeletons; - after the
:of a type hint, the type names; - inside a doc block, its tags after
@, and after@paramthe names of the parameters it documents.
A name another module exports is offered with the import it needs: a module
of the standard library or a package as import json and json.encode, and
a module of the project as import .shapes { Point }.
Signature help follows a call while its arguments are written. It marks
the parameter the cursor is on and shows its documentation, follows a
constructor to its @new, and stays on a variadic parameter for every
argument it takes.
Hover shows a declaration as a line of Zuri, its documentation laid out the way the standard library reference lays it out, the type a variable is inferred to hold, the value of a constant, and a module’s overview.
Inlay hints name the parameter each literal argument goes to, and, when turned on, the type a variable is inferred to hold.
Finding Your Way
Go to definition follows an imported name back to the module that declares it, and a built-in to the stub that documents it. Go to declaration stops at the import that brings a name into the file, and go to type definition goes to the class of the value a name holds. Go to implementation lists the methods of subclasses that override a method, or the subclasses of a class.
Find references and the highlights of the current name go by what each name
resolves to, never by its spelling, so a distance method of one class is
never confused with another class’s.
The call hierarchy shows what calls a function and what it calls; the type
hierarchy, the classes above and below a class. Each file has an outline,
and the workspace symbol search finds a declaration by the letters of its
name in order, so hrq finds HttpRequest. Each import links to the file
it loads, and each class, function and method shows how often it is used.
Diagnostics
Every error the server reports is either an error the compiler reports or a failure the runtime is certain to raise when the code runs:
- a syntax error, with the recovered rest of the file still checked;
- an error from the compiler, such as
breakoutside a loop; - an import of a module that does not exist, or of a name its module does not have;
module.namefor a name the module does not have;- a name nothing declares;
self.name = valueoutside@new, for a field no class in the chain declares;- a literal argument of the wrong type, or a missing one, for a typed parameter, with the runtime’s own message.
A member that an instance of a fully known class does not have is a
warning, since an extension the server has not read could add it. Unused
imports and declarations, and code after a return, raise, break or
continue, are shown faded. A use of anything documented @deprecated is
struck through.
The server never guesses. A value whose type cannot be told, a module whose wildcard imports lead somewhere unread, and a class whose superclass cannot be found are never diagnosed. Arity is not checked either, because Zuri does not enforce it.
Quick fixes add the import a missing name needs, correct a misspelt name or
member, remove an unused import, and declare a field set outside @new.
Changing Code
Rename changes a declaration and every use of it across the workspace.
A { name } dictionary entry keeps its key and becomes { name: renamed }.
A rename is refused, with the reason, when the new name is a keyword, when
it is already declared where the declaration or a use of it would see it,
when it would make private something used from outside, and for what the
standard library or a package declares. A member used through a receiver
whose type cannot be told is left alone, and the editor is told where.
Extract and inline refactorings are offered only where they keep what the program does:
- Extract variable moves an expression into a
varjust before its statement, never out of a loop’s condition, the right ofandoror, a branch of?:, awhenlabel or anelse ifcondition, and never past something with a side effect that runs first. Every identical copy in the block can share the variable when the expression has no side effect. - Extract function turns whole statements of one block into a function.
The outer locals they use become its parameters, and what they set and is
read afterwards comes back: one value directly, several in a dictionary.
Statements that use
selforparentbecome a private method. - Inline variable replaces each use of a
varset once and never again with its value. - Inline function replaces a call of a function whose body is a single
returnwith that expression, and removes the function once every call is inlined.
Formatting runs zuri fmt on a file or a selection and sends back only
the lines that differ. A file with a syntax error is left as it is.
Renaming or moving a file updates the relative imports that reach it, in the files that import it and in the moved file itself.
Running Tests
The server finds the tests every file of the workspace declares with the
test module: each describe and it, and their _only, _skip, _todo
and _each forms, whose name is written as a string. A lens above each one
runs it, and one above the first runs the whole file.
A test runs as zuri test would run it: zuri run on its file, in a process
of its own, from the root of the project. Each result reaches the editor as
it happens, and each failed assertion is shown on the line it failed at
until the file changes or its tests run again.
Highlighting
Semantic highlighting colours each name by what it resolves to: a class, a function, a method, a parameter, a variable, a field, a module or a decorator, with whatever the standard library defines marked as such. It works with any editor theme that colours semantic tokens.
How the Server Reads a Project
The server reads every .zu file of the workspace folders the editor opens,
in the background, on isolates of its own, so it answers while it reads.
Inside a git work tree it reads what git tracks and what git would track,
so .gitignore is honoured; anything under .git, .zuri or
node_modules is left out, and so is whatever zuri.exclude names.
Imports resolve exactly as the runtime resolves them: a relative import
against the importing file, then the project’s .zuri/libs, then the
standard library, then the native modules, then the packages installed for
the user in ZURI_HOME/libs. The standard library is read from the Zuri
installation the server belongs to, or from zuri.root, or from
ZURI_ROOT. A module the workspace imports is read the first time it is
needed.
Settings
The server reads its settings from the zuri section of the editor’s
configuration, as nested objects or dotted names, and follows changes as
they are made.
| Setting | Default | What it does |
|---|---|---|
root | '' | The installation whose standard library is read; empty uses ZURI_ROOT, or the installation beside zuri |
exclude | [] | Globs, relative to a workspace folder, of files never read |
index.workers | 2 | Isolates that read the workspace in the background |
diagnostics.enable | true | Whether problems are reported at all |
diagnostics.scope | 'openFiles' | 'openFiles', or 'workspace' for every file of the workspace |
diagnostics.delay | 300 | Milliseconds after a change before a file is checked |
diagnostics.unusedHints | true | Whether unused declarations are shown faded |
inlayHints.parameterNames | 'literals' | Which arguments are named: 'none', 'literals' or 'all' |
inlayHints.variableTypes | false | Whether inferred variable types are shown |
codeLens.references | true | Whether declarations show how often they are used |
completion.autoImport | true | Whether names from modules not imported yet are offered |
completion.callParentheses | false | Whether completing a function writes its parentheses |
testing.codeLens | true | Whether suites and tests show lenses that run them |
Starting the Server
Editors start the server themselves, over its standard input and output:
zuri lsp [--stdio] [--log <path>] [--log-level <level>]
--stdio names the only way the server talks, and is accepted because
editors pass it. --log appends the server’s log to a file as well as
sending it to the editor. --log-level sets how much is logged: error,
warn, info (the default), debug, or trace, which also writes every
message the editor and the server exchange to the log file.
The server exits with 0 when the editor shuts it down in order, and with
1 when the connection ends without that, or when the editor that started
it has exited.
Visual Studio Code
Install the Zuri extension from the marketplace, or from a downloaded package:
code --install-extension zuri-vscode-0.2.0.vsix
The extension starts the server for every workspace with a .zu file. Its
settings are the server’s, under zuri., plus two of its own: zuri.path,
the zuri executable to run, and zuri.trace.server, which records the
messages between the editor and the server in the Zuri Language Server
output. Tests appear in the Test Explorer, and the commands Zuri: Restart
Language Server, Zuri: Show Language Server Output and Zuri: Check
Zuri Installation are in the command palette.
Neovim
Neovim 0.11 configures language servers itself:
vim.filetype.add({ extension = { zu = 'zuri' } })
vim.lsp.config('zuri', {
cmd = { 'zuri', 'lsp' },
filetypes = { 'zuri' },
root_markers = { 'project.toml', '.git' },
settings = {
zuri = {
inlayHints = { variableTypes = true },
},
},
})
vim.lsp.enable('zuri')
Inlay hints are shown with vim.lsp.inlay_hint.enable(), and a code lens
runs with vim.lsp.codelens.run().
Helix
Add the language and its server to languages.toml:
[[language]]
name = "zuri"
scope = "source.zuri"
file-types = ["zu"]
roots = ["project.toml"]
comment-token = "#"
block-comment-tokens = { start = "/*", end = "*/" }
indent = { tab-width = 2, unit = " " }
language-servers = ["zuri-lsp"]
[language-server.zuri-lsp]
command = "zuri"
args = ["lsp"]
[language-server.zuri-lsp.config.zuri]
inlayHints = { parameterNames = "literals" }
Sublime Text
With the LSP package installed, add a client to its settings:
{
"clients": {
"zuri": {
"enabled": true,
"command": ["zuri", "lsp"],
"selector": "source.zuri",
"settings": {
"zuri.diagnostics.scope": "openFiles"
}
}
}
}
The selector names the syntax Zuri files open with. The Visual Studio Code
extension’s grammar, zuri.tmLanguage.json, is a TextMate grammar whose
scope is source.zuri; PackageDev’s Convert command turns it into a
.tmLanguage file for Sublime Text.
Emacs
Eglot, built into Emacs 29, needs a mode for Zuri files and the command that starts the server:
(define-derived-mode zuri-mode prog-mode "Zuri"
"A major mode for Zuri source."
(setq-local comment-start "# "))
(add-to-list 'auto-mode-alist '("\\.zu\\'" . zuri-mode))
(with-eval-after-load 'eglot
(add-to-list 'eglot-server-programs '(zuri-mode "zuri" "lsp")))
(add-hook 'zuri-mode-hook #'eglot-ensure)
Settings go in eglot-workspace-configuration:
(setq-default eglot-workspace-configuration
'(:zuri (:inlayHints (:variableTypes t))))
When Something Goes Wrong
The server sends what it logs at its --log-level and above, info by
default, to the editor’s log for language servers, and while the editor asks
for a trace, everything down to debug as well. Starting it with --log and
--log-level debug writes all of it to a file; --log-level trace adds
every message exchanged.
When nothing works at all, the editor cannot start zuri: running
zuri --version in a terminal shows whether it is on PATH, and an editor
that takes a path, as zuri.path does in Visual Studio Code, can be given
the executable directly.
When the standard library cannot be found, the server says so in its log.
Setting zuri.root, or ZURI_ROOT for the editor, to the Zuri installation
fixes it.
Appendix A: Keywords
Zuri reserves thirty-one words. None of them can be used as a variable, function, class or parameter name.
| Keyword | What it does | Covered in |
|---|---|---|
and | logical conjunction, short-circuiting | Operators |
as | binds an error in catch, or renames an import | Errors, Modules |
assert | raises AssertError when its condition is falsy | Control Flow |
break | leaves the innermost loop | Control Flow |
catch | runs a block, intercepting anything it raises | Errors |
class | declares a class, or with > an extension to one | Classes, Extensions |
const | declares a name that cannot be reassigned | Variables |
continue | skips to the next iteration | Control Flow |
def | declares a function, named or anonymous | Functions |
default | the fall-through branch of a using | Control Flow |
do | begins a do/while loop | Control Flow |
echo | prints a value and a newline | Hello, World! |
else | the alternative branch of an if | Control Flow |
false | the boolean false | Data Types |
for | iterates over anything iterable | Control Flow |
if | conditional branch | Control Flow |
import | loads a module | Modules |
in | separates a for loop’s variables from its iterable | Control Flow |
iter | the counting loop | Control Flow |
nil | the absence of a value | Data Types |
or | logical disjunction, short-circuiting | Operators |
parent | the superclass constructor, or a superclass method | Inheritance |
raise | raises an error | Errors |
return | leaves the current function | Control Flow |
self | the current instance, inside a method | Classes |
static | puts a field or method on the class, not the instance | Classes |
true | the boolean true | Data Types |
using | multi-way branch on one subject | Control Flow |
var | declares a variable | Variables |
when | one branch of a using | Control Flow |
while | the conditional loop | Control Flow |
Names That Are Not Keywords
These are ordinary globals, not reserved words, so nothing stops you from shadowing one. Doing so is a good way to confuse the next reader:
time sum bytes file instance_of typeof
delprop getprop hasprop setprop id print
rand is_bigint is_bool is_bytes is_callable is_class
is_dict is_file is_function is_instance is_int is_iterable
is_list is_number is_object is_string
The built-in error classes are also globals: Error, TypeError,
ValueError, NumericError, ArgumentError, NotImplementedError,
RangeError, AccessError, AssertError, PropertyError,
UndefinedError and ModuleNotFoundError.
Reserved by Convention
Two module-level names are provided by the runtime rather than declared by
you: __file__ and __root__. See Modules.
Names beginning with $ are never produced by the lexer, which is how the
compiler synthesises loop variables that cannot collide with yours.
A leading underscore marks something private, both for class members and for module members. The compiler enforces it in both cases.
Appendix B: Operators and Precedence
Precedence
Tightest first. Operators on the same row bind equally and associate left
to right, except **, which associates right to left.
| Level | Operators | Notes |
|---|---|---|
| 1 | literals, (...), [...], {...}, self, parent | |
| 2 | .. | binds primaries only |
| 3 | . () [] | member, call, index and slice |
| 4 | ++ -- | postfix only |
| 5 | ** | right-associative; the exponent may carry a unary operator |
| 6 | ! - ~ | unary |
| 7 | * / // % | |
| 8 | + - | |
| 9 | << >> >>> | |
| 10 | & | |
| 11 | ^ | |
| 12 | | | |
| 13 | < <= > >= == != | |
| 14 | and | |
| 15 | or | |
| 16 | ? : | |
| 17 | = and every compound assignment |
Consequences worth remembering:
2 ** 3 ** 2is512.**is right-associative.2 * 3 ** 2is18.**outranks*.-2 ** 2is-4.**binds tighter than unary minus; write(-2) ** 2to raise a negative number.2 ** -1is0.5. The exponent may carry its own sign.1 + 2..5is1 + (2..5). Parenthesise ranges with computed endpoints.
Arithmetic
| Operator | Meaning | Decorator |
|---|---|---|
+ | addition; string and list concatenation | @add |
- | subtraction | @sub |
* | multiplication; string and list repetition | @mul |
/ | division, always floating point | @div |
// | floor division, rounds toward negative infinity | @floordiv |
% | remainder, keeps the sign of the left operand | @mod |
** | exponentiation | @pow |
- (unary) | negation | @neg |
Comparison
| Operator | Meaning | Decorator |
|---|---|---|
== | equal, by value for numbers, strings, lists, dicts; by identity otherwise | @eq |
!= | not equal, the negation of == | @eq |
< | less than, numbers only | @lt |
<= | less than or equal | @lte |
> | greater than | @gt |
>= | greater than or equal | @gte |
@eq runs only when both operands are objects, so x == nil and x == 5
never call it.
Logic
| Operator | Meaning |
|---|---|
and | both truthy; returns the operand, not a bool |
or | either truthy; returns the operand, not a bool |
! | logical negation, always a bool |
? : | conditional expression |
Bitwise
| Operator | Meaning | Decorator |
|---|---|---|
& | and | @and |
| | or | @or |
^ | xor | @xor |
~ | complement | @not |
<< | left shift | @lshift |
>> | arithmetic right shift | @rshift |
>>> | logical right shift, zero filling | @urshift |
Assignment
| Operator | Equivalent to |
|---|---|
= | assignment |
+= -= *= /= //= **= %= | x = x op y |
&= |= ^= ~= <<= >>= >>>= | x = x op y |
++ -- | increment, decrement; postfix only, evaluates to the new value |
Access
| Syntax | Meaning |
|---|---|
x.name | member of an object, dictionary or module |
x[key] | index with a computed key |
x[a, b] | slice, from a up to but not including b |
x[, b] | slice from the start |
x[a, ] | slice to the end |
a..b | a range value |
...x | variadic parameter, in a parameter list only |
Negative indices count back from the end, for strings, lists and bytes.
Truthiness
Falsy: false, nil, 0, 0.0, -0.0, NaN, 0n, '', and
bytes(0).
Truthy: everything else, including every negative number, [], {} and
'0'.
Operators Zuri Does Not Have
No in for membership; use contains(). No ?. optional chaining. No
?? null coalescing; or covers it, with the truthiness caveat above. No
comma operator. No prefix ++/--.
Appendix C: Decorated Methods
A method whose name begins with @ is called by the runtime when a piece
of syntax is applied to an instance of its class. They are ordinary
methods otherwise: inherited, overridable, and callable by name.
See Decorated Methods for the guided treatment, and Class Extensions for adding one to a class you did not write.
Construction
| Decorator | Called by | Signature |
|---|---|---|
@new | ClassName(...) | @new(...args) |
@new is the only place self.x = value may declare a field that was not
declared with var.
Arithmetic
| Decorator | Operator | Signature |
|---|---|---|
@add | a + b | @add(other) |
@sub | a - b | @sub(other) |
@mul | a * b | @mul(other) |
@div | a / b | @div(other) |
@floordiv | a // b | @floordiv(other) |
@mod | a % b | @mod(other) |
@pow | a ** b | @pow(other) |
@neg | -a | @neg() |
Bitwise
| Decorator | Operator | Signature |
|---|---|---|
@and | a & b | @and(other) |
@or | a | b | @or(other) |
@xor | a ^ b | @xor(other) |
@not | ~a | @not() |
@lshift | a << b | @lshift(other) |
@rshift | a >> b | @rshift(other) |
@urshift | a >>> b | @urshift(other) |
@not is bound to ~, the bitwise complement. ! is logical negation and
is not overridable: an instance is always truthy, so !instance is always
false.
Comparison
| Decorator | Operator | Signature |
|---|---|---|
@eq | a == b, a != b | @eq(other) |
@lt | a < b | @lt(other) |
@lte | a <= b | @lte(other) |
@gt | a > b | @gt(other) |
@gte | a >= b | @gte(other) |
!= is the negation of @eq. @eq runs only when the right operand is
an object too, so x == nil never calls it, and it must return a bool.
Lists, dictionaries, contains() and index_of() compare instances by
identity whether or not the class defines it.
Iteration
| Decorator | Called by | Signature |
|---|---|---|
@key | for ... in | @key(previous) |
@value | for ... in | @value(key) |
@key(previous) receives the previous key, starting from nil, and
returns the next one or nil when the sequence is finished. @value(key)
returns what is stored at that key.
Defining both is what makes is_iterable() return true for the class.
Display
| Decorator | Called by | Signature |
|---|---|---|
@to_string | echo, print() | @to_string() |
@to_string() returns the text shown for the instance, including when it
sits inside a list or dictionary being shown, as a key or as a value. It
must return a string; anything else raises a TypeError. Without it, an
instance shows as <instance of ClassName>.
Serialisation
| Decorator | Called by | Signature |
|---|---|---|
@to_json | json.encode() | @to_json() |
Returns whatever should be encoded in the instance’s place, which is where you decide what does and does not cross the wire.
Not a Decorator
to_string() has no @. It is a real method every value already carries,
and a class may override it.
Nothing calls it implicitly. String interpolation and + render an
instance as <instance of ClassName>, and echo and print() use
@to_string().
Resolution
The left operand decides. a + b looks for @add on a’s class only; if
a is a number and b is your instance, the operation is a TypeError
rather than a call to b’s @add.
An operator with no matching decorator raises a TypeError naming the
exact signature:
operator '+' not defined for call signature (nil, number)
Appendix D: Built-in Functions
These functions are available in every file with no import. They are the parts of the language that happen to be spelled as calls rather than as syntax.
| Function | Returns | Summary |
|---|---|---|
time() | number | Returns the current epoch time to the microseconds resolution. |
sum(...values: list) | number | Calculates the sum of all the elements passed as arguments. |
bytes(x: number|list) | bytes|any | If x is a number, this function returns a new bytes object with length x having all its bytes set to 0x0. |
file(path: string, mode: ?string) | file | Returns an open file handle to the file specified in the path in the specified mode. |
instance_of(x, y) | boolean | Returns true if x is an instance of the given class y or false otherwise. |
typeof(x) | string | Returns the type of the given value as a string. |
delprop(object: instance, name: string) | void | Deletes the property name from the given instance of object. |
getprop(object: instance, name: string) | any|nil | Returns the value of the property name from the given instance of object. |
hasprop(object: instance, name: string) | boolean | Returns true if the property name exists in the given instance of object. |
setprop(obj: instance, prop: string, value) | boolean | Sets the value of the object’s property with the matching name to the given value. |
id(x) | number | Returns the unique identifier of value x within the system. |
print(...values: list) | void | Prints the given arguments to standard output. |
rand(x: ?number, y: ?number) | number | If no argument is given, returns a random number between 0 and 1. |
is_bigint(x) | boolean | Returns true if x is a bigint or false otherwise. |
is_bool(x) | boolean | Returns true if x is a boolean or false otherwise. |
is_callable(x) | boolean | Returns true if x is a callable or false otherwise. |
is_class(x) | boolean | Returns true if x is a class or false otherwise. |
is_dict(x) | boolean | Returns true if x is a dictionary or false otherwise. |
is_function(x) | boolean | Returns true if x is a function or false otherwise. |
is_instance(x) | boolean | Returns true if x is an instance of any class or false otherwise. |
is_int(x) | boolean | Returns true if x is an integer or false otherwise. |
is_list(x) | boolean | Returns true if x is a list or false otherwise. |
is_number(x) | boolean | Returns true if x is a number or false otherwise. |
is_object(x) | boolean | Returns true if x is an object or false otherwise. |
is_string(x) | boolean | Returns true if x is a string or false otherwise. |
is_bytes(x) | boolean | Returns true if x is bytes or false otherwise. |
is_file(x) | boolean | Returns true if x is a file or false otherwise. |
is_iterable(x) | boolean | Returns true if x is an iterable object or false otherwise. |
time()
time() -> number
Returns the current epoch time to the microseconds resolution.
Example:
%> time()
1686787200.123456
The time is returned as a floating point number where the integer part represents the number of seconds since the epoch and the fractional part represents the microseconds.
Returns number
Note: The epoch time is the number of seconds that have elapsed since January 1, 1970 (midnight UTC/GMT).
sum()
sum(...values: list) -> number
Calculates the sum of all the elements passed as arguments. Returns 0
when no argument is passed in.
Example:
%> math.sum([1, 2, [3, 4, [5, 6]]])
21
Parameters
values(...number)
Returns number
bytes()
bytes(x: number|list) -> bytes|any
If x is a number, this function returns a new bytes object with length
x having all its bytes set to 0x0.
If x is a list, it returns a new bytes object whose contents are the
bytes specified in the list.
%> bytes(5)
(00 00 00 00 00)
%> bytes([65, 66, 67, 68, 69])
(41 42 43 44 45)
Parameters
x(number|list) — The number or list to convert to bytes.
Returns bytes|any
Note: If x is a list, then the list must only contain valid bytes which can be any number between 0 and 255.
file()
file(path: string, mode: ?string) -> file
Returns an open file handle to the file specified in the path in the specified mode. If the mode is not specified, the file will be opened in the read only mode.
Valid modes include:
%> file('sample.txt', 'r')
<file at sample.txt in mode r>
%> file('sample.txt', 'w')
<file at sample.txt in mode w>
%> file('sample.txt', 'a')
<file at sample.txt in mode a>
%> file('sample.txt', 'r+')
<file at sample.txt in mode r+>
%> file('sample.txt', 'w+')
<file at sample.txt in mode w+>
%> file('sample.txt', 'a+')
<file at sample.txt in mode a+>
%> file('sample.lock', 'x')
<file at sample.lock in mode x>
x and x+ create the file and fail when anything already exists at
the path. The check and the creation are one step, so of two programs
creating the same path in x mode, exactly one succeeds. That makes it
the mode for lock files. The failure comes on the first read, write or
open(), since creating the handle touches nothing.
Parameters
path(string) — The path to the file to open.mode(?string) — The mode to open the file in.
Returns file
instance_of()
instance_of(x, y) -> boolean
Returns true if x is an instance of the given class y or false
otherwise.
Parameters
x(any) — The value to check.y(class) — The class to check for.
Returns boolean
typeof()
typeof(x) -> string
Returns the type of the given value as a string.
Parameters
x(any) — The value to check.
Returns string
delprop()
delprop(object: instance, name: string) -> void
Deletes the property name from the given instance of object.
Parameters
object(instance) — The instance to delete the property from.name(string) — The name of the property to delete.
Returns void
getprop()
getprop(object: instance, name: string) -> any|nil
Returns the value of the property name from the given instance of
object. If the object has no such property, nil is returned.
Parameters
object(instance) — The instance to get the property from.name(string) — The name of the property to get.
Returns any|nil
hasprop()
hasprop(object: instance, name: string) -> boolean
Returns true if the property name exists in the given instance of
object. If the object has no such property, false is returned.
Parameters
object(instance) — The instance to check for the property.name(string) — The name of the property to check.
Returns boolean
setprop()
setprop(obj: instance, prop: string, value) -> boolean
Sets the value of the object’s property with the matching name to the
given value. If the property already exists, it overwrites it and
returns true, otherwise it returns false.
Parameters
obj(instance) — The object to set the property of.prop(string) — The property to set.value(any) — The value to set the property to.
Returns boolean
id()
id(x) -> number
Returns the unique identifier of value x within the system. This value is also equivalent to the current address of object x in memory.
Parameters
x(any) — The value to get the identifier of.
Returns number
print()
print(...values: list) -> void
Prints the given arguments to standard output.
Unlike echo (which always appends a newline and only ever prints one
value), print() writes every argument back-to-back with no separator
and no trailing newline. It also critically writes a bytes object as
RAW bytes rather than its Display text. That raw-byte path is what
lets a script stream binary output (e.g. a PBM/PNG image body one
scanline at a time).
Parameters
values(...any) — Any number of arguments to print
Returns void
Note: In the REPL, it also appends a newline at the end.
rand()
rand(x: ?number, y: ?number) -> number
If no argument is given, returns a random number between 0 and 1. If x is given, returns a random number between 0 and x. If y is given, returns a random number between x and y.
Parameters
x(?number) — The lower bound of the random number.y(?number) — The upper bound of the random number.
Returns number
is_bigint()
is_bigint(x) -> boolean
Returns true if x is a bigint or false otherwise. A bigint is a
distinct type from number created either with the n literal suffix
(123n) or by an operation whose result overflows what a regular
number can represent exactly. is_number(x) and is_int(x) are both
false for a bigint even though it holds an integer value; check
is_bigint(x) separately when a value might be either.
Parameters
x(any) — The value to check.
Returns boolean
is_bool()
is_bool(x) -> boolean
Returns true if x is a boolean or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_callable()
is_callable(x) -> boolean
Returns true if x is a callable or false otherwise. Callables
includes classes, functions, methods and closures.
Parameters
x(any) — The value to check.
Returns boolean
is_class()
is_class(x) -> boolean
Returns true if x is a class or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_dict()
is_dict(x) -> boolean
Returns true if x is a dictionary or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_function()
is_function(x) -> boolean
Returns true if x is a function or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_instance()
is_instance(x) -> boolean
Returns true if x is an instance of any class or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_int()
is_int(x) -> boolean
Returns true if x is an integer or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_list()
is_list(x) -> boolean
Returns true if x is a list or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_number()
is_number(x) -> boolean
Returns true if x is a number or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_object()
is_object(x) -> boolean
Returns true if x is an object or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_string()
is_string(x) -> boolean
Returns true if x is a string or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_bytes()
is_bytes(x) -> boolean
Returns true if x is bytes or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_file()
is_file(x) -> boolean
Returns true if x is a file or false otherwise.
Parameters
x(any) — The value to check.
Returns boolean
is_iterable()
is_iterable(x) -> boolean
Returns true if x is an iterable object or false otherwise.
Iterables includes lists, dictionaries, strings, bytes, and instances of
any class that defines both @key() and @value() decorator functions.
Parameters
x(any) — The value to check.
Returns boolean
Module Variables
Every file also has these variables, which describe the file itself rather than anything it declares.
__file__
The path of the file it is read in. In the file a program was started from, it is the path the program was started with; in an imported module, it is the module’s full path.
Type string
__root__
The path of the file the program was started from, the same in every
module of one run and in every isolate, so __root__ == __file__ holds
in that one file alone. That is how a file serves both as a module
others import and as a program of its own:
def main() {
echo 'started directly'
}
if __root__ == __file__ {
main()
}
__root__ and __file__ are both set when a file runs; the REPL, which
runs no file, defines neither.
Type string
Appendix E: Built-in Type Methods
Every method on every built-in type, with its signature, what it returns, and its edge cases. This is the reference; the chapters in Text, Numbers and Collections are the introduction.
Methods are called with a dot, on the value itself:
echo 'zuri'.upper()
echo 255.hex()
echo [3, 1, 2].sort()
ZURI
ff
[1, 2, 3]
| Type | Methods | Page |
|---|---|---|
string | 41 | String Methods |
number | 43 | Number Methods |
bigint | 26 | Bigint Methods |
bool | 1 | Boolean Methods |
list | 40 | List Methods |
dict | 21 | Dictionary Methods |
range | 9 | Range Methods |
bytes | 25 | Bytes Methods |
file | 27 | File Methods |
function | 6 | Function Methods |
What Is Not Listed Here
@key and @value exist on every iterable built-in type, which is
what makes for ... in work on them. They are documented as a protocol
in Appendix C rather than repeated on every
page.
to_string() exists on every value, nil included.
Class instances carry whatever their class declares, plus
to_string(). See Classes and Objects.
Module members are not methods. See Appendix F.
Conventions in These Pages
A parameter written name: ?type is optional. A parameter written
...name is variadic. A parameter with no type shown takes any value.
A method that mutates its receiver says so. Where a type has both,
the distinction matters: list.sort() mutates and returns the list,
while list.reverse() returns a new list and leaves the original
alone.
String Methods
Every method on the built-in string type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
length() | number | Returns the length of a string. |
upper() | string | Returns a copy of the string with all the cased characters converted to uppercase. |
lower() | string | Return a copy of the string with all the cased characters converted to lowercase. |
is_alpha() | boolean | Returns true if all the characters in the string are all alphabets and the string is not empty., otherwise returns false. |
is_alnum() | boolean | Returns true if all the characters in the string are either alphabets or numbers and the string is not empty, otherwise returns false. |
is_number() | Returns true if all the characters in the string are all digits and the string is not empty, otherwise returns false. | |
is_lower() | boolean | Returns true if at least one character in the string is cased, all cased characters are lower cased and the string is not empty. |
is_upper() | boolean | Returns true if at least one character in the string is cased, all cased characters are upper cased and the string is not empty. |
is_space() | boolean | Returns true if there are only whitespace characters in the string and the string is not empty. |
ord() | number | Returns the Unicode code point of the string, which must be exactly one character long. |
trim(chars: ?string) | string | Returns a copy of the string with characters stripped from both ends. |
ltrim(chars: ?string) | string | Returns a copy of the string with characters stripped from its start only. |
rtrim(chars: ?string) | string | Returns a copy of the string with characters stripped from its end only. |
join(string: string) | string | Returns a string which is a concatenation of the items in the iterable using the string as the separator. |
split(delimiter: string) | list | Returns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found. |
index_of(str: string, start_index: ?number) | number | Returns the index position of the first occurrence of the string str in the string string. |
last_index_of(str: string, end_index: ?number) | number | Returns the index position of the last occurrence of the string str in the string string, searching from the end. |
starts_with(str: string) | boolean | Returns true if the string begins with the string or character specified in str, otherwise it returns false. |
ends_with(str: string) | boolean | Returns true if the string ends with the string or character specified in str, otherwise it returns false. |
count(str: string) | number | Returns the number of non-overlapping occurrences of the substring str in the string. |
to_number(base) | number | Returns the first numeric value contained in the string if any exists or 0 if the string contains no numeric value. |
to_bigint(base) | bigint | Returns the integer value of the string as a bigint, or 0n if the string does not spell one. |
to_list() | list | Returns a list whose elements consists of every character contained in the string in order of appearance. |
to_bytes() | bytes | Returns the content of the string as a stream of bytes. |
lpad(width: number, fill: ?string) | string | Returns the string left justified in a string of length width. |
rpad(width: number, fill: ?string) | string | Returns the string right justified in a string of length width. |
match(str: string) | boolean|dictionary | If the string str is a regular string, this method returns true if the string contains a substring str. |
matches(reg: string) | dictionary | Returns a dictionary containing every match of the given regular expression reg in the source string. |
replace(str: string, replacement: string, use_regex: ?bool) | string | Returns a copy of the string with all occurrences or matches of str replaced by the replacement string. |
replace_with(regex: string, callback: function) | string | Returns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred. |
ascii() | string | Reinterprets the string as a raw byte view: each byte of its UTF-8 encoding becomes its own character (a codepoint between 0 and 255, i.e. |
case_fold() | string | Returns a copy of the string case-folded for case-insensitive comparison, using full Unicode case folding rather than plain lowercasing. |
compare(other: string) | number | Compares the string with another string. |
is_empty() | boolean | Returns true if the string is empty, false otherwise. |
contains(str: string) | boolean | Returns true if the string contains the specified substring, false otherwise. |
lines() | list | Returns the lines of the string as an list as it would be if split on newline characters. |
each_line(callback: function) | void | Iterates over each line of the string, calling the provided callback function with the line and its index. |
each(callback: function) | void | Iterates over each character of the string, calling the provided callback function with the character and its index. |
capitalize() | string | Returns a new string with the first character capitalized and the rest in lowercase. |
title() | string | Returns a new string with each word capitalized. |
to_string() | string | Returns the string itself. |
length()
length() -> number
Returns the length of a string. Note that this method is UTF-8
compatible and will return the UTF-8 length for the string if the string
contains UTF-8 characters whether written directly or via the \u or
\U escapes.
For example:
%> 'This is a pretty long string'.length()
28
%> 'उनका एक समय'.length()
11
%> 'This text mixes English and 粵語'.length()
30
Returns number
upper()
upper() -> string
Returns a copy of the string with all the cased characters converted to
uppercase. Note that the result of this method may return false when
tested with is_upper() of the string contains Unicode characters
that are not case folded.
For example:
%> 'zuri'.upper()
'ZURI'
Returns string
lower()
lower() -> string
Return a copy of the string with all the cased characters converted to
lowercase.
For example:
%> 'Zuri Is Bae'.lower()
'zuri is bae'
Returns string
is_alpha()
is_alpha() -> boolean
Returns true if all the characters in the string are all alphabets and
the string is not empty., otherwise returns false.
For example:
%> 'abracadabra'.is_alpha()
true
%> 'my tooth aches'.is_alpha()
false
%> ''.is_alpha()
false
Returns boolean
is_alnum()
is_alnum() -> boolean
Returns true if all the characters in the string are either alphabets
or numbers and the string is not empty, otherwise returns false. This
method is the same as string.is_alpha() or string.is_number().
For example:
%> '3Idiots'.is_alnum()
true
%> 'Three Idiots'.is_alnum()
false
%> '3 Idiots'.is_alnum()
false
%> '3'.is_alnum()
true
%> 'idiots'.is_alnum()
true
%> ''.is_alnum()
false
Returns boolean
is_number()
is_number()
Returns true if all the characters in the string are all digits and
the string is not empty, otherwise returns false.
For example:
%> '123.5'.is_number()
false
%> '1970'.is_number()
true
%> '1980s'.is_number()
false
is_lower()
is_lower() -> boolean
Returns true if at least one character in the string is cased, all
cased characters are lower cased and the string is not empty. Otherwise,
it returns false.
For example:
%> 'all'.is_lower()
true
%> 'all...123'.is_lower()
true
%> 'All...123'.is_lower()
false
%> ''.is_lower()
false
Returns boolean
is_upper()
is_upper() -> boolean
Returns true if at least one character in the string is cased, all
cased characters are upper cased and the string is not empty. Otherwise,
it returns false.
For example:
%> 'ALL'.is_upper()
true
%> 'ALL...123'.is_upper()
true
%> 'All...123'.is_upper()
false
%> ''.is_upper()
false
Returns boolean
is_space()
is_space() -> boolean
Returns true if there are only whitespace characters in the string and
the string is not empty. Otherwise, it returns empty.
For example:
%> '. '.is_space()
false
%> '\r\n'.is_space()
true
%> '\t '.is_space()
true
Returns boolean
ord()
ord() -> number
Returns the Unicode code point of the string, which must be exactly one character long.
%> 'A'.ord()
65
%> 'AB'.ord()
Unhandled Error: ord() must be called on a single character, got AB
StackTrace:
<repl>:1 -> @.script()
Returns number
Raises Error if the string is not exactly one character long.
trim()
trim(chars: ?string) -> string
Returns a copy of the string with characters stripped from both ends.
With no argument, whitespace is stripped: space, tab (\t), line feed
(\n), vertical tab, form feed and carriage return (\r). Other
Unicode spaces, such as a no-break space, are kept.
Given chars, every character in it is stripped instead, in any order
and any number of times, until a character not in chars is reached
at each end. chars is a set of characters, not a prefix or suffix:
'xyax'.trim('xy') is 'a'. An empty chars strips nothing.
The string itself is never changed. A string with nothing to strip comes back as an equal copy.
For example:
%> ' example '.trim()
'example'
%> '\t example \r\n'.trim()
'example'
%> ' example '.trim('e')
' example '
%> 'example'.trim('e')
'xampl'
%> '--==example==--'.trim('-=')
'example'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
ltrim()
ltrim(chars: ?string) -> string
Returns a copy of the string with characters stripped from its start
only. The characters stripped are chosen exactly as they are for
trim(): whitespace when chars is not given, or every character of
chars when it is.
For example:
%> ' example '.ltrim()
'example '
%> 'example'.ltrim('e')
'xample'
%> '0012'.ltrim('0')
'12'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
rtrim()
rtrim(chars: ?string) -> string
Returns a copy of the string with characters stripped from its end only.
The characters stripped are chosen exactly as they are for trim():
whitespace when chars is not given, or every character of chars
when it is.
For example:
%> ' example '.rtrim()
' example'
%> 'example'.rtrim('e')
'exampl'
%> 'line\r\n'.rtrim('\r\n')
'line'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
join()
join(string: string) -> string
Returns a string which is a concatenation of the items in the iterable using the string as the separator. If the iterable contains just one item or the string is empty, the original element is returned. If the iterable contains non-string items, the items are converted to their string representation before joining.
Bytes are the only non supported iterables.
For example:
%> ','.join(['ok', 1, true])
'ok,1,true'
%> '--'.join('name')
'n--a--m--e'
%> ','.join('a')
'a'
Parameters
string(string) — The string to join the items in the iterable.
Returns string
split()
split(delimiter: string) -> list
Returns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found.
If the delimiter is an empty string, the resultant list will contain the individual characters of the string in the order in which they appear in the original string. Consecutive delimiters are not grouped together and are deemed to delimit empty strings. Splitting an empty string with a specified separator returns an empty list.
This method has full UTF-8 support.
For example:
%> 'name'.split('')
[n, a, m, e]
%> '1<>2<>3'.split('<>')
[1, , 2, , 3]
%> '1,2,3'.split(',')
[1, 2, 3]
%> ''.split(',')
[]
%> '地点'.split('')
[地, 点]
%> 'who is in the garden'.split('/\s/')
[who, is, in, the, garden]
Parameters
delimiter(string) — The delimiter to use the split the string.
Returns list
index_of()
index_of(str: string, start_index: ?number) -> number
Returns the index position of the first occurrence of the string str
in the string string. If the str cannot be found anywhere in
string, it returns -1. If the start_index parameter is given, it
will start scanning from the given index.
For example:
%> 'hello, world'.index_of(' ')
6
%> 'hello, world'.index_of('e')
1
%> 'hello, world'.index_of('q')
-1
%> 'hello, world'.index_of('o')
4
%> 'hello, world'.index_of('o', 5) # next index of `o` starting from index 5.
8
Parameters
str(string) — The string to search for.start_index(?number) — The index to start the search from.
Returns number
last_index_of()
last_index_of(str: string, end_index: ?number) -> number
Returns the index position of the last occurrence of the string str
in the string string, searching from the end. If str cannot be
found anywhere in string, it returns -1.
If the end_index parameter is given, only a match that begins at or
before that index counts. That is the same thing index_of()’s own
second parameter bounds, so for any index n, index_of(str, n) and
last_index_of(str, n) are the first and last matches of the two halves
n splits the string into.
An empty str returns -1, matching index_of().
For example:
%> 'hello, world'.last_index_of('o')
8
%> 'hello, world'.last_index_of('l')
10
%> 'hello, world'.last_index_of('q')
-1
%> 'hello, world'.last_index_of('o', 7) # last `o` starting at or before index 7.
4
Splitting a path on its final separator is the usual reason to reach for it:
%> var path = 'a/b/c'
%> path.last_index_of('/')
3
%> path[path.last_index_of('/') + 1, path.length()]
'c'
Parameters
str(string) — The string to search for.end_index(?number) — The highest index a match may start at.
Returns number
starts_with()
starts_with(str: string) -> boolean
Returns true if the string begins with the string or character
specified in str, otherwise it returns false.
For example:
%> 'hello, world'.starts_with('hello')
true
%> 'hello, world'.starts_with('hellios')
false
Parameters
str(string) — The string to search for.
Returns boolean
ends_with()
ends_with(str: string) -> boolean
Returns true if the string ends with the string or character specified
in str, otherwise it returns false.
For example:
%> 'gumtree'.ends_with('tree')
true
%> 'gumtree'.ends_with('mree')
false
Parameters
str(string) — The string to search for.
Returns boolean
count()
count(str: string) -> number
Returns the number of non-overlapping occurrences of the substring str in the string.
For those coming from Python who may consider this method similar to Python’s own, this method differs in that it does not allow specifying a start and end region for the operation. Zuri considers this unnecessary as the same can be accomplished by slicing the string.
For example:
%> 'Hallelujah'.count('l')
3
%> 'ding dong'.count('ng')
2
%> 'ding dong'[2,7].count('ng') # setting region to search for counts - 'ng do'
1
Parameters
str(string) — The string to search for.
Returns number
to_number()
to_number(base) -> number
Returns the first numeric value contained in the string if any exists or
0 if the string contains no numeric value. Floating numbers that have
the same value as their integer counterparts will return the integer
value.
For example:
%> '123.0 hell'.to_number()
123
%> '427 and 12'.to_number()
427
%> '96.3 of 31'.to_number()
96.3
%> 'error'.to_number()
0
Parameters
base(number) — The base the digits are in, from 2 to 36. Defaults to10. A fractional part is only read in base 10, since no other base spells one.
Returns number
Raises RangeError if base is outside 2 to 36.
to_bigint()
to_bigint(base) -> bigint
Returns the integer value of the string as a bigint, or 0n if the
string does not spell one.
This is to_number() for integers too large to be a number. A number is
exact only up to 2^53; past that, digits are lost, and an id or a
BIGINT UNSIGNED read from a database routinely runs past it. Every
digit survives here however long the run.
%> '9007199254740993'.to_bigint()
9007199254740993n
%> '9007199254740993'.to_number() # rounded down by one
9007199254740992
%> '-42'.to_bigint()
-42n
%> 'ff'.to_bigint(16)
255n
%> 'row 427 of 12'.to_bigint()
427n
%> 'error'.to_bigint()
0n
The number is found exactly as to_number() finds it: the first one
written in the string, with any text around it ignored. No fractional
part is read, since this produces an integer, so '12.5'.to_bigint() is
12n.
Parameters
base(number) — The base the digits are in, from 2 to 36. Defaults to10.
Returns bigint
Raises RangeError if base is outside 2 to 36.
to_list()
to_list() -> list
Returns a list whose elements consists of every character contained in
the string in order of appearance. Characters that repeat in the string
will have different entries in the same index as they appear in the
string.
For example:
%> 'Zuri'.to_list()
[Z, u, r, i]
%> 'Plantation'.to_list()
[P, l, a, n, t, a, t, i, o, n]
Returns list
to_bytes()
to_bytes() -> bytes
Returns the content of the string as a stream of bytes.
The Zuri REPL may truncate long bytes data when printing to console/terminal.
For example:
%> 'Zuri'.to_bytes()
(42 6c 61 64 65)
%> 'Plantation'.to_bytes()
(50 6c 61 6e 74 61 74 69 6f 6e)
Returns bytes
lpad()
lpad(width: number, fill: ?string) -> string
Returns the string left justified in a string of length width. Padding
is done using the specified character fill if given of a space (' ')
if a fill is not specified. The original string is returned if width
is less than string.length().
For example:
%> 'cat'.lpad(5)
' cat'
%> 'cat'.lpad(5, '-')
'--cat'
%> 'cat'.lpad(2, '-')
'cat'
Parameters
width(number) — The length of the string after padding.fill(?string) — The character to use for padding.
Returns string
rpad()
rpad(width: number, fill: ?string) -> string
Returns the string right justified in a string of length width.
Padding is done using the specified character fill if given of a space
(' ') if a fill is not specified. The original string is returned if
width is less than string.length().
For example:
%> 'Hmm'.rpad(6)
'Hmm '
%> 'Hmm'.rpad(6, '.')
'Hmm...'
%> 'Hmm'.rpad(3, '.')
'Hmm'
Parameters
width(number) — The length of the string after padding.fill(?string) — The character to use for padding.
Returns string
match()
match(str: string) -> boolean|dictionary
If the string str is a regular string, this method returns true if
the string contains a substring str. Otherwise, it returns false.
If the string str contains a valid regular
expression (we’ll get to that shortly below), it
returns false if a match for the regex str cannot be found in the
string. Otherwise, it returns a dictionary of the
first match: the whole match under key 0, each capture group under its
number, and a named group under its name as well. A group that takes no
part in the match is nil. Groups that share a name, under the J
modifier, give the name to the one that took part.
If the offset argument is specified, it becomes the offset in the string at which to start matching.
For example:
%> 'gorilla'.match('go') # regular string match
true
%> 'gorilla'.match('gox') # regular string non-match
false
%> 'gorilla'.match('/?gox/') # regular expression match
{0: go}
%> 'gorilla'.match('/gox\d/') # regular expression non-match
false
%> '2024-01'.match('/(?<year>\d+)-(\d+)/')
{0: 2024-01, 1: 2024, 2: 01, year: 2024}
Parameters
str(string) — The string to match.
Returns boolean|dictionary
matches()
matches(reg: string) -> dictionary
Returns a dictionary containing every match of the given regular expression reg in the source string. If no match is found, an empty dictionary is returned.
If the offset argument is specified, it becomes the offset in the string at which to start matching.
For example:
%> '123 dollars'.matches('/[a-z]+|\d+/')
{0: [123, dollars]}
%> 'who is in the garden'.matches('/\w+/')
{0: [who, is, in, the, garden]}
Parameters
reg(string) — The regular expression to match.
Returns dictionary
replace()
replace(str: string, replacement: string, use_regex: ?bool) -> string
Returns a copy of the string with all occurrences or matches of str replaced by the replacement string.
In the replacement string, if str is a regular expression, then
capture groups can be referenced using the syntax $index. Taking as an
example, capture group 0 contains the entire match and can be used in
the replacement string as $0.
To escape the
$sign in the replacement string, use the double backslashes (\\).
For example:
%> 'lady friend'.replace('d', 'z') # non-regex
'lazy frienz'
%> 'John is 26 years old'.replace('/(\d+)/', '1$1') # regex example
'John is 126 years old'
%> 'John is 26 years old'.replace('/(\d+)/', '1\\$2')
'John is 1$2 years old'
Parameters
str(string) — The string to match.replacement(string) — The replacement string.use_regex(?bool) — Whether to use the regular expression or the string string as the match string (default = true).
Returns string
Note: When the third parameter
use_regexis set to false, str will never be treated as a regular expression even if it contains a valid regular expression.
replace_with()
replace_with(regex: string, callback: function) -> string
Returns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred.
The callback function is defined as follows:
def replacer(match, p1, p2, /* …, */ pN, offset, string) {
return replacement
}
The arguments to the function are as follows:
-
match: The matched substring. (Corresponds to$0.) -
p1, p2, …, pN: The nth string found by a capture group (including named capturing groups) corresponds to$1,$2, etc. For example, if the pattern is/(\a+)(\b+)/, thenp1is the match for\a+, andp2is the match for\b+. If the group is part of a disjunction (e.g."abc".replace_with('/(a)|(b)/', replacer)), the unmatched alternative will benil. -
offset: The offset of the matched substring within the whole string being examined. For example, if the whole string was'abcd', and the matched substring was'bc', then this argument will be1. -
string: The whole string being examined.
The exact number of arguments depends on how many capture groups are contained in the regex.
For example:
%> echo 'name'.replace_with('/m/', @(match, offset) {
.. return match + '-'
.. })
'nam-e'
Below is another example that uses a capture group:
%> var text = 'all is well'
%>
%> echo text.replace_with('/([a-z]+)/', @(match, val) {
.. if val == 'is' return 'is not'
.. return 'will be'
.. })
'will be is not will be'
Parameters
regex(string) — The regular expression to match.callback(function) — The callback function to invoke for each match.
Returns string
ascii()
ascii() -> string
Reinterprets the string as a raw byte view: each byte of its UTF-8
encoding becomes its own character (a codepoint between 0 and 255,
i.e. a Latin-1-style one-byte-per-character mapping), rather than the
decoded sequence of Unicode characters that length(), each(), and
indexing otherwise operate on.
The result is still a valid string (every codepoint between 0 and
255 is valid UTF-8), so it can be used anywhere a normal string can.
It just no longer round-trips back through the original multi-byte
characters if the string had any, and its length() now reports the
original BYTE count of the string rather than its original CHARACTER
count.
This is meant for the rare case where code needs to walk a string byte-for-byte instead of character-by-character, e.g. one that originated from a byte stream where the bytes were never meant to be decoded as Unicode at all.
%> 'café'.length()
4
%> 'café'.ascii().length()
5
Returns string
case_fold()
case_fold() -> string
Returns a copy of the string case-folded for case-insensitive
comparison, using full Unicode case folding rather than plain
lowercasing. This matters for characters whose fold is not just their
lowercase form: for example, the German ß folds to ss.
Two strings that are considered equal ignoring case will always produce
identical output from case_fold(), which makes it the correct method
to use for case-insensitive comparisons; lower() is not a substitute
for it.
%> 'HELLO World'.case_fold()
'hello world'
%> 'Straße'.case_fold()
'strasse'
Returns string
compare()
compare(other: string) -> number
Compares the string with another string.
Parameters
other(string) — The other string to compare with.
Returns number — A negative number if the string is less than the
other string. - Zero if the strings are equal. - A positive number if
the string is greater than the other string.
Raises Error if the other string is not a string.
is_empty()
is_empty() -> boolean
Returns true if the string is empty, false otherwise.
Returns boolean
contains()
contains(str: string) -> boolean
Returns true if the string contains the specified substring, false otherwise.
Parameters
str(string) — The substring to search for.
Returns boolean
lines()
lines() -> list
Returns the lines of the string as an list as it would be if split on newline characters.
Returns list
each_line()
each_line(callback: function) -> void
Iterates over each line of the string, calling the provided callback function with the line and its index.
Parameters
callback(function) — A function that takes two arguments: the line and its index.
Returns void
Raises Error if the callback is not a function.
each()
each(callback: function) -> void
Iterates over each character of the string, calling the provided callback function with the character and its index.
Parameters
callback(function) — A function that takes two arguments: the character and its index.
Returns void
Raises Error if the callback is not a function.
capitalize()
capitalize() -> string
Returns a new string with the first character capitalized and the rest in lowercase.
Returns string
title()
title() -> string
Returns a new string with each word capitalized.
Returns string
to_string()
to_string() -> string
Returns the string itself.
Returns string
Number Methods
Every method on the built-in number type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
to_string() | string | Returns the string representation of the number. |
to_bool() | boolean | Converts the number to a boolean, by the same rule if uses. |
to_bigint() | bigint | Converts the number to a bigint, the counterpart to bigint.to_number(). |
abs() | number | Returns the absolute value of the number. |
chr() | string | Returns the Unicode character whose code point is equal to the number. |
bin() | string | Converts the number to its binary string representation. |
hex() | string | Converts the number to its hexadecimal string representation. |
oct() | string | Converts the number to its octal string representation. |
int() | number | Truncates the number down to its integer part, discarding anything after the decimal point. |
max(other: number) | number | Returns the larger of the number and other. |
min(other: number) | number | Returns the smaller of the number and other. |
factorial() | number | Returns the factorial of the number, i.e. |
sin() | number | Returns the sine of the number, taken to be in radians. |
cos() | number | Returns the cosine of the number, taken to be in radians. |
tan() | number | Returns the tangent of the number, taken to be in radians. |
sinh() | number | Returns the hyperbolic sine of the number. |
cosh() | number | Returns the hyperbolic cosine of the number. |
tanh() | number | Returns the hyperbolic tangent of the number. |
asin() | number | Returns the arcsine (inverse sine) of the number, in radians. |
acos() | number | Returns the arccosine (inverse cosine) of the number, in radians. |
atan() | number | Returns the arctangent (inverse tangent) of the number, in radians. |
atan2(x: number) | number | Returns the four-quadrant arctangent of the number and x, in radians. |
asinh() | number | Returns the inverse hyperbolic sine of the number. |
acosh() | number | Returns the inverse hyperbolic cosine of the number. |
atanh() | number | Returns the inverse hyperbolic tangent of the number. |
exp() | number | Returns e (Euler’s number) raised to the power of the number. |
expm1() | number | Returns e raised to the power of the number, minus 1. |
log() | number | Returns the natural logarithm (base e) of the number. |
log2() | number | Returns the base-2 logarithm of the number. |
log10() | number | Returns the base-10 logarithm of the number. |
log1p() | number | Returns the natural logarithm of 1 plus the number. |
cbrt() | number | Returns the cube root of the number. |
sqrt() | number | Returns the square root of the number. |
sign() | number | Returns the sign of the number: 1 if it is positive, -1 if it is negative, and 0 (with its own original sign preserved) if it is zero. |
ceil() | number | Returns the smallest whole number greater than or equal to the number. |
round() | number | Rounds the number to the nearest whole number. |
floor() | number | Returns the largest whole number less than or equal to the number. |
is_nan() | boolean | Returns true if the number is NaN (not a number, e.g. |
is_inf() | boolean | Returns true if the number is positive or negative infinity, false otherwise. |
is_finite() | boolean | Returns true if the number is neither infinite nor NaN, false otherwise. |
trunc() | number | Truncates the number towards zero, discarding anything after the decimal point. |
fraction() | number | Returns the digits after the number’s decimal point, read as a whole number rather than a fraction. |
fixed(n) | Returns the number rounded to n decimal places, with a half rounding away from zero the same way round() does. |
to_string()
to_string() -> string
Returns the string representation of the number.
%> 5.to_string()
'5'
Returns string
to_bool()
to_bool() -> boolean
Converts the number to a boolean, by the same rule if uses. Zero
(either sign) and NaN are false, and every other number, negative
ones included, is true.
%> 5.to_bool()
true
%> 0.to_bool()
false
%> (-5).to_bool()
true
%> (0 / 0).to_bool()
false
Returns boolean
to_bigint()
to_bigint() -> bigint
Converts the number to a bigint, the counterpart to
bigint.to_number().
%> 12345.to_bigint()
12345n
%> 2.to_bigint() ** 100.to_bigint()
1267650600228229401496703205376n
Returns bigint
Raises RangeError if the number is not an exact integer.
Note: Only an exact integer has a
bigintform, so a fractional number,InfinityandNaNall raise rather than being rounded or clamped. Numbers above2^53are already imprecise as doubles, so converting one yields the exact integer the double holds, not the decimal literal it was written as.
abs()
abs() -> number
Returns the absolute value of the number.
%> (-5).abs()
5
%> 5.abs()
5
Returns number
chr()
chr() -> string
Returns the Unicode character whose code point is equal to the number.
%> 65.chr()
'A'
Returns string
bin()
bin() -> string
Converts the number to its binary string representation. The number is truncated to an integer first.
%> 10.bin()
'1010'
Returns string
Note: A negative number always returns
'0'; there is no signed or two’s-complement form.
hex()
hex() -> string
Converts the number to its hexadecimal string representation. The number is truncated to an integer first.
%> 255.hex()
'ff'
Returns string
Note: A negative number always returns
'0'; there is no signed or two’s-complement form.
oct()
oct() -> string
Converts the number to its octal string representation. The number is truncated to an integer first.
%> 8.oct()
'10'
Returns string
Note: A negative number always returns
'0'; there is no signed or two’s-complement form.
int()
int() -> number
Truncates the number down to its integer part, discarding anything after
the decimal point. Unlike floor(), this rounds towards zero rather
than towards negative infinity, so the result for a negative number
differs from floor().
%> 3.9.int()
3
%> (-3.9).int()
-3
Returns number
max()
max(other: number) -> number
Returns the larger of the number and other.
A NaN never wins: when one of the two is NaN, the other is returned,
and only two NaNs give NaN. -0 counts as smaller than 0, so
(-0.0).max(0) is 0 whichever side the zeros are on.
%> 5.max(9)
9
%> 9.max(5)
9
%> (0 / 0).max(3)
3
%> (-0.0).max(0)
0
Parameters
other(number) — The number to compare against.
Returns number
min()
min(other: number) -> number
Returns the smaller of the number and other.
A NaN never wins: when one of the two is NaN, the other is returned,
and only two NaNs give NaN. -0 counts as smaller than 0, so
0.min(-0.0) is -0 whichever side the zeros are on.
%> 5.min(9)
5
%> 9.min(5)
5
%> (0 / 0).min(3)
3
%> 0.min(-0.0)
-0
Parameters
other(number) — The number to compare against.
Returns number
factorial()
factorial() -> number
Returns the factorial of the number, i.e. the product of every positive
integer less than or equal to it. 0.factorial() is 1, matching the
standard mathematical definition.
%> 5.factorial()
120
%> 0.factorial()
1
Returns number
Raises Error if the number is negative or not a whole number.
sin()
sin() -> number
Returns the sine of the number, taken to be in radians.
Returns number
cos()
cos() -> number
Returns the cosine of the number, taken to be in radians.
Returns number
tan()
tan() -> number
Returns the tangent of the number, taken to be in radians.
Returns number
sinh()
sinh() -> number
Returns the hyperbolic sine of the number.
Returns number
cosh()
cosh() -> number
Returns the hyperbolic cosine of the number.
Returns number
tanh()
tanh() -> number
Returns the hyperbolic tangent of the number.
Returns number
asin()
asin() -> number
Returns the arcsine (inverse sine) of the number, in radians.
Returns number
Note: Only defined for a receiver between
-1and1inclusive; outside that range, this returnsNaNrather than raising an error.
acos()
acos() -> number
Returns the arccosine (inverse cosine) of the number, in radians.
Returns number
Note: Only defined for a receiver between
-1and1inclusive; outside that range, this returnsNaNrather than raising an error.
atan()
atan() -> number
Returns the arctangent (inverse tangent) of the number, in radians.
Returns number
atan2()
atan2(x: number) -> number
Returns the four-quadrant arctangent of the number and x, in radians.
The receiver is treated as the y-coordinate and x as the x-coordinate,
matching the conventional atan2(y, x) signature: y.atan2(x).
%> 1.0.atan2(1.0)
0.7853981633974483
Parameters
x(number) — The x-coordinate.
Returns number
asinh()
asinh() -> number
Returns the inverse hyperbolic sine of the number.
Returns number
acosh()
acosh() -> number
Returns the inverse hyperbolic cosine of the number.
Returns number
Note: Only defined for a receiver greater than or equal to
1; below that, this returnsNaNrather than raising an error.
atanh()
atanh() -> number
Returns the inverse hyperbolic tangent of the number.
Returns number
Note: Only defined for a receiver between
-1and1exclusive; outside that range, this returnsNaNrather than raising an error.
exp()
exp() -> number
Returns e (Euler’s number) raised to the power of the number.
Returns number
expm1()
expm1() -> number
Returns e raised to the power of the number, minus 1. For a number
close to zero, this is more numerically accurate than computing n.exp() - 1
directly.
Returns number
log()
log() -> number
Returns the natural logarithm (base e) of the number.
%> 1.0.log()
0
Returns number
log2()
log2() -> number
Returns the base-2 logarithm of the number.
%> 8.0.log2()
3
Returns number
log10()
log10() -> number
Returns the base-10 logarithm of the number.
%> 100.0.log10()
2
Returns number
log1p()
log1p() -> number
Returns the natural logarithm of 1 plus the number. For a number close
to zero, this is more numerically accurate than computing (1 + n).log()
directly.
Returns number
cbrt()
cbrt() -> number
Returns the cube root of the number.
%> 27.0.cbrt()
3
Returns number
sqrt()
sqrt() -> number
Returns the square root of the number.
%> 16.sqrt()
4
Returns number
Note: For a negative number this returns
NaNrather than raising an error; there is nobigint-style promotion into complex numbers. Checkis_nan()on the result, or the sign of the receiver beforehand, if that distinction matters to the caller.
sign()
sign() -> number
Returns the sign of the number: 1 if it is positive, -1 if it is
negative, and 0 (with its own original sign preserved) if it is zero.
%> 7.sign()
1
%> (-7).sign()
-1
%> 0.sign()
0
Returns number
ceil()
ceil() -> number
Returns the smallest whole number greater than or equal to the number.
%> 3.14159.ceil()
4
Returns number
round()
round() -> number
Rounds the number to the nearest whole number. A value exactly halfway between two whole numbers rounds away from zero.
%> 3.14159.round()
3
%> 3.6.round()
4
Returns number
floor()
floor() -> number
Returns the largest whole number less than or equal to the number.
%> 3.14159.floor()
3
Returns number
is_nan()
is_nan() -> boolean
Returns true if the number is NaN (not a number, e.g. the result of
0/0), false otherwise.
Returns boolean
is_inf()
is_inf() -> boolean
Returns true if the number is positive or negative infinity, false
otherwise.
Returns boolean
is_finite()
is_finite() -> boolean
Returns true if the number is neither infinite nor NaN, false
otherwise.
Returns boolean
trunc()
trunc() -> number
Truncates the number towards zero, discarding anything after the decimal
point. For values that fit in a 64-bit integer this matches int();
unlike int(), trunc() stays a floating-point result rather than
going through an integer cast, so it does not overflow for numbers
larger than a 64-bit integer can hold.
%> (-3.9).trunc()
-3
Returns number
fraction()
fraction() -> number
Returns the digits after the number’s decimal point, read as a whole
number rather than a fraction. Note that this is NOT the same as (n - n.int()):
1.92.fraction() is 92, not 0.92.
%> 1.92.fraction()
92
%> 1.5.fraction()
5
%> 5.fraction()
0
Returns number
fixed()
fixed(n)
Returns the number rounded to n decimal places, with a half rounding
away from zero the same way round() does.
A number already shorter than n places is returned unchanged, and so
are NaN and the infinities. Beyond 17 places an f64 has no digits
left to round, so a larger n behaves as 17.
%> 1.554576852757686786786.fixed(9)
1.554576853
%> 1.554576852757686786786.fixed(8)
1.55457685
%> 1.554576852757686786786.fixed(1)
1.6
%> 1.554576852757686786786.fixed(0)
2
%> (-2.5).fixed(0)
-3
Bigint Methods
Every method on the built-in bigint type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
to_string(radix) | string | Returns the decimal digits of the bigint, with a leading - when it is negative and no trailing n. |
to_number() | number | Converts the bigint to a number. |
to_bool() | boolean | Converts the bigint to a boolean, by the same rule if uses. |
to_bytes(order) | bytes | Returns the two’s-complement byte representation, which carries the sign and so round-trips back to the same value. |
bin() | string | Returns the base-2 digits, equivalent to to_string(2). |
hex() | string | Returns the base-16 digits in lowercase, equivalent to to_string(16). |
oct() | string | Returns the base-8 digits, equivalent to to_string(8). |
abs() | bigint | Returns the absolute value. |
sign() | number | Returns the sign as a plain number: 1 when positive, -1 when negative and 0 for zero. |
max(other) | bigint | Returns the larger of the two bigints. |
min(other) | bigint | Returns the smaller of the two bigints. |
pow(exponent) | bigint | Raises the bigint to exponent, the method form of **. |
sqrt() | bigint | Returns the integer square root, truncated towards zero, so 145n.sqrt() is 12n rather than 12.04.... |
cbrt() | bigint | Returns the integer cube root, truncated towards zero. |
nth_root(n) | bigint | Returns the integer nth root, truncated towards zero. |
gcd(other) | bigint | Returns the greatest common divisor of the two bigints. |
lcm(other) | bigint | Returns the least common multiple of the two bigints. |
modpow(exponent, modulus) | bigint | Returns (self ** exponent) % modulus without ever building the full power, which is what makes it usable for the huge exponents cryptography needs. |
modinv(modulus) | bigint|nil | Returns the modular multiplicative inverse: the x solving self * x == 1 (mod modulus). |
bits() | number | Returns how many bits the magnitude occupies, ignoring the sign. |
bit(index) | boolean | Returns whether the bit at index is set, counting from the least significant bit at index 0. |
set_bit(index, value) | bigint | Returns a new bigint with the bit at index set or cleared. |
trailing_zeros() | number|nil | Returns the count of least-significant zero bits, which is the largest power of two dividing the bigint. |
is_zero() | boolean | Returns whether the bigint is zero. |
is_even() | boolean | Returns whether the bigint is even. |
is_odd() | boolean | Returns whether the bigint is odd. |
to_string()
to_string(radix) -> string
Returns the decimal digits of the bigint, with a leading - when it is
negative and no trailing n. Pass radix to render in another base
instead, using lowercase letters for digit values above nine.
%> 255n.to_string()
'255'
%> 255n.to_string(16)
'ff'
%> (-255n).to_string(16)
'-ff'
Parameters
radix(number) — base to render in, from 2 to 36. Defaults to 10.
Returns string
Raises RangeError if radix is outside 2 to 36.
Note: The
nthatechoand string interpolation show is part of the repr, not the conversion;to_string()never includes it.
to_number()
to_number() -> number
Converts the bigint to a number.
%> 6n.to_number()
6
Returns number
Note: This is lossy for anything past
2^53: the result is the nearest double, and a value past the double range becomesinfor-infrather than wrapping or reading as zero. Checkbits()beforehand when that matters.
to_bool()
to_bool() -> boolean
Converts the bigint to a boolean, by the same rule if uses. 0n is
false, and every other bigint, negative ones included, is true.
%> 5n.to_bool()
true
%> 0n.to_bool()
false
%> (-5n).to_bool()
true
Returns boolean
to_bytes()
to_bytes(order) -> bytes
Returns the two’s-complement byte representation, which carries the sign
and so round-trips back to the same value. The result is the shortest
byte string that can hold it, and is never empty: zero is a single
0x00 byte.
%> 258n.to_bytes()
(01 02)
%> 258n.to_bytes('little')
(02 01)
%> (-1n).to_bytes()
(ff)
Parameters
order(string) —'big'or'little'. Defaults to'big'.
Returns bytes
Raises RangeError if order is neither 'big' nor 'little'.
bin()
bin() -> string
Returns the base-2 digits, equivalent to to_string(2).
%> 255n.bin()
'11111111'
Returns string
Note: This is the sign-and-magnitude form, so a negative bigint comes back with a leading
-rather than as two’s complement. Useto_bytes()for the two’s-complement view.
hex()
hex() -> string
Returns the base-16 digits in lowercase, equivalent to to_string(16).
%> 255n.hex()
'ff'
Returns string
oct()
oct() -> string
Returns the base-8 digits, equivalent to to_string(8).
%> 255n.oct()
'377'
Returns string
abs()
abs() -> bigint
Returns the absolute value.
%> (-5n).abs()
5n
Returns bigint
sign()
sign() -> number
Returns the sign as a plain number: 1 when positive, -1 when
negative and 0 for zero.
%> (-9n).sign()
-1
%> 0n.sign()
0
Returns number
max()
max(other) -> bigint
Returns the larger of the two bigints.
%> 3n.max(7n)
7n
Parameters
other(bigint)
Returns bigint
Raises TypeError if other is not a bigint.
min()
min(other) -> bigint
Returns the smaller of the two bigints.
%> 3n.min(7n)
3n
Parameters
other(bigint)
Returns bigint
Raises TypeError if other is not a bigint.
pow()
pow(exponent) -> bigint
Raises the bigint to exponent, the method form of **.
%> 2n.pow(100)
1267650600228229401496703205376n
Parameters
exponent(number|bigint) — a integer from 0 to 2^32 - 1. Negative exponents have no integral answer and are rejected rather than truncated to zero.
Returns bigint
Raises RangeError if exponent is negative, fractional or too
large.
sqrt()
sqrt() -> bigint
Returns the integer square root, truncated towards zero, so
145n.sqrt() is 12n rather than 12.04....
%> 144n.sqrt()
12n
%> 145n.sqrt()
12n
Returns bigint
Raises RangeError if the bigint is negative.
cbrt()
cbrt() -> bigint
Returns the integer cube root, truncated towards zero. Negatives are
fine here, unlike sqrt().
%> (-27n).cbrt()
-3n
Returns bigint
nth_root()
nth_root(n) -> bigint
Returns the integer nth root, truncated towards zero.
%> 1000000n.nth_root(3)
100n
Parameters
n(number|bigint) — a integer from 1 to 2^32 - 1.
Returns bigint
Raises RangeError if n is zero, negative, fractional or too
large, or if n is even and the bigint is negative.
gcd()
gcd(other) -> bigint
Returns the greatest common divisor of the two bigints. The result is
always non-negative regardless of either sign, and 0n.gcd(0n) is 0n.
%> 48n.gcd(18n)
6n
Parameters
other(bigint)
Returns bigint
Raises TypeError if other is not a bigint.
lcm()
lcm(other) -> bigint
Returns the least common multiple of the two bigints. The result is
always non-negative, and is 0n when either side is zero.
%> 48n.lcm(18n)
144n
Parameters
other(bigint)
Returns bigint
Raises TypeError if other is not a bigint.
modpow()
modpow(exponent, modulus) -> bigint
Returns (self ** exponent) % modulus without ever building the full
power, which is what makes it usable for the huge exponents cryptography
needs.
%> 4n.modpow(13n, 497n)
445n
Parameters
exponent(bigint)modulus(bigint) — must not be zero.
Returns bigint
Raises TypeError if either argument is not a bigint.
Raises RangeError if modulus is zero, or if exponent is
negative and no modular inverse exists.
Note: The remainder is floored rather than truncated, so the result carries the sign of
modulus, not of the receiver. A negativeexponentis allowed only when the receiver is invertible modulomodulus.
modinv()
modinv(modulus) -> bigint|nil
Returns the modular multiplicative inverse: the x solving self * x == 1 (mod modulus).
%> 3n.modinv(11n)
4n
%> 4n.modinv(8n)
nil
Parameters
modulus(bigint) — must not be zero.
Returns bigint|nil
Raises TypeError if modulus is not a bigint.
Raises RangeError if modulus is zero.
Note: Returns
nilrather than raising when the receiver andmodulusare not coprime, since having no inverse is an ordinary answer and not a caller mistake. The result carries the sign ofmodulus.
bits()
bits() -> number
Returns how many bits the magnitude occupies, ignoring the sign. Zero occupies none.
%> 255n.bits()
8
%> 0n.bits()
0
Returns number
bit()
bit(index) -> boolean
Returns whether the bit at index is set, counting from the least
significant bit at index 0.
%> 5n.bit(0)
true
%> 5n.bit(1)
false
Parameters
index(number) — a non-negative integer.
Returns boolean
Raises RangeError if index is negative or fractional.
Note: The bigint is read as two’s complement, so a negative receiver reports
truefor every index above its magnitude rather than running out of bits.
set_bit()
set_bit(index, value) -> bigint
Returns a new bigint with the bit at index set or cleared. The
receiver is left untouched.
%> 5n.set_bit(1, true)
7n
Parameters
index(number) — a non-negative integer.value(boolean)
Returns bigint
Raises RangeError if index is negative or fractional.
trailing_zeros()
trailing_zeros() -> number|nil
Returns the count of least-significant zero bits, which is the largest power of two dividing the bigint.
%> 40n.trailing_zeros()
3
%> 0n.trailing_zeros()
nil
Returns number|nil
Note: Returns
nilfor zero, which has no largest such power and would otherwise have to report an arbitrary number.
is_zero()
is_zero() -> boolean
Returns whether the bigint is zero.
%> 0n.is_zero()
true
Returns boolean
is_even()
is_even() -> boolean
Returns whether the bigint is even. Zero is even.
%> 4n.is_even()
true
Returns boolean
is_odd()
is_odd() -> boolean
Returns whether the bigint is odd.
%> 5n.is_odd()
true
Returns boolean
Boolean Methods
Every method on the built-in bool type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
to_string() | string | Returns the string representation of the boolean. |
to_string()
to_string() -> string
Returns the string representation of the boolean.
%> true.to_string()
'true'
%> false.to_string()
'false'
Returns string
List Methods
Every method on the built-in list type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
length() | number | Returns the number of items in the list. |
append(value) | list | Adds the given value x to the end of the list. |
clear() | Removes all items from the list. | |
clone() | list | Returns a new list containing all items from the list. |
count(value) | number | Returns the number of times item x occurs in the list. |
extend(list: list) | list | Updates the content of the list by appending all the contents of list x to the end of the original list in exact order. |
index_of(value, start_index: ?int) | number | Returns the zero-based index of the first occurrence of the value x in the list starting from the given start_index or -1 if the list does not contain the value x. |
last_index_of(value, end_index: ?int) | number | Returns the zero-based index of the last occurrence of the value x in the list, searching from the end, or -1 if the list does not contain the value x. |
insert(value, index: int) | list | Inserts the item x into the list at the specified index. |
pop() | any | Removes the last item in a list and returns the value of that item. |
shift(count: ?int) | any | Removed the specified count of items from the beginning of the list and returns it. |
remove_at(index: int) | any | Removes the item at the specified index in the list and returns it. |
remove(value) | any | Removes the first occurrence of item x from the list. |
reverse() | list | Returns a new list containing the items in the original list in reverse order. |
sort(comparator: ?function) | list | Sorts the items in the list in-place and returns the sorted list. |
contains(value) | boolean | Returns true if the list contains the item x or false otherwise. |
delete(start: int, end: int) | number | Deletes a range of items from the list starting from the start to the end limit and returns the number of items removed. |
first() | any | Returns the first item in the list or nil if the list is empty. |
last() | any | Returns the last item in the list or nil if the list is empty. |
is_empty() | boolean | Returns true if the list is empty or false otherwise. |
take(n: int) | list | Returns a new list containing the first n items in the list or a new copy of the list if n greater than or equals to the list.length(). |
get(index: int) | any | Returns the value at the specified index in the list. |
compact() | list | Returns a new list containing the items in the original list but with all nil values removed. |
unique() | list | Returns a new list containing the unique values from the original list. |
zip(...lists: list) | list | Returns a list that contains the items in the original list merged with corresponding items from the individual arguments. |
zip_from(list: list) | list | The same as list.zip() except that instead of accepting an arbitrary list or arguments, it accepts a single list that should contain other lists. |
to_dict() | dict | Returns a number indexed dictionary representing the list. |
each(callback: function) | Iterates over each element in the list, calling the provided callback function with the current element as an argument. | |
map(callback: function) | list | Creates a new list populated with the results of calling a provided function on every element in the calling list. |
filter(callback: function) | list | Creates a new list with all elements that pass the test implemented by the provided function. |
reduce(callback: function, initial) | any | Applies a function against an accumulator and each element in the list (from left to right) to reduce it to a single value and returns the accumulated result of the callback function. |
some(callback: function) | boolean | Tests whether at least one element in the list passes the test implemented by the provided function. |
every(callback: function) | boolean | Tests whether all elements in the list pass the test implemented by the provided function. |
find(callback: function) | any | Returns the value of the first element in the list that satisfies the provided testing function. |
find_index(callback: function) | number | Returns the index of the first element in the list that satisfies the provided testing function. |
find_last(callback: function) | any | Returns the value of the last element in the list that satisfies the provided testing function. |
find_last_index(callback: function) | number | Returns the index of the last element in the list that satisfies the provided testing function. |
find_all(callback: function) | list | Returns a new list containing all elements of the calling list that satisfy the provided testing function. |
partition(callback: function) | list | Returns an list containing two lists: the first with elements that satisfy the provided testing function, and the second with elements that do not satisfy the testing function. |
to_string() | string | Returns the string representation of the list. |
length()
length() -> number
Returns the number of items in the list.
For example:
%> ['A', 'B', 'C'].length()
3
Returns number
append()
append(value) -> list
Adds the given value x to the end of the list.
For example:
%> var a = [1,2,3]
%> a.append(4)
%> a
[1, 2, 3, 4]
Parameters
value(any)
Returns list
clear()
clear()
Removes all items from the list.
For example:
%> var a = [1,2,3,4,5]
%> a
[1, 2, 3, 4, 5]
%> a.clear()
%> a
[]
clone()
clone() -> list
Returns a new list containing all items from the list. The new list is
a shallow copy of the original list. This is equivalent to list[,].
For example:
%> var a = [1, 2, 3]
%> var b = a.clone()
%> a.append(4)
%> a
[1, 2, 3, 4]
%> b
[1, 2, 3]
Returns list
count()
count(value) -> number
Returns the number of times item x occurs in the list.
For example:
%> [1, 2, 1, 3, 2, 1, 1].count(1)
4
Parameters
value(any)
Returns number
extend()
extend(list: list) -> list
Updates the content of the list by appending all the contents of list
x to the end of the original list in exact order. This is equivalent
to list + x.
For example:
%> var a = [1, 2, 3]
%> var b = [4, 5, 6]
%> a.extend(b)
%> a
[1, 2, 3, 4, 5, 6]
%> b
[4, 5, 6]
Parameters
list(list)
Returns list
index_of()
index_of(value, start_index: ?int) -> number
Returns the zero-based index of the first occurrence of the value x in
the list starting from the given start_index or -1 if the list does
not contain the value x.
For example:
%> [1,2].index_of(3)
-1
%> [4,5,6,5].index_of(5)
1
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].index_of('a')
0
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].index_of('a', 1)
3
Parameters
value(any)start_index(?int)
Returns number
last_index_of()
last_index_of(value, end_index: ?int) -> number
Returns the zero-based index of the last occurrence of the value x in
the list, searching from the end, or -1 if the list does not contain
the value x.
If end_index is given, only a match at or before that index counts.
That is the same position index_of()’s own second parameter bounds, so
for any index n, index_of(x, n) and last_index_of(x, n) are the
first and last matches of the two halves n splits the list into.
Values are compared the way index_of() compares them, by value rather
than by identity, so two separate dictionaries holding the same entries
match each other.
For example:
%> [1,2].last_index_of(3)
-1
%> [4,5,6,5].last_index_of(5)
3
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].last_index_of('a')
5
%> ['a', 'b', 'r', 'a', 'h', 'a', 'm'].last_index_of('a', 4)
3
Parameters
value(any)end_index(?int)
Returns number
insert()
insert(value, index: int) -> list
Inserts the item x into the list at the specified index. By
specifying an index of zero (list.insert(x, 0)), one can prepend the
list and list.insert(x, list.length()) is equivalent to
list.append(x). If the index specified is greater than
list.length(), the list will be padded with nil up till the index
preceding the specified index.
For example:
%> var a = [1,2,3]
%> a.insert(4, 0)
%> a
[4, 1, 2, 3]
%> a.insert(5, a.length())
%> a
[4, 1, 2, 3, 5]
%> a.insert(6, 3)
%> a
[4, 1, 2, 6, 3, 5]
%> a.insert(7, 11)
%> a
[4, 1, 2, 6, 3, 5, nil, nil, nil, nil, nil, 7]
Parameters
value(any)index(int)
Returns list
pop()
pop() -> any
Removes the last item in a list and returns the value of that item.
For example:
%> var a = [4, 5, 6]
%> a.pop()
6
%> a
[4, 5]
Returns any
shift()
shift(count: ?int) -> any
Removed the specified count of items from the beginning of the list and returns it. If count is not specified, count defaults to 1. If one item is shifted, the method returns that item. If more than one item is shifted, the method returns a list containing the shifted items.
The square brackets (
[]) around thecount: numberin the method definition indicates that the parameter is optional and does not mean you have to type the square brackets.
If the number of items required to be shifted exceeds the size of the
list, the list is cleared and nil is returned.
For example:
%> var a = [9, 8, 7, 6, 5, 4, 3, 2, 1, 0]
%> a.shift()
9
%> a
[8, 7, 6, 5, 4, 3, 2, 1, 0]
%> a.shift(3)
[8, 7, 6]
%> a
[5, 4, 3, 2, 1, 0]
%> a.shift(10)
%> a
[]
Parameters
count(?int)
Returns any
remove_at()
remove_at(index: int) -> any
Removes the item at the specified index in the list and returns it. If
the index is less than 0 or greater than list.length() - 1, an Error
is raised.
For example:
%> var a = [1, 2, 3, 4, 5]
%> a.remove_at(3)
4
%> a
[1, 2, 3, 5]
%> a.remove_at(6)
Unhandled Error: list index 6 out of range at remove_at()
StackTrace:
<repl>:1 -> @.script()
%> a.remove_at(-1)
Unhandled Error: list index -1 out of range at remove_at()
StackTrace:
<repl>:1 -> @.script()
Parameters
index(int)
Returns any
remove()
remove(value) -> any
Removes the first occurrence of item x from the list.
For example:
%> var a = ['Kirk', 'Tasha', 'Emily', 'Kirk']
%> a.remove('Kirk')
%> a
[Tasha, Emily, Kirk]
Notice that only the first occurrence of Kirk was removed.
Parameters
value(any)
Returns any
reverse()
reverse() -> list
Returns a new list containing the items in the original list in reverse order.
For example:
%> var a = ['apple', 'mango', 'banana', 'orange', 'peach']
%> a.reverse()
[peach, orange, banana, mango, apple]
Returns list
sort()
sort(comparator: ?function) -> list
Sorts the items in the list in-place and returns the sorted list.
Sorting in Lists follows are strict set of precedence based on the
object type. The order for sorting is as follows in ascending
orders:
nil, boolean, numbers, strings, ranges, lists, dictionaries, file,
bytes, functions, classes and modules.
When the corresponding items in the list are of the same type, they are
sorted based on their respective values according to the type. For
example, the number 5 is less than 8 and as such will appear first
in the sort.
For example:
%> var a = ['A', 5, false, nil, [21, 13, 46]]
%> a.sort()
%> a
[nil, false, 5, A, [13, 21, 46]]
Notice how the boolean value precedes the number and how the number in turn precedes the string and the strings in turn, precedes the list in the result. Also, note that the items of the inner list is sorted.
Sorting by something else
Pass a comparator to decide the order yourself. It is given two items and returns a negative number to put the first one first, a positive number to put the second one first, and zero to leave them as they are:
%> [3, 1, 2].sort(@(a, b) => b - a)
[3, 2, 1]
%> ['pear', 'fig', 'banana'].sort(@(a, b) => a.length() - b.length())
['fig', 'pear', 'banana']
The sort is stable, so items the comparator calls equal keep the order they were already in. That is what lets a list be sorted by one thing and then another to order by both:
people.sort(@(a, b) => a.name.compare(b.name))
people.sort(@(a, b) => a.age - b.age)
leaves people of the same age in name order.
Parameters
comparator(function)
Returns list
Note: A comparator sorts only the list it is given. The inner lists that
sort()sorts on its own are left alone, since only the comparator knows what the order is meant to be.
Note: A comparator that contradicts itself produces some order rather than an error; there is no arrangement that satisfies it.
contains()
contains(value) -> boolean
Returns true if the list contains the item x or false otherwise.
For example:
%> ['dog', 'cat', 'wolf', 'tiger'].contains('cat')
true
%> ['dog', 'cat', 'wolf', 'tiger'].contains('giraffe')
false
Parameters
value(any)
Returns boolean
delete()
delete(start: int, end: int) -> number
Deletes a range of items from the list starting from the start to the
end limit and returns the number of items removed. If the start and end
are the same, this will be equivalent to list.remove_at(start).
For example:
%> var a = [1, 2, 3, 4, 5, 6, 7, 8, 9]
%> a.delete(3, 6)
4
%> a
[1, 2, 3, 8, 9]
%> a.delete(1,1) # equal start and end
1
%> a
[1, 3, 8, 9]
Parameters
start(int)end(int)
Returns number
first()
first() -> any
Returns the first item in the list or nil if the list is empty.
For example:
%> ['c', 'd', 'a', 'b'].first()
'c'
Returns any
last()
last() -> any
Returns the last item in the list or nil if the list is empty.
For example:
%> ['c', 'd', 'a', 'b'].last()
'b'
Returns any
is_empty()
is_empty() -> boolean
Returns true if the list is empty or false otherwise.
For example:
%> [1, 2].is_empty()
false
%> [].is_empty()
true
Returns boolean
take()
take(n: int) -> list
Returns a new list containing the first n items in the list or a new
copy of the list if n greater than or equals to the list.length().
If n < 0, returns list.take(list.length() - n).
For example:
%> var a = [10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20]
%> a.take(4)
[10, 11, 12, 13]
%> a.take(11) # taking more than the size of the list
[10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20]
%> a.take(-5) # taking n < 0
[10, 11, 12, 13, 14, 15]
Parameters
n(int)
Returns list
get()
get(index: int) -> any
Returns the value at the specified index in the list. If index is
outside the boundary of the list indexes (0..(list.length() - 1)), an
Error is thrown. This method is equivalent to list[index].
For example:
%> [13, 14, 15, 16].get(1)
14
%> [13, 14, 15, 16].get(6)
Unhandled Error: list index 6 out of range at get()
StackTrace:
<repl>:1 -> @.script()
Parameters
index(int)
Returns any
compact()
compact() -> list
Returns a new list containing the items in the original list but with
all nil values removed.
For example:
%> [21, nil, 14, 'age', nil, nil, [], 11].compact()
[21, 14, age, [], 11]
Returns list
unique()
unique() -> list
Returns a new list containing the unique values from the original list.
For example:
%> [1, 1, 3, 5].unique()
[1, 3, 5]
Returns list
zip()
zip(...lists: list) -> list
Returns a list that contains the items in the original list merged with corresponding items from the individual arguments. This generates a list of length equal to the length of the original argument.
If the size of any of the arguments is less than the size of the
original list, it’s corresponding entry will be nil.
For example:
%> var a = [4, 5, 6]
%> var b = [7, 8, 9]
%> [1, 2, 3].zip(a, b)
[[1, 4, 7], [2, 5, 8], [3, 6, 9]]
%> [1, 2].zip(a, b)
[[1, 4, 7], [2, 5, 8]]
%> a.zip([1, 2], [8])
[[4, 1, 8], [5, 2, nil], [6, nil, nil]]
%> [1, 2].zip([3])
[[1, 3], [2, nil]]
%> [1].zip([10, 11], [12, 13, 14])
[[1, 10, 12]]
%> [[1, 2], [3]].zip(a, b)
[[[1, 2], 4, 7], [[3], 5, 8]]
Parameters
lists(...list)
Returns list
zip_from()
zip_from(list: list) -> list
The same as list.zip() except that instead of accepting an arbitrary
list or arguments, it accepts a single list that should contain other
lists.
For example:
%> [1, 2].zip_from([[3, 4]])
[[1, 3], [2, 4]]
Parameters
list(list)
Returns list
to_dict()
to_dict() -> dict
Returns a number indexed dictionary representing the list.
For example:
%> ['English', 'French', 'Spanish'].to_dict()
{0: English, 1: French, 2: Spanish}
Returns dict
each()
each(callback: function)
Iterates over each element in the list, calling the provided callback function with the current element as an argument.
Example:
['A', 'B', 'C'].each(@(r) {
echo r
})
# Output: A B C
The
eachmethod does not return a new list; it simply executes the callback for each element. If you want to create a new list based on the original, consider using themapmethod instead.
Parameters
callback(function) — The function to execute for each element in the list.
Raises Error if the callback is not a function.
map()
map(callback: function) -> list
Creates a new list populated with the results of calling a provided function on every element in the calling list.
Example:
echo [1, 2, 3].map(@(x) {
return x * 2
})
# Output: [2, 4, 6]
Parameters
callback(function) — The function to execute on each element in the list. It receives the current element and its index as arguments.
Returns list
Raises Error if the callback is not a function.
filter()
filter(callback: function) -> list
Creates a new list with all elements that pass the test implemented by the provided function.
It returns a new list with the elements that pass the test. If no elements pass the test, an empty list will be returned.
Example:
echo [1, 2, 3].filter(@(x) {
return x % 2 == 0
})
# Output: [2]
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns list
Raises Error if the callback is not a function.
reduce()
reduce(callback: function, initial) -> any
Applies a function against an accumulator and each element in the list (from left to right) to reduce it to a single value and returns the accumulated result of the callback function.
Example:
echo [1, 2, 3].reduce(@(acc, x) {
return acc + x
})
# Output: 6
Parameters
callback(function) — The function to execute on each element in the list. It receives the current element and its index as arguments.initial— The initial value to use as the accumulator. If no initial value is provided, the first element of the list will be used as the initial accumulator, and the iteration will start from the second element.
Returns any
Raises Error if the callback is not a function.
some()
some(callback: function) -> boolean
Tests whether at least one element in the list passes the test implemented by the provided function.
Example:
echo [1, 2, 3].some(@(x) {
return x % 2 == 0
})
# Output: true
The some method returns true if the callback function returns a
truthy value for at least one element in the list. If the callback
function returns a falsy value for all elements, some will return
false. If the list is empty, some will return false by default.
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns boolean
Raises Error if the callback is not a function.
every()
every(callback: function) -> boolean
Tests whether all elements in the list pass the test implemented by the provided function.
Example:
echo [1, 2, 3].every(@(x) {
return x > 0
})
# Output: true
The every method returns true if the callback function returns a
truthy value for every element in the list. If the callback function
returns a falsy value for any element, every will return false. If
the list is empty, every will return true by default.
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns boolean
Raises Error if the callback is not a function.
find()
find(callback: function) -> any
Returns the value of the first element in the list that satisfies the
provided testing function. If no elements satisfy the testing function,
find returns nil.
Example:
echo [1, 2, 3].find(@(x) {
return x % 2 == 0
})
# Output: 2
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns any — The first element that satisfies it, or nil.
Raises Error if the callback is not a function.
find_index()
find_index(callback: function) -> number
Returns the index of the first element in the list that satisfies the
provided testing function. If no elements satisfy the testing function,
find_index returns -1.
Example:
echo [1, 2, 3].find_index(@(x) {
return x % 2 == 0
})
# Output: 1
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns number — The index of the first element that satisfies it,
or -1.
Raises Error if the callback is not a function.
find_last()
find_last(callback: function) -> any
Returns the value of the last element in the list that satisfies the
provided testing function. If no elements satisfy the testing function,
find_last returns nil.
Example:
echo [1, 2, 3].find_last(@(x) {
return x % 2 == 0
})
# Output: 2
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns any — The last element that satisfies it, or nil.
Raises Error if the callback is not a function.
find_last_index()
find_last_index(callback: function) -> number
Returns the index of the last element in the list that satisfies the
provided testing function. If no elements satisfy the testing function,
find_last_index returns -1.
Example:
echo [1, 2, 3].find_last_index(@(x) {
return x % 2 == 0
})
# Output: 1
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns number — The index of the last element that satisfies it,
or -1.
Raises Error if the callback is not a function.
find_all()
find_all(callback: function) -> list
Returns a new list containing all elements of the calling list that satisfy the provided testing function.
Example:
echo [1, 2, 3].find_all(@(x) {
return x % 2 == 0
})
# Output: [2]
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns list
Raises Error if the callback is not a function.
partition()
partition(callback: function) -> list
Returns an list containing two lists: the first with elements that satisfy the provided testing function, and the second with elements that do not satisfy the testing function.
Example:
echo [1, 2, 3].partition(@(x) {
return x % 2 == 0
})
# Output: [[2], [1, 3]]
Parameters
callback(function) — The function to test each element of the list. It receives the current element and its index as arguments.
Returns list
Raises Error if the callback is not a function.
to_string()
to_string() -> string
Returns the string representation of the list.
%> [1, 'two', 3].to_string()
'[1, two, 3]'
Returns string
Dictionary Methods
Every method on the built-in dict type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
length() | number | Returns the length of the dictionary. |
add(key: string, value) | Adds a new key-value pair to the dictionary with the given key and value. | |
set(key: string, value) | Sets the value of the given key to the given value in the dictionary. | |
clear() | Clears the content of the dictionary. | |
clone() | dict | Returns a new dictionary which is a deep copy of the original dictionary. |
compact() | dict | Returns a new dictionary that contains every key-value pair in the original dictionary except for keys whose associated value is nil. |
contains(key: string) | boolean | Returns true if any of the keys in the dictionary is equal to x, false otherwise. |
extend(dict: dict) | Adds all key-value pairs in dictionary x to the original dictionary. | |
get(key: string, default_value) | any|nil | Returns the value of the given key in the dictionary. |
keys() | list | Returns a list containing the keys in the dictionary. |
values() | list | Returns a list containing the value of all keys in the dictionary. |
remove(key) | any|nil | Removes a given key and it’s corresponding value from the dictionary and returns the value of the key. |
is_empty() | boolean | Returns true if the dictionary is empty, otherwise returns false. |
find_key(value) | string|nil | Returns the key whose value is equal to x in the dictionary or nil if no key has the value x. |
to_list() | list | Returns a list that contains a list of key and a list of values from the dictionary. |
each(callback: function) | void | Iterates over each key-value pair in the dictionary, calling the provided callback function with the value and key as arguments. |
filter(callback: function) | dict | Creates a new dictionary containing only the key-value pairs for which the provided callback function returns true. |
some(callback: function) | boolean | Tests whether at least one key-value pair in the dictionary passes the test implemented by the provided callback function. |
every(callback: function) | boolean | Tests whether all key-value pairs in the dictionary pass the test implemented by the provided callback function. |
reduce(callback: function, initial) | any | Reduces the dictionary to a single value by iteratively combining each key-value pair using the provided callback function. |
to_string() | string | Returns the string representation of the dictionary. |
length()
length() -> number
Returns the length of the dictionary. The length of a Zuri dictionary is
equal to the number of keys it contains. i.e. dict.length() == dict.keys().length().
For example:
%> {name: 'Zuri', version: 1}.length()
2
Returns number
add()
add(key: string, value)
Adds a new key-value pair to the dictionary with the given key and
value.
For example:
%> var dict = {}
%> dict.add('name', 'Zuri')
%> dict
{name: Zuri}
Parameters
key(string)value(any)
set()
set(key: string, value)
Sets the value of the given key to the given value in the dictionary. If
there is no existing entry for the key in the dictionary, a new entry
will be added.
For example:
%> dict.set('name', 'New Zuri')
%> dict
{name: New Zuri}
%> dict.set('version', 1)
%> dict
{name: New Zuri, version: 1}
@note:
dict.set(x, y)is equivalent to the following Zuri code.%> if dict.contains(x) { .. dict[x] = 1 .. } else { .. dict.add(x, 1) .. }
Parameters
key(string)value(any)
clear()
clear()
Clears the content of the dictionary.
For example:
%> var a = {name: 'Zuri'}
%> a
{name: Zuri}
%> a.clear()
%> a
{}
clone()
clone() -> dict
Returns a new dictionary which is a deep copy of the original
dictionary.
For example:
%> var new_dict = dict.clone()
%> new_dict
{name: New Zuri, version: 1}
Returns dict
compact()
compact() -> dict
Returns a new dictionary that contains every key-value pair in the
original dictionary except for keys whose associated value is nil.
For example:
%> var dict2 = {name: 'James', age: 20, address: nil, country: nil}
%> dict2.compact()
{name: James, age: 20}
Returns dict
contains()
contains(key: string) -> boolean
Returns true if any of the keys in the dictionary is equal to x,
false otherwise.
For example:
%> dict2.contains('name')
true
%> dict2.contains('street')
false
Parameters
key(string)
Returns boolean
extend()
extend(dict: dict)
Adds all key-value pairs in dictionary x to the original
dictionary.
For example:
%> var dict = {name: 'Zuri'}
%> dict.extend({version: 1})
%> dict
{name: Zuri, version: 1}
Parameters
dict(dict)
get()
get(key: string, default_value) -> any|nil
Returns the value of the given key in the dictionary. If the given key
is not defined in the dictionary and the default value is given, the
default value will be returned. Otherwise, nil is returned.
For example:
%> dict.get('version') # value exists
1
%> dict.get('age') # value does not exist
%> dict.get('age', 6) # value does not exist, but default is given
6
%> dict.get('version', 1.1) # value exists and default is given
1
Parameters
key(string)default_value(any|nil)
Returns any|nil
keys()
keys() -> list
Returns a list containing the keys in the dictionary.
For example:
%> dict.keys()
[name, version]
Returns list
values()
values() -> list
Returns a list containing the value of all keys in the dictionary.
For example:
%> dict.values()
[Zuri, 1]
Returns list
remove()
remove(key) -> any|nil
Removes a given key and it’s corresponding value from the dictionary and returns the value of the key.
For example:
%> dict = {username: 'james', email: 'a@b.c', active: true}
%> dict.remove('active')
true
%> dict
{username: james, email: a@b.c}
Parameters
key(string)
Returns any|nil
is_empty()
is_empty() -> boolean
Returns true if the dictionary is empty, otherwise returns
false.
For example:
%> dict.is_empty()
false
%> {}.is_empty()
true
Returns boolean
find_key()
find_key(value) -> string|nil
Returns the key whose value is equal to x in the dictionary or nil
if no key has the value x.
For example:
%> dict.find_key('james')
'username'
%> dict.find_key('camel')
Parameters
value(any)
Returns string|nil
to_list()
to_list() -> list
Returns a list that contains a list of key and a list of values from the
dictionary.
For example:
%> var dict = {username: 'james', email: 'a@b.c'}
%> dict.to_list()
[[username, email], [james, a@b.c]]
Returns list
each()
each(callback: function) -> void
Iterates over each key-value pair in the dictionary, calling the provided callback function with the value and key as arguments.
Example:
var myDict = {a: 1, b: 2, c: 3}
myDict.each(@(value, key) {
echo '${key}: ${value}'
})
# Output:
# a: 1
# b: 2
# c: 3
Parameters
callback(function) — The function to call for each key-value pair.
Returns void
Raises Error If the callback is not a function.
filter()
filter(callback: function) -> dict
Creates a new dictionary containing only the key-value pairs for which the provided callback function returns true. The callback function is called with the value and key as arguments.
Example:
var myDict = {a: 1, b: 2, c: 3}
var filteredDict = myDict.filter(@(value, key) {
return value > 1
})
echo filteredDict
# Output: {'b': 2, 'c': 3}
Parameters
callback(function) — The function to test each key-value pair. It should return true to keep the pair, or false to exclude it.
Returns dict
Raises Error If the callback is not a function.
some()
some(callback: function) -> boolean
Tests whether at least one key-value pair in the dictionary passes the test implemented by the provided callback function. The callback function is called with the value and key as arguments. The method returns true if the callback returns true for any key-value pair, otherwise it returns false.
Example:
var myDict = {a: 1, b: 2, c: 3}
var hasGreaterThanTwo = myDict.some(@(value, key) {
return value > 2
})
echo hasGreaterThanTwo
# Output: true
Parameters
callback(function) — The function to test each key-value pair. It should return true to indicate a passing pair, or false to indicate a failing pair.
Returns boolean
Raises Error If the callback is not a function.
every()
every(callback: function) -> boolean
Tests whether all key-value pairs in the dictionary pass the test implemented by the provided callback function. The callback function is called with the value and key as arguments. The method returns true if the callback returns true for every key-value pair, otherwise it returns false.
Example:
var myDict = {a: 1, b: 2, c: 3}
var allGreaterThanZero = myDict.every(@(value, key) {
return value > 0
})
echo allGreaterThanZero
# Output: true
Parameters
callback(function) — The function to test each key-value pair. It should return true to indicate a passing pair, or false to indicate a failing pair.
Returns boolean
Raises Error If the callback is not a function.
reduce()
reduce(callback: function, initial) -> any
Reduces the dictionary to a single value by iteratively combining each key-value pair using the provided callback function. The callback function is called with the accumulator, value, key, and the dictionary itself as arguments. The method returns the final accumulated value after processing all key-value pairs in the dictionary.
Example:
var myDict = {a: 1, b: 2, c: 3}
var sum = myDict.reduce(@(accumulator, value, key) {
return accumulator + value
}, 0)
echo sum
# Output: 6
Parameters
callback(function) — The function to execute on each key-value pair in the dictionary. It should return the updated accumulator value after processing the pair.initial(any) — The initial value to use as the first argument to the first call of the callback function.
Returns any
Raises Error If the callback is not a function.
to_string()
to_string() -> string
Returns the string representation of the dictionary.
%> {a: 1, b: 2}.to_string()
'{a: 1, b: 2}'
Returns string
Range Methods
Every method on the built-in range type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
lower() | number | Returns the lower limit of the range. |
upper() | number | Returns the upper limit of the range. |
range() | number | Returns a number equal to the numbers between the range. |
within(value: number) | boolean | Returns true if the given number falls somewhere within the or false otherwise. |
step(size: int) | range | Sets the step size of the range. |
get_step() | number | Returns the step size of the range. |
loop(callback: function) | void | Iterates over each number in the range, calling the provided callback function with the number, its index. |
to_list() | list | Returns the range as a list of its individual numbers, stepping from the lower limit to the upper limit (exclusive), or in reverse when the range descends. |
to_string() | string | Returns the string representation of the range. |
lower()
lower() -> number
Returns the lower limit of the range.
For example:
%> (10..100).lower()
10
Returns number
upper()
upper() -> number
Returns the upper limit of the range.
For example:
%> (20..30).upper()
30
Returns number
range()
range() -> number
Returns a number equal to the numbers between the range.
For example:
%> (21..93).range()
72
The result of stays the same irrespective of the direction of the range. For example, swapping the upper and lower limit of our previous still returns the same result.
%> (21..93).range()
72
Returns number
within()
within(value: number) -> boolean
Returns true if the given number falls somewhere within the or false otherwise.
For example:
%> (93..21).within(103)
false
%> (93..21).within(57)
true
Parameters
value(number)
Returns boolean
step()
step(size: int) -> range
Sets the step size of the range.
For example:
%> var a = (10..100).step(20)
%> a
<range 10..100, step=20>
%> for i in a {
.. echo i
.. }
10
30
50
70
90
Parameters
size(int) — The step size of the range.
Returns range
get_step()
get_step() -> number
Returns the step size of the range.
Returns number
loop()
loop(callback: function) -> void
Iterates over each number in the range, calling the provided callback function with the number, its index.
Example:
var r = 0..5 # 0, 1, 2, 3, 4
r.loop(@(num, index) {
echo 'Number at index ${index}: ${num}'
})
# Output:
# Number at index 0: 0
# Number at index 1: 1
# Number at index 2: 2
# Number at index 3: 3
# Number at index 4: 4
Parameters
callback(function) — A function that takes two arguments: the number, its index.
Returns void
Raises Error if the callback is not a function.
to_list()
to_list() -> list
Returns the range as a list of its individual numbers, stepping from the lower limit to the upper limit (exclusive), or in reverse when the range descends.
%> (1..5).to_list()
[1, 2, 3, 4]
%> (5..1).to_list()
[5, 4, 3, 2]
Returns list
to_string()
to_string() -> string
Returns the string representation of the range.
%> (1..5).to_string()
'1..5'
Returns string
Bytes Methods
Every method on the built-in bytes type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
length() | number | Returns the number of bytes in the byte stream. |
is_empty() | boolean | Returns true if the byte stream holds no bytes at all, and false otherwise. |
append(n: int) | bytes | Adds an item to the top of a byte stream. |
clone() | bytes | Returns a deep clone of the byte stream. |
extend(n: bytes) | bytes | Extends the byte stream with the bytes from the given byte stream. |
index_of(byte: int, start_index: ?number) | number | Returns the index of the first occurrence of the given byte in the byte stream. |
last_index_of(byte: int, end_index: ?number) | number | Returns the index of the last occurrence of the given byte in the byte stream, searching from the end, or -1 when the byte is not there. |
pop() | number | Removes the last item in a byte stream and returns it. |
remove(index: number) | bytes | Removes the item at the specified index in the byte stream and return the previous value at the specified index. |
reverse() | bytes | Reverses the items in the byte stream. |
first() | number | Returns the first item in the byte stream or nil if the byte stream is empty. |
last() | number | Returns the last item in the byte stream or nil if the byte stream is empty. |
get(index: number) | number | Returns the item at the specified index in the byte stream. |
take(n: int) | bytes | Returns a new byte stream containing the first n items in the bytes or a new copy of the bytes if n greater than or equals to the bytes.length(). |
split(delimiter: bytes) | list | Splits the content of a byte stream based on the specified delimiter. |
dispose() | Due to the nature of byte stream and their use-case (especially streaming data), it is easy for the system memory to get filled up with data in the byte stream. | |
is_alpha() | boolean | Returns true if the byte stream only contains alpha characters, false otherwise. |
is_alnum() | boolean | Returns true if the byte stream only contains alpha characters and numbers, false otherwise. |
is_number() | boolean | Returns true if the byte stream only contains numbers, false otherwise. |
is_lower(n) | boolean | Returns true if the byte stream only contains lower case characters, false otherwise. |
is_upper() | boolean | Returns true if the byte stream only contains upper case characters, false otherwise. |
is_space() | boolean | Returns true if the byte stream only contains space characters, false otherwise. |
to_list() | list | Returns the byte stream as a list of bytes. |
to_string() | string | Returns the byte stream as a string, reading it as UTF-8. |
each(callback: function) | void | Iterates over each byte of the bytes object, calling the provided callback function with the byte and its index. |
length()
length() -> number
Returns the number of bytes in the byte stream.
%> bytes([25, 57]).length()
2
Returns number
is_empty()
is_empty() -> boolean
Returns true if the byte stream holds no bytes at all, and false
otherwise. Equivalent to testing length() == 0, and unaffected by
whatever the bytes happen to contain: a stream of zero bytes is not
empty.
%> bytes(0).is_empty()
true
%> bytes(3).is_empty()
false
%> 'hi'.to_bytes().is_empty()
false
Returns boolean
append()
append(n: int) -> bytes
Adds an item to the top of a byte stream.
For example,
%> var a = bytes([0x40, 0x75])
%> a.append(0x16)
%> echo a
(40 75 16)
Parameters
n(int) — The byte to add.
Returns bytes
clone()
clone() -> bytes
Returns a deep clone of the byte stream.
For example,
%> bytes([19, 11]).clone()
(13 b)
Returns bytes
extend()
extend(n: bytes) -> bytes
Extends the byte stream with the bytes from the given byte stream.
For example,
%> var a = bytes([33, 91, 126])
%> var b = bytes([119, 42])
%> a
(21 5b 7e)
%> b
(77 2a)
%> a.extend(b)
(21 5b 7e 77 2a)
%> a
(21 5b 7e 77 2a)
Parameters
n(bytes) — The byte stream to extend with.
Returns bytes
Note:
extend()is an in-place action so the original byte stream will be modified.
index_of()
index_of(byte: int, start_index: ?number) -> number
Returns the index of the first occurrence of the given byte in the byte stream.
%> bytes([25, 57, 25]).index_of(57)
1
%> bytes([25, 57, 25]).index_of(25, 1)
2
Parameters
byte(int) — The byte to search for.start_index(?number) — The index to start the search from. Defaults to 0.
Returns number
last_index_of()
last_index_of(byte: int, end_index: ?number) -> number
Returns the index of the last occurrence of the given byte in the byte
stream, searching from the end, or -1 when the byte is not there.
If end_index is given, only a match at or before that index counts.
That is the same position index_of()’s own second parameter bounds, so
for any index n, index_of(b, n) and last_index_of(b, n) are the
first and last occurrences in the two halves n splits the stream into.
%> bytes([25, 57, 25]).last_index_of(25)
2
%> bytes([25, 57, 25]).last_index_of(25, 1)
0
Parameters
byte(int) — The byte to search for.end_index(?number) — The highest index a match may sit at. Defaults to the last byte.
Returns number
pop()
pop() -> number
Removes the last item in a byte stream and returns it.
%> var a = bytes([79, 43, 9])
%> a.pop()
9
%> a
(4f 2b)
Returns number
remove()
remove(index: number) -> bytes
Removes the item at the specified index in the byte stream and return the previous value at the specified index.
%> var a = bytes([25, 57, 25])
%> a.remove(1)
57
%> a
(25 25)
Parameters
index(number) — The index to remove.
Returns bytes
reverse()
reverse() -> bytes
Reverses the items in the byte stream.
%> bytes([5, 4, 3, 2, 1]).reverse()
(1 2 3 4 5)
Returns bytes
first()
first() -> number
Returns the first item in the byte stream or nil if the byte stream is
empty.
%> bytes([25, 57, 42]).first()
25
Returns number
last()
last() -> number
Returns the last item in the byte stream or nil if the byte stream is
empty.
%> bytes([25, 57, 42]).last()
42
Returns number
get()
get(index: number) -> number
Returns the item at the specified index in the byte stream.
Parameters
index(number) — The index to get the item from.
Returns number
take()
take(n: int) -> bytes
Returns a new byte stream containing the first n items in the bytes or
a new copy of the bytes if n greater than or equals to the
bytes.length(). If n < 0, returns bytes.take(bytes.length() - n).
For example:
%> var a = bytes([10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20])
%> a.take(4)
(0a 0b 0c 0d)
%> a.take(11) # taking more than the size of the bytes
(0a 0b 0c 0d 0e 0f 10 11 12 13 14)
%> a.take(-5) # taking n < 0
(0a 0b 0c 0d 0e 0f)
Parameters
n(int)
Returns bytes
split()
split(delimiter: bytes) -> list
Splits the content of a byte stream based on the specified delimiter.
For example,
%> bytes(0).split(bytes(0))
[]
%> echo 'test'.to_bytes().split(bytes(0))
[(74), (65), (73), (74)]
Parameters
delimiter(bytes) — The delimiter to split on.
Returns list
dispose()
dispose()
Due to the nature of byte stream and their use-case (especially streaming data), it is easy for the system memory to get filled up with data in the byte stream. The method allows users to reset a byte stream and empty it.
This method allows a fine-grained control on manual memory management of byte stream.
For example,
%> var a = bytes([13, 36])
%> a.dispose()
%> a
()
is_alpha()
is_alpha() -> boolean
Returns true if the byte stream only contains alpha characters,
false otherwise.
%> bytes([65, 66, 67]).is_alpha()
true
%> bytes([65, 66, 67, 128]).is_alpha()
false
Returns boolean
is_alnum()
is_alnum() -> boolean
Returns true if the byte stream only contains alpha characters and
numbers, false otherwise.
%> bytes([65, 66, 67, 48, 49, 50]).is_alnum()
true
%> bytes([65, 66, 67, 48, 49, 50, 8]).is_alnum()
false
Returns boolean
is_number()
is_number() -> boolean
Returns true if the byte stream only contains numbers, false
otherwise.
%> bytes([48, 49, 50]).is_number()
true
%> bytes([48, 49, 50, 68]).is_number()
false
Returns boolean
is_lower()
is_lower(n) -> boolean
Returns true if the byte stream only contains lower case characters,
false otherwise.
%> bytes([97, 98, 99]).is_lower()
true
%> bytes([97, 98, 99, 68]).is_lower()
false
Returns boolean
is_upper()
is_upper() -> boolean
Returns true if the byte stream only contains upper case characters,
false otherwise.
%> bytes([65, 66, 67]).is_upper()
true
%> bytes([65, 66, 67, 98]).is_upper()
false
Returns boolean
is_space()
is_space() -> boolean
Returns true if the byte stream only contains space characters,
false otherwise.
%> bytes([32, 32, 32]).is_space()
true
%> bytes([32, 32, 32, 68]).is_space()
false
Returns boolean
to_list()
to_list() -> list
Returns the byte stream as a list of bytes.
%> bytes([0x31, 0x55, 0xe9, 0x21]).to_list()
[49, 85, 233, 33]
Returns list
to_string()
to_string() -> string
Returns the byte stream as a string, reading it as UTF-8. Each sequence that is not valid UTF-8 reads as U+FFFD, the replacement character, so the string of such bytes does not encode back to the same bytes.
%> bytes([65, 66, 67, 68, 69]).to_string()
'ABCDE'
Returns string
each()
each(callback: function) -> void
Iterates over each byte of the bytes object, calling the provided callback function with the byte and its index.
Example:
var data = bytes([0x48, 0x65, 0x6C, 0x6C, 0x6F]) # "Hello" in bytes
data.each(def(byte, index) {
echo 'Byte at index ${index}: ${byte}'
})
# Output:
# Byte at index 0: 72
# Byte at index 1: 101
# Byte at index 2: 108
# Byte at index 3: 108
# Byte at index 4: 111
Parameters
callback(function) — A function that takes two arguments: the byte and its index.
Returns void
Raises Error if the callback is not a function.
File Methods
Every method on the built-in file type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
exists() | boolean | Returns true if a file exists or false otherwise. |
close() | void | Closes the stream to an opened file. |
open() | void | Opens the stream to a file for the operation originally specified on the file object during creation. |
read(length: ?int) | string|bytes | Reads the content of an opened file up to the specified length and returns it as string or bytes if the file was opened in the binary mode. |
gets(length: ?int) | string|bytes | Same as read(), but doesn’t open or close the file automatically. |
write(data: string|bytes) | string|bytes | Writes a string or bytes to an opened file at the current insertion point. |
puts(data: string|bytes) | string|bytes | Same as write(), but doesn’t open or close the file automatically. |
number() | int | Returns the integer file descriptor number that is used by the underlying implementation to request I/O operations from the operating system. |
is_tty() | boolean | Returns true if the file is connected to a TTY like device or false otherwise. |
is_open() | boolean | Returns true if the file is open for reading or writing and false otherwise. |
is_closed() | boolean | Returns true if the file is closed for reading or writing and false otherwise. |
flush() | void | Flushes the buffer held by a file. |
stats() | dict | Returns the statistics or details of a file. |
symlink() | boolean | Creates a symbolic link for the original file at the specified path. |
delete() | boolean | Deletes a file. |
rename(new_name: string) | boolean | Renames a file to to new_name. |
path() | string | Returns the path to the file. |
abs_path() | string | Returns the absolute path to the file. |
copy(path: string) | boolean | Copies a file from the path specified in the original file to the given path. |
truncate(length: ?number) | boolean | Truncates the entire file if length is not given or truncates the file such that only length number of bytes is left in it. |
chmod(mode: int) | boolean | Changes the permission on the file to the one specified in the number given. |
set_times(atime: number, mtime: number) | boolean | Sets the last access time and last modified time of the file. |
seek(offset: number, seek_type: int) | boolean | Sets the position of a file reader or writer in a file. |
tell() | number | Returns the current position of the reader/writer in a file. |
mode() | string | Returns the mode in which the current file was opened. |
name() | string | Returns the name of the current file. |
to_string() | string | Returns the file handle as a string, naming its path and the mode it was opened in. |
exists()
exists() -> boolean
Returns true if a file exists or false otherwise.
For example:
%> file('sample.txt').exists()
true
Returns boolean
close()
close() -> void
Closes the stream to an opened file. You’ll rarely ever need to call this method yourself in most use cases.
For example:
%> var f = file('sample.txt')
%> f.close()
Returns void
open()
open() -> void
Opens the stream to a file for the operation originally specified on the file object during creation. You may need to call this method after a call to read() if the length isn’t specified or write() if you wish to read or write again as the file will already be closed.
For example:
%> f.open()
Returns void
read()
read(length: ?int) -> string|bytes
Reads the content of an opened file up to the specified length and returns it as string or bytes if the file was opened in the binary mode. If the length is not specified, the file will be read to the end.
In text mode the bytes read must be valid UTF-8; anything else raises
rather than being silently replaced. Open the file in a binary mode
('rb') to read arbitrary bytes instead. Note that io.stdin is
already binary.
This method requires that the file be opened in the read mode (default
mode) or a mode that supports reading. If you aren’t reading the full
length of the file, you’ll need to call the close() method to free the
file for further reading, otherwise, the close() method will be
automatically called for you.
An example has been given above.
Parameters
length(?int)
Returns string|bytes
gets()
gets(length: ?int) -> string|bytes
Same as read(), but doesn’t open or close the file automatically.
Parameters
length(?int)
Returns string|bytes
write()
write(data: string|bytes) -> string|bytes
Writes a string or bytes to an opened file at the current insertion
point. When the file is opened with the a mode enabled, write will
always start from the end of the file. If the seek() method has been
previously called, write will begin from the seeked position, otherwise
it will start at the beginning of the file.
An example has been given above.
Parameters
data(string|bytes)
Returns string|bytes
puts()
puts(data: string|bytes) -> string|bytes
Same as write(), but doesn’t open or close the file automatically.
Parameters
data(string|bytes)
Returns string|bytes
number()
number() -> int
Returns the integer file descriptor number that is used by the underlying implementation to request I/O operations from the operating system. This can be very useful for low-level interfaces that uses or act as file descriptors.
For example:
%> file('sample.txt').number()
6
A standard stream reports the descriptor it is named by, 0, 1 or
2, rather than the private duplicate the runtime holds open for it, so
number() is what tells io.stdout apart from io.stderr.
Returns int
Note:
-1on a platform with no file descriptors, and on any handle that is not currently open.
is_tty()
is_tty() -> boolean
Returns true if the file is connected to a TTY like device or false
otherwise.
For example:
%> file('sample.txt').is_tty()
false
%> import io
%> io.stdout.is_tty() # io.stdin is a file...
true
Returns boolean
is_open()
is_open() -> boolean
Returns true if the file is open for reading or writing and false
otherwise.
@note:
stdfiles are always open.
For example:
%> file('sample.txt').is_open()
true
Returns boolean
is_closed()
is_closed() -> boolean
Returns true if the file is closed for reading or writing and false
otherwise.
For example:
%> file('sample.txt').is_closed()
false
Returns boolean
flush()
flush() -> void
Flushes the buffer held by a file. This could be useful for writable files as file writes are buffered.
For example:
%> w.flush()
Returns void
stats()
stats() -> dict
Returns the statistics or details of a file.
For example:
%> file('sample.txt').stats()
{is_readable: true, is_writable: true, is_executable: false,
is_symbolic: false, size: 72, mode: 33188, dev: 16777230,
ino: 4865113, nlink: 1, uid: 501, gid: 20, mtime: 1631395239,
atime: 1631395271, ctime: 1631395239, blocks: 8, blksize: 4096}
Every key above is present on every platform, so reading size or
mtime needs no check of which one you are on.
Returns dict
Note: Windows keeps a different set of facts about a file, and the ones it has no answer for read as
0:dev,ino,uidandgid, withnlinkalways1.modeis assembled from the file’s type and its read-only attribute, so it carries the right file-type bits and either0o444or0o666, widened by0o111for a directory or a namePATHEXTsays the shell would run.ctimeis the file’s creation time there, Windows having no equivalent of a Unix inode-change time.
symlink()
symlink() -> boolean
Creates a symbolic link for the original file at the specified path.
For example:
%> file('sample.txt').symlink('sample2.txt')
true
Returns boolean
Note: Windows decides at creation time whether a link stands for a file or a directory, so the original is inspected first; one pointing at something that does not exist yet is made as a file link. Creating any symbolic link there is privileged, and fails unless the machine is in developer mode or the process is elevated.
delete()
delete() -> boolean
Deletes a file.
For example:
%> file('test-2.zu').delete()
true
Returns boolean
Note: If the file is opened by one or more processes or threads outside of the current process or thread, the file will not be deleted until the last process frees it.
Note: This method throws Error on failure.
rename()
rename(new_name: string) -> boolean
Renames a file to to new_name. The new name can be a full path in
another location in which case the file will be moved.
For example:
%> file('sample copy.txt').rename('sample-2.txt')
true
Parameters
new_name(string)
Returns boolean
Note: The new name cannot be empty
Note: This method throws Error on failure.
path()
path() -> string
Returns the path to the file.
For example:
%> file('sample.txt').path()
'sample.txt'
Returns string
abs_path()
abs_path() -> string
Returns the absolute path to the file.
For example:
%> file('sample.txt').abs_path()
'C:\Users\username\zuri-docs\sample.txt'
Returns string
copy()
copy(path: string) -> boolean
Copies a file from the path specified in the original file to the given path.
For example:
%> file('./sample.txt').copy('samp.txt')
true
Parameters
new_name(string)
Returns boolean
truncate()
truncate(length: ?number) -> boolean
Truncates the entire file if length is not given or truncates the file such that only length number of bytes is left in it.
For example:
%> file('./samp.txt').truncate()
true
Parameters
length(?number)
Returns boolean
chmod()
chmod(mode: int) -> boolean
Changes the permission on the file to the one specified in the number given.
@note: The number is required to be an octal number. e.g. 0c755
For example:
%> file('sample.txt').chmod(0c755)
true
Parameters
mode(int)
Returns boolean
Note: Windows stores one read-only attribute where Unix stores nine permission bits, so the owner-write bit decides it and the rest are dropped:
0o755and0o700are the same instruction there. A mode with no owner-write bit marks the file read-only.
set_times()
set_times(atime: number, mtime: number) -> boolean
Sets the last access time and last modified time of the file.
@note: Time is expected in UTC seconds
@note: set argument -1 to leave the current value.
For example:
%> file('sample.txt').set_times(time(), time())
true
%> file('sample.txt').stats()
{is_readable: true, is_writable: true, is_executable: true,
is_symbolic: false, size: 72, mode: 33261, dev: 16777230,
ino: 4865113, nlink: 1, uid: 501, gid: 20, mtime: 1631477099,
atime: 1631477100, ctime: 1631477099, blocks: 8, blksize: 4096}
Parameters
atime(number)mtime(number)
Returns boolean
seek()
seek(offset: number, seek_type: int) -> boolean
Sets the position of a file reader or writer in a file. The position
must be within the range of the file size. seek_type must be on of
SEEK_SET, SEEK_CUR or SEEK_END from the io package.
For example:
%> f.seek(5, io.SEEK_SET)
true
Parameters
offset(number)seek_type(int)
Returns boolean
tell()
tell() -> number
Returns the current position of the reader/writer in a file.
For example:
%> import io
%> var f = file('sample.txt')
%> f.seek(5, io.SEEK_SET)
true
%> f.tell()
5
Returns number
mode()
mode() -> string
Returns the mode in which the current file was opened.
For example:
%> file('sample.txt').mode()
'r'
Returns string
name()
name() -> string
Returns the name of the current file.
For example:
%> file('./sample.txt').name()
'sample.txt'
Returns string
to_string()
to_string() -> string
Returns the file handle as a string, naming its path and the mode it was opened in.
%> file('sample.txt', 'w').to_string()
'<file at sample.txt in mode w>'
Returns string
Function Methods
Every method on the built-in function type, with its signature, what
it returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
name() | string | The function’s declared name. |
arity() | number | How many parameters the function declares, counting a variadic one as a single parameter. |
is_variadic() | bool | Whether the last parameter is variadic (...name). |
call(...args: list) | any | Calls the function with the given arguments and returns its result. |
apply(args: list) | any | Calls the function with the arguments in a list, and returns its result. |
to_string() | string | The function rendered for display, as <function NAME(ARITY)>, with a trailing ... on the arity when the function is variadic. |
name()
name() -> string
The function’s declared name. An anonymous function is named @anon
followed by a number, counted in the order the compiler met it; a bound
method reports the method’s own name, not its class’s.
%> def named(a, b, ...c) {}
%> named.name()
'named'
%> print.name()
'print'
Returns string
arity()
arity() -> number
How many parameters the function declares, counting a variadic one as a single parameter.
%> def named(a, b, ...c) {}
%> named.arity()
3
A method read off an instance counts the instance as its first
parameter, so a method declaring two parameters reports 3. A foreign
function from ffi has no receiver and reports its C parameter list.
Returns number
is_variadic()
is_variadic() -> bool
Whether the last parameter is variadic (...name).
%> def named(a, b, ...c) {}
%> named.is_variadic()
true
Returns bool
call()
call(...args: list) -> any
Calls the function with the given arguments and returns its result. The same as calling it directly; useful when the function is held in a variable and the call site reads better spelled out.
%> def add(a, b) { return a + b }
%> add.call(2, 3)
5
Parameters
args(...any) — The arguments to call with.
Returns any — Whatever the function returns.
Raises anything the called function raises.
apply()
apply(args: list) -> any
Calls the function with the arguments in a list, and returns its result.
call() takes them written out; this one takes them in a list.
%> def add(a, b) { return a + b }
%> add.apply([2, 3])
5
%> def collect(first, ...rest) { return [first, rest] }
%> collect.apply([1, 2, 3])
[1, [2, 3]]
A list shorter than the function’s arity leaves the remaining parameters
nil, exactly as calling it directly with too few arguments does; a
longer one overflows into a variadic parameter, or is discarded when
there is none.
Parameters
args(list) — The arguments to call with, in order.
Returns any — Whatever the function returns.
Raises TypeError when args is not a list, and anything the
called function raises.
to_string()
to_string() -> string
The function rendered for display, as <function NAME(ARITY)>, with a
trailing ... on the arity when the function is variadic.
%> def named(a, b, ...c) {}
%> named.to_string()
'<function named(3...)>'
Returns string
Appendix F: The Standard Library Index
Every module in the standard library, what it is for, and where the book
covers it. All of them are reachable with a bare import, with no package
manager and no dependency to add.
Each module’s source is in libs/, and every public function in it carries
a doc block stating its parameters, its defaults and its edge cases.
Data and Serialisation
| Module | What it is for | Book |
|---|---|---|
json | encode, decode, read and write JSON | 13 |
yaml | parse YAML, with anchors, tags and multi-document streams | 13 |
toml | parse and write TOML, and edit one without disturbing its layout | 13 |
csv | read and write CSV, with dialect detection | 13 |
struct | pack and unpack binary layouts | 10 |
base64 | Base64 encode and decode | 10 |
convert | base, hex, binary, octal and unicode conversions | 13 |
compress | deflate, zlib, gzip, zstd, lz4, bzip2, brotli, tar, zip, checksums | 10 |
Text and Markup
| Module | What it is for | Book |
|---|---|---|
html | a WHATWG-conformant parser, a DOM and CSS selectors | 13 |
wire | templating, with directives as HTML attributes | 14 |
url | parse, build and percent-encode URLs | 13 |
mime | detect a media type from a name or from content | 13 |
colors | ANSI terminal colour, with graceful degradation | 13 |
Time
| Module | What it is for | Book |
|---|---|---|
date | dates, times, formatting, parsing and IANA time zones | 13 |
Cryptography and Identity
| Module | What it is for | Book |
|---|---|---|
hash | digests and HMACs, plus PBKDF2 | 10 |
bcrypt | password hashing | 13 |
crypto | RSA signing and HKDF | 13 |
uuid | UUID versions 1, 3, 4, 5, 6, 7 and 8 | 13 |
jwt | sign, verify and decode JSON Web Tokens, with JWKS | 13 |
Structure and Validation
| Module | What it is for | Book |
|---|---|---|
validate | a fluent schema builder | 13 |
types | checked coercion between types | 13 |
set | an ordered set with the usual algebra | 13 |
enum | named constants from a list or a dictionary | 13 |
array | typed fixed-width numeric arrays, Int8 through Double | 13 |
The System
| Module | What it is for | Book |
|---|---|---|
os | processes, filesystem, paths, environment, signals | 9 |
env | a .env file into the environment, and typed values back out | 19 |
io | standard streams, the terminal, and in-memory files | 9, 10 |
stat | the S_IS* predicates over a file mode | 13 |
args | a command-line parser with subcommands and --help | 20 |
log | levelled, structured logging with pluggable transports | 13 |
isolate | OS-thread concurrency, channels and broadcasts | 11 |
ffi | calling C and Rust libraries, callbacks, and linking static libraries | 26 |
Testing
| Module | What it is for | Book |
|---|---|---|
test | suites, matchers, test doubles, snapshots and reports | 23 |
Databases
| Module | What it is for | Book |
|---|---|---|
sql | one contract for every relational database, with SQLite, PostgreSQL and MySQL adapters | 17 |
| Module | What it is for | Book |
|---|---|---|
mail | messages, SMTP, IMAP and POP3, both server ends, and DKIM | 18 |
The Network
| Module | What it is for | Book |
|---|---|---|
net | TCP, UDP, unix sockets, TLS, DTLS, addresses and polling | 12 |
http | an HTTP/1.1 and HTTP/2 client and server | 15 |
rpc | JSON-RPC 2.0 over HTTP, WebSockets, sockets and streams, both ends | 28 |
Graphics
| Module | What it is for | Book |
|---|---|---|
imagine | decode, draw, filter and encode images | 16 |
Numbers
| Module | What it is for | Book |
|---|---|---|
math | the mathematical constants | 4 |
Metaprogramming
| Module | What it is for | Book |
|---|---|---|
zuri | lexing, parsing, compiling and runtime reflection | 21 |
Packages and Their Submodules
Most of the larger modules are packages: a directory whose index.zu
re-exports what its parts make public. import http reaches almost all of
http without naming a submodule. Import a submodule directly when you
want only that part of it, or when a name would otherwise collide.
array
One module per element width, each exporting a single class.
| Submodule | Class | Element |
|---|---|---|
array.int8 | Int8Array | signed 8-bit |
array.uint8 | UInt8Array | unsigned 8-bit |
array.int16 | Int16Array | signed 16-bit |
array.uint16 | UInt16Array | unsigned 16-bit |
array.int32 | Int32Array | signed 32-bit |
array.uint32 | UInt32Array | unsigned 32-bit |
array.int64 | Int64Array | signed 64-bit |
array.uint64 | UInt64Array | unsigned 64-bit |
array.float | FloatArray | 32-bit float |
array.double | DoubleArray | 64-bit float |
compress
| Submodule | What it is for |
|---|---|
compress.deflate | raw DEFLATE streams |
compress.zlib | DEFLATE with a zlib header |
compress.gzip | DEFLATE with a gzip header and trailer |
compress.zstd | Zstandard, levels 1 to 22 |
compress.lz4 | LZ4, block and frame formats |
compress.bzip2 | bzip2 |
compress.brotli | Brotli |
compress.tar | reading and writing TAR archives |
compress.zip | reading and writing ZIP archives |
compress.checksum | CRC32, CRC32C, Adler-32 |
ffi
| Submodule | What it is for |
|---|---|
ffi.errors | every error the module raises, under FfiError |
ffi.types | Type, and the struct, union and enum builders |
ffi.pointer | Pointer: native memory and access through it |
ffi.library | Library: a loaded shared library |
ffi.declare | Declarations: C and Rust source read into types and signatures |
ffi.callback | Callback: a Zuri function behind a C function pointer |
html
| Submodule | What it is for |
|---|---|
html.tokenizer | the WHATWG tokenizer: text in, tokens out |
html.parser | tree construction: tokens in, a document out |
html.node | the document tree and everything you can do to it |
html.selector | finding nodes with CSS selectors |
html.serialize | writing a document back out, minified or pretty |
html.entities | named character references, both directions |
html.elements | the element tables tree construction consults |
html.namespaces | the five namespace URIs the parser deals in |
http
| Submodule | What it is for |
|---|---|
http.client | HttpClient, the request side |
http.server | HttpServer, the listening side |
http.worker | serving across several isolates |
http.router | matching a method and path to a handler |
http.middleware | CORS, access logging, security headers, and the rest |
http.request | the request object handlers receive |
http.response | the response object handlers return |
http.headers | Headers, with the field-name rules of RFC 9110 |
http.cookies | Cookie and CookieJar |
http.session | server-side sessions, and the stores that keep them |
http.session.sql | keeping sessions in a relational database |
http.body | reading and writing message bodies |
http.multipart | multipart/form-data, including file uploads |
http.files | serving files from disk, with ranges and caching |
http.stream | chunked and streaming transfers |
http.sse | server-sent events |
http.websocket | the WebSocket protocol |
http.proxy | forwarding requests to another server |
http.negotiate | parsing Accept-style headers |
http.status | the IANA status codes and their reason phrases |
http.h1 | the HTTP/1.1 wire format |
http.errors | the module’s error hierarchy |
imagine
| Submodule | What it is for |
|---|---|
imagine.image | Image, the pixel buffer everything else operates on |
imagine.canvas | drawing: lines, shapes, fills, text |
imagine.color | Color, and conversion between colour spaces |
imagine.filters | blur, sharpen, convolution, and the rest |
imagine.font | loading and measuring fonts |
imagine.strokefont | the built-in stroke font, with no file to load |
imagine.formats | decoding and encoding PNG, JPEG, GIF, WebP and more |
imagine.animation | multi-frame images |
imagine.constants | the named constants the module understands |
imagine.errors | the module’s error hierarchy |
io
| Submodule | What it is for |
|---|---|
io.bytesio | BytesIO, a file-shaped object backed by memory |
io.tty | terminal control: raw mode, size, cursor |
isolate
| Submodule | What it is for |
|---|---|
isolate.channel | bounded multi-producer, multi-consumer queues |
isolate.broadcast | one-to-many publish and subscribe |
isolate.error | IsolateError |
jwt
| Submodule | What it is for |
|---|---|
jwt.core | encode(), decode(), sign(), verify() |
jwt.signer | Signer, a reusable configured signer |
jwt.verifier | Verifier, a reusable configured verifier |
jwt.token | the Token object a complete decode returns |
jwt.jwks | resolving a signing key from a JSON Web Key Set |
jwt.codec | algorithm identifiers and the low-level encoding |
jwt.errors | the module’s error hierarchy |
log
| Submodule | What it is for |
|---|---|
log.logger | the module-level info(), warn(), error() and friends |
log.level | the LogLevel enum and the default level |
log.transport | Transport, the base class every sink extends |
log.console | ConsoleTransport, the default |
log.file | FileTransport, with size-based rotation |
log.dispatch | configuring which transports receive what |
net
| Submodule | What it is for |
|---|---|
net.tcp | TcpSocket and TcpStream |
net.udp | UdpSocket |
net.unix | UnixStream, over a path rather than an address |
net.tls | TLS over a TCP stream |
net.dtls | DTLS over a UDP socket |
net.ip | parsing, formatting and classifying IP addresses |
net.addr | SocketAddrV4 and SocketAddrV6 |
net.poll | asking which of a set of sockets is ready |
os
| Submodule | What it is for |
|---|---|
os.path | joining, resolving and comparing path strings |
os.fs | directories, permissions, symlinks, globbing |
os.env | reading, writing and listing environment variables |
os.process | process identity, subprocesses, signals |
os.system | facts about the process, the runtime and the machine |
os.tempfile | the temporary directory, and scratch files in it |
rpc
| Submodule | What it is for |
|---|---|
rpc.message | Request, Notification, Response, and reading and writing them as JSON |
rpc.error | RpcError, the codes the specification defines, and the module’s other errors |
rpc.service | Service and the Context its handlers are given |
rpc.http | http_handler() and HttpClient, JSON-RPC over HTTP |
rpc.framing | HeaderFraming, LineFraming and MessageFraming, telling messages apart |
rpc.transport | the transports, WebSockets among them, and pipe() between isolates |
rpc.endpoint | Endpoint, a service on a connection |
rpc.server | serve(), an endpoint for every connection to a listening socket |
sql
| Submodule | What it is for |
|---|---|
sql.driver | the contract an adapter implements, and the capability flags |
sql.errors | every error a database raises, under one root |
sql.params | rewriting ? and :name into whatever an engine wants |
sql.types | how Zuri values and database values correspond |
sql.decimal | Decimal, for a column a float must not hold |
sql.result | ResultSet and ExecResult |
sql.cursor | reading a result a row at a time |
sql.statement | a statement compiled once and run many times |
sql.transaction | Transaction, and the savepoints inside it |
sql.connection | the Connection a program holds |
sql.crud | building the four statements that are always the same |
sql.schema | asking a database what is in it |
sql.pool | keeping connections open and lending them out |
sql.sqlite | the SQLite adapter, and its blobs, backups and hooks |
sql.postgres | the PostgreSQL adapter, and LISTEN/NOTIFY |
sql.mysql | the MySQL and MariaDB adapter |
mail
| Submodule | What it is for |
|---|---|
mail.errors | every error the mail stack raises, under one root |
mail.address | reading and writing the addresses in a header |
mail.headers | the header block, in order and without regard to case |
mail.encoding | the encodings a header and a body use |
mail.content | Content-Type and Content-Disposition |
mail.message | a message, its MIME tree, and building one |
mail.dkim | signing a message and checking a signature |
mail.sasl | the authentication mechanisms all three protocols share |
mail.stream | a line-oriented connection, and negotiating TLS over one |
mail.smtp | the sending and receiving ends of SMTP |
mail.imap | the client and server ends of IMAP, and where mail is kept |
mail.imap.parser | the IMAP grammar |
mail.imap.store | MailStore, MaildirStore and MemoryStore |
mail.pop3 | the client end of POP3 |
mail.pool | running a mail server on more than one connection at once |
test
| Submodule | What it is for |
|---|---|
test.expect | Expect and every matcher on it |
test.runner | collecting the declarations and running them |
test.reporter | Reporter, and the seven built-in ones |
test.result | Case, Suite, Failure and Summary |
test.mock | Mock, mock() and spy_on() |
test.snapshot | the snapshot store and its file format |
test.conduct | discovering and running a directory of test files |
test.diff | structural equality, and rendering what differs |
test.format | rendering any value for a failure message |
test.source | reading a stack trace back to the failing line |
test.style | terminal colour, symbols and width |
test.error | AssertionError and TestSetupError |
test.context | what is true while one test is running |
validate
| Submodule | What it is for |
|---|---|
validate.validators | the one-line entry points, one per rule |
validate.validator | Validator, the fluent builder |
validate.schema | Schema, validating a whole dictionary at once |
validate.rule | Rule, the base class custom rules extend |
validate.rules | every built-in rule |
wire
| Submodule | What it is for |
|---|---|
wire.compile | turning a parsed template into an instruction tree |
wire.render | walking a compiled template and writing the page |
wire.expression | the language between {{ and }} |
wire.filters | the filters every template starts with |
wire.escape | context-aware escaping |
wire.loader | resolving the path in an x-include |
wire.normalize | rewriting the pseudo elements before parsing |
wire.values | how a template reads the values it is given |
wire.constants | the directive names Wire reserves |
wire.errors | the module’s errors, and the locations they carry |
zuri
| Submodule | What it is for |
|---|---|
zuri.token | tokenize(), and the Token type it returns |
zuri.ast | parse(), and the Node type it returns |
zuri.compile | compile(), and the Instr type it returns |
zuri.reflect | inspecting a live function, class, module or instance |
Shadowing a Module
A file in ./.zuri/libs/ shadows a standard library module of the same
name, because that directory is searched first. See
The Module System.
Appendix G: The Error Hierarchy
Every error in Zuri is an instance of a class, and every one of them
inherits from Error. They are declared in ordinary Zuri and go through
the same class machinery user code does, which is why subclassing one
behaves exactly like subclassing anything else.
Error
├── TypeError
├── ValueError
├── NumericError
├── ArgumentError
├── NotImplementedError
├── RangeError
├── AccessError
├── AssertError
├── PropertyError
├── UndefinedError
└── ModuleNotFoundError
Fields
Every error carries three:
| Field | What it holds |
|---|---|
message | the text, defaulting to 'An unexpected error has occurred' |
type | the class name, as a string |
stacktrace | a list of frames, innermost first |
catch {
raise ValueError('bad input')
} as e {
echo e.type
echo e.message
echo e.stacktrace
}
ValueError
bad input
[/path/to/main.zu:2 -> @.script()]
When Each One Is Raised
| Class | Raised when |
|---|---|
Error | the base class; a general failure with nothing more specific to say |
TypeError | an operation received the wrong type: an undefined operator signature, a method on nil, an annotated parameter given the wrong thing |
ValueError | the type was right and the value was not |
NumericError | an arithmetic operation failed |
ArgumentError | a call passed the wrong number of arguments to a native function |
NotImplementedError | a method meant to be overridden was not |
RangeError | an index or a bound fell outside what the value allows |
AccessError | a permission or access check failed |
AssertError | an assert condition was falsy |
PropertyError | a member that does not exist was read: a missing dictionary key, an undeclared field, a module member that was not exported |
UndefinedError | an undefined global was read |
ModuleNotFoundError | an import could not be resolved |
Catching
catch catches everything inside its block. To handle one kind and let the
rest through, test and re-raise:
catch {
load_config()
} as e {
if !instance_of(e, ModuleNotFoundError) {
raise e
}
echo 'no config, using defaults'
}
instance_of() walks the whole chain, so a test against Error matches
everything.
A parameter annotated Error accepts any of them, which is the readable
way to write a handler:
def report(e: Error) {
echo '${e.type}: ${e.message}'
}
Subclassing
class HttpError < Error {
@new(message, status) {
parent(message)
self.type = 'HttpError'
self.status = status
}
}
Two things make this work well. Call parent(message) so the base
constructor sets message and the stack trace is captured. Set
self.type so the class name appears in logs and in the uncaught-error
banner.
Carry whatever the handler needs. An error class exists precisely so it can
hold more than a string; the capstone’s TaskError carries the name of the
field that failed validation, and that is what lets an API answer
{"error": "...", "field": "title"}.
Uncaught
An error nobody catches prints its type, its message, the source around the
failure, and the stack trace, then exits 1:
Unhandled ValueError: bottomed out
--> /path/to/main.zu:3
1 | def recurse(n) {
2 | if n <= 0 {
> 3 | raise ValueError("bottomed out")
4 | }
5 | recurse(n - 1)
Stack trace (most recent call last):
at recurse() /path/to/main.zu:3
at recurse() /path/to/main.zu:5
... 19 more frames ...
at recurse() /path/to/main.zu:5
A deep stack is truncated in the middle. The top and the bottom are the parts that tell you anything.
There Is No finally
Code after a catch statement runs whether the block raised or not,
because the handler either recovers or re-raises. See
Error Handling for the patterns that replace
it.
Appendix H: Coming From Another Language
The places Zuri will surprise you are the places it looks most familiar. This is that list.
Everyone
| You expect | Zuri does |
|---|---|
x++ as a statement only | x++ is postfix only, and in an expression it evaluates to the new value. ++x does not parse. |
[] to be falsy | [] and {} are truthy. Use is_empty(). |
finally | There is none. Code after the catch statement runs either way. |
try | The keyword is catch, and it takes the block that might fail: catch { ... } as e { ... }. |
new Thing() | Call the class: Thing(). |
switch/case with fall-through | using/when, first match only, no break. |
a main function | A file’s top level is the program. |
| declaration hoisting | None. A def must appear above the top-level line that calls it. |
| overloading by arity | None. Two defs of one name in one scope is a compile error. |
| reopening a class | class Ext > Target adds methods to an existing class, globally. See Extensions. |
string ordering with < | < is numbers only. Use compare(), which returns -1, 0 or 1. |
x in collection | No membership operator. Use contains(). |
?. and ?? | Neither exists. or covers the common case, with the truthiness caveat above. |
an eval() | There is none, deliberately. See Metaprogramming. |
From Python
- Blocks are braces, not indentation, and every control-flow body takes one.
defdeclares functions andclassdeclares classes, but there is noselfparameter:selfis implicit inside a method and required for every field access.- The constructor is
@new, not__init__. Dunder methods are@-prefixed decorated methods:@add,@lt,@key,@to_json. __eq__is@eq, and!=is always its negation. It runs only when the other side is an object too, sox == nilnever calls it. Without it, instances compare by identity.- No list comprehensions.
map(),filter()andreduce()are methods on the list. len(x)isx.length().str(x)isx.to_string().int(x)isx.int()orx.to_number().- Slicing is
s[a, b], with a comma, nots[a:b]. elifiselse if.- Modules run once and are cached, as in Python. Circular imports behave the same way, and have the same caveat.
if __name__ == '__main__'isif __root__ == __file__.
From JavaScript
varis block-scoped and behaves likelet.constprevents rebinding and is enforced in local scopes.==does no coercion.'1' == 1isfalse. There is no===.- Arrow functions are
@(x) => x * 2ordef(x) => x * 2. The@is the common spelling. - There is no
thisrebinding to worry about.selfis the instance, always. - Objects and dictionaries are the same thing, and a class is not one.
Classes are sealed: you cannot add a property to an instance, though a
class Ext > Targetdeclaration can add a method to the class. nullandundefinedare bothnil.- No
async/awaitand no event loop. Concurrency is isolates: real threads with separate heaps, communicating by copying. JSON.stringifyisjson.encode, andcompactdefaults totrue.- Template literals are
'${expr}', in ordinary single or double quotes.
From Ruby
- No implicit returns. A function without
returnyieldsnil. - No blocks or
yield. Pass an anonymous function. nilis the only nil-like value, andfalseis separate from it.- Methods do not end in
?or!. Predicates are namedis_*and mutation is documented rather than punctuated. eachhands the callback value first, index second.forhands you key first, value second.- There is no
method_missing. Adding methods to an existing class is possible, but through an explicitclass Ext > Targetdeclaration rather than by reopening the class. - Modules are files, not a language construct. There is no
includeorextend.
From Go
- Dynamically typed, with optional annotations on parameters that are checked at every call.
- Errors are raised and caught, not returned. There is no
err != nilpattern. - Isolates are not goroutines. They are OS threads with separate heaps, and values crossing between them are copied. That is the whole concurrency model, and it is why there are no mutexes.
- Channels are bounded queues and behave the way you expect,
select()included. - No interfaces. A parameter typed
Erroraccepts any subclass, and that is the whole of the polymorphism story alongside inheritance. deferhas no equivalent. Close what you opened, on both paths.
From Java or C#
- No static typing, no generics, no interfaces, no packages-as-namespaces. A module is a file.
- Single inheritance, and no
abstractkeyword. A base method that raisesNotImplementedErroris the idiom. public/privateis a leading underscore, enforced at compile time for both class members and module members.- There is no overloading. One name, one method, and a second declaration of either is a compile error.
equals()is@eq, and==calls it.toString()isto_string(), and nothing calls it for you. Whatechoshows for an instance comes from@to_string().
From C
- Numbers are doubles. There is no integer type, and
/never truncates.//is floor division. %keeps the sign of the left operand, as in C.//rounds toward negative infinity, which C’s/does not.- No pointers, no manual memory management. A generational collector owns the heap.
bytesis the buffer type, andstructis how you read and write binary layouts.switchisusing, with no fall-through.
Things That Will Save You an Hour
Interpolation does not call to_string(). '${thing}' is
<instance of Thing>. Call the method. echo shows an instance through
@to_string(), a separate method.
list.sort() mutates and returns; list.reverse() does neither to the
original. That asymmetry is the most common list bug in Zuri code.
A def scopes like a var. At the top level of a file it binds a
module-level name; anywhere else it is a local of the block it is written
in, and it goes away with that block.
import is local by default. If your module imports something and a
third file cannot see it through you, add the @.
A same-directory import is import .sibling, not the full path from
the project root.
A conditional expression breaks across lines either way. ? and :
may each end a line or begin the next, so cond ? on one line and a
line starting with ? or : both parse.
Appendix I: Custom Commands
zuri fmt, zuri init and zuri test are Zuri programs. The runtime
ships them in a cmds directory beside itself, and when the first word
after zuri is not run, it looks that word up there and runs the
script it finds. A project adds commands of its own the same way, in
its .zuri/cmds directory, and they are run, listed and documented
exactly like the ones that ship.
This is the place for the scripts a project keeps running by hand: a
release checklist, a data import, a code generator. As a command, each
one is found by name from anywhere in the checkout’s root, shows up in
zuri --help with a line saying what it does, and parses its own
arguments with the same args module as everything else.
Where Commands Are Found
zuri <name> looks in these places, in this order, and runs the first
match:
| Where | What it holds |
|---|---|
$ZURI_ROOT/cmds, or cmds beside the executable when ZURI_ROOT is unset | the commands the runtime ships |
.zuri/cmds in the project | the project’s own commands |
cmds in each package in the project’s .zuri/libs | commands the project’s packages provide |
cmds in each package in $ZURI_HOME/libs | commands the packages installed for your user provide |
The project is the nearest directory above the working directory that
holds a project.toml, so a project’s commands run from anywhere inside
it. With no project, .zuri in the working directory is used.
The shipped commands come first, so a project cannot replace one: a
project command named test is never run, and zuri --help leaves it
out of the listing. A project’s own commands come before anything a
package provides, and a project’s packages before your user’s.
Two packages in the same place providing the same command is refused
rather than settled by chance. zuri <name> then names both packages
and runs neither, and zuri --help marks the command as claimed twice.
zuri install refuses to create the clash in the first place.
zuri init writes a .gitignore that ignores everything under .zuri
except .zuri/cmds, so a project’s commands are committed with the
rest of its code while the tools that keep state in .zuri do not
leave it behind in the repository.
Writing One
A command is a single .zu file, or a directory with an index.zu:
.zuri/
cmds/
greet.zu
release/
index.zu
notes.zu
tests/
zuri greet runs greet.zu, and zuri release runs
release/index.zu. When a directory and a file share a name, the
directory wins. The directory form is for a command that has grown past
one file: index.zu imports its siblings relatively, as any package
does, with import .notes.
The name a command answers to is its file or directory name. It is a
single name, never a path: zuri tools/greet is refused as an unknown
command before anything is looked up.
A name starting with _ is private, the way an identifier starting with
one is. _notes.zu and _shared/index.zu are never commands: zuri --help leaves them out and zuri _notes is an unknown command. That is
the place for code several commands share, which each of them imports
relatively, as import .._shared.notes from release/index.zu.
Naming and Describing It
The first doc block in the file introduces the command. @command
states the name it answers to, and @description is the one line
zuri --help shows beside it:
/**
* @command release
* @description Tags a release and writes its notes from the commits
* since the last one.
*
* Run it from a clean checkout on the main branch:
*
* zuri release 1.4.0
*/
import .notes
The block has to open within the first 16 KiB of the file, which in
practice means at the top. A description longer than a line carries on
over the indented lines beneath the tag, and ends at a blank line or
the next tag. A command without a description is still listed, with
its name alone. @command is what a reader of the file sees first, and
it names the file’s own command; the runtime lists and runs a command
by its file or directory name, so keep the two the same.
Arguments and the Exit Status
Everything after the command’s name belongs to the command, --help
included, and a command reads it with the args
module. Build a parser named after the command and call parse() with
nothing: it reads the real command line, converts and validates each
value, answers --help from the declarations, and refuses anything
undeclared with a message and exit status 1. That is what gives every
command, shipped or not, the same flags, the same help layout and the
same errors:
/**
* @command greet
* @description Greets whoever is named, as often as asked.
*/
import args
def parser() {
var p = args.Parser('greet', false)
p.description = 'Greets whoever is named, as often as asked.'
p.add_index('name', 'Who to greet', { required: true })
p.add_option('times', 'How many greetings', { short_name: 't', type: args.INT, value: 1 })
return p
}
var parsed = parser().parse()
var name = parsed.indexes[0]
iter var i = 0; i < parsed.options.times; i++ {
echo 'Hello, ${name}!'
}
The raw list is there as well. os.args has the same shape for a
command as for zuri run: the executable, the command’s own file, then
what the user typed, so the command’s arguments are os.args[2,].
Reading that list by hand is how the checks and the help text drift
apart, which is exactly what the parser exists to prevent, so a command
should reach for args and leave os.args alone.
A command ends with status 0 when it runs to the end. An uncaught error
prints its trace and ends it with status 1, and os.exit() ends it
with any other status. A command that checks something, the way
zuri fmt --dry-run checks formatting, should exit with a non-zero exit
code when the check fails, so a shell script or a CI job can act on it.
Where It Runs
A command runs in the directory zuri was started in, so os.cwd()
is the user’s directory, and for a project command that is the
project’s root. __file__ is the command’s own file, which is how a
command reaches something shipped beside it:
import os
var template = os.join_paths(os.dir_name(__file__), 'notes.template')
Seeing It Work
This example builds a throwaway project with one command, runs it the
way a user would, asks it for its help message, and reads back the
listing zuri --help gives:
import os
var project = os.create_temp_dir('zuri-commands-')
var commands = os.join_paths(project, '.zuri', 'cmds')
os.create_dir(commands, 0c755, true)
var source = file(os.join_paths(commands, 'greet.zu'), 'w')
source.write(
'/**\n' +
' * @command greet\n' +
' * @description Greets whoever is named.\n' +
' */\n' +
'import args\n' +
'var parser = args.Parser("greet", false)\n' +
'parser.description = "Greets whoever is named."\n' +
'parser.add_index("name", "Who to greet", { value: "world" })\n' +
'echo "Hello, " + parser.parse().indexes[0] + "!"\n'
)
source.close()
# Runs `zuri` in the project with `arguments`, and returns what it printed.
def zuri(arguments) {
var child = os.spawn(os.exe_path, arguments, { cwd: project, stdin: 'null' })
child.wait()
return child.read_stdout().to_string()
}
echo zuri(['greet', 'Ada']).trim('\n')
echo zuri(['greet', '--help']).trim('\n')
# The part of the listing this project adds.
var listing = zuri(['--help'])
var start = listing.index_of('PROJECT COMMANDS')
echo listing[start, listing.index_of('\n\n', start)]
os.remove_dir(project, true)
Hello, Ada!
Usage: greet [OPTIONS] [name]
Greets whoever is named.
POSITIONAL ARGUMENTS:
[name] Who to greet (default: world)
OPTIONS:
-h, --help Show this help message and exit
PROJECT COMMANDS:
greet Greets whoever is named.
Testing a Command
Build the parser in a function, as the greet command above does, and
the argument handling can be tested by handing parse() a list, the
seam Chapter 20 describes.
Everything else is ordinary Zuri: keep the command’s logic in the
functions and files beside index.zu, and test those.
The shipped commands keep their tests in a tests directory inside the
command, and a project command can do the same. zuri test runs a
directory anywhere in the project:
zuri test .zuri/cmds/release/tests