Metaprogramming and Reflection
Zuri’s own compiler is available to Zuri programs. The zuri module hands
you the lexer, the parser and the bytecode compiler as ordinary functions,
alongside a reflection API over live functions, classes, modules and
instances.
That is an unusual amount of the language to expose, so it is worth being clear about what it is for. This is the machinery behind documentation generators, linters, plugin loaders, serialisers, debuggers and editor tooling — programs whose subject is other programs. It is not machinery for ordinary code, and the chapter ends with the one thing it deliberately cannot do.
Two Halves That Cannot Answer for Each Other
The module divides cleanly, and confusing the halves is the most common mistake:
| Reflection | The compiler API | |
|---|---|---|
| Subject | objects that exist right now | source text |
| Entry points | zuri.reflect.* | tokenize(), parse(), compile() |
| Can tell you | a function’s arity, a class’s fields, an instance’s values | a doc block, a comment, a line number, what the compiler emitted |
| Cannot tell you | anything about the source it came from | anything about a running value |
A function’s arity is a runtime fact: the function object carries it, and reflection reads it. That same function’s doc block is a source fact: it exists only in the file, and only the parser can find it. There is no call that crosses the gap, which is why a documentation generator parses files rather than importing them.
One Shape for Everything
tokenize(), parse() and compile() all return the same shape of
result: a flat list of small, uniformly tagged records.
| Function | Returns | Tag | Payload |
|---|---|---|---|
tokenize(source) | Token list | kind | value, text |
parse(source) | Node list | kind | fields |
compile(source) | Instr list | op | fields |
There is one Token class, one Node class and one Instr class — not
one class per token kind, grammar rule or opcode. The grammar has around
fifty shapes and the instruction set around sixty opcodes; a class for each
would be a hundred-odd near-identical classes to keep in step with the
compiler forever.
The cost of that choice is real: nothing tells you which fields a given
kind carries. You cannot autocomplete your way to a Binary node’s
left, op and right. So the module leans on a different habit — run
it and look:
import zuri
echo zuri.parse('var total = a + 1')[0].dump()
Stmt@1:1
statement:
Var@1:1
name: 'total'
name_span: 1:5-1:10
value:
Binary@1:13
left:
Identifier@1:13
name: 'a'
op: 'Plus'
right:
Integer@1:17
value: 1
type_hint: nil
is_constant: false
Everything about that node is in front of you: it is a Stmt wrapping a
Var, the initialiser is a Binary whose op is the string 'Plus', and
the literal 1 is an Integer node rather than a generic literal. Each
header names where that node starts, and name_span says exactly where the
name total is. Nothing had to be looked up.
dump() is the fastest way to learn any node’s shape, and it is the first
thing to reach for whenever you are unsure. zuri.dump_file(path) does the
same for a whole file.
Runtime Reflection
Reflection answers questions about a value you are holding. Every entry
point is under zuri.reflect.
What Kind of Thing Is This?
import zuri
def add(a, b) {
return a + b
}
class Marker {}
echo zuri.reflect.kind(add)
echo zuri.reflect.kind(Marker)
echo zuri.reflect.kind(Marker())
echo zuri.reflect.kind(42)
echo zuri.reflect.kind(print)
function
class
instance
number
function
kind() is typeof()’s more literal cousin. Note the last line: a
built-in native function reports 'function', as does a closure and a
bound method, because from the language’s point of view all three are
interchangeably callable.
Functions
import zuri
def add(a, b) {
return a + b
}
var info = zuri.reflect.function_info(add)
echo info.name
echo info.arity
echo info.variadic
echo info.is_method
echo info.owning_class_name
add
2
false
false
nil
function_info() returns { name, arity, variadic, is_method, owning_class_name, source_path }. The last one is the only field that
reaches back toward the source, and it is a path, not content — to read
the function’s doc block you still have to parse that file.
Classes
import zuri
class Shape {
static var sides = 0
@new(name) {
self.name = name
}
area() {
return 0
}
_secret() {
return 'hidden'
}
@to_json() {
return { name: self.name }
}
}
var info = zuri.reflect.class_info(Shape)
echo info.name
echo info.superclass_name
echo info.fields
echo info.statics
echo info.methods.keys()
Shape
nil
[name]
[sides]
[@new, area, @to_json, _secret]
Four things in that output are worth pausing on.
fields lists name, which was never declared with var — it was
created by self.name = name inside @new. Reflection reports the class’s
real shape, not just what the var lines said.
statics is separate from fields, because a static belongs to the
class and a field belongs to each instance.
Decorated methods appear under their @ names. @new and @to_json
are in methods alongside ordinary ones.
_secret is listed. Reflection sees private members; more on that
below.
Each entry in methods is a full function_info, so
info.methods.area.arity works without another call.
superclass_name is a string or nil, not a class:
import zuri
class Base {}
class Derived < Base {}
echo zuri.reflect.class_info(Derived).superclass_name
echo zuri.reflect.class_info(Base).superclass_name
Base
nil
Instances
import zuri
class Point {
@new(x, y) {
self.x = x
self.y = y
}
distance() {
return (self.x ** 2 + self.y ** 2).sqrt()
}
}
var p = Point(3, 4)
echo zuri.reflect.get_props(p)
echo zuri.reflect.has_prop(p, 'x')
echo zuri.reflect.get_prop(p, 'x')
zuri.reflect.set_prop(p, 'x', 10)
echo p.x
[x, y]
true
3
10
get_props() lists the field names an instance actually has.
get_prop(), set_prop(), has_prop() and del_prop() read and write
them by computed name — the reflective equivalents of the getprop family
of built-ins from Chapter 6.
They cannot create a field. A class is sealed, so set_prop() on a name
the class never declared returns false and changes nothing.
Methods by Name
Three calls turn a method name into something callable:
import zuri
class Greeter {
@new(name) {
self.name = name
}
greet() {
return 'hello ${self.name}'
}
@to_json() {
return { name: self.name }
}
}
var g = Greeter('ada')
echo zuri.reflect.has_method(g, 'greet')
echo zuri.reflect.has_decorator(g, 'to_json')
echo typeof(zuri.reflect.get_method(g, 'greet'))
echo zuri.reflect.bind_method(g, 'greet')()
true
true
function
hello ada
The distinction between the last two matters. get_method() returns the
unbound function; calling it needs the instance supplied yourself.
bind_method() returns it already attached to the instance, so it can
be stored, passed around and called with nothing extra — which is what
makes a dispatch table of methods possible:
import zuri
class Counter {
@new() {
self.n = 0
}
up() {
self.n++
return self.n
}
down() {
self.n--
return self.n
}
}
var counter = Counter()
var actions = {
up: zuri.reflect.bind_method(counter, 'up'),
down: zuri.reflect.bind_method(counter, 'down'),
}
echo actions['up']()
echo actions['up']()
echo actions['down']()
echo counter.n
1
2
1
1
has_decorator() and get_decorator() do the same for @-prefixed
methods, taking the name without the @.
Reflection Sees Private Members
Everything above ignores the leading-underscore rule:
import zuri
class Vault {
var _combination = '1234'
@new() {}
}
echo zuri.reflect.get_props(Vault())
echo zuri.reflect.get_prop(Vault(), '_combination')
[_combination]
1234
This is deliberate, and it is the same decision getprop() makes. A
serialiser has to see every field or it writes an incomplete record; a
debugger has to see every field or it shows a lie. Hiding them would defeat
the only purpose these functions have.
The rule to take from it: reflection is infrastructure, not a way around
encapsulation. Code that reaches for get_prop() to read a private field
it could not otherwise read has not found a loophole; it has written
something the next reader will not expect.
Modules
import zuri
import math
var info = zuri.reflect.module_info(math)
echo info.name
echo info.members.contains('PI')
math
true
module_info() reports a module’s name, its path and the members it
exports. That is enough for a plugin loader: import a directory of modules,
ask each whether it has a known function, and call the ones that do.
The Collector
The garbage collector runs on its own as a program allocates. gc() runs
a full collection on the spot:
import zuri
var scratch = [1, 2, 3]
scratch = nil
zuri.reflect.gc()
echo 'collected'
collected
Every object nothing reaches is freed before gc() returns, and anything
that releases a resource when collected releases it then: an ffi pointer
taken over with own() runs its destructor. The memory it frees goes back
to the system before it returns too. No program needs gc() to
stay correct. It is for the moments timing matters, such as releasing
native resources at a known point, or a test checking what collection
does. A full collection visits every live object, so calling it in a loop
is slow.
Tokens
tokenize() is the first stage: text in, a flat list of lexical tokens
out.
import zuri
for token in zuri.tokenize('var x = 1 # note') {
echo '${token.kind} at ${token.line}:${token.column} -> ${token.text}'
}
Var at 1:1 -> var
Identifier at 1:5 -> x
Equal at 1:7 -> =
Integer at 1:9 -> 1
Comment at 1:11 -> # note
Eof at 1:17 ->
Nothing is filtered. Comments, doc blocks and newlines are all real
tokens, and the list always ends with one Eof. The parser’s grammar skips
trivia when it runs; tokenize() reports everything the lexer saw, which
is exactly what a formatter or a syntax highlighter needs.
Each Token carries:
| Field | What it is |
|---|---|
kind | the lexer’s own variant name — 'Identifier', 'Plus', 'Comment' |
line, column | 1-indexed position; column counts characters, not bytes |
start, end | character offsets spanning the token’s exact text |
text | the exact source text, delimiters included |
value | the payload, where there is one — a literal’s value, an identifier’s name, a comment’s content |
start and end are what make edits possible: they let you recover or
replace a token’s original text without re-lexing, which is how a rename
tool or an automatic formatter works.
is_trivia() is true for comments and doc blocks, so filtering them is one
call:
import zuri
var tokens = zuri.tokenize('var x = 1 # note')
echo tokens.filter(@(t) => t.is_trivia()).map(@(t) => t.kind)
echo tokens.filter(@(t) => !t.is_trivia()).length()
[Comment]
5
Lexing Never Raises
A malformed input does not throw. It produces a token of kind 'Error'
carrying what went wrong:
import zuri
var tokens = zuri.tokenize("var s = 'unterminated")
echo tokens.map(@(t) => t.kind)
echo tokens.find(@(t) => t.kind == 'Error') != nil
[Var, Identifier, Equal, Error, Eof]
true
tokenize() is therefore total over any input at all, which is what
makes it safe to point at a file you did not write, or at a half-typed
buffer in an editor. Neither parse() nor compile() has that property —
both raise on bad input, because neither can produce a meaningful result
from it.
The Syntax Tree
parse() is the second stage: tokens become a tree.
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
echo tree[0].kind
echo tree[0].fields.keys()
echo tree[0].fields.name
echo tree[0].line
Function
[name, name_span, parameters, body, is_variadic]
double
1
A Node has a kind, a position and a fields dictionary. The fields hold
more nodes, plain lists, plain dictionaries or scalars and nothing else, so
walking the tree is uniform no matter which node you are looking at.
Where Each Node Is
A node’s position covers the whole of the source it was read from.
start and end are character offsets, the first character and the one
just past the last, so slicing the source with them gives back the node’s
own text. line and col are where it starts and end_line and end_col
where it ends, all counted from 1:
import zuri
var source = 'def double(n) {\n return n * 2\n}'
var function = zuri.parse(source)[0]
var product = function.fields.body.fields.statements[0].fields.value
echo source[product.start, product.end]
echo '${function.line}:${function.col} to ${function.end_line}:${function.end_col}'
echo source[function.fields.name_span.start, function.fields.name_span.end]
n * 2
1:1 to 3:2
double
A declaration starts at its keyword, so the function starts at def, and
its name has a name_span of its own. That is everything a tool needs to
underline a mistake, jump to a definition or select a whole statement.
Walking It
walk_nodes(tree, visitor) visits every node, depth first:
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
var kinds = []
zuri.walk_nodes(tree, @(node) {
kinds.append(node.kind)
})
echo kinds
[Function, Argument, TypeHint, Any, Block, Return, Binary, Identifier, Integer]
That output repays a second look, because it shows how much the parser
makes explicit. The single unannotated parameter n still produced an
Argument wrapping a TypeHint wrapping an Any — the absence of an
annotation is represented in the tree, not omitted from it. A tool walking
this never has to special-case “no type was written”.
The visitor may also be a dictionary from kind to handler, in which case only matching nodes are called:
import zuri
var tree = zuri.parse('def a() {}
def b() {}
var c = 1')
var names = []
zuri.walk_nodes(tree, {
Function: @(node) {
names.append(node.fields.name)
},
})
echo names
[a, b]
find_nodes(tree, kind) is the shortcut when you want one kind and nothing
else:
import zuri
var tree = zuri.parse('def double(n) { return n * 2 }')
echo zuri.find_nodes(tree, 'Return').length()
echo zuri.find_nodes(tree, 'Binary')[0].fields.op
1
Multiply
Comments Survive
This is the property that makes documentation tooling possible. Comments and doc blocks are kept in the tree, in their original position, as their own nodes:
import zuri
var source = "# a note
def add(a, b) {
return a + b
}
"
echo zuri.parse(source).map(@(n) => n.kind)
[Comment, Function]
A doc block sits as a sibling immediately before whatever it documents, so pairing them is one pass with one variable of state — which is exactly what the worked example below does.
Most parsers throw comments away. Keeping them is what separates a parser you can build a formatter or a documentation generator on from one you can only build an interpreter on.
Bytecode
compile() is the third stage: the tree becomes VM instructions.
import zuri
for instr in zuri.compile('var a = 1 + 2') {
echo '${instr.op} from line ${instr.line}'
}
LoadConst from line 0
AddImm from line 1
SetGlobal from line 1
LoadNil from line 1
Return from line 1
An Instr has an op, a line and a fields dictionary — the same shape
a Node has — and walk_instrs() and find_instrs() mirror the AST
walkers exactly.
Notice AddImm. The compiler folded the constant 2 into the add
instruction rather than loading it separately. That is the sort of question
only the bytecode can answer, and reading it is the most direct way to find
out what the compiler actually made of something:
import zuri
echo zuri.compile('var a = 2 * 3 ** 2').map(@(i) => i.op)
[LoadConst, LoadConst, LoadConst, Pow, Mul, SetGlobal, LoadNil, Return]
The power comes before the multiply, which is ** binding tighter than
*. One line of bytecode settles an argument that reading the expression
does not.
Resolved Operands
Several opcodes carry only a raw index into the compiler’s constant pool, which on its own tells you nothing. Rather than make every caller fetch and correlate a constants table, each such field arrives with its resolved value alongside the index:
import zuri
var load = zuri.find_instrs(zuri.compile('var a = 42'), 'LoadConst')[0]
echo load.fields.keys()
echo load.fields.value
[dst, const_idx, value]
42
const_idx is the raw index, value is what it points at, and dst is
the register the result lands in. The same
pairing appears on GetGlobal, GetField, Invoke and the rest.
A Closure instruction’s resolved value is a nested function prototype
({ name, arity, variadic, instructions }) whose own instructions are
wrapped the same way, recursively — so compiling a file with functions in
it exposes their bodies too, not only the code around them.
Line Numbers Are Statement-Grained
Every instruction carries the source line it came from, but never a
column. The compiler’s own tracking is line-only. A Token and a Node
both have real column information from the lexer; an Instr does not.
Checking Source Without Running It
parse() and compile() raise on the first sign of bad source, which is
right for a tool that needs the whole program. A tool working on source as
it is being written, half a statement at a time, needs the opposite:
everything that can be read, and an exact account of what is wrong.
parse_partial() reads as much as it can. Where a statement fails, the
parser reports it, skips to where the next statement starts and carries
on, so one mistake costs one statement and is reported once:
import zuri
var result = zuri.parse_partial('var a = 1
var = 2
var c = 3')
echo result.is_clean()
echo result.errors
echo result.nodes.map(@(n) => n.fields.statement.fields.name)
false
[2:5: Variable name expected.]
[a, , c]
Each error is a Diagnostic: a message and the same six-part position a
node has, so it can be underlined exactly. The statement that failed is
still in the tree, its missing name an empty string sitting where the name
belongs.
check() goes one stage further. Source that parses goes on to the
compiler, which catches what the grammar cannot see, such as a break
with no loop around it or a name declared twice in one scope:
import zuri
var source = 'def total(items) {
var sum = 0
for item in items {
sum += item
}
break
}'
for problem in zuri.check(source) {
echo problem
echo source[problem.start, problem.end]
}
6:3: 'break' used outside of a loop
break
An empty list from check() means the source compiles. Neither function
runs anything or raises on bad source, so both can be pointed at whatever
an editor’s buffer holds at the time.
Reading a File
Each of these has a _file counterpart that reads the path first:
| Source string | File |
|---|---|
tokenize(source) | tokenize_file(path) |
parse(source) | parse_file(path) |
compile(source) | compile_file(path) |
parse_partial(source) | parse_partial_file(path) |
check(source) | check_file(path) |
There is a fourth with no string equivalent: dump_file(path) reads,
parses and returns the dump() of every top-level node. It is usually the
fastest possible answer to “what is actually in this file”.
A Worked Example: A Documentation Extractor
Here is the pattern the book’s own reference appendices are built on. It takes source text and returns every documented function in it, pairing each declaration with the doc block above it.
The key fact is that zuri.parse() keeps comments in the tree. A doc block
comes back as its own DocBlock node, sitting as a sibling immediately
before whatever it documents. zuri.doc.attach() walks a list of siblings
and pairs each one with the doc block right above it, and zuri.doc reads
the inside of each block into its prose and its @tags:
import zuri
def documented_functions(text) {
var found = []
for attached in zuri.doc.attach(zuri.parse(text)) {
var node = attached.node
if node.kind == 'Function' and attached.documented {
var params = node.fields.parameters.map(@(p) => p.fields.name)
found.append({ name: node.fields.name, params, doc: attached.block })
}
}
return found
}
var source = "/**\n * Adds two numbers.\n *\n * @param number a\n * @param number b\n */\n" +
"def add(a: number, b: number) {\n return a + b\n}\n\n" +
"def undocumented(x) {\n return x\n}\n"
for fn in documented_functions(source) {
var takes = fn.doc.all('param').map(@(tag) => '${tag.name}: ${tag.type}')
echo '${fn.name}(${', '.join(fn.params)})'
echo ' ' + fn.doc.summary()
echo ' takes ' + ', '.join(takes)
}
add(a, b)
Adds two numbers.
takes a: number, b: number
undocumented is absent from the output: attach() hands it back with
documented false, because no doc block sits above it.
Three things make this approach worth preferring over scanning the text yourself.
The parser knows what a declaration is. A function named def_handler,
a def inside a string literal, a doc block inside a block comment — all of
them fool a text scan and none of them fool the parser.
Parameter names and types come from the real nodes. A parameter’s
type_hint node carries its types and whether it is nullable, so a
generated signature matches what the runtime will actually enforce rather
than what the source happened to look like.
It cannot drift. When the grammar gains something, the parser gains it too, and a tool built this way keeps working.
The tags are read for you. zuri.doc joins a tag’s wrapped lines,
reads its type and the name it documents, and treats @return as
@returns. It is the same reading the book’s own reference is generated
with, so a block reads the same here as it does there.
A Second Example: A Linter
The other half of the module’s use is checking rather than generating. Here is a rule of the kind a team accumulates: flag every function whose parameter list is longer than some limit.
import zuri
def long_signatures(source, limit) {
var offenders = []
zuri.walk_nodes(zuri.parse(source), {
Function: @(node) {
var count = node.fields.parameters.length()
if count > limit {
offenders.append({ name: node.fields.name, line: node.line, count })
}
},
})
return offenders
}
var source = "def small(a, b) {}\n" +
"def large(a, b, c, d, e) {}\n" +
"def also_large(a, b, c, d) {}\n"
for problem in long_signatures(source, 3) {
echo 'line ${problem.line}: ${problem.name}() takes ${problem.count}'
}
line 2: large() takes 5
line 3: also_large() takes 4
Three properties make this worth doing with the parser rather than with a regular expression over the text.
It cannot be fooled by text that looks like code. A function named
def_handler, the word def inside a string literal, a commented-out
declaration — none of them are Function nodes, so none of them are
counted.
It reports the real line. node.line comes from the lexer, so the
message points at the declaration whatever the formatting around it.
It keeps working. When the grammar gains something, the parser gains it too, and a rule written this way does not quietly stop matching.
The dictionary form of the visitor is doing real work here: only Function
nodes reach the handler, so there is no if node.kind == ... and no chance
of matching a kind you did not mean to.
Serialising Without a @to_json
Reflection covers the case where you would otherwise write the same method on every class:
import zuri
import json
class User {
@new(name, age) {
self.name = name
self.age = age
}
}
class Product {
@new(title, price) {
self.title = title
self.price = price
}
}
def to_record(instance) {
var record = {}
for field in zuri.reflect.get_props(instance) {
record[field] = zuri.reflect.get_prop(instance, field)
}
return record
}
echo json.encode(to_record(User('ada', 36)))
echo json.encode(to_record(Product('desk', 120)))
{"name":"ada","age":36}
{"title":"desk","price":120}
One function, every class, no per-class method to keep in step with the
fields. Note that this reads private fields too, which is right for a
debugging dump and wrong for an API response — for the latter, filter on
the leading underscore, or write a real @to_json and let the class decide
what it exposes.
What Else This Is Good For
Plugin systems. Load modules from a directory, ask each whether it has
the function your host expects, and call the ones that do. module_info()
answers the question without importing blindly and hoping.
Editor tooling. tokenize(), parse_partial() and check() never
raise, so they can run against a half-typed buffer. Every token and every
node carries start and end, so a rename or a reformat can rewrite exact
spans, and every problem check() finds can be underlined where it is.
Understanding the compiler. zuri.compile() on a snippet is faster
than reading the compiler’s source, and it can never go out of date with
the compiler you are actually running.
Answering questions about this book. Appendices D and E are generated
from libs/ with exactly the techniques above, and the audit that checks
those stubs is written the same way.
What This Is Not Good For
There is no eval(). You can compile source to bytecode and inspect it;
you cannot execute a string as code.
Zuri leaves it out for security. eval() erases the line between data and
code, and every program that has one eventually runs a string it did not
mean to: a form field, a query parameter, a config value, a webhook
payload. The moment any of those reaches an eval(), whoever supplied it
is running code with your program’s full permissions. It reads your files,
opens your sockets and sends your secrets anywhere it likes. This is the
single most damaging vulnerability class in dynamic languages, and it keeps
happening because the dangerous call looks harmless in review: one
function, three characters of input, and nothing in the language warns you.
So Zuri does not provide one, and every job eval() is usually reached for
has a better answer here:
| Instead of | Use |
|---|---|
| evaluating a user-supplied expression | a dictionary of handlers keyed by the input |
| calling a method whose name you computed | zuri.reflect.bind_method() |
| varying behaviour at runtime | pass a function in |
| turning text into data | json.decode() |
Each of those does the job with a fixed, auditable set of things that can happen. That is the difference between a program that handles input and a program that obeys it.
Classes are sealed, so reflection cannot add a method to one at runtime. Patterns from other languages that rely on monkey-patching do not translate, and the alternative is the one the language wants: express the variation as a subclass, or as a function you pass in.