Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

zuri.token

import zuri

Everything here is re-exported by zuri, so import zuri is enough and the names are called as zuri.*. Importing zuri.token on its own works too and reaches the same definitions.

The Token type zuri.tokenize() returns: one entry per lexical token the lexer found in a source file, in source order, with nothing filtered out. Comments, doc blocks, and newlines are all ordinary tokens here, exactly as the lexer itself produces them. The parser’s own grammar quietly skips those when it runs, but tokenize() reports everything the lexer actually saw.

import zuri

for t in zuri.tokenize('var x = 1 + 2') {
  echo '${t.kind} ${t.line}:${t.column} ${t.text}'
}

Functions

tokenize()

zuri.tokenize(source) -> list[Token]

Lexes source in full and returns every token it contains, in source order, as a list of Token.

Nothing is filtered: Comment, DocBlock, and Newline tokens are all included, and the list always ends with one Eof token.

A lexer-level problem in source, an unterminated string, an unbalanced doc block, an unexpected character, does not raise; it surfaces as an ordinary token with kind 'Error' instead, so this function is total over any input, including malformed source.

Parameters

  • source (string) — The Zuri source to lex

Returns list[Token]

tokenize_file()

zuri.tokenize_file(path) -> list[Token]

Reads the file at path and lexes its contents; exactly tokenize(file(path).read()), for the common case of lexing a real file rather than a source string you’ve already got in hand.

Parameters

  • path (string) — Path to the Zuri source file to lex

Returns list[Token]

Raises Error if path can’t be opened or read

Classes

Token

class zuri.Token

One lexical token from a Zuri source file.

Every instance has:

  • kind (string): the token’s kind, e.g. 'Identifier', 'Plus', 'Comment', 'Eof', the exact variant name from the lexer’s own token kind enum.

  • line / column (number): 1-indexed source position this token starts at; column counts characters, not bytes.

  • start / end (number): character offsets into the source string spanning this token’s exact text.

  • text (string): this token’s exact source text, source[start, end], delimiters included: a string literal’s quotes, a doc block’s opening and closing markers, and so on.

  • value (any): this token’s own payload, where it has one, a literal, identifier, comment, or doc block’s string content, a number, or a bigint. nil for every fixed symbol or keyword, and for Newline/Eof. An Error-kind token instead carries { message, line, offset } describing what went wrong.

  • printable — has a @to_string(), so echo and print() show something useful

Constructor

zuri.Token(data)

Returns a new instance of a Token. Ordinary Zuri code never needs to call this directly, since zuri.tokenize() builds these for you.

Parameters

  • data (dict) — { kind, line, column, start, end, text, value }, the token’s fields

Token.is_trivia()

zuri.Token.is_trivia() -> bool

Is this token a comment or a doc block? Neither ever reaches the parser’s own grammar, but both are real entries in zuri.tokenize()’s output.

Returns bool

Token.to_string()

zuri.Token.to_string() -> string

This token as kind@line:column, e.g. 'Identifier@1:5'.

Returns string


2026, Richard Ore and Zuri contributors