Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Binary Data and Streams

Text is convenient. Protocols, file formats, images and checksums are not text, and this chapter is about the four tools Zuri gives you for them: the bytes type, the struct module, io.BytesIO, and compress.

The bytes Type

bytes is a mutable buffer of 8-bit values.

var b = bytes(3)
var c = bytes([1, 2, 3, 4, 5])
var d = 'Hi'.to_bytes()

echo b
echo c
echo d
(00 00 00)
(01 02 03 04 05)
(48 69)

bytes(n) allocates n zero bytes. bytes(list) builds one from numbers, each of which must be in 0..256. to_bytes() on a string gives you its UTF-8 encoding.

The hexadecimal-in-parentheses rendering is how bytes always prints, which makes it impossible to confuse one with a list at a glance.

Reading

var b = bytes([1, 2, 3, 4, 5])

echo b.length()
echo b[0]
echo b[1, 3]
echo b.first()
echo b.last()
echo b.get(2)
echo b.index_of(3)
echo b.last_index_of(3)
echo b.to_list()
echo 'Hi'.to_bytes().to_string()
5
1
(02 03)
1
5
3
2
2
[1, 2, 3, 4, 5]
Hi

Indexing gives a number. Slicing gives bytes. to_string() decodes as UTF-8, to_list() gives numbers.

index_of() and last_index_of() both return -1 when the byte is not there, and both take a second argument bounding where a match may sit, so they search the two halves either side of one index.

Slicing

b[a, b] takes the bytes from a up to but not including b, and returns a new byte stream:

var b = bytes([10, 20, 30, 40, 50])

echo b[1, 3]
echo b[, 3]
echo b[3, ]
echo b[-2, ]
(14 1e)
(0a 14 1e)
(28 32)
(28 32)

The rules are exactly the list’s. Either bound may be omitted: b[, n] starts at the beginning and b[n, ] runs to the end. Negative bounds count back from the end, so b[-2, ] is the last two bytes.

Note the difference between an index and a slice, because for bytes the two return different types:

var b = bytes([10, 20, 30])

echo b[0]
echo typeof(b[0])
echo b[0, 1]
echo typeof(b[0, 1])
10
number
(0a)
bytes

One index gives you the numeric value of that byte. A slice of length one gives you a byte stream containing it. Reaching for b[0] when you meant b[0, 1] is the most common slip here, and it shows up as a number where a bytes was expected rather than as an error at the slicing site.

Bounds Are Checked

A slice that runs past the end raises rather than returning what it can:

var b = bytes([10, 20, 30, 40, 50])

catch {
  echo b[1, 99]
} as e {
  echo '${e.type}: ${e.message}'
}
RangeError: slice bounds 1..99 out of range (length 5)

length() itself is always a legal upper bound, because the bound is exclusive, and an empty slice is legal rather than an error:

var b = bytes([10, 20, 30, 40, 50])

echo b[0, b.length()]
echo b[2, 2]
echo b[2, 2].is_empty()
(0a 14 1e 28 32)
()
true

An empty byte stream prints as ().

A Slice Is a Copy

Slicing allocates a new stream, so writing through one does not disturb the original:

var original = bytes([10, 20, 30])
var part = original[0, 2]

part[0] = 99

echo original
echo part
(0a 14 1e)
(63 14)

That matters when you are parsing a buffer. Pulling a header out with frame[0, 4] gives you something you can modify freely, and the frame you are still reading from is untouched. It also means slicing in a loop copies every time, so a parser that walks a large buffer should carry an offset and slice once per field rather than re-slicing the remainder each step.

Slice, Then Decode

The common shape when a buffer holds text with a known extent:

var b = bytes([72, 101, 108, 108, 111, 33])

echo b[0, 5].to_string()
echo b.to_string()
Hello
Hello!

to_string() decodes the whole stream it is called on, so the slice is what limits the extent. Slicing on a byte boundary in the middle of a multi-byte character produces a stream that is not valid UTF-8; decode whole units, or keep the tail for the next read.

There Is No Slice Assignment

A slice can be read but not written to:

var b = bytes([10, 20, 30])

b[0, 2] = bytes([1, 1])
SyntaxError: invalid assignment target

Assign to one index at a time, or rebuild the stream with extend(). A single index does accept assignment, and the value has to be a real byte:

var b = bytes([72, 101, 108, 108, 111, 33])

b[0] = 74

echo b.to_string()

catch {
  b[1] = 300
} as e {
  echo '${e.type}: ${e.message}'
}
Jello!
NumericError: bytes element must be an integer in 0..=255, got 300

Note the difference from bytes([300]), which wraps rather than raising. Construction is lenient; assignment is not.

Writing

var b = bytes([1, 2, 3])

b.append(4)
b.extend(bytes([9]))
echo b
echo b.pop()
echo b.reverse()
(01 02 03 04 09)
9
(04 03 02 01)

Unlike a list, bytes.reverse() mutates in place. So do append, extend, pop and remove.

dispose() releases the buffer’s memory immediately rather than waiting for the collector, which matters when you have just finished with something very large.

Splitting

split() takes a bytes delimiter, not a number:

echo bytes([1, 2, 3, 4, 5]).split(bytes([3]))
[(01 02), (04 05)]

Walking

A byte stream is iterable, and every form that works on a list works here. What you get out is always a number between 0 and 255, never a one-character string.

var b = bytes([72, 105])

for value in b {
  echo value
}
72
105

Two variables give you the index first and the value second:

var b = bytes([72, 105])

for index, value in b {
  echo '${index}: ${value}'
}
0: 72
1: 105

iter is the form to use when you are decoding a structure and the position drives the walk — reading a two-byte length, then skipping that many bytes, then reading the next field:

var b = bytes([72, 105, 33])

iter var i = 0; i < b.length(); i += 2 {
  echo '${i}: ${b[i]}'
}
0: 72
2: 33

And each() takes a function, with the value first and the index second as everywhere else:

bytes([72, 105]).each(@(value, index) {
  echo '${index}=${value}'
})
0=72
1=105

When you want the characters rather than the numbers, convert first: b.to_string() decodes the whole stream as UTF-8, and b.to_list() gives you the numbers as an ordinary list.

Binary Files

Append b to the file mode and reads give you bytes instead of a string, and writes accept bytes:

var out = file('data.bin', 'wb')
out.write(bytes([1, 2, 3]))
out.close()

echo file('data.bin', 'rb').read()
(01 02 03)

Reading a non-UTF-8 file without the b fails rather than silently producing replacement characters, which is the behaviour you want: a decoding failure is information.

struct: Packing and Unpacking

struct converts between Zuri values and a fixed binary layout. A format string is a sequence of /-separated fields, each CODE[COUNT][:NAME].

import struct

var header = struct.pack('N:magic/n:version/Z8:name/C:flags',
  0x5A555249, 1, 'task', 3)

echo header
echo header.length()
echo struct.unpack('N:magic/n:version/Z8:name/C:flags', header)
(5a 55 52 49 00 01 74 61 73 6b 00 00 00 00 03)
15
{magic: 1515541065, version: 1, name: task, flags: 3}

unpack() hands back a dictionary keyed by the :NAME you gave each field. Fields with no name get numbered:

echo struct.unpack('N/N', struct.pack('N2', 7, 8))
{1: 7, 2: 8}

A count greater than one under a single name numbers the keys:

echo struct.unpack('n3:vals', struct.pack('n3', 1, 2, 3))
{vals1: 1, vals2: 2, vals3: 3}

The Format Codes

Strings:

CodeSizeMeaning
acountNUL-padded string
Acountspace-padded string
ZcountNUL-padded and NUL-terminated, like C
h / Hceil(count/2)hex string, low or high nibble first

Integers. The letter tells you the width and the byte order:

CodeSizeMeaning
c / C1signed / unsigned 8-bit
?1boolean
s / S2signed / unsigned 16-bit, native order
n / v2unsigned 16-bit, big- / little-endian
i l / I L4signed / unsigned 32-bit, native order
N / V4unsigned 32-bit, big- / little-endian
q / Q8signed / unsigned 64-bit, native order
J / P8unsigned 64-bit, big- / little-endian
u / U16signed / unsigned 128-bit, little-endian

Floats:

CodeSizeMeaning
f / g / G4float: native / little / big endian
d / e / E8double: native / little / big endian
w / W2half-precision: little / big endian

Padding:

CodeMeaning
xwrite count NUL bytes, consumes no argument
Xback up count bytes
@seek to absolute position count

n, N and J are the network-order codes, which is what you want for almost every wire protocol.

A count of * means “the rest”.

Precision

Zuri numbers are doubles, so they hold integers exactly only up to 2^53. Every 64-bit and 128-bit code (q, Q, J, P, u, U) automatically produces a bigint when the unpacked value falls outside that range, rather than quietly losing digits. Packing accepts either.

The Rest of the Module

calcsize(format) gives the fixed byte size of a format with no * in it. pack_into(buffer, offset, format, ...) and unpack_from(format, buffer, offset) work on an existing buffer at a position, which is how you build one large frame without allocating and concatenating per field.

io.BytesIO

BytesIO is a file-shaped object backed by memory. It implements the same interface file() handles do, which means anything that takes a file takes a BytesIO too:

import io

var buffer = io.BytesIO(bytes(0), 'w')

buffer.write('hello ')
buffer.write('world')

echo buffer.source.to_string()
hello world

The constructor takes a bytes source and a mode. It has read, gets, write, puts, seek, tell, flush, close and stats, so a function written against files needs no changes to work in memory.

That is the useful part: test a function that writes a file without touching the disk, and parse an in-memory buffer with code written for a stream.

Compression

The compress module has a submodule per format, each with compress() and decompress():

import compress

var raw = ('the quick brown fox ' * 20).to_bytes()

echo raw.length()
echo compress.gzip.compress(raw).length()
echo compress.deflate.compress(raw).length()
echo compress.zstd.compress(raw).length()
echo compress.brotli.compress(raw).length()
echo compress.bzip2.compress(raw).length()
echo compress.lz4.compress(raw).length()
400
47
29
41
31
75
37

The formats are deflate, zlib, gzip, zstd, lz4, bzip2 and brotli. zlib is re-exported at the top level, so compress.compress() and compress.decompress() are the zlib pair:

echo compress.decompress(compress.compress(raw)).to_string() == raw.to_string()
true

Which to reach for: gzip when something else has to read it, zstd when you want the best ratio-to-speed trade, lz4 when speed is the only thing that matters, brotli for text you will serve over HTTP.

Archives

compress.tar and compress.zip read and write archives, and compress.checksum has crc32() and adler32():

import compress

echo compress.checksum.crc32('hello'.to_bytes())
907060870

Hashing and Encoding

base64 moves binary through text channels:

import base64

var encoded = base64.encode('hello'.to_bytes())

echo encoded
echo base64.decode(encoded).to_string()
aGVsbG8=
hello

hash covers the digest algorithms:

import hash

echo hash.sha256('hello')
echo hash.hmac_sha256('key', 'hello')
2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
9307b3b915efb5171ff14d8cb55fbcc798c6c0ef1456d66ded1a6aa723a58b7b

Every function takes an optional final as_bytes argument; pass true to get raw bytes instead of a hex string.

The algorithms are md2, md4, md5, sha1, sha224, sha256, sha384, sha512, the sha3 family, shake128/256, blake2b512, blake2s256, ripemd160, whirlpool and gost, each with an hmac_ counterpart, plus pbkdf2() for key derivation. hash.hash(algorithm, data) selects by name at runtime.

For password storage, use bcrypt rather than any of these. A fast hash is the wrong tool for a password, and Chapter 13 covers the right one.

A Worked Example: A Length-Prefixed Frame

Most binary protocols are a header followed by a payload. Here is the whole round trip:

import struct
import compress

var HEADER = 'N:length/C:compressed'

def encode_frame(payload) {
  var body = payload.to_bytes()
  var compressed = 0

  if body.length() > 128 {
    body = compress.gzip.compress(body)
    compressed = 1
  }

  var frame = struct.pack(HEADER, body.length(), compressed)
  frame.extend(body)

  return frame
}

def decode_frame(frame) {
  var size = struct.calcsize('N/C')
  var header = struct.unpack(HEADER, frame[0, size])
  var body = frame[size, size + header.length]

  if header.compressed == 1 {
    body = compress.gzip.decompress(body)
  }

  return body.to_string()
}

var frame = encode_frame('ping')
echo frame
echo decode_frame(frame)

echo decode_frame(encode_frame('the quick brown fox ' * 20)).length()
(00 00 00 04 00 70 69 6e 67)
ping
400

calcsize() tells you where the header ends, slicing gives you the two halves, and the compressed flag is one byte because a protocol that has to guess is a protocol that breaks.