Files and the Filesystem
Two things do the work here. The built-in file() function gives you a
handle to one file. The os module covers everything else: directories,
paths, globs, temporary files and the environment.
Opening a File
var handle = file('notes.txt')
var writer = file('notes.txt', 'w')
file(path) defaults to read mode. The second argument is the mode:
| Mode | Meaning |
|---|---|
r | read; the file must exist |
w | write; creates the file, truncates an existing one |
a | append; writes always go to the end, creates the file |
r+ | read and update; the file must exist |
w+ | read and update; creates the file, does not truncate |
a+ | read and append; creates the file |
x | write; creates the file, fails if the path already exists |
x+ | read and write; creates the file, fails if the path already exists |
Append b to any of them for binary mode: 'rb', 'wb', 'ab'.
x is the one to reach for when two programs might create the same file.
The existence check and the creation happen in a single step, so exactly
one of them succeeds and every other one fails, which is what a lock file
needs.
Creating a handle does not touch the disk. Nothing happens until you read,
write or call open().
Reading a Whole File
file('notes.txt', 'w').write('It works!')
echo file('notes.txt').read()
It works!
read() with no argument opens the file, reads all of it, and closes it
again. That is the one-liner for “give me this file’s contents”, and it is
what you want most of the time.
In text mode you get a string, decoded strictly as UTF-8. In binary mode
you get bytes.
Reading in Chunks
read(length) reads at most that many bytes and leaves the handle open, so
you can call it again. This is how you process a file too large to hold in
memory at once:
file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')
var handle = file('notes.txt')
handle.open()
var chunks = 0
while true {
var chunk = handle.read(6)
if chunk.is_empty() {
break
}
chunks++
}
handle.close()
echo chunks
3
Six bytes at a time is a demonstration; in real code the chunk is tens of kilobytes. The shape is what matters: open once, read until you get an empty result, close once.
Note that chunks fall wherever the byte count lands, not on line boundaries. A chunked reader that needs whole lines has to keep the tail of each chunk and join it to the front of the next.
Reading Lines
For text you want line by line, the simplest form reads the file and splits it:
file('notes.txt', 'w').write('alpha\nbeta\ngamma\n')
for line in file('notes.txt').read().lines() {
echo '[${line}]'
}
[alpha]
[beta]
[gamma]
lines() handles both \n and \r\n, and drops the trailing empty
piece a final newline would otherwise produce.
Writing
file('notes.txt', 'w').write('It works!')
Like read(), write() opens the handle if it is closed, writes, flushes,
and closes it again. One call, one complete file.
That auto-close has a consequence worth being precise about. Two
consecutive write() calls on a closed handle in w mode each truncate
the file, so only the last one survives. When you are writing more than
once, open the handle yourself:
var handle = file('notes.txt', 'w')
handle.open()
handle.write('first line\n')
handle.write('second line\n')
handle.close()
echo file('notes.txt').read()
first line
second line
The explicit open() is what keeps the handle open across both writes.
Without it, each write() would open, truncate, write and close, and only
second line would survive.
An already-open handle is written to where it stands, so the sequence above does exactly what it reads like.
puts() writes without ever opening or closing. It requires an open
handle, and it is the method to reach for inside a loop.
Reading and Writing at a Position
var handle = file('notes.txt')
handle.open()
handle.seek(6, 0)
echo handle.tell()
echo handle.read(4)
handle.close()
seek(offset, whence) takes 0 for the start of the file, 1 for the
current position and 2 for the end. The io module names them:
import io
handle.seek(0, io.SEEK_SET)
handle.seek(-10, io.SEEK_END)
tell() reports the current offset.
Asking About a File
var handle = file('notes.txt')
echo handle.exists()
echo handle.path()
echo handle.abs_path()
echo handle.name()
echo handle.mode()
echo handle.is_open()
echo handle.is_closed()
stats() returns a dictionary of metadata about the file on disk:
file('notes.txt', 'w').write('alpha\nbeta\n')
var info = file('notes.txt').stats()
echo info.size
echo info.is_readable
echo info.keys()
11
true
[is_readable, is_writable, is_executable, is_symbolic, size, mode, dev, ino, nlink, uid, gid, mtime, atime, ctime, blocks, blksize]
size is in bytes. mtime, atime and ctime are epoch seconds, ready
to hand to the date module. mode is the raw permission-and-type word,
and the stat module is what turns it into an answer:
import stat
file('notes.txt').chmod(0c644)
var info = file('notes.txt').stats()
echo stat.S_ISREG(info.mode)
echo stat.S_ISDIR(info.mode)
echo stat.file_mode(info.mode)
true
false
-rw-r--r--
S_ISREG, S_ISDIR, S_ISLNK and the rest of the family each answer one
question about the kind of entry. file_mode() renders the permission bits
the way ls -l does.
Managing Files
file('notes.txt').copy('backup.txt')
file('backup.txt').rename('archive.txt')
file('archive.txt').delete()
There is also truncate(length), chmod(mode), set_times(access, modify) and symlink(target).
Always Close What You Opened
read() and write() clean up after themselves. Anything you opened with
open() is yours to close, and the cleanest way to guarantee it is a
catch that closes on the way out:
var handle = file(path, 'w')
handle.open()
catch {
write_everything(handle)
} as e
handle.close()
if e {
raise e
}
For something that has to be cleaned up however the program ends, rather
than however one block ends, register it with os.at_exit():
import os
var scratch = file('scratch.txt', 'w')
scratch.open()
os.at_exit(@{
scratch.close()
scratch.delete()
})
scratch.write('working notes')
echo scratch.is_open()
true
Handlers run last registered first, and they run whether the program
reached the end of its script, called os.exit(), or stopped on an
uncaught error. A handler that raises is reported on standard error and
the rest still run, so one failed cleanup cannot cancel the others.
Directories
import os
os.create_dir('sub/deep', nil, true)
The three arguments are the path, the permission bits, and whether to
create intermediate directories. It returns false when the directory
already existed.
echo os.dir_exists('sub')
echo os.is_dir('sub')
echo os.remove_dir('sub', true)
remove_dir’s second argument makes it recursive.
Listing and Globbing
Everything in this section works on a real tree, so build one first:
import os
os.create_dir('tree/sub', nil, true)
file('tree/top.txt', 'w').write('a')
file('tree/sub/nested.txt', 'w').write('b')
file('tree/sub/notes.md', 'w').write('c')
read_dir() lists one directory:
echo os.read_dir('tree')
echo os.read_dir('tree', true)
[., .., sub, top.txt]
[., .., sub, sub/nested.txt, sub/notes.md, top.txt]
Three things to notice. . and .. are included, so a loop over the result
almost always wants to skip them. The recursive form returns nested entries
as paths relative to the directory you asked about, not as bare names,
which is what makes them usable directly. And entries come back sorted by
name, with a directory’s contents following immediately after it, so the
listing reads the same on every machine and filesystem.
glob() is usually what you actually want:
echo os.glob('*.txt', 'tree')
echo os.glob('**/*.txt', 'tree')
[top.txt]
[sub/nested.txt]
* matches within one path segment; ** matches across segments. The
second argument is the base directory, and results come back relative to
it.
Read those two results together, because the distinction catches people
out. *.txt found the file at the top and not the nested one. **/*.txt
found the nested one and not the top-level one, because **/ means
“in a subdirectory”. Neither pattern finds both.
To match at every depth, glob ** and filter:
echo os.glob('**', 'tree')
echo os.glob('**', 'tree').filter(@(p) => p.ends_with('.txt'))
[sub, sub/nested.txt, sub/notes.md, top.txt]
[sub/nested.txt, top.txt]
** on its own matches every entry at every depth, directories included,
which is why the filter is doing real work in the second line. glob()
walks with read_dir() underneath, so matches arrive in that same sorted
order.
Clean up when you are done:
echo os.remove_dir('tree', true)
true
Paths
Every path function is pure string manipulation except where noted:
echo os.join_paths('sub', 'a.txt')
echo os.base_name('sub/a.txt')
echo os.dir_name('sub/a.txt')
echo os.real_path('sub')
echo os.relative_path(os.cwd(), '/full/path/to/sub')
sub/a.txt
a.txt
sub
/home/you/project/sub
sub
real_path() resolves symlinks and requires the path to exist.
abs_path() does not. expand_user() turns a leading ~ into the home
directory. path_contains(base, candidate) answers whether one path is
inside another, which is the check you need before serving a file a user
named.
echo os.cwd()
echo os.home_dir()
os.change_dir('/some/where')
Temporary Files
echo os.temp_dir()
var path = os.create_temp_file('report-', '.csv')
var dir = os.create_temp_dir('build-')
Both create the thing and hand you its path. Clean them up yourself when you are done.
Environment Variables
echo os.get_env('HOME')
echo os.get_env('NOPE', 'fallback')
os.set_env('ZURI_BOOK', '1')
os.unset_env('ZURI_BOOK')
echo os.environ()
echo os.expand_vars('$HOME/projects')
get_env() takes a fallback. environ() gives the whole set as a
dictionary.
Locating Files Relative to Your Code
The current working directory is wherever the user ran zuri from, which
is not where your source lives. Use __file__:
import os
var HERE = os.dir_name(__file__)
var templates = os.join_paths(HERE, 'templates')
Doing this in every module that reads a file next to itself is the difference between a program that works and one that works only from the project root.
A Worked Example
Counting words across every text file in a directory tree, start to finish:
import os
def text_files(root) {
return os.glob('**', root).filter(@(p) => p.ends_with('.txt'))
}
def word_count(root) {
var counts = {}
for path in text_files(root) {
var text = file(os.join_paths(root, path)).read()
for word in text.lower().split('/\W+/') {
if word.is_empty() {
continue
}
counts.set(word, counts.get(word, 0) + 1)
}
}
return counts
}
os.create_dir('wc/sub', nil, true)
file('wc/a.txt', 'w').write('Hello world, hello!')
file('wc/sub/b.txt', 'w').write('World of Zuri.')
file('wc/sub/skip.md', 'w').write('not counted')
echo word_count('wc')
os.remove_dir('wc', true)
{hello: 2, world: 2, of: 1, zuri: 1}
Four things in that function are worth pointing at.
glob('**', root) returns paths relative to the root, so they have to
be joined back onto it before they can be opened. Forgetting that join is
the most common mistake in code that globs.
The .md file is absent from the result because text_files() filtered it
out, which is the filter doing the job ** alone cannot.
split('/\W+/') is a regular expression, which is why punctuation does not
end up in the keys — world, and hello! became world and hello. A
plain split(' ') would have kept both.
And counts.get(word, 0) + 1 supplies the starting value for a key that
does not exist yet. Without the fallback, the first sighting of every word
would raise.
Where to Go Next
Binary files, byte streams and the io module’s in-memory files are
Chapter 10. Reading a file over the network is
Chapter 12. The complete list of methods a file
handle carries is Appendix E.