Reference
API reference
The protocol steward’s daemon speaks, every method it answers, the shapes of its results, its events and its errors.
Transport
The daemon listens on two Unix stream sockets in
$XDG_RUNTIME_DIR/steward/ (a 0700 directory;
without XDG_RUNTIME_DIR, a private directory under
/tmp):
| socket | for | offers |
|---|---|---|
content.socket |
applications | catalog, content, events, invalidate,
status |
api.socket |
the user’s own tools | everything on content.socket, plus maintenance and settings |
The protocol is JSON-RPC 2.0, one JSON object per
line (UTF-8, \n terminated), in both directions.
→ {"jsonrpc":"2.0","id":1,"method":"stat","params":{"path":"/etc"}}
← {"jsonrpc":"2.0","id":1,"result":{"path":"/etc","kind":"dir", …}}
- Every request with an
idgets exactly one response with thatid. A request without one is a notification and gets none. paramsis an object (named parameters). Methods without parameters accept it absent,nullor{}.- Requests on one connection are handled
concurrently; responses may come back in a different
order. Match them by
id. - A connection that has subscribed also
receives
eventandgapnotifications between responses. - Any request may ask for progress: with
"progress": truein its params, slow work sendsprogressnotifications before the response,{"id": <the request's id>, "stage": "…", "message": "…", …}.messageis for people;stageand the other fields are for programs. Clients that don’t ask get none.locatereportsrescanning,waiting(for a scan already running) andrescanned.
Conventions
- Paths
-
Absolute, without
~, in requests and answers alike. Every path is a JSON string, including names that aren’t valid UTF-8; see File names that aren’t UTF-8. - Content ids
-
btv2:followed by 64 lowercase hex digits, the file’s BEP 52 pieces root. Requests also accept bare hex and either case. - Sizes
-
Bytes.
sizeis the apparent size;allocis space on disk (blocks × 512). - Times
-
Entry.mtimeis in seconds since the Unix epoch. Fields ending in_nsare nanoseconds. Eventtime,started,finishedand the like are seconds as floating point.
File names that aren’t UTF-8
Linux file names are bytes, and some aren’t valid UTF-8: a Latin-1
® (0xAE) from an old Windows share, a file
made on purpose by a test. JSON strings are Unicode. steward sends every
path as a JSON string anyway, and keeps the exact bytes, using the same
convention as Python’s file APIs (“surrogateescape”, PEP 383):
- valid UTF-8 travels as itself (all but a handful of names, on any real system);
- each byte that isn’t part of valid UTF-8 travels as the lone
surrogate U+DC80 + byte, written in JSON as an escape:
0xAEis\udcae.
on disk: /home/me/Shorts/Arch Deluxe\xae.doc
on wire: "/home/me/Shorts/Arch Deluxe\udcae.doc"
The same rule applies to paths you send:
{"path": "/home/me/Shorts/Arch Deluxe\udcae.doc"} names
that file. Every path field in every request, answer and event follows
it.
In Python it needs no work: json.loads turns
\udcae into a lone surrogate in a str,
open(path) and os.stat(path) use it directly,
and os.fsencode(path) returns the exact bytes. The bundled
client accepts paths as str or bytes. In
JavaScript, JSON.parse keeps lone surrogates in strings
too. Elsewhere, decode \udc80–\udcff escapes
to the bytes 0x80–0xff yourself; a JSON parser
that rejects lone surrogates, or replaces them (as jq does,
printing �), loses the original bytes. Never assume a path
is printable.
Errors
{"jsonrpc":"2.0","id":7,"error":{"code":-32000,"message":"/srv is not indexed: no configured root covers it","data":{"type":"not_indexed"}}}data.type is stable and meant for programs.
message is for people and may change.
| code | data.type |
meaning |
|---|---|---|
| -32700 | parse_error |
the line wasn’t JSON (the response has id: null) |
| -32600 | invalid_request |
not a JSON-RPC 2.0 request |
| -32601 | method_not_found |
no such method |
| -32602 | invalid_params |
missing or ill-typed parameters, a malformed content id, a relative path, an unusable piece size |
| -32000 | forbidden |
an administration method sent to content.socket |
| -32000 | not_indexed |
the path is not in the index, or no root covers it |
| -32000 | not_under_root |
the path is outside every configured root (for methods that scan or read it) |
| -32000 | unknown_content |
no such content id |
| -32000 | no_layer |
the content has no stored verification layer |
| -32000 | changing |
the file kept changing while being read |
| -32000 | failed |
anything else; see message |
Types
Entry
One indexed path.
| field | type | |
|---|---|---|
path |
string | absolute |
kind |
string | file, dir, symlink or
other |
mode |
integer | permission bits (0o7777 mask) |
uid, gid |
integer | owner |
size |
integer | apparent size |
alloc |
integer | space on disk |
mtime |
integer | modification time, seconds |
total_size, total_alloc |
integer | subtree totals for directories; the entry’s own figures otherwise |
total_files, total_dirs |
integer | files and directories beneath (a directory counts itself) |
total_items |
integer | every entry beneath, the directory included: files, directories, symlinks, the rest; 1 for non-directories |
tags |
string[] | e.g. classify:repo; omitted when empty |
category |
string | files only: image, video,
audio, document, source,
archive, object, disk-image,
torrent; omitted when none |
content_id |
string | when a current content id is stored; omitted otherwise |
Resolution
What steward knows of one content id (result of resolve).
| field | type | |
|---|---|---|
id |
string | normalised content id |
size |
integer | null | the content’s size, when known |
state |
string | present, offline, absent,
unknown or mismatch |
layer |
boolean | a verification layer is stored (always true up to 1 MiB) |
observations |
Observation[] | current copies, reachable ones first |
Observation: path (string),
inode (string,
"<fsid hex>:<inode>"; equal for hard links),
online (boolean), offline_at (string | null:
the unmounted directory holding it), mtime_ns
(integer).
Inspected
One path after inspect:
path, kind (file,
dir, symlink, other, or null on
error), id (string | null: null for directories, empty
files and non-files), size, error (null, or
{type, message} with type not_under_root,
not_found, changing, unreadable
or invalid_params).
Verdict
What verify found: id
(the id you claimed), path, state
(unchanged, changed, gone,
unreadable, not_file), current
(the id the file holds now, or null).
ScanReport
root, dirs_read, dirs_trusted,
entries_seen, inserted, updated,
deleted, errors (unreadable entries),
millis, offline (directories found on
unmounted volumes), load_ms, write_ms,
totals_ms.
Catalog
Read-only lookups in the index. Both sockets.
status
The daemon’s state. No parameters.
| field | |
|---|---|
db |
index file |
configured |
the configured roots and their policies |
indexed |
an Entry for each indexed root |
scanning |
whether a scan is running |
hashing |
the hashing job, or null: path,
files_total, files_done,
bytes_total, bytes_done (includes bytes read
of files in flight), started |
activity |
scan (path, kind,
secs, or null), reading (files being hashed:
path, size, read,
secs), hash_queue, connections,
subscribers, event_seq |
daemon |
version, pid, started,
uptime_secs, epoch, hash_threads,
db, db_bytes, api_socket,
content_socket, managed |
recent_scans |
the last 50 scans, newest first: ScanReport plus kind
and finished, or root, error,
kind, finished |
schedule |
path and next (seconds since epoch) for
each root |
problems |
the last 200 warnings and errors, newest first: time,
level, context, message |
stat
{path} → Entry, with tags inherited from its
ancestors.
children
{path} → Entry[], the directory’s children, largest
total_alloc first. Each child carries only its own
tags.
locate
{pattern, limit = 1000, mode = "auto", ignore_case = false, kind?, check = "none"}
→ string[] of paths whose final name component matches.
| param | |
|---|---|
mode |
auto (a case-sensitive glob if pattern
contains *, ? or [, otherwise a
substring), substring, exact (the whole name),
glob (the whole name), regex (anywhere in the
name; Rust regex syntax) |
ignore_case |
for exact, glob and regex;
substrings always ignore ASCII case |
kind |
file, dir, symlink or
other |
check |
none: answer from the index. exists:
lstat each result. rescan: also rescan the
folder of each result that’s gone (or its nearest existing parent; at
most 64 at once, the rest queued) and search again |
Results are sorted by their bytes. A rescan has to wait for any scan
already running, such as a root’s scheduled full rescan; with
"progress": true the request reports that as it
happens:
← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"rescanning","message":"3 of 41 results are gone from disk; rescanning 2 folder(s)","stale":3,"folders":2,"queued":0}}
← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"waiting","message":"waiting for the full scan of /home/me in progress (41 s so far) before rescanning /home/me/src","for":"the full scan of /home/me","secs":41}}
← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"rescanned","message":"rescanned /home/me/src in 38 ms","folder":"/home/me/src","ms":38.2}}
← {"jsonrpc":"2.0","id":7,"result":{"paths":[…], …}}
A client can also do the checking itself, with nothing but
content.socket: locate without a check;
lstat each result; pass the ones that are gone to inspect, which rescans each one’s
folder (its own listing, trusting the folders below) and waits; then
locate again. That is exactly what
check: "rescan" does in one request.
With check other than none the result is an
object instead of a list:
{paths, stale, rescanned, queued, limited, search_ms, check_ms, rescan_ms}:
the results that exist, the indexed results that are gone, the folders
rescanned now and those queued (past 64), whether the search stopped at
limit, and the milliseconds spent searching, checking and
rescanning. A malformed glob or regular expression is
invalid_params. exact, substring
and case-sensitive glob are answered by the index;
regex and case-insensitive glob read every
name, which takes a few seconds on a very large index.
find_content
{id} → string[]: the indexed paths currently holding
this content. resolve says more
(reachability, state, sizes).
duplicates
{path, limit = 1000} → groups of identical hashed files
with at least two paths under path, most wasted space
first:
[{"id":"btv2:…","size":1048576000,"wasted":2097152000,"paths":["/a/x.mkv","/b/x.mkv","/c/x (1).mkv"]}]content_summary
{path} → content-id coverage under path:
files, bytes, hashed_files,
hashed_bytes, unhashed_files,
unhashed_bytes, distinct_ids,
duplicate_groups, duplicate_files,
wasted_bytes, truncated (true if it stopped at
2,000,000 files).
Content
Both sockets.
content_id
{path} → content id or null (empty file). Hashes the
file if it has no current id. The path must be a regular file. Prefer inspect, which also updates the index
for the path and takes many at once.
resolve
{contents: [{id, size?}], recheck = false} →
Resolution[], in request order.
size: the size you expect. A different known size givesstate: "mismatch"and no observations.recheck: re-stat each copy before answering. Copies that changed or vanished are dropped, and their directories are queued for a rescan. A copy that can’t be reached because its volume isn’t mounted is reported offline, withoffline_at.- Paths being verified are left out.
→ {"jsonrpc":"2.0","id":2,"method":"resolve","params":{"contents":[{"id":"btv2:1d8e…","size":3145728}],"recheck":true}}
← {"jsonrpc":"2.0","id":2,"result":[{"id":"btv2:1d8e…","size":3145728,"state":"present","layer":true,
"observations":[{"path":"/home/me/a.bin","inode":"9f3c…:1442","online":true,"offline_at":null,"mtime_ns":1790800000000000000}]}]}
inspect
{paths: [string]} → Inspected[], in request order. These
paths may have changed: bring the index up to date for them now, and
give each regular file’s content id.
- A file’s parent directory is rescanned (its own listing read even if its times are unchanged); then the stored id is used if the file’s size and modification time are unchanged, or the file is hashed now.
- A directory is rescanned in full. Files under it are hashed only if
a root’s
contentidpolicy covers them. - Paths must be absolute and under a configured root. Problems are reported per path; the call itself succeeds.
verify
{id, path, reason = ""} → Verdict. There is reason to
believe path no longer holds id. steward
withholds the path from resolve, rescans its directory,
rereads the file in full and records what it holds:
unchanged: it holdsid.changed: it holdscurrent(null for an empty file). Paths that were listed foridgetcontent.lostwith reasonchanged.gone,not_file: nothing hashable is there.unreadable: reading failed. The inode’s stored id is dropped, and its paths getcontent.lostwith reasonunreadable.
reason is logged, never interpreted. Errors:
invalid_params, not_under_root,
changing (still changing after three reads).
piece_layer
{id, piece_size} →
{id, size, piece_size, layer}. The content’s BEP 52 Merkle
layer at piece_size (a power of two, at least 1 MiB),
derived from the stored 1 MiB layer without reading the file.
layer is the concatenated 32-byte hashes in hex, one per
piece, and empty for content no bigger than one piece. Errors:
unknown_content, no_layer,
invalid_params.
invalidate
{path} → "queued". Something under
path changed. The daemon gathers invalidations for two
seconds, then rescans the nearest existing directory of each (in
full).
Events
subscribe
{since?, ids?} → {epoch, seq, complete}.
From now on this connection also receives events.
since: first replay the backlog’s events after thisseq.ids: delivercontent.*events only for these content ids.storage.*events always arrive.epochchanges each time the daemon starts;seqis the last event’s number in it.completeis false when events aftersincewere missed: the backlog (the last 4,096 events) no longer reaches back that far, orsinceis ahead ofseq. The daemon can’t tell asincefrom an earlier run, so compareepochwith the one you saw before: a different epoch means you missed everything in between.
Subscribing again replaces the filter. The response comes before any replayed event.
unsubscribe
{} → true. Stop delivering events on this
connection.
Notifications
{"jsonrpc":"2.0","method":"event","params":{"seq":42,"time":1790812345.12,"name":"content.moved","data":{"id":"btv2:…","from":"/a/x.mkv","to":"/b/x.mkv"}}}
{"jsonrpc":"2.0","method":"gap","params":{"after":41}}name |
data |
meaning |
|---|---|---|
content.observed |
{id, path} |
the content was seen at this path (new, renamed onto, or hashed) |
content.moved |
{id, from, to} |
the same inode moved from one path to another |
content.lost |
{id, path, reason} |
the path no longer holds this content: deleted,
changed or unreadable |
storage.offline |
{path} |
this known directory’s volume is not mounted |
storage.online |
{path} |
it is mounted again |
storage.unindexed |
{path} |
this root was removed from the configuration |
A gap notification means this subscriber fell behind and
events after after were dropped for it. After a gap, an
incomplete subscribe, or a new epoch, re-resolve whatever you track.
Maintenance
api.socket only.
scan
{path, trust_dir_mtime = false} → ScanReport. Rescan
path (a root or anything under one) now and wait.
trust_dir_mtime skips directories whose times show nothing
changed. Error: not_under_root.
classify
{path} → {scanned, tagged}. Re-run
classification, from the enclosing repository if there is one.
hash_tree
{path} → {stale, hashed}. Hash every file
under path without a current content id, and wait. For a
media folder, this can take hours.
export_qdirstat
{path, out} → {entries, out}. Write
path’s subtree as a gzipped qdirstat 2.0 cache file at
out. Refuses to overwrite an existing file.
Settings
api.socket only. Changes are validated, written into
settings.toml (comments and formatting kept), and applied
at once.
settings
No parameters →
{file, db, roots, scanning, hashing, hash_threads, hash_threads_default}.
Each of roots is
{settings, indexed, offline, fs}: the root’s policy, an
Entry for it (or null before its first scan), whether its volume is
offline, and the capacity of the filesystem it is on (null when
offline):
fs field |
|
|---|---|
type, mount |
filesystem type and mount point |
bytes_total, bytes_free |
space, as df reports it |
inodes_total, inodes_free |
inodes, as df -i reports them; null where inodes are
allocated on demand (btrfs) |
metadata_total, metadata_used |
btrfs only: metadata space allocated, and used. Many small files
exhaust this while df still shows free space |
put_root
{root} → reload result. Add a root, or replace the one
with the same path. root has the configuration keys:
path is required, the rest default. Validation: the path
must be an existing absolute directory, interval_minutes at
least 1, exclude patterns valid, content-id folders inside the root.
remove_root
{path} → reload result. Stop indexing a root and drop
its entries.
reload
No parameters → {added, changed, removed}. Re-read
settings.toml, start new roots, rescan changed ones, drop
removed ones.