--- title: Overview hero: true hero-eyebrow: Per-user file index service · Linux hero-title: Know your files. Every path, every copy. hero-lede: steward keeps one current picture of your filesystem, covering every path, how much space it takes, what kind of thing it is, and a content id that follows each file wherever it moves. It shares that picture with your applications over a local socket and never changes a byte of what it indexes. hero-links: - label: Install href: "#install" kind: primary - label: Get started href: getting-started.html kind: secondary - label: Build on steward href: applications.html kind: secondary description: steward is a per-user file index service for Linux. It indexes paths, sizes, classes and BitTorrent v2 content ids, and serves them to applications over a Unix socket. --- ## What steward does Many programs on a desktop walk the same directories over and over: the disk-usage viewer, the file search tool, the backup tool, the photo library, the deduplicator. Each builds its own partial, soon-stale picture. steward does the walking once, keeps the result current, and lets every application ask. ::: cards ::: card [Layer 1]{.label} ### The index Every path under your roots with its `lstat` fields, plus subtree totals for each directory: size, space on disk, file and directory counts. Name search like `locate`, and disk usage like `qdirstat`, from one index. ::: ::: card [Layer 2]{.label} ### Classification What things *are*: git repositories, ignored build output, dependency folders, virtualenvs, caches and trash, plus a category for each file (image, video, audio, document, source, archive…). ::: ::: card [Layer 3]{.label} ### Content ids A content id for each file: the BitTorrent v2 (BEP 52) Merkle root. It names bytes, not paths, so it survives renames and moves, finds duplicates anywhere, and matches the id other software computes for the same bytes. ::: ::: ## Install One command, no root, Linux on x86_64 or ARM64: ```sh curl -fsSL https://toppk.github.io/steward/install.sh | sh steward service install # run the daemon now and at every login ``` The script downloads the latest release for your machine, verifies its SHA-256 checksums and installs `steward` (and, on desktops, `steward-ui`) into `~/.local/bin`. It says whether it installed, upgraded or found the same version already there. It never replaces a program that isn't steward, and doesn't touch your shell configuration. Later, `steward upgrade` does the same and restarts the daemon. Prefer to look first? [Read install.sh](install.sh), or download a binary from [GitHub Releases](https://github.com/toppk/steward/releases/latest) (`steward_linux_amd64` and its `.sha256`, for example), check it with `sha256sum -c`, make it executable and put it on your `PATH`. [Getting started](getting-started.html) covers configuration and building from source. ## What it promises ::: safety **steward only reads what it indexes.** It never creates, modifies, moves or deletes your files. The only things it writes are its own index, its sockets, its settings file when you change settings through it, and export files you ask for (which it will not overwrite). ::: - **Unplugged is not deleted.** When a disk is unmounted, everything that was on it is reported *offline*, not gone. Nothing is dropped, and nothing is rescanned as if 600,000 files had vanished. - **Identity follows the bytes.** A renamed or moved file keeps its content id without being read again. A file whose bytes change loses it. - **Applications get capabilities, not control.** Applications talk to `content.socket`, which offers lookups and content primitives. Changing what is indexed, or how, is administration, on a separate socket. - **No inotify.** steward keeps up with periodic scans that skip unchanged directories, and with applications telling it what they changed. That scales to tens of millions of files with no kernel watch limits. ## How it fits together ::: stack ::: tier [File manager]{.box} [Disk usage viewer]{.box} [Photo library]{.box} [Your application]{.box} ::: [JSON-RPC 2.0 over Unix sockets]{.wire} ::: tier [content.socket [lookups · resolve · inspect · verify · events]{.small}]{.box .socket} [api.socket [roots · scans · settings · maintenance]{.small}]{.box .socket} ::: ::: tier [steward daemon [scheduler · scanner · classifier · hasher]{.small}]{.box .daemon} ::: ::: tier [index [SQLite format, via turso]{.small}]{.box .store} [settings.toml [your roots and their policies]{.small}]{.box .store} ::: ::: The daemon (`steward daemon`) runs as your user, as a systemd user service, at low CPU and I/O priority. The rest of the `steward` command and the `steward-ui` desktop app are clients like any other. ## Who this documentation is for ::: cards ::: card ### People running steward [Getting started](getting-started.html), [Configuration](configuration.html), the [command line](cli.html) and the [desktop app](gui.html). ::: ::: card ### Application developers [Building on steward](applications.html) explains the patterns. The [API reference](api.html) lists every method, result and event. ::: ::: card ### AI agents [For AI agents](agents.html) is a compact operating guide. Every page is also available as Markdown, and [llms.txt](llms.txt) indexes them. ::: ::: ## A first look Once it is installed and running: ```sh steward status # roots, scans, hashing, activity steward tree ~ -d 2 # where the space goes steward locate '*.iso' # find by name, from the index steward inspect ~/Downloads/film.mkv # its content id, now steward dups ~ # identical files, most wasted space first ``` On the author's machine, a first scan of a 15.7-million-entry home directory takes about a minute and a half. A daily rescan of an unchanged `/usr` takes 69 ms, and the daemon runs in about half a gigabyte of memory. --- title: Getting started eyebrow: Start lede: Build steward, run it as a user service, point it at your files and ask it some questions. Ten minutes, most of them spent waiting for the first scan. description: Build, install and configure steward, then run your first queries. --- ## Install steward runs on Linux, on x86_64 and ARM64. Install the latest release into `~/.local/bin`, without root: ```sh curl -fsSL https://toppk.github.io/steward/install.sh | sh ``` The installer downloads the release for your machine, verifies its SHA-256 checksums, and installs: | program | what it is | |---|---| | `steward` | the command line, and the daemon (`steward daemon`) that scans, classifies, hashes and answers requests | | `steward-ui` | the desktop app, installed where the X11/Wayland keyboard libraries are present (`STEWARD_UI=1` or `0` decides) | It only ever replaces earlier steward builds, and never an unrelated program of the same name. It reports whether it installed, upgraded, or found the same version already there, and doesn't change your shell configuration (it warns if `~/.local/bin` isn't on your `PATH`). Set `STEWARD_INSTALL_DIR` to install somewhere else. To inspect first, [read install.sh](install.sh), or download directly from [GitHub Releases](https://github.com/toppk/steward/releases/latest). Each release has `steward_linux_amd64`, `steward_linux_arm64`, `steward-ui_linux_amd64` and `steward-ui_linux_arm64`, each with a `.sha256` file: ```sh curl -fLO https://github.com/toppk/steward/releases/latest/download/steward_linux_amd64 curl -fLO https://github.com/toppk/steward/releases/latest/download/steward_linux_amd64.sha256 sha256sum -c steward_linux_amd64.sha256 install -m 755 steward_linux_amd64 ~/.local/bin/steward ``` Then run the daemon as a systemd **user** service, now and at every login: ```sh steward service install steward service logs -f # watch it work ``` `steward service install` writes `~/.config/systemd/user/steward.service` for the installed `steward`, enables it and starts it. The service runs at `Nice=10` with idle I/O priority, so scans and hashing yield to everything else. To keep it running while you're logged out, enable lingering (`loginctl enable-linger`). **Upgrading** is `steward upgrade`: it installs the latest release the same way and restarts the service. `steward version` shows the version of both the command and the running daemon. ### From source You need Rust 1.97 or newer, [just](https://github.com/casey/just) and, for the desktop app, the GPUI build dependencies (`just deps` installs them on Fedora). Then: ```sh git clone https://github.com/toppk/steward cd steward just install # builds, installs to ~/.local/bin, runs `steward service install` ``` To try it without installing, run `just daemon -v` in one terminal and `just cli …` in another. ## Choose what to index With no settings file, steward indexes your home directory. To choose, write `~/.config/steward/settings.toml`. `just init-config` installs a fully commented example. A typical file: ```toml [[root]] path = "~" exclude = ["/.cache/", "node_modules/"] [[root]] path = "/home/media" classify = false # no git repositories to find here contentid = ["Movies", "TV"] # compute content ids for these folders ``` Each `[[root]]` is a directory tree with its own policy: how often to rescan, what to exclude, whether to classify, which folders get content ids. Apply changes with `steward reload` (or `kill -HUP` the daemon). The desktop app's Settings tab edits the same file and applies it at once. [Configuration](configuration.html) lists every key. ## Watch the first scan The first scan of a root reads every directory under it. Ask for progress: ```sh steward status # JSON: roots, scan in progress, hashing, activity steward settings # each root's policy and index totals ``` or open the desktop app (`steward-ui`) and its **Daemon** tab. When a root has been scanned once, later rescans are quick: by default steward rescans each root once a day, and skips directories whose times show nothing changed inside them. ## Ask some questions ```sh steward tree ~ -d 2 # where the space goes, two levels deep steward ls ~/Downloads # children, largest first steward stat ~/src/steward # one entry, with its tags steward locate '*.iso' # by name: glob, or substring without * ? [ steward locate invoice -l 20 # substring match, first 20 ``` Every answer comes from the index, without touching the disk. `stat` shows tags such as `classify:repo` or `classify:build-output`, inherited from the directory that earned them. ## Content ids Folders listed under `contentid` are hashed in the background after each scan. Any file can also be hashed on demand: ```sh steward inspect ~/Downloads/debian-13.iso # [{ "path": "...", "kind": "file", "id": "btv2:5b1f…", "size": 702545920, "error": null }] steward resolve btv2:5b1f… # where is this content now? steward dups /home/media # identical files, most wasted space first steward content-summary /home/media # how much of it has content ids ``` A content id names the bytes, so it stays the same when the file is renamed or moved, even onto another disk. [Concepts](concepts.html#content-ids) explains what it is and when it changes. ## Without the daemon Every `steward` command can also work directly on an index file, with no daemon involved, using `--db PATH` (or `STEWARD_DB`). This is handy for experiments and for one-off indexes: ```sh steward --db /tmp/usr.db scan /usr steward --db /tmp/usr.db tree /usr -d 1 steward --db /tmp/usr.db export-qdirstat /usr -o /tmp/usr.cache.gz ``` The last command writes a cache file that [qdirstat](https://github.com/shundhammer/qdirstat) opens directly. ## Uninstall ```sh steward service uninstall # stop, disable and remove the unit rm ~/.local/bin/steward ~/.local/bin/steward-ui ``` Your settings (`~/.config/steward/`) and the index (`~/.local/state/steward/`) stay; delete them yourself if you want them gone. --- title: Concepts eyebrow: Start lede: The handful of ideas that explain everything else. These are roots and the index, how steward stays current without inotify, what offline means, classification, content ids, and the observations and events built on them. description: steward's concepts — roots, the index, scans, offline volumes, classification, content ids, observations and events. --- ## Roots and the index A **root** is a directory tree you ask steward to index, together with its policy: how often to rescan it, what to exclude, whether to classify it, which of its folders get content ids. All roots share **one index** and one namespace of absolute paths, so a query about `/home/media/TV` and one about `~/src` go to the same place. If roots are nested, a path belongs to the deepest root that contains it. For every path under a root, the index keeps what `lstat` reports: | kept | not kept | |---|---| | name and parent, type, permissions, owner and group | file contents | | size, and space on disk (blocks × 512) | extended attributes, ACLs | | link count, filesystem id, inode number | access times | | modification time (and change time, for directories) | anything inside archives | Each directory also carries **subtree totals**: total size, total space on disk, and the number of files, directories and **items** beneath it. Items counts every entry: files, directories, symlinks, sockets and the rest. It is what uses up a filesystem's inodes (or, on btrfs, its metadata space), so ranking by items finds where millions of small files pile up. They are kept current as the tree changes, so "how big is this folder" is a lookup, not a walk. A file with several hard links contributes its share of the bytes to each directory that holds a link, so totals don't count the same bytes twice. Items count names, so each link counts once: like `du --inodes -l`, not plain `du --inodes`, which counts a shared inode only once. The index is a single database file in SQLite format, by default `~/.local/state/steward/index.db`. Its size grows with the number of paths, not the size of the files: about 130 bytes per path. ## Staying current without inotify steward does not watch the filesystem with inotify. Watching tens of millions of paths needs one kernel watch per directory, raised system limits, and a full rescan after every overflow or restart anyway. Instead, steward keeps up in four ways: Periodic rescans : Each root is rescanned on its own schedule, once a day by default (`interval_minutes`). Writes are limited to what changed. Trusting rescans : A directory whose modification and change times match the index has the same names in it as last time, so a *trusting* rescan does not read it again. It only descends into its subdirectories. This is the trick `updatedb` uses, and it makes a rescan of an unchanged `/usr` take milliseconds. It cannot see a file rewritten in place (the directory's times don't change), so every `full_every`th rescan stats every file. With the default `full_every = 1`, every scheduled rescan is full. Invalidation : An application that changed files can say so (`invalidate`). steward waits two seconds to gather related changes, then rescans the nearest directory that still exists. Inspection : An application that needs the answer now asks steward to `inspect` specific paths. They are rescanned and, for files, hashed before the call returns. A directory modified within a second of a scan is stored as "untrusted", so the next trusting scan reads it again. A change landing in the same clock tick as the scan cannot be missed for good. ## Filesystems and offline volumes steward records each path's **filesystem id**, which Linux derives from the filesystem's UUID (and the subvolume, on btrfs). It does not use the device number, which can change between boots. When a scan finds a known directory on a *different* filesystem than it was indexed on, that directory's volume isn't mounted: the scan sees the empty mount point instead. steward then reports the directory as **offline**, leaves everything under it exactly as indexed, and does not read or change it. Unplugging a disk is not mistaken for deleting everything on it, and plugging it back in brings everything back, still identified. Offline status shows up in `settings`, in `resolve` results, in `storage.offline` and `storage.online` events, and in the desktop app. By default (`one_filesystem = true`) a root does not descend into other filesystems mounted beneath it. Make another root for each one you want indexed. ## Classification Classification records what things *are*, using names and a little context and never contents. It runs after each scan of a root that has `classify = true`. **Directory classes** are tags on the topmost directory they apply to, and every path beneath inherits them. They are prefixed `classify:`. | tag | meaning | |---|---| | `repo` | the top of a git repository (it contains `.git`) | | `vcs-metadata` | the `.git` directory itself | | `ignored` | a path the repository's `.gitignore` rules ignore | | `build-output` | `target/` beside `Cargo.toml`, `build/` beside a CMake, Gradle or Meson file, `dist/`, `zig-out/`, `.next/`, `CMakeFiles/` … | | `dependencies` | `node_modules/` beside `package.json`, `.terraform/` | | `venv` | a Python virtual environment (`.venv/` or `venv/` containing `pyvenv.cfg`) | | `cache` | `__pycache__/`, `.mypy_cache/`, `.pytest_cache/`, `.gradle/`, `.cache/` … | | `trash` | the desktop trash (`~/.local/share/Trash`, `.Trash-` on volumes) | Names that would be ambiguous alone need a neighbour to confirm them: a folder called `build` is only build output if a build file sits next to it. **File categories** come from the file name alone and are computed when asked, not stored: `image`, `video`, `audio`, `document`, `source`, `archive`, `object`, `disk-image`, `torrent`. ## Content ids A content id names a file's **bytes**, not its path. steward uses the BitTorrent v2 format (BEP 52) and writes it as `btv2:` followed by 64 hex digits: ``` btv2:1d8e17666fc6c23760eb2d8ec0883756d1a0dd34b6cf71292c6f37b38567434f ``` It is the root of a SHA-256 Merkle tree over the file's 16 KiB blocks: the same value BEP 52 calls the file's *pieces root*. Any software that reads v2 torrents, or computes the same tree, gets the same id for the same bytes, and the id doesn't depend on a torrent's piece size. Empty files have no content id, as in BEP 52. ### When an id is computed - In the background, for every file under a root's `contentid` folders, after each scan of that root. - On demand, for any file, through `inspect`, `content_id` or `verify`, or `hash_tree` for a whole folder. Hashing runs on its own thread pool (`hash_threads`, at most four by default) so it never slows scans, and saves its progress as it goes. ### When an id stays and when it goes An id is stored per **inode**, with the size and modification time it was computed under. It is reported only while both still match. - **Renames and moves keep it.** They change neither size nor modification time, so the id moves with the file, and a scan that finds the file at its new path links it at once without reading it again. A move onto another filesystem creates a new inode, which has to be read again before it has an id (the same one), either by its folder's policy or because an application asks. - **Hard links share it.** All paths to one inode show the same id. - **Writes drop it.** Any write changes the modification time, so the old id stops being reported, and the file is hashed again. - **One gap:** bytes changed while size and modification time were restored (`touch -d`, some sync tools) go unnoticed until something reads the file. `verify` exists for exactly this. It rereads the file and records what it really holds. ### The verification layer Alongside each id of a file over 1 MiB, steward keeps the hash of every 1 MiB section of the file: 32 bytes per MiB, under 1 GB for a 28 TiB library. From it steward derives the tree's layer for any power-of-two piece size of 1 MiB or more (`piece_layer`) without reading the file again. That is the same data a v2 torrent stores in its `piece layers` field. ## Observations steward's catalog is a set of **observations**: "content X was seen at path P on filesystem F, inode I". `resolve` turns a content id into its current observations and a **state**: | state | meaning | |---|---| | `present` | at least one copy is reachable now | | `offline` | copies are known, but all of them are on volumes that aren't mounted | | `absent` | steward has seen this content before, but holds no current copy | | `unknown` | steward has never seen this content | | `mismatch` | the caller said to expect a size, and this content has a different one | Copies with the same id are byte-identical, so "which copy" is a choice of convenience (reachable, nearby, on a fast disk), not of correctness. ## Events Applications can subscribe to changes in the catalog instead of polling: | event | when | |---|---| | `content.observed` | content was seen at a new path (a scan linked a hashed inode, or a file was hashed) | | `content.moved` | the same inode was found at a new path in one scan: a rename or move | | `content.lost` | a path stopped holding content: `deleted`, `changed`, or `unreadable` | | `storage.offline` | a known directory's volume is no longer mounted | | `storage.online` | it is back | | `storage.unindexed` | a root was removed from the configuration | Events are numbered, and the daemon keeps the most recent 4,096, so a client that reconnects can resume where it stopped. When that isn't possible (the daemon restarted, or the client fell too far behind) the client is told so explicitly and should re-resolve what it tracks. Without inotify, an event arrives when steward notices the change: at the next scan, invalidation, inspection or verification. ## Two sockets steward listens on two Unix sockets in `$XDG_RUNTIME_DIR/steward/`, and the split is by **capability**: - `content.socket` is for applications: catalog lookups, the content primitives (`resolve`, `inspect`, `verify`, `piece_layer`), and events. Nothing on it changes what is indexed or how. - `api.socket` is administration: everything above, plus adding and removing roots, forcing scans, classification and hashing jobs, settings, and qdirstat exports. The command line tool and the desktop app use `api.socket`. Most applications only need `content.socket`, and that is the one to grant to a sandboxed or less trusted program. --- title: Configuration eyebrow: Use lede: One TOML file says what steward indexes and how. It is yours to edit, and steward edits it too (keeping your comments) when you change roots from the desktop app or the API. description: Every key in steward's settings.toml, and the environment variables steward reads. --- ## Where it lives `~/.config/steward/settings.toml`, or `$XDG_CONFIG_HOME/steward/settings.toml`. Set `STEWARD_CONFIG` to use another file. A missing file is fine: steward then indexes your home directory with default settings. Unknown keys are errors, so a typo is reported instead of silently ignored. ## Applying changes After editing the file, any of these applies it without restarting: ```sh steward reload # or: systemctl --user reload steward # the service maps reload to SIGHUP ``` A reload starts scanning new roots, rescans roots whose settings changed, and **removes roots you deleted from the index** (the files themselves are untouched, of course). The reload result lists what was added, changed and removed. The desktop app's Settings tab, `steward put-root` / `steward remove-root`, and the `put_root` / `remove_root` API methods validate the change, write it into this file with your comments and formatting kept, and apply it at once. ## Top-level keys | key | default | meaning | |---|---|---| | `db` | `~/.local/state/steward/index.db` | The index file. `~` is expanded; `$XDG_STATE_HOME` moves the default. | | `hash_threads` | half the CPUs, at most 4 | Files hashed at once for content ids. Hashing is mostly disk-bound, and a few sequential readers beat many seeking ones on spinning disks. A change applies to the next hashing job. | | `[[root]]` | your home directory | One table per root, below. With no `[[root]]` at all, steward indexes `$HOME`; `root = []` means no roots. | ## Root keys | key | default | meaning | |---|---|---| | `path` | *required* | The directory to index: absolute, or starting with `~`. | | `interval_minutes` | `1440` (a day) | Time between scheduled rescans. At least 1. | | `full_every` | `1` | Every Nth scheduled rescan is full; the rest are trusting. `0` and `1` both mean always full. See [Staying current](concepts.html#staying-current-without-inotify). | | `one_filesystem` | `true` | Don't descend into other filesystems mounted below `path`. | | `exclude` | `[]` | Paths not to index at all, as gitignore patterns relative to `path`. | | `classify` | `true` | Tag repositories, ignored files, build output, caches and the like. | | `contentid` | `[]` | Folders whose files get content ids automatically, relative to `path` or absolute (inside the root). | ### Exclude patterns `exclude` uses `.gitignore` syntax, anchored at the root's `path`: ```toml exclude = [ "/Downloads/incomplete", # leading slash: this exact path under the root "*.iso", # any file or directory with this name, anywhere "node_modules/", # trailing slash: directories only "!/keep/node_modules/", # negation re-includes ] ``` Excluded paths are not scanned, not counted in totals and never hashed. ### Content id folders Hashing reads every byte, so choose the folders where content identity is useful (a media library, downloads, photo archives) rather than all of your home directory: ```toml [[root]] path = "/home/media" classify = false contentid = ["Movies", "TV", "/home/media/Music"] ``` After each scan of the root, files in these folders without a current content id are queued for hashing. Files elsewhere get one only when an application or the command line asks (`inspect`, `cid`, `hash`). ## A complete example ```toml # ~/.config/steward/settings.toml hash_threads = 2 [[root]] path = "~" exclude = ["/.cache/", "/.local/share/Steam/", "node_modules/"] [[root]] path = "/home/media" interval_minutes = 720 # twice a day classify = false contentid = ["Movies", "TV"] [[root]] path = "/mnt/archive" # an external disk; offline when unplugged contentid = ["/mnt/archive"] # everything on it ``` ## Environment | variable | used by | meaning | |---|---|---| | `STEWARD_CONFIG` | daemon, CLI | Path of the settings file. | | `XDG_CONFIG_HOME` | daemon, CLI | Base for the default settings path. | | `XDG_STATE_HOME` | daemon | Base for the default index path. | | `XDG_RUNTIME_DIR` | everything | Where the sockets live (`$XDG_RUNTIME_DIR/steward/`). Without it, a private directory under `/tmp`. | | `STEWARD_DB` | CLI | Work directly on this index file instead of through the daemon (same as `--db`). | | `RUST_LOG` | everything | Override log filtering, e.g. `RUST_LOG=steward_index=trace`. See [Operations](operations.html#logging). | --- title: Command line eyebrow: Use lede: "`steward` asks the daemon questions and gives it instructions. Most answers are JSON, ready for jq. A few commands print a human-readable tree or list instead." description: Every steward command line subcommand and option. --- ## Usage ``` steward [--db PATH] [-v…] [arguments] ``` `steward` is both the daemon and its client. As a client it talks to the running daemon over its administration socket (`$XDG_RUNTIME_DIR/steward/api.socket`). Errors go to stderr as `steward: error: …` with exit status 1. | option | meaning | |---|---| | `--db PATH` | Work directly on this index file, without the daemon (also `STEWARD_DB`). Settings still come from `settings.toml`. | | `-v`, `-vv`, `-vvv` | Log more to stderr: info, debug, trace. Mostly useful with `--db`, where the engine runs inside the command. | Paths may be relative; `steward` makes them absolute before sending them. File names that aren't valid UTF-8 are handled exactly: `locate` prints each path as its raw bytes (so its output can be fed back to `stat`, `xargs -d '\n'` and the like), human views such as `tree` and `ls` show a stray byte as `\xAE`, and JSON output carries it as `\udcae` ([details](api.html#file-names-that-arent-utf-8)). ## Running steward `steward daemon` : Run the daemon in the foreground: what the service runs. `-v` and `-vv` log more. `steward service install | uninstall | start | stop | restart | status | logs [-f]` : Manage the systemd user service that runs `steward daemon`. `install` writes `~/.config/systemd/user/steward.service` for this `steward`, enables it and (re)starts it; `uninstall` stops, disables and removes it. See [Operations](operations.html#the-service). `steward upgrade` : Install the latest release over this one, verified as the installer does, and restart the service. `steward version` : This command's version and the running daemon's (`steward --version` prints just the first). Release builds report their tag, such as `v0.1.0`; builds from source report `dev`. `steward ui [FOLDER]` : Open the desktop app, `steward-ui`. ## Looking around `steward status` : The daemon's state as JSON: configured and indexed roots, whether a scan is running, the hashing job, live activity, recent scans, the schedule and recent warnings. See [Operations](operations.html#status). `steward settings` : The settings file, each root's policy, its index totals and whether it is offline. `steward ls PATH [--by space|items]` : The children of a directory, largest on disk first, with size, share, file count and tags. `steward tree PATH [-d DEPTH] [-t TOP] [--by space|items]` : A qdirstat-style tree, largest first: `DEPTH` levels (default 2), the `TOP` largest children at each level (default 10). ``` $ steward tree /usr/share -d 1 -t 4 13.6G 100.0% ########## 318801 share/ 3.9G 28.6% ### 23851 kicad/ 1.4G 9.9% # 2779 virtio-win/ 987.0M 7.1% # 17019 cursor/ 970.1M 7.0% # 2594 code/ 6.5G ... 469 more ``` The columns are space on disk, share of the parent, a bar, and the number of files beneath. `--by items` ranks by **items** instead: every entry beneath (files, directories, symlinks and the rest), which is what uses up inodes (or btrfs metadata). The first column is then the item count and the last the space on disk: ``` $ steward tree /usr/share -d 1 -t 3 --by items 390.3k 100.0% ########## 13.6G share/ 58.4k 15.0% # 172.9M icons/ 43.3k 11.1% # 655.4M doc/ 37.9k 9.7% # 183.5M help/ 250.7k ... 470 more ``` `steward stat PATH` : One entry as JSON: type, mode, owner, size, space on disk, modification time, subtree totals, tags (including inherited ones), file category and current content id. `steward locate PATTERN [-x | -g | -r] [-i] [-t f|d|l|o] [--check MODE] [-l LIMIT] [-q]` : Paths whose final name component matches. By default, like classic `locate`: a case-insensitive substring, or a case-sensitive glob over the whole name if the pattern has `*`, `?` or `[`. At most `LIMIT` results (default 1000). | option | matches | |---|---| | `-x`, `--exact` | the whole name, literally: `locate -x .git` | | `-g`, `--glob` | the whole name as a glob: `locate -g '*.iso'` | | `-r`, `--regex` | a regular expression anywhere in the name: `locate -r '^IMG_\d{4}\.CR3$'` | | `-i`, `--ignore-case` | ignore case (substrings always do) | | `-t`, `--type` | only files (`f`), directories (`d`), symlinks (`l`) or other (`o`) | `--check` decides what to do about results the index has but the disk doesn't (any more): | `--check` | | |---|---| | `rescan` *(default)* | check each result; rescan the folders of the ones that are gone (each one's folder, or its nearest surviving parent) as one job, then search again, so renamed files show up under their new names | | `warn` | check, leave gone results out and say how many | | `skip` | trust the index; fastest | Results are sorted by their bytes, as `LC_ALL=C sort` would. The paths alone go to stdout. Rescans are reported on stderr as they happen, including any wait for a scan already running (such as a root's daily full rescan), so a slow answer says why: ``` steward: 3 of 41 results are gone from disk; rescanning 2 folder(s) steward: waiting for the full scan of /home/me in progress (41 s so far) before rescanning /home/me/src steward: rescanned /home/me/src in 38 ms ``` When it's done, a report follows: how many results, how long the search, the check on disk and any rescans took, which folders were rescanned or queued, and whether the limit was reached. `-q` turns the report off: ``` $ steward locate joystick.xml /home/me/.config/game/joystick.xml steward: 1 result in 1.62 s: search 1.59 s, check on disk 0.1 ms ``` ```sh steward locate -x .git -t d # every git checkout's .git directory steward locate -g '*.cr3' -i # raw photos, any case steward locate -r '^v\d+\.\d+' --check skip ``` ## Content ids `steward inspect PATH…` : Bring the index up to date for these paths now, and give each regular file's content id, hashing it if needed. Directories are rescanned. `steward cid PATH` : One file's content id, hashing it if needed. `steward resolve ID… [--recheck]` : Where copies of each content id are, whether they are reachable, and the content's state. `--recheck` re-stats each copy first. `steward find ID` : The paths currently holding this content, as a JSON array. `steward dups PATH [-l LIMIT]` : Groups of identical files under `PATH` (among hashed files), most wasted space first. Default limit 50. `steward content-summary PATH` : How many files and bytes under `PATH` have content ids, distinct ids, duplicate groups and reclaimable bytes. `steward hash PATH` : Hash every file under `PATH` that has no current content id. Waits until done: for a large folder, hours. `steward piece-layer ID [-p PIECE_SIZE]` : The BEP 52 piece layer of this content for a power-of-two piece size of at least 1 MiB (default 1 MiB), as hex, without reading the file. ## Keeping current `steward scan PATH [--trust]` : Rescan `PATH` now and wait for it. `--trust` skips directories whose times show nothing changed. The path must be under a configured root. `steward invalidate PATH` : Tell the daemon something under `PATH` changed. It rescans a couple of seconds later; the command returns at once. `steward classify PATH` : Re-run classification under `PATH`. `steward events [--since SEQ] [--id ID]…` : Print content and storage events as they happen, one JSON object per line. `--since` first replays the daemon's backlog after that sequence number; `--id` limits content events to these ids. ## Roots and settings `steward put-root PATH [options]` : Add a root, or replace the one at `PATH`, in `settings.toml` and apply it. Options: `--interval MINUTES` (default 1440), `--full-every N` (default 1), `--cross-filesystems`, `--no-classify`, `--exclude PATTERN` and `--contentid FOLDER` (both repeatable). `steward remove-root PATH` : Remove a root from `settings.toml` and its entries from the index. `steward reload` : Re-read `settings.toml`: start new roots, rescan changed ones, prune removed ones. ## Exporting `steward export-qdirstat PATH -o FILE` : Write `PATH`'s subtree as a gzipped qdirstat 2.0 cache file, which qdirstat opens directly. Refuses to overwrite an existing file. ## Anything else `steward raw METHOD [PARAMS]` : Call any API method with JSON parameters and print the result: ```sh steward raw stat '{"path": "/etc"}' steward raw resolve '{"contents": [{"id": "btv2:1d8e…", "size": 3145728}], "recheck": true}' ``` ## With jq ```sh steward status | jq .activity # what is running now steward settings | jq '.roots[] | {path: .settings.path, offline}' steward dups ~ | jq -r '.[] | "\(.wasted)\t\(.paths[0])"' steward resolve "$id" | jq -r '.[0].observations[] | select(.online) | .path' ``` ## Without the daemon With `--db PATH`, every command runs the engine inside `steward` itself, against that index file. Nothing else needs to be running, and it is a convenient way to build a one-off index of something: ```sh steward --db /tmp/usr.db scan /usr steward --db /tmp/usr.db dups /usr ``` `events` needs the daemon. Don't point `--db` at the index a running daemon is using: the file is locked. --- title: Desktop app eyebrow: Use lede: steward-ui is a window onto the daemon. It shows where your space goes, what is duplicated, how each root is configured, and what the daemon is doing right now. It never touches the disk itself. Everything it shows comes from the daemon. description: steward-ui, the desktop app. Its tabs, its keys, and what each view shows. --- ## Starting it ```sh steward ui # opens your home directory (or run steward-ui directly) steward ui /home/media/TV ``` The path picks the root to show first and the folder to open in it. The app needs a running daemon. Warnings it shows in colour are also written to stderr (`steward-ui: warning: …`), so they can be copied. Across the top are the **roots**, one button each. Hover for a root's totals and how full its filesystem is. A root whose volume isn't mounted says *(offline)* and shows the index as it was last scanned. On the right, badges appear while the daemon is **scanning** or **hashing**. Click either for the details in the Daemon tab. ## Tree A qdirstat-style view of the current root: every directory with its space on disk, its share of the parent, a bar, its file count, and tags such as `classify:repo` or `classify:build-output` (build output, caches and dependencies in amber, trash in red). Children are sorted largest first. **Rank by Space | Items** switches what the tree sorts, bars and percentages by. Items counts every entry beneath (files, directories, symlinks), which is what fills a filesystem's inodes or btrfs metadata. Expanded folders and the selection are kept. Selecting an entry shows its details: path, mode, owner, space on disk and apparent size, counts, and its content id if it has one. **All locations** opens the Content ids tab on every path holding the same bytes. | key | action | |---|---| | ↑ ↓ or k j | move the selection | | PgUp PgDn Home End | move further | | → or l | expand; again to step into the first child | | ← or h | collapse; again to go to the parent | | Enter or Space | expand or collapse | | g | make the selected directory the top of the view | | Backspace or u | go up a level | | r or F5 | rescan the selected directory | | i | rank by space or by items | | / or Ctrl+F | locate: find by name across the index | Locate finds names across the whole index. Beside the field, choose how the pattern matches: **Contains** (a substring, or a glob if it has `* ? [`), **Exact**, **Glob** or **Regex**, and **Aa** to ignore case. Results are checked on disk as they come back: the folders of any that have gone are rescanned and the search repeated, so a file renamed since the last scan shows under its new name. Move through the results with the arrow keys, Enter to reveal one in the tree, Esc to return. ## Content ids For the current root: - **Coverage** of each content-id folder: how many files and bytes have content ids, distinct ids, duplicate groups and reclaimable space, with a **Hash now** button to fill the gaps. - **Look up** any `btv2:` id and list every path holding it. Click one to show it in the tree. - **Duplicates**: groups of identical files, most wasted space first. Click a group for its paths. ## Settings Every root and its policy, edited in place: path, rescan interval, full-scan cadence, one filesystem, classification, exclude patterns and content-id folders. **Save** sends the change to the daemon, which validates it, writes it into `settings.toml` (your comments are kept) and applies it at once. **Add root** creates one, **Remove root** drops a root and its index entries (after you confirm; your files are not touched). ## Daemon The daemon's internal state, as a snapshot. **Refresh** takes a new one, and **Refresh every 2 s** keeps it live. Daemon : Version, process id, uptime, the index file and its size on disk, hashing threads, both socket paths, connected clients and event subscribers. Now : The scan in progress and how long it has run; the hashing job with a progress bar, rate and estimated time left; every file being read at this moment, with how far along it is; and the folders waiting to be hashed. Warnings and errors : The last 200 the daemon logged, newest first, each with where it happened, e.g. `hash{path=/home/media/TV}: …`. Roots : Each root's state (indexed, offline, not scanned yet), its totals, its rescan interval, when the next scan is due, and its filesystem: how full, free inodes where the filesystem has a fixed number, and btrfs metadata use. Anything above 95% is shown in amber. Recent scans : The last 50 scans: when, what, full or trusting, how long, how many entries, and what changed (added, updated, deleted, unreadable, offline). Recent events : The latest content and storage events: files observed, moved and lost; volumes going offline and coming back. ## Everywhere | key | action | |---|---| | Ctrl+1 … Ctrl+4 | Tree, Content ids, Settings, Daemon | | Ctrl++ Ctrl+− Ctrl+0 | larger, smaller, normal size (the keypad keys work too) | | Ctrl+Q | quit | --- title: Operations eyebrow: Use lede: Running steward day to day. This covers the service, what it logs, how to see what it is doing, what it costs, and what to do when something looks wrong. description: Running steward's daemon as a service, logging, status and diagnostics, resource use and troubleshooting. --- ## The service `steward service install` writes a systemd user unit, `~/.config/systemd/user/steward.service`, that runs `steward daemon` from wherever `steward` is installed, then enables and starts it: ```sh steward service install # write, enable, (re)start steward service status # systemd's view: running since, recent lines steward service logs -f # the journal, following steward service restart # also: start, stop steward service uninstall # stop, disable, remove the unit ``` The unit runs the daemon with `Nice=10` and `IOSchedulingClass=idle`, so scans and hashing only use the disk when nothing else wants it. It is a per-user service: one daemon per user, indexing what that user can read, answering only that user (the sockets are in a `0700` directory). It starts with your first login session; `loginctl enable-linger` keeps it running when you're logged out. `steward service` only ever rewrites or removes a unit it wrote itself. `steward daemon` runs the daemon in the foreground, which is the quickest way to watch a first scan (`-v` for debug, `-vv` for trace): ```sh steward daemon -v ``` ### Upgrading ```sh steward upgrade ``` That installs the latest release over the current one (verifying its checksums, as the installer does) and restarts the service, so the daemon and its clients always speak the same protocol. `steward version` shows both versions. If you installed from source, `just install` does the same. ## Logging steward's programs log one line per event to stderr: ``` 21:44:02.118 steward: scan{path=/home/me kind=full}: done in 4210 ms: read 182203 dirs, … 21:44:09.540 steward: warning: hash{path=/home/media/TV}: saving content ids: database snapshot is stale…; retrying (1 of 9) 21:45:11.003 steward: error: hash /home/media/Movies: … ``` The part in braces is the **span**: the scan, hashing job or client connection the line belongs to. Warnings and errors are labelled, and coloured on a terminal. | level | the daemon shows it | what's there | |---|---|---| | error | always | a scan, hashing job, reload or invalidation that failed | | warning | always | retries, unreadable files, volumes found offline, verifications that found other bytes | | info | always | startup, each scheduled scan's summary, hashing results | | debug | with `-v` | roots and their policies, each phase of scans and hashing, progress every minute | | trace | with `-vv` | every file hashed, every request, every event | `RUST_LOG` overrides the levels with `tracing`'s filter syntax, for example `RUST_LOG=steward_index=trace` for the scanner alone. Under systemd, when stderr is the journal, each line carries its syslog priority, so `journalctl --user -u steward -p warning` shows just warnings and errors. Lines have no timestamp there, because the journal adds its own. The `steward` command logs only warnings and errors unless given `-v`; `steward-ui` logs the messages it shows. ## Status `steward status` returns the daemon's whole state as JSON. The parts most worth knowing: | field | what it holds | |---|---| | `activity.scan` | the scan in progress: path, kind, seconds | | `hashing` | the hashing job: path, files and bytes done and total (including bytes read so far of files in flight), start time | | `activity.reading` | every file being read for hashing right now: path, size, bytes read, seconds | | `activity.hash_queue` | folders waiting to be hashed | | `activity.connections`, `activity.subscribers` | clients connected, and those subscribed to events | | `daemon` | version, pid, uptime, index path and size on disk, hashing threads, socket paths | | `recent_scans` | the last 50 scans with timings, changes and errors | | `schedule` | when each root is scanned next | | `problems` | the last 200 warnings and errors, newest first | ```sh watch -n2 'steward status | jq .activity' steward status | jq '.problems[:5]' ``` `systemctl --user kill -s USR2 steward` writes the `activity` part to the log. That helps when you are reading the journal and don't have a client handy. The desktop app's [Daemon tab](gui.html#daemon) shows all of this. ## What it costs Measured on the author's machine (NVMe, 15.7 million paths in the home directory, a 44 TiB media library): | | | |---|---| | first scan of the home directory | about 1½ minutes | | rescan of an unchanged `/usr` (750,000 paths) | 69 ms trusting, 1 s full | | memory while scanning | about 0.5 GB resident | | index size | about 130 bytes per path (1.9 GB for 15.7 million) | | content ids | one read of each file, at disk speed; 32 bytes per MiB kept | Scans are I/O on metadata only. Hashing reads every byte of the files under `contentid` folders once, then only files that change. ## The index file The index lives at `~/.local/state/steward/index.db` (see [`db`](configuration.html#top-level-keys)), with a write-ahead log next to it. It holds nothing that can't be rebuilt from the filesystem, except content ids, which cost a full read to recompute. To start over: stop the daemon, delete `index.db` and the files beside it named `index.db-*`, start the daemon. Every root is scanned again, and content ids are recomputed. Removing a root from the configuration removes its entries from the index at the next reload. ## Troubleshooting `… is not indexed: no configured root covers it` : The path isn't under any `[[root]]`. Add one (`steward put-root PATH`) or check for a typo. `… is not indexed yet: root … has not finished its first scan` : Wait; `steward status | jq .activity.scan` shows the scan's progress. `… is not in the index: nothing by that name under root …` : The root is indexed but has no such path: it was created after the last scan, or excluded. `steward inspect PATH` (or `invalidate`) updates it. `… is not under a configured root` (type `not_under_root`) : `scan`, `inspect` and `verify` only work inside configured roots. A root says *(offline)* : The volume it lives on isn't mounted. Mount it; the next scan (`steward scan ROOT`) brings it back online with everything still identified. Hashing seems stuck : `steward status | jq '.activity.reading'` shows each file being read and how far along it is. A single large file on a slow disk can take minutes. Check `problems` for errors. `steward: error: connecting to the steward daemon at …` : The daemon isn't running (`steward service status`), or `XDG_RUNTIME_DIR` differs between the daemon and the client (common in `sudo` or `ssh` sessions). Out of inodes, or "No space left on device" with space free : Something has created a great many small files. Find where: `steward tree ~ -d 3 --by items`, or **Rank by Items** in the desktop app. On btrfs there is no inode limit; small files exhaust *metadata* space instead, and `steward settings | jq '.roots[].fs'` (or the Daemon tab) shows its use. To survey a disk your user can't read, such as a backup volume, index it once as root into a scratch file and explore that: `steward --db /root/backup.db scan /backup`, then `steward --db /root/backup.db tree /backup -d 3 --by items`. Something is changed on disk, but steward hasn't noticed : steward doesn't watch for changes. The next scheduled scan will notice; `steward scan PATH` or `steward inspect PATH` notices now. Files rewritten in place, keeping their directory's times, are only noticed by full scans, which are the default. --- title: Building on steward eyebrow: Build lede: How an application should use steward. That means asking about paths, learning what a file's bytes are, finding content wherever it now lives, telling steward what you changed, and following changes as they happen. It also means keeping steward optional. description: Patterns for applications using steward's content.socket, with Python examples. --- ## Principles **Treat steward as enrichment, not a dependency.** Your application should work, perhaps more slowly or with less to show, when steward isn't installed, isn't running, or doesn't cover the path in question. Check for the socket and fall back gracefully. Nothing in steward is something your application can't, in principle, do itself; steward just did it already, for everyone. **Use `content.socket`.** It has everything an application needs. Lookups, the content primitives and events are all there, and none of them can change what is indexed or how. Leave `api.socket` to the user's own tools. **Expect the picture to be recent, not instant.** steward has no inotify. What you read reflects the last scan of that path. When it matters, ask steward to look now (`inspect`), or recheck what it tells you (`resolve` with `recheck`). **You report, steward decides.** Applications tell steward that paths may have changed, or that a file may not hold what steward thinks. steward looks for itself and records what it finds. No call lets an application assert facts into the catalog, and none writes to a file. ## Connecting The socket is `$XDG_RUNTIME_DIR/steward/content.socket`. The protocol is JSON-RPC 2.0, one JSON object per line. Any language can speak it: see the [API reference](api.html). For Python, `python/steward_client.py` in the repository is a single file with no dependencies beyond the standard library: copy it into your project. It has a blocking `Client` and an asyncio `AsyncClient` with the same methods, typed results and errors that carry a stable `type`. ```python from steward_client import Client, StewardError, ConnectionLost def steward(): """A steward connection, or None when steward isn't available.""" try: return Client(timeout=30) # content.socket by default except ConnectionLost: return None ``` ## Paths are bytes File names on Linux are bytes and need not be UTF-8. steward never mangles them: such bytes travel as `\udcXX` escapes ([details](api.html#file-names-that-arent-utf-8)), which Python decodes natively. Treat every path you get as an opaque value to hand to the filesystem (or back to steward), and only make it printable at the moment you show it: ```python p = c.locate("Deluxe")[0] # str, possibly with lone surrogates open(p, "rb") # works raw = os.fsencode(p) # the exact bytes on disk print(p.encode("utf-8", "backslashreplace").decode()) # safe to print ``` ## Ask about paths ```python with Client() as c: e = c.stat("~/Pictures") print(e.total_alloc, e.total_files, e.tags) for child in c.children("~/Pictures"): # largest on disk first print(child.name, child.total_alloc, child.category, child.content_id) raws = c.locate("*.CR3", limit=5000) # glob on the file name ``` `stat` includes tags inherited from parent directories, so a file inside a git checkout's build output carries `classify:repo` and `classify:build-output`. `children` reports each child's own tags. ## Learn what a file's bytes are When you need a file's content id, ask steward to **inspect** it: ```python [f] = c.inspect(["~/Downloads/talk.mp4"]) if f.ok: print(f.id) # btv2:…, or None for an empty file else: print(f.error["type"], f.error["message"]) ``` `inspect` updates the index for that path first (it may be new, renamed or rewritten), then returns the stored id if the file's size and modification time are unchanged since it was hashed, or hashes it now. It takes many paths at once and answers each separately, in order. Errors are per path (`not_under_root`, `not_found`, `changing`, `unreadable`) and never fail the whole call. Hashing reads the whole file, at disk speed. Ask for the files you actually need, and batch them. ## Find content wherever it is now Your application remembers a content id, and later wants the bytes. Ask steward to **resolve** it: ```python [r] = c.resolve([(content_id, expected_size)], recheck=True) if r.state == "present": path = r.online[0].path # any online copy: they are identical elif r.state == "offline": where = r.observations[0].offline_at # e.g. /mnt/archive: ask the user to plug it in elif r.state in ("absent", "unknown"): ... # no copy steward knows of elif r.state == "mismatch": ... # your size and steward's disagree: not the content you meant ``` Passing the size you expect is a cheap guard. `recheck=True` makes steward `lstat` each copy before answering, and drop any that changed or vanished. Use it when you are about to open the file. Copies sharing an `inode` are hard links. Several ids can be resolved in one call, and the results come back in order. ## Tell steward what you changed When your application writes, moves or deletes files under steward's roots, say so. There are two ways, depending on whether you need an answer: ```python # You need the content ids of what you just wrote: results = c.inspect(written_paths) # You don't need anything back; steward rescans soon (about 2 s later): c.invalidate(directory) ``` Either way, other applications see the change sooner than the next scheduled scan, and receive the events. ## When bytes don't match Your application may find that a file steward listed as content X doesn't hold X: a checksum failed, or a decoder choked. Ask steward to **verify** it: ```python v = c.verify(content_id, path, reason="checksum mismatch at offset 4 MiB") if v.state == "unchanged": ... # the file does hold content_id; the problem is elsewhere elif v.state == "changed": ... # it holds v.current now; steward has stopped listing it for content_id elif v.state in ("gone", "unreadable", "not_file"): ... ``` steward rereads the whole file regardless of its stat (verification exists for changes that kept size and time), records what it holds, and from then on reports it accordingly. `reason` is only logged. While a verification runs, that path is left out of `resolve` results. Afterwards, `resolve` again to find another copy. ## Follow changes Rather than polling, subscribe to events for the content you care about. `AsyncClient.events` gives a stream on its own connection that reconnects by itself, resumes where it left off, and tells you when it couldn't: ```python import asyncio from steward_client import AsyncClient, Event, Gap async def follow(tracked: dict[str, str]): # content id -> path async with AsyncClient() as c: async with c.events(ids=tracked) as events: async for e in events: if isinstance(e, Gap): # Missed events (daemon restart, or we fell behind): re-resolve everything. for r in await c.resolve(list(tracked), recheck=True): tracked[r.id] = r.online[0].path if r.online else None elif e.name == "content.moved": tracked[e.data["id"]] = e.data["to"] elif e.name == "content.lost": tracked[e.data["id"]] = None # and resolve to find another copy elif e.name == "storage.offline": ... # paths under e.data["path"] are unreachable for now ``` The `ids` filter applies to `content.*` events. `storage.*` events always arrive, and carry only a directory: compare it with the paths you hold. `await events.watch(new_ids)` changes the filter without losing anything. Ids in results and events are always in the form `btv2:` plus 64 lowercase hex digits, so use that form as your keys. Events arrive when steward notices a change, at the next scan, `inspect`, `invalidate` or `verify`. They aren't instant. They are numbered, and a `Gap` is the only way to miss one, so your view never silently drifts. ## Piece layers For each content id of a file over 1 MiB, steward keeps the hash of every 1 MiB section. From it, steward derives the content's BEP 52 Merkle layer for any power-of-two piece size of at least 1 MiB, without reading the file: ```python layer = c.piece_layer(content_id, piece_size=4 << 20) # bytes: 32 per piece ``` That is exactly a v2 torrent's `piece layers` entry for the file. It is useful wherever partial verification of large files matters. The result is empty for files no bigger than one piece. ## Errors Every error has a stable `type` for programs and a message for people: | type | means | do | |---|---|---| | `not_indexed` | the path isn't in the index (yet) | `inspect` it, or treat as unknown | | `not_under_root` | the path is outside every configured root | steward can't help with this path | | `invalid_params` | a malformed id, a relative path, a bad piece size | fix the call | | `unknown_content` | no such content id | treat as unknown | | `no_layer` | the content predates verification layers | `inspect` a copy | | `changing` | the file kept changing while being read | retry later | | `forbidden` | an administration method on `content.socket` | use the user's tools for that | | `transport` | the connection failed (`ConnectionLost` in Python) | reconnect; every method is safe to repeat | ```python try: e = c.stat(path) except StewardError as err: if err.type == "not_indexed": [f] = c.inspect([path]) ``` ## Concurrency and connections Requests on one connection are handled concurrently and matched by `id`. A long `inspect` does not hold up a quick `stat` sent after it. `AsyncClient` does this for you. A subscribed connection carries events as well as responses: `AsyncClient.events` opens its own. If the daemon restarts, calls in flight fail with `ConnectionLost` and the next call reconnects. Every steward method is safe to repeat. ## Other languages Any language with Unix sockets and JSON will do: ```sh printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"stat","params":{"path":"/etc"}}' | socat - UNIX-CONNECT:"$XDG_RUNTIME_DIR/steward/content.socket" ``` In Rust, the `steward-proto` crate has the request types and a small blocking client (`Client::connect_to(&steward_proto::content_socket_path())`). ## Checklist - Works without steward; detects the socket and degrades gracefully. - Uses `content.socket`. - Passes absolute paths (the clients resolve `~` and relative paths for you). - Passes expected sizes to `resolve`, and `recheck=True` before opening. - Calls `inspect` or `invalidate` after writing under steward's roots. - Calls `verify` when bytes disagree with an id, then resolves again. - Treats a `Gap`, or a new `epoch`, as "re-resolve everything". --- title: For AI agents eyebrow: Build lede: A compact operating guide for agents that help someone use steward, or that write software on top of it. Facts first, then recipes, then the mistakes to avoid. description: Operating guide for AI agents using or integrating steward. Detection, safe calls, intent-to-command recipes and pitfalls. --- ::: agent Every page of this site is also plain Markdown: replace `.html` with `.md`. [llms.txt](llms.txt) indexes them, and [llms-full.txt](llms-full.txt) has them all in one file. The [API reference](api.md) is the authority on methods and shapes. ::: ## Facts - steward is a per-user Linux service (`steward daemon`) that indexes the paths under configured **roots**: `lstat` fields, directory subtree totals, classification tags, and content ids for some files. - It **only reads** the files it indexes. It never modifies, moves or deletes them. - Its answers come from the index and reflect the **last scan** of each path (there is no inotify), not necessarily this instant. - A **content id** is `btv2:` plus 64 hex digits: the BitTorrent v2 (BEP 52) Merkle root of the file's bytes. Equal ids mean identical bytes. Renames keep the id; writes change it. Empty files have none. - Clients connect to `$XDG_RUNTIME_DIR/steward/content.socket` (applications) or `api.socket` (administration) and speak JSON-RPC 2.0, one JSON object per line. - The `steward` command is a client for `api.socket`. Most subcommands print JSON; `tree`, `ls` and `locate` print text. ## Is steward here? ```sh test -S "$XDG_RUNTIME_DIR/steward/content.socket" && echo running steward status | jq '{roots: [.configured[].path], scanning, hashing: .hashing.path}' ``` If the socket is missing, steward isn't running. `steward service status` says whether it is installed as a service; `steward version` shows the command's and the daemon's versions. Don't start or install it without the user's agreement. ## Ground rules 1. **Read freely; change only when asked.** Lookups are harmless. Adding or removing roots, editing `settings.toml`, forcing hashing of large folders, and exports are the user's decisions. Ask first. 2. **steward reports; it doesn't decide.** A duplicate list says which files hold identical bytes, not which copy to delete. Never delete or move the user's files because of something steward said, unless the user tells you exactly what to do. 3. **Check freshness when it matters.** Before acting on a path steward returned, confirm it (`stat` it yourself, or use `resolve --recheck`). After changing files yourself, tell steward (`steward inspect PATHS` or `steward invalidate DIR`). 4. **Offline is not gone.** A root or observation marked offline is on an unmounted volume. Its files still exist. 5. **Hashing costs a full read.** `inspect` and `cid` on a few files are cheap. `hash` on a media folder can take hours. Don't start one casually. ## Recipes | the user wants | do | |---|---| | to know what uses the space | `steward tree PATH -d 2` (text), or `steward raw children '{"path":"/abs/path"}'` (JSON) | | the size of a folder | `steward stat PATH \| jq '{total_alloc, total_size, total_files}'` | | to find files by name | `steward locate 'pattern'`: substring, or glob with `* ? [`; `-x` exact, `-g` glob, `-r` regex, `-t d` directories only; names only, not contents. Results are confirmed on disk and stale folders rescanned (`--check rescan`, the default) | | the content id of a file | `steward inspect PATH \| jq -r '.[0].id'` | | other copies of a file | `id=$(steward inspect PATH \| jq -r '.[0].id'); steward find "$id"` | | where some content is now | `steward resolve ID --recheck`: check `state` and `observations[].online` | | duplicates | `steward dups PATH` (only hashed files count; `content-summary` shows coverage) | | what's in a git checkout that is ignored or built | `steward stat PATH \| jq .tags`; tags like `classify:build-output`, `classify:ignored` | | why steward seems behind or busy | `steward status \| jq '{activity, hashing, problems: .problems[:5]}'` | | steward to notice a change now | `steward inspect PATHS` (files: rescan + id), or `steward scan DIR` | | to index another folder | with consent: `steward put-root PATH` (add `--contentid FOLDER` for ids) | Paths can be relative or use `~`. The CLI makes them absolute. The raw API needs absolute paths. ## Reading results - Sizes are bytes. `alloc` and `total_alloc` are space on disk, and `size` and `total_size` are apparent size. Report the one the user means. For "how much space", use `total_alloc`. - `total_dirs` counts the directory itself. - `tags` on `stat` include inherited ones. On `children`, only each child's own. - `resolve` states: `present`, `offline` (all copies on unmounted volumes), `absent` (known, no copy now), `unknown` (never seen), `mismatch` (your size disagrees). - Errors carry `data.type`. The command line prints `steward: error: ` and exits 1. The common ones are `not_indexed` (not scanned yet or no such path), `not_under_root` (outside every root), `invalid_params`. ## Writing software that uses steward Read [Building on steward](applications.md), then the [API reference](api.md). In short: - Make steward optional. The program must work without it. - Connect to `content.socket`. In Python, vendor `python/steward_client.py` (stdlib only): `Client()` / `AsyncClient()`. - To identify files, `inspect(paths)` (batch). To locate content, `resolve([(id, size)], recheck=True)`. After writing files, `inspect` or `invalidate`. When bytes disagree, `verify(id, path)` and resolve again. - To follow changes, `AsyncClient.events(ids=…)`. Handle `Gap` by re-resolving everything. - Every method is safe to repeat after `ConnectionLost`. ```python from steward_client import Client, ConnectionLost try: with Client(timeout=30) as c: [f] = c.inspect(["/home/me/Downloads/file.iso"]) copies = c.resolve([(f.id, f.size)], recheck=True)[0].online if f.id else [] except ConnectionLost: copies = [] # steward isn't available: carry on without it ``` --- title: API reference eyebrow: Reference lede: The protocol steward's daemon speaks, every method it answers, the shapes of its results, its events and its errors. description: steward's JSON-RPC 2.0 API — transport, sockets, every method, result types, events and errors. --- ## Transport The daemon listens on two Unix stream sockets in `$XDG_RUNTIME_DIR/steward/` (a `0700` directory; without `XDG_RUNTIME_DIR`, a private directory under `/tmp`): | socket | for | offers | |---|---|---| | `content.socket` | applications | [catalog](#catalog), [content](#content), [events](#events), `invalidate`, `status` | | `api.socket` | the user's own tools | everything on `content.socket`, plus [maintenance](#maintenance) and [settings](#settings) | The protocol is **JSON-RPC 2.0**, one JSON object per line (UTF-8, `\n` terminated), in both directions. ``` → {"jsonrpc":"2.0","id":1,"method":"stat","params":{"path":"/etc"}} ← {"jsonrpc":"2.0","id":1,"result":{"path":"/etc","kind":"dir", …}} ``` - Every request with an `id` gets exactly one response with that `id`. A request without one is a notification and gets none. - `params` is an object (named parameters). Methods without parameters accept it absent, `null` or `{}`. - Requests on one connection are handled **concurrently**; responses may come back in a different order. Match them by `id`. - A connection that has [subscribed](#subscribe) also receives `event` and `gap` notifications between responses. - Any request may ask for **progress**: with `"progress": true` in its params, slow work sends `progress` notifications before the response, `{"id": , "stage": "…", "message": "…", …}`. `message` is for people; `stage` and the other fields are for programs. Clients that don't ask get none. `locate` reports `rescanning`, `waiting` (for a scan already running) and `rescanned`. ## Conventions Paths : Absolute, without `~`, in requests and answers alike. Every path is a JSON string, including names that aren't valid UTF-8; see [File names that aren't UTF-8](#file-names-that-arent-utf-8). Content ids : `btv2:` followed by 64 lowercase hex digits, the file's BEP 52 pieces root. Requests also accept bare hex and either case. Sizes : Bytes. `size` is the apparent size; `alloc` is space on disk (blocks × 512). Times : `Entry.mtime` is in seconds since the Unix epoch. Fields ending in `_ns` are nanoseconds. Event `time`, `started`, `finished` and the like are seconds as floating point. ## File names that aren't UTF-8 Linux file names are bytes, and some aren't valid UTF-8: a Latin-1 `®` (`0xAE`) from an old Windows share, a file made on purpose by a test. JSON strings are Unicode. steward sends every path as a JSON string anyway, and keeps the exact bytes, using the same convention as Python's file APIs ("surrogateescape", PEP 383): - valid UTF-8 travels as itself (all but a handful of names, on any real system); - each byte that isn't part of valid UTF-8 travels as the lone surrogate U+DC80 + byte, written in JSON as an escape: `0xAE` is `\udcae`. ``` on disk: /home/me/Shorts/Arch Deluxe\xae.doc on wire: "/home/me/Shorts/Arch Deluxe\udcae.doc" ``` The same rule applies to paths you send: `{"path": "/home/me/Shorts/Arch Deluxe\udcae.doc"}` names that file. Every path field in every request, answer and event follows it. In Python it needs no work: `json.loads` turns `\udcae` into a lone surrogate in a `str`, `open(path)` and `os.stat(path)` use it directly, and `os.fsencode(path)` returns the exact bytes. The bundled client accepts paths as `str` or `bytes`. In JavaScript, `JSON.parse` keeps lone surrogates in strings too. Elsewhere, decode `\udc80`–`\udcff` escapes to the bytes `0x80`–`0xff` yourself; a JSON parser that rejects lone surrogates, or replaces them (as `jq` does, printing `�`), loses the original bytes. Never assume a path is printable. ## Errors ```json {"jsonrpc":"2.0","id":7,"error":{"code":-32000,"message":"/srv is not indexed: no configured root covers it","data":{"type":"not_indexed"}}} ``` `data.type` is stable and meant for programs. `message` is for people and may change. | code | `data.type` | meaning | |---|---|---| | -32700 | `parse_error` | the line wasn't JSON (the response has `id: null`) | | -32600 | `invalid_request` | not a JSON-RPC 2.0 request | | -32601 | `method_not_found` | no such method | | -32602 | `invalid_params` | missing or ill-typed parameters, a malformed content id, a relative path, an unusable piece size | | -32000 | `forbidden` | an administration method sent to `content.socket` | | -32000 | `not_indexed` | the path is not in the index, or no root covers it | | -32000 | `not_under_root` | the path is outside every configured root (for methods that scan or read it) | | -32000 | `unknown_content` | no such content id | | -32000 | `no_layer` | the content has no stored verification layer | | -32000 | `changing` | the file kept changing while being read | | -32000 | `failed` | anything else; see `message` | ## Types ### Entry One indexed path. | field | type | | |---|---|---| | `path` | string | absolute | | `kind` | string | `file`, `dir`, `symlink` or `other` | | `mode` | integer | permission bits (`0o7777` mask) | | `uid`, `gid` | integer | owner | | `size` | integer | apparent size | | `alloc` | integer | space on disk | | `mtime` | integer | modification time, seconds | | `total_size`, `total_alloc` | integer | subtree totals for directories; the entry's own figures otherwise | | `total_files`, `total_dirs` | integer | files and directories beneath (a directory counts itself) | | `total_items` | integer | every entry beneath, the directory included: files, directories, symlinks, the rest; 1 for non-directories | | `tags` | string[] | e.g. `classify:repo`; omitted when empty | | `category` | string | files only: `image`, `video`, `audio`, `document`, `source`, `archive`, `object`, `disk-image`, `torrent`; omitted when none | | `content_id` | string | when a current content id is stored; omitted otherwise | ### Resolution What steward knows of one content id (result of [`resolve`](#resolve)). | field | type | | |---|---|---| | `id` | string | normalised content id | | `size` | integer \| null | the content's size, when known | | `state` | string | `present`, `offline`, `absent`, `unknown` or `mismatch` | | `layer` | boolean | a verification layer is stored (always true up to 1 MiB) | | `observations` | Observation[] | current copies, reachable ones first | **Observation**: `path` (string), `inode` (string, `":"`; equal for hard links), `online` (boolean), `offline_at` (string \| null: the unmounted directory holding it), `mtime_ns` (integer). ### Inspected One path after [`inspect`](#inspect): `path`, `kind` (`file`, `dir`, `symlink`, `other`, or null on error), `id` (string \| null: null for directories, empty files and non-files), `size`, `error` (null, or `{type, message}` with type `not_under_root`, `not_found`, `changing`, `unreadable` or `invalid_params`). ### Verdict What [`verify`](#verify) found: `id` (the id you claimed), `path`, `state` (`unchanged`, `changed`, `gone`, `unreadable`, `not_file`), `current` (the id the file holds now, or null). ### ScanReport `root`, `dirs_read`, `dirs_trusted`, `entries_seen`, `inserted`, `updated`, `deleted`, `errors` (unreadable entries), `millis`, `offline` (directories found on unmounted volumes), `load_ms`, `write_ms`, `totals_ms`. ## Catalog Read-only lookups in the index. Both sockets. ### status The daemon's state. No parameters. | field | | |---|---| | `db` | index file | | `configured` | the configured roots and their policies | | `indexed` | an Entry for each indexed root | | `scanning` | whether a scan is running | | `hashing` | the hashing job, or null: `path`, `files_total`, `files_done`, `bytes_total`, `bytes_done` (includes bytes read of files in flight), `started` | | `activity` | `scan` (`path`, `kind`, `secs`, or null), `reading` (files being hashed: `path`, `size`, `read`, `secs`), `hash_queue`, `connections`, `subscribers`, `event_seq` | | `daemon` | `version`, `pid`, `started`, `uptime_secs`, `epoch`, `hash_threads`, `db`, `db_bytes`, `api_socket`, `content_socket`, `managed` | | `recent_scans` | the last 50 scans, newest first: ScanReport plus `kind` and `finished`, or `root`, `error`, `kind`, `finished` | | `schedule` | `path` and `next` (seconds since epoch) for each root | | `problems` | the last 200 warnings and errors, newest first: `time`, `level`, `context`, `message` | ### stat `{path}` → Entry, with tags inherited from its ancestors. ### children `{path}` → Entry[], the directory's children, largest `total_alloc` first. Each child carries only its own tags. ### locate `{pattern, limit = 1000, mode = "auto", ignore_case = false, kind?, check = "none"}` → string[] of paths whose final name component matches. | param | | |---|---| | `mode` | `auto` (a case-sensitive glob if `pattern` contains `*`, `?` or `[`, otherwise a substring), `substring`, `exact` (the whole name), `glob` (the whole name), `regex` (anywhere in the name; Rust `regex` syntax) | | `ignore_case` | for `exact`, `glob` and `regex`; substrings always ignore ASCII case | | `kind` | `file`, `dir`, `symlink` or `other` | | `check` | `none`: answer from the index. `exists`: `lstat` each result. `rescan`: also rescan the folder of each result that's gone (or its nearest existing parent; at most 64 at once, the rest queued) and search again | Results are sorted by their bytes. A rescan has to wait for any scan already running, such as a root's scheduled full rescan; with `"progress": true` the request reports that as it happens: ``` ← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"rescanning","message":"3 of 41 results are gone from disk; rescanning 2 folder(s)","stale":3,"folders":2,"queued":0}} ← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"waiting","message":"waiting for the full scan of /home/me in progress (41 s so far) before rescanning /home/me/src","for":"the full scan of /home/me","secs":41}} ← {"jsonrpc":"2.0","method":"progress","params":{"id":7,"stage":"rescanned","message":"rescanned /home/me/src in 38 ms","folder":"/home/me/src","ms":38.2}} ← {"jsonrpc":"2.0","id":7,"result":{"paths":[…], …}} ``` A client can also do the checking itself, with nothing but `content.socket`: `locate` without a check; `lstat` each result; pass the ones that are gone to [`inspect`](#inspect), which rescans each one's folder (its own listing, trusting the folders below) and waits; then `locate` again. That is exactly what `check: "rescan"` does in one request. With `check` other than `none` the result is an object instead of a list: `{paths, stale, rescanned, queued, limited, search_ms, check_ms, rescan_ms}`: the results that exist, the indexed results that are gone, the folders rescanned now and those queued (past 64), whether the search stopped at `limit`, and the milliseconds spent searching, checking and rescanning. A malformed glob or regular expression is `invalid_params`. `exact`, `substring` and case-sensitive `glob` are answered by the index; `regex` and case-insensitive `glob` read every name, which takes a few seconds on a very large index. ### find_content `{id}` → string[]: the indexed paths currently holding this content. [`resolve`](#resolve) says more (reachability, state, sizes). ### duplicates `{path, limit = 1000}` → groups of identical hashed files with at least two paths under `path`, most wasted space first: ```json [{"id":"btv2:…","size":1048576000,"wasted":2097152000,"paths":["/a/x.mkv","/b/x.mkv","/c/x (1).mkv"]}] ``` ### content_summary `{path}` → content-id coverage under `path`: `files`, `bytes`, `hashed_files`, `hashed_bytes`, `unhashed_files`, `unhashed_bytes`, `distinct_ids`, `duplicate_groups`, `duplicate_files`, `wasted_bytes`, `truncated` (true if it stopped at 2,000,000 files). ## Content Both sockets. ### content_id `{path}` → content id or null (empty file). Hashes the file if it has no current id. The path must be a regular file. Prefer [`inspect`](#inspect), which also updates the index for the path and takes many at once. ### resolve `{contents: [{id, size?}], recheck = false}` → Resolution[], in request order. - `size`: the size you expect. A different known size gives `state: "mismatch"` and no observations. - `recheck`: re-stat each copy before answering. Copies that changed or vanished are dropped, and their directories are queued for a rescan. A copy that can't be reached because its volume isn't mounted is reported offline, with `offline_at`. - Paths being [verified](#verify) are left out. ``` → {"jsonrpc":"2.0","id":2,"method":"resolve","params":{"contents":[{"id":"btv2:1d8e…","size":3145728}],"recheck":true}} ← {"jsonrpc":"2.0","id":2,"result":[{"id":"btv2:1d8e…","size":3145728,"state":"present","layer":true, "observations":[{"path":"/home/me/a.bin","inode":"9f3c…:1442","online":true,"offline_at":null,"mtime_ns":1790800000000000000}]}]} ``` ### inspect `{paths: [string]}` → Inspected[], in request order. These paths may have changed: bring the index up to date for them now, and give each regular file's content id. - A file's parent directory is rescanned (its own listing read even if its times are unchanged); then the stored id is used if the file's size and modification time are unchanged, or the file is hashed now. - A directory is rescanned in full. Files under it are hashed only if a root's `contentid` policy covers them. - Paths must be absolute and under a configured root. Problems are reported per path; the call itself succeeds. ### verify `{id, path, reason = ""}` → Verdict. There is reason to believe `path` no longer holds `id`. steward withholds the path from `resolve`, rescans its directory, rereads the file in full and records what it holds: - `unchanged`: it holds `id`. - `changed`: it holds `current` (null for an empty file). Paths that were listed for `id` get `content.lost` with reason `changed`. - `gone`, `not_file`: nothing hashable is there. - `unreadable`: reading failed. The inode's stored id is dropped, and its paths get `content.lost` with reason `unreadable`. `reason` is logged, never interpreted. Errors: `invalid_params`, `not_under_root`, `changing` (still changing after three reads). ### piece_layer `{id, piece_size}` → `{id, size, piece_size, layer}`. The content's BEP 52 Merkle layer at `piece_size` (a power of two, at least 1 MiB), derived from the stored 1 MiB layer without reading the file. `layer` is the concatenated 32-byte hashes in hex, one per piece, and empty for content no bigger than one piece. Errors: `unknown_content`, `no_layer`, `invalid_params`. ### invalidate `{path}` → `"queued"`. Something under `path` changed. The daemon gathers invalidations for two seconds, then rescans the nearest existing directory of each (in full). ## Events ### subscribe `{since?, ids?}` → `{epoch, seq, complete}`. From now on this connection also receives events. - `since`: first replay the backlog's events after this `seq`. - `ids`: deliver `content.*` events only for these content ids. `storage.*` events always arrive. - `epoch` changes each time the daemon starts; `seq` is the last event's number in it. - `complete` is false when events after `since` were missed: the backlog (the last 4,096 events) no longer reaches back that far, or `since` is ahead of `seq`. The daemon can't tell a `since` from an earlier run, so compare `epoch` with the one you saw before: a different epoch means you missed everything in between. Subscribing again replaces the filter. The response comes before any replayed event. ### unsubscribe `{}` → `true`. Stop delivering events on this connection. ### Notifications ```json {"jsonrpc":"2.0","method":"event","params":{"seq":42,"time":1790812345.12,"name":"content.moved","data":{"id":"btv2:…","from":"/a/x.mkv","to":"/b/x.mkv"}}} {"jsonrpc":"2.0","method":"gap","params":{"after":41}} ``` | `name` | `data` | meaning | |---|---|---| | `content.observed` | `{id, path}` | the content was seen at this path (new, renamed onto, or hashed) | | `content.moved` | `{id, from, to}` | the same inode moved from one path to another | | `content.lost` | `{id, path, reason}` | the path no longer holds this content: `deleted`, `changed` or `unreadable` | | `storage.offline` | `{path}` | this known directory's volume is not mounted | | `storage.online` | `{path}` | it is mounted again | | `storage.unindexed` | `{path}` | this root was removed from the configuration | A `gap` notification means this subscriber fell behind and events after `after` were dropped for it. After a gap, an incomplete subscribe, or a new epoch, re-resolve whatever you track. ## Maintenance `api.socket` only. ### scan `{path, trust_dir_mtime = false}` → ScanReport. Rescan `path` (a root or anything under one) now and wait. `trust_dir_mtime` skips directories whose times show nothing changed. Error: `not_under_root`. ### classify `{path}` → `{scanned, tagged}`. Re-run classification, from the enclosing repository if there is one. ### hash_tree `{path}` → `{stale, hashed}`. Hash every file under `path` without a current content id, and wait. For a media folder, this can take hours. ### export_qdirstat `{path, out}` → `{entries, out}`. Write `path`'s subtree as a gzipped qdirstat 2.0 cache file at `out`. Refuses to overwrite an existing file. ## Settings `api.socket` only. Changes are validated, written into `settings.toml` (comments and formatting kept), and applied at once. ### settings No parameters → `{file, db, roots, scanning, hashing, hash_threads, hash_threads_default}`. Each of `roots` is `{settings, indexed, offline, fs}`: the root's policy, an Entry for it (or null before its first scan), whether its volume is offline, and the capacity of the filesystem it is on (null when offline): | `fs` field | | |---|---| | `type`, `mount` | filesystem type and mount point | | `bytes_total`, `bytes_free` | space, as `df` reports it | | `inodes_total`, `inodes_free` | inodes, as `df -i` reports them; null where inodes are allocated on demand (btrfs) | | `metadata_total`, `metadata_used` | btrfs only: metadata space allocated, and used. Many small files exhaust this while `df` still shows free space | ### put_root `{root}` → reload result. Add a root, or replace the one with the same path. `root` has the [configuration keys](configuration.html#root-keys): `path` is required, the rest default. Validation: the path must be an existing absolute directory, `interval_minutes` at least 1, exclude patterns valid, content-id folders inside the root. ### remove_root `{path}` → reload result. Stop indexing a root and drop its entries. ### reload No parameters → `{added, changed, removed}`. Re-read `settings.toml`, start new roots, rescan changed ones, drop removed ones. # Design notes steward is a per-user service that owns one index of the filesystem and serves it to applications over `$XDG_RUNTIME_DIR/steward/content.socket` (administration goes through `api.socket` next to it). It is deliberately the opposite of baloo/tracker: the core only knows *structure* (paths, stat fields, well-known classes, content identity). Anything content-specific — EXIF, audio tags, full text — belongs in the application that cares, keyed by the content ids steward provides. ``` apps (file manager, photo app, backup tool, qdirstat-style viewer) │ JSON-RPC 2.0 over content.socket / api.socket ┌──────┴───────────────────────────────────────────────────────┐ │ stewardd │ │ scheduler ─ invalidation queue ─ request handlers │ │ ┌─────────────┐ ┌──────────────────┐ ┌──────────────────┐ │ │ │ L1 index │→ │ L2 classify │ │ L3 contentid │ │ │ │ scan + diff │ │ tags, gitignore │ │ BEP 52 merkle │ │ │ └──────┬──────┘ └────────┬─────────┘ └────────┬─────────┘ │ │ └────────── turso (SQLite format) ───────┘ │ └──────────────────────────────────────────────────────────────┘ ``` ## Crates | crate | role | |---|---| | `steward-index` | L1: parallel walker, diffing writer, queries, storage for tags and content ids | | `steward-classify` | L2: directory classes and gitignore state → tags; file categories by extension | | `steward-contentid` | L3: BitTorrent v2 pieces root, no I/O policy, no DB | | `steward-proto` | wire types, socket paths, a small blocking client | | `steward-log` | logging shared by the daemon, CLI and GUI (`tracing`) | | `stewardd` | the service: config, scheduling, invalidation, socket server | | `steward-cli` | `steward` command: `tree`, `ls`, `locate`, `dups`, `cid`, … | | `steward-ui` | GPUI qdirstat-style tree with locate; a plain socket client like any other app | ## Layer 1: the index One `entries` row per path: `parent`, `name` (raw bytes, so non-UTF-8 names survive), `kind`, `mode`, `uid`, `gid`, `size`, `alloc` (`st_blocks*512`), `nlink`, `dev`, `ino`, `mtime_ns`, `ctime_ns`, plus subtree totals `t_size`, `t_alloc`, `t_files`, `t_dirs` on directories. Roots have `parent = 0` and their absolute path as `name`. `UNIQUE(parent, name)` serves both tree navigation and path resolution; it is the only secondary index. Kept lean on purpose (measured on a 15.7M-entry home, 1.9 GB): files store ctime and totals as 0, which take no bytes, because only directories' ctime (trusting rescans) and totals are used; a file's share of its directory's totals is computed from its own size, alloc and nlink. Two further options were measured and deferred until size matters: keying the name index on a hash (about 270 MB, but the database would no longer enforce one entry per name), and compressing timestamps. **Scan pipeline.** A rayon walker (the dust/disktree shape) stats every entry and sends one `Listing` per directory to a single async writer. A listing is always sent before its subdirectories are walked, so the writer already has the parent's row id. The writer loads the stored children of that directory (one indexed query), diffs by name, and writes only inserts, updates and subtree deletes, committing every 20k changes so readers see progress. Totals are then recomputed only for directories that changed and their ancestors, deepest first. **Hardlinks** contribute `1/nlink` of their size to every directory that holds a link, which keeps totals local (no global de-dup pass) and exact whenever all links are inside the subtree being viewed. **Scanning a subpath** updates that subtree and re-aggregates its ancestors. Scanning a parent of an existing root adopts the old root instead of duplicating it. ### Why turso / SQLite Measured on this machine (turso 0.8.1, release build): | operation | time | |---|---| | insert 1M rows, one transaction | 1.4 s | | build `parent` index over 1M rows | 0.7 s | | `name LIKE '%x%'` over 1M rows | 110 ms | | 1000 child listings (20k rows) | 11 ms | | `/usr` first scan, 753k entries | 2.7 s | | `/usr` full rescan, nothing changed | 0.95 s | | `/usr` trusting rescan | 0.07 s | | `steward tree /usr -d 2`, `steward locate` | < 0.1 s | Readers are unaffected by a running scan. That is fast enough that an in-memory tree in the daemon isn't needed; the database is the single source of truth, which is what lets other processes and layers share it. All SQL lives in `steward-index`, so swapping stores later touches one crate. ## Invalidation without inotify inotify needs a watch per directory, overflows, and costs kernel memory proportional to the tree. steward instead layers cheap, bounded checks: Periodic rescans default to once a day, each one full; the checks below are what shorter intervals (`interval_minutes`, `full_every`) choose between. 1. **Trusting rescan.** A directory's mtime/ctime changes whenever an entry is added, removed or renamed in it, so a directory whose stamps match the index is not `readdir`'d; the walk descends through its known subdirectories. This is the mlocate trick and is why a rescan of `/usr` costs 70 ms. It misses in-place size changes of files in unchanged directories. 2. **Full rescan** (every `full_every`th round). Stats every file; still only writes what differs. 3. **Client invalidation.** Apps that change files call `invalidate` (or `inspect`, below, to have them handled now); the daemon debounces for 2 s, drops paths covered by another, and rescans the nearest surviving ancestor. The applications doing the writing know best what they touched. 4. **Correctness rule:** a directory's stored mtime is only ever updated by its *own* listing, never by its parent's diff. Otherwise an interrupted scan could store a new mtime without the new children, and every later trusting scan would believe the stale listing. Future, all still inotify-free: - **Lazy validation on read.** Before answering `children`, `lstat` the directory and rescan it inline if its stamps moved (one syscall per query), so interactive views are never stale. - **btrfs generations** (`BTRFS_IOC_TREE_SEARCH` / `find-new`) to list changed inodes since the last scan's transid without walking. - **fanotify with `FAN_MARK_FILESYSTEM`** as an opt-in for system installs; one mark per filesystem, but it needs `CAP_SYS_ADMIN`. ## Layer 2: classification Runs in-process after every scan and writes `tags(entry, source, tag)` with `source = "classify"`. A tag goes on the *topmost* entry it applies to and is inherited; `stat` returns effective (inherited) tags, `children` returns each child's own. - `repo` on a directory containing `.git`; `vcs-metadata` on the `.git`. - `ignored` for entries matched by the repo's `.gitignore` files and `.git/info/exclude` (the `ignore` crate's matcher, applied over the index, not a second filesystem walk; only the ignore files themselves are read). - Well-known directories, trusted only with evidence where the name is too generic: `target` only beside `Cargo.toml`, `node_modules` beside `package.json`, `.venv` only with `pyvenv.cfg`; plus caches (`__pycache__`, `.mypy_cache`, `.cache`, …), build output (`.next`, `CMakeFiles`, `zig-out`, …), and trash. - File categories (image, video, audio, archive, document, source, …) come from the extension and are computed on read, never stored. Classification of a changed subpath restarts from its enclosing repository so ignore rules from above still apply, and is skipped below an already classified directory. Other classifiers are meant to be separate clients that own their own `source` namespace (not yet exposed over the socket). ## Layer 3: content ids The id is the BitTorrent v2 (BEP 52) **pieces root**: SHA-256 of each 16 KiB block, a binary merkle tree padded to a power of two with zero leaves. It is independent of piece length, so it is the same value any BEP 52 software computes for the same bytes, whatever piece size it uses. Verified identical to libtorrent 2.0.11's `pieces root` for files from 730 B to 214 MB. Empty files have no id, as in BEP 52. Stored per `(dev, ino)` with the `size` and `mtime_ns` it was computed under; an id is only returned while both still match. Renames and moves within a filesystem keep the id: they change the inode's ctime, which is deliberately not compared (a rename must not re-read a film). Any write changes mtime and invalidates the id; restoring an old mtime after writing (`touch -d`) would go unnoticed, the same trade rsync's quick check makes. Paths are found through `content_entries` (hashed inodes only, one row per hard link). A scan that meets a new or changed file whose inode was hashed under its current size and mtime links it at once, so a renamed or moved file keeps its id without being read, and the scan reports the change (see events below). It looks inodes up one at a time, switching to loading every hashed inode once a scan meets more than 2,048 new files. A file that changes while being hashed is discarded and retried next pass. Only subtrees listed under `contentid` in the config are hashed automatically; `steward cid` and `steward hash` work anywhere. Queries: `find_content` (id → current paths), `resolve` (below), `duplicates` (by wasted bytes). ### Verification layer With each content id the hashing pass keeps a **1 MiB verification layer**: the root of every 64-block (1 MiB) subtree of the file's BEP-52 tree, 32 bytes per MiB (about 0.9 GB for 28 TiB), keyed by content id. Any piece layer for a power-of-two piece size of 1 MiB or more derives from it by hashing pairs upward, so no layer ever needs the file read again; `piece_layer` returns exactly the bytes of a v2 torrent's `piece layers` entry for the file (verified against libtorrent at 1, 4 and 16 MiB). Files of 1 MiB or less have none, as in BEP 52. ## Filesystems and offline volumes Stored filesystem identity is `f_fsid` from statvfs, which Linux derives from the filesystem UUID (plus the subvolume on btrfs), never `st_dev`: btrfs assigns `st_dev` at mount time, so it can change with mount order across reboots. A known directory found on a different filesystem than it was indexed on belongs to a volume that is not mounted: it is reported **offline** and nothing under it is read or changed. An unmounted media disk must never read as "630,000 files deleted". Steward only ever reads indexed files. The only files it writes are its own index and socket, `settings.toml` when a root change is requested, and qdirstat exports, which refuse to overwrite an existing file. ## Content primitives What applications build on, all on `content.socket`. The test for adding one: it must make sense for any application, not one consumer. Steward knows content and where it was observed; what a consumer does with that (sharing it, editing it, deduplicating it) stays in the consumer. - **`resolve {contents: [{id, size?}], recheck?}`**: where each content is. Per id, in request order: `state` is `present` (a copy is reachable), `offline` (copies are known, all on unmounted volumes), `absent` (seen before, no copy now), `unknown` (never seen) or `mismatch` (the given size differs from the content's); `observations` lists `{path, inode, online, offline_at, mtime_ns}`, reachable ones first (hard links share `inode`); `layer` says whether the verification layer is stored. `recheck` re-stats every copy first: changed or vanished copies are dropped and their directories queued for rescanning, a copy that fails to stat is offline if the nearest existing ancestor is an indexed directory now on another filesystem. - **`inspect {paths}`**: these paths may have changed; bring the catalog up to date for them now. For a file, the parent directory is rescanned (its own listing read even when its times say nothing changed, since in-place edits leave them alone; subdirectories trusted as usual), then its content id established: the stored id if size and mtime still match, else a hash on the hashing pool. Directories are rescanned in full; files below them are hashed only by the configured policy. Answers per path `{path, kind, id, size, error}`, errors typed (`not_under_root`, `not_found`, `changing`, `unreadable`). - **`verify {id, path, reason}`**: an application has reason to think `path` no longer holds `id` (it found bytes that disagree). The path is withheld from `resolve` while steward rescans its directory and rereads the file in full, whatever the stat says; it records what the file holds and answers `unchanged`, `changed` (with `current`), `gone`, `unreadable` (the inode's id is dropped) or `not_file`. `reason` is logged, never interpreted. This is how silent changes behind an unchanged size and mtime are caught. - **`piece_layer {id, piece_size}`**: a stored part of the content's own hash tree (above). None of them writes to a file. The worst a client can do is make steward read and hash files under its roots. ### Events `subscribe {since?, ids?}` turns on events for a connection; `unsubscribe` stops them. Each arrives as an `event` notification `{seq, time, name, data}`: | name | data | from | |---|---|---| | `content.observed` | `{id, path}` | a scan linking a hashed inode at a new path, or a hash | | `content.moved` | `{id, from, to}` | one scan losing and gaining the same hashed inode | | `content.lost` | `{id, path, reason}` | `deleted`, `changed` (new stat, or a reread disagreed), `unreadable` | | `storage.offline` | `{path}` | a scan finding a known directory's volume unmounted | | `storage.online` | `{path}` | a scan finding it back | | `storage.unindexed` | `{path}` | a root removed from the configuration | `ids` filters `content.*` events; `storage.*` always arrive, and carry the directory only. Events are as timely as what produces them: a periodic scan, an `invalidate`, an `inspect` or a `verify` (there is no inotify). The daemon numbers events and keeps the last 4,096 in memory. `since` replays those after it; the subscribe result `{epoch, seq, complete}` says whether the replay reached back that far. `epoch` changes on every daemon start, when earlier `seq`s mean nothing. A subscriber that falls more than 1,024 events behind gets a `gap` notification `{after}`. After any gap a client resolves what it tracks again. ## Protocol JSON-RPC 2.0, one message per line, on two Unix sockets in `$XDG_RUNTIME_DIR/steward/` (directory `0700`). The split is by capability, not by client: - `content.socket`, for applications: `status`, `stat`, `children`, `locate`, `content_id`, `find_content`, `duplicates`, `content_summary`, `piece_layer`, `resolve`, `inspect`, `verify`, `invalidate`, `subscribe`, `unsubscribe`. - `api.socket`, administration: all of those, plus `scan`, `classify`, `hash_tree`, `export_qdirstat` and settings management (`settings`, `put_root`, `remove_root`, `reload`). Settings changes are validated by the daemon, written into settings.toml with comments kept (`toml_edit`), and applied at once. `{"jsonrpc":"2.0","id":1,"method":"stat","params":{"path":"/home/me"}}` → `{"jsonrpc":"2.0","id":1,"result":{…}}`. Requests on one connection run concurrently; match responses by `id`. Errors are `{code, message, data: {type}}`, where `type` is stable for programs: `parse_error`, `invalid_request`, `method_not_found`, `invalid_params`, `forbidden` (an administration method on `content.socket`), `not_indexed`, `not_under_root`, `unknown_content`, `no_layer`, `changing`, `failed`. Debug with `steward raw METHOD '{…}'` or `socat - UNIX-CONNECT:$XDG_RUNTIME_DIR/steward/content.socket`. `status` reports the daemon's internal state: - `activity`: the scan in progress (path, mode, seconds), every file being read for hashing (path, size, bytes read so far, seconds), the hash queue, open connections and subscribers, and the last event's `seq`; - `hashing`: the job, whose `bytes_done` includes what has been read of files still in flight, so progress moves during a 50 GB film; - `daemon`: version, pid, start time and uptime, index size on disk, hashing threads, socket paths; - `recent_scans`: the last 50 scans, with timings, changes and errors; - `schedule`: when each root is scanned next; - `problems`: the last 200 warnings and errors logged, with their spans. `kill -USR2` writes `activity` to the daemon's log. The GUI's Daemon tab (ctrl-4) shows all of it, plus the latest events replayed from the backlog, as a snapshot with a Refresh button and optional refresh every two seconds; the header's scanning and hashing badges open it. `python/steward_client.py` is a standard-library client: blocking and asyncio, typed results for the content primitives, and an event stream that reconnects, resumes from the last `seq` and yields a `Gap` when it can't. ## Configuration `$XDG_CONFIG_HOME/steward/settings.toml` (or `$STEWARD_CONFIG`). Roots are subtrees of the one `/` namespace, stored in one index; each carries its own policy, so layers can run on much less than layer 1 covers: ```toml [[root]] path = "~" exclude = ["/down/big", "*.iso"] # gitignore syntax, relative to the root [[root]] path = "/home/media" classify = false contentid = ["Movies", "TV"] # content ids only here ``` Also per root: `interval_minutes`, `full_every`, `one_filesystem`. With no roots the daemon indexes `$HOME`. `steward reload` (or SIGHUP) re-reads the file: new roots start scanning, roots whose settings changed get a full rescan (so a new exclude drops its entries at once), and indexed roots no configured root covers are deleted, also at startup. The daemon refuses to scan outside its roots, so the index only ever holds what the settings describe; `steward --db` works ad hoc. A single database for all roots is deliberate for now: one namespace means one answer to "where else is this content". Splitting per root (independent write locks, removal by deleting a file) stays possible because all SQL is in `steward-index`. ## Paths on the wire Names are stored as raw bytes. On the wire every path is a JSON string, and a byte that isn't part of valid UTF-8 travels as the lone surrogate U+DC80 + byte (`\udcae`), Python's surrogateescape. Rust strings can't hold lone surrogates, so inside steward such a byte is the private-use character U+E000 plus two hex digits (a real U+E000 is doubled), and `steward_proto::wire` translates at exactly two places: the daemon's socket and the Rust client's. Paths are built into answers with `wire::path` and read from requests with `wire::serde_path`; nothing serializes a `PathBuf` directly, which would fail on such names. ## Logging stewardd, `steward` and `steward-ui` log through `tracing` (the `steward-log` crate), one line per event on stderr: `stewardd: warning: hash{path=/home/media/TV}: saving content ids: …; retrying`. Errors and warnings are labelled (coloured on a terminal); spans name the scan, hashing job or client connection a line belongs to. The daemon logs at info by default, `-v` adds debug (each phase of scans and hashing), `-vv` trace (every file, request and event); the CLI is quiet unless something is wrong, and its `-v` steps up from warnings. `RUST_LOG` overrides both (`RUST_LOG=steward_index=trace`). When stderr is the systemd journal, lines carry their syslog priority and no timestamp; otherwise the daemon prefixes local time. ## Known gaps - `locate` matches the final name component only; full-path and faster substring search want a trigram/FTS index (turso has an `fts` feature). - Classification reloads the whole subtree after every scan (1.4 s for `/usr`); it should follow the scan's dirty set instead. - Per-user only. A system-wide instance would need to filter results by the caller's permissions (`SO_PEERCRED`). - The event backlog lives in memory: a daemon restart is always a gap. - Offline handling is tested by faking a directory's stored fsid, not with real mounts.