Skip to main content

Backup and restore

A mesh already survives losing nodes — claims are replicated, and a node that disappears is covered by its replicas. That protects you from hardware. It does not protect you from a bad deploy, a wrong DELETE, or a mapping change you want to undo, because the mesh replicates those faithfully.

Snapshots are for that: a point-in-time copy in storage you control, which a node can put back.

Every route on this page requires the admin role.

The shape of it​

A repository is where backups live — a directory, or an S3-compatible bucket. You register it once per node.

A snapshot is one capture of one index or namespace into a repository. Snapshots inside a repository share segments, so ten daily snapshots of a mostly-unchanged index cost far less than ten copies.

Snapshot, restore and cleanup are jobs. They run in the background and the CLI waits for them by default; --no-wait returns a job id instead.

Register a repository​

A directory, which is the right choice when it is a mount you already back up:

gnarl repo add nightly --path /mnt/backups/gnarl

Or an S3-compatible bucket — AWS, GCS, R2, MinIO:

gnarl repo add offsite \
--endpoint https://s3.us-east-1.amazonaws.com \
--bucket gnarl-backups \
--region us-east-1 \
--prefix production/ \
--access-key-id AKIAEXAMPLE \
--secret-access-key "$AWS_SECRET_ACCESS_KEY"

--prefix lets one bucket hold several repositories. --virtual-host switches from path-style addressing (/bucket/key) to virtual-host style; MinIO and most self-hosted stores need path style, which is the default.

Registration opens and lists the repository before accepting it, so a wrong bucket or a stale credential fails here — at the command you are watching — rather than inside a scheduled backup that nobody reads the output of.

gnarl repo list
gnarl repo remove nightly

repo remove forgets the registration. It does not delete the backups. Deleting someone's backups as a side effect of tidying a config is not something you can undo.

Encryption keys are not recoverable

--encryption-key takes a 32-byte hex key and encrypts the repository at rest:

gnarl repo add offsite --path /mnt/backups/gnarl --encryption-key "$GNARL_REPO_KEY"

The key is never returned by any endpoint — not by repo list, not masked, not at all. Nothing in the product can recover it, and a repository encrypted with a key you no longer have is indistinguishable from random bytes.

Store it wherever you keep credentials you cannot re-issue, and store it somewhere other than the machine being backed up. A key that only exists on the node it protects is not a backup key.

Take a snapshot​

gnarl snapshot create nightly 2026-09-22 --index products
gnarl snapshot list nightly
gnarl snapshot show nightly 2026-09-22

snapshot show reports signature_verified — whether this node checked the snapshot's signature against a key it independently trusts. That matters at restore time, below.

Namespaces​

A namespace is logical, so how it gets captured depends on how it is stored, and the job result tells you which happened as kind.

gnarl snapshot create nightly acme-2026-09-22 --namespace acme

A promoted namespace has a dedicated physical index and is captured exactly: segments copied and pinned per claim.

A pooled namespace shares its segments with every other namespace that hashes into the same pool, so it has no physical subset of its own. Its documents are exported instead (kind: logical). That capture is exact in content but has two properties worth knowing:

  • It is not a point-in-time cut. It walks a live index, so a document written during the walk may or may not be included.
  • It restores by re-ingesting, not by putting bytes back.

Promote the namespace before snapshotting when exactness matters.

A keyed namespace exports with its _source still sealed. The tenant's key is never needed to take the backup and never reaches it.

Snapshotting a __pool_N index by name is different again: it captures every namespace stored in that pool, which is why it requires --allow-shared-pool. Read that flag as "yes, I mean every tenant in this pool".

Restore​

gnarl snapshot restore nightly 2026-09-22

Restoring a namespace snapshot also routes the namespace back at its dedicated index. Without that the documents would be restored and invisible, because namespace queries route by promotion state rather than by what is on disk.

Two flags that turn a refusal into data loss

Restore refuses some things on purpose. Both overrides exist because there are legitimate reasons to proceed — and both are irreversible when you are wrong.

--allow-overwrite-live-index rolls a live index back to the snapshot. Without it, restoring over an index that already exists is refused. With it, everything written since the snapshot is gone. Restore to a new name first if you want to look before you commit.

--allow-unverified-signer accepts a snapshot this node cannot attribute to itself or to a known mesh peer. By default that is refused, because an unverifiable snapshot is one you cannot prove came from your own fleet. Prefer naming the key instead:

gnarl snapshot restore nightly 2026-09-22 --signer-public-key <hex>

--signer-public-key must be the key the snapshot names as its signer, so it is an assertion you can check, not a bypass.

Reclaim space​

gnarl snapshot delete nightly 2026-09-22

That deletes the descriptor only. Segments are shared between snapshots, so reclaiming them is cleanup's job — judged against every snapshot that survives, never a side effect of deleting one.

gnarl snapshot cleanup nightly
gnarl snapshot cleanup nightly --grace-seconds 3600

Cleanup marks from every descriptor, then sweeps. Two safeguards are worth knowing because they explain why it sometimes frees less than you expect:

  • If any descriptor cannot be read, the sweep is abandoned and nothing is deleted. One unreadable descriptor makes everything it references look unreferenced, and deleting on that basis would destroy live backups.
  • Objects younger than the grace period are never collected (default 24 hours). A snapshot writes its segments before its descriptor, so a sweep that ignored age would delete the segments of a snapshot still in flight.

Jobs​

gnarl snapshot jobs
gnarl snapshot jobs --id <job-id>

One snapshot job runs at a time per node; a second returns 409.

History is durable across restarts, with one thing to watch: a job that was running when the node stopped is reopened as failed, because whether it finished is unknown. That is a statement about our knowledge, not about the data — list the repository to see what it actually wrote.

HTTP API​

The CLI is a client of these. Every route requires admin.

MethodPathPurpose
GET/v1/repositoriesRegistered repositories; credentials never returned
PUT GET DELETE/v1/repositories/{repo}Register, inspect, unregister
POST/v1/repositories/{repo}/_cleanupMark and sweep unreferenced segments
GET/v1/repositories/{repo}/snapshotsSnapshots in a repository
PUT GET DELETE/v1/repositories/{repo}/snapshots/{snapshot}Create, describe, delete
POST/v1/repositories/{repo}/snapshots/{snapshot}/_restoreRestore
GET/v1/snapshot_jobsJobs, newest first
GET/v1/snapshot_jobs/{id}One job

Register a repository:

curl -X PUT http://localhost:8080/v1/repositories/nightly \
-H 'Content-Type: application/json' \
-d '{"type": "fs", "location": "/mnt/backups/gnarl"}'

Take a snapshot — 202 with a job id, because the work outlives the request:

curl -X PUT http://localhost:8080/v1/repositories/nightly/snapshots/2026-09-22 \
-H 'Content-Type: application/json' \
-d '{"index": "products"}'

Restore, naming the signer rather than waiving the check:

curl -X POST http://localhost:8080/v1/repositories/nightly/snapshots/2026-09-22/_restore \
-H 'Content-Type: application/json' \
-d '{"signer_public_key": "<hex>"}'

Poll the job:

curl http://localhost:8080/v1/snapshot_jobs

Exit codes​

A failed restore exits 1, and so does an unverifiable signature.

That is worth stating plainly because it is a limitation: a scheduled job cannot currently distinguish "this backup is not ours" from "the network was down" by exit code alone. ExitCode::VerificationFailed is reserved as 4 but is not emitted yet. Until it is, check the job's error and the descriptor's signature_verified rather than branching on the status. The full table is in the CLI reference.

A workable routine​

  1. Register a repository off the node — a different disk at minimum, a bucket in a different account for anything you would miss.
  2. Snapshot on a schedule, named by date.
  3. cleanup on a slower schedule; let the grace period do its job.
  4. Rehearse a restore into a new index name. A backup nobody has restored is a hypothesis. --allow-overwrite-live-index should not be the first time you run the command.
  5. Keep the encryption key somewhere the node cannot take with it.