Skip to content

Bootstrap and recovery

A new node does not replay the chain from genesis. It could not: no live peer serves the data a genesis replay would need, and post-quantum signature volume makes full replay impractical anyway.

Instead it verifies a certified checkpoint chain, syncs its execution layer to the network’s head, and restores a consensus snapshot. It lands on the pruned tier, with the sync pivot as its history cutoff; deeper history comes from the cold archive with frost-archiver restore-el.

genesis committee digest already trusted — it is in the genesis block
↓ verifies
checkpoint chain hash-linked, quorum-signed by checkpoint roots
↓ authorises
block-key schedule each delegation checked against the roots
↓ anchors
verified state snapshot

The joining node, configured with bootstrap_from and starting with no local state:

  1. Fetches the checkpoint chain from the peer’s ingress over the open, paged frost_getCheckpointChain and frost_getJournalTail methods, and verifies it against the committee digest recorded in genesis — not against the peer’s word. The result is a certified anchor: a commit index, the block number and hash it closed, and a state root.
  2. Snap-syncs its execution layer over devp2p to the peer’s journal tail. The state is hash-verified against headers that chain to the anchor. This step is 8141-geth only: frost-reth has no snap sync, so checkpoint bootstrap of a fresh frost-reth node is not available today.
  3. Downloads the peer’s consensus snapshot from GET /bootstrap/snapshot, presenting the fleet’s bearer token, capped at bootstrap_from.max_snapshot_bytes.
  4. Verifies the staged snapshot against the anchor: the chain must re-verify and extend what step 1 proved, the journal must record the anchor and the synced head, the store’s anchor commit must carry the certified digest, and the block-key schedule must hash to the certified root.
  5. Installs it, journal last — the journal is the “installed” marker, so a crash anywhere earlier leaves a restart to redo steps 3–5 idempotently.
  6. Follows normally from there.

The critical property is step 1. Trust bottoms out in the chain definition the node had to have in order to be on this network at all. The serving peer is not trusted — it is a data source, and everything it provides is checked.

This is a proof-of-stake chain, so the standard caveat applies and it is worth stating plainly rather than burying.

A node joining from a checkpoint cannot independently verify the entire history back to genesis, because it does not have the signatures on old blocks. It verifies that a quorum of the genesis-pinned committee certified a chain of checkpoints leading to the state it is adopting. That is a strong guarantee, and it is not the same as replaying everything.

The residual assumption is that no committee checkpoint-root key set has ever been compromised. Root signatures do not expire, so stolen root keys could certify a competing history — the classic long-range attack.

Block keys are outside this assumption. They expire and rotate, so a stolen block key is a bounded, revocable exposure rather than a permanent one. That narrowing is one of the main reasons the key tiers are split.

Restart after downtime. The write-ahead journal reconciles against the execution layer’s head and the node resumes. Blocks the execution layer lost are re-driven from replayed commits. See Running frost-node.

Below the commit floor. Where operators set min_commits, validators keep at least that many commits behind the head whatever the byte clamp says, so a node that returns within that window re-joins by commit sync at full speed.

Beyond the retention window. If a node has been down longer than peers retain consensus history, it cannot catch up incrementally. Re-bootstrap it: frost-node run --config node.yaml --force-bootstrap wipes the consensus state — the consensus database, driver journal, checkpoint files and prune floor, never the execution-layer datadir — and re-runs the join from the configured bootstrap_from peer. It is refused without one. A follower in this state suspends its stream and reports frost_observer_stranded = 1 rather than churning.

The node halted on a hash mismatch. This is not a “restart and hope” situation. It means a rebuild produced a different block hash than the journal recorded, which means nondeterministic execution. Investigate before restarting: check that the binary matches the recorded digest for the pinned commit, and that the consensus parameters compiled into it are the network’s.

A node running different consensus-critical parameters does not degrade gracefully — it splits. See Requirements.

Nodes follow the chain only while the checkpoint chain remains fresh. Staleness is counted in block-producing commits since certification last advanced: past 3 checkpoint intervals the node alarms, past 10 it halts. At the network’s 600-commit interval and its observed commit rate, that is roughly 1.6 and 5.4 minutes.

Followers additionally have a stall watchdog (follower_stall_warn_secs, default 120): a follower whose handled commit index stops advancing warns, repeating while the stall lasts, and says when it resumes.

The reasoning: a stale checkpoint chain means block-key delegation authority is no longer being confirmed. Continuing to follow blindly, on the authority of delegations nobody is currently certifying, is precisely the failure mode this guard prevents.

If you see freshness alarms, the problem is usually network-wide rather than local — check whether checkpoints are being certified at all.

A state-archive node is rebuilt mechanically rather than synced: replay full-witness blocks from the cold object-store archive in archive mode.

Everything from the archive is hash-verified against the chain’s own commitments before use. The storage is treated as untrusted, so corruption or loss is detected rather than believed.

This is a rebuild-from-scratch operation, not a repair. Budget time for it.

Keys — especially checkpoint root keys. They cannot be regenerated, and losing them means losing the ability to certify checkpoints.

Manifest and config — small, and annoying to reconstruct.

Not chain data. Bootstrap exists precisely so that chain data is reconstructible from the network with full verification.