# Monitoring and incident response One page: what is watched, who answers, and what the answer is for each thing that can go wrong. Written 2026-09-11 for the Trail of Bits maturity roadmap item 3; it is the operating counterpart of `docs/MINT-SECURITY.md` (the site threat model) and `SAFE-SETUP.md` (key custody). ## Who answers One operating entity, DISORDERLY LLC, one person on call: the founder. The second Safe signer is the escalation for anything that needs two signatures. Acknowledgement target: within 4 hours for an alert, next business day for a notice. Every acknowledgement and action is a dated line in `cycles/notices.log` on the droplet (the supervisors already append there). ## What is watched | watcher | runs | what it covers | how it reaches a person | |---|---|---|---| | `scripts/watch-events.js` (unit `disorderly-watch`) | every minute on the droplet | every state-changing event of the four contracts and the three Safes (treasury, reserve, ops: owners, threshold, modules, guard, executions), classified alert / notice / digest (table below) | `NOTIFY_WEBHOOK` (the Discord channel a signer reads), `NOTIFY_LOG`, journal | | governance and settlement supervisors | continuous | the human steps they reach: a batch to sign, a record ready, a commitment anchored, a close batch | same webhook | | `deploy/preflight.sh --server` | before every deploy and once a day by hand during the mint window | DNS and mail auth, TLS, unit health, backups, pending reboot, mint-page drift (pinned address vs `mint/config.json`) | terminal; a FAIL blocks the deploy | | nightly backup (`disorderly-backup`) | 03:00 UTC | database dump and runtime file archive, 14 days | `deploy/test-restore.sh` proves a restore; run it monthly | | UptimeRobot | every 5 minutes | the site and the API answer | email to hello@ | | `deploy/cloudflare-only.sh` (monthly cron on the droplet) | monthly, and by hand after install | ports 80 and 443 accept connections from Cloudflare's published ranges only, so the origin address is useless to anyone who learns it; SSH untouched | the log at `/var/log/cloudflare-ips.log`; a change in the published ranges is applied before old ones are removed | The watcher's pointer is `server/data/watch-events.json`; it only advances after a window was read in full, so an RPC outage delays alerts rather than losing them. Delivery is at least once as of 2026-09-15: every alert and notice is written to that file the moment it is read from the chain, before any post, and stays queued until the webhook has accepted it, retried on every poll with an attempt count in the title. Every payload carries a stable id, `:` (the `key` field), so an event is never queued twice and a receiver can drop the rare duplicate that a crash between acceptance and acknowledgement produces. The state file is replaced atomically and the previous copy kept as `.bak`; a corrupt file is recovered from the `.bak` or the watcher refuses to start, it is never read as empty. The webhook post times out after `NOTIFY_TIMEOUT_MS` (10 s). Before this, the pointer advanced past an alert the webhook had rejected and the alert survived only in the local log. Without `NOTIFY_WEBHOOK` set, console and log are the channels there are and the queue drains to them, which the startup line says. ## Event classes | class | events | expected pattern | |---|---|---| | alert | `CouncilRootUpdated`, `OperatorRootUpdated`, `RoyaltyReceiverUpdated`, `OwnershipTransferStarted`, `OwnershipTransferred` (all four owned contracts), `ProposalAmended`, `Skimmed`, `PaymentDeferred`, `ValuationDeferred`, Safe `AddedOwner`, `RemovedOwner`, `ChangedThreshold`, `EnabledModule`, `DisabledModule`, `ChangedGuard`, `ExecutionFailure` (each of the three Safes) | never without a Safe transaction the signers recognise and, for roots and the receiver, a published reason. A module or guard change is the way to move a Safe's funds without the owners' signatures, so it is always an incident unless planned. Any alert that does not match a planned action is an incident. | | notice | `MintOpened`, `MintFinalized`, `RevealCommitted`, `StartingIndexSet`, `TierOffsetsSet`, `BaseURIUpdated`, `UnrevealedURIUpdated`, `MetadataLocked`, `Withdrawn`, `CycleOpened`, `Swept`, `DeliberationCommitted`, `ProposalPublished`, `CyclePublished`, `ThresholdReached`, `BacklogValued`, `OwedPaid`, `TokenSwept`, `WethUnwrapped`, Safe `ExecutionSuccess` | the scheduled steps; each should match a runbook step or a supervisor notice that preceded it | | digest | `CouncilMinted`, `OperatorMinted`, `Claimed`, `Received`, `Released`, `Accrued` | routine traffic, counted and posted hourly | ## Playbook Each entry: the signal, the first check, the action, and what must never be done. "Pause" means stop the clocks (`systemctl stop disorderly-supervise disorderly-settle`); the API stays up so holders can still read. **Signer device lost, stolen or suspected compromised.** Signal: the person knows, or an unplanned Safe alert. First check: the Safe's owner list on Etherscan. Action: with the other two keys, `removeOwner` on all three Safes and `addOwner` for a freshly initialised device per SAFE-SETUP.md, threshold stays 2; record the rotation in ADDRESSES.md; pause the clocks until done so no batch waits on the old key. Never: reuse the seed, or lower the threshold to move faster. **Deployer key between deploy and `acceptOwnership`.** Signal: the window itself, open until the Safe accepts. Action: keep it under one hour; `transferOwnership` to the Safe in the deploy transaction batch and have both signers ready to `acceptOwnership` immediately; the watcher's `OwnershipTransferStarted` and `OwnershipTransferred` notices are the confirmation. If an unexpected `OwnershipTransferStarted` appears during the window, the deployer key is compromised: do not open the mint, redeploy from a new deployer, and abandon the first contract before any mint reaches it. **Mint page drift.** Signal: `preflight.sh` reports the served `mint/config.json` or the pinned address block differs from the repo, or a holder reports a different address. Action: treat as a site compromise per MINT-SECURITY.md: put the site into maintenance from the Cloudflare side, redeploy from the dev machine, rotate the SSH key and any web credentials, then compare the served page byte for byte. The contract address is published on the site, in ADDRESSES.md, on Etherscan and in the launch thread; a mismatch anywhere is the signal. Never: patch the live file by hand on the droplet. **Reveal window about to expire.** Signal: `RevealCommitted` notice, then no `StartingIndexSet` within about 200 blocks of the target (the hash is readable for 256 blocks after it). Action: anyone can call `setStartingIndex`; the founder does it from any wallet with gas. If the window is missed, `commitReveal` again (permissionless once the old target has aged out) and publish why. Never: attempt to time the commit against a block; the delay exists so nobody can. **Stalled Safe gate.** Signal: a supervisor notice (a batch to sign) with no matching Safe `ExecutionSuccess` inside the cadence, or a `close-batch` notice repeated. Action: the second signer signs; if a signer is unreachable for more than the claim window's slack, the batch waits, and the ledger is left alone until it executes (SAFE-SIGNING.md: never re-prepare over a pending batch). Never: execute from a single key, which the Safe refuses anyway. **Deferred router leg (`PaymentDeferred`).** Signal: the alert names the destination that refused ETH. First check: is the destination still the intended Safe, and does it hold code on this chain. Action: fix the destination (a Safe that cannot receive is a Safe module or guard problem), then anyone calls `pushOwed(destination)`; `OwedPaid` confirms. The other two legs were paid; nothing is stuck. Never: repoint the royalty receiver to route around it before the cause is known. **Feed unusable (`ValuationDeferred`).** Signal: the alert; the ETH went to the reserve unvalued and is remembered. Action: check Chainlink's status for ETH/USD; nothing to do on chain, the next healthy `release()` values the backlog (`BacklogValued`). If the feed is deprecated before the threshold is met, the split never starts; that is the accepted design and the reserve keeps everything. Never: replace the router on a single stale reading. **Amended record (`ProposalAmended`).** Signal: the alert with the amendment index and reason hash. Action: the published reason must exist at the reason hash before the amendment; if the amendment was not planned, pause the clocks and treat the Safe as compromised (only the Safe can amend). Never: amend without the published reason. **Unplanned root or receiver change, or an owner transfer.** Signal: any of the five alerts without a matching planned Safe transaction. Action: pause the clocks, verify the Safe's owner list, and if the Safe itself executed it, one of the two signing keys is compromised: rotate per the first entry. Before `finalize`, a hostile root can be replaced by the true one from the archived list; after `finalize` roots cannot change. The receiver can be repointed back. Never: assume it was a mistake. **RPC failing (`INCOMPLETE` in the recorder, watcher retries).** Signal: supervisor log lines about receipts or logs not returned; the watcher's "poll failed, will retry". Action: nothing is corrupted, recording fails closed and retries; switch `RPC_URL` to the fallback provider in `.env` and restart the units if it persists past an hour. Never: record a payout by hand because the RPC is slow. **API credits exhausted (402 from the model).** Signal: a supervisor stage stalls with a credit error. Action: top up the prepaid balance under the LLC (OPS-SPEND.md); the stage resumes on the next pass. Never: switch to an auto-reload card. **Droplet down.** Signal: UptimeRobot. Action: DigitalOcean console; if the host is gone, restore from the nightly backup per `deploy/test-restore.sh` onto a fresh droplet with `deploy/setup-server.sh`, re-push the env with `deploy/push-env-keys.sh`, and re-enable the units. The chain is the record; nothing on the droplet is the only copy of a signed anchor. ## Rehearsal Rehearsed once before mainnet and dated here: the founder runs `watch-events.js --once` against a Sepolia block range with known events, executes a planned `setRoyaltyReceiver` on the rehearsal 721 and confirms the alert arrives in the channel within a minute, then walks the signer-rotation entry on the Sepolia Safes with both signers. | rehearsed | date | notes | |---|---|---| | watcher decodes known events | 2026-09-11 | Sepolia blocks 11670814-11670818 (commit) and 11671080-11671085 (publish) | | watcher installed on the droplet as `disorderly-watch` | 2026-09-15 | started by `enable-supervisors.sh`; log line names all five contracts (721, distributor, registry, router, treasury Safe) on chain 11155111; `ROYALTY_ROUTER` and `TREASURY_SAFE` pushed with `push-env-keys.sh` | | planned alert arrives in the channel | not yet | | | signer rotation on the Sepolia Safes | not yet | |