Appendix E - Operational Runbooks and Checklists
Ingestion degradation
Symptoms
- increased
5xxor retryable responses; - rising database acquisition latency;
- old mobile queues reported by clients;
- commit latency above target.
Immediate actions
- Confirm user impact and affected roles.
- Check database CPU, I/O, WAL, locks, and pool saturation.
- Check recent migrations or deployments.
- Reduce non-critical worker concurrency.
- Protect ingestion by pausing reports and exports.
- Scale API replicas only if the database has headroom.
- Communicate retry guidance and monitor recovery burst.
Integrity checks
- accepted batch count versus stored point count;
- duplicate rate;
- outbox lag;
- latest-state freshness;
- mobile acknowledgement behavior.
WebSocket reconnect storm
- Verify whether a deployment or proxy timeout triggered disconnects.
- Confirm clients use jittered backoff.
- Observe authentication and snapshot load.
- Reduce optional realtime detail and increase coalescing.
- Keep critical alert messages prioritized.
- Drain or roll replicas gradually.
- Run a post-incident reconnect benchmark.
Disk pressure
- Identify filesystem and growth source.
- Protect PostgreSQL from reaching zero free space.
- Pause non-essential exports and large reports.
- Verify WAL archive health and replication slots.
- Add capacity or safely remove verified expired temporary data.
- Do not manually delete PostgreSQL data files.
- Review retention and capacity alerts.
Failed backup
- Treat stale backup coverage as an incident.
- Preserve the last known good backup and keys.
- Diagnose storage, credentials, capacity, and WAL archive state.
- Retry through the documented tool, not an improvised copy.
- Run integrity verification.
- Schedule an immediate restore drill when coverage is restored.
Suspected credential compromise
- Identify credential class and scope.
- Revoke or rotate it.
- Invalidate related sessions or token families.
- Search security and audit events.
- Check for false points, exports, public links, or commands.
- Preserve evidence with access controls.
- Notify affected parties according to policy.
- Correct derived data through versioned workflows.
Release readiness checklist
- All migrations tested from the supported previous release.
- Fresh database bootstrap passes.
- Cross-tenant suite passes with runtime role.
- Ingestion duplicate and partial-response tests pass.
- WebSocket reconnect and slow-client tests pass.
- Mobile field tests cover supported OS versions.
- Protocol fuzz corpus passes.
- Backup and PITR drill is within target.
- Image digests, SBOM, and signatures are recorded.
- Alerts and runbooks are enabled.
- Rollback has been rehearsed.
- Known limitations are approved and communicated.
Daily operator checklist
- Ingest success and latency within objectives.
- Active-session freshness within objective.
- Critical job backlog age acceptable.
- Database disk and WAL archive healthy.
- Future partitions exist.
- Backup completed and verified.
- No unexpected dead-letter growth.
- No abnormal invalid GPS frame spike.
Monthly resilience checklist
- Restore a backup into an isolated environment.
- Verify RLS and application smoke tests on the restore.
- Exercise one failure runbook.
- Review capacity forecast and retention.
- Rotate one non-emergency credential in a controlled drill.
- Review support and emergency access.
- Review public links and stale API clients.
- Re-run representative mobile battery tests after major OS updates.