Impact: All platform API requests failed, so the console, CLI and integrations
were unusable. The website and marketing pages were unaffected.
What happened: The primary database server's storage filled completely, which
prevented PostgreSQL from starting and left the cluster with no primary. Space
was freed and the database recovered immediately. The issue is fully resolved
and no data was lost.
What we're doing to prevent recurrence:
- Alerting on database storage headroom, so we act well before a volume
approaches full.
- Monitoring replication health directly. A replica that had silently stopped
keeping up is what consumed the space, with no outward sign.
- Restoring and routinely verifying our standby database, so a failure of this
kind is handled automatically rather than manually.