Lattice Database Partial Outage

Incident Report for Orbee

Resolved

The issue is now fully resolved, and monitoring has shown all operations has gone back to normal.

Normal database maintenance and optimization operations in the background combined with a node reboot caused some replica nodes to enter a long, blocking startup phase that required cleanup of the incomplete processes as well as normal startup operations. During this time, the node would still accept queries but not have the capacity to execute them, causing connections and queries to hang. These queries would persist, queued, and cause the node to fail performance-wise following startup, causing another reboot. This then triggered the process all over again.

Once the affected nodes were quarantined, query operations resumed normally on healthy nodes. We also paused the running of schema and ingestion tasks, which queued instead. This allowed the nodes to complete the startup and optimization processes without impedance.

Once the nodes were operational they could rejoin the cluster and resume accepting queries. Once the cluster was fully healed, all operations could resume. The ingestion + schema changes ran into no roadblocks and query performance continued to perform normally.

Alerts and quarantine automation have been created to handle this situation in the future with much faster resolution times and little to no noticeable changes to performance and capabilities in the Platform.
Posted Aug 14, 2026 - 20:49 PDT

Monitoring

All nodes are now functional and operating normally. There is no inconsistent behavior anymore, and all queued schema changes and data ingestion are running normally.

We'll continue to monitor to verify that the nodes remain healthy and that everything wraps up accordingly.
Posted Aug 14, 2026 - 19:57 PDT

Update

50% of the affected replication nodes are now healthy; we are continuing to work on getting the remaining nodes back into an operational state. Querying data continues to not be affected.
Posted Aug 14, 2026 - 18:00 PDT

Identified

We have identified degraded performance in a portion of our customer data warehouse system, causing the nodes to fall behind in their replication process. We've isolated the affected nodes and are working on getting them back up-and-running.

At this time, customers can read their data using Data Studio and the associated Record views. Ingesting new data and changing schemas for native objects are queued but not being processed at this time.

We'll provide another update as soon as we get the affected nodes back up-and-running.
Posted Aug 14, 2026 - 15:59 PDT
This incident affected: CDP.