Block storage not reachable

Incident Report for Tilaa

Postmortem

On 31 July 2026, the SSD Block Storage cluster in Amsterdam experienced an outage of 56 minutes, starting at 15:32 CEST and solved by Tilaa engineers at 16:28 CEST.

The issue was caused by a storage manager node that promoted itself to primary manager, after it deemed the other storage manager nodes of the cluster to be unavailable due to time synchronisation issues.

However, the system clock of this now-primary manager was the one that actually had a clock offset of about 6 seconds compared to the other nodes.

The faulty storage manager rejected all other manager nodes in the cluster due to these time synchronisation issues, causing 27 networked storage disks to be reported unavailable.

The issues were reported by the SSD Block Storage cluster software as warnings, and not as critical events, which is why we only became aware of this issue at 16:02 CEST.

After investigation and manual intervention, the Ceph cluster returned to a healthy state, and all storage nodes in the Ceph cluster became available again around 16:30 CEST for client disk operations.

We have made the following improvements:

  • The reachability of our support team has been improved with better call rotation
  • Our monitoring and alerting systems will report any of these events as critical, immediately sending out alerts to our on-duty engineers
Posted Aug 04, 2026 - 16:54 CEST

Resolved

The incident has been resolved, and we got confirmation from customers that everything is back as it should
Posted Jul 31, 2026 - 18:19 CEST

Monitoring

We have identified the root cause of the issue that caused our Block Storage (Big Disk) to become unavailable.
We are currently seeing IOPS returning to normal levels on the Amsterdam SSDs, and all affected services are actively recovering.

We will continue to monitor the situation closely.
Posted Jul 31, 2026 - 17:05 CEST

Investigating

We are currently investigating an issue affecting our Block Storage (previously known as Big Disk).

At this time, the impact appears to be isolated to SSD storage in Amsterdam.
We are working to identify the root cause and will provide an update as soon as more information becomes available.
Posted Jul 31, 2026 - 16:16 CEST
This incident affected: Data Center (Amsterdam) and Services (Big Disk).