On 31 July 2026, the SSD Block Storage cluster in Amsterdam experienced an outage of 56 minutes, starting at 15:32 CEST and solved by Tilaa engineers at 16:28 CEST.
The issue was caused by a storage manager node that promoted itself to primary manager, after it deemed the other storage manager nodes of the cluster to be unavailable due to time synchronisation issues.
However, the system clock of this now-primary manager was the one that actually had a clock offset of about 6 seconds compared to the other nodes.
The faulty storage manager rejected all other manager nodes in the cluster due to these time synchronisation issues, causing 27 networked storage disks to be reported unavailable.
The issues were reported by the SSD Block Storage cluster software as warnings, and not as critical events, which is why we only became aware of this issue at 16:02 CEST.
After investigation and manual intervention, the Ceph cluster returned to a healthy state, and all storage nodes in the Ceph cluster became available again around 16:30 CEST for client disk operations.
We have made the following improvements: