Things have been stable and we've implemented temporary measures to compensate for the failed PSUs, with replacement and additional spare parts being on their way with express shipments. We're closing this incident.
Posted Aug 28, 2026 - 15:03 CEST
Monitoring
The cluster has recovered and we have performed an initial analysis:
Our team was performing routine maintenance work and shut down a storage server to check for a suspected defect power supply (PSU). While reattaching the PSU after inspecting it, this triggered a circuit breaker at around 11:30 and loss of one phase in a single rack.
Unfortunately a neighbouring storage server also exposed a hidden defect power supply connected to the redundant power phase and thus caused the redundancy in the storage cluster to drop below safe levels for serving IO and thus started blocking requests.
We quickly rebooted the affected servers and IO resumed after around 10 minutes at 11:40. A few lagging requests needed to be unblocked manually which finally resolved around 12:10.
In-depth review will happen in the next days and we'll revisit our processes for potential improvements.
Posted Aug 28, 2026 - 12:33 CEST
Identified
Due to an electrical issue during maintenance work, we had triggered an overcharge protection, setting multiple storage nodes out of service. Most hosts came back online, the cluster has recovered and service is restored. We are working on getting the remaining host back.
Posted Aug 28, 2026 - 12:01 CEST
Investigating
There was a timeout in the storage cluster causing some VMs to block on storage access.
Posted Aug 28, 2026 - 11:46 CEST
This incident affected: RZOB (production) (VM storage cluster).