Resolved -
Things have been stable and we've implemented temporary measures to compensate for the failed PSUs, with replacement and additional spare parts being on their way with express shipments. We're closing this incident.
Aug 28, 15:03 CEST
Monitoring -
The cluster has recovered and we have performed an initial analysis:
Our team was performing routine maintenance work and shut down a storage server to check for a suspected defect power supply (PSU). While reattaching the PSU after inspecting it, this triggered a circuit breaker at around 11:30 and loss of one phase in a single rack.
Unfortunately a neighbouring storage server also exposed a hidden defect power supply connected to the redundant power phase and thus caused the redundancy in the storage cluster to drop below safe levels for serving IO and thus started blocking requests.
We quickly rebooted the affected servers and IO resumed after around 10 minutes at 11:40. A few lagging requests needed to be unblocked manually which finally resolved around 12:10.
In-depth review will happen in the next days and we'll revisit our processes for potential improvements.
Aug 28, 12:33 CEST
Identified -
Due to an electrical issue during maintenance work, we had triggered an overcharge protection, setting multiple storage nodes out of service.
Most hosts came back online, the cluster has recovered and service is restored. We are working on getting the remaining host back.
Aug 28, 12:01 CEST
Investigating -
There was a timeout in the storage cluster causing some VMs to block on storage access.
Aug 28, 11:46 CEST