Incident on a compute node (PAR01)
Affected components : Virtual servers (Partial outage)
- Resolved
Post-Mortem of the 05/11/2023 Incident Affecting a Compute Node at the Paris-Saint Denis Data Center:
At 00:17, our operational team identified an incident on one of our racks within the Paris-Saint Denis data center, resulting in the loss of one of the dual power feeds.
This did not affect our production as all our infrastructure equipment is equipped with redundant power on two separate circuits.
At 00:35, we engaged the data center's maintenance staff to investigate this loss and to reset the involved circuit breaker if it was necessary.
At 01:26, amidst the maintenance efforts, we recorded the failure of one of our compute nodes (Compute Server), which coincided with the restoration of the auxiliary power supply, returning the power to normal levels.
We were immediately concerned that an error in handling the auxiliary breaker reset might have led to the disconnection of one or more power cables in the rack.
At 01:57, without a swift resolution, we opted to migrate the hosted VMs to alternate nodes within the cluster, ensuring the continuation of service for our impacted clients.
At 02:26, we had successfully switched and restored all VMs affected by the incident and dispatched team members on-site to diagnose the still-out-of-service node.
At 03:01, our team on-site verified our hardware and found that several electrical plugs appeared to be incorrectly inserted into the PDUs, after which they adjusted them.
At 03:11, our system team regained control over the impacted node and reintegrated it back into our resource pool. The entire rack was checked to ensure all connections were secure and there were no further power issues.
We confirm that this incident has been fully resolved. A service disruption was noted for some clients from 01:26 to 01:57, with total service restoration for all by 02:26.
We apologize for any inconvenience this may have caused and are available for any queries through our ticket support system.
- Identified
All VMs have been transferred to the operational nodes within the hosting cluster and have been restored at this stage.
If your VPS is still not operational despite this, please initiate a reboot from the client area or contact technical support.
Concurrently, we are dispatching a member of our team to perform a more comprehensive diagnostic of the hosting node that remains unavailable despite the restoration of the auxiliary power path.
- Identified
We have just switched the majority of VMs to alternate nodes within the architecture to restore service for the majority of clients.
A small number of VMs remain unavailable, and our teams are continuing to work on restoring them.
The next update will be provided in 15 minutes.
- Identified
This incident occurred following a non-impact event on one of our racks at the Paris-Saint Denis data center, which resulted in a power path loss at 00:17 AM.
We mobilized the data center teams at 00:35 AM to restore the auxiliary power path.
At this stage, we believe human error may have led to the unplugging of an electrical cord from one of our compute servers during efforts to re-establish the auxiliary path, consequently severing the last functioning power supply to this server.
We are waiting for an update from the data center team.
The next update will be provided in 15 minutes.
- Investigating
We are currently experiencing an issue with a compute node in the Paris-Saint Denis datacenter; our teams are investigating.
Some VMs are currently unreachable.
Next update in 15 minutes.