-
Incident
EPYC-501 defective mainboard
ResolvedDue to a hardware defect on the motherboard, this system is currently unstable and reboots intermittently. We are already in contact with our dealer. Whenever the system shuts down, we restart it immediately. We realize this isn't a clean solution, but we have to wait for our dealer to resolve the issue.
-
Identified (1 Monat ago)
The server crashed again at 11:31 a.m. due to a "fatal hardware error." We restarted it immediately. The server should be back online in a few minutes.
Since our motherboard is end-of-life, we are currently waiting for a shipment from AsRock in Taiwan to receive a replacement. As soon as it arrives in Germany, we will announce a maintenance window to replace it.
-
Identified (1 Monat ago)
The server was restarted after a crash. We are already on our way to Frankfurt with replacement hardware for maintenance.
-
Resolved (1 Monat ago)
Replacing the motherboard did not resolve the issue: Neither the replacement board nor the originally installed board could be operated stably. We have therefore migrated all VMs from EPYC-501 to our 3rd-generation EPYC host systems. They have been running stably there ever since, and all data and configurations were fully migrated.
EPYC-501 has thus been deactivated and is being sent to the vendor for inspection. The recurring reboots and “fatal hardware error” crashes have finally come to an end for all affected servers—we are closing this incident.
A full review of yesterday’s maintenance, as well as the 14-day runtime credit for all customers with an EPYC 5th Gen server, can be found in the postmortem for the maintenance window. As things stand, we plan to complete the rollback to 5th-Gen hardware by the end of August and will announce it in a timely manner as a separate maintenance window.
If your server is behaving differently than usual or is unavailable, please contact Support briefly.
Affected: Host System - AMD EPYC 5th Gen Partial outage -