Skip to main content
Skip to main content
Finished

[HPC] Cluster Outage Due to a Cooling System Malfunction

Aufgrund einer Störung der Kühlanlage am Freitagnachmittag musste der HPC-Cluster ungeplant vollständig abgeschaltet werden.

Die Störung der Kühlanlage konnte inzwischen behoben werden. Der HPC-Cluster wird derzeit schrittweise wieder in Betrieb genommen.

Severity:
High

Due to a malfunction in the cooling system on Friday afternoon, the HPC cluster had to be shut down completely on an unplanned basis.

The cooling system malfunction has since been resolved. The HPC cluster is currently being brought back online in stages.

 

Jobs that were running at the time of the shutdown were interrupted by the outage. Affected jobs will, where possible, be automatically restarted (requeued) by Slurm. Users are asked to check the status of their jobs and, if necessary, resubmit any failed jobs.
Jobs that were already pending remain in the queue and will be started automatically as soon as the required resources are available again.

 

Ongoing file transfers may also have been interrupted by the outage. Please check the status and completeness of your transferred data as needed.

In the event of any further unforeseen issues, we will, of course, inform you immediately, as always.

Everything at a glance
Type
Incident
Start date
25.09.2026 15:30 o'clock
End date
26.09.2026 08:00 o'clock
Severity
High

Contact person

Error

Mustafa André Schmidt

Employees