[HPC] Cluster Outage Due to a Cooling System Malfunction
Due to a malfunction in the cooling system on Friday afternoon, the HPC cluster had to be shut down completely on an unplanned basis.
The cooling system malfunction has since been resolved. The HPC cluster is currently being brought back online in stages.
Jobs that were running at the time of the shutdown were interrupted by the outage. Affected jobs will, where possible, be automatically restarted (requeued) by Slurm. Users are asked to check the status of their jobs and, if necessary, resubmit any failed jobs.
Jobs that were already pending remain in the queue and will be started automatically as soon as the required resources are available again.
Ongoing file transfers may also have been interrupted by the outage. Please check the status and completeness of your transferred data as needed.
In the event of any further unforeseen issues, we will, of course, inform you immediately, as always.