chip SLURM Job Queueing Issue
Hi everyone,
DoIT staff are aware of an issue regarding SLURM job scheduling on the chip cluster, where jobs are indefinitely stuck in the pending state with a reason of “(None)”. This boiled down to the SLURM scheduler being overwhelmed with over 9000 concurrent job submissions this weekend.
As a temporary work around, we have intervened to get all submitted jobs scheduled. We will be monitoring the queue to ensure future job submissions run smoothly. Finally, we are developing a more permanent fix to increase the job-handling capacity of SLURM, and to make it more resilient when handling many jobs.
We apologize for any inconvenience caused by this issue. To be clear, no jobs were lost as a result of the underlying issue or our intervention. However, some users may see that their jobs are in an “administratively held” state as we wrap up our intervention.
As always, please submit a descriptive support ticket via the link below if you have any questions or notice other issues: https://rtforms.umbc.edu/rt_authenticated/doit/DoIT-support.php?auto=Research%20Computing
Gregory Ballantine
HPC System Administrator
See also: Over 9000 in “popular” media (YouTube)