August 25th chip and RRStor Downtime Completed
To all chip users,
The downtime for the chip cluster scheduled for August 25th has been successfully completed, and the cluster is now back online and accessible. We appreciate your patience while we installed security patches and performed upgrades on RRStor.
Note that some cluster nodes are still draining for a reboot. Any nodes marked as “draining” or “drng” will reboot once their running jobs have finished.
Notable changes
In order to reduce confusion regarding job terminations upon preemption, there will now be a notification message written to the Slurm error file of a preempted job. If a job is preempted, the associated Slurm error file will see a message like “[SYSTEM NOTIFICATION] Job XYZ was PREEMPTED by Slurm.” along with a timestamp and other information.
We will add more reports such as this that will help users figure out what resources were used by their jobs.
--
If you encounter any issues, please submit a support ticket here: https://rtforms.umbc.edu/rt_authenticated/doit/DoIT-support.php?auto=Research%20Computing
For additional information, you can check our documentation at: https://umbc.atlassian.net/wiki/spaces/faq/pages/1082589207/UMBC+HPCF+-+chip
Thank you,
Gregory Ballantine
HPC Specialist