<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><link rel="alternate" type="text/html" href="https://um-grex.github.io/status/"/><title>NFS /Home on Status of the Grex HPC system</title><link>https://um-grex.github.io/status/affected/nfs-/home/</link><description>Incident history</description><generator>github.com/cstate</generator><language>en-us</language><lastBuildDate>2026-09-02T08:30:00+00:00</lastBuildDate><updated>2026-09-02T08:30:00+00:00</updated><copyright>The MIT License (MIT) Copyright © 2025 UM-Grex</copyright><atom:link href="https://um-grex.github.io/status/affected/nfs-/home/index.xml" rel="self" type="application/rss+xml"/><item><title>[Resolved] Planned Grex outage - updating SLURM and home</title><link>https://um-grex.github.io/status/issues/2026-09-20-slurm-storage-planned/</link><pubDate>Wed, 02 Sep 2026 08:30:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2026-09-20-slurm-storage-planned/</guid><category>2026-09-08 23:10:00</category><description>&lt;h4 id="sept-8-outage-complete"&gt;Sept 8, outage complete&lt;/h4&gt;
&lt;p&gt;All the /project migration jobs were completed. Grex is fully available.&lt;/p&gt;
&lt;h4 id="status-update-as-of-evening-of-sept-4"&gt;Status update as of evening of Sept 4&lt;/h4&gt;
&lt;p&gt;We have to extend the ongoing outage past the planned Sept.4 ending. The tentative ETA is September 8, 2026.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have successfully updated Grex’s SLURM scheduler to a current version to address a CVE.&lt;/li&gt;
&lt;li&gt;We have also migrated all user /home directories to the new storage server. This should be transparent for all users.&lt;/li&gt;
&lt;li&gt;We are still in the process of migration of /project directories to the new /project filesystem appliance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to some issues we have encountered during the storage, our data migration to the new /project filesystem took longer than the expected outage window. Thus we have to extend the Grex outage to September 8. Some functionality like OpenOnDemand and Nextcloud is not available until the end of the outage.&lt;/p&gt;</description><content type="html">&lt;h4 id="sept-8-outage-complete"&gt;Sept 8, outage complete&lt;/h4&gt;
&lt;p&gt;All the /project migration jobs were completed. Grex is fully available.&lt;/p&gt;
&lt;h4 id="status-update-as-of-evening-of-sept-4"&gt;Status update as of evening of Sept 4&lt;/h4&gt;
&lt;p&gt;We have to extend the ongoing outage past the planned Sept.4 ending. The tentative ETA is September 8, 2026.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have successfully updated Grex’s SLURM scheduler to a current version to address a CVE.&lt;/li&gt;
&lt;li&gt;We have also migrated all user /home directories to the new storage server. This should be transparent for all users.&lt;/li&gt;
&lt;li&gt;We are still in the process of migration of /project directories to the new /project filesystem appliance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to some issues we have encountered during the storage, our data migration to the new /project filesystem took longer than the expected outage window. Thus we have to extend the Grex outage to September 8. Some functionality like OpenOnDemand and Nextcloud is not available until the end of the outage.&lt;/p&gt;
&lt;p&gt;However, we are able to open Grex for SSH access now, to groups whose /project had been migrated, so that they could log in, access their data, run their jobs and thus test the system. Access to Grex is blocked for the groups whose data are still in the process of migration. As of now, about 95% of all projects were migrated. We apologize for the delay.&lt;/p&gt;
&lt;h4 id="a-planned-grex--outage"&gt;A planned Grex outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned, full outage of the Grex HPC system, starting at 8:30AM on Tuesday, September 2 2026
We expect the outage to last until end of the day on Friday, September 4.&lt;/p&gt;
&lt;p&gt;We will perform update of the SLURM controller, and NFS storage, as well as a minor Linux update on all the compute nodes.
During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage - updating storage</title><link>https://um-grex.github.io/status/issues/2025-12-08-planned-storage-outage/</link><pubDate>Mon, 08 Dec 2025 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-12-08-planned-storage-outage/</guid><category>2025-12-08 12:00:00</category><description>&lt;h4 id="planned-project-storage-outage-resolved"&gt;Planned /project storage outage resolved&lt;/h4&gt;
&lt;p&gt;The project outage is resolved.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-outage"&gt;A planned /project storage outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned outage of the Grex HPC system, starting at 8AM on Monday, December 8, 2025. We expect the outage to last one working day.&lt;/p&gt;
&lt;p&gt;We will perform hardware maintenance on our Lustre storage controller.&lt;/p&gt;
&lt;p&gt;During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</description><content type="html">&lt;h4 id="planned-project-storage-outage-resolved"&gt;Planned /project storage outage resolved&lt;/h4&gt;
&lt;p&gt;The project outage is resolved.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-outage"&gt;A planned /project storage outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned outage of the Grex HPC system, starting at 8AM on Monday, December 8, 2025. We expect the outage to last one working day.&lt;/p&gt;
&lt;p&gt;We will perform hardware maintenance on our Lustre storage controller.&lt;/p&gt;
&lt;p&gt;During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage, updating storage and patching operating system</title><link>https://um-grex.github.io/status/issues/2025-09-24-planned-updates-outage/</link><pubDate>Wed, 24 Sep 2025 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-09-24-planned-updates-outage/</guid><category>2025-09-26 17:50:00</category><description>&lt;h4 id="the-outage-completed"&gt;The outage completed&lt;/h4&gt;
&lt;p&gt;Grex is operating normally. Queued jobs were not affected and are now running. Thank you for your patience!
If you notice any issues or abnormalites, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-and-compute-outage"&gt;A planned /project storage and Compute outage&lt;/h4&gt;
&lt;p&gt;A /project storage outage to update Lustre filesystem software is in progress on Grex!
This requires a quiet filesystem so the system is drained from running jobs and access throgh login nodes and OpenOndemand is closed to users.
We will also use the outage to apply various security patches to the Linux OS on Grex , as well as performing other software updates.
ETA for the outage end is Friday , Sept 26. Thank you for your patience!&lt;/p&gt;</description><content type="html">&lt;h4 id="the-outage-completed"&gt;The outage completed&lt;/h4&gt;
&lt;p&gt;Grex is operating normally. Queued jobs were not affected and are now running. Thank you for your patience!
If you notice any issues or abnormalites, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-and-compute-outage"&gt;A planned /project storage and Compute outage&lt;/h4&gt;
&lt;p&gt;A /project storage outage to update Lustre filesystem software is in progress on Grex!
This requires a quiet filesystem so the system is drained from running jobs and access throgh login nodes and OpenOndemand is closed to users.
We will also use the outage to apply various security patches to the Linux OS on Grex , as well as performing other software updates.
ETA for the outage end is Friday , Sept 26. Thank you for your patience!&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned power outage in HPCC, Grex down</title><link>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</link><pubDate>Wed, 23 Apr 2025 09:10:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</guid><category>2025-04-23 12:40:00</category><description>&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</description><content type="html">&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-power-outage-happened-in-hpcc-centre"&gt;A power outage happened in HPCC Centre&lt;/h4&gt;
&lt;p&gt;A power outage in Grex&amp;rsquo;s datacentre happened , with a complete loss of power at about 9:10 AM Winnipeg time.
The system is down. The reason for the outage is a problem at Manitoba Hydro, our electricity provider.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://account.hydro.mb.ca/Portal/outeroutage.aspx"&gt;https://account.hydro.mb.ca/Portal/outeroutage.aspx&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We are waiting for the power to be restored. Thank you for your patience!&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned storage outage on Grex</title><link>https://um-grex.github.io/status/issues/2025-03-16-unplanned-home/</link><pubDate>Sun, 16 Mar 2025 10:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-03-16-unplanned-home/</guid><category>2025-03-16 17:10:00</category><description>&lt;h4 id="the-filesystem-came-back"&gt;The filesystem came back&lt;/h4&gt;
&lt;p&gt;As of now, it is working. We will investigate the cause and update the notice. If you see any problem with your data on /home , please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="failure-of-the-nfs-home-filesystem"&gt;Failure of the NFS /home filesystem&lt;/h4&gt;
&lt;p&gt;The NFS storage had failed around 10AM Sunday, March 16. The /home filesystem is currently unavailable.
Jobs, even those that are running from the /project filesystem, may lack access to the local software stack.
Thus we have placed a SLURM reservation to prefent new jobs from starting.&lt;/p&gt;</description><content type="html">&lt;h4 id="the-filesystem-came-back"&gt;The filesystem came back&lt;/h4&gt;
&lt;p&gt;As of now, it is working. We will investigate the cause and update the notice. If you see any problem with your data on /home , please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="failure-of-the-nfs-home-filesystem"&gt;Failure of the NFS /home filesystem&lt;/h4&gt;
&lt;p&gt;The NFS storage had failed around 10AM Sunday, March 16. The /home filesystem is currently unavailable.
Jobs, even those that are running from the /project filesystem, may lack access to the local software stack.
Thus we have placed a SLURM reservation to prefent new jobs from starting.&lt;/p&gt;
&lt;p&gt;We are working on resolving the issue. Thank you for your patience!&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC datacentre power outage</title><link>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</link><pubDate>Sun, 23 Feb 2025 07:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</guid><category>2025-02-24 17:10:00</category><description>&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</description><content type="html">&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC/Grex outage for electrical and cooling work.</title><link>https://um-grex.github.io/status/issues/2024-07-16-planned-hpcc-outage/</link><pubDate>Tue, 16 Jul 2024 06:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-07-16-planned-hpcc-outage/</guid><category>2024-07-18 18:00:00</category><description>&lt;h4 id="update-3-on-july-18"&gt;Update 3 on July 18&lt;/h4&gt;
&lt;p&gt;Cooling in HPCC is working, so the outage is over and Grex is operational.&lt;/p&gt;
&lt;h4 id="update-2-on-july-17"&gt;Update 2 on July 17&lt;/h4&gt;
&lt;p&gt;Unfortunately, due to heat outside of the datacentre, and the work inside the datacentre, we were unable to keep the environment cool enough to run even the storage and login nodes. So Grex is fully powered down again.&lt;/p&gt;
&lt;h4 id="update-1-on-july-17"&gt;Update 1 on July 17&lt;/h4&gt;
&lt;p&gt;The electrical part of the update is over. We have powered up Grex login nodes and storage, so the users can SSH in and access their data.
What does not yet work is compute, because water cooling is being worked on. So no jobs will get started, and OOD Web portal on zebu is also down.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-3-on-july-18"&gt;Update 3 on July 18&lt;/h4&gt;
&lt;p&gt;Cooling in HPCC is working, so the outage is over and Grex is operational.&lt;/p&gt;
&lt;h4 id="update-2-on-july-17"&gt;Update 2 on July 17&lt;/h4&gt;
&lt;p&gt;Unfortunately, due to heat outside of the datacentre, and the work inside the datacentre, we were unable to keep the environment cool enough to run even the storage and login nodes. So Grex is fully powered down again.&lt;/p&gt;
&lt;h4 id="update-1-on-july-17"&gt;Update 1 on July 17&lt;/h4&gt;
&lt;p&gt;The electrical part of the update is over. We have powered up Grex login nodes and storage, so the users can SSH in and access their data.
What does not yet work is compute, because water cooling is being worked on. So no jobs will get started, and OOD Web portal on zebu is also down.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="outage-started-on-july-16"&gt;Outage started on July 16&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now.
During this outage, Physical Plant will work on HPCC power and cooling, and the entire Grex system will be powered down.&lt;/p&gt;
&lt;p&gt;Users will not have access to any Grex services (compute and storage and Web portal) during the outage. During the outage, all running jobs will be terminated.&lt;/p&gt;
&lt;p&gt;Should you have any questions about the upcoming Grex outage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; ! Thank you for your patience,&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for Lustre FS and Major Linux update</title><link>https://um-grex.github.io/status/issues/2024-05-07-planned-lustre-outage/</link><pubDate>Tue, 07 May 2024 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-05-07-planned-lustre-outage/</guid><category>2024-05-10 13:00:00</category><description>&lt;h4 id="final-update-on-may-10"&gt;Final update on May 10&lt;/h4&gt;
&lt;p&gt;The OS and Lustre update outage is over!&lt;/p&gt;
&lt;p&gt;The OS on Grex was upgraded to Alma Linux on all CPU, GPU and login nodes {except for bison, tatanka and compute partition}. Now, Grex is available for running jobs. Please note the changes about software stack:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="update-on-may-9"&gt;Update on May 9&lt;/h4&gt;
&lt;p&gt;The outage extended into May 10. The /home filesystem had been updated; update of the Linux OS and HPC software is still in progress.
Sorry about the inconvenience and thank you for your patience!&lt;/p&gt;</description><content type="html">&lt;h4 id="final-update-on-may-10"&gt;Final update on May 10&lt;/h4&gt;
&lt;p&gt;The OS and Lustre update outage is over!&lt;/p&gt;
&lt;p&gt;The OS on Grex was upgraded to Alma Linux on all CPU, GPU and login nodes {except for bison, tatanka and compute partition}. Now, Grex is available for running jobs. Please note the changes about software stack:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="update-on-may-9"&gt;Update on May 9&lt;/h4&gt;
&lt;p&gt;The outage extended into May 10. The /home filesystem had been updated; update of the Linux OS and HPC software is still in progress.
Sorry about the inconvenience and thank you for your patience!&lt;/p&gt;
&lt;h4 id="update-on-may-8"&gt;Update on May 8&lt;/h4&gt;
&lt;p&gt;The outage continues. The /project filesystem appliance had been updated.
We are working on updating the /home filesystem, Linux OS and HPC software stacks. The system is still closed for users.&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-may-7---9-2024"&gt;Grex system outage on May 7 - 9 2024&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now .
We are performing a major update of the DDN Lustre storage controller. We will also do a major Linux OS update from CentOS 7 to AlmaLinux 8.
We will reboot and reinstall all of Grex compute and login nodes. All running and queued jobs will be deleted.
Grex login nodes , compute and storage will be unavailable during the outage window. Thank you for your patience!
Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for SLURM and minor OS update</title><link>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</link><pubDate>Mon, 18 Dec 2023 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</guid><category>2023-12-19 18:00:00</category><description>&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;</description><content type="html">&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;
&lt;p&gt;We will eboot and reinstall all of Grex compute and login nodes and to migrate the SLURM job database.
Thus Jobs that are still running by the time of the outage will be lost, and Grex login nodes will be unavailable during the outage window.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Two Planned Grex outages for HPCC transformer work</title><link>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</link><pubDate>Tue, 12 Sep 2023 18:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</guid><category>2023-09-19 21:00:00</category><description>&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;</description><content type="html">&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;
&lt;p&gt;To avoid the failure of jobs, we have made two reservations to avoid any longer jobs (that cannot be finished before the beginning of the outage) from starting. For more information, run the following command from any login node: &amp;ldquo;scontrol show res&amp;rdquo;. To take advantage of the cluster before, and between, the outages, we recommend users submit short jobs that can finish by the time the upcoming outage begins.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Planned power outage Aug 31 to Sept 1</title><link>https://um-grex.github.io/status/issues/2021-08-31-planed-power-outage/</link><pubDate>Tue, 31 Aug 2021 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2021-08-31-planed-power-outage/</guid><category>2021-09-01 12:00:00</category><description>&lt;h2 id="grex-is-back-online"&gt;Grex is back online&lt;/h2&gt;
&lt;p&gt;Grex is back online, accepting CPU jobs. GPU nodes will take a bit more to reinstall NVIDIA updates, but will be online by end of today.&lt;/p&gt;
&lt;p&gt;The previous jobs on the queue were lost since the outage afftected all
compute nodes that are rebooted after restoring the power. Please check
your data and re-submit the jobs that are not done before the outage.&lt;/p&gt;
&lt;h2 id="grex-will-be-down-for-a-campus-power-maintenance"&gt;Grex will be down for a campus power maintenance&lt;/h2&gt;
&lt;p&gt;Physical Plant has a planning a power outage for power breakers maintenance. This will affect Grex building, so and all the nodes and strage systems will be shut-down. The system will be unavailable for users, and all running jobs will be lost. We will be performing some work on Grex during the maintenance window as well. ETA for Grex availabinity is noon, September 1.&lt;/p&gt;</description><content type="html">&lt;h2 id="grex-is-back-online"&gt;Grex is back online&lt;/h2&gt;
&lt;p&gt;Grex is back online, accepting CPU jobs. GPU nodes will take a bit more to reinstall NVIDIA updates, but will be online by end of today.&lt;/p&gt;
&lt;p&gt;The previous jobs on the queue were lost since the outage afftected all
compute nodes that are rebooted after restoring the power. Please check
your data and re-submit the jobs that are not done before the outage.&lt;/p&gt;
&lt;h2 id="grex-will-be-down-for-a-campus-power-maintenance"&gt;Grex will be down for a campus power maintenance&lt;/h2&gt;
&lt;p&gt;Physical Plant has a planning a power outage for power breakers maintenance. This will affect Grex building, so and all the nodes and strage systems will be shut-down. The system will be unavailable for users, and all running jobs will be lost. We will be performing some work on Grex during the maintenance window as well. ETA for Grex availabinity is noon, September 1.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item></channel></rss>