<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><link rel="alternate" type="text/html" href="https://um-grex.github.io/status/"/><title>Lustre /Project on Status of the Grex HPC system</title><link>https://um-grex.github.io/status/affected/lustre-/project/</link><description>Incident history</description><generator>github.com/cstate</generator><language>en-us</language><lastBuildDate>2026-09-02T08:30:00+00:00</lastBuildDate><updated>2026-09-02T08:30:00+00:00</updated><copyright>The MIT License (MIT) Copyright © 2025 UM-Grex</copyright><atom:link href="https://um-grex.github.io/status/affected/lustre-/project/index.xml" rel="self" type="application/rss+xml"/><item><title>[Resolved] Planned Grex outage - updating SLURM and home</title><link>https://um-grex.github.io/status/issues/2026-09-20-slurm-storage-planned/</link><pubDate>Wed, 02 Sep 2026 08:30:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2026-09-20-slurm-storage-planned/</guid><category>2026-09-08 23:10:00</category><description>&lt;h4 id="sept-8-outage-complete"&gt;Sept 8, outage complete&lt;/h4&gt;
&lt;p&gt;All the /project migration jobs were completed. Grex is fully available.&lt;/p&gt;
&lt;h4 id="status-update-as-of-evening-of-sept-4"&gt;Status update as of evening of Sept 4&lt;/h4&gt;
&lt;p&gt;We have to extend the ongoing outage past the planned Sept.4 ending. The tentative ETA is September 8, 2026.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have successfully updated Grex’s SLURM scheduler to a current version to address a CVE.&lt;/li&gt;
&lt;li&gt;We have also migrated all user /home directories to the new storage server. This should be transparent for all users.&lt;/li&gt;
&lt;li&gt;We are still in the process of migration of /project directories to the new /project filesystem appliance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to some issues we have encountered during the storage, our data migration to the new /project filesystem took longer than the expected outage window. Thus we have to extend the Grex outage to September 8. Some functionality like OpenOnDemand and Nextcloud is not available until the end of the outage.&lt;/p&gt;</description><content type="html">&lt;h4 id="sept-8-outage-complete"&gt;Sept 8, outage complete&lt;/h4&gt;
&lt;p&gt;All the /project migration jobs were completed. Grex is fully available.&lt;/p&gt;
&lt;h4 id="status-update-as-of-evening-of-sept-4"&gt;Status update as of evening of Sept 4&lt;/h4&gt;
&lt;p&gt;We have to extend the ongoing outage past the planned Sept.4 ending. The tentative ETA is September 8, 2026.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We have successfully updated Grex’s SLURM scheduler to a current version to address a CVE.&lt;/li&gt;
&lt;li&gt;We have also migrated all user /home directories to the new storage server. This should be transparent for all users.&lt;/li&gt;
&lt;li&gt;We are still in the process of migration of /project directories to the new /project filesystem appliance.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Due to some issues we have encountered during the storage, our data migration to the new /project filesystem took longer than the expected outage window. Thus we have to extend the Grex outage to September 8. Some functionality like OpenOnDemand and Nextcloud is not available until the end of the outage.&lt;/p&gt;
&lt;p&gt;However, we are able to open Grex for SSH access now, to groups whose /project had been migrated, so that they could log in, access their data, run their jobs and thus test the system. Access to Grex is blocked for the groups whose data are still in the process of migration. As of now, about 95% of all projects were migrated. We apologize for the delay.&lt;/p&gt;
&lt;h4 id="a-planned-grex--outage"&gt;A planned Grex outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned, full outage of the Grex HPC system, starting at 8:30AM on Tuesday, September 2 2026
We expect the outage to last until end of the day on Friday, September 4.&lt;/p&gt;
&lt;p&gt;We will perform update of the SLURM controller, and NFS storage, as well as a minor Linux update on all the compute nodes.
During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage - updating storage</title><link>https://um-grex.github.io/status/issues/2025-12-08-planned-storage-outage/</link><pubDate>Mon, 08 Dec 2025 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-12-08-planned-storage-outage/</guid><category>2025-12-08 12:00:00</category><description>&lt;h4 id="planned-project-storage-outage-resolved"&gt;Planned /project storage outage resolved&lt;/h4&gt;
&lt;p&gt;The project outage is resolved.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-outage"&gt;A planned /project storage outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned outage of the Grex HPC system, starting at 8AM on Monday, December 8, 2025. We expect the outage to last one working day.&lt;/p&gt;
&lt;p&gt;We will perform hardware maintenance on our Lustre storage controller.&lt;/p&gt;
&lt;p&gt;During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</description><content type="html">&lt;h4 id="planned-project-storage-outage-resolved"&gt;Planned /project storage outage resolved&lt;/h4&gt;
&lt;p&gt;The project outage is resolved.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-outage"&gt;A planned /project storage outage&lt;/h4&gt;
&lt;p&gt;We will be performing a planned outage of the Grex HPC system, starting at 8AM on Monday, December 8, 2025. We expect the outage to last one working day.&lt;/p&gt;
&lt;p&gt;We will perform hardware maintenance on our Lustre storage controller.&lt;/p&gt;
&lt;p&gt;During the outage, OpenOnDemand, Storage, and Login nodes will not be available, and all running jobs will be stopped. The outage will not affect the users’ data.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage, updating storage and patching operating system</title><link>https://um-grex.github.io/status/issues/2025-09-24-planned-updates-outage/</link><pubDate>Wed, 24 Sep 2025 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-09-24-planned-updates-outage/</guid><category>2025-09-26 17:50:00</category><description>&lt;h4 id="the-outage-completed"&gt;The outage completed&lt;/h4&gt;
&lt;p&gt;Grex is operating normally. Queued jobs were not affected and are now running. Thank you for your patience!
If you notice any issues or abnormalites, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-and-compute-outage"&gt;A planned /project storage and Compute outage&lt;/h4&gt;
&lt;p&gt;A /project storage outage to update Lustre filesystem software is in progress on Grex!
This requires a quiet filesystem so the system is drained from running jobs and access throgh login nodes and OpenOndemand is closed to users.
We will also use the outage to apply various security patches to the Linux OS on Grex , as well as performing other software updates.
ETA for the outage end is Friday , Sept 26. Thank you for your patience!&lt;/p&gt;</description><content type="html">&lt;h4 id="the-outage-completed"&gt;The outage completed&lt;/h4&gt;
&lt;p&gt;Grex is operating normally. Queued jobs were not affected and are now running. Thank you for your patience!
If you notice any issues or abnormalites, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-planned-project-storage-and-compute-outage"&gt;A planned /project storage and Compute outage&lt;/h4&gt;
&lt;p&gt;A /project storage outage to update Lustre filesystem software is in progress on Grex!
This requires a quiet filesystem so the system is drained from running jobs and access throgh login nodes and OpenOndemand is closed to users.
We will also use the outage to apply various security patches to the Linux OS on Grex , as well as performing other software updates.
ETA for the outage end is Friday , Sept 26. Thank you for your patience!&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned project storage outage in HPCC</title><link>https://um-grex.github.io/status/issues/2025-07-18-storage-outage/</link><pubDate>Fri, 18 Jul 2025 07:10:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-07-18-storage-outage/</guid><category>2025-07-18 17:30:00</category><description>&lt;h4 id="update-1730-central-time-system-is-available"&gt;Update 17:30 Central Time: System is available&lt;/h4&gt;
&lt;p&gt;The storage is working. If you notice any issues with the storage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="update-1700-central-time-system-is-available-with-conditions"&gt;Update 17:00 Central Time: System is available with conditions&lt;/h4&gt;
&lt;p&gt;The failure was traced to a hardware issue on one of the storage controllers.
We have cleared the controller state and restarted the /project storage appliance.
However, the storage targets are not yet properly balanced across the storage servers, so there can be some performance issues. We are working with the storage vendor to address these issues.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-1730-central-time-system-is-available"&gt;Update 17:30 Central Time: System is available&lt;/h4&gt;
&lt;p&gt;The storage is working. If you notice any issues with the storage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="update-1700-central-time-system-is-available-with-conditions"&gt;Update 17:00 Central Time: System is available with conditions&lt;/h4&gt;
&lt;p&gt;The failure was traced to a hardware issue on one of the storage controllers.
We have cleared the controller state and restarted the /project storage appliance.
However, the storage targets are not yet properly balanced across the storage servers, so there can be some performance issues. We are working with the storage vendor to address these issues.&lt;/p&gt;
&lt;p&gt;As of now, OpenOnDemand and /project are now available, and Grex system would accept new jobs.&lt;/p&gt;
&lt;h4 id="a-project-storage-outage-happened-on-grex"&gt;A /project storage outage happened on Grex&lt;/h4&gt;
&lt;p&gt;A /project storage outage happened on Grex this morning! The /project filesystem became unavailable.
This affectes running jobs and access to the OpenOnDemand portal as well. We are investigating the issue.&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned power outage in HPCC, Grex down</title><link>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</link><pubDate>Wed, 23 Apr 2025 09:10:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</guid><category>2025-04-23 12:40:00</category><description>&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</description><content type="html">&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-power-outage-happened-in-hpcc-centre"&gt;A power outage happened in HPCC Centre&lt;/h4&gt;
&lt;p&gt;A power outage in Grex&amp;rsquo;s datacentre happened , with a complete loss of power at about 9:10 AM Winnipeg time.
The system is down. The reason for the outage is a problem at Manitoba Hydro, our electricity provider.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://account.hydro.mb.ca/Portal/outeroutage.aspx"&gt;https://account.hydro.mb.ca/Portal/outeroutage.aspx&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We are waiting for the power to be restored. Thank you for your patience!&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC datacentre power outage</title><link>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</link><pubDate>Sun, 23 Feb 2025 07:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</guid><category>2025-02-24 17:10:00</category><description>&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</description><content type="html">&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC/Grex outage for electrical and cooling work.</title><link>https://um-grex.github.io/status/issues/2024-08-26-planned-hpcc-outage/</link><pubDate>Mon, 26 Aug 2024 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-08-26-planned-hpcc-outage/</guid><category>2024-09-10 16:00:00</category><description>&lt;h4 id="update-sept--10"&gt;Update Sept 10&lt;/h4&gt;
&lt;p&gt;The outage is over. Grex is fully online and available to users.
There are many important changes made on the Grex system. Please check them out at:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;
&lt;h4 id="update-sept--6"&gt;Update Sept 6&lt;/h4&gt;
&lt;p&gt;Due to a delay with deployment of the new water cooling system, Grex&amp;rsquo;s outage is extended until Wednesday, Sept. 11.
At this point, the cooling for new row of racks cannot be fully enabled. Thus, the partial availability of Grex continues.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-sept--10"&gt;Update Sept 10&lt;/h4&gt;
&lt;p&gt;The outage is over. Grex is fully online and available to users.
There are many important changes made on the Grex system. Please check them out at:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;
&lt;h4 id="update-sept--6"&gt;Update Sept 6&lt;/h4&gt;
&lt;p&gt;Due to a delay with deployment of the new water cooling system, Grex&amp;rsquo;s outage is extended until Wednesday, Sept. 11.
At this point, the cooling for new row of racks cannot be fully enabled. Thus, the partial availability of Grex continues.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)&lt;/li&gt;
&lt;li&gt;Home and Project file systems are online.&lt;/li&gt;
&lt;li&gt;OpenOnDemand portal (&lt;a href="https://zebu.hpc.umanitoba.ca"&gt;https://zebu.hpc.umanitoba.ca&lt;/a&gt;, Simplified Desktop) is online&lt;/li&gt;
&lt;li&gt;Running jobs of short duration (must end before September 9, 2024) on skylake and GPU partitions would work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thank you for your patience!&lt;/p&gt;
&lt;h4 id="update-aug-30"&gt;Update Aug 30&lt;/h4&gt;
&lt;p&gt;We have completed the migration of all of the storage systems, and most of the compute servers into the new datacentre racks.
However, the cooling system installation and acceptance is due next week, so the Grex system is not yet fully online.&lt;/p&gt;
&lt;p&gt;During the long weekend, users have access to the following Grex services or systems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)&lt;/li&gt;
&lt;li&gt;Home and Project file systems are online.&lt;/li&gt;
&lt;li&gt;OpenOnDemand portal (&lt;a href="https://zebu.hpc.umanitoba.ca"&gt;https://zebu.hpc.umanitoba.ca&lt;/a&gt;, Simplified Desktop) is online&lt;/li&gt;
&lt;li&gt;Running jobs of short duration (must end before September 3, 2024) on skylake and some of the GPU partitions would work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following systems or services are as of now offline and unavailable: Old login nodes tatanka and bison are decommissioned and unavailable. grex.hpc.umanitoba.ca is now a yak alias. Old compute partition is decommissioned and unavailable. Most new GPU and CPU partitions are offline because the cooling system is yet to be completed in HPCC.&lt;/p&gt;
&lt;h4 id="update-as-of-aug-28"&gt;Update as of Aug 28&lt;/h4&gt;
&lt;p&gt;The First phase: Aug 26 - Aug 28, 2024 is done. We have migrated our storage, login and management nodes to the final location.
Grex is now partially open for users with limitted services:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt; - Use the login nodes and OOD portal
- Access to storage {home and project} if you need to access your data.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Please note that users can not yet submit jobs as the migration of the compute nodes is not done yet, pending completion of the new cooling systems. We may also experience intermittent interruptions with access to the storage and the login nodes as we are continue with the outage.&lt;/p&gt;
&lt;h4 id="outage-started-on-aug-26"&gt;Outage started on Aug 26&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now.&lt;/p&gt;
&lt;p&gt;During this outage, Physical Plant will work on HPCC power and cooling, and the entire Grex system will be powered down. Then, the system will be migrated to our new water cooled rack infrastructure.&lt;/p&gt;
&lt;p&gt;Users will not have access to any Grex services (compute, storage and the OOD Web portal) during the fist stage of the outage that is expected to last at least three days (until Aug 29).&lt;/p&gt;
&lt;p&gt;We will be updating this page as the work in HPCC progresses.&lt;/p&gt;
&lt;p&gt;Should you have any questions about the upcoming Grex outage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; ! Thank you for your patience,&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;</content></item></channel></rss>