<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><link rel="alternate" type="text/html" href="https://um-grex.github.io/status/"/><title>Lustre /Global/Scratch on Status of the Grex HPC system</title><link>https://um-grex.github.io/status/affected/lustre-/global/scratch/</link><description>Incident history</description><generator>github.com/cstate</generator><language>en-us</language><lastBuildDate>2024-07-16T06:00:00+00:00</lastBuildDate><updated>2024-07-16T06:00:00+00:00</updated><copyright>The MIT License (MIT) Copyright © 2025 UM-Grex</copyright><atom:link href="https://um-grex.github.io/status/affected/lustre-/global/scratch/index.xml" rel="self" type="application/rss+xml"/><item><title>[Resolved] Planned HPCC/Grex outage for electrical and cooling work.</title><link>https://um-grex.github.io/status/issues/2024-07-16-planned-hpcc-outage/</link><pubDate>Tue, 16 Jul 2024 06:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-07-16-planned-hpcc-outage/</guid><category>2024-07-18 18:00:00</category><description>&lt;h4 id="update-3-on-july-18"&gt;Update 3 on July 18&lt;/h4&gt;
&lt;p&gt;Cooling in HPCC is working, so the outage is over and Grex is operational.&lt;/p&gt;
&lt;h4 id="update-2-on-july-17"&gt;Update 2 on July 17&lt;/h4&gt;
&lt;p&gt;Unfortunately, due to heat outside of the datacentre, and the work inside the datacentre, we were unable to keep the environment cool enough to run even the storage and login nodes. So Grex is fully powered down again.&lt;/p&gt;
&lt;h4 id="update-1-on-july-17"&gt;Update 1 on July 17&lt;/h4&gt;
&lt;p&gt;The electrical part of the update is over. We have powered up Grex login nodes and storage, so the users can SSH in and access their data.
What does not yet work is compute, because water cooling is being worked on. So no jobs will get started, and OOD Web portal on zebu is also down.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-3-on-july-18"&gt;Update 3 on July 18&lt;/h4&gt;
&lt;p&gt;Cooling in HPCC is working, so the outage is over and Grex is operational.&lt;/p&gt;
&lt;h4 id="update-2-on-july-17"&gt;Update 2 on July 17&lt;/h4&gt;
&lt;p&gt;Unfortunately, due to heat outside of the datacentre, and the work inside the datacentre, we were unable to keep the environment cool enough to run even the storage and login nodes. So Grex is fully powered down again.&lt;/p&gt;
&lt;h4 id="update-1-on-july-17"&gt;Update 1 on July 17&lt;/h4&gt;
&lt;p&gt;The electrical part of the update is over. We have powered up Grex login nodes and storage, so the users can SSH in and access their data.
What does not yet work is compute, because water cooling is being worked on. So no jobs will get started, and OOD Web portal on zebu is also down.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="outage-started-on-july-16"&gt;Outage started on July 16&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now.
During this outage, Physical Plant will work on HPCC power and cooling, and the entire Grex system will be powered down.&lt;/p&gt;
&lt;p&gt;Users will not have access to any Grex services (compute and storage and Web portal) during the outage. During the outage, all running jobs will be terminated.&lt;/p&gt;
&lt;p&gt;Should you have any questions about the upcoming Grex outage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; ! Thank you for your patience,&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for Lustre FS and Major Linux update</title><link>https://um-grex.github.io/status/issues/2024-05-07-planned-lustre-outage/</link><pubDate>Tue, 07 May 2024 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-05-07-planned-lustre-outage/</guid><category>2024-05-10 13:00:00</category><description>&lt;h4 id="final-update-on-may-10"&gt;Final update on May 10&lt;/h4&gt;
&lt;p&gt;The OS and Lustre update outage is over!&lt;/p&gt;
&lt;p&gt;The OS on Grex was upgraded to Alma Linux on all CPU, GPU and login nodes {except for bison, tatanka and compute partition}. Now, Grex is available for running jobs. Please note the changes about software stack:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="update-on-may-9"&gt;Update on May 9&lt;/h4&gt;
&lt;p&gt;The outage extended into May 10. The /home filesystem had been updated; update of the Linux OS and HPC software is still in progress.
Sorry about the inconvenience and thank you for your patience!&lt;/p&gt;</description><content type="html">&lt;h4 id="final-update-on-may-10"&gt;Final update on May 10&lt;/h4&gt;
&lt;p&gt;The OS and Lustre update outage is over!&lt;/p&gt;
&lt;p&gt;The OS on Grex was upgraded to Alma Linux on all CPU, GPU and login nodes {except for bison, tatanka and compute partition}. Now, Grex is available for running jobs. Please note the changes about software stack:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;ul&gt;
&lt;li&gt;If you have questions or concerns, please don&amp;rsquo;t hesitate to contact us at: support@tech.alliancecan.ca&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="update-on-may-9"&gt;Update on May 9&lt;/h4&gt;
&lt;p&gt;The outage extended into May 10. The /home filesystem had been updated; update of the Linux OS and HPC software is still in progress.
Sorry about the inconvenience and thank you for your patience!&lt;/p&gt;
&lt;h4 id="update-on-may-8"&gt;Update on May 8&lt;/h4&gt;
&lt;p&gt;The outage continues. The /project filesystem appliance had been updated.
We are working on updating the /home filesystem, Linux OS and HPC software stacks. The system is still closed for users.&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-may-7---9-2024"&gt;Grex system outage on May 7 - 9 2024&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now .
We are performing a major update of the DDN Lustre storage controller. We will also do a major Linux OS update from CentOS 7 to AlmaLinux 8.
We will reboot and reinstall all of Grex compute and login nodes. All running and queued jobs will be deleted.
Grex login nodes , compute and storage will be unavailable during the outage window. Thank you for your patience!
Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for SLURM and minor OS update</title><link>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</link><pubDate>Mon, 18 Dec 2023 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</guid><category>2023-12-19 18:00:00</category><description>&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;</description><content type="html">&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;
&lt;p&gt;We will eboot and reinstall all of Grex compute and login nodes and to migrate the SLURM job database.
Thus Jobs that are still running by the time of the outage will be lost, and Grex login nodes will be unavailable during the outage window.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Two Planned Grex outages for HPCC transformer work</title><link>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</link><pubDate>Tue, 12 Sep 2023 18:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</guid><category>2023-09-19 21:00:00</category><description>&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;</description><content type="html">&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;
&lt;p&gt;To avoid the failure of jobs, we have made two reservations to avoid any longer jobs (that cannot be finished before the beginning of the outage) from starting. For more information, run the following command from any login node: &amp;ldquo;scontrol show res&amp;rdquo;. To take advantage of the cluster before, and between, the outages, we recommend users submit short jobs that can finish by the time the upcoming outage begins.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned power outage in HPCC datacentre</title><link>https://um-grex.github.io/status/issues/2023-05-30-unplanned-power-outage/</link><pubDate>Tue, 30 May 2023 07:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-05-30-unplanned-power-outage/</guid><category>2023-05-30 09:30:00</category><description>&lt;h4 id="update-outage-concluded"&gt;UPDATE: outage concluded&lt;/h4&gt;
&lt;p&gt;Grex compute and storage are back up after 1 1/2 hour of downtime. Please restart
your jobs! If you notice any further malfuncions, please do not hesitate
to contact us at the &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning &amp;ldquo;Grex&amp;rdquo; in the subject line.&lt;/p&gt;
&lt;h4 id="unplanned-power-loss-at-hpcc-datacentre"&gt;Unplanned power loss at HPCC datacentre&lt;/h4&gt;
&lt;p&gt;NOTICE: On May 30. 7AM to 8:14AM there was an unplanned power outage in HPCC
that rebooted all the Grex compute nodes, storage, and management systems.
Running jobs were lost. We apologize for the inconvenience. We work on restarting
and testing of the system.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-outage-concluded"&gt;UPDATE: outage concluded&lt;/h4&gt;
&lt;p&gt;Grex compute and storage are back up after 1 1/2 hour of downtime. Please restart
your jobs! If you notice any further malfuncions, please do not hesitate
to contact us at the &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning &amp;ldquo;Grex&amp;rdquo; in the subject line.&lt;/p&gt;
&lt;h4 id="unplanned-power-loss-at-hpcc-datacentre"&gt;Unplanned power loss at HPCC datacentre&lt;/h4&gt;
&lt;p&gt;NOTICE: On May 30. 7AM to 8:14AM there was an unplanned power outage in HPCC
that rebooted all the Grex compute nodes, storage, and management systems.
Running jobs were lost. We apologize for the inconvenience. We work on restarting
and testing of the system.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for HPCC work and storage update</title><link>https://um-grex.github.io/status/issues/2023-02-22-planned-grex-outage/</link><pubDate>Wed, 22 Feb 2023 14:30:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-02-22-planned-grex-outage/</guid><category>2023-02-27 3:00:00</category><description>&lt;h4 id="update-outage-concluded"&gt;UPDATE: outage concluded&lt;/h4&gt;
&lt;p&gt;During the outage of Feb 22, few changes have been made on Grex:
We reinstalled all login and compute nodes with a new image to fix the errors with UCX.
We have added a new storage “project” with similar structure as for Compute Canada.
All the data from scratch has been moved to the project file system.&lt;/p&gt;
&lt;p&gt;For an overview of the changes, please have a look to the documentation page:&lt;/p&gt;</description><content type="html">&lt;h4 id="update-outage-concluded"&gt;UPDATE: outage concluded&lt;/h4&gt;
&lt;p&gt;During the outage of Feb 22, few changes have been made on Grex:
We reinstalled all login and compute nodes with a new image to fix the errors with UCX.
We have added a new storage “project” with similar structure as for Compute Canada.
All the data from scratch has been moved to the project file system.&lt;/p&gt;
&lt;p&gt;For an overview of the changes, please have a look to the documentation page:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/docs/lustre/"&gt;https://um-grex.github.io/grex-docs/docs/lustre/&lt;/a&gt;&lt;/p&gt;
&lt;h4 id="planned-grex-outage"&gt;Planned Grex outage&lt;/h4&gt;
&lt;p&gt;The Grex storage outage is about to start. Please save your interactive work! ETA for the outage&amp;rsquo;s end is Feb 23 or Friday, Feb 24, 2023.&lt;/p&gt;
&lt;p&gt;During the outage we are planning to reboot and reinstall all of Grex compute and login nodes, connect the new Project storage, and perform some re-cabling of the Infiniband fabric of the cluster.
Jobs that are still running by the time of the outage will be lost, and Grex login nodes will be unavailable during the outage window.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! If you have questions or concerns regarding the outage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;mailto:support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Electric failure, loss of power to management rack</title><link>https://um-grex.github.io/status/issues/2023-01-26-electric-failure-datacentre/</link><pubDate>Thu, 26 Jan 2023 14:50:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-01-26-electric-failure-datacentre/</guid><category>2023-01-26 18:00:00</category><description>&lt;h4 id="update-datacentre-change-rolled-back-systems-operational"&gt;Update: datacentre change rolled back, systems operational&lt;/h4&gt;
&lt;p&gt;It appears that some of the running jobs continued running , and login nodes&amp;rsquo; access nor storage systems were affected.
Grex is now operational. In case you notice any ongoing issue, please let us know by email to &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, Subject line containing Grex.&lt;/p&gt;
&lt;h4 id="faulty-electrical-work-loss-of-power-to-management-rack"&gt;Faulty electrical work, loss of power to management rack&lt;/h4&gt;
&lt;p&gt;We have an unplanned outage on Grex due to a failed electrical work that affected its management rack at around 2:50 PM, Jan 26, 2023 .&lt;/p&gt;</description><content type="html">&lt;h4 id="update-datacentre-change-rolled-back-systems-operational"&gt;Update: datacentre change rolled back, systems operational&lt;/h4&gt;
&lt;p&gt;It appears that some of the running jobs continued running , and login nodes&amp;rsquo; access nor storage systems were affected.
Grex is now operational. In case you notice any ongoing issue, please let us know by email to &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, Subject line containing Grex.&lt;/p&gt;
&lt;h4 id="faulty-electrical-work-loss-of-power-to-management-rack"&gt;Faulty electrical work, loss of power to management rack&lt;/h4&gt;
&lt;p&gt;We have an unplanned outage on Grex due to a failed electrical work that affected its management rack at around 2:50 PM, Jan 26, 2023 .&lt;/p&gt;
&lt;p&gt;Running jobs may be lost, and access to Grex login nodes may be degraded or unavailable.&lt;/p&gt;
&lt;p&gt;We are working on resolving the issue. Sorry about the inconvenience it may have caused.&lt;/p&gt;</content></item><item><title>[Resolved] One of the /global/scratch filesystem servers crashed, FS runs in degraded mode</title><link>https://um-grex.github.io/status/issues/2022-05-13-filesystem-issues/</link><pubDate>Fri, 13 May 2022 19:30:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2022-05-13-filesystem-issues/</guid><category>2022-05-13 22:00:00</category><description>&lt;h2 id="lustre-filesystem-problems"&gt;Lustre filesystem problems&lt;/h2&gt;
&lt;p&gt;One of the /global/scratch filesystem servers crashed, FS runs in degraded mode. We are investigating the issue.
Most running jobs will continiue to run, albeit slower, and may stuck in I/O operations while Lustre servers are unresponsive.&lt;/p&gt;
&lt;h2 id="update"&gt;Update:&lt;/h2&gt;
&lt;p&gt;After restarting of Lustre servers, the /global/scratch filesystem is back to normal operation.&lt;/p&gt;</description><content type="html">&lt;h2 id="lustre-filesystem-problems"&gt;Lustre filesystem problems&lt;/h2&gt;
&lt;p&gt;One of the /global/scratch filesystem servers crashed, FS runs in degraded mode. We are investigating the issue.
Most running jobs will continiue to run, albeit slower, and may stuck in I/O operations while Lustre servers are unresponsive.&lt;/p&gt;
&lt;h2 id="update"&gt;Update:&lt;/h2&gt;
&lt;p&gt;After restarting of Lustre servers, the /global/scratch filesystem is back to normal operation.&lt;/p&gt;</content></item><item><title>[Resolved] Planned power outage Aug 31 to Sept 1</title><link>https://um-grex.github.io/status/issues/2021-08-31-planed-power-outage/</link><pubDate>Tue, 31 Aug 2021 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2021-08-31-planed-power-outage/</guid><category>2021-09-01 12:00:00</category><description>&lt;h2 id="grex-is-back-online"&gt;Grex is back online&lt;/h2&gt;
&lt;p&gt;Grex is back online, accepting CPU jobs. GPU nodes will take a bit more to reinstall NVIDIA updates, but will be online by end of today.&lt;/p&gt;
&lt;p&gt;The previous jobs on the queue were lost since the outage afftected all
compute nodes that are rebooted after restoring the power. Please check
your data and re-submit the jobs that are not done before the outage.&lt;/p&gt;
&lt;h2 id="grex-will-be-down-for-a-campus-power-maintenance"&gt;Grex will be down for a campus power maintenance&lt;/h2&gt;
&lt;p&gt;Physical Plant has a planning a power outage for power breakers maintenance. This will affect Grex building, so and all the nodes and strage systems will be shut-down. The system will be unavailable for users, and all running jobs will be lost. We will be performing some work on Grex during the maintenance window as well. ETA for Grex availabinity is noon, September 1.&lt;/p&gt;</description><content type="html">&lt;h2 id="grex-is-back-online"&gt;Grex is back online&lt;/h2&gt;
&lt;p&gt;Grex is back online, accepting CPU jobs. GPU nodes will take a bit more to reinstall NVIDIA updates, but will be online by end of today.&lt;/p&gt;
&lt;p&gt;The previous jobs on the queue were lost since the outage afftected all
compute nodes that are rebooted after restoring the power. Please check
your data and re-submit the jobs that are not done before the outage.&lt;/p&gt;
&lt;h2 id="grex-will-be-down-for-a-campus-power-maintenance"&gt;Grex will be down for a campus power maintenance&lt;/h2&gt;
&lt;p&gt;Physical Plant has a planning a power outage for power breakers maintenance. This will affect Grex building, so and all the nodes and strage systems will be shut-down. The system will be unavailable for users, and all running jobs will be lost. We will be performing some work on Grex during the maintenance window as well. ETA for Grex availabinity is noon, September 1.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item></channel></rss>