<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><link rel="alternate" type="text/html" href="https://um-grex.github.io/status/"/><title>Network on Status of the Grex HPC system</title><link>https://um-grex.github.io/status/affected/network/</link><description>Incident history</description><generator>github.com/cstate</generator><language>en-us</language><lastBuildDate>2025-09-28T06:00:00+00:00</lastBuildDate><updated>2025-09-28T06:00:00+00:00</updated><copyright>The MIT License (MIT) Copyright © 2025 UM-Grex</copyright><atom:link href="https://um-grex.github.io/status/affected/network/index.xml" rel="self" type="application/rss+xml"/><item><title>[Resolved] Grex login failure, login nodes and OOD.</title><link>https://um-grex.github.io/status/issues/2025-09-28-grex-login-failure/</link><pubDate>Sun, 28 Sep 2025 06:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-09-28-grex-login-failure/</guid><category>2025-09-29 22:00:00</category><description>&lt;h1 id="login-failure-resolved"&gt;Login failure resolved&lt;/h1&gt;
&lt;p&gt;The issue was a temporary lock-up of the /home NFS server. After restarting it, the system operates normally.&lt;/p&gt;
&lt;h1 id="login-failure"&gt;Login failure&lt;/h1&gt;
&lt;p&gt;Access to Grex login nodes and OpenOnDemand are disrupted. We are investigating.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</description><content type="html">&lt;h1 id="login-failure-resolved"&gt;Login failure resolved&lt;/h1&gt;
&lt;p&gt;The issue was a temporary lock-up of the /home NFS server. After restarting it, the system operates normally.&lt;/p&gt;
&lt;h1 id="login-failure"&gt;Login failure&lt;/h1&gt;
&lt;p&gt;Access to Grex login nodes and OpenOnDemand are disrupted. We are investigating.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Unplanned power outage in HPCC, Grex down</title><link>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</link><pubDate>Wed, 23 Apr 2025 09:10:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-04-23-unplanned-power-outage/</guid><category>2025-04-23 12:40:00</category><description>&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</description><content type="html">&lt;h4 id="power-to-the-hpcc-centre-restored"&gt;Power to the HPCC Centre restored&lt;/h4&gt;
&lt;p&gt;Manitoba Hydro had restored power to Campus, and Grex is back online.
All running and queued jobs were lost during the outage.
We have used the opportunity to update SLURM scheduler to the current major version 24.11 .
All Grex subsystems (compute, storage, login nodes and Web portal) are operational.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h4 id="a-power-outage-happened-in-hpcc-centre"&gt;A power outage happened in HPCC Centre&lt;/h4&gt;
&lt;p&gt;A power outage in Grex&amp;rsquo;s datacentre happened , with a complete loss of power at about 9:10 AM Winnipeg time.
The system is down. The reason for the outage is a problem at Manitoba Hydro, our electricity provider.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://account.hydro.mb.ca/Portal/outeroutage.aspx"&gt;https://account.hydro.mb.ca/Portal/outeroutage.aspx&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We are waiting for the power to be restored. Thank you for your patience!&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC datacentre power outage</title><link>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</link><pubDate>Sun, 23 Feb 2025 07:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2025-02-23-hpcc-poweroutage/</guid><category>2025-02-24 17:10:00</category><description>&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</description><content type="html">&lt;h4 id="hpcc-planned-power-outage-update"&gt;HPCC planned power outage update&lt;/h4&gt;
&lt;p&gt;The outage started on Feb 23 is over. Grex is operational. Some of
the GPU compute nodes may still be unavailable, and will be in production shortly.&lt;/p&gt;
&lt;h4 id="hpcc-planned-power-outage"&gt;HPCC planned power outage&lt;/h4&gt;
&lt;p&gt;Physical Plant had informed us that it plans to shut down power feed for transformer work, to several buildings including HPCC where Grex is located.
For this reason, we have a complete Grex outage starting from 7 AM, Feb 23, 2025. The planned end of the outage is by the end of the day on Monday, Feb 24.
The system is completely be unavailable to the users. Login nodes and data are not be accessible during the outage.&lt;/p&gt;</content></item><item><title>[Resolved] Planned HPCC/Grex outage for electrical and cooling work.</title><link>https://um-grex.github.io/status/issues/2024-08-26-planned-hpcc-outage/</link><pubDate>Mon, 26 Aug 2024 08:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2024-08-26-planned-hpcc-outage/</guid><category>2024-09-10 16:00:00</category><description>&lt;h4 id="update-sept--10"&gt;Update Sept 10&lt;/h4&gt;
&lt;p&gt;The outage is over. Grex is fully online and available to users.
There are many important changes made on the Grex system. Please check them out at:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;
&lt;h4 id="update-sept--6"&gt;Update Sept 6&lt;/h4&gt;
&lt;p&gt;Due to a delay with deployment of the new water cooling system, Grex&amp;rsquo;s outage is extended until Wednesday, Sept. 11.
At this point, the cooling for new row of racks cannot be fully enabled. Thus, the partial availability of Grex continues.&lt;/p&gt;</description><content type="html">&lt;h4 id="update-sept--10"&gt;Update Sept 10&lt;/h4&gt;
&lt;p&gt;The outage is over. Grex is fully online and available to users.
There are many important changes made on the Grex system. Please check them out at:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://um-grex.github.io/grex-docs/updates/"&gt;https://um-grex.github.io/grex-docs/updates/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;
&lt;h4 id="update-sept--6"&gt;Update Sept 6&lt;/h4&gt;
&lt;p&gt;Due to a delay with deployment of the new water cooling system, Grex&amp;rsquo;s outage is extended until Wednesday, Sept. 11.
At this point, the cooling for new row of racks cannot be fully enabled. Thus, the partial availability of Grex continues.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)&lt;/li&gt;
&lt;li&gt;Home and Project file systems are online.&lt;/li&gt;
&lt;li&gt;OpenOnDemand portal (&lt;a href="https://zebu.hpc.umanitoba.ca"&gt;https://zebu.hpc.umanitoba.ca&lt;/a&gt;, Simplified Desktop) is online&lt;/li&gt;
&lt;li&gt;Running jobs of short duration (must end before September 9, 2024) on skylake and GPU partitions would work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thank you for your patience!&lt;/p&gt;
&lt;h4 id="update-aug-30"&gt;Update Aug 30&lt;/h4&gt;
&lt;p&gt;We have completed the migration of all of the storage systems, and most of the compute servers into the new datacentre racks.
However, the cooling system installation and acceptance is due next week, so the Grex system is not yet fully online.&lt;/p&gt;
&lt;p&gt;During the long weekend, users have access to the following Grex services or systems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)&lt;/li&gt;
&lt;li&gt;Home and Project file systems are online.&lt;/li&gt;
&lt;li&gt;OpenOnDemand portal (&lt;a href="https://zebu.hpc.umanitoba.ca"&gt;https://zebu.hpc.umanitoba.ca&lt;/a&gt;, Simplified Desktop) is online&lt;/li&gt;
&lt;li&gt;Running jobs of short duration (must end before September 3, 2024) on skylake and some of the GPU partitions would work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following systems or services are as of now offline and unavailable: Old login nodes tatanka and bison are decommissioned and unavailable. grex.hpc.umanitoba.ca is now a yak alias. Old compute partition is decommissioned and unavailable. Most new GPU and CPU partitions are offline because the cooling system is yet to be completed in HPCC.&lt;/p&gt;
&lt;h4 id="update-as-of-aug-28"&gt;Update as of Aug 28&lt;/h4&gt;
&lt;p&gt;The First phase: Aug 26 - Aug 28, 2024 is done. We have migrated our storage, login and management nodes to the final location.
Grex is now partially open for users with limitted services:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt; - Use the login nodes and OOD portal
- Access to storage {home and project} if you need to access your data.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Please note that users can not yet submit jobs as the migration of the compute nodes is not done yet, pending completion of the new cooling systems. We may also experience intermittent interruptions with access to the storage and the login nodes as we are continue with the outage.&lt;/p&gt;
&lt;h4 id="outage-started-on-aug-26"&gt;Outage started on Aug 26&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect now.&lt;/p&gt;
&lt;p&gt;During this outage, Physical Plant will work on HPCC power and cooling, and the entire Grex system will be powered down. Then, the system will be migrated to our new water cooled rack infrastructure.&lt;/p&gt;
&lt;p&gt;Users will not have access to any Grex services (compute, storage and the OOD Web portal) during the fist stage of the outage that is expected to last at least three days (until Aug 29).&lt;/p&gt;
&lt;p&gt;We will be updating this page as the work in HPCC progresses.&lt;/p&gt;
&lt;p&gt;Should you have any questions about the upcoming Grex outage, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; ! Thank you for your patience,&lt;/p&gt;
&lt;p&gt;Your Grex HPC team.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Grex outage for SLURM and minor OS update</title><link>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</link><pubDate>Mon, 18 Dec 2023 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-12-18-planned-outage/</guid><category>2023-12-19 18:00:00</category><description>&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;</description><content type="html">&lt;h4 id="grex-system-outage-completed-on-dec-19-2023"&gt;Grex system outage completed on Dec 19 2023&lt;/h4&gt;
&lt;p&gt;The SLURM scheduler had been updated. Linux OS also had a minor update. The Grex system is now fully operational.
If you encounter any problems, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;
&lt;h4 id="grex-system-outage-on-december-18-2023"&gt;Grex system outage on December 18 2023&lt;/h4&gt;
&lt;p&gt;There is a planned outage on Grex in effect. We are performing a major update of the SLURM scheduler, communication libraries, and minor Linux OS updates for security patching.&lt;/p&gt;
&lt;p&gt;We will eboot and reinstall all of Grex compute and login nodes and to migrate the SLURM job database.
Thus Jobs that are still running by the time of the outage will be lost, and Grex login nodes will be unavailable during the outage window.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Two Planned Grex outages for HPCC transformer work</title><link>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</link><pubDate>Tue, 12 Sep 2023 18:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-09-12_18-planned-grex-outages/</guid><category>2023-09-19 21:00:00</category><description>&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;</description><content type="html">&lt;h4 id="both-power-outages-are-now-complete"&gt;Both power outages are now complete&lt;/h4&gt;
&lt;p&gt;Grex system is open to the users. Queued jobs were not affected and seems to be running now.
There were no major upgrades or changes done during the outage.&lt;/p&gt;
&lt;h4 id="planned-grex-power-outages-in-september-2023"&gt;Planned Grex power outages in September 2023&lt;/h4&gt;
&lt;p&gt;The Physical Plant is about to perform some electrical works on the transformer that feeds, amongst other things on campus, the HPCC data center that hosts Grex. The outage will start for 6 PM on September 12 and September 19. These outages would require a complete power shutdown in HPCC for about an hour, which means the system would be completely inaccessible to the users, and all running jobs would be terminated.&lt;/p&gt;
&lt;p&gt;To avoid the failure of jobs, we have made two reservations to avoid any longer jobs (that cannot be finished before the beginning of the outage) from starting. For more information, run the following command from any login node: &amp;ldquo;scontrol show res&amp;rdquo;. To take advantage of the cluster before, and between, the outages, we recommend users submit short jobs that can finish by the time the upcoming outage begins.&lt;/p&gt;
&lt;p&gt;Thank you for your patience! Should you have questions or concerns, please do not hesitate to contact us at &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt; !&lt;/p&gt;</content></item><item><title>[Resolved] Electric failure, loss of power to management rack</title><link>https://um-grex.github.io/status/issues/2023-01-26-electric-failure-datacentre/</link><pubDate>Thu, 26 Jan 2023 14:50:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2023-01-26-electric-failure-datacentre/</guid><category>2023-01-26 18:00:00</category><description>&lt;h4 id="update-datacentre-change-rolled-back-systems-operational"&gt;Update: datacentre change rolled back, systems operational&lt;/h4&gt;
&lt;p&gt;It appears that some of the running jobs continued running , and login nodes&amp;rsquo; access nor storage systems were affected.
Grex is now operational. In case you notice any ongoing issue, please let us know by email to &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, Subject line containing Grex.&lt;/p&gt;
&lt;h4 id="faulty-electrical-work-loss-of-power-to-management-rack"&gt;Faulty electrical work, loss of power to management rack&lt;/h4&gt;
&lt;p&gt;We have an unplanned outage on Grex due to a failed electrical work that affected its management rack at around 2:50 PM, Jan 26, 2023 .&lt;/p&gt;</description><content type="html">&lt;h4 id="update-datacentre-change-rolled-back-systems-operational"&gt;Update: datacentre change rolled back, systems operational&lt;/h4&gt;
&lt;p&gt;It appears that some of the running jobs continued running , and login nodes&amp;rsquo; access nor storage systems were affected.
Grex is now operational. In case you notice any ongoing issue, please let us know by email to &lt;a href="mailto:support@tech.alliancecan.ca"&gt;support@tech.alliancecan.ca&lt;/a&gt;, Subject line containing Grex.&lt;/p&gt;
&lt;h4 id="faulty-electrical-work-loss-of-power-to-management-rack"&gt;Faulty electrical work, loss of power to management rack&lt;/h4&gt;
&lt;p&gt;We have an unplanned outage on Grex due to a failed electrical work that affected its management rack at around 2:50 PM, Jan 26, 2023 .&lt;/p&gt;
&lt;p&gt;Running jobs may be lost, and access to Grex login nodes may be degraded or unavailable.&lt;/p&gt;
&lt;p&gt;We are working on resolving the issue. Sorry about the inconvenience it may have caused.&lt;/p&gt;</content></item><item><title>[Resolved] Planned Network outage of the login nodes</title><link>https://um-grex.github.io/status/issues/2022-09-09-planned-login-outage/</link><pubDate>Fri, 09 Sep 2022 11:45:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2022-09-09-planned-login-outage/</guid><category>2022-08-14 12:00:00</category><description>&lt;h1 id="update-the-switch-complete"&gt;Update: the switch complete!&lt;/h1&gt;
&lt;p&gt;Grex&amp;rsquo;s internet connection to all login nodes is now migrated to the UManitoba network IP space. New login node names:&lt;/p&gt;
&lt;p&gt;grex.hpc.umanitoba.ca (instead of grex.westgrid.ca; same for the ones below)
bison.hpc.umanitoba.ca
tatanka.hpc.umanitoba.ca
yak.hpc.umanitoba.ca
aurochs.hpc.umanitoba.ca&lt;/p&gt;
&lt;h1 id="switching-grex-network-from-bcnet-to-umanitoba-network-"&gt;SWITCHING GREX NETWORK FROM BCNET TO UMANITOBA NETWORK. &amp;lt;==&lt;/h1&gt;
&lt;p&gt;On Friday, Sep 09 between 12:00 and 2:00 PM, we will switch the network
from BCNET {westgrid.ca} to UManitoba network. During this process, access
to Grex via old DNS names {grex, bison, tatanka} will be disrupted. We expect
less disruption for yak.hpc.umanitoba.ca&lt;/p&gt;</description><content type="html">&lt;h1 id="update-the-switch-complete"&gt;Update: the switch complete!&lt;/h1&gt;
&lt;p&gt;Grex&amp;rsquo;s internet connection to all login nodes is now migrated to the UManitoba network IP space. New login node names:&lt;/p&gt;
&lt;p&gt;grex.hpc.umanitoba.ca (instead of grex.westgrid.ca; same for the ones below)
bison.hpc.umanitoba.ca
tatanka.hpc.umanitoba.ca
yak.hpc.umanitoba.ca
aurochs.hpc.umanitoba.ca&lt;/p&gt;
&lt;h1 id="switching-grex-network-from-bcnet-to-umanitoba-network-"&gt;SWITCHING GREX NETWORK FROM BCNET TO UMANITOBA NETWORK. &amp;lt;==&lt;/h1&gt;
&lt;p&gt;On Friday, Sep 09 between 12:00 and 2:00 PM, we will switch the network
from BCNET {westgrid.ca} to UManitoba network. During this process, access
to Grex via old DNS names {grex, bison, tatanka} will be disrupted. We expect
less disruption for yak.hpc.umanitoba.ca&lt;/p&gt;
&lt;p&gt;The process will not affect users&amp;rsquo; data.&lt;/p&gt;
&lt;p&gt;OpenOnDemand interface will also be disabled and hopefully the service will
be restored sometime after the outage when new SSL certificates are in place.&lt;/p&gt;
&lt;p&gt;In the meantime, you can use yak to connect to Grex:&lt;/p&gt;
&lt;p&gt;ssh &lt;a href="mailto:your-user-name@yak.hpc.umanitoba.ca"&gt;your-user-name@yak.hpc.umanitoba.ca&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Replace &lt;your-user-name&gt; by your Alliance (Compute Canada) user name.&lt;/p&gt;
&lt;p&gt;Please note that this node has avx512 architecture and if you have to use it
for compiling your codes, they may not run on “compute” partition. Other than
that, it should behave as any other old login node.&lt;/p&gt;</content></item><item><title>[Resolved] CANARIE outage affects Grex external network and Legacy login nodes</title><link>https://um-grex.github.io/status/issues/2022-08-16-canarie-outage/</link><pubDate>Tue, 16 Aug 2022 09:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2022-08-16-canarie-outage/</guid><category>2022-08-16 14:00:00</category><description>&lt;h1 id="update-connectivity-restored"&gt;Update: connectivity restored&lt;/h1&gt;
&lt;p&gt;All systems should be functioning normally, as the connection is restored.&lt;/p&gt;
&lt;h1 id="a-planned-canarie-outage-disables-grex-legacy-network-to-bcnet"&gt;a planned CANARIE outage disables Grex legacy network to BCNet&lt;/h1&gt;
&lt;p&gt;Thus, network connection to Grex login nodes in .westgrid.ca domain login nodes (bison and tatanka, grex.westgrid.ca ,aurochs.login.ca) is not available at the moment.&lt;/p&gt;
&lt;p&gt;A workaround is to use to the new login node yak.hpc.umanitoba.ca for logging in and transferring files.&lt;/p&gt;
&lt;p&gt;Most running jobs are unaffected; however, commercial licenses for software like MATLAB and ANSYS are also not reacheable due to the network unavailability.&lt;/p&gt;</description><content type="html">&lt;h1 id="update-connectivity-restored"&gt;Update: connectivity restored&lt;/h1&gt;
&lt;p&gt;All systems should be functioning normally, as the connection is restored.&lt;/p&gt;
&lt;h1 id="a-planned-canarie-outage-disables-grex-legacy-network-to-bcnet"&gt;a planned CANARIE outage disables Grex legacy network to BCNet&lt;/h1&gt;
&lt;p&gt;Thus, network connection to Grex login nodes in .westgrid.ca domain login nodes (bison and tatanka, grex.westgrid.ca ,aurochs.login.ca) is not available at the moment.&lt;/p&gt;
&lt;p&gt;A workaround is to use to the new login node yak.hpc.umanitoba.ca for logging in and transferring files.&lt;/p&gt;
&lt;p&gt;Most running jobs are unaffected; however, commercial licenses for software like MATLAB and ANSYS are also not reacheable due to the network unavailability.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Grex has a problem with external network and login nodes</title><link>https://um-grex.github.io/status/issues/2022-06-22_network_and_login_nodes/</link><pubDate>Thu, 23 Jun 2022 19:00:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2022-06-22_network_and_login_nodes/</guid><category>2022-06-24 9:30:00</category><description>&lt;h1 id="update-jun-24-10am"&gt;Update Jun 24, 10AM&lt;/h1&gt;
&lt;p&gt;Finally, all Ethernet switches are powered up, and all Grex login nodes are available. Running jobs and storage were not affected during the outage. Grex should be fully operational now.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h2 id="update-jun-24-8am"&gt;Update Jun 24, 8AM&lt;/h2&gt;
&lt;p&gt;The reason for this partial outage is a faulty UPS that fed some of the Grex network switches. As of now, the power to most of the switches is re-routed, so jobs run normally, but only yak.hpc.umanitoba.ca works for users to connect to.&lt;/p&gt;</description><content type="html">&lt;h1 id="update-jun-24-10am"&gt;Update Jun 24, 10AM&lt;/h1&gt;
&lt;p&gt;Finally, all Ethernet switches are powered up, and all Grex login nodes are available. Running jobs and storage were not affected during the outage. Grex should be fully operational now.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;
&lt;h2 id="update-jun-24-8am"&gt;Update Jun 24, 8AM&lt;/h2&gt;
&lt;p&gt;The reason for this partial outage is a faulty UPS that fed some of the Grex network switches. As of now, the power to most of the switches is re-routed, so jobs run normally, but only yak.hpc.umanitoba.ca works for users to connect to.&lt;/p&gt;
&lt;p&gt;Legacy login nodes of grex.westgrid.ca are on, but external network to them is still unavailable. Please use Yak to connect for now.&lt;/p&gt;
&lt;h2 id="grex-network-management-vms-and-login-nodes-are-down"&gt;Grex network, management VMs and login nodes are down&lt;/h2&gt;
&lt;p&gt;We are investigating the issue. Access to Grex is not possible, but running jobs and storage seems to be largely unaffected.&lt;/p&gt;
&lt;p&gt;If you have questions or concerns, please don’t hesitate to contact us at: &lt;a href="mailto:support@computecanada.ca"&gt;support@computecanada.ca&lt;/a&gt; , mentioning Grex in the subject line.&lt;/p&gt;</content></item><item><title>[Resolved] Westgrid network failure</title><link>https://um-grex.github.io/status/issues/2021-network-failure/</link><pubDate>Wed, 09 Jun 2021 11:35:00 +0000</pubDate><guid>https://um-grex.github.io/status/issues/2021-network-failure/</guid><category>2021-07-28 12:10:00</category><description>&lt;h2 id="this-is-an-example-of-notification-the-problem-was-there-and-was-resolved-the-dates-are-arbitrary"&gt;this is an example of notification. The problem was there and was resolved, the dates are arbitrary&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Update&lt;/em&gt; : The issue with network / packet loss issue is resolved
&lt;span class="faded"&gt;(11:35 UTC — Jul 28)&lt;/span&gt;
.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Problem&lt;/em&gt;:&lt;/p&gt;
&lt;p&gt;Login nodes of Grex are experiencing a heavy packet loss. This leads to connection slowness and intermittent failures.
Sorry about the inconvenience. We are working on resloving the issue
&lt;span class="faded"&gt;(11:35 UTC — Jun 9)&lt;/span&gt;
.&lt;/p&gt;</description><content type="html">&lt;h2 id="this-is-an-example-of-notification-the-problem-was-there-and-was-resolved-the-dates-are-arbitrary"&gt;this is an example of notification. The problem was there and was resolved, the dates are arbitrary&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Update&lt;/em&gt; : The issue with network / packet loss issue is resolved
&lt;span class="faded"&gt;(11:35 UTC — Jul 28)&lt;/span&gt;
.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Problem&lt;/em&gt;:&lt;/p&gt;
&lt;p&gt;Login nodes of Grex are experiencing a heavy packet loss. This leads to connection slowness and intermittent failures.
Sorry about the inconvenience. We are working on resloving the issue
&lt;span class="faded"&gt;(11:35 UTC — Jun 9)&lt;/span&gt;
.&lt;/p&gt;</content></item></channel></rss>