{"is":"issue","title":"Planned HPCC/Grex outage for electrical and cooling work.","body":"\u003ch4 id=\"update-sept--10\"\u003eUpdate Sept  10\u003c/h4\u003e\n\u003cp\u003eThe outage is over. Grex is fully online and available to users.\nThere are many important changes made on the Grex system. Please check them out at:\u003c/p\u003e\n\u003cp\u003e\u003ca href=\"https://um-grex.github.io/grex-docs/updates/\"\u003ehttps://um-grex.github.io/grex-docs/updates/\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eYour Grex HPC team.\u003c/p\u003e\n\u003ch4 id=\"update-sept--6\"\u003eUpdate Sept  6\u003c/h4\u003e\n\u003cp\u003eDue to a delay with deployment of the new water cooling system, Grex\u0026rsquo;s outage is extended until Wednesday, Sept. 11.\nAt this point, the cooling for new row of racks cannot be fully enabled. Thus, the partial availability of Grex continues.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eSSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)\u003c/li\u003e\n\u003cli\u003eHome and Project file systems are online.\u003c/li\u003e\n\u003cli\u003eOpenOnDemand portal (\u003ca href=\"https://zebu.hpc.umanitoba.ca\"\u003ehttps://zebu.hpc.umanitoba.ca\u003c/a\u003e, Simplified Desktop) is online\u003c/li\u003e\n\u003cli\u003eRunning jobs of short duration (must end before September 9, 2024) on skylake and GPU partitions would work.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThank you for your patience!\u003c/p\u003e\n\u003ch4 id=\"update-aug-30\"\u003eUpdate Aug 30\u003c/h4\u003e\n\u003cp\u003eWe have completed the migration of all of the storage systems, and most of the compute servers into the new datacentre racks.\nHowever, the cooling system installation and acceptance is due next week, so the Grex system is not yet fully online.\u003c/p\u003e\n\u003cp\u003eDuring the long weekend, users have access to the following Grex services or systems:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eSSH to Login nodes (yak.hpc.umanitoba.ca; grex.hpc.umanitoba.ca is now a yak alias)\u003c/li\u003e\n\u003cli\u003eHome and Project file systems are online.\u003c/li\u003e\n\u003cli\u003eOpenOnDemand portal (\u003ca href=\"https://zebu.hpc.umanitoba.ca\"\u003ehttps://zebu.hpc.umanitoba.ca\u003c/a\u003e, Simplified Desktop) is online\u003c/li\u003e\n\u003cli\u003eRunning jobs of short duration (must end before September 3, 2024) on skylake and some of the GPU partitions would work.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThe following systems or services are as of now offline and unavailable: Old login nodes tatanka and bison are decommissioned and unavailable.  grex.hpc.umanitoba.ca is now a yak alias. Old compute partition is decommissioned and unavailable. Most new GPU and CPU partitions are offline because the cooling system is yet to be completed in HPCC.\u003c/p\u003e\n\u003ch4 id=\"update-as-of-aug-28\"\u003eUpdate as of Aug 28\u003c/h4\u003e\n\u003cp\u003eThe First phase: Aug 26 - Aug 28, 2024 is done. We have migrated our storage, login and management nodes to the final location.\nGrex is now partially open for users with limitted services:\u003c/p\u003e\n\u003cpre\u003e\u003ccode\u003e  - Use the login nodes and OOD portal\n  - Access to storage {home and project} if you need to access your data.\n\u003c/code\u003e\u003c/pre\u003e\n\u003cp\u003ePlease note that users can not yet submit jobs as the migration of the compute nodes is not done yet, pending completion of the new cooling systems. We may also experience intermittent interruptions with access to the storage and the login nodes as we are continue with the outage.\u003c/p\u003e\n\u003ch4 id=\"outage-started-on-aug-26\"\u003eOutage started on Aug 26\u003c/h4\u003e\n\u003cp\u003eThere is a planned outage on Grex in effect now.\u003c/p\u003e\n\u003cp\u003eDuring this outage, Physical Plant will work on HPCC power and cooling, and the entire Grex system will be powered down. Then, the system will be migrated to our new water cooled rack infrastructure.\u003c/p\u003e\n\u003cp\u003eUsers will not have access to any Grex services (compute, storage and the OOD Web portal) during the fist stage of the outage that is expected to last at least three days (until Aug 29).\u003c/p\u003e\n\u003cp\u003eWe will be updating this page as the work in HPCC  progresses.\u003c/p\u003e\n\u003cp\u003eShould you have any questions about the upcoming Grex outage, please do not hesitate to contact us at \u003ca href=\"mailto:support@tech.alliancecan.ca\"\u003esupport@tech.alliancecan.ca\u003c/a\u003e ! Thank you for your patience,\u003c/p\u003e\n\u003cp\u003eYour Grex HPC team.\u003c/p\u003e\n","createdAt":"2024-08-26 08:00:00 +0000 UTC","lastMod":"2024-08-26 08:00:00 +0000 UTC","permalink":"https://um-grex.github.io/status/issues/2024-08-26-planned-hpcc-outage/","severity":"down","resolved":true,"informational":false,"resolvedAt":"2024-09-10 16:00:00","affected":["Compute nodes","Network","Lustre /project"],"filename":"2024-08-26-planned-hpcc-outage.md"}