On Monday, September 21, 2026, the Alpha cluster will undergo scheduled maintenance. Key changes:
Brief interruptions to job scheduling may occur during the maintenance window. No action is required for most users; if you've been submitting jobs to an institution-specific partition by name, switch to alpha or grace going forward.
Rack 2 is complete. One node is failed out due to a hardware issue, for which we have an RMA from nvidia on the way.
Rack 3 drain / reservation is now starting.
Following our meeting yesterday, we have have raised the minimum GPU request for the beta phase from 1 to 4 GPUs. Researchers will now be required to request 4 GPUs to ensure alignment with the workload designed for the NVL72 system. We are currently discussing the creation of a dedicated partition for testing jobs that require only 1 GPU on the NVL72. This partition will feature a significantly lower wall time, allowing for quicker turnaround on tests. We will provide further updates as they become available.
Researchers may encounter the following error message if they attempt to request only 1 GPU:
srun/sbatch: error: QOSMinGRES
srun/sbatch: error: Unable to allocate resources: Job violates accounting/QOS policy (job submit limit, user's size and/or time limits)
If you have any questions or concerns, please don't hesitate to reach out to us. Researchers can also submit a ticket for assistance.
Issues with rack 1 have been fully resolved. All nodes are back in production.
Resolved after 2h 46m
We will enable our One GPU Job Governance tomorrow at 9 AM. Initial tests will be conducted, after which the RTX nodes will be made available to the community. Please note that the system will be fully functional during this time; however, there may be some issues when submitting single GPU jobs. We apologize for any inconvenience this may cause.
The Beta control plane (login endpoint, Slurm controller) and Grace/Grace cluster remained up and stable throughout the weekend.
The Empire AI team is working closely with our BMS vendor for root-cause analysis and prevention.
Resolved after 3d
The NVLink switch issues on Rack 1 have been fully resolved. As part of the remediation, the entire rack — switches and compute nodes — has been updated to the latest supported firmware and operating system versions.
All systems have been verified healthy (fabric connectivity, node health checks, and management connectivity) and Rack 1 is back in normal service, ready to resume serving traffic.
Resolved after 21d
Draining rack 1 in order to update firmware and verify IMEX / NVLink fabric
Maintenance on rack 2 completed
DGX control plane racks have been effectively isolated / mitigated from this incident since leaving Grace/Grace offline and disabling leak-detect shunt trip on that system
Resolved after 1m
Grace-Grace is again operational
Resolved after 7d
No incidents reported