Back

Incident history

September
Alpha Cluster - Institutional Partition Changes
Scheduled for September 18 at 09:00 AM EDT - September 21 at 05:00 PM EDT
Completed

On Monday, September 21, 2026, the Alpha cluster will undergo scheduled maintenance. Key changes:

  • Institution-specific partitions (e.g. -p columbia, -p nyu, etc.) are being retired. All jobs will run through the shared alpha and grace partitions going forward — no action needed if you already use these.
  • Job submissions will now require an explicit --account flag.
  • Researchers without an active project will need to use the burst queue, lower-priority queue instead of being blocked from submitting jobs.

Brief interruptions to job scheduling may occur during the maintenance window. No action is required for most users; if you've been submitting jobs to an institution-specific partition by name, switch to alpha or grace going forward.

September 18 at 09:00 AM EDT
NVL72 Fleet Remediation — Phased Rack Rollout
Scheduled for September 9 at 05:13 PM EDT - September 24 at 05:13 PM EDT
Scheduled

Rack 2 is complete. One node is failed out due to a hardware issue, for which we have an RMA from nvidia on the way.

Rack 3 drain / reservation is now starting.

September 9 at 05:13 PM EDT
Minimum GPU Requirement on Beta
Scheduled for September 1 at 09:00 AM EDT - September 8 at 05:00 PM EDT
Completed

Following our meeting yesterday, we have have raised the minimum GPU request for the beta phase from 1 to 4 GPUs. Researchers will now be required to request 4 GPUs to ensure alignment with the workload designed for the NVL72 system. We are currently discussing the creation of a dedicated partition for testing jobs that require only 1 GPU on the NVL72. This partition will feature a significantly lower wall time, allowing for quicker turnaround on tests. We will provide further updates as they become available.

Researchers may encounter the following error message if they attempt to request only 1 GPU:
srun/sbatch: error: QOSMinGRES
srun/sbatch: error: Unable to allocate resources: Job violates accounting/QOS policy (job submit limit, user's size and/or time limits)

If you have any questions or concerns, please don't hesitate to reach out to us. Researchers can also submit a ticket for assistance.

September 1 at 09:00 AM EDT
August
Beta Rack 1 Fully Restored
Resolved

Issues with rack 1 have been fully resolved. All nodes are back in production.

August 25 at 01:12 PM EDT

Resolved after 2h 46m

One GPU Job Governance Implementation
Scheduled for August 25 at 09:00 AM EDT - August 25 at 10:00 AM EDT
Completed

We will enable our One GPU Job Governance tomorrow at 9 AM. Initial tests will be conducted, after which the RTX nodes will be made available to the community. Please note that the system will be fully functional during this time; however, there may be some issues when submitting single GPU jobs. We apologize for any inconvenience this may cause.

August 25 at 09:00 AM EDT
Friday 08/21 Incident Resolved - Grace/Grace + Beta Control Plane are Online
Resolved

The Beta control plane (login endpoint, Slurm controller) and Grace/Grace cluster remained up and stable throughout the weekend.

The Empire AI team is working closely with our BMS vendor for root-cause analysis and prevention.

August 21 at 02:17 PM EDT

Resolved after 3d

Rack 1 back in service - entire beta cluster looks healthy
Resolved

The NVLink switch issues on Rack 1 have been fully resolved. As part of the remediation, the entire rack — switches and compute nodes — has been updated to the latest supported firmware and operating system versions.

All systems have been verified healthy (fabric connectivity, node health checks, and management connectivity) and Rack 1 is back in normal service, ready to resume serving traffic.

August 9 at 06:08 PM EDT

Resolved after 21d

Draining rack 1 for maintenance
Scheduled for August 9 at 11:58 AM EDT - August 16 at 11:58 AM EDT
In Progress

Draining rack 1 in order to update firmware and verify IMEX / NVLink fabric

August 9 at 11:58 AM EDT
Maintenance on rack 2 (b2-*) on the beta partition for maintenance to resolve an NVLink fabric issue
Scheduled for August 8 at 10:00 AM EDT - August 10 at 09:00 AM EDT
Completed

Maintenance on rack 2 completed

August 8 at 10:00 AM EDT
Beta control plane taken offline by Grace-Grace shunt trip
Resolved

DGX control plane racks have been effectively isolated / mitigated from this incident since leaving Grace/Grace offline and disabling leak-detect shunt trip on that system

August 4 at 08:14 PM EDT

Resolved after 1m

Power Event: Grace/Grace Rack
Resolved

Grace-Grace is again operational

August 4 at 03:12 PM EDT

Resolved after 7d

July

No incidents reported