Update - Power has been restored, but nodes will remain down until we can reaccess power tomorrow.
Aug 26, 2026 - 19:13 PDT
Identified - One of the power breakers that powers parts of both Hive and Farm tripped. Data Center operators are calling facilities.
Aug 26, 2026 - 17:23 PDT
Cluster: Hive Degraded Performance
90 days ago
100.0 % uptime
Today
Cluster: Farm Operational
90 days ago
100.0 % uptime
Today
Cluster: Franklin Operational
90 days ago
100.0 % uptime
Today
Login Node Operational
90 days ago
99.35 % uptime
Today
Compute Nodes Partial Outage
90 days ago
98.95 % uptime
Today
GPU Nodes Partial Outage
90 days ago
98.95 % uptime
Today
Network Operational
90 days ago
99.35 % uptime
Today
Storage Operational
90 days ago
98.88 % uptime
Today
Quobyte Parallel File System Operational
90 days ago
98.69 % uptime
Today
Home Directories Operational
90 days ago
99.28 % uptime
Today
Legacy Storage Operational
90 days ago
98.68 % uptime
Today
Module System and Software Operational
90 days ago
99.35 % uptime
Today
Hippo User Portal Operational
90 days ago
99.54 % uptime
Today
OnDemand Operational
90 days ago
99.37 % uptime
Today
Scheduler (Slurm) Operational
90 days ago
100.0 % uptime
Today
Operational
Degraded Performance
Partial Outage
Major Outage
Maintenance
Major outage
Partial outage
No downtime recorded on this day.
No data exists for this day.
had a major outage.
had a partial outage.

Scheduled Maintenance

Emergency Slurm update September 2nd Sep 2, 2026 08:00-18:00 PDT

Update - We will be undergoing scheduled maintenance during this time.
Aug 24, 2026 - 15:38 PDT
Scheduled - SchedMD, the company that develops the Slurm job scheduler, has privately announced several security vulnerabilities. These vulnerabilities have not been disclosed at this time pending remediation, but have been reported to allow escalated privileges within the clusters.

SchedMD is scheduled to publish a fix on September 2, 2026. As per our commitment to follow UCOP security guidelines, we will be upgrading Slurm once the fix has been released. As a result, we will be implementing a maintenance reservation that will prevent jobs from starting if they overlap the September 2 restart. Currently running jobs that are scheduled to run past September 2 may be subject to cancellation depending on our assessed severity of the vulnerabilities once they are fully disclosed.

Aug 24, 2026 - 15:32 PDT
Aug 27, 2026

No incidents reported today.

Aug 26, 2026

Unresolved incident: Hive: Partial power outage in the Data Center.

Aug 25, 2026

No incidents reported.

Aug 24, 2026

No incidents reported.

Aug 23, 2026

No incidents reported.

Aug 22, 2026

No incidents reported.

Aug 21, 2026
Resolved - We had the Data Center staff rebalance power across the 3 feeds they provide us, which resolved this issue.
Aug 21, 15:40 PDT
Update - Quobyte, as well as most nodes, were brought back into service this morning. Admins are monitoring the situation.
Jul 6, 17:21 PDT
Monitoring - We have brought up the Quobyte file system and are working on maintenance tasks regarding the outage.
Jul 6, 09:01 PDT
Update - Facilities were able to restore power to the data center, but we are not able to bring up the Quobyte file system. Data center operators are not available to assist until Monday morning.
Jul 5, 20:12 PDT
Update - Admins are on-site waiting for campus facilities.
Jul 5, 15:32 PDT
Identified - Power has partially failed in the Data Center, which has brought Hive down again. Staff are in contact with the Data Center Operators and Facilities is being scheduled to investigate the power issue.
Jul 5, 10:16 PDT
Aug 20, 2026

No incidents reported.

Aug 19, 2026

No incidents reported.

Aug 18, 2026

No incidents reported.

Aug 17, 2026

No incidents reported.

Aug 16, 2026

No incidents reported.

Aug 15, 2026

No incidents reported.

Aug 14, 2026

No incidents reported.

Aug 13, 2026

No incidents reported.