← The Signal Report Work with me

Report 059 · Energy Storage

Why data centers drop off the grid, and why the grid's own repair is what triggers it

On a July evening in 2024, a lightning arrestor failed on a 230 kV line and about 1,500 megawatts of Virginia data center load left the grid inside two minutes. No utility shed that load. The data centers disconnected themselves, correctly, according to protection settings that were working exactly as designed. What tripped them was the grid's standard procedure for fixing the fault.

There is a specific kind of failure I spent a career around, and it is not a component breaking. It is two systems, each performing exactly as specified, discovering a coupling that nobody modeled. In electronic warfare we called one version of it fratricide: your jammer is not malfunctioning when it takes down your own radios, it is transmitting, which is its job. The damage comes from the interaction, not the defect.

The grid just produced a textbook example, twice, and the second one was ten days ago.

What happened, from NERC's incident review

At approximately 7:00 p.m. Eastern on July 10, 2024, a lightning arrestor failed on a 230 kV transmission line in the Eastern Interconnection. The failure produced a permanent fault, and the line eventually locked out.

That much is routine. What followed was not. Coincident with the disturbance, the same local area saw roughly 1,500 megawatts of load reduction, and NERC's own review is precise about who did it:

"None of this load was disconnected from the system by utility equipment; rather, the load was disconnected on the customer side by customer protection and controls. It was determined that the 1,500 MW of load reduction was exclusively data center-type load."

Read that twice, because it inverts the usual grid story. Nobody shed this load. Nobody ordered it off. The grid did not fail to deliver power to these facilities. The facilities decided, in milliseconds, that they would rather not be connected, and left.

The physical consequences were modest and NERC says so plainly: frequency rose to a high of 60.047 Hz and settled back to 60.0 Hz in about four minutes, voltage rose to 1.07 per unit, and operators removed shunt capacitor banks in the local area to bring voltages back to normal. Nothing broke. I want to be exact about that, because the interesting part of this story is the mechanism, not a disaster that did not happen.

The six dips

Here is the part that makes it a systems problem rather than an equipment problem.

When a transmission line faults, the overwhelming majority of faults are transient: a branch, a bird, a lightning strike, an arc that clears itself the moment the line de-energizes. So protection engineers have used automatic reclosing for about a century. Trip the line, wait, close it back in. If the fault is gone, service is restored in under a second and nobody notices. If it is still there, trip again and try once more.

On this line the auto-reclosing control was configured for three reclose attempts, staggered at each end. That configuration produced six successive voltage depressions, and NERC tabulates every one of them:

Depression 1 (initial arrestor failure), 19:00:23.351, lasting 42 milliseconds. Depression 2 (simultaneous automatic reclose at both terminals), 19:00:23.883, 66 ms. Depression 3 (second reclose at terminal A), 19:00:39.211, 58 ms. Depression 4 (second reclose at terminal B), 19:00:44.630, 50 ms. Depression 5 (third reclose at terminal A), 19:01:39.600, 66 ms. Depression 6 (third reclose at terminal B), 19:01:45.016, 59 ms.

Six voltage sags, none longer than 66 milliseconds, spread over about 82 seconds. To the transmission system that is a fault being correctly and repeatedly attacked by its own recovery logic. It is the grid trying to heal.

What a data center hears

A data center is, electrically, a very large pile of power electronics that hates voltage sags. NERC's review explains the equipment plainly: to ride through disturbances, facilities use uninterruptible power supplies that "will instantaneously take over providing power to the data center equipment when a grid disturbance occurs."

Three architectures show up. A centralized design puts static UPS units at the load-center level, typically 2 to 5 MW each, using power electronics to switch the load onto a battery bank. A decentralized design uses many small rack-level UPSs, typically 3 to 4 kW. And a dynamic rotary UPS, a DRUPS, uses a flywheel plus a clutch that spins up a diesel engine when the grid misbehaves.

Those architectures behave differently as seen from the grid, and the difference is the whole ballgame. A static battery UPS takes the load for the length of the sag and hands it back when voltage recovers, so the grid sees a brief dropout that returns quickly. A DRUPS transfers to the flywheel and the engine, and per NERC, "transferring the load back to the grid from the DRUPS system must be done manually." Same disturbance, two utterly different shapes of load loss, decided by a purchasing choice made years earlier.

None of this is a defect. A UPS that switches to battery during a voltage sag is a UPS doing precisely the thing it was bought to do.

The counter

Then NERC describes the control scheme that actually caused the damage, and it is worth quoting because it is the hinge of the entire event:

"The scheme detects and counts voltage disturbances on the grid. If a certain number of voltage disturbances are seen within a certain time, the data center will transfer its load to the backup system, and it will remain there until it is manually reconnected to the grid. The typical number of voltage disturbances that trigger this scheme is three, and a typical time is one minute."

Three dips in one minute, and the facility stops trusting the grid entirely. It goes to backup and stays there until a human puts it back.

Now put the two designs side by side. The utility's reclosing scheme is built to try three times. The data center's protection scheme is built to quit after three tries. The utility's recovery procedure generates, as its normal signature, exactly the pattern the customer's equipment reads as "this grid is coming apart."

NERC's finding lands where you would now expect: most of the sustained load reduction happened at the third voltage depression, when approximately 1,260 megawatts dropped off and, in the review's words, "did not return for hours." The cause is attributed to "the interaction between the automatic reclosing sequence on the faulted transmission line and the data center's protection/control scheme that counts the number of voltage disturbances within a specified period of time."

Neither device malfunctioned. Neither engineering team was careless. The reclose logic is correct. The disturbance counter is correct. They were designed by different companies, for different owners, against different failure models, and they had never been simulated together because the grid had no model of what was behind that meter.

The measurement nobody had

Which brings me to the part I find hardest to look away from, because it is the same problem I write about in laboratories.

On May 4, 2026, NERC issued a Level 3 Alert, its most serious category, with seven Essential Actions for computational loads. Read what Essential Action #1 actually asks transmission planners to go collect from data centers: expected minimum and maximum megawatts and power factor by season; dynamic parameters including "uninterruptable power supply (UPS) settings and configurations"; the percentage split of IT load versus cooling load; expected maximum ramp rate up and down; and the settings of protective devices, including "the reconnecting voltage and timing for when the computational load facility would be expected to reconnect to the System."

Essential Action #6 asks transmission owners to install dynamic fault recorders at these facilities to actually capture what they do during disturbances.

You do not write that list if you already have the data. Every item on it is a parameter that determines how hundreds of megawatts behave in the first 100 milliseconds of a fault, and the alert exists because planners were modeling these facilities without any of it. The grid was carrying, in some areas, gigawatts of a load class whose dynamic behavior had never been measured, only assumed. It is the same failure as an error bar nobody defined or a reagent nobody validated: the number in the model was never wrong, it was never a measurement in the first place.

NERC even names the model it wants used, PERC1, for Power Electronic Reconnecting and Ceasing. The name is the admission. The defining behaviors of this load are ceasing and reconnecting.

It happened again, at twice the size

On the morning of July 23, 2026, a transmission line fault in Northern Virginia produced the same class of event. Roughly 3 gigawatts of data center load came off the PJM system between 7:55 and 8:00 a.m., about 3 percent of system demand at the time.

Again, the load was not shed. Dominion Energy's spokesperson Jeremy Slayton said the data centers "elected to transfer load off the system," that "no load was shed and Dominion Energy did not disconnect area data centers," and that "Our system protections acted exactly as they were intended to support the reliable operation of the grid." PJM, per Reuters, said the disconnection "produced a measurable frequency change but no reliability impacts to the bulk power system."

Both statements are, as far as I can tell, entirely true. That is what makes it worth writing about. Two years apart, on the largest grid in the United States, the same mechanism produced a load loss roughly twice as large, and everyone's equipment performed as intended both times. A grid operator plans contingencies around losing its single largest generator. Three gigawatts leaving in seconds is larger than almost any generator on the system, and it is a contingency that arrives from the demand side, triggered by a routine fault.

What is actually mandatory, and what is not

Here the coverage went wrong in a way worth correcting, because several outlets reported that NERC's Level 3 Alert "mandates action." The alert says otherwise, in its own words:

"This Level 3 NERC Alert is not the same as a Reliability Standard, and it does not create a mandatory obligation to take the Essential Actions. Your organization will not be subject to penalties for failure to implement the Essential Actions."

What the alert does require is a response. Registered entities had to acknowledge it by May 11, 2026 and report on their activities by August 3, 2026, which is two days from now. And the response form is revealing. It asks entities to rate each Essential Action as low effort, significant effort, or "cumbersome workload," and asks planners whether their requirements already comply, with three choices: yes; no, but we plan to; or "No, and we have no plans to modify our requirements."

That is a survey with teeth in the reporting, not in the doing. It is a serious, well-aimed survey, and treating it as a rule misreads how the process works.

The binding piece arrived separately. On July 16, 2026, in Docket RD26-7-000, FERC issued an order directing NERC to file new or modified mandatory Reliability Standards for computational loads, along with Rules of Procedure and registry changes, by December 31, 2026, with a Phase II work plan due March 1, 2027. The draft registry criterion in the alert is concrete: loads of 20 MW and greater, connected at 60 kV, containing more than 1 MW of IT load. If that holds, a facility that only consumes electricity becomes a registered entity with enforceable reliability obligations, which is a genuine change in what the grid regulates.

Worth noticing: NERC's January 2025 incident review closed by listing open questions, the first being "Should large loads be a NERC registered entity?" It took the regulator eighteen months and a second, larger event to answer it.

The battery is the mechanism

For anyone who works on storage, there is an uncomfortable symmetry here. We spend most of our time arguing about how much energy a system holds and how long it lasts. This event turned on none of that. The UPS batteries did their job for a few dozen milliseconds and were never remotely stressed.

What decided the outcome was the control law: the voltage threshold at which the system stops trusting the grid, how many events it tolerates before it latches, and whether coming back is automatic or requires a person. Those are lines of logic, not kilowatt-hours, and they are usually set by a vendor default that nobody revisits after commissioning. In a facility, the default is invisible. Aggregated across an interconnection, the default is a contingency.

That is also why the same building can be an asset instead of a liability. A data center with a large storage system, coordinated ride-through settings, and an agreed reconnection ramp is a grid resource. The same building with three-strikes logic and manual reconnection is a 1,260 MW hole that lasts hours. The hardware is nearly identical. The difference is whether anyone measured and negotiated the behavior, which is exactly the gap NERC is now trying to close.

Getting that behavior right, the cycling and control logic rather than the nameplate, is the work I do on the energy side, and I'll be exact about my role: I help design the AI battery-cycling systems for a veteran-owned (HUBZone) energy-storage integrator. I don't own that company and earn nothing from this link; I flag it because it's a field I build in, not just write about. Full policy here.

The signal

Data centers are already reshaping how storage gets built and interconnected. This is the other half of that story: they are also reshaping what a contingency is.

The lesson generalizes past the grid. When two protective systems meet for the first time in production, the dangerous question is never "will each one work?" Both worked here. The question is whether either one's normal output looks, to the other, like an emergency. A reclose sequence is a repair. A disturbance counter reads it as a collapse. Both are right, and the interaction cost 1,260 megawatts for hours.

Three questions recover most of this for any large load. What does its protection do on the third disturbance, not the first? Does it come back by itself, or does someone have to drive in? And has anyone actually instrumented it, or is the model built from a datasheet nobody tested?

The grid spent a century learning to plan for generators that fail. It is now learning that the customer can be the contingency, and the customer's equipment does not have to break to cause one.

Sources

  1. North American Electric Reliability Corporation, "Incident Review: Considering Simultaneous Voltage-Sensitive Load Reductions," published 8 January 2025. (PRIMARY. Downloaded and text-extracted locally. Source for the July 10, 2024 event: the ~7:00 p.m. Eastern timing, the 230 kV lightning arrestor failure and lockout in the Eastern Interconnection, the three-attempt staggered reclosing configuration, all six voltage depression timestamps and durations (42–66 ms), the ~1,500 MW figure and its attribution exclusively to data center-type load, the verbatim sentence that no load was disconnected by utility equipment, the 60.047 Hz frequency peak and ~4-minute recovery, the 1.07 per-unit voltage and the shunt capacitor bank removal, the UPS architecture descriptions and the 2–5 MW / 3–4 kW / DRUPS figures, the verbatim manual-transfer sentence for DRUPS, the verbatim disturbance-counting scheme passage with its typical three-in-one-minute values, the ~1,260 MW that did not return for hours, the attribution of the loss to the reclose/counter interaction, and the closing open question on whether large loads should be a NERC registered entity.)
  2. North American Electric Reliability Corporation, "Essential Action to Industry: Computational Load Modeling, Studies, Instrumentation, Commissioning, Operations, Protection, and Control" (Level 3 NERC Alert), initial distribution 4 May 2026. (PRIMARY. Downloaded and text-extracted locally, 15 pages. Source for the May 4 2026 issuance, the seven Essential Actions, the May 11 acknowledgement and August 3 2026 reporting deadlines, the verbatim passage stating the alert is not a Reliability Standard and carries no penalties, the Essential Action #1 data-collection list including UPS settings and reconnecting voltage and timing, the PERC1 model and its expansion, Essential Action #6 on dynamic fault recording, the response-form workload options and the "no plans to modify" option, and the draft registry criterion of 20 MW and greater, connected at 60 kV, with more than 1 MW of IT load.)
  3. Federal Energy Regulatory Commission, "E-1 | RD26-7-000," dated 16 July 2026. (Opened. Confirms the docket number and order date on FERC's own site; the page is a stub carrying no order text.)
  4. Darrell Proctor, "FERC Orders Mandatory NERC Reliability Standards for Data Center and Other Computational Loads," POWER Magazine, July 2026. (Opened and read. Source for the substance of the FERC order: the December 31, 2026 deadline for filing standards plus Rules of Procedure and registry criteria changes, and the March 1, 2027 Phase II work plan deadline.)
  5. Michael Brooks, "Line Fault Causes 3 GW in Data Center Load to Drop in Virginia," RTO Insider, 23 July 2026. (Opened. Source for the July 23 2026 event, the 7:55–8:00 a.m. window, the ~3 GW figure and the transmission line fault cause.)
  6. Shane Snider, "Fault in Data Center Alley Triggered 3 GW Load Drop on PJM," Data Center Knowledge, 23 July 2026. (Opened. Source for the Ashburn location, the ~3 percent of system demand figure, the three verbatim quotes from Dominion Energy spokesperson Jeremy Slayton, and PJM's statement as reported by Reuters that the event produced a measurable frequency change but no reliability impacts.)

Note on the July 2026 date: RTO Insider and Data Center Knowledge both place the event on 23 July 2026 and both published that day; at least one other account dates it 22 July. This report uses 23 July, the date given by the two sources actually opened, and flags the discrepancy rather than resolving it silently. Note on scope: neither event caused a reported reliability failure, and this report does not claim one. Separately, NERC's incident review states that similar incidents have occurred in other Interconnections involving cryptocurrency mining and oil and gas loads; no figures for those are cited here because none were opened.

Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports