Cover illustration for the article Nobody's Watching: The Physical, Cyber, and Human Failures Hiding Inside Critical Infrastructure

Nobody’s Watching: The Physical, Cyber, and Human Failures Hiding Inside Critical Infrastructure

A blueberry farmer, a Chinese hacking unit, and a retiring SCADA engineer walk into the same pipeline. Nobody’s laughing.


On November 11, 2025, a blueberry farmer in Snohomish County, Washington, noticed a petroleum sheen in a drainage ditch running through his field. He called BP. That phone call (from a civilian standing in mud, not from any sensor, algorithm, or control room display) was the first indication that the Olympic Pipeline, a 400-mile artery carrying 90 percent of the Pacific Northwest’s transportation fuel, was hemorrhaging gasoline into the earth.

What followed was a cascade that touched nearly a million lives. BP shut down both parallel lines minus 16-inch and 20-inch, to determine which was leaking. A partial restart on November 16 failed when the line leaked again. Washington Governor Bob Ferguson and Oregon Governor Tina Kotek declared states of emergency. An estimated 900,000 Thanksgiving travelers faced disruption at Seattle-Tacoma International Airport, which lost its primary jet fuel supply. Airlines scrambled to tanker fuel from as far as Blaine, Washington. Over 200 feet of pipeline had to be excavated before the source was found on November 24: a breach in the 20-inch gasoline line.

The leak detection system (a Computational Pipeline Monitoring (CPM) system, the industry-standard technology mandated by regulators) did not catch it. A farmer did.

This was not an isolated incident. It was the second major leak on the Olympic Pipeline in less than two years. In December 2023, alarms sounded multiple times on the same system near Conway, Washington. Pipeline staff investigated. They found nothing. Alarms sounded again the next day. BP shut down the line. A field technician then found 21,000 gallons of gasoline overflowing from a concrete vault, 4,000 gallons of which had already reached a fish-bearing stream. The cause: a corroded carbon-steel nut on a 3/8-inch monitoring assembly, a nut that should never have been used due to the corrosive potential of combining dissimilar metals. BP was fined $3.8 million. The same pipeline was fined $100,000 for a 2020 diesel spill. And in 1999, a catastrophic rupture on the Olympic system spilled 277,200 gallons of gasoline into Whatcom Creek in Bellingham, creating a fireball that killed two ten-year-old boys and an eighteen-year-old man.

One pipeline. Four significant failures across 26 years. A pattern that did not self-correct.

Senator Maria Cantwell put it bluntly: “The fact that a blueberry farmer, not BP, first identified the spill, and that it is still not known for certain which of the two pipelines is leaking, raises significant concerns about the capabilities of the Olympic Pipeline’s leak detection systems and the adequacy of your inspection and maintenance programs”.

That sentence contains the entire thesis of this article: the systems that critical infrastructure operators rely upon to detect failure, prevent intrusion, and maintain operational continuity are not performing at the level the public, or the regulators, assume.

And the problem is not one problem. It is three, converging simultaneously, with compound effects that no single technology, regulation, or hiring initiative can address in isolation.


Body One: The Physical, Your CPM is Lying to You

Pipeline incidents in the United States are not rare events. They are a daily occurrence, averaging 1.7 per day since 2010, according to PHMSA data. In 2024 alone, 531 incidents were reported, producing 65 fires, 22 explosions, 10 fatalities, 26 injuries, 2,554 evacuees, and $139 million in property damage. Texas accounted for 37 percent of incidents, three times the national average even when normalized by pipeline mileage.

And these are the reported numbers. A FracTracker Alliance audit of four high-profile 2024 incidents found that three out of four significantly misrepresented one or more key metrics, every single misrepresentation in the same direction: minimizing impacts. Injuries undercounted. Property damage listed as $20 when a home and two lives were destroyed. Spill volumes underestimated.

The Texas Railroad Commission’s own data tells an even more alarming story: 12,303 pipeline incidents in that state alone in 2022, more than an order of magnitude above the PHMSA national total. The true daily incident rate across the country may be far higher than anyone can definitively calculate, because the regulatory landscape is fragmented and the data is self-reported by the operators themselves.

The False Positive Trap

At the heart of the detection problem is a paradox that every pipeline operator recognizes but few have solved: alarm fatigue.

A CPM system generates alerts based on mass-balance calculations, pressure differentials, and flow-rate analysis. When the data feeding those calculations is noisy, incomplete, or miscalibrated, which it frequently is on aging infrastructure, the system generates false positives. Operators who receive 40 false alarms per month do what humans inevitably do: they stop trusting the alarms. Response times lengthen. Real anomalies get buried in noise. The system designed to protect the operator becomes the system the operator works around.

This is not a technology failure. It is a data quality failure.

The algorithms themselves are often adequate. What is inadequate is the data flowing into them. SCADA telemetry from legacy sensors drifts. Meter errors compound. Transient operational states (batch changes, pump starts, valve actuations) create data artifacts that mimic leak signatures. Without a rigorous, continuous process for validating the quality of the data before it enters the detection algorithm, even the most sophisticated AI-driven system will either cry wolf or miss the real threat entirely.

The operators who have solved this problem have done so not by replacing their CPM systems, but by layering a data quality validation method upstream of their detection logic, enriching raw SCADA data with meaningful context labels, validating reliability and relevance at every decision point, and feeding only cleaned, verified data into the detection engine. The results are measurable: up to 90 percent reduction in false positives, leak detection speeds improved by four to eight hours, and broad applicability across both gathering and transmission line geometries.

This is the unsexy work that prevents the next blueberry-farmer moment. It does not make headlines. But its absence does.

The Market Is Moving

The global leak detection market was valued at $1.99 billion in 2025 and is projected to reach $3.28 billion by 2034, driven by AI-assisted analytics, fiber-optic sensing, and real-time monitoring integration with SCADA and IoT platforms. The AI-based segment captured the largest market share in 2025. But here is the catch: AI-based leak detection is only as reliable as the data it trains on. Feed a machine-learning model dirty data and you get a confident, fast, wrong answer.

The organizations leading in this space are not simply buying new algorithms, they are building the data governance layer that makes those algorithms trustworthy. They are deploying continuous validation processes that monitor data behavior and availability in real time, ensuring that alarms represent genuine anomalies, not instrumentation noise. The competitive advantage is no longer the algorithm itself. It is the quality of the data flowing through it.


Body Two: The Cyber, They’re Already Inside

Three days ago, February 17, 2026, Dragos, the preeminent OT cybersecurity intelligence firm, released its 9th Annual OT/ICS Cybersecurity Year in Review. The findings should alarm every operator of critical infrastructure in America.

Dragos tracked 119 ransomware groups targeting industrial organizations in 2025, up 49 percent from 80 groups in 2024. Collectively, these groups impacted 3,300 industrial organizations. Manufacturing accounted for more than two-thirds of victims. The average dwell time for ransomware in OT environments was 42 days.

But here is the number that should be printed on every control-room wall in the country: organizations with comprehensive OT visibility detected and contained ransomware incidents in an average of 5 days. Five versus forty-two. That is not an incremental improvement. That is the difference between a manageable disruption and a business-threatening catastrophe.

The Adversaries Have Names

Dragos identified three new OT-specific threat groups in 2025, bringing the total tracked to 26 worldwide, 11 of which were active during the year.

  • AZURITE explicitly targets engineering workstations, the machines where operators change controller logic and interact with physical processes. It rapidly weaponizes publicly available proof-of-concept code, exploiting the lag between disclosure and patching. It exfiltrates alarm data, configuration files, and operational intelligence useful only for disrupting operations.
  • PYROXENE conducts multi-year supply-chain campaigns using social engineering. Fake LinkedIn profiles posing as recruiters target operations personnel. In June 2025, the group deployed custom wiper malware against Israeli targets during regional conflict.
  • SYLVANITE operates as an initial access provider, rapidly weaponizing edge device vulnerabilities and handing off established footholds to VOLTZITEthe group tracked by the U.S. government as Volt Typhoon, for deeper OT intrusions. Dragos directly observed this handoff.

The sophistication here is systemic. These groups are operating as a coordinated ecosystem, with specialized roles: access development, reconnaissance, and operational execution.

KAMACITE: Mapping America’s Control Loops

Perhaps the most chilling revelation in the Dragos report concerns KAMACITE, the access-development group that enables ELECTRUM, the actor tied to Ukraine’s 2015 and 2016 power outages.

Between March and July 2025, KAMACITE conducted sustained reconnaissance against internet-exposed industrial devices across the United States. The scanning was not random. It focused on specific components (Schneider Electric Altivar variable frequency drives, Smart HMIs, Accuenergy AXM metering modules, and Sierra Wireless AirLink cellular gateways) in a sequence that suggested deliberate mapping of entire control loops, from operator interface to actuator to metering point to remote gateway.

This is not espionage. This is targeting. Control-loop mapping removes the key barrier between unauthorized access and physical impact. An attacker who understands the control loop no longer needs to guess how a process behaves.

Separately, VOLTZITE compromised Sierra Wireless AirLink cellular gateways connected to U.S. midstream pipeline operations and pivoted to engineering workstations. In a case study involving Littleton Electric Light and Water Department in Massachusetts, Volt Typhoon was discovered to have been inside the utility’s systems for over 300 days before detection, exfiltrating data related to OT operating procedures and spatial layout of energy grid operations.

As of February 19, 2026, yesterday, researchers warned that Volt Typhoon remains embedded in U.S. utilities and some breaches may never be found.

ELECTRUM: From Ukraine to Poland to You

On December 29, 2025, ELECTRUM executed the first major coordinated cyber attack targeting distributed energy resources at scale, hitting roughly 30 wind farms, solar installations, and combined heat and power facilities across Poland’s electrical grid. The attack disrupted communication and control systems, disabled key equipment beyond repair at affected sites, and caused loss of view, loss of control, and denial-of-service conditions.

While no power outages occurred, Dragos assessed that the affected facilities could have collectively produced about 1.5 gigawatts of electricity, and a sudden, simultaneous loss of that generation would have caused noticeable frequency instability, a factor linked to cascading failures and major blackouts.

The attack represented a strategic shift: from targeting centralized control systems (as in Ukraine’s 2015 and 2016 grid attacks) to targeting the distributed edge of the grid. Wind farms. Solar arrays. The exact infrastructure the United States is building at breakneck speed to power AI data centers.

The Field Findings: Where the Gaps Are

Dragos’s field findings from incident response cases, penetration tests, and assessments paint an unflinching picture of the defensive landscape:

  • 30 percent of incident response cases began with unexplained operational issues, irregular events asset owners couldn’t diagnose.
  • 82 percent of organizations lack clear criteria for when operational anomalies should trigger cyber investigations.
  • 88 percent of tabletop exercises revealed degraded detection capabilities.
  • 56 percent of penetration tests successfully abused living-off-the-land tools without triggering alerts.
  • 81 percent of assessments identified poor IT/OT segmentation.
  • 73 percent of all-time IR cases involved compromised VPN or jumphost credentials.
  • Only 46 percent of assessments found adequate OT network monitoring deployed.

That 30 percent figure, where cases began not with an alert but with a human operator noticing “something weird”, is the single most important data point in this entire report. It means that nearly one-third of the organizations that experienced a cybersecurity incident had no automated mechanism to detect it. The human was the sensor. The person on shift, staring at a display, noticed something didn’t look right.

What happens when that person retires?


Body Three: The Workforce, The Knowledge That Walks Out the Door

The SANS 2024 State of ICS/OT Cybersecurity survey, based on inputs from more than 530 professionals across critical infrastructure sectors, found that over 50 percent of the ICS/OT cybersecurity workforce has fewer than five years of experience. A staggering 51 percent hold no industry-specific certifications. Only 34 percent prepare for cyber incidents using ICS/OT-specific tools and range environments.

Meanwhile, only 25 percent of organizations allocate significant budget toward workforce training, recruitment, and retention, while 52 percent pour their budgets into technology investments. Organizations recognize that people are their greatest risk. They just don’t fund them accordingly.

The problem extends beyond cybersecurity into core operations. Experienced SCADA engineers are retiring in every sector, energy, water, pipelines, chemical. They carry institutional knowledge that exists nowhere in any documented system: the quirks of aging infrastructure, regional load behavior, the workaround that keeps a 1990s-era controller from faulting, the switching sequence at a specific substation that isn’t written down anywhere because the person who memorized it has been doing it for 30 years.

“The risk of retiring talent isn’t just about labor shortages,” one utility workforce analysis noted. “It’s about institutional knowledge walking out the door. From switching sequences at substations to workarounds for legacy system issues, much of what makes operations efficient today lives in the heads of experienced employees and, worryingly, nowhere else”.

This is why knowledge digitization (the structured capture of operational expertise from veteran personnel into retrievable, searchable, actionable formats) is not a nice-to-have initiative. It is a survival mechanism. Organizations that integrate diverse data sources, from SCADA systems to ERP tickets, and structure that data to build domain-specific knowledge bases are preserving the institutional memory that makes safe operations possible. Those that do not are gambling that the next generation will be able to reverse-engineer what the last generation carried in their heads.

The Tabletop Gap

The Dragos finding that 88 percent of tabletop exercises revealed degraded detection capabilities is not just a cybersecurity statistic. It is a workforce readiness statistic. Tabletop exercises expose two things simultaneously: the quality of the response plan and the quality of the people executing it. When half your OT workforce has been on the job for fewer than five years and holds no ICS-specific certifications, the exercise isn’t just testing the plan. It is the training.

Operators who run regular, OT-specific tabletop exercises calibrated to their facility-specific control systems, threat landscape, and operational processes report something that no technology purchase alone can produce: alignment. IT and OT teams that have practiced together communicate faster during real incidents. Operators who have rehearsed manual mode can actually execute it under pressure. And critically, exercises identify the assumptions nobody questioned, the remote access pathway nobody documented, the VPN credential nobody rotated, the alarm that was silenced three years ago and never re-enabled.


The Regulatory Paradox

One might expect that a rising threat environment would produce a rising regulatory response. It has, on paper. The TSA has issued and updated its Pipeline Security Directives six times since the Colonial Pipeline ransomware attack in May 2021, most recently with SD02F in May 2025. The November 2024 Notice of Proposed Rulemaking would formalize these requirements into permanent regulation, mandating comprehensive Cybersecurity Risk Management Programs that integrate IT and OT, risk-based assessments of critical control systems, supply chain security measures, and 24-hour incident reporting.

But enforcement is moving in the opposite direction.

In 2025, PHMSA initiated 111 enforcement cases minus 44 percent below the 2024 total and roughly half the historical average of 217 per year. In the first three months of the Trump administration’s second term, only five pipeline safety enforcement actions were initiated, compared to 68 in the same period of the first Trump administration, a 92 percent decrease. Senator Cantwell’s December 2025 letter to PHMSA revealed that proposed enforcement penalties assessed on pipeline companies declined 98 percent.

PHMSA spokesperson Emily Wong called the characterization “disingenuous,” pointing to the largest single civil penalty in PHMSA history issued against one operator. Industry leaders contend the agency is shifting priorities, not reducing enforcement. But the effect on the ground is unmistakable: the federal referee is stepping back.

This produces a paradox that every pipeline and utility operator must internalize: regulatory requirements are tightening while enforcement is loosening. The rules are harder. The likelihood of getting caught for breaking them is lower. The temptation to defer investment is higher. But the threat environment (physical, cyber, and human) is more dangerous than at any point in the modern history of critical infrastructure.

The operators who thrive in this environment will not be those who do the minimum to avoid a fine. They will be those who build internal capability (defensible architectures, validated detection systems, trained and exercised workforces, and OT-specific cybersecurity programs) because they recognize that in a world of 42-day ransomware dwell times and blueberry-farmer leak detections, self-reliance is not optional.


The Three Characteristics of Resilient Operators

Across every data point examined in this article (from the Dragos Year in Review, to the SANS survey, to the PHMSA incident data, to the Olympic Pipeline failures) a clear profile emerges of what separates resilient operators from vulnerable ones.

1. They treat data quality as infrastructure. Resilient operators do not assume their SCADA data is accurate. They validate it continuously, checking reliability, accuracy, and relevance at every point in the decision-making chain. They enrich raw data with context before it reaches detection algorithms, simulation models, or digital twins. They understand that a digital twin built on bad data is a digital hallucination, and that an AI model trained on noisy telemetry will produce confident, fast, catastrophically wrong answers.

2. They exercise, not just plan. Resilient operators do not write an incident response plan, file it, and call it compliance. They test it, regularly, with OT-specific scenarios calibrated to their actual facilities, their actual control systems, and the actual threat groups targeting their sector. They use tabletop exercises not just to validate the plan, but to cross-train their IT and OT teams, transfer knowledge from veteran operators to newer staff, and identify the invisible assumptions that create blind spots. The SANS data is unambiguous: organizations that test their IRP quarterly have exercised ICS network outages at a 75 percent rate and can operate in manual mode at 72 percent. Monthly testers reach nearly 90 percent.

3. They capture what their people know before those people leave. Resilient operators have structured programs for digitizing institutional knowledge, not just creating SOPs, but integrating diverse data sources from SCADA historians, maintenance records, ERP tickets, and field observations into domain-specific knowledge bases that new operators can query, learn from, and act on. They understand that the most expensive data loss is not a ransomware encryption. It is the retirement of the person who knew that the valve on Line 14 sticks when the ambient temperature drops below 28 degrees, and that the alarm on Compressor 7 has been unreliable since the firmware update in 2019, and that the manual bypass procedure for Station 12 requires a specific sequence that isn’t documented anywhere.

None of this is theoretical. It is operational. And the clock is running.


The author is an independent analyst covering critical infrastructure, geopolitics, and institutional risk.


Related reading

Scott Ortkiese

Scott Ortkiese

President and CEO of Faulkner Capital Holdings. He writes on geopolitics, energy markets, structured finance and American decline, and is the author of the forthcoming book The Decline of the American Empire.

About/so@throughlinesynthesis.com/LinkedIn/Substack