
A predictive maintenance platform flags a centrifugal pump on critical duty: 78% probability of failure inside thirty days. The reliability engineer pulls the asset and finds a probability, a trend line, and a contributing-factors panel listing three channels by importance.
That is enough information to act on. It is not enough information to defend, and the difference between those two becomes the whole question as soon as somebody with a budget asks why the pump came out.
The four askings, and what each one costs
The same four-word question, what did the model see, gets asked at four escalating prices.
The Orientation Asking
An engineer checking a flagged asset on a Tuesday morning wants to know whether this is the bearing she has been watching since June or something new. A contributing-factors panel answers this adequately.
The Production Asking
The pump comes out, the teardown finds a lightly scored race that would plausibly have run another year, and eleven hours of output went with it. The superintendent wants to know why he lost a shift. The panel still shows three channels ranked by importance. It does not identify which sensor readings were sampled and when, that channel two’s displacement probe had drifted past its calibration date in April, or that the model had been retrained in July and this was its first call on the asset class since.
The Programme Asking
After a handful of those, somebody senior asks whether predictive maintenance is earning its keep. This is where programmes die, and they rarely die from a bad model: they die because six months after the fact nobody can produce the evidence that separates a model error from a drifted sensor from a perfectly correct call on an asset that was genuinely marginal, and in the absence of that evidence the cheapest interpretation available to a sceptical budget holder is that the platform has been guessing expensively.
The Legal Asking
An asset fails, somebody gets hurt or a release occurs, and the maintenance record stops being an internal document and becomes evidence that opposing counsel will read line by line with the benefit of hindsight and no obligation to be charitable. At that point the record either reconstructs the decision or it does not.
Why a Conventional Programme Survives the Same Scrutiny
Time-based maintenance answers all four askings trivially, which is worth dwelling on because it is the reason a demonstrably worse engineering method keeps outliving better ones in regulated environments. The pump came out because the interval said so, the interval came from the manufacturer or from site history, and the paperwork names who signed it. The reasoning is mediocre engineering and excellent evidence.
Condition monitoring with human interpretation sits in between. An analyst reviewed a spectrum, wrote a finding, and signed it. Somebody can be asked what they saw.
Predictive platforms improved the engineering and, without anybody choosing it, degraded the evidence. The decision moved inside software that does not write down its reasoning in a form a maintenance file can hold.
What Changed in the Last Five Years
Two developments make this sharper than it was.
The first is regulatory. NIST’s AI Risk Management Framework and ISO/IEC 42001 both ask whether an organisation can demonstrate control over its AI use, rather than whether a given model is capable. The Texas Responsible AI Governance Act has had an Attorney General’s complaint portal open since 1 September 2026, which for Gulf Coast operators is a live consideration rather than a forecast.
The second is quieter. Cority’s 2026 survey of 2,000 environment, health, safety and sustainability leaders, run with Censuswide, found 95% said their teams or frontline workers are using AI outside approved systems. In a maintenance context that means failure narratives drafted by public assistants and anomalous readings interpreted by whatever tool was open. The people doing it are clearing a backlog, not cutting corners. But material of unknown provenance ends up in a record that gets read under subpoena, and that exposure stays invisible for years.
The Four Fields That Answer all Four Askings
The record is not a product and it is not complicated. At the moment a recommendation is made, four things get written down.
1. Input custody. Which sensors fed the recommendation, and when each was last calibrated. This is the field that separates a model error from an instrument error, and it is the one most often missing.
2. Model version. Which version produced the output. Platforms retrain on vendor schedules; a recommendation from the July model is not a recommendation from the April model.
3. Operating envelope. What the asset’s operating state was, with an explicit flag when it ran outside the conditions the model was trained on. A confident prediction about an unfamiliar state is the failure mode worth catching.
4. Human disposition. Who accepted, overrode or deferred the recommendation, and one line on why.
Those four map onto a framework for oil and gas companies in Houston, TX and other operators facing the same problem on long-lived, high-consequence assets. Reliability engineering has never lacked documented-reasoning discipline. It has not yet extended that discipline to decisions where the reasoning moved inside software.
Where the Approach is Still Arguable
Writing four fields at every recommendation is a real data commitment, and most recommendations are never contested. An operator could reasonably argue for capturing the full record only on critical assets, or only where a recommendation triggers an intervention above some cost threshold.
The counter-argument is that the asking which matters most is the one nobody predicted, and a record built selectively is a record with gaps exactly where the unusual events sit. Where that line belongs is a judgment each site makes with its own risk appetite, and one worth making deliberately rather than by default.
