Manufacturing AI

Predictive Maintenance AI: When It Works and When It Does Not

Predictive maintenance is the AI project manufacturers ask for first and the one that most often disappoints. Here is an honest guide to when it pays off, when it is premature, and what to do instead.

9 min read
Predictive Maintenance AI: When It Works and When It Does Not

Predictive maintenance is the AI project most manufacturers ask about first. The pitch is irresistible: stop fixing machines on a calendar and start fixing them just before they break, cutting both unplanned downtime and wasted preventive work. It is also the project that most often stalls, overruns, or quietly gets shelved after a disappointing pilot. The technology is rarely the problem. The conditions around it usually are.

This article is deliberately not a sales page. Most of the value below is in helping you recognise when predictive maintenance is premature for your plant, because launching it too early is the single most reliable way to burn budget and goodwill. If your data, sensors, and processes are ready, predictive maintenance can be one of the highest-return uses of manufacturing AI. If they are not, there are cheaper, faster alternatives that get you most of the benefit while you build toward the real thing.

Why predictive maintenance often fails

The failures rarely look like a model that cannot learn. They look like a model that learns something useless, or a model that nobody trusts enough to act on. A few patterns repeat across plants:

  • No labelled failures to learn from. A predictive model needs to have seen things break in order to recognise the run-up to a breakdown. Many machines simply do not fail often enough, or their failures were never recorded in a structured way, so there is nothing to train against.
  • Alerts with nowhere to go. A model that flags an asset three weeks before failure is only useful if maintenance can be scheduled, parts are in stock, and someone is accountable for acting. Without that, predictions become noise people learn to ignore.
  • Alarm fatigue. A model tuned to catch every failure will also cry wolf. After a handful of false alarms that triggered an unnecessary teardown, technicians stop trusting it, and the project dies socially long before it dies technically.
  • The wrong asset. Teams often start with the machine that broke down most spectacularly last year rather than the one where prediction is both feasible and economically worthwhile. Those are seldom the same asset.

None of these are AI problems. They are readiness problems, and they are knowable in advance.

Required data quality and event history

Predictive maintenance is a supervised problem in disguise: you are trying to predict an event, and to do that the event has to have happened, repeatedly, in a form a model can read. The hard requirement is not just sensor data. It is sensor data paired with a trustworthy history of what went wrong and when.

Concretely, you need maintenance and failure records that are time-stamped, attributed to a specific asset, and consistent enough to be aligned with the sensor stream. A maintenance log full of free-text notes like "fixed it" with no timestamp is almost worthless for training. You also need enough failure events of the same type to learn a pattern; a single dramatic breakdown does not generalise.

Just as important is the rhythm of your data. If a bearing degrades over weeks but your vibration readings are recorded once a shift, you may be sampling too coarsely to see the run-up. If your records have long gaps, were reset when a control system was replaced, or mix several machine configurations under one asset ID, the history is unreliable in ways that are easy to miss until a model is confidently wrong. Honest data archaeology, asking how each field was actually captured and how often it lies, is unglamorous but decisive.

Sensor coverage

You cannot predict a failure mode you cannot observe. Each type of failure has tell-tale signals: bearing wear shows up in vibration and sometimes temperature, electrical faults in current signatures, lubrication problems in temperature and acoustic emissions. If the relevant sensor is not installed, or is installed in the wrong place, no amount of modelling recovers the missing signal.

This is where plants discover that their existing instrumentation was built for control and safety, not for prognostics. A temperature sensor that trips a shutdown is not the same as one that captures the slow thermal drift preceding a fault. Before committing to a model, map the failure modes you actually care about and check, mode by mode, whether you are measuring the quantity that precedes them, at a useful sampling rate, with sensors placed close to the source.

Sometimes the right move is to add instrumentation and collect a few months of data before any modelling begins. That is a legitimate and often unavoidable phase. Treating it as wasted time is a mistake; treating it as optional is a bigger one.

A baseline maintenance process must exist first

Predictive maintenance is a layer on top of a functioning maintenance organisation, not a substitute for one. If you do not yet have a reliable asset register, a computerised maintenance management system that people actually use, defined failure codes, and a parts and scheduling process that can respond to a warning, then a prediction has no path to action.

This is the most common reason to slow down. A model that predicts a failure two weeks out is worthless if it takes four weeks to get the part and there is no slot to do the work. The discipline of running good preventive maintenance, recording outcomes consistently, and closing the loop between detection and repair is exactly the discipline that later produces the clean event history a model needs. Plants that invest in that baseline first tend to find their eventual predictive project both easier and far more valuable, because the surrounding machine is ready to use what the model produces.

Put bluntly: if maintenance today is reactive and undocumented, the first project is not AI. It is process.

The economic threshold for ROI

Predictive maintenance only pays when the cost of an unexpected failure is high enough to justify the cost of predicting it. That sounds obvious, yet many pilots target assets where it simply is not true.

The economics tilt in your favour when several conditions line up: unplanned downtime on the asset is genuinely expensive, whether through lost production, scrap, safety exposure, or knock-on effects across a line; failures are not so rare that you will never gather data, nor so frequent and cheap that a simple schedule already handles them well; and the failure gives enough warning to act before it happens. An asset that fails instantly and catastrophically offers no lead time to exploit, no matter how good the model.

It also matters how many similar assets you have. Instrumenting and modelling a one-of-a-kind machine carries the full engineering cost for a single payoff. A fleet of identical pumps lets you spread that cost and reuse the model, which is why fleets are usually where predictive maintenance earns its keep first. Run this arithmetic before, not after, you build anything. If the numbers do not clear the threshold, the right answer is to choose a different asset or a different approach.

Pilot design

When the conditions are met, a good pilot is designed to be honest rather than impressive. The point is to learn whether prediction works on your assets, under your operating conditions, with your people in the loop.

  • Pick one or two failure modes on a well-instrumented, economically meaningful asset. Resist the urge to predict everything at once.
  • Define success up front in operational terms. Decide what lead time is useful, what false-alarm rate technicians will tolerate, and what action a prediction triggers, before you see any results.
  • Run in shadow mode first. Let the model make predictions while maintenance continues as usual, and compare what it would have caught against what actually happened. This builds trust without risking production.
  • Keep a human in the loop. A prediction should inform a maintenance planner, not silently dispatch a work order. Trust is earned by being right in front of the people who will act on it.
  • Plan for the boring outcome. A pilot that shows prediction is not yet feasible because of data or sensor gaps is a success: it saved you a far more expensive failure at full scale.

An AI Readiness Audit before the pilot answers most of the readiness questions on paper, so you spend pilot budget testing the genuine uncertainty rather than discovering basic gaps the hard way.

Alternatives when predictive maintenance is premature

One alternative sits outside this list because it is a different project rather than a smaller version of this one: computer vision quality inspection. It often clears the readiness bar sooner, because defect images are easier to collect than labelled failure histories, and it is frequently the higher-return first pilot on the same shop floor.

If the readiness checks come back negative, do not abandon the goal of fewer breakdowns. Step down to an approach that fits the data and process you have today, and use it to build toward prediction later.

  • Scheduled (preventive) maintenance. The default for a reason. Servicing on a sensible interval, informed by manufacturer guidance and your own failure history, eliminates a surprising share of breakdowns with no model at all. It is also how you generate the consistent records a future model will need.
  • Condition-based maintenance. Act on simple, interpretable thresholds: vibration above a level, temperature past a limit, runtime hours since last service. This captures much of the practical benefit of prediction, is easy for technicians to trust because the trigger is visible, and requires no failure history to start.
  • Anomaly detection. When you have sensor data but few or no labelled failures, an unsupervised model can flag readings that deviate from an asset's normal behaviour without claiming to predict a specific failure. It is less precise than true prediction, but it works with the data you have and surfaces problems worth a human look.

These are not consolation prizes. For many plants, condition-based and anomaly-based approaches deliver the bulk of the available value at a fraction of the cost and risk, while the instrumentation and records they generate quietly assemble the foundation that genuine predictive maintenance will one day stand on. The honest sequence is almost always: get the process right, observe the assets, then predict. That sequencing question is exactly what an honest readiness assessment is for.

Frequently asked questions

How many historical failures do you need before predictive maintenance is viable?

Enough examples of the specific failure mode you want to predict, which usually means dozens rather than a handful. A machine that has failed twice in five years cannot support a model, no matter how much sensor data surrounds it. Rare failures are a reliability engineering problem, not a modelling one.

Is condition monitoring the same as predictive maintenance?

No, and conflating them causes most of the disappointment. Condition monitoring reports the current state against thresholds and is often the right answer on its own. Predictive maintenance forecasts a future failure, which needs failure history that most mid-sized plants have not accumulated.

What should we do first if predictive maintenance turns out to be premature?

Start recording failures properly, with cause, date, machine, and downtime. That record is the asset the model will eventually need, it takes a year or two to accumulate, and building it costs almost nothing beyond discipline. Meanwhile threshold alerting captures much of the near-term value.

Does predictive maintenance pay off if the line is not capacity constrained?

Usually not. Avoided downtime only converts to money if the line was going to be running and selling. On a line with slack, the saving is maintenance efficiency rather than output, which is a much smaller number and often below the threshold that justifies the project.

Is predictive maintenance your right first project?

A Manufacturing AI Readiness Audit checks whether your data and process are ready, and finds a higher-ROI pilot if they are not.

Book a Manufacturing AI Readiness Audit
SAI

Sitnik AI

Applied AI consultancy for healthcare and manufacturing teams. Led by a PhD computer scientist and former CTO, with research in medical imaging and production AI systems.

Ready to Get Started?

Book a free consultation to discuss your AI project.