Every manufacturer we talk to has heard the predictive maintenance pitch: an AI bot watches your machines and tells you a bearing is about to fail before it takes down the line. Almost none of them have asked the harder question first — what does that bot actually need to see to make that call? Predictive maintenance isn't a model you bolt onto an existing ERP; it's a data pipeline, and the pipeline is where most projects quietly fail before the AI ever gets involved.
The uncomfortable truth is that a predictive-maintenance bot is only as good as the signal it's fed, and most shop floors in India aren't yet generating that signal in a form worth feeding. Before anyone commits budget to "AI-powered maintenance," it's worth being specific about what the bot needs, where that data actually comes from, and what breaks when the inputs are missing.
What "predictive" actually requires — and what most plants have instead
Predictive maintenance works by comparing a machine's current behavior against its own historical pattern and flagging drift before it becomes failure — vibration creeping up, temperature running hotter for the same load, cycle time slowing on the same job. That comparison needs three things: continuous sensor or usage data (not a monthly service log), a maintenance history tied to the specific asset, and enough historical failures logged with cause to teach the model what "about to fail" actually looks like on this exact machine.
Most plants we work with have the opposite: a paper or spreadsheet maintenance log updated after the fact, no continuous sensor feed, and asset history scattered across whichever technician happened to note it down. That's not a predictive-maintenance gap — it's a data-capture gap, and it has to close first. This is exactly the distinction we walk clients through under industry-specific ERP work: a generic "asset register" field isn't the same as structured, timestamped condition data a model can actually learn from.
Where the data has to live before the bot can use it
A predictive bot can't reason over data that's still sitting in a technician's notebook or a disconnected SCADA screen. It needs the maintenance history, the work orders, and the usage data joined to the same asset record inside the ERP — otherwise the model is guessing from a fraction of the picture. In practice this means machine-usage counters or IoT sensors feeding readings on a schedule tight enough to catch drift (hourly, not monthly), every completed and skipped maintenance task logged against the asset, and every past failure tagged with a root cause rather than just "fixed."
Once that's in place, the same infrastructure that feeds a real-time dashboard is what feeds the maintenance bot — they're pulling from the identical live data, not two separate systems. That's also why bolting predictive maintenance onto a legacy ERP that only refreshes reports overnight rarely works: by the time the batch job runs, the drift the bot needed to catch already happened.
What the bot does with a false positive — and why that matters more than accuracy
The metric everyone asks about is accuracy — how often the bot is right. The metric that actually determines whether the system survives its first year is what happens when it's wrong. A predictive-maintenance bot that flags a machine for inspection and turns out to be wrong costs an inspection. A bot that's silent when it should have flagged something costs an unplanned line stoppage. Those aren't symmetric costs, which is why the threshold for raising a flag should be tuned toward "cheap to check" rather than "rarely wrong" — the same guardrail logic that applies to any Agentic AI system making a call with real operational consequences.
In practice, that means the bot doesn't get to shut down a line or schedule downtime on its own — it raises a flag, a maintenance supervisor reviews the evidence (the specific reading that triggered it, the trend line, the comparable past incidents), and the human decides the response. The AI's job is catching the signal early enough that there's still a choice to make; the decision itself stays with someone who knows the machine and the production schedule.
Start with one machine, not the whole plant
The manufacturers who get real value from predictive maintenance almost never start plant-wide. They pick the one or two assets where unplanned downtime is genuinely expensive — a bottleneck machine, one with a history of costly failures — get the sensor and logging pipeline solid there, and prove the model catches real drift before it becomes a breakdown. Only then does it make sense to extend the same pipeline to the rest of the floor. Rolling it out everywhere at once, before the data pipeline is proven on a single asset, is the most common way these projects stall: the model gets blamed for bad predictions when the actual problem was thin, inconsistent input data across too many machines at once.
If your maintenance data still lives in a notebook or a spreadsheet nobody reconciles against the ERP, that's the real starting point — not a model, not a dashboard, not a bot. Get the asset history and the usage data structured and current first, and the predictive layer becomes a much smaller, much more honest project than the pitch decks make it sound.