ARTICLE 06 · MAINTENANCE · 2026-04-26

Let AI warn you before machines break down: how to cut downtime and emergency repair costs

Technicians don't need more charts. They need a warning that arrives early enough, gives enough reasons and says what to check next. That is the line between predictive maintenance and an expensive alarm.

Let AI warn you before machines break down: how to cut downtime and emergency repair costs
The short version
  • Emergency repairs cost several times more than planned ones, because they come bundled with line stoppages and late deliveries.
  • Start with a handful of failure types you can “see coming” before you try to predict everything at once.
  • A warning that nobody is responsible for following up is worthless. It has to become a work order with a real person and a real date and time.

Start with failures you can “see coming and act on”

Not every failure can be predicted. Some parts break instantly with no signal at all. The place to start is the kind of failure your technicians can already call: “When it sounds like this, it'll go within two weeks.” That is exactly the kind of knowledge a system can learn from.

For the development team · Technical detail

Rank asset criticality by safety, quality, throughput and cost, then choose failure modes that have early symptoms and enough lead time to act on, such as bearing vibration, high temperature or changes in motor current. Failures that happen instantly without any signal may be better handled with redundancy or spare stock than with AI.

Do a practical FMEA: what is the symptom, where is it measured and under what load, what needs checking when a warning fires, who decides to stop the machine, and how long does it take to get the spare part? Don't choose an asset because it has lots of sensors. Choose it because a warning can change the work plan.

The goal: turn emergency work into planned work. Predicting failure dates that look precise on a dashboard is beside the point.

Get condition data and maintenance data to speak the same language

For the development team · Technical detail

Sensor data needs timestamps, units and context such as speed, load, product, start/stop and ambient temperature. Work order history has to record the symptom, cause, findings, parts and actual time spent. A note that just says “repaired” can't be used to teach the system. Synchronize the clocks of the PLC, gateway and CMMS before any analysis.

DataQuestion it answersCommon mistake
Vibration/TemperatureWhen did the condition change?Not linked to load
Alarm/PLC StateWhat mode was the machine in?Mismatched timestamps
Work OrderWhat was actually found and fixed?No failure code
Production ContextWhich product was running when the change happened?Changeovers not recorded

If failures are rare, use anomaly detection to flag deviations from normal conditions and have a technician inspect them. Don't translate a score into a failure date without evidence. Show the trend and the context so people can make the call.

Design alerts that turn into work orders

An alert that nobody is responsible for following up becomes background noise within two weeks. Every time the system issues a warning, it has to be clear who must do what, by when, and what happens if they don't.

The result you want: fewer emergency repairs, more planned maintenance and shorter stoppages
The result you want: fewer emergency repairs, more planned maintenance and shorter stoppages

Split alerts into Watch, Inspect and Act, each with its own criteria and SLA. An alert must state the asset, the symptom, when it started, how severe it is, the evidence and how to inspect it. When a technician confirms or rejects it, record the reason so the threshold can be adjusted. The system has to connect to the CMMS, because a message dropped into a group chat simply disappears.

For the development team · Technical detail

Start in shadow mode and count every alert and every breakdown, including the ones the system didn't warn about. When precision is low, the team burns out on alarms; when recall is low, the risk is still there; when lead time is too short, nobody can plan. So you have to look at all three together with downtime and maintenance cost.

Before alerts are allowed to affect the production plan

  • Technicians understand the evidence and can link it to an inspection
  • Production can see the cost of stopping against the cost of the risk
  • There is a clear final decision-maker
  • There is a procedure for when a sensor drops out or gives skewed readings
  • Spare parts and resources are actually available to respond

Measure reliability, with model accuracy as only one part of it

For the development team · Technical detail

Track Unplanned Downtime, Planned Work Ratio, MTBF, MTTR, Emergency Purchases, Overtime, False Alarms, Misses and Actionable Lead Time. Compare against similar assets or a baseline adjusted for running hours, and include the cost of checking alerts and maintaining sensors in your costs.

For the development team · Technical detail

When you scale, build a template for each failure mode instead of copying one threshold to every machine. Check calibration, data drift and changes in the operating envelope every month at first, then according to risk once the system is stable. A new model version has to pass a test on historical data and a shadow run before it replaces the old one.

Predictive maintenance succeeds when technicians have fewer emergency jobs and can plan better, even if the model never once predicts “failure in 7 days”. A signal that can be explained and checked in time is worth more than a flashy prediction the team doesn't believe.

Set an alert budget before you set thresholds

Agree with the team how many alerts needing inspection they can handle in a week. If the system sends more than that, it will be ignored even if it is accurate enough. Start with fairly strict thresholds, record the misses, and adjust gradually based on the cost of failure. This keeps trust far better than firing off alarms and asking the technicians to be patient.

Choose failures that give enough advance warning to act in time
Choose failures that give enough advance warning to act in time
Hypothetical case: A group of motors showed high vibration every time they ramped up speed. The original model sent warnings every morning, so the technicians stopped looking. Once operating state was added and comparisons were limited to periods of steady load, the alerts dropped to just a few, and one of them caught an alignment problem developing before it forced an emergency stop. The turning point was getting the context right. A new model had nothing to do with it.

An alert card needs 5 things

The asset, the symptom, the trend evidence, the operating context and the next inspection, with when it should be done. Without that last item, the alert still isn't connected to maintenance work.

Tip: Every time a technician marks a false alarm, give them four to six reasons to choose from instead of a long text box. You'll get consistent data for tuning the system without piling on paperwork.

DNA MAKER · SOLUTION BLUEPRINT

From knowledge to a problem-solving system that works in practice

The underlying problem

A predictive model is worthless if its alerts don't arrive within the lead time, the spare parts aren't there, or the alerts aren't linked to the work orders technicians actually use.

A step-by-step approach

  1. Choose assets and failure modes based on criticality and actionable lead time
  2. Align the data between sensors, operating state and repair history
  3. Build alert cards, thresholds, feedback and the CMMS workflow before scaling

Get alerts all the way to the technician as work that can be handled

Predictive maintenance doesn't end when the model finds an anomaly, because technicians still need to know what to inspect, when, and how far the data can be trusted. DNA Maker works with the client's reliability engineers and maintenance team to turn failure modes, operating context and response steps into alert cards and workflows that get used in practice. Diagnosing the machines stays with the experts. We help make sure their knowledge arrives together with the sensor data and work orders, which cuts down on switching between screens and on alerts that nobody owns.

The solution might be a condition monitoring portal, a mobile inspection app or an AI agent that connects time-series data to the CMMS, creates inspections, collects feedback from technicians and tracks the health of the data and the model. DNA Maker helps design the data flow, shop-floor UX, APIs, edge/cloud architecture, software development and model monitoring together with the client's OT and IT teams. If your organization already has machine data that still isn't turning into maintenance work, we're ready to help map the gaps from signal to decision and build a prototype technicians can try before you invest in scaling up.

Software engineering glossary

This table is a shared vocabulary for executives, process owners and the development team, so nobody reads a term differently. There's no need to memorize it. Read the meaning, the example and the question on the right, because those questions often reveal scope, risks and hidden costs before development begins.

TermWhat it isA simple exampleWhat to ask the development team
Anomaly DetectionFinding values that differ from the normal patternSpotting abnormal vibration compared with the same load rangeWhich kinds of anomaly actually lead to action?
CMMSA system for managing maintenance workCreating a work order from an alertHow will alerts become work orders without creating duplicate jobs?
ThresholdThe boundary value that triggers an alertTemperature above the limit for 10 minutes straightAre thresholds set by risk or by averages, and who approves them?
Time-series DataData arranged in time orderVibration readings every secondDo the times, units and operating states line up?
Model MonitoringTracking whether a model is still performing wellAlerting when the pattern of sensor data changesWho gets notified when model or data quality drops?
Try this tomorrow: Pick one group of critical assets, write down the failure modes and the actionable lead time, and then check which data supports them. Don't start by installing more sensors.