- Emergency repairs cost several times more than planned ones, because they come bundled with line stoppages and late deliveries.
- Start with a handful of failure types you can “see coming” before you try to predict everything at once.
- A warning that nobody is responsible for following up is worthless. It has to become a work order with a real person and a real date and time.
Start with failures you can “see coming and act on”
Not every failure can be predicted. Some parts break instantly with no signal at all. The place to start is the kind of failure your technicians can already call: “When it sounds like this, it'll go within two weeks.” That is exactly the kind of knowledge a system can learn from.
For the development team · Technical detail
Rank asset criticality by safety, quality, throughput and cost, then choose failure modes that have early symptoms and enough lead time to act on, such as bearing vibration, high temperature or changes in motor current. Failures that happen instantly without any signal may be better handled with redundancy or spare stock than with AI.
Do a practical FMEA: what is the symptom, where is it measured and under what load, what needs checking when a warning fires, who decides to stop the machine, and how long does it take to get the spare part? Don't choose an asset because it has lots of sensors. Choose it because a warning can change the work plan.
Get condition data and maintenance data to speak the same language
For the development team · Technical detail
Sensor data needs timestamps, units and context such as speed, load, product, start/stop and ambient temperature. Work order history has to record the symptom, cause, findings, parts and actual time spent. A note that just says “repaired” can't be used to teach the system. Synchronize the clocks of the PLC, gateway and CMMS before any analysis.
| Data | Question it answers | Common mistake |
|---|---|---|
| Vibration/Temperature | When did the condition change? | Not linked to load |
| Alarm/PLC State | What mode was the machine in? | Mismatched timestamps |
| Work Order | What was actually found and fixed? | No failure code |
| Production Context | Which product was running when the change happened? | Changeovers not recorded |
If failures are rare, use anomaly detection to flag deviations from normal conditions and have a technician inspect them. Don't translate a score into a failure date without evidence. Show the trend and the context so people can make the call.
Design alerts that turn into work orders
An alert that nobody is responsible for following up becomes background noise within two weeks. Every time the system issues a warning, it has to be clear who must do what, by when, and what happens if they don't.

Split alerts into Watch, Inspect and Act, each with its own criteria and SLA. An alert must state the asset, the symptom, when it started, how severe it is, the evidence and how to inspect it. When a technician confirms or rejects it, record the reason so the threshold can be adjusted. The system has to connect to the CMMS, because a message dropped into a group chat simply disappears.
For the development team · Technical detail
Start in shadow mode and count every alert and every breakdown, including the ones the system didn't warn about. When precision is low, the team burns out on alarms; when recall is low, the risk is still there; when lead time is too short, nobody can plan. So you have to look at all three together with downtime and maintenance cost.
Before alerts are allowed to affect the production plan
- Technicians understand the evidence and can link it to an inspection
- Production can see the cost of stopping against the cost of the risk
- There is a clear final decision-maker
- There is a procedure for when a sensor drops out or gives skewed readings
- Spare parts and resources are actually available to respond
Measure reliability, with model accuracy as only one part of it
For the development team · Technical detail
Track Unplanned Downtime, Planned Work Ratio, MTBF, MTTR, Emergency Purchases, Overtime, False Alarms, Misses and Actionable Lead Time. Compare against similar assets or a baseline adjusted for running hours, and include the cost of checking alerts and maintaining sensors in your costs.
For the development team · Technical detail
When you scale, build a template for each failure mode instead of copying one threshold to every machine. Check calibration, data drift and changes in the operating envelope every month at first, then according to risk once the system is stable. A new model version has to pass a test on historical data and a shadow run before it replaces the old one.
Predictive maintenance succeeds when technicians have fewer emergency jobs and can plan better, even if the model never once predicts “failure in 7 days”. A signal that can be explained and checked in time is worth more than a flashy prediction the team doesn't believe.
RELIABILITY TIP · MAKE EVERY ALERT COUNT
Set an alert budget before you set thresholds
Agree with the team how many alerts needing inspection they can handle in a week. If the system sends more than that, it will be ignored even if it is accurate enough. Start with fairly strict thresholds, record the misses, and adjust gradually based on the cost of failure. This keeps trust far better than firing off alarms and asking the technicians to be patient.

An alert card needs 5 things
The asset, the symptom, the trend evidence, the operating context and the next inspection, with when it should be done. Without that last item, the alert still isn't connected to maintenance work.
Tip: Every time a technician marks a false alarm, give them four to six reasons to choose from instead of a long text box. You'll get consistent data for tuning the system without piling on paperwork.
From knowledge to a problem-solving system that works in practice
The underlying problem
A predictive model is worthless if its alerts don't arrive within the lead time, the spare parts aren't there, or the alerts aren't linked to the work orders technicians actually use.
A step-by-step approach
- Choose assets and failure modes based on criticality and actionable lead time
- Align the data between sensors, operating state and repair history
- Build alert cards, thresholds, feedback and the CMMS workflow before scaling
Get alerts all the way to the technician as work that can be handled
Predictive maintenance doesn't end when the model finds an anomaly, because technicians still need to know what to inspect, when, and how far the data can be trusted. DNA Maker works with the client's reliability engineers and maintenance team to turn failure modes, operating context and response steps into alert cards and workflows that get used in practice. Diagnosing the machines stays with the experts. We help make sure their knowledge arrives together with the sensor data and work orders, which cuts down on switching between screens and on alerts that nobody owns.
The solution might be a condition monitoring portal, a mobile inspection app or an AI agent that connects time-series data to the CMMS, creates inspections, collects feedback from technicians and tracks the health of the data and the model. DNA Maker helps design the data flow, shop-floor UX, APIs, edge/cloud architecture, software development and model monitoring together with the client's OT and IT teams. If your organization already has machine data that still isn't turning into maintenance work, we're ready to help map the gaps from signal to decision and build a prototype technicians can try before you invest in scaling up.
SOFTWARE ENGINEERING GLOSSARY
Software engineering glossary
This table is a shared vocabulary for executives, process owners and the development team, so nobody reads a term differently. There's no need to memorize it. Read the meaning, the example and the question on the right, because those questions often reveal scope, risks and hidden costs before development begins.
| Term | What it is | A simple example | What to ask the development team |
|---|---|---|---|
| Anomaly Detection | Finding values that differ from the normal pattern | Spotting abnormal vibration compared with the same load range | Which kinds of anomaly actually lead to action? |
| CMMS | A system for managing maintenance work | Creating a work order from an alert | How will alerts become work orders without creating duplicate jobs? |
| Threshold | The boundary value that triggers an alert | Temperature above the limit for 10 minutes straight | Are thresholds set by risk or by averages, and who approves them? |
| Time-series Data | Data arranged in time order | Vibration readings every second | Do the times, units and operating states line up? |
| Model Monitoring | Tracking whether a model is still performing well | Alerting when the pattern of sensor data changes | Who gets notified when model or data quality drops? |
