Building a Condition Monitoring Programme: From Criticality Ranking to the First Catch
Author
Yousif Atabani
Date Published

Disclaimer: Research and analysis by the engineering team. Sources referenced below.
Most condition monitoring programmes do not fail because the technology is wrong. They fail because someone bought the technology first and designed the programme second.
The pattern is depressingly consistent. A plant suffers an expensive failure, management approves a budget, an analyser arrives, data collection begins on everything within reach, and eighteen months later there is a hard drive full of readings, no documented catch, and a finance director asking what the money bought. The instrument was never the problem. The missing pieces were a criticality ranking, a match between technique and failure mode, a data discipline that makes trends trustworthy, someone competent to read the results, and a loop that turns findings into work orders.
A condition monitoring programme is those five pieces, in that order; the hardware comes last. This article walks through building one from scratch, from ranking your assets through technique selection, data discipline, staffing and KPIs to a realistic twelve-month rollout.
The prize is worth the discipline. The US Department of Energy's Operations and Maintenance Best Practices guide puts the savings of a functional predictive programme at 8% to 12% over a purely preventive one and 30% to 40% over run-to-failure, with typical results including a 25% to 30% reduction in maintenance costs and 35% to 45% less downtime. Averages across many facilities, but the direction is not in dispute.
Criticality Ranking: The Foundation Everything Sits On
You cannot monitor everything, and you should not try. The first deliverable of any programme is a ranked asset list, built in workshops simple enough to run in a week: score each asset for consequence of failure and likelihood of failure, and multiply the two.
Consequence is what actually happens when the machine stops: does production halt or does a standby take over, is there a safety or environmental exposure, what does the repair cost, and what is the lead time on the long-delivery parts. A boiler feed pump with no installed spare and a five-month lead time on a rotor scores at the top; a fully spared transfer pump scores near the bottom, however large it is. Likelihood is how prone the asset is to failing: age, duty, service history, operating environment and known bad actors.
Multiply and sort. The output usually surprises somebody in the room, which is the point: the workshops surface disagreements between operations, maintenance and engineering that have run silently for years. For a typical industrial site the top tier comes out at a few dozen assets, perhaps 20 to 50 machines whose failure stops the plant or endangers people. That tier is where the programme starts. Not the full register of four hundred machines; the top few dozen, done properly.
Keep the failure data behind the likelihood scores honest. ISO 14224 defines a standard taxonomy for equipment boundaries, failure modes and maintenance records, and borrowing its structure, even loosely, means the ranking rests on comparable data rather than on whoever shouts loudest in the workshop.
Matching Technique to Failure Mode
The second design decision is which technique watches which asset, and each technique sees a different family of failure modes. Choosing by failure mode, rather than buying one technology and applying it everywhere, is what ISO 17359, the general framework standard for condition monitoring and diagnostics, formalises: identify the likely failure modes first, then select parameters that actually change as those failures develop.
Vibration analysis is the primary tool for rotating machinery: unbalance, misalignment, looseness, bearing defects and gear problems all announce themselves weeks or months before functional failure. We have covered the technique in depth, including what each fault looks like in a spectrum, in our guide to vibration analysis for rotating equipment, so this article will not repeat the spectral detail. For programme design the point is simpler: every rotating asset in your top tier gets vibration monitoring, at an interval shorter than the time its dominant failure mode takes to develop.
Infrared thermography covers failure modes that do not vibrate: loose and corroded electrical connections, overloaded conductors, failing breakers, blocked cooling passages, passing steam traps and valve leak-through. A thermographic survey of the electrical distribution is often the fastest payback in the whole programme, because a hot connection found early costs a torque wrench and an hour, while the same connection left alone becomes a burned busbar or a switchgear fire. Field practice grades severity by temperature rise over a similar component under similar load: a few degrees is logged and rechecked, tens of degrees is scheduled promptly, and around 15 degrees or more is treated as urgent.
Ultrasound listens in the 20 kHz to 100 kHz range where the earliest bearing distress lives. Its two headline uses are compressed air and gas leak detection, where surveys routinely find leakage worth 20% to 30% of compressor output, and bearing monitoring, where a rise of around 8 dB over baseline typically means a lubrication problem and larger rises the start of physical damage. Ultrasound-guided greasing, stopping when the sound level drops, prevents both under- and over-lubrication; over-greasing kills at least as many bearings as neglect does.
Motor current signature analysis reads the motor's supply current as a diagnostic signal and finds broken rotor bars, air-gap eccentricity and certain stator problems that are hard to distinguish mechanically. It needs access only to the panel, which suits submerged, sealed and otherwise unreachable drives.
Performance monitoring is the least glamorous technique and the only one that catches efficiency failures: a pump whose wear ring has opened up, a heat exchanger fouling towards its pressure-drop limit, a compressor drifting off its curve. Flow, head, differential pressure, temperatures and power draw, trended against the original curve, show degradation that produces no vibration, no heat signature and no debris. On energy-intensive assets this is frequently where the money is.
Oil analysis deserves its own section, because it is the technique most often bought and least often understood.
What an Oil Analysis Report Actually Tells You
An oil analysis programme is cheap, a modest per-sample lab fee, and answers three questions: is the oil fit for service, is something contaminating it, and is the machine wearing abnormally. A standard report carries four families of numbers.
Wear metals, in parts per million from spectrometric analysis, identify which component is shedding material. Iron points to gears, shafts and cylinder walls; copper to bushings, thrust washers and cooler cores; lead and tin to white-metal journal bearing overlays; chromium to rolling element bearings and piston rings; aluminium to pistons and certain thrust components. One elevated reading means little. A rising trend in one metal, sample after sample, is a component announcing its decline.
Viscosity is the oil's most important property, and a change of more than about 10% from grade nominal in either direction is a flag: climbing usually means oxidation, falling usually means fuel dilution or shear breakdown of the viscosity improver. Either way the oil film the bearings depend on is no longer the one the designer assumed.
Water is quietly lethal. Even a few hundred parts per million measurably shortens rolling element bearing life, and free water strips additives, feeds rust and can flash to steam in a hot bearing. Know your system's limit and treat any upward trend as an ingress investigation, not an oil change.
Particle count, reported as an ISO cleanliness code such as 18/16/13, counts particles in three size bands per millilitre and is the leading indicator in hydraulic and circulating systems: counts climb before wear metals do, because the particles doing the damage arrive before the damage. A system running two codes above target is not dirty oil, it is a wear rate multiplier.
The design lesson: oil analysis only trends if sampling is disciplined. Same point, same method, machine running and warm, drawn from mid-flow rather than the sump bottom. A pristine sample from the wrong point tells you about the wrong oil.
Not sure which techniques your critical assets actually justify? Our maintenance and asset management team builds criticality rankings and technique matrices as a scoped exercise, and will tell you plainly where monitoring is not worth the money.

Techniques are chosen against failure modes, not bought as a set. Each sees something the others cannot. Source: MIMAH engineering practice; ISO 17359:2018.
The Data Discipline That Makes a Trend Worth Trusting
Every technique above produces numbers, and the numbers are only evidence if they are comparable over time. Four disciplines make them comparable.
Baselines first. No reading means anything in isolation. The first months of the programme exist to establish what normal looks like for each machine: its vibration signature at known load, its oil chemistry when healthy, its thermal image on an ordinary afternoon.
Fixed points and fixed conditions. Measurements are taken at permanently marked locations, with the same mounting and method, at recorded and comparable operating conditions. A machine measured at half load and at full load is two different data sets, and mixing them manufactures phantom trends.
Intervals set by criticality. The top tier earns monthly attention or permanent sensors, the middle tier quarterly, the remainder annual or on condition. The test for any interval is brutal and simple: can the failure mode develop from first detectable to functional failure inside the gap between two readings? If yes, the interval is wrong or the asset needs online monitoring.
Alarm bands per machine, not per textbook. Generic limits produce false alarms on some machines and silence on degrading ones. Standards supply the framework and starting values, and we have written separately on how the ISO 20816 vibration zone limits actually work, but the mature programme sets alert and alarm thresholds from each machine's own baseline and escalates on rate of change, not just absolute level.
Who Reads the Data: Staffing, Training and the Outsourcing Question
The instrument is the cheap part of the programme. The expensive part, and the part that actually produces catches, is the person who looks at a trend and makes a call.
There is a recognised competence ladder for this. ISO 18436-2 defines four categories of vibration condition monitoring personnel: Category I collects data correctly on established routes, Category II sets up measurements and performs standard spectrum diagnosis, Category III designs programmes and handles the difficult diagnostics, and Category IV covers advanced analysis a plant rarely needs on staff. Parallel schemes exist for thermography, lubricant analysis and ultrasound under the same standard family.
For a plant starting from scratch, the realistic model is a hybrid. Train one or two of your own technicians to Category I or II so that collection is owned in-house and done consistently, because an outsourced collector is always the first budget line cut. Then buy the analysis: a contract analyst reviewing your data remotely costs a fraction of a Category III salary and has seen a thousand machines rather than forty. Bring analysis in-house only when machine count and finding rate justify a full-time role, which for many mid-sized plants is never, and that is a perfectly sound end state.
The one non-negotiable is a named owner. A programme that belongs to everyone belongs to no one, and the data stops being looked at within a year.

The ISO 18436-2 competence ladder, and the hybrid staffing model that works for a plant starting from scratch: own the collection, buy the analysis. Source: ISO 18436-2:2014; MIMAH engineering practice.
Closing the Loop: Finding, Job, Verification
A finding that does not become a work order is a rumour. The work-order loop is short and it must be complete.
The finding is written against the asset with the evidence attached: the trend, the spectrum, the thermal image, the lab report. It states a suspected fault, a severity, and a recommended window, for example "outer race bearing defect, early stage, replace within eight weeks at next planned stop".
The job is planned and executed inside that window, through the normal planning process. Condition monitoring does not bypass planning; it feeds planning with better lead time than any other input the planner has. This is the practical mechanism behind the shift we described in preventive versus predictive maintenance: scheduled intrusive work shrinks because condition data keeps proving it unnecessary, and the work that remains targets real faults.
The verification is the step most programmes skip and the step that builds all the credibility. After the repair, measure again. Either the fault signature is gone, confirming the diagnosis, or it is not, meaning the diagnosis was wrong or the repair introduced a new problem, and both are things you want to know within days rather than at the next failure. Every verified catch becomes a documented case: what was found, what was done, what a run-to-failure would have cost. That file is the programme's defence at budget time.
KPIs That Prove the Programme Is Working
Condition monitoring is one of the easier maintenance investments to defend, provided you count the right things from day one.
Confirmed catches is the headline number: findings verified as real faults at repair. Track the count and the hit rate; a programme calling faults that teardowns keep failing to find has an analysis problem.
Avoided cost per catch is the number finance actually reads: for each catch, estimate the delta between the planned repair as executed and the plausible unplanned failure, including secondary damage, downtime at contribution margin, freight and overtime. Be conservative and show the working. Even counted cautiously, programmes routinely show the order-of-magnitude returns the DOE guide reports.
The PM-to-CM shift shows the programme changing how the plant works: the share of maintenance hours spent on time-based intrusive tasks should fall as condition-based tasks replace them, and unplanned corrective hours should fall with it. Route compliance, the percentage of scheduled readings taken on time, is the leading indicator that predicts all the others: when compliance slips, catches stop about six months later.
Emergency work order rate and monitored-asset availability, trended over a couple of years, close the case. None of this needs sophisticated software; a spreadsheet maintained honestly beats a dashboard nobody reconciles.
A Realistic Twelve-Month Rollout
Months 1 and 2: rank and design. Run the criticality workshops, agree the top tier, and map technique to failure mode for each asset in it. Decide the staffing model. Buy nothing yet beyond the lab contract.
Months 3 and 4: baseline. First vibration surveys at recorded operating conditions, first oil samples to the lab, first thermographic survey of the electrical distribution. Mark measurement points permanently. Expect no findings; expect a baseline library.
Months 5 and 6: second pass and limits. Repeat the rounds, set alert and alarm bands per machine from the two data sets, and start Category I training for your collectors. The first genuine findings usually surface here, often from thermography and oil before the vibration trends mature.
Months 7 to 9: run the loop. Routes on schedule, findings written as work orders, repairs verified by post-repair measurement. Start the catch file. Ultrasound and motor current analysis join where the failure-mode map calls for them.
Months 10 to 12: prove and expand. Compile the KPI pack, present it, then extend to the second criticality tier and move inaccessible or fast-failing points to permanent sensors. By month twelve a well-run programme on a few dozen critical assets has typically paid its first year's cost out of one or two catches.
Inherited a plant with no monitoring history at all? Our industrial engineering services team carries out baseline condition surveys, criticality studies and root cause analysis on rotating and electrical plant, and hands you the asset ranking and technique matrix as working documents rather than a report that sits on a shelf.

A realistic twelve-month rollout. The usual failure is starting too large: four hundred machines, six months of data, no baselines and no findings before the budget is questioned. Source: MIMAH engineering practice.
Frequently Asked Questions
How many assets should a new condition monitoring programme cover? Start with the top criticality tier only, for most sites somewhere between twenty and fifty machines. The limiting resource is not instruments but attention. A tight programme on forty machines that produces documented catches will win the budget to expand; a shallow one on four hundred will produce noise and then be cancelled.
What does a condition monitoring programme cost to start? For a portable-instrument programme on a few dozen assets, the entry cost is a mid-range vibration analyser, a thermal camera, an ultrasound gun, a per-sample lab contract and training for one or two technicians, which typically lands well under the cost of a single unplanned failure of one top-tier machine. Permanent online systems cost more per point and are justified selectively, on assets that are critical, unreachable or fast-failing.
Should we do the analysis in-house or outsource it? Collect in-house, analyse outside, at least for the first years. In-house collection keeps the routes on schedule and builds ownership; contracted analysis buys experience across thousands of machines that no new programme can grow internally. Revisit the split when your finding rate would keep a qualified analyst busy full time.
Which technique should we buy first? The one your criticality ranking and failure modes point to, which for most plants means vibration for the rotating tier plus a lab-based oil programme, since together they cover the majority of high-consequence mechanical failure modes for a modest outlay. A thermographic survey of the electrical system is usually the quickest single payback and needs no permanent infrastructure.
How long before the programme shows results? Baselines take two to three months, trustworthy trends take two readings beyond that, and the first verified catches typically arrive between months five and nine. Anyone promising catches in the first quarter is selling something. The twelve-month mark is the fair point to judge the programme, on catches, avoided cost and route compliance.
Start Narrow, Measure Honestly, Prove It
A condition monitoring programme is not a technology purchase. It is a decision to stop being surprised by your own machines, executed as a ranking, a technique map, a data discipline, a competence plan and a closed loop, in that order. Plants that start narrow and deep build programmes that survive their third budget cycle; plants that start wide and shallow generally do not.
Our team brings four decades of experience in turbine, generator and rotating equipment work to this problem, from steam turbine overhauls in Nigeria to root cause analysis at White Nile Sugar in Sudan, and the lesson from the failures we have investigated is the one this article started with: the machines that surprise people are the machines nobody ranked, baselined or measured.
Ready to build the programme rather than just buy the instruments? Talk to our engineering team. We will help you rank your assets, match techniques to failure modes, set up the routes and the lab contract, and leave you with a programme your own people run.
Sources:
