Electronics Guide

Predictive and Preventive Methods

Predictive and preventive maintenance methods shift maintenance from a reactive activity that responds to failures to a planned activity that anticipates them. These methods combine monitoring technologies, analytical models, and systematic maintenance planning to maximize equipment availability while controlling maintenance cost and unplanned downtime.

The approaches differ in what triggers an intervention. Preventive maintenance works on fixed intervals of calendar time, operating hours, or usage cycles, servicing or replacing components before they typically wear out. Condition-based maintenance acts on the present measured state of an asset. Predictive maintenance extends that idea by trending the measurements and scheduling work only when the trend shows that service is genuinely needed. Prognostics goes one step further and estimates how much useful life remains, converting a warning into a planning horizon.

This part of the reliability-engineering body of knowledge gathers the monitoring technologies, analytical models, and management practices that make proactive maintenance work, together with the warranty and field-service analysis that closes the loop between fielded performance and future maintenance decisions. The same logic that governs a pump bearing governs a power converter or a server rack; only the measured parameters and the failure physics change.

Articles in This Category

The Maintenance Strategy Spectrum

Maintenance strategies form a spectrum defined by what prompts action. Reactive, or run-to-failure, maintenance defers all work until a component breaks. It carries no monitoring cost but the highest exposure to collateral damage, secondary failures, overtime labor, and unplanned downtime. Preventive maintenance trades some of that exposure for predictability by acting on a schedule, at the cost of servicing components that may still have useful life. Condition-based maintenance acts on measured state, and predictive maintenance and prognostics use trends in that state to forecast a failure and act just in time.

Scheduled replacement helps only when failure is age-related. If the hazard rate of a failure mode is constant, replacing a component at a fixed age exchanges an old item for a new one with exactly the same instantaneous risk, adding cost and introducing fresh infant-mortality risk without improving reliability. In the language of life-data analysis, an age-based replacement task can be justified only when the Weibull shape parameter exceeds one, meaning the hazard rate is rising. The bathtub curve and the probability distributions used to fit life data make this test explicit.

How often that test is passed surprised the industry that first ran it. The 1978 study by F. Stanley Nowlan and Howard F. Heap of United Airlines, prepared for the U.S. Department of Defense and now regarded as the founding document of reliability-centered maintenance, sorted commercial-aircraft failure modes into six hazard-rate patterns. Only three of them, accounting for 4 percent, 2 percent, and 5 percent of the items studied, showed a distinct age-related wear-out region; the remaining 89 percent did not, and 68 percent of items followed an infant-mortality pattern in which the hazard rate was highest when the item was new. The practical conclusion was not that scheduled overhaul is useless but that it is applicable to a minority of failure modes, and that applying it indiscriminately consumes resources while sometimes injecting the very defects it was meant to prevent.

Most mature programs therefore blend strategies, reserving instrumentation-intensive techniques for critical or expensive assets and applying simpler time-based plans, or deliberate run-to-failure, where failures are inexpensive and consequences are contained. The U.S. Department of Energy's Operations and Maintenance Best Practices guide offers one widely quoted target mix for a mature program: less than 10 percent reactive work, 25 to 35 percent preventive, and 45 to 55 percent predictive. The proportions are guidance rather than a specification, and the right balance depends on asset criticality, failure physics, and the cost of monitoring.

Selecting Tasks with Reliability-Centered Maintenance

Reliability-centered maintenance supplies the decision logic that turns the spectrum into a program. It begins not with equipment but with function: what the asset must do, in its operating context, and to what standard. From there it identifies functional failures, the failure modes that cause them, and the effects and consequences of each. SAE JA1011 sets the evaluation criteria that any process must satisfy to be called RCM, framed as seven questions covering functions, functional failures, failure modes, failure effects, failure consequences, proactive tasks and their intervals, and default actions when no suitable proactive task exists. SAE JA1012 expands those criteria into a practitioner guide, and IEC 60300-3-11 provides an international application guide within the dependability-management series.

Consequence drives task selection. Failure modes with safety or environmental consequences must be managed to a tolerable risk regardless of cost, whereas modes with only economic consequences must justify their task on cost grounds. The decision logic then evaluates candidate tasks in order: an on-condition or condition-based task if a measurable degradation indicator exists, scheduled restoration or scheduled discard if the mode is age-related, a failure-finding inspection if the failure is hidden, and no scheduled maintenance if none of these is both applicable and effective. When the logic exhausts every option on a mode with severe consequences, the correct answer is a design change rather than a maintenance task, which is where this discipline meets design for reliability.

Hidden failures deserve particular attention in electronic systems, because protective and standby functions fail silently. A redundant power supply, a watchdog timer, a ground-fault interrupter, or a battery-backed alarm can be dead for months without any operational symptom, and the loss only becomes visible when the protected function is called upon. Failure-finding tasks exist to reveal these failures, and the interval is set from the failure rate of the hidden item and the availability required of the protected function: the shorter the interval, the smaller the mean fraction of time the protection is unavailable. Because these tasks find failures rather than prevent them, their value is measured in avoided multiple failures, not in extended component life.

The P-F Interval and Detection Lead Time

Condition monitoring is feasible only when degradation is detectable before function is lost. The interval between the point at which an incipient failure first becomes detectable and the point of functional failure is the P-F interval, and it governs everything about the inspection program. The inspection interval must be shorter than the P-F interval; common practice sets it at roughly half, so that at least two opportunities to detect the defect fall inside the warning window. What remains after detection, the net P-F interval, is the time actually available to diagnose the problem, order parts, schedule an outage, and complete the work.

P-F intervals vary by orders of magnitude across failure mechanisms, and the spread explains why one technology suits an asset while another does not. Bearing spalling announces itself in vibration spectra weeks or months before seizure. Insulation degradation in a motor winding can be trended over years. Electrolyte loss in an aluminum electrolytic capacitor develops over thousands of hours of operation at elevated temperature. Against these, a latch-up event, an electrostatic-discharge failure, or a dielectric breakdown in a gate oxide can carry a device from healthy to destroyed in microseconds, leaving no interval to exploit. Failure modes at that end of the range must be addressed by derating, redundancy, protection circuits, or screening rather than by inspection, a point developed further in physics of failure.

Choosing a monitoring technology therefore starts with a failure mode, not with a sensor catalog. A failure modes and effects analysis identifies the modes worth monitoring; the P-F interval for each mode determines whether periodic inspection, continuous online monitoring, or no monitoring at all is the right answer. Continuous monitoring earns its cost when the P-F interval is short relative to a practical inspection round, or when the consequence of a missed detection is severe.

Condition Monitoring for Electronic Systems

Classical condition monitoring grew up around rotating machinery, where vibration, lubricant chemistry, and acoustic emission provide rich degradation signals. Purely electronic assemblies offer fewer of those signals, so the monitored parameters are usually electrical and thermal precursors that shift measurably as a component degrades.

Aluminum electrolytic capacitors are the classic example. Manufacturers rate them for an endurance life at a maximum core temperature, commonly with a model in which useful life roughly doubles for every 10 degrees Celsius of reduction in operating temperature, and end of life is conventionally declared at a capacitance loss of about 20 percent or a doubling of equivalent series resistance. Both quantities can be measured in circuit through ripple-voltage or impedance techniques, giving a genuine degradation trend rather than a pass-fail test. Power semiconductor modules degrade through bond-wire lift-off and solder-layer fatigue under thermal cycling, and both mechanisms are observable as a rise in on-state voltage or in junction-to-case thermal resistance. Power MOSFETs show increasing on-resistance and shifting gate-threshold voltage. Lithium-ion batteries are trended through capacity fade, conventionally reaching end of life at 80 percent of rated capacity, and through the growth of internal resistance. Cooling fans report tachometer speed and current draw, which drift as bearings wear.

Several techniques supplement direct parameter measurement. Thermographic inspection of boards, connectors, and switchgear locates high-resistance joints and overloaded conductors before they fail, and ultrasonic detection of corona and partial discharge finds insulation and connection defects in cabling and switchgear; both transfer directly from rotating-machinery practice to electrical equipment. Built-in self-test and boundary-scan structures already present for manufacturing test can be repurposed to monitor interconnect integrity in service. Canary devices, deliberately weakened components placed in the same thermal and electrical environment as the circuit they protect, are designed to fail first and provide advance warning. Performance trending at the system level, watching efficiency, output ripple, error rates, retry counts, or margin in a communications link, often reveals degradation that no single component measurement isolates.

Whatever the sensor, alert thresholds decide whether the program is useful. Thresholds too tight generate false alarms that erode trust and drive unnecessary work; thresholds too loose miss the degradation they were installed to catch. Threshold setting normally starts from manufacturer limits or the statistical distribution of measurements from healthy units, then is refined against confirmed findings, and the confirmations come from disciplined root cause analysis of the assets that are opened up.

From Diagnostics to Prognostics

Diagnostics answers what is wrong now; prognostics answers how long the asset will continue to meet its function. Three modeling approaches are in general use. Physics-of-failure models apply a known damage law, such as a thermal-cycling fatigue model for solder joints or an Arrhenius relation for temperature-activated chemical degradation, and consume measured stresses to accumulate damage. Data-driven models learn the mapping from sensor features to remaining useful life from run-to-failure histories, and dominate where the physics is intractable or the asset is complex. Hybrid approaches use a physical model to structure the problem and data to calibrate its parameters, which is often the most practical option when failure data are scarce.

A remaining-useful-life estimate is worthless without an honest statement of its uncertainty. Sensor noise, unit-to-unit variation, uncertain future load profiles, and model error all widen the prediction interval, and a maintenance decision should be made against a lower confidence bound rather than a point estimate. Prognostic performance is judged by metrics developed for this purpose, including prognostic horizon, the lead time before failure at which predictions first settle within an acceptable error band; alpha-lambda accuracy, which requires the prediction to stay within a shrinking error cone as failure approaches; and measures of relative accuracy and convergence. Reporting these alongside the estimate keeps a prognostic system accountable in the way that a reliability prediction is accountable to field data.

Prognostics pays off only when the estimate changes a decision. The useful outputs are a scheduling recommendation, a spare-parts order, a load or duty-cycle adjustment that extends life until the next planned outage, or a decision to keep operating. Digital twins, which pair a physical asset with a continuously updated model driven by its own telemetry, increasingly serve as the vehicle for those decisions, and appear again in smart factory reliability and in internet of things reliability.

Standards, Data Architecture, and Skills

A monitoring program produces value only if its data flows cleanly from sensor to decision, and several standards define that path. ISO 17359 gives general guidelines for establishing a condition monitoring and diagnostics program, from criticality assessment through measurement selection to diagnosis and prognosis. ISO 13374 defines an open architecture for processing condition-monitoring data, decomposing the chain into data acquisition, data manipulation, state detection, health assessment, prognostic assessment, and advisory generation; the MIMOSA Open System Architecture for Condition-Based Maintenance implements that model and supports interoperability between instruments and analysis software from different vendors. Measurement standards make the readings comparable: the ISO 20816 series governs the measurement and evaluation of machine vibration and consolidates the earlier ISO 10816 and ISO 7919 series, and the ISO 4406 cleanliness code expresses particulate contamination in a lubricant or hydraulic fluid as a three-number rating. None of those measurements applies to an electronic assembly, which has neither bearings nor lubricant; for that case IEEE 1856-2017, the IEEE Standard Framework for Prognostics and Health Management of Electronic Systems, supplies the corresponding terminology, sensor-selection guidance, and capability-classification framework.

Competence is as much a prerequisite as instrumentation. The ISO 18436 series defines the requirements for training and certification of condition-monitoring personnel by technique and by category, and the distinction between a technician who collects data and an analyst who interprets it is where many programs succeed or fail. Data governance matters equally: measurements must be tied to a stable asset identity in the computerized maintenance management system, work orders must record what was actually found, and the findings must be fed back to the thresholds and models. Without that loop, a monitoring program accumulates data instead of knowledge. The same closed-loop discipline that governs failure reporting and corrective action applies here.

Program Economics and Metrics

The transition from reactive to predictive maintenance requires investment in instrumentation, analysis capability, training, and organizational change, and the case for it must be made in financial terms. The U.S. Department of Energy's Operations and Maintenance Best Practices guide reports that preventive maintenance is generally estimated to save 12 to 18 percent over a reactive program, and that a properly functioning predictive maintenance program can save a further 8 to 12 percent over preventive maintenance alone, with larger opportunities where a facility currently relies heavily on reactive work. The same guide cites industrial average results from surveys of organizations that established functional predictive programs: a tenfold return on investment, a 25 to 30 percent reduction in maintenance costs, elimination of 70 to 75 percent of breakdowns, a 35 to 45 percent reduction in downtime, and a 20 to 25 percent increase in production.

These figures describe surveyed averages, mostly from industrial plant rather than electronic equipment, and they should be read as evidence that the direction is right rather than as a forecast for any particular program. Results vary widely with asset criticality, the observability of the dominant failure modes, instrumentation cost, and the maturity of the maintenance organization. A defensible business case models the specific assets in question: the cost of a monitoring channel against the expected cost of the failures it prevents, weighted by their probability and consequence, over the remaining life of the asset. That calculation belongs to reliability economics and, at portfolio scale, to asset management integration.

Programs are steered by a small set of measures rather than by anecdote. Availability, mean time between failures, and mean time to repair describe the outcome; the proportion of planned to unplanned work, schedule and preventive-maintenance compliance, and the ratio of confirmed findings to alarms describe the process. A monitoring program that raises confirmed findings while holding false alarms down is working; one whose alarm volume rises without a corresponding fall in unplanned failures is generating noise. Definitions for the outcome measures are set out in key reliability metrics, and the statistical treatment of the underlying failure data in statistical methods for reliability.

Why Predictive and Preventive Methods Matter

Effective maintenance strategy affects organizational performance through three channels: equipment availability, operating cost, and safety. Organizations that master these methods achieve higher uptime, lower maintenance cost per unit of output, and fewer equipment-related safety incidents. The benefits compound, because each confirmed finding refines a threshold, each false alarm sharpens a model, and each teardown adds to the institutional record of how a particular fleet actually fails.

The methods also close a loop that reaches back into engineering. Field failure data gathered through maintenance and field reliability and warranty analysis are the most direct measurement of how a design behaves in its real environment, and they correct the assumptions embedded in reliability predictions, derating rules, and design guidelines. A maintenance program that reports what it finds makes the next product better; one that merely replaces parts does not.

The objective is not the most sophisticated monitoring available but the strategy each failure mode warrants. Some modes justify continuous instrumentation and a prognostic model. Many justify a periodic inspection. Some justify running to failure and keeping a spare on the shelf. The subcategories above supply the technologies, the analytical methods, and the management practices required to tell these cases apart and to act on the difference.