Electronics Guide

Reliability Testing and Qualification

Reliability testing and qualification convert reliability requirements into empirical evidence. Analysis and prediction methods estimate how long a product should last and where it is likely to fail, but only physical testing reveals the failure modes that analysis cannot foresee and confirms that the design, the components, and the manufacturing process actually deliver the required dependability. The discipline applies controlled stresses that accelerate or precipitate failure mechanisms, so that engineers can validate a design, screen out defective units, and demonstrate compliance within practical development schedules rather than waiting years for field data.

Two complementary objectives run through this work. Qualification establishes, once, that a product design is capable of meeting its reliability targets under the conditions it will encounter; it is repeated only when the design, the materials, or the process changes. Ongoing reliability testing then monitors production, verifying that units coming off the line continue to match the qualified baseline and catching latent defects before they reach customers. The topics in this category address both halves of that program, from the statistics that size a test to the chambers and fixtures that run it.

The work is heavily standardized, and for good reason. A qualification result is a commercial document as much as a technical one: customers, certification bodies, and regulators accept it because it was produced by a recognized method with a recognized sample plan. Standards from JEDEC, the International Electrotechnical Commission, the United States Department of Defense, IPC, and the Automotive Electronics Council supply those methods, and the sections below map how they fit together.

Articles in This Category

Three Distinct Purposes

Reliability tests look superficially alike, because most of them place hardware in a chamber and apply heat, humidity, or vibration. Their purposes differ sharply, and treating one result as though it answered another question is a persistent source of error.

Qualification asks whether a design is fit to be released. It applies a fixed battery of stresses to a modest sample drawn from representative production material and requires that no unit fail. Its output is a verdict on a design, not a failure rate. Screening asks whether an individual manufactured unit carries a latent defect. It applies stress to every unit, or to a sampled fraction, and its output is a sorting decision. Demonstration asks whether a stated numerical requirement, such as a reliability of 0.99 at one year or a mean time between failures of 50,000 hours, is met at a stated confidence. Its output is a statistical statement with an explicit risk level attached.

Two further categories serve development rather than release. Discovery testing, of which highly accelerated life testing is the best-known example, drives hardware past its specification to find weak links and measure margin; it produces corrective actions, not predictions. Characterization testing measures how a mechanism behaves as a function of stress so that an acceleration model can be fitted. Neither yields a pass or fail verdict, and neither should be presented as one.

Standards and Test Frameworks

Component and Semiconductor Qualification

JEDEC standard JESD47, Stress-Test-Driven Qualification of Integrated Circuits, defines the baseline set of stress tests, conditions, and sample sizes used to qualify a new integrated circuit, a product family, or a process change. It references the JESD22 series for the individual test methods: JESD22-A108 for high-temperature operating life, JESD22-A104 for temperature cycling, JESD22-A101 for steady-state temperature-humidity bias, JESD22-A110 for highly accelerated stress testing, and JESD22-A102 for unbiased autoclave exposure. The IEC 60749 series covers much the same ground for international use, with many methods deliberately harmonized with their JEDEC counterparts.

Military and space procurement follows a parallel track. MIL-STD-883 supplies test methods for microcircuits, including Method 1015 for burn-in, together with the screening and quality-conformance procedures that distinguish class levels; MIL-PRF-38535 governs the qualified manufacturers listing under which those parts are supplied. Automotive parts follow the Automotive Electronics Council specifications: AEC-Q100 for integrated circuits, AEC-Q101 for discrete semiconductors, AEC-Q102 for optoelectronics, AEC-Q104 for multichip modules, and AEC-Q200 for passive components. These are more demanding than the commercial baseline, chiefly through wider temperature grades, larger samples, and a dedicated early-life failure rate test.

Environmental Test Methods

The IEC 60068-2 series is the international catalog of environmental test methods, and each part defines one stress in enough procedural detail that two laboratories can produce comparable results: 60068-2-1 for cold, 60068-2-2 for dry heat, 60068-2-6 for sinusoidal vibration, 60068-2-11 for salt mist, 60068-2-14 for change of temperature, 60068-2-27 for mechanical shock, 60068-2-30 for cyclic damp heat, 60068-2-64 for broadband random vibration, and 60068-2-78 for steady-state damp heat, among many others. A test specification selects parts, severities, and durations from that catalog rather than reinventing procedures.

MIL-STD-810 takes a different philosophy. Rather than prescribing severities, it directs the engineer through a tailoring process that derives test conditions from the platform's actual life-cycle environment, then supplies methods to execute them, including Method 501 for high temperature, Method 502 for low temperature, Method 503 for temperature shock, Method 507 for humidity, Method 509 for salt fog, Method 514 for vibration, and Method 516 for shock. The current issue is MIL-STD-810H, released in 2019 and amended by Change 1 in May 2022. Because tailoring is central to the standard, a claim that a product is "MIL-STD-810 compliant" is meaningless without the specific methods and severities that were applied.

Sector-Specific Frameworks

Regulated and safety-critical industries layer their own frameworks on top of the general methods. Road vehicle electronics follow ISO 16750, whose parts address electrical loads, mechanical loads, climatic loads, and chemical loads separately, defining the conditions a module must survive in an engine bay, a passenger compartment, or an exterior mounting. Airborne equipment follows RTCA DO-160, which combines temperature and altitude, vibration, fluid susceptibility, lightning-induced transients, and electromagnetic compatibility into a single qualification document that certification authorities recognize. Telecommunications equipment is commonly qualified against Telcordia GR-63-CORE for the physical environment and GR-1089-CORE for electrical protection.

The automotive sector also contributed the robustness-validation approach, which reframes qualification around the mission profile. Instead of asking only whether a part survived a fixed stress battery, robustness validation compares the demonstrated capability of the part against the loads its specific application will impose, and reports the margin between them. That framing exposes a weakness of pass-or-fail qualification: a part that barely survives a standard test and a part that survives it with an order of magnitude of margin receive the same certificate.

Board and Assembly Level Testing

Component qualification says nothing about how a part behaves once it is soldered to a board, where the mismatch between the coefficients of thermal expansion of silicon, mold compound, solder, and laminate governs interconnect life. IPC-9701 supplies the reference thermal cycling method for characterizing the fatigue life of surface mount solder attachments, with defined cycling conditions such as 0 to 100 degrees Celsius and −55 to 100 degrees Celsius. Its later revision narrowed the scope explicitly to characterization: the method generates fatigue-life data for comparison and for feeding life models, not a pass-or-fail qualification verdict.

Mechanical robustness of assemblies is addressed separately. JESD22-B111 defines the board-level drop test used for handheld products, mounting a test board on a rigid fixture and subjecting it to a controlled half-sine pulse while interconnect resistance is monitored continuously, because a solder joint may open for microseconds during an impact and close again afterward. Cyclic bend testing addresses the flexure that occurs during board assembly, connector insertion, and enclosure handling, which is a common cause of cracked ceramic capacitors and lifted pads.

Planning a Qualification Program

From Mission Profile to Test Plan

A defensible plan starts with the mission profile: the temperature distribution, humidity exposure, vibration spectrum, power cycling pattern, duty cycle, and required service life that the product will actually see. That profile, combined with a physics-of-failure review of the expected mechanisms, determines which stresses matter. A wall-powered indoor instrument that never cycles power is dominated by different mechanisms than a battery-powered outdoor sensor that cycles thermally twice a day, and running the same test battery on both wastes effort on one and under-tests the other.

Standard test batteries remain valuable as a floor, because they are recognized and comparable, but they are calibrated to a generic use case. Where the application is more severe than that assumption, the plan must extend duration, widen the temperature range, or add a stress the standard omits. Where the application is milder, the standard battery may be retained simply because customers expect it. Recording the reasoning in the test plan matters as much as the result, since a later reviewer must be able to see why each test was chosen and what field condition it represents.

Sample Size and Statistical Confidence

Qualification samples must come from material that represents production, which is why standards require units drawn from multiple production lots rather than from a single engineering build. Multiple lots capture the lot-to-lot variation that a single build conceals, and three lots is the customary minimum in semiconductor practice.

The sample sizes look arbitrary until the underlying statistics are exposed. A zero-failure, or success-run, plan requires n units to survive a test with no failures, where n equals the natural logarithm of one minus the confidence divided by the natural logarithm of the reliability to be shown. Demonstrating 99 percent reliability at 90 percent confidence therefore takes 230 units, which is why the widely used convention of 77 units drawn from each of three lots, 231 units in total, appears throughout JESD47 and AEC-Q100: passing it with zero failures corresponds to a lot tolerance percent defective of 1 percent at 90 percent confidence. The arithmetic also exposes the method's limits. Sample sizes climb steeply as the reliability target tightens, and a zero-failure result establishes a bound without revealing anything about how or when the product would eventually fail.

Preconditioning and Stress Sequencing

Order matters, because damage accumulates. Surface mount components are preconditioned before reliability stress under JESD22-A113, which subjects them to a bake, a controlled moisture soak, and multiple simulated reflow passes so that the subsequent reliability test evaluates a part in the condition a customer will actually assemble, not a pristine one. Skipping preconditioning produces optimistic results, since the assembly process itself inflicts thermal and hygroscopic damage.

Moisture sensitivity is handled by a related pair of standards. IPC/JEDEC J-STD-020 classifies plastic packages into moisture sensitivity levels according to the floor life they tolerate after removal from a dry pack, and J-STD-033 specifies the matching dry-pack, humidity-indicator, and bake-out practices that reset accumulated floor life. Where lead-free reflow peaks near 260 degrees Celsius, absorbed moisture can flash to steam and crack a package, so a qualification program that ignores handling controls may fail for reasons that have nothing to do with the design.

Beyond preconditioning, sequences are usually structured so that non-destructive checks precede destructive ones, and so that a single sample group carries a realistic combination of stresses rather than one stress in isolation. Combined-environment testing, in which temperature and vibration are applied together, is more representative than sequential single-stress testing but harder to interpret when a failure occurs, and the choice between them is a deliberate trade-off rather than a matter of convenience.

Demonstrating a Numerical Requirement

When a contract states a reliability figure, testing must produce a statistical statement rather than a qualitative pass. Where a fixed mission or warranty period is specified, the success-run plan described above answers the question directly. Where the requirement is expressed as a mean time between failures for a repairable system with an approximately constant failure rate, the demonstration is framed in accumulated operating hours instead of units. Under those assumptions, a test that finishes with zero failures supports a lower confidence bound of the total test time divided by 2.30 at 90 percent confidence, so demonstrating a 10,000-hour mean time between failures requires roughly 23,000 accumulated failure-free hours. Tests that end with failures use the corresponding chi-square bound, which widens as the failure count grows.

Fixed-length plans of this kind are simple to administer but can be long. Sequential plans, such as the probability ratio sequential test, evaluate the accumulated evidence continuously and stop as soon as it favors acceptance or rejection decisively, which on average shortens the test substantially for products that are clearly good or clearly bad, at the cost of an indeterminate schedule. Both approaches make consumer and producer risk explicit: the consumer risk is the probability of accepting a product that does not meet the requirement, and the producer risk is the probability of rejecting one that does. The discrimination ratio between the acceptable and unacceptable reliability levels, together with those two risks, determines the length of the test, and negotiating them is a commercial decision with a direct schedule cost.

Bayesian demonstration offers a third route, formally combining prior evidence, such as results from a predecessor product or from component-level qualification, with the new test data. It can reduce test time materially when the prior is genuinely informative and defensible, but it shifts the argument onto the justification of the prior, which is why customers frequently insist on classical plans instead.

Screening and Ongoing Reliability Testing

Qualification ends when the design is released; the manufacturing risk does not. Screening addresses the early-life portion of the failure-rate curve, where latent defects such as marginal solder joints, contaminated die attach, weak wire bonds, and partially damaged oxides fail quickly under stress but might otherwise survive final test and fail at a customer. Classical burn-in applies elevated temperature with bias for a fixed duration; environmental stress screening and highly accelerated stress screening substitute thermal cycling and vibration, which precipitate mechanical defects that steady-state burn-in leaves untouched.

Every screen consumes useful life, so the decision to apply one is economic. Where the defect rate is low and the units are cheap, screening every unit costs more than the field failures it prevents. Where the escape cost is high, as in an implanted medical device, a satellite payload, or a part that is inaccessible once installed, the calculation reverses decisively. Mature processes therefore graduate from full screening to sampled screening, and a rising screen fallout rate becomes a process alarm rather than a routine yield loss.

Ongoing reliability testing samples finished production continuously and subjects it to a subset of the qualification stresses, typically life test and temperature cycling. Its purpose is not to requalify the design but to detect drift: a changed mold compound, a new bond wire supplier, a fab process tweak, or a shift in a solder paste profile. Because the sample from any one interval is small, ongoing reliability testing is analyzed as a trend across intervals rather than as a series of independent verdicts, and its statistical power comes from accumulating device-hours over time.

Interpreting, Documenting, and Maintaining Qualification

Every failure during qualification warrants failure analysis, whatever the sample plan says. A zero-failure requirement makes a single failure fatal to the test, which creates pressure to reclassify it as a test artifact, a handling error, or an equipment fault. Sometimes that is genuinely true, and the standards permit a documented rejection on those grounds, but the burden of proof belongs on the party seeking the exclusion, supported by physical evidence of the mechanism. A failure dismissed without analysis is a field failure deferred.

The qualification report is the deliverable that survives the program. It should record the material tested and its lot traceability, the test plan and the rationale behind it, the equipment and its calibration status, the actual conditions achieved rather than the nominal ones, the readout points, every anomaly with its disposition, and the analysis of any failure. Customers in regulated sectors frequently witness the testing or audit the report, and automotive supply agreements fold the results into the broader production part approval process.

Qualification is also a perishable claim. A die shrink, a change of wafer fab or assembly site, a new mold compound or lead frame, a substrate supplier change, a design revision, or a shift in the solder alloy all trigger a review, and the standards define which of them demand full requalification and which permit a reduced set of tests. Qualification by similarity and the use of generic data from a related product family can legitimately shorten that work, but only when the supplier demonstrates that the two products share the relevant construction and process, and that the family carries no common failure mechanism that the substitution would hide.

Common Pitfalls

The most frequent error is treating a passed qualification as a measured reliability. A zero-failure result at a stated confidence bounds the failure probability; it does not estimate a failure rate, predict a service life, or say anything about mechanisms that the chosen stresses failed to activate. A related error runs in the opposite direction: quoting a mean time between failures derived from a handful of accelerated failures to three significant figures, as though the extrapolation carried no uncertainty.

Over-stress is the second recurring problem. Pushing a test above a physical transition, such as a mold compound's glass transition temperature or a dielectric's breakdown field, produces failures that the product would never experience in service, and each such failure corrupts the conclusion rather than strengthening it. Under-stress fails more quietly: a test whose severity falls short of the real environment passes reliably and proves nothing.

Finally, programs fail on execution as often as on design. Chambers drift out of calibration, thermocouples measure air rather than the component, functional monitoring is absent so that intermittent failures go unrecorded, fixtures resonate and change the vibration actually delivered to the unit, and the sample turns out to have come from a single engineering lot built by hand. Each of these invalidates a result that looks entirely respectable on paper, which is why test-condition verification and instrumented monitoring belong in the plan from the outset.

Why Reliability Testing Matters

Reliability testing is where reliability engineering meets reality. Predictions rest on models and historical data; testing supplies the direct measurement that validates those estimates, exposes design and process weaknesses, and provides the documented basis on which products are released and certified. A field failure discovered after shipment is expensive to diagnose and costly to a manufacturer's reputation, whereas the same weakness found on a shaker table or in a temperature-humidity chamber is comparatively inexpensive to correct.

Designing an effective program is itself an engineering trade-off. Tests must apply stresses severe enough to precipitate genuine wear-out and defect mechanisms, yet not so severe that they trigger failures the product would never experience in service and waste effort chasing artifacts. Sample sizes, stress levels, and test durations must be chosen so that the result carries statistical meaning while the program still fits the schedule and budget. The standards and methods covered here provide proven frameworks for striking that balance across diverse product types and operating environments, from consumer electronics to aerospace and industrial systems.

Related Topics