Electronics Guide

Physics of Failure Approaches

Physics of failure (PoF), also called reliability physics, is a reliability engineering methodology that uses knowledge of physical failure mechanisms and their root causes to predict, prevent, and mitigate failures in electronic systems. Rather than relying solely on statistical analysis of historical failure data, PoF examines the physical, chemical, and mechanical processes that cause components and assemblies to degrade and ultimately fail under operational stresses.

The approach treats failure as the deterministic outcome of identifiable processes acting on real materials and geometries, with statistical scatter arising from variation in material properties, dimensions, and applied loads. By modeling those processes, engineers design products that address reliability concerns at their source, develop qualification tests that accelerate the mechanisms that actually limit life, and extrapolate test results to use conditions with a defensible physical basis. Physics of failure is now standard practice in semiconductor process qualification, and it is widely applied in aerospace, defense, automotive, medical, and telecommunications electronics.

This article covers physics of failure as a method: how a mechanism is modeled, how an acceleration factor is derived from it, and how the result feeds design and test. The mechanisms themselves are catalogued in Failure Mechanism Understanding; the survey below is deliberately brief because that page carries the depth.

Foundations of Physics of Failure

Physics of failure rests on the principle that understanding failure mechanisms at their root cause enables better design decisions, more meaningful testing, and more accurate life predictions than purely empirical approaches.

From Empirical to Physics-Based Approaches

Understanding the evolution from traditional reliability methods clarifies the advantages of PoF:

  • Traditional empirical methods: Rely on historical field failure data and fitted statistical distributions; treat failure as a random process characterized by a failure rate
  • Handbook prediction limitations: MIL-HDBK-217 and similar handbooks assign empirically derived, largely constant failure rates by part type. The handbook's last revision, MIL-HDBK-217F Notice 2, dates from 1995, and its part-count and part-stress models do not represent the wear-out mechanisms of technologies developed since
  • Constant-rate assumption: Handbook models imply an exponential time-to-failure distribution, which obscures the increasing hazard rate that characterizes most degradation mechanisms
  • Physics-based foundation: PoF identifies specific failure mechanisms and models their progression using material properties, geometry, and applied stress
  • Root cause focus: Understanding why failures occur enables design solutions that address underlying causes rather than symptoms
  • Technology transfer: A mechanism model calibrated for one material system applies to another when the same mechanism operates, which supports assessment of new technologies with no field history

University and consortium research drove the adoption of the approach in electronics, most visibly at the Center for Advanced Life Cycle Engineering at the University of Maryland, and standards bodies have since formalized the shift. IEEE Std 1413 defines a framework for documenting the assumptions, inputs, and uncertainties behind any hardware reliability prediction, and the JEDEC publication JEP148 describes qualification of semiconductor devices based on physics-of-failure risk assessment rather than fixed test matrices. PoF does not replace statistical methods; it supplies the mechanistic structure that statistical models then quantify.

Stress-Strength Interference

The stress-strength model provides a fundamental framework for understanding failure:

  • Basic concept: Failure occurs when applied stress exceeds material or component strength; both stress and strength are distributions, so reliability equals the probability that strength exceeds stress
  • Stress sources: Environmental factors (temperature, humidity, vibration, shock), electrical loading (voltage, current, power), and mechanical loading (forces, pressures, board flexure)
  • Strength characteristics: Material properties, component ratings, and design margins that resist applied stresses
  • Interference region: The overlap between the two distributions represents the probability of failure; widening either distribution increases risk even when the mean margin is unchanged
  • Time dependence: Strength typically degrades as damage accumulates while stress may vary, so interference grows with time and produces the wear-out region of the hazard curve

A capacitor rated at 50 V and operated at 25 V illustrates the point. The nominal margin is large, but a wide dielectric-strength distribution from process variation, combined with occasional line transients, can still place part of the population in the interference region. Derating widens the separation between the distributions rather than eliminating overlap outright.

Overstress and Wear-Out Mechanisms

Physics of failure divides mechanisms into two classes that demand different treatment:

  • Overstress mechanisms: A single event exceeds strength and causes immediate failure, as in electrostatic discharge, electrical overstress, dielectric rupture, mechanical fracture, or thermal runaway
  • Wear-out mechanisms: Damage accumulates gradually until a parameter crosses a failure threshold, as in electromigration, dielectric breakdown, solder fatigue, corrosion, and bias temperature instability
  • Different design responses: Overstress is controlled by margin, protection devices, and load limiting; wear-out is controlled by derating, material selection, and thermal design
  • Different test strategies: Overstress is characterized by step-stress testing to destruct limits; wear-out is characterized by time-to-failure testing at elevated stress
  • Different statistics: Overstress failures follow the tail of the load distribution; wear-out failures follow a life distribution such as the lognormal or Weibull with an increasing hazard rate

Confusing the two leads to misdirected effort, such as attempting to derate against an electrostatic discharge event or adding transient protection against a fatigue problem.

Failure Mechanism Identification

Systematic identification of applicable failure mechanisms is essential to PoF:

  • Mechanism categories: Chemical (corrosion, oxidation, intermetallic growth), mechanical (fatigue, wear, creep, fracture), electrical (electromigration, dielectric breakdown, charge trapping), and thermal (interdiffusion, stress relaxation)
  • Life cycle load analysis: Characterize every stress the product experiences during manufacture, shipping, storage, and operation, not just the nominal operating condition
  • Material and geometry mapping: Susceptibility depends on the specific material set and structure, so mechanism identification proceeds from a bill of materials and physical layout
  • Stress interaction: Simultaneous stresses may accelerate failure beyond the sum of their individual effects, as when temperature cycling cracks a passivation layer and admits moisture
  • Life cycle phasing: Different mechanisms dominate in different phases; reflow soldering imposes the harshest thermal excursion many parts ever see

JEDEC publication JEP122 catalogs the recognized failure mechanisms and models for semiconductor devices and serves as a common reference for this step.

Failure Mechanism Models

Mathematical models describe how failure mechanisms progress:

  • Rate equations: Express damage accumulation rate as a function of stress variables and material properties
  • Activation energy: Thermally activated mechanisms follow Arrhenius behavior, with rate proportional to exp(−Ea/kT), where k is the Boltzmann constant (8.617 × 10−5 eV/K) and T is absolute temperature
  • Acceleration factors: Ratios relating mechanism rates at different stress levels, which make accelerated testing quantitative
  • Damage thresholds: Criteria that separate initiation, propagation, and functional failure, since a mechanism may operate for years before it affects performance
  • Multi-physics models: Coupled thermal, mechanical, electrical, and chemical analysis for mechanisms that no single-domain model captures

A single activation energy applies only to a single mechanism. Assigning one composite activation energy to an entire assembly is a common error, because different mechanisms accelerate at different rates and the dominant mechanism can change with temperature.

Common Electronic Failure Mechanisms

Electronic systems exhibit characteristic failure mechanisms determined by their materials, structures, and operating conditions. The models below are the quantitative core of PoF practice: each relates an observable stress to a time to failure, and each carries assumptions that limit its valid range.

Electromigration

Electromigration is the transport of metal atoms by momentum transfer from conducting electrons:

  • Mechanism physics: High current density produces an electron wind that displaces metal atoms along the direction of electron flow, with mass transport concentrated on the fastest diffusion path
  • Failure manifestation: Void formation where atoms are depleted, causing resistance increase and eventual opens, and hillock or extrusion formation where atoms accumulate, potentially causing shorts between adjacent lines
  • Critical factors: Current density, typically becoming a concern above roughly 105 A/cm2, along with temperature, conductor geometry, grain structure, and interface quality
  • Black's equation: MTF = A × j−n × exp(Ea/kT), where j is current density. The exponent n is approximately 1 when void growth limits life and approximately 2 when void nucleation limits it
  • Activation energies: Roughly 0.5 to 0.7 eV for aluminum, where grain-boundary diffusion dominates, and roughly 0.8 to 1.0 eV for copper, where interface and surface diffusion dominate; the higher barrier is one reason copper interconnects tolerate greater current density
  • Blech effect: Below a critical product of current density and line length, back-stress from accumulated atoms balances the electron wind and the line becomes effectively immortal, which designers exploit through short segments between vias
  • Stress-induced voiding: A related but distinct mechanism, in which tensile stress from thermal expansion mismatch in confined metal drives vacancies together into voids with no current flowing at all. It is qualified by separate high-temperature storage tests, because a design that satisfies electromigration limits can still fail this way
  • Design mitigation: Larger conductor cross-sections, refractory barrier and cap layers, copper metallization with engineered interfaces, reservoir geometries at via ends, and reduced operating temperature

Electromigration remains a primary constraint in integrated circuits, because shrinking interconnect cross-sections raise current density faster than allowable currents fall. Package-level and board-level conductors are far less susceptible, since their cross-sections are orders of magnitude larger.

Time-Dependent Dielectric Breakdown

Gate oxide and interlayer dielectric breakdown limits integrated circuit reliability:

  • Mechanism physics: The electric field across a thin dielectric generates traps in the bulk and at interfaces; when traps form a continuous percolation path, a conductive filament shorts the dielectric
  • Failure progression: Gradual leakage increase, then soft breakdown in which a localized path carries limited current, then hard breakdown with thermal runaway along the filament
  • Critical factors: Field strength, dielectric thickness, temperature, defect density from processing, and the area under stress, since larger areas contain more weak spots
  • E-model: Time to breakdown falls exponentially with increasing field. This thermochemical model fits low-field, elevated-temperature data and yields the more conservative extrapolation to use conditions
  • 1/E model: Time to breakdown depends exponentially on the reciprocal of the field. This anode-hole-injection model fits high-field data, where Fowler-Nordheim tunneling dominates
  • Model selection: Published comparisons place the transition between the two regimes in the mid-field range of roughly 5 to 7 MV/cm, and the two extrapolations diverge by orders of magnitude by the time they reach use conditions. Neither model describes oxides thinner than about 4 to 5 nm well, where direct tunneling dominates and power-law models in gate voltage are used instead
  • Design approaches: Voltage derating, area-scaled lifetime budgets, control of process defect density, and thicker physical dielectrics where performance permits

Because the choice between models changes the extrapolated lifetime by orders of magnitude, disclosing the model and the stress range used is an essential part of any TDDB qualification claim.

Hot Carrier Injection

Energetic carriers damage transistor channels and gate dielectrics:

  • Mechanism physics: Carriers gain energy in the high-field drain region; a fraction acquires enough energy to create interface states or to be injected into the gate dielectric, where they become trapped charge
  • Device impact: Threshold voltage shift, transconductance and drive current degradation, and increased subthreshold slope, which together erode timing margin in digital circuits and matching in analog circuits
  • Critical factors: Drain voltage, channel length, dielectric thickness, and switching activity, since damage in digital circuits accumulates mainly during transitions
  • Acceleration models: Classical lucky-electron models scale degradation with substrate current, which peaks at intermediate gate bias. In short-channel and fin-based devices, carrier-carrier scattering and self-heating make drain current and local temperature better predictors
  • Design mitigation: Lightly doped drain and graded junction structures, supply voltage reduction, channel and halo engineering, and limits on sustained high-field operation

Supply voltage scaling reduced hot carrier degradation in core logic, but the mechanism remains a first-order concern for input and output devices, high-voltage transistors, and analog circuits that hold devices in saturation for long periods.

Bias Temperature Instability

Negative bias temperature instability (NBTI) affects PMOS transistors under negative gate bias at elevated temperature:

  • Mechanism physics: Field and temperature drive depassivation of silicon-hydrogen bonds at the channel interface and charge trapping in the dielectric, both of which shift threshold voltage
  • Partial recovery: A significant fraction of the shift recovers within microseconds to seconds once stress is removed, so slow measurements systematically understate the damage and on-the-fly or fast-pulse techniques are required
  • Device impact: Threshold voltage magnitude increases and mobility falls, degrading drive current and skewing circuit timing over the product life
  • Time dependence: Degradation follows an approximate power law in time, with reported exponents commonly between 0.15 and 0.25 depending on the dominant process and the measurement delay
  • Critical factors: Gate field, temperature, duty cycle, and process-dependent interface and dielectric quality
  • Counterpart mechanism: Positive bias temperature instability degrades NMOS devices with high-permittivity gate stacks, so modern reliability models treat NBTI and PBTI together as bias temperature instability

Bias temperature instability grew in importance as oxide fields rose with scaling and as high-permittivity gate stacks introduced additional trap populations. Circuit-level reliability simulators now age transistor models to project timing margin at end of life.

Corrosion and Electrochemical Migration

Moisture and ionic contamination drive several distinct mechanisms:

  • Electrochemical corrosion: Galvanic action between dissimilar metals proceeds where an electrolyte film bridges them, consuming the more active metal
  • Electrochemical migration: Under bias and adsorbed moisture, metal dissolves at the anode, migrates, and plates out as dendrites that grow toward the cathode until they short adjacent conductors
  • Conductive anodic filament formation: Within laminate, metal filaments propagate along delaminated glass-fiber-to-resin interfaces between plated holes, a mechanism that fine-pitch, high-voltage designs must consider
  • Stress corrosion cracking: Combined mechanical stress and a corrosive environment crack materials at stress concentrations well below their static strength
  • Environmental factors: Relative humidity, ionic residues from flux and handling, temperature, applied bias, and conductor spacing
  • Protection approaches: Passivation and moisture-barrier layers, conformal coating, cleanliness control verified by ion chromatography and surface insulation resistance testing, adequate spacing, and hermetic or near-hermetic packaging for severe environments

Corrosion mechanisms accelerate strongly with humidity and temperature, which is why temperature-humidity-bias and highly accelerated stress testing are standard qualification steps for non-hermetic parts.

Solder Joint Fatigue

Thermal cycling causes fatigue failure of solder interconnections:

  • Mechanism physics: Mismatch in coefficient of thermal expansion between component and substrate imposes cyclic shear strain on the joint; solder alloys creep at ordinary operating temperatures, so the damage is combined creep and fatigue
  • Damage accumulation: Cyclic inelastic strain coarsens the microstructure, initiates cracks near the highest-strain interface, and propagates them across the joint until the connection opens, often intermittently at temperature extremes first
  • Coffin-Manson relation: Cycles to failure scale with the inelastic strain range raised to a negative power. With fatigue ductility exponents near 0.5 to 0.7, the resulting strain-range exponent falls near 1.4 to 2
  • Norris-Landzberg modification: Adds cycling frequency and maximum-temperature terms to the Coffin-Manson form, which matters because longer dwells allow more creep damage per cycle
  • Intermetallic growth: A tin-copper or tin-nickel intermetallic layer forms at each pad interface and thickens by diffusion with time and temperature. An excessively thick and brittle layer, along with Kirkendall voiding beneath it, moves the crack path to the interface and degrades resistance to drop and bend loading
  • Critical factors: Temperature range and mean, dwell time, distance to neutral point (which makes large packages worse), solder alloy, joint standoff and shape, and board and component stiffness
  • Design mitigation: Compliant leads or interposers, closer expansion matching, underfill or corner bonding, larger standoff, and placement that limits the distance from the package center to the outermost joints

Tin-lead and tin-silver-copper alloys respond differently to the same profile, so acceleration factors derived for one alloy family must not be reused for another without recalibration. Solder fatigue is frequently the life-limiting mechanism for surface mount assemblies in thermally cycling environments.

Passive Component and Board-Level Wear-Out

Mechanisms outside the die and the solder joint often govern assembly life:

  • Electrolytic capacitor dry-out: Electrolyte escapes through the seal at a rate set by core temperature, so capacitance falls and equivalent series resistance rises. Manufacturers publish life at a rated temperature and apply an approximate doubling of life for every 10 degrees Celsius of reduction in core temperature, with end of life commonly defined as a 20 percent capacitance loss or a doubling of the initial equivalent series resistance
  • Ripple current self-heating: Ripple current heats the capacitor core above ambient, so any life estimate that uses ambient temperature in place of core temperature overstates life substantially
  • Ceramic capacitor flex cracking: Board flexure during depaneling, in-circuit test, connector insertion, or fastening drives cracks from the termination into the ceramic body, producing degraded insulation resistance, intermittent shorts, and in severe cases thermal runaway
  • Insulation resistance degradation: Under sustained direct-current bias and elevated temperature, oxygen vacancies migrate within a ceramic dielectric and leakage current grows with time, which is the wear-out mechanism that highly accelerated life testing of ceramic capacitors targets
  • Tin whiskers: Compressive stress in pure tin finishes drives filament growth that can bridge fine-pitch conductors. Nickel underlayers, post-plate annealing, matte rather than bright deposits, and alloying additions reduce the risk but no single measure eliminates it
  • Contact degradation: Fretting corrosion under micro-motion, stress relaxation of contact springs, and loss of contact normal force raise connector and relay contact resistance well before any visible damage appears

These mechanisms dominate system life in equipment whose semiconductors are lightly stressed, such as power supplies and industrial controllers. A physics of failure assessment that stops at the die is therefore incomplete.

Degradation Modeling and Life Prediction

Physics of failure enables quantitative life prediction by modeling degradation processes and their progression to functional failure.

Degradation Path Analysis

Tracking parameter degradation reveals the failure trajectory:

  • Degradation metrics: Measurable parameters that change as a mechanism progresses, such as leakage current, threshold voltage shift, contact resistance, capacitance loss, or equivalent series resistance
  • Failure criteria: Parameter thresholds beyond which the device no longer meets specification, defined before testing begins to avoid post hoc adjustment
  • Degradation models: Functions describing parameter change with time under specified stress, commonly linear, power law, or exponential in form
  • Extrapolation: Projection of measured paths to the failure threshold, which yields life estimates without waiting for actual failures
  • Population variability: Unit-to-unit variation in initial value and degradation rate, modeled with random-effects or hierarchical formulations

Degradation analysis extracts far more information per test unit than pass-fail testing, which makes it valuable when sample sizes are small or when a test must end before most units fail.

Damage Accumulation Models

Cumulative damage models track mechanism progression under variable loading:

  • Miner's rule: Linear damage accumulation, with failure predicted when the sum of cycle fractions reaches one; simple, widely used, and known to be approximate
  • Known limitations: The linear rule ignores load sequence effects and the interaction between damage states, so observed damage sums at failure often differ from one
  • Cycle counting: Rainflow counting reduces an irregular temperature or vibration history to equivalent constant-amplitude cycles that damage models can consume
  • Multi-mechanism damage: Several mechanisms may progress simultaneously, sometimes synergistically, as when creep damage accelerates fatigue crack growth
  • Remaining life estimation: Accumulated damage computed from measured loads forms the basis of prognostics and condition-based maintenance

Damage accumulation converts a realistic mission profile into a life estimate, which is what distinguishes a usable prediction from a constant-stress laboratory result.

Acceleration Models

Acceleration models relate mechanism rates at different stress levels:

  • Arrhenius model: Temperature acceleration for thermally activated mechanisms; AF = exp[(Ea/k) × (1/Tuse − 1/Ttest)], with temperatures in kelvins
  • Eyring model: A generalized form that combines a temperature term with additional stress terms such as voltage or humidity
  • Inverse power law: Voltage, current, or vibration acceleration; AF = (Stest/Suse)n, with the exponent determined experimentally for the specific mechanism and material
  • Coffin-Manson and Norris-Landzberg: Thermal cycling acceleration driven by temperature range, with frequency and maximum-temperature corrections for creep-dominated solder
  • Peck model: Temperature-humidity acceleration combining an Arrhenius term with an inverse power law in relative humidity; reported humidity exponents cluster near 2.5 to 3 and activation energies near 0.7 to 0.9 eV for aluminum metallization corrosion

Every acceleration factor is mechanism specific and valid only within the stress range over which it was calibrated. Testing above that range risks activating a mechanism that never occurs in the field, which produces optimistic or simply irrelevant results.

Life Prediction Methodology

Systematic life prediction integrates mechanism models with application conditions:

  • Operating profile definition: Characterize expected thermal, electrical, mechanical, and environmental conditions across the full life cycle, including storage and transport
  • Mechanism ranking: Identify the mechanisms plausibly active for the specific materials, geometry, and load set
  • Model application: Apply the appropriate mechanism models with application-specific parameters, obtained from measurement or supplier data rather than defaults
  • Life calculation: Compute predicted life for each mechanism; the shortest governs, and the ranking identifies where design effort pays
  • Uncertainty quantification: Propagate parameter and load uncertainty, typically by Monte Carlo simulation, to obtain a life distribution with confidence bounds rather than a single number

A physics-based prediction is auditable: each input is a stated material property, dimension, or load, and each can be challenged, measured, and improved. That traceability is often more valuable than the point estimate itself.

Design for Reliability Using PoF

Physics of failure principles guide design decisions that improve reliability by addressing failure mechanisms at their source.

Derating Strategies

Operating components below rated limits reduces failure mechanism rates:

  • Voltage derating: Reduces electric field stress, slowing dielectric breakdown, hot carrier damage, and electrochemical migration
  • Current derating: Reduces joule heating and current density, mitigating electromigration and thermal degradation
  • Temperature derating: Reduces thermally activated rates exponentially and is frequently the single most effective measure
  • Power derating: Limits self-heating so that junction temperature stays within the range where mechanism models were validated
  • Mechanism-specific guidelines: Derating factors should follow the sensitivity of the governing mechanism rather than a uniform percentage applied to every parameter

Derating also has limits and costs. Excessive voltage derating on electrolytic and tantalum capacitors, for example, can be counterproductive, and oversizing conductors consumes routing area. Effective derating is mechanism informed, not maximal.

Material Selection

Material choices fundamentally determine susceptibility to failure mechanisms:

  • Conductor materials: Copper tolerates higher current density than aluminum because of its higher electromigration activation energy; barrier and cap layers control the fast interface diffusion path
  • Dielectric materials: High-permittivity gate dielectrics allow a physically thicker film at the same equivalent oxide thickness, which sharply reduces direct-tunneling leakage. Breakdown behavior is not automatically better, since these stacks introduce trap populations that drive PBTI and require their own reliability characterization
  • Solder alloys: Tin-silver-copper and tin-lead alloys differ in creep behavior, stiffness, and fatigue response, so alloy substitution requires requalification rather than reuse of prior acceleration factors
  • Substrate materials: Closer matching of thermal expansion coefficients reduces interconnect strain; glass transition temperature and out-of-plane expansion govern plated-hole reliability during assembly
  • Encapsulants and finishes: Moisture barriers, stress-buffer layers, and whisker-mitigating surface finishes address environmental and metallurgical mechanisms

Material substitutions made for cost or availability reasons frequently change the dominant failure mechanism, which is why PoF review belongs in the change control process and not only in initial design.

Thermal Design

Temperature profoundly affects most failure mechanisms:

  • Junction temperature: Thermally activated mechanisms accelerate exponentially with junction temperature, so a modest reduction can yield a large life gain
  • Thermal resistance: Minimize junction-to-ambient resistance through die attach, package selection, board copper, thermal vias, and heat sinking
  • Hot spot management: Distribute dissipating components and manage on-die power density to avoid local extremes that average temperature measurements hide
  • Thermal cycling: Limit the amplitude and dwell of temperature excursions to reduce solder fatigue, wire bond flexure, and delamination
  • Thermal simulation: Computational fluid dynamics and finite element analysis identify gradients and transients before hardware exists, provided the models are validated against measurement

Given the strong temperature dependence of most mechanisms, thermal design typically returns more reliability improvement per unit of engineering effort than any other single measure.

Geometric Design Optimization

Component and interconnect geometry sets local stress levels:

  • Conductor sizing: Adequate cross-section keeps current density below electromigration limits, with attention to worst-case rather than average current
  • Via design: Multiple parallel vias reduce current crowding, and via placement avoids regions of peak thermomechanical strain
  • Solder joint geometry: Standoff height, pad definition, and fillet shape redistribute strain away from crack-initiation sites
  • Stress concentration avoidance: Smooth transitions, generous radii, and avoidance of abrupt stiffness changes reduce peak stress
  • Clearances and spacings: Adequate conductor spacing suppresses electrochemical migration and limits field stress in surface and internal dielectrics
  • Board-level placement: Keeping large packages away from board edges, mounting holes, and connector insertion paths limits flexure-induced pad cratering and joint cracking

Geometry changes guided by mechanism understanding often improve reliability at no unit cost, which makes them the most economical class of PoF-driven design action.

Physics-Based Testing

Physics of failure principles make reliability testing more effective by ensuring that tests stress the mechanisms of interest at defensible acceleration levels.

Accelerated Test Design

Tests should follow from failure mechanism understanding:

  • Mechanism targeting: Choose stress types and levels that accelerate the specific mechanisms expected to limit life
  • Acceleration factor validation: Confirm that the acceleration model holds across the test range, ideally by testing at two or more stress levels and checking consistency
  • Failure mode verification: Analyze test failures physically to confirm they match field failure modes; matching times to failure without matching mechanisms proves nothing
  • Stress level selection: High enough for practical duration, low enough to avoid melting, phase changes, or mechanism shifts that never occur in service
  • Sample size and censoring: Plan the number of units and test duration against the confidence needed, recognizing that most qualification tests end with many units unfailed
  • Combined stresses: Sequential or simultaneous stresses may be required to reproduce field interactions such as moisture ingress following thermal cycling

Physics-based test design avoids the central pitfall of accelerated testing, which is producing failures that are real but irrelevant.

Highly Accelerated Life Testing

HALT applies extreme stresses to expose design weaknesses quickly:

  • Step stress approach: Increase stress in steps until failures occur, mapping the margin between specification and actual capability
  • Operating limits: Determine the temperature and vibration levels at which function ceases but recovers when stress is removed
  • Destruct limits: Determine the levels that cause permanent damage, which identify the weakest element in the design
  • Combined stresses: Simultaneous thermal cycling and broadband random vibration reveal interactions that single-stress testing misses
  • Rapid feedback: Results arrive in days, early enough to change a design rather than merely document its limits

HALT is qualitative. Its stresses generally lie outside the range where acceleration models were calibrated, so HALT results establish margin and expose weaknesses but do not yield life predictions.

Mechanism-Specific Tests

Standardized tests target individual failure mechanisms:

  • Electromigration testing: Constant high current density at elevated temperature on dedicated test structures, with continuous resistance monitoring to detect void growth
  • Dielectric breakdown testing: Constant voltage, constant current, or ramped stress on capacitor structures, with time or charge to breakdown as the measured response
  • Hot carrier and BTI testing: Bias stress with periodic or on-the-fly parameter measurement, using fast measurement to capture recoverable components
  • Temperature cycling: Repeated excursions between temperature extremes with defined ramp rates and dwells, monitored by daisy-chain resistance for interconnect fatigue
  • Humidity testing: Temperature-humidity-bias and highly accelerated stress testing, the latter using pressurized steam to reach high humidity above the boiling point and shorten test time
  • Preconditioning: Moisture soak followed by reflow simulation, applied before other tests so that assembly-induced damage such as delamination is represented

JEDEC, IPC, and AEC standards define these methods so that results are comparable across suppliers, but each standard still requires the user to select conditions appropriate to the intended application.

Failure Analysis Integration

Failure analysis validates mechanism assumptions and supplies model parameters:

  • Mechanism confirmation: Physical and electrical analysis verifies which mechanism produced each test failure
  • Model refinement: Measured crack paths, void locations, and trap densities calibrate and correct mechanism models
  • Unexpected mechanisms: Analysis often reveals mechanisms that the test plan never anticipated, which is among the most valuable outputs of a qualification program
  • Root cause depth: Analysis distinguishes design, process, material, and application causes, which determines who must act on the finding
  • Feedback loop: Findings drive design changes, supplier corrective action, and revision of the next test plan

Failure analysis closes the PoF loop. Without it, accelerated testing measures time to an unidentified event and cannot support extrapolation.

Advanced PoF Methods

Advanced physics of failure applications address complex systems and support sophisticated reliability assessment.

Multi-Physics Simulation

Computational tools enable coupled analysis across physical domains:

  • Thermal-electrical coupling: Joule heating raises temperature, which raises resistance and dissipation, requiring iterative solution and revealing thermal runaway risk
  • Thermal-mechanical coupling: Temperature distribution drives thermal stress, and temperature-dependent material properties feed back into the stress solution
  • Electrical-chemical coupling: Bias and moisture drive the electrochemical reactions behind corrosion and migration
  • Finite element analysis: Detailed models of packages, boards, and joints resolve local strain that closed-form models cannot, at the cost of substantial meshing and material characterization effort
  • Model validation: Simulated strain, temperature, and life must be checked against measurement, since constitutive models for solder creep in particular carry large uncertainty

Simulation supports virtual qualification, in which candidate designs are ranked before hardware exists. Its accuracy depends entirely on the quality of the material data and boundary conditions supplied.

Prognostics and Health Management

PoF supports real-time reliability assessment and remaining-life estimation:

  • Sensor integration: Temperature, humidity, vibration, and electrical monitors capture the loads the unit actually experiences
  • Canary and precursor structures: Deliberately weakened structures or monitored parameters fail or drift before the functional circuit does, providing advance warning
  • Damage accumulation tracking: Sensed loads feed damage models continuously, converting field exposure into consumed life
  • Remaining useful life prediction: Projection of future damage under an assumed mission profile, with uncertainty bounds that widen as the horizon lengthens
  • Condition-based maintenance: Maintenance scheduled on measured condition rather than fixed intervals, which reduces both unnecessary removals and in-service failures

Prognostics turns PoF from a design-phase analysis into an operational capability, and it is most valuable where unscheduled failure is expensive or hazardous.

Competing Mechanism Analysis

Multiple mechanisms compete to cause failure in real systems:

  • Dominant mechanism identification: Determine which mechanism reaches its failure criterion first under the specific operating conditions
  • Condition dependence: Rankings shift with temperature, voltage, humidity, and duty cycle, so a single dominant mechanism rarely holds across an entire product line
  • Competing risks models: Statistical frameworks that combine several time-to-failure distributions and handle censoring by other causes
  • Design trade-offs: Mitigating one mechanism can worsen another, as when stiffer underfill improves fatigue life but raises die stress
  • Application-specific assessment: The mission profile, not the component data sheet, determines which mechanism governs

Understanding competition prevents over-engineering a non-limiting mechanism while the governing one remains untouched.

Statistical Integration

PoF and statistical methods combine for comprehensive reliability assessment:

  • Physics-informed priors: Mechanism models supply physically reasonable prior distributions for Bayesian analysis, which is valuable when field data are sparse
  • Degradation path modeling: Statistical models fitted to physics-based degradation functions rather than arbitrary curves
  • Population variability: Distributions of dimensions, material properties, and defect densities across a production population explain the observed spread in life
  • Uncertainty propagation: Monte Carlo simulation carries input uncertainty through the physics model to a life distribution
  • Model selection: Likelihood-based comparison of competing physical models against experimental data, rather than selection by convention

Physics supplies the functional form and statistics supplies the uncertainty. Together they produce assessments that are both more accurate and more defensible than either approach alone.

Limitations and Practical Challenges

Physics of failure is powerful but not universal, and honest practice acknowledges where it does not reach.

  • Defect-driven early failures: Mechanism models describe intrinsic wear-out. Infant mortality caused by manufacturing defects, contamination, or handling damage follows the defect population, which requires screening, process control, and statistical methods rather than physics models
  • Parameter availability: Models need material properties, dimensions, and process details that suppliers often treat as proprietary, so users frequently substitute literature values with unquantified error
  • Mission profile uncertainty: Predicted life is only as good as the assumed load history, and real usage often differs substantially from the design assumption
  • Model validity ranges: Every model has a stress range, geometry range, and material system over which it was calibrated; extrapolation beyond those bounds is unsupported
  • Software and system failures: PoF addresses hardware degradation and says nothing about design errors, specification gaps, or software faults, which dominate failure reports in many modern systems
  • Effort and expertise: Detailed mechanism analysis costs time and specialized skill, so it should be focused on the components and mechanisms that actually govern life

These limits argue for a combined program: physics of failure for wear-out and design margin, process control and screening for defects, and field data feedback to validate both.

Industry Applications

Physics of failure has been adopted across industries where reliability is critical and empirical methods prove inadequate.

Semiconductor Industry

PoF is fundamental to semiconductor reliability engineering:

  • Technology qualification: Each process node is qualified through mechanism-specific stress testing on dedicated structures before product design begins
  • Design rules: Electromigration current limits, dielectric field limits, and antenna rules translate mechanism models into constraints that designers can check automatically
  • Product qualification: JEDEC standards define the stress tests, conditions, and sample plans used for product release
  • Failure analysis: Root cause determination drives process and design corrective action across large volumes
  • Reliability simulation: Aging-aware circuit simulators apply BTI and hot carrier models to project timing and analog margin at end of life

The industry has built extensive PoF infrastructure because reliability constraints, not just lithography, increasingly limit how far a technology can be scaled and how aggressively it can be operated.

Aerospace and Defense

Long life requirements and harsh environments drive PoF adoption:

  • Extended life assessment: Aircraft, satellites, and ground systems must function for decades, far beyond any feasible life test, so extrapolation must rest on physical models
  • Extreme environments: Wide temperature ranges, vibration, vacuum, and radiation require mechanism-specific understanding rather than generic failure rates
  • Parts obsolescence: When an original component is discontinued, PoF supports assessment of substitutes and of the risk from counterfeit or refurbished parts
  • Fleet management: Usage monitoring and damage accumulation tracking support inspection intervals and life extension decisions across a fleet
  • Prediction frameworks: Recognition that handbook failure rates poorly represent modern hardware pushed defense programs toward documented, assumption-explicit prediction frameworks of the kind IEEE Std 1413 describes

Aerospace applications demonstrate the value of PoF precisely where empirical data cannot exist, namely for long service lives and small production quantities.

Automotive Electronics

Automotive reliability requirements increasingly rely on PoF methods:

  • Very low defect targets: Volumes in the millions mean that even single-digit parts-per-million failure rates produce significant field returns, which demands root cause understanding rather than aggregate statistics
  • Harsh environment: AEC-Q100 defines ambient temperature grades matched to mounting location, running from Grade 3 at −40 to +85 degrees Celsius through Grade 1 at −40 to +125 degrees Celsius to Grade 0 at −40 to +150 degrees Celsius
  • Long service life: Warranty exposure and service-life expectations measured in many years require physics-based extrapolation from qualification tests lasting weeks
  • Qualification standards: The AEC-Q100, AEC-Q101, and AEC-Q200 documents define mechanism-specific stress tests for integrated circuits, discrete semiconductors, and passive components
  • Mission profile testing: Qualification conditions derived from measured vehicle usage, including engine-off soak, cold start, and vibration spectra specific to the mounting location
  • Safety-related systems: Electrification and driver assistance place electronics in safety-critical roles, where quantified hardware failure rates feed functional safety analysis

Automotive electronics reliability has improved substantially through the application of physics of failure principles to design, supplier qualification, and mission-profile-based testing.

Summary

Physics of failure shifts reliability assessment from fitting curves to failure counts toward modeling the physical processes that make electronic systems fail. Identifying mechanisms at their root cause lets engineers design against the specific process that limits life, build tests that accelerate that process rather than an unrelated one, and extrapolate to use conditions on a physical rather than a statistical basis.

The dominant mechanisms in electronics are well characterized. Electromigration follows Black's equation with activation energies near 0.5 to 0.7 eV for aluminum and 0.8 to 1.0 eV for copper. Dielectric breakdown is described by the E-model at low fields and the 1/E model at high fields, with neither serving the thinnest oxides. Bias temperature instability follows a recoverable power law in time. Solder fatigue follows Coffin-Manson behavior with creep corrections. Away from the die, electrolyte loss in aluminum electrolytic capacitors and flex cracking in ceramic capacitors frequently set assembly life. Each model yields an acceleration factor, and each acceleration factor is valid only within its calibrated range.

Physics of failure does not replace statistical methods, and it does not address defect-driven early failures, specification errors, or software faults. Its proper role is to supply mechanistic structure for wear-out, which statistical analysis then quantifies and field data then validate. As electronics operate in more demanding environments under longer service-life expectations, that structure becomes increasingly necessary for achieving and demonstrating required reliability.

Related Topics