Component Reliability
This article is organized by component class - semiconductors, passives, connectors, and electromechanical parts - and by how each is selected and characterized. The same subject cut by failure mechanism and process, covering wear-out, solder joints, moisture, ESD, counterfeits, obsolescence, and radiation, is treated in Electronics Reliability.
Component reliability forms the foundation of electronic system dependability. Every electronic system, regardless of complexity, ultimately depends on the performance of individual parts: semiconductors, passive devices, connectors, and electromechanical elements. Understanding how these parts fail, which mechanisms drive their degradation, and how to select and apply them for maximum reliability is essential knowledge for reliability engineers and circuit designers alike.
Electronic components fail through mechanisms that depend on component type, materials, construction, and operating conditions. Semiconductor devices wear out through electromigration, hot carrier injection, bias temperature instability, and time-dependent dielectric breakdown. Passive components degrade through material aging, thermal stress, and environmental exposure. Connectors and solder joints fail through fatigue, corrosion, fretting, and intermetallic growth. Each mechanism has a characteristic dependence on stress and temperature, and that dependence determines both when a component will fail and what can be done about it.
Two ideas organize the material that follows. The first is that failure is physical: every failure mode traces back to a material or interface responding to a stress, and the stress-life relationship for that mechanism is what makes prediction, acceleration, and derating possible. The second is that most component failures in fielded equipment are not inherent wear-out but the consequence of misapplication, overstress, or a defect that screening failed to catch. Sound component engineering therefore combines mechanism knowledge with disciplined selection, derating, and qualification.
This article covers component reliability from fundamental failure mechanisms through practical application guidelines. It addresses the reliability characteristics of the major component families, discusses failure physics at the material and device level, and presents strategies for selection, derating, and qualification. Whether the product is consumer electronics, industrial equipment, or a safety-critical system, the same body of knowledge applies; only the required margins change.
Semiconductor Device Reliability
Integrated Circuit Failure Mechanisms
Modern integrated circuits contain billions of transistors, kilometers of interconnect, and many material interfaces, each a potential failure site. The dominant intrinsic wear-out mechanisms are electromigration, time-dependent dielectric breakdown, hot carrier injection, and bias temperature instability. JEDEC publication JEP122, Failure Mechanisms and Models for Semiconductor Devices, is the industry reference that defines these mechanisms and the acceleration models and activation energies used to extrapolate accelerated test data to use conditions.
Electromigration occurs when current flow through a metal interconnect transfers momentum from conducting electrons to metal atoms, driving those atoms in the direction of electron flow. Voids form where atoms depart and hillocks form where they accumulate; voids grow into open circuits, and hillocks can bridge adjacent conductors. The mean time to failure follows Black's equation, in which lifetime is inversely proportional to current density raised to a power near two and falls exponentially with temperature through an Arrhenius term. Reported activation energies are near 0.7 electron volts for aluminum interconnect and higher for copper, which is one reason the industry's transition to copper damascene metallization improved electromigration margin. Design rules cap the current density allowed in each metal layer, and layout tools check those limits automatically.
Time-dependent dielectric breakdown affects the gate dielectric of MOS transistors. Under electric field stress, defects accumulate in the dielectric until a percolation path forms and the insulator breaks down. Lifetime depends strongly on the field across the dielectric and on temperature, but the correct functional form is debated: field-driven and current-driven exponential models and a power-law voltage model each fit portions of the data, and the choice materially affects extrapolated lifetime at use voltage. As geometries scaled, dielectric thickness fell and operating fields rose, making dielectric breakdown a leading constraint on supply voltage. High-permittivity dielectrics with metal gates restored physical thickness at a given equivalent oxide thickness and improved breakdown margin.
Hot carrier injection occurs when carriers accelerated by the lateral field near the drain gain enough energy to be injected into the gate dielectric, where they are trapped or create interface states. The result is a threshold voltage shift and loss of drive current. In long-channel devices the worst-case stress condition occurs near the gate bias that maximizes substrate current, roughly half the drain voltage; in heavily scaled devices the mechanism shifts toward multiple-carrier, energy-driven degradation, and the classic substrate-current monitor no longer predicts it well. Reduced supply voltages, lightly doped drain structures, and careful selection of operating points mitigate hot carrier damage.
Bias temperature instability degrades threshold voltage under gate bias at elevated temperature. Negative bias temperature instability affects PMOS devices under negative gate bias and has been a first-order concern since the 130-nanometer generation; positive bias temperature instability emerged as a comparable concern for NMOS devices once high-permittivity gate stacks were introduced. Both involve a mix of charge trapping and interface-state generation. Degradation recovers partially and quickly when stress is removed, which means measurement delays of even milliseconds understate the damage. Modern characterization therefore uses on-the-fly or fast measure-stress-measure techniques, and circuit-level models account for duty cycle rather than assuming continuous stress.
Not all integrated circuit failures are wear-out. Single-event upsets caused by alpha particles emitted from package materials and by secondary particles from cosmic-ray neutrons flip stored bits without damaging the device. Soft error rate is quoted in failures in time per megabit and is measured with the accelerated methods of the JESD89 series. Because the mechanism is transient, mitigation is architectural rather than physical: error-correcting codes, parity, scrubbing, and redundancy. Confusing soft errors with hard failures during field return analysis is a common and expensive mistake.
Discrete Semiconductor Reliability
Discrete semiconductors, including diodes, transistors, thyristors, and power devices, have failure mechanisms that differ from integrated circuits because their geometries are larger and their power densities higher. Thermal management, current handling, and packaging dominate their reliability rather than transistor-level wear-out.
Die attach fatigue develops when thermal cycling repeatedly stresses the bond between the die and its substrate. The thermal expansion mismatch between silicon and common substrate materials produces shear strain that accumulates cycle by cycle, producing voids and cracks in the attach layer. The immediate symptom is rising thermal resistance, which raises junction temperature, which accelerates every other thermally activated mechanism, including further attach degradation. This positive feedback makes die attach fatigue a classic runaway failure. Soft solder attach has largely given way to silver sintering and diffusion soldering in demanding power applications because the sintered joint has a far higher melting point relative to its service temperature and creeps much less.
Wire bond fatigue affects both discrete devices and integrated circuits. Bond wires flex as thermal expansion moves the die relative to the package, and repeated flexing initiates fatigue cracks at the bond heel that end in wire lift or fracture. A second and slower mechanism is metallurgical: gold wire bonded to aluminum pads forms a sequence of gold-aluminum intermetallic phases, and diffusion imbalance produces Kirkendall voiding at elevated temperature, raising resistance and weakening the joint. Copper wire, now widespread for cost reasons, forms intermetallics more slowly but demands tighter bonding process control because it is harder than gold and can crater the pad.
Power devices are additionally characterized by power cycling rather than passive thermal cycling. In power cycling the die heats itself during each load pulse, so the junction swings through a large temperature range while the package base stays comparatively cool. Manufacturers of power modules publish power-cycling capability curves relating cycles to failure against junction temperature swing; bond wire lift-off and die attach fatigue are the two mechanisms those curves capture. Reducing the junction temperature swing is usually far more effective than reducing the mean temperature.
Power MOSFETs face several device-specific limits. Gate oxide damage accumulates when transient overvoltage stresses the thin gate dielectric, so gate drive design, clamping, and layout that minimizes gate loop inductance all protect it. Unclamped inductive switching drives the device into avalanche, and the single-pulse avalanche energy rating bounds how much energy the die can absorb without destruction. Silicon carbide MOSFETs deserve particular attention: the lower conduction band offset between silicon carbide and silicon dioxide and the higher defect density at that interface make gate oxide reliability and threshold voltage instability active engineering concerns, and body-diode conduction can expand basal plane dislocations in some material grades. Manufacturers address these with specific gate voltage limits and qualification data that designers should read rather than assume.
Insulated gate bipolar transistors and thyristors face challenges rooted in their bipolar conduction. Older planar IGBT designs could latch a parasitic thyristor and lose gate control, though modern trench and field-stop designs are essentially latch-up immune within their rated operating area. Safe operating area limits must still be respected to prevent secondary breakdown, in which current concentrates in a local filament and drives thermal runaway. Short-circuit withstand time, typically specified in microseconds, defines how long protection has to act before the device is destroyed.
Semiconductor Package Reliability
The package provides mechanical protection, electrical connection, and a thermal path. It is frequently the limiting factor for device reliability, particularly where thermal cycling, humidity, or mechanical stress dominate the use environment.
Package delamination occurs when interfaces within the package separate because of moisture absorption, thermal stress, or inadequate adhesion. Delamination breaks wire bonds, raises thermal resistance, and opens paths for moisture ingress. Its most dramatic form is the popcorn effect, in which moisture absorbed by a plastic package flashes to steam during reflow soldering and cracks the package. The industry controls this through moisture sensitivity levels: J-STD-020 classifies a package by the floor life it can tolerate at factory ambient before reflow, and J-STD-033 defines the corresponding handling, bagging, and bake requirements. Lead-free reflow made the problem worse, because peak package temperatures rose from roughly 220 degrees Celsius to between 245 and 260 degrees Celsius depending on package thickness and volume.
Plastic package reliability depends heavily on the molding compound. The compound must adhere to the die, leadframe, and bond wires while resisting moisture and conducting heat. Expansion mismatch between the compound and the other package materials generates stress that can crack the compound, shear the passivation, or shift the electrical parameters of stress-sensitive circuits. Low-stress, low-ionic-contamination formulations and die coat layers have substantially improved plastic package reliability, to the point that plastic parts now serve in applications once reserved for hermetic packages.
Area array packages such as ball grid arrays concentrate the reliability question at the solder interconnect, which serves as both electrical connection and mechanical attachment. Thermal cycling fatigues these joints, with the highest strain and the earliest failures at the balls farthest from the neutral point of the package. Underfill redistributes that strain across the whole die area and can extend thermal cycling life by an order of magnitude; corner staking and edge bonding provide a lower-cost partial benefit aimed mainly at mechanical shock and drop.
Hermetic packages of ceramic or metal provide the best protection against moisture and are still required for many space, military, and high-temperature applications. Their characteristic failure is loss of seal. MIL-STD-883 Method 1014 defines fine and gross leak testing, with helium tracer gas the usual fine-leak medium, and Method 2020 defines particle impact noise detection to find loose conductive particles sealed inside the cavity. Hermeticity is not absolute: a cavity package with an internal getter or a low-moisture seal ambient still admits moisture slowly through seals and glass-to-metal feedthroughs over long service lives.
Semiconductor Quality and Screening
Semiconductor quality programs combine process control, testing, and screening to deliver consistent and predictable reliability. The goal is twofold: hold the intrinsic capability of the process, and remove defective units before they reach a customer.
Wafer fabrication quality rests on process control, contamination management, and systematic defect reduction. Statistical process control monitors critical parameters and triggers action when they drift. Cleanroom discipline and rigorous chemical and gas purity limit particle and metallic contamination. In-line defect inspection and yield-to-reliability correlation identify the defect types that survive test but shorten life.
Electrical testing at wafer probe screens functional failures and parametric outliers. Beyond simple specification limits, statistical screening removes parts that are within specification but far from their lot population, on the reasoning that an outlier is evidence of an unusual defect. The automotive industry formalized this as part average testing and statistical yield analysis, and applies it aggressively because it catches latent defects that functional test cannot see.
Burn-in applies elevated temperature and voltage to precipitate infant mortality failures, following the approach codified in MIL-STD-883 Method 1015. Dynamic burn-in exercises device functionality during stress and detects more failure modes than static bias alone. The economics have shifted, however: as defect densities fell and die cost rose, blanket burn-in became difficult to justify for many mature commercial products, and manufacturers increasingly substitute targeted stress, statistical screening, and improved process control. Burn-in remains standard where the cost of an early field failure is high.
Qualification and screening for military, aerospace, and medical applications add requirements beyond commercial practice. MIL-PRF-38535 defines the qualified manufacturers list system for monolithic microcircuits, with quality assurance classes that culminate in the space-level class, and MIL-PRF-19500 performs the same role for discrete semiconductors. Screening flows drawn from MIL-STD-883 add temperature cycling, constant acceleration, seal testing, particle impact noise detection, and radiographic inspection. These screens raise unit cost substantially, and they are justified by consequence rather than by quantity.
Passive Component Reliability
Resistor Reliability
Resistors are among the most reliable electronic components when properly selected and applied, but they are not failure free. Their failure modes divide into drift, which degrades circuit accuracy, and catastrophic open or short circuits, which are usually the result of overstress.
Thick film chip resistors dominate surface mount designs on cost. Their resistive element is a fired ruthenium oxide glass composite, trimmed to value with a laser cut that leaves a local stress concentration where cracks can start under thermal cycling or board flex. Two application-specific weaknesses deserve attention. First, thick film elements tolerate short pulses poorly relative to their average power rating, because the pulse energy is dissipated in a very thin film; pulse-withstanding constructions exist for applications such as inrush limiting and snubbers. Second, the silver-bearing inner termination of a standard chip resistor reacts with atmospheric sulfur to form non-conductive silver sulfide, producing open circuits in equipment exposed to industrial or agricultural atmospheres. Sulfur-resistant constructions verified by flowers-of-sulfur exposure testing address this, and they should be specified deliberately rather than discovered after a field failure.
Thin film resistors offer superior stability, lower temperature coefficient, and lower noise than thick film types, and their sputtered nickel-chromium or tantalum nitride films are less affected by moisture than thick film pastes. The trade is robustness: thin film elements are more sensitive to electrostatic discharge and to transient overstress, and their very thin cross section leaves little thermal mass. Tantalum nitride films are specified where humidity resistance matters most, because the film self-passivates.
Wirewound resistors handle high power and pulse energy but bring their own limits. The wound element has significant series inductance that disqualifies it from high-frequency and fast-switching applications unless a non-inductive winding is used. Thermal cycling fatigues the wire and its end connections, and the resistance alloy oxidizes if the protective coating or cement is compromised. Surface temperatures in wirewound power resistors routinely exceed 200 degrees Celsius at rated dissipation, which constrains everything mounted nearby.
Carbon composition resistors, now largely obsolete for new designs, illustrate the principles by counterexample. Their bulk carbon element absorbs moisture, drifts irreversibly with humidity and heat, and can open outright. Those characteristics drove the industry toward film technologies. Their one surviving advantage, high pulse energy absorption in a bulk element with no film to vaporize, keeps them in a few niche applications.
Across all types, the dominant reliability lever is temperature. Resistor life and stability specifications are quoted at a defined ambient and load; the standard power derating curve holds full rated power to roughly 70 degrees Celsius and then falls linearly to zero at the maximum rated ambient. Operating on that curve is a rating limit, not a reliability target.
Capacitor Reliability
Capacitors span a wider reliability range than any other passive family. Some technologies serve for decades without measurable degradation while others have a defined, temperature-driven end of life. Failure modes include short circuits, open circuits, capacitance loss, rising equivalent series resistance, and increased leakage.
Multilayer ceramic capacitors dominate modern electronics. Their behavior depends on dielectric class. Class 1 dielectrics such as C0G, also written NP0, have a temperature coefficient near zero within a few tens of parts per million per degree Celsius, negligible aging, and no meaningful voltage coefficient; they are the choice for timing, filters, and precision analog. Class 2 dielectrics such as X7R, rated from minus 55 to plus 125 degrees Celsius within 15 percent, and X5R, rated from minus 55 to plus 85 degrees Celsius within the same tolerance, deliver far higher capacitance density at the cost of pronounced voltage and temperature coefficients. The effective capacitance of a small-case Class 2 part can fall well below half its nominal value at rated direct voltage, a loss that catches designers who size decoupling or bulk capacitance from the marked value. Class 2 dielectrics also age logarithmically, losing a few percent of capacitance per decade of hours after firing, and recover their initial value only if heated above the Curie point.
The characteristic mechanical failure of ceramic capacitors is flex cracking. Board bending during depaneling, connector insertion, test fixturing, or mounting propagates a crack from the termination into the dielectric, producing a short, an intermittent, or a slow leakage rise that can end in thermal runaway. Mitigations are well understood: orient parts away from high-strain axes and board edges, keep them clear of depaneling routes and mounting hardware, use smaller case sizes, and specify flexible-termination parts whose compliant conductive polymer layer decouples the ceramic from board strain.
Aluminum electrolytic capacitors have a defined wear-out life governed by loss of electrolyte through the end seal. As electrolyte escapes, capacitance falls and equivalent series resistance rises, and the part eventually fails to perform its filtering function long before it fails electrically. Life follows an Arrhenius relationship commonly expressed as a doubling for every 10 degree Celsius reduction in core temperature within the rated range, so a capacitor rated 2,000 hours at 105 degrees Celsius offers roughly 32,000 hours at 65 degrees. The temperature that matters is the internal core temperature, which includes self-heating from ripple current in the equivalent series resistance, not the board ambient. Electrolytics also require voltage derating, and parts stored unpowered for years may need controlled reforming before full voltage is applied.
Solid tantalum capacitors with manganese dioxide cathodes offer excellent volumetric efficiency and stability but can fail energetically. A defect in the tantalum pentoxide dielectric produces local leakage; the manganese dioxide normally self-heals by converting to an insulating lower oxide, but if the circuit can supply enough current the site instead heats until the tantalum ignites. This is why classic practice derates manganese dioxide tantalum capacitors to half of rated voltage and why series resistance, typically expressed as a minimum ohms-per-volt of circuit impedance, is specified in low-impedance positions. Polymer-cathode tantalum capacitors replace the manganese dioxide with a conductive polymer that fails benignly rather than igniting, and manufacturers accordingly permit substantially less aggressive voltage derating.
Film capacitors offer the best combination of reliability and stability for high-voltage, high-ripple, and precision applications. Polypropylene and polyester dielectrics have low dissipation factor and stable capacitance, and metallized constructions self-heal by vaporizing the thin electrode metallization around a fault, isolating it. The design consequence is graceful degradation: a metallized film capacitor loses a small amount of capacitance with each clearing event rather than failing abruptly, and end of life is conventionally defined as a capacitance loss of a few percent. Film capacitors dominate direct-current link, snubber, and alternating-current applications where an electrolytic would not survive the ripple or a ceramic would not survive the voltage transients.
Inductor and Transformer Reliability
Inductors and transformers combine a magnetic core, a winding, and an insulation system, and each element has its own limits. High current, elevated temperature, and magnetic saturation all degrade reliability if the component is misapplied.
Insulation defines the thermal life of a wound component. IEC 60085 classifies insulation systems by their permitted hot-spot temperature, with the common classes rated at 130, 155, and 180 degrees Celsius. Thermal aging of organic insulation follows an Arrhenius relationship, so a persistent hot spot shortens life sharply even when the average temperature looks acceptable. Uneven current distribution, poor impregnation, and proximity to a hot semiconductor all create such hot spots. In high-voltage transformers, partial discharge in voids within the insulation erodes material progressively until breakdown occurs, which is why impregnation quality and partial discharge testing matter for those designs.
Core saturation causes inductance to collapse, and the resulting current spike can destroy switching devices within microseconds. The trap is temperature dependence: the saturation flux density of manganese-zinc power ferrite falls substantially between room temperature and operating temperature, with common power grades near 490 millitesla at 25 degrees Celsius and near 390 millitesla at 100 degrees Celsius. A converter verified on a cold bench can therefore saturate once it is hot. Direct-current bias, transformer flux walking from asymmetric drive, and load transients all reduce the available margin further.
Core losses generate heat that must be removed. Hysteresis loss scales with frequency and with flux swing raised to a power somewhat above two, and eddy current loss scales with the square of frequency for a given flux density, which is why laminations are thinned and ferrites are used as frequency rises. Losses also depend on temperature, and some ferrite grades have a loss minimum near their intended operating temperature; operating far from that point wastes margin. Combined core and winding losses set the temperature rise, and the temperature rise sets the insulation life, closing the loop back to the previous paragraph.
Mechanical robustness matters in vibration and shock environments. Wire breakage at the termination, core cracking from mechanical or thermal stress, and gap fretting in gapped cores all appear in the field. Magnetostriction in the core also produces audible noise and mechanical stress at switching frequencies within the audible band. Varnish impregnation, mechanical staking, and proper mounting address these problems.
Passive Component Derating
Derating means operating a component below its maximum ratings so that the applied stress leaves margin against both variation and degradation. For passive components, derating addresses power dissipation, voltage, current, and temperature. The appropriate depth of derating depends on the reliability requirement and on the specific failure mechanism the derating is intended to suppress.
Resistor derating primarily targets power dissipation and the resulting temperature rise. A widely used commercial guideline limits dissipation to half of the rated value at the maximum expected ambient, applied on top of the manufacturer's derating curve rather than instead of it. High-reliability practice derates further. Voltage derating matters separately for high-value resistors, where the maximum working voltage rather than the power rating is the binding limit.
Capacitor derating addresses voltage and temperature together, and the appropriate factor is strongly technology dependent. Aluminum electrolytics are commonly held to roughly 80 percent of rated voltage and operated well below rated temperature, because temperature directly buys life. Manganese dioxide tantalum capacitors are conventionally derated to 50 percent of rated voltage because of their ignition failure mode. Class 2 ceramics require attention to a subtler point: the practical derating is driven less by breakdown than by the loss of capacitance under direct bias, so the designer must derate the capacitance value as well as the voltage. Film capacitors in alternating-current or pulse service are limited by root-mean-square current and by the rate of voltage change rather than by working voltage alone.
Inductor derating considers direct current, root-mean-square current, and temperature. Operating below the saturation current rating preserves inductance under transients, and operating below the thermal current rating limits temperature rise. These are two different ratings with two different failure consequences, and datasheets list both. Switching frequency must also be accounted for, because core loss grows faster than linearly with frequency and consumes thermal budget that would otherwise be available for winding loss.
Derating rules should be documented in an organizational standard rather than left to individual judgment, so that decisions are consistent, reviewable, and inherited by the next project. Waivers are legitimate when analysis justifies them, and they should be recorded with that analysis attached.
Connector and Interconnect Reliability
Connector Contact Reliability
Connectors provide separable electrical connections and are inherently less reliable than permanent joints. Contact reliability is governed by conditions at the contact interface: normal force, contact geometry, surface finish, plating, and the environment. Connector performance is characterized by the EIA-364 series of test procedures, which define the mating, durability, vibration, thermal, and corrosion tests used to qualify a design.
Contact resistance arises from the constriction of current through the small asperities where the mating surfaces actually touch, plus the resistance of any film covering them. Adequate normal force is what makes the interface work: it deforms asperities to create metallic contact, provides wiping action during mating that displaces films, and maintains the pressure that keeps contaminants out. Insufficient normal force is the root cause behind a large share of connector problems, and it can result from spring relaxation at elevated temperature as easily as from an original design deficiency.
Fretting corrosion is the dominant failure mechanism for tin-plated separable contacts. Small relative motions, on the order of micrometers and driven by thermal expansion or vibration, repeatedly break the tin surface and expose fresh metal, which oxidizes. The oxide debris accumulates in the interface and drives resistance up in an erratic, intermittent pattern that is notoriously hard to diagnose. Higher normal force, contact geometry that resists micromotion, contact lubricants, and strain relief that keeps the connector from moving all mitigate fretting; gold-on-gold interfaces largely avoid it because gold does not form an insulating oxide.
Plating selection is therefore a reliability decision rather than a cost decision alone. Gold over a nickel underplate provides stable low resistance and corrosion resistance; durable gold thicknesses on the order of 0.75 micrometers suit high mating-cycle or harsh-environment service, while thin flash gold is appropriate only for low-cycle, benign applications. ASTM B488 classifies gold plating by type, grade, and thickness for this purpose. Tin is economical and performs well in stable, low-cycle applications with adequate normal force, but is fretting-prone and is generally not used for low-level signal contacts. Palladium and palladium-nickel with a gold flash occupy the middle ground. Mixing plating types across a mating pair is poor practice, because it creates a galvanic couple and gives the harder surface an opportunity to wear through the softer one.
Environment closes the picture. Humidity enables electrochemical corrosion, particularly where dissimilar metals meet. Elevated temperature accelerates oxidation, promotes intermetallic growth at plating interfaces, and relaxes contact springs. Industrial atmospheres bearing sulfur and chlorine compounds attack silver and copper alloys and creep corrosion products across surfaces. Mixed flowing gas testing reproduces these atmospheres in accelerated form and is the standard means of qualifying a connector for a corrosive environment. Sealed connectors rated to an appropriate ingress protection level are required where liquid or dust exposure is expected.
Printed Circuit Board Interconnect
The printed circuit board is the interconnection substrate for nearly every electronic assembly, and it is a reliability item in its own right. Board reliability depends on laminate properties, copper integrity, and plated hole quality, and is specified by acceptance classes: IPC-6012 defines Class 2 for general dedicated service and Class 3 for high-reliability equipment where continued performance is essential.
Conductive anodic filament formation occurs when copper ions migrate along a degraded glass-to-resin interface under a voltage gradient in the presence of moisture, forming a conductive filament that shorts adjacent features. It appears most often between closely spaced plated holes, since that is where the field is strongest and the fiber path shortest. Susceptibility depends on resin chemistry, glass weave and finish, drilling quality, and hole-to-hole spacing; poor drilling that damages the glass bundle is a frequent contributor. CAF-resistant laminates and CAF testing per the IPC-TM-650 test method series address the risk, and spacing rules provide the design margin.
Plated through-hole reliability is governed by the ability of the copper barrel to accommodate the mismatch between the low in-plane expansion of the laminate and its much larger out-of-plane expansion, which is especially severe above the glass transition temperature. Repeated excursions above that temperature during assembly and rework, and repeated thermal cycling in service, crack barrels and separate inner layer connections. IPC-6012 sets minimum average copper barrel thickness at 20 micrometers for Class 2 and 25 micrometers for Class 3, and interconnect stress testing or thermal shock coupons verify that production boards achieve the intended cycling capability. High glass transition temperature and low expansion laminates were adopted largely because lead-free assembly temperatures made this mechanism worse.
Delamination separates laminate layers under moisture, thermal stress, or inadequate lamination. Absorbed moisture is the usual trigger, flashing to steam during reflow and driving the layers apart. Bare boards absorb moisture in storage exactly as packaged components do, and IPC-1601 provides the corresponding handling, storage, and bake guidance; moisture sensitivity level classification under J-STD-020, by contrast, applies to packaged components rather than to bare boards. Conflating the two leaves a real exposure unmanaged.
Electrochemical migration can occur between adjacent surface conductors when moisture, ionic contamination, and a voltage gradient coincide. Metal dissolves at the anode, migrates through the moisture film, and plates as a dendrite at the cathode until the dendrite bridges the gap. Fine-pitch assemblies with no-clean flux residues are particularly exposed, since residues that are benign when fully activated can be hygroscopic and ionic when they are not. Surface insulation resistance testing is the standard means of qualifying a materials and process combination, and cleanliness control and conformal coating are the standard mitigations.
Solder Joint Reliability
Solder joints make the permanent connection between components and the board, and in many products they are the single largest contributor to wear-out failure. Their reliability depends on joint geometry, alloy properties, and the thermal and mechanical stress history of the assembly. The transition from tin-lead to lead-free alloys changed the relevant material properties, and its long-term consequences are still being characterized.
Thermal cycle fatigue dominates in most applications. Expansion mismatch between component and board imposes a shear strain on the joint, and because solder creeps readily at ordinary service temperatures, much of that strain is inelastic and accumulates as damage. Cycles to failure follow a Coffin-Manson relationship in which life falls as a power of the inelastic strain range, with the Norris-Landzberg formulation adding the frequency and temperature dependence needed to relate accelerated tests to field conditions. Practical consequences follow directly: large, stiff, leadless packages fail sooner than small or compliant ones; the balls or terminations farthest from the neutral point fail first; and dwell time at temperature matters as much as the temperature extremes, because creep needs time. IPC-9701 is the standard thermal cycling test method for surface mount solder attachments and defines the cycle profiles used for comparison.
Intermetallic compound growth is a slower mechanism operating in parallel. Copper-tin intermetallics form during soldering and continue to thicken by diffusion throughout service, faster at higher temperature. A thin intermetallic layer is necessary for a sound joint; a thick one is brittle, cracks under mechanical shock, and consumes the copper of the pad or termination. This is the mechanism behind pad cratering and brittle interfacial fracture in drop testing, and the reason that thermal aging is included in mechanical shock qualification.
Lead-free alloys, principally the tin-silver-copper family typified by the composition with 3.0 percent silver and 0.5 percent copper, melt near 217 to 220 degrees Celsius against 183 degrees for eutectic tin-lead. They are stronger and stiffer but less ductile, which makes them better in some thermal cycling regimes and worse in mechanical shock, and their higher stiffness transmits more stress into the component and the board. Lower-silver alloys were introduced to improve drop performance and reduce cost, trading away some thermal cycling capability. No single alloy is best across all stress types, so alloy selection should follow the dominant field stress.
Two tin-specific phenomena deserve mention. Tin pest is the transformation of ordinary white tin to a brittle gray allotrope below 13.2 degrees Celsius, accompanied by a large volume increase; it is thermodynamically possible but rarely observed in practice, because nucleation is very slow and the alloying elements in ordinary solders suppress it. Tin whiskers are a genuine and current concern: pure tin finishes spontaneously grow conductive single-crystal filaments that can short closely spaced conductors, a mechanism that has caused documented failures in high-reliability systems. Whisker susceptibility is evaluated under JESD201, and mitigation for aerospace and other high-performance systems is managed under GEIA-STD-0005-2, with control plans that variously specify a nickel underlayer, matte rather than bright tin, annealing, conformal coating, or the deliberate addition of lead.
Joint design remains a first-order lever. Standoff height, pad geometry, solder volume, and fillet shape all determine how strain distributes through the joint. Non-solder-mask-defined pads generally outperform solder-mask-defined pads in fatigue because the stress concentration moves away from the mask edge. Compliant leads accommodate strain that a leadless termination transmits directly into the solder. These choices are made at layout time and are difficult to correct afterward.
Electromechanical Component Reliability
Relay Reliability
Electromechanical relays provide galvanic isolation, very low on-state resistance, and load switching capability that solid-state devices match only at a cost. Their moving parts bring wear-out mechanisms that place a finite bound on life. Relay datasheets accordingly quote two separate life figures: mechanical life with no load, often tens of millions of operations, and electrical life at a defined load, frequently one hundred thousand operations or fewer. The gap between the two shows how completely the load dominates relay lifetime.
Contact erosion occurs as arcs at make and break melt and vaporize contact material and transfer it between contacts. Inductive loads sustain the arc with stored energy and are far more damaging than resistive loads of the same current; capacitive loads and lamp loads punish the contacts on make instead, with inrush currents that can weld them. Arc suppression across the load or the contacts reduces arc energy and extends life substantially, at the cost of slower load turn-off.
Contact material is selected against the load. Silver tin oxide has largely replaced silver cadmium oxide under restrictions on cadmium and resists welding well under inrush. Silver nickel and fine silver serve general purpose loads, and refractory alloys handle very high currents. Low-level signal switching is a different problem entirely: a contact that has never carried enough current to burn through surface films may fail to conduct at millivolt and microampere levels, so gold-plated or bifurcated contacts are specified for dry circuits, and such contacts must never be used at high current, which would destroy the plating.
Mechanical wear limits life even when the electrical load is trivial. Pivot and bearing wear, spring relaxation, and armature fatigue accumulate with operations. Contact contamination is a parallel path to failure: organic vapors from adjacent materials can polymerize on contacts under arcing to form insulating films, and metallic debris from erosion can bridge or jam the mechanism. Sealed relays exclude external contamination but cannot remove what is generated internally.
Coil reliability depends on the insulation system and the operating temperature. Coil dissipation raises the winding temperature above ambient, and the insulation ages accordingly, ending in turn-to-turn shorts. A simple flyback diode across the coil is the usual drive protection, but it extends the release time and prolongs the arc at the contacts, so a resistor or Zener element in series with the diode is often preferred where contact life matters more than electrical quiet.
Switch Reliability
Manual switches and pushbuttons share the contact physics of relays and add the demands of a human interface. Their reliability is set by contact material, actuation mechanism, and exposure, and their published life ratings assume a specified load.
Contact bounce produces several make-and-break events during a single actuation, typically over a few milliseconds, which a digital input will register as multiple transitions unless the signal is debounced in hardware or software. Bounce duration generally worsens as the switch wears and contact surfaces roughen. Bounce is also a reliability matter, not merely a nuisance, because each bounce event arcs and erodes the contacts.
Actuation mechanisms wear predictably. Sliding contacts abrade, return springs take a permanent set, detents round over, and snap-action mechanisms lose their over-center crispness and dwell in an ambiguous state. Membrane and dome switches fail through tactile dome fatigue and through cracking of printed silver traces at flex points. Higher-grade materials and tighter tolerances buy longer actuation life, and the datasheet life rating should be compared against the expected number of actuations over the product lifetime, which is frequently underestimated.
Environmental exposure is usually more severe for switches than for other components, because switches must reach the outside of the enclosure. Panel-mounted controls meet dust, liquids, cleaning agents, and mechanical abuse. Sealed constructions and switches with an appropriate ingress protection rating under IEC 60529 address this, trading tactile feedback and cost for protection. Sealing the switch does not seal the panel, and the panel cutout is a common leak path that the switch specification does not cover.
Crystal and Oscillator Reliability
Quartz crystals and crystal oscillators provide the frequency references on which digital systems, communications, and instrumentation depend. Their reliability question is unusual, because the dominant concern is not catastrophic failure but slow, specified drift.
Aging is that drift. Mass transfer of contamination onto or off the resonator surface and stress relaxation in the mounting structure shift the resonant frequency gradually over time. The aging rate is highest immediately after manufacture and decreases roughly logarithmically thereafter, so datasheets quote a first-year figure, commonly a few parts per million for standard commercial crystals, with much lower rates for oven-controlled oscillators. Where long-term accuracy matters, designers either pre-age the part, budget for the cumulative drift over the service life, or provide for periodic calibration or disciplining from an external reference.
Activity dips are narrow increases in equivalent series resistance and frequency perturbation that appear at particular temperatures, caused by coupling between the desired thickness-shear mode and an unwanted mode whose temperature coefficient differs. A severe dip can stop an oscillator at one temperature while it runs correctly on either side of it, which makes the resulting field failures look intermittent and temperature-dependent rather than component-related. Crystal design and screening minimize dips, and adequate oscillator loop gain margin tolerates the ones that remain.
Package seal integrity is critical. A leak admits moisture and contamination onto the resonator surface, causing frequency shifts far larger than normal aging and eventually stopping oscillation. Hermetic ceramic or metal packages with seam-welded or glass seals provide the necessary protection, and seal testing verifies integrity in manufacturing. Crystals are also sensitive to mechanical shock, which can fracture the blank or shift the mount, and to board stress transmitted through the mounting pads.
The oscillator circuit is part of the reliability equation. Excessive drive level heats the resonator locally, accelerates aging, and can fracture the blank, so small surface mount AT-cut crystals are typically rated for drive on the order of tens of microwatts and a series resistor is often added to limit it. Insufficient loop gain leaves the circuit unable to start reliably at cold temperature or with a crystal at the high end of its resistance tolerance; the usual design rule requires the circuit's negative resistance to be at least five times the crystal's maximum equivalent series resistance across all conditions. Load capacitance must also match the crystal's specification, or the oscillator will run off frequency regardless of how good the crystal is.
Component Selection for Reliability
Reliability Data Sources
Component selection for reliability requires reliability data, and every available source has both a use and a limitation. Interpreting the data correctly matters more than obtaining it.
Manufacturer data is the most direct source. Datasheets may quote failure rates in failures in time with a stated confidence level, and reliability reports document qualification results, high-temperature operating life data, and the activation energy assumed in the extrapolation. Application notes frequently contain the most useful material of all, because they describe the failure modes the manufacturer has actually seen. The quality and completeness of this data varies widely, and a failure rate quoted without its test conditions, sample size, confidence level, and activation energy cannot be compared with any other number.
Handbook prediction methods provide consistent, repeatable estimates and are often contractually required, but their currency varies. MIL-HDBK-217 remains widely cited even though its last released revision, Notice 2 of revision F, is dated 28 February 1995; a revision G has been under discussion for many years without being published, so the models in current use predate essentially all modern semiconductor technology. Telcordia SR-332, whose Issue 4 was released in 2016, provides a parts-count and parts-stress method oriented toward telecommunications equipment and can incorporate a manufacturer's own laboratory and field data. The FIDES Guide, developed in France and issued in a 2022 edition, is explicitly built around physics-of-failure reasoning and mission profiles and accounts for the process and manufacturing quality of the equipment, not only the parts. IEC 61709 supplies reference conditions and stress models for converting failure rates between operating conditions. 217Plus, developed from the earlier PRISM methodology, updates the military handbook approach with process-grading factors.
Every handbook method shares a structural weakness: each assumes a constant failure rate and produces a single number that looks more authoritative than it is. Handbook predictions are useful for comparing design alternatives, allocating requirements, and sizing spares. They are unreliable as absolute forecasts of field performance, and they cannot represent wear-out mechanisms at all.
Field failure data from an organization's own products is the most relevant source available, because it reflects the actual use environment, duty cycle, and manufacturing process. Its limitations are practical rather than conceptual: returns data undercounts failures that users tolerate or discard, the lag between shipment and return distorts trend analysis, and root cause is often never determined because the returned unit is scrapped. A closed-loop failure reporting and corrective action system is what converts field returns from an expense into information.
Qualification and characterization testing generates data specific to the application. Accelerated life testing produces failure rate and lifetime estimates within a practical schedule, and mechanism-specific testing quantifies the particular wear-out modes that matter. Such data is expensive to generate and is justified where the component is novel, the application stresses are outside the manufacturer's qualification envelope, or the consequences of failure are severe.
Component Qualification
Component qualification verifies that a part meets reliability requirements for its intended application. Qualification may proceed by test, by analysis, by similarity to an already qualified part, or by a combination of these. The appropriate approach depends on application criticality, the novelty of the component, and the data already available.
Industry-standard qualification flows provide the baseline. JESD47 defines stress-test-driven qualification for integrated circuits, drawing its stress conditions from the JESD22 test method series and its acceleration models from JEP122. The Automotive Electronics Council standards impose considerably more demanding flows: AEC-Q100 for integrated circuits, AEC-Q101 for discrete semiconductors, AEC-Q102 for optoelectronics, and AEC-Q200 for passive components, each with temperature grades running from minus 40 to plus 85, 105, 125, and 150 degrees Celsius. Selecting an automotive-qualified part is often the most economical way to obtain a documented, deeply tested component for an industrial or high-reliability application.
New component qualification applies where no relevant experience base exists. The plan should verify performance across the full operating envelope and demonstrate life against the specific failure mechanisms that the component's materials and construction make credible. A qualification plan assembled without a mechanism analysis tends to test what is easy rather than what matters.
Qualification by similarity leverages experience with a comparable part, restricting the test effort to the differences. This is efficient and legitimate, and it fails when the assessment of similarity is superficial. A die shrink, a change of assembly site, a new molding compound, or a change of lead finish can each invalidate the comparison while leaving the part number and datasheet unchanged. Product change notification agreements with suppliers exist precisely so that these changes are visible.
Second-source qualification applies when an alternative supplier is added. Parts that are electrically interchangeable are frequently not interchangeable in reliability, because materials, die design, and assembly processes differ. The second source requires its own qualification against the same application conditions, and its lots should be tracked separately in the field so that a difference in performance can be detected.
Ongoing lot verification maintains confidence through production. Periodic reliability monitoring, lot acceptance testing where required, and statistical sampling plans balance cost against confidence. The essential requirement is a mechanism that detects a change in the incoming population before it reaches the customer, and an escalation path when it does.
Derating Standards and Guidelines
Derating standards convert the general principle of stress margin into specific, checkable limits for each component type. Formal derating documents come principally from the space and defense sectors, where the cost of failure justifies the effort of writing and enforcing them.
MIL-STD-1547, Electronic Parts, Materials, and Processes for Space and Launch Vehicles, is the classic military derating source and specifies derating requirements along with end-of-life limits, mounting, and process requirements. Its revision B was issued in December 1992, cancelled in 1997, and reinstated by the Defense Standardization Council in 2008; the companion MIL-HDBK-1547 carries the supporting technical guidance. NASA's EEE-INST-002, Instructions for EEE Parts Selection, Screening, Qualification, and Derating, provides derating tables tied to defined parts levels and is accompanied by a derating analysis workbook. In Europe, ECSS-Q-ST-30-11C sets derating requirements for electrical, electronic, and electromechanical components, with the current revision issued in June 2021, and scales the requirements against mission duration and mean temperature. For naval electronics, the NAVSEA parts derating requirements manual TE000-AB-GTP-010 supplies derating curves for the most commonly used part families.
Commercial and automotive practice is less centralized. No general-purpose commercial derating standard has the standing that the space and defense documents have in their domains; instead, most commercial organizations write their own derating policy, often adapting the space and defense tables and relaxing them where the application permits. Automotive practice approaches the same objective from a different direction, specifying a mission profile and requiring the supplier to qualify the part against it under the applicable AEC-Q document, so that the margin is demonstrated by test rather than asserted by a percentage. Component manufacturers also publish derating guidance specific to their technologies, and for mechanisms such as tantalum ignition or ceramic capacitance loss under bias, that guidance is more useful than any general table.
Company-specific standards tailor these sources to an organization's products, environments, and field experience. They should be reviewed periodically, because component technology changes: guidance written for through-hole parts in a ventilated chassis does not transfer unmodified to a dense surface mount assembly in a sealed enclosure, and derating rules that ignore the direct-current bias behavior of modern ceramic capacitors will mislead.
Derating analysis verifies compliance. Each component is evaluated against the applicable limits at worst-case combinations of supply voltage, load, tolerance, and ambient temperature, using thermal analysis or measurement to establish the actual component temperature rather than the board ambient. Circuit simulation and automated stress analysis tools handle much of the arithmetic. Violations should be resolved by component change or circuit redesign, and accepted only through a documented waiver that records the analysis and the residual risk.
Critical Component Management
Critical components are those whose failure would produce unacceptable consequences: a safety hazard, a mission failure, or a major economic loss. They warrant additional scrutiny across the whole product lifecycle, from selection through obsolescence.
Identification begins during design. Failure modes and effects analysis and fault tree analysis identify the components whose failure propagates to a severe system effect, and those components are designated critical. Additional parts may be designated for supply reasons rather than technical ones, including sole-sourced items, long-lead items, and parts from suppliers with limited financial or manufacturing stability.
Enhanced controls scale with the consequence. They may include a source-controlled drawing that fixes the acceptable configuration, incoming inspection or lot acceptance testing, supplier audits, lot traceability from wafer or date code through the finished product, and parts-per-million quality agreements where standard acceptable quality level sampling is too coarse to be meaningful. Counterfeit avoidance is a specific and serious element of this control set: SAE AS5553 defines a counterfeit avoidance, detection, mitigation, and disposition system for electronic parts, AS6081 addresses purchases from independent distributors, and AS6171 provides the test methods for suspect parts. Purchasing critical components through authorized distribution remains the single most effective control.
Lifecycle management addresses the certainty that components disappear before products do. Obsolescence monitoring, formal management of diminishing manufacturing sources and material shortages, qualification of replacements, and last-time-buy analysis all belong to this activity. A last-time buy requires a forecast of remaining demand including spares and warranty, and it commits capital and storage while creating its own risks, since stored parts can degrade, absorb moisture, and develop solderability problems. Redesign to eliminate a critical component is sometimes the cheaper answer.
Documentation makes the whole system auditable. A critical components list records each designated part with the rationale for its designation, qualification records demonstrate that requirements are met, and change control ensures that no substitution or supplier change occurs without review. In regulated industries this documentation is a certification requirement; in every industry it is what allows the next engineer to understand why a particular part number is not negotiable.
Reliability Testing and Characterization
Component Reliability Testing
Component reliability testing generates the data used for prediction, qualification, and mechanism understanding. Tests may target a specific mechanism or assess overall life, and the test approach should follow from the question being asked.
Accelerated life testing raises stress to compress time. Temperature acceleration through the Arrhenius relationship is the most common approach, since most degradation mechanisms are thermally activated, and high-temperature operating life testing per JESD22-A108 is the standard semiconductor implementation; JESD85 gives the methods for converting the resulting data into a failure rate in failures in time. Voltage, current, humidity, and temperature cycling accelerate other mechanisms. The essential requirement, and the most common source of error, is that the acceleration model must be correct for the mechanism actually operating. An assumed activation energy applied to the wrong mechanism, or a stress level high enough to activate a mechanism that never occurs in the field, invalidates the extrapolation entirely.
Highly accelerated life testing takes the opposite approach. It steps temperature and vibration well beyond specification limits to find operating and destruct limits and to expose design weaknesses quickly. HALT makes no claim to predict field reliability, and treating its results as a lifetime estimate is a misuse. Its value is that it finds weak links during development while design changes are still affordable, and its production counterpart, highly accelerated stress screening, applies a calibrated subset of the same stresses to precipitate latent defects in shipped units.
Environmental testing evaluates behavior under the stresses of the intended use environment. Temperature cycling per JESD22-A104 exercises thermal fatigue mechanisms in packages, die attach, and solder. Temperature-humidity-bias testing, classically at 85 degrees Celsius and 85 percent relative humidity per JESD22-A101, evaluates moisture-driven corrosion and migration, with highly accelerated stress testing per JESD22-A110 using pressurized humidity to compress the same evaluation into a fraction of the time. Combined and sequential stress testing often reveals interactions that single-stress tests miss, such as moisture ingress through a delamination that thermal cycling created.
Mechanism-specific testing supports physics-of-failure modeling. Electromigration test structures characterize interconnect lifetime, dielectric breakdown testing characterizes gate stack lifetime, and package-level tests evaluate die attach, wire bond, and seal integrity independently. These focused tests yield model parameters rather than pass-or-fail results, and those parameters are what allow reliability to be predicted for a new design or a new operating condition.
Failure Analysis for Components
Failure analysis determines why a component failed, which is the prerequisite for effective corrective action. The approach depends on component type, failure mode, and available analytical resources, and the discipline of the sequence matters as much as the sophistication of the instruments.
Non-destructive analysis comes first, because it preserves evidence. Electrical characterization establishes how the part's behavior departs from specification and often localizes the fault to a pin or a function. External and optical examination documents physical condition, handling damage, and environmental exposure. X-ray inspection reveals wire bond geometry, voiding, and cracks without opening the package, and scanning acoustic microscopy detects delamination at internal interfaces. Curve tracing and thermal or emission microscopy narrow the fault location before any irreversible step.
Destructive analysis follows, trading the sample for detail. Decapsulation exposes the die for optical and electron microscopy. Cross-sectioning through the identified site reveals the geometry of the failure, and focused ion beam milling does the same at submicrometer scale. Chemical and elemental analysis identifies contamination, corrosion products, and material anomalies. Each destructive step should be planned against a hypothesis, because a poorly chosen cross-section plane destroys the evidence it was meant to expose.
Root cause analysis extends beyond the physical failure site to the reason the failure occurred. A cracked die may reflect a handling event, a mounting stress, an assembly process deviation, or a design that placed the part where the board flexes. The physical mechanism answers what failed; root cause answers why, and only the second supports corrective action. The distinction is also what separates blaming a supplier from fixing a problem, since a large share of components returned as failed are found to have been overstressed or misapplied in the using system.
Reporting closes the loop. A useful report states the failure description and history, the analysis performed, the findings, the root cause with the evidence supporting it, and the recommended corrective and preventive actions. A searchable failure analysis database converts individual investigations into organizational knowledge and makes trend analysis possible, which is frequently where systemic problems become visible.
Reliability Characterization
Reliability characterization builds a quantitative description of how a component fails over time. That description is what makes prediction, design optimization, and defensible application limits possible.
Failure distribution characterization identifies the statistical model that describes the failure times. The Weibull distribution is the workhorse because its shape parameter carries physical meaning: a shape parameter below one indicates a decreasing hazard rate and points to infant mortality and defect populations, a value of one gives the constant hazard rate that handbook methods assume, and a value above one indicates wear-out. The lognormal distribution fits many semiconductor degradation mechanisms, including electromigration, better than the Weibull does. Parameter estimation must handle censored data, since most life tests end with survivors, and goodness-of-fit assessment guards against reading a trend into a distribution that does not describe the data. Mixed populations, which produce a characteristic bend in a probability plot, are common and should not be forced into a single distribution.
Acceleration factor determination links test conditions to use conditions. Testing at several stress levels reveals the stress-life relationship rather than assuming it. Arrhenius analysis extracts an activation energy for thermally driven mechanisms, inverse power law models describe voltage and current acceleration, Coffin-Manson relationships describe cyclic fatigue, and Peck's model combines humidity and temperature for moisture-driven mechanisms. The acceleration factor is usually the largest single source of error in a reliability estimate, because it is raised to a power in the extrapolation; a modest error in an assumed activation energy becomes an order-of-magnitude error in predicted life.
Physics-of-failure modeling grounds the description in materials and stresses rather than in curve fitting. A physical model predicts life from geometry, material properties, and applied stress, which allows extrapolation to designs and conditions that have never been tested. Such models must be validated against test data, and their honest limitation is that they require material property data and mechanism understanding that are not always available.
Degradation analysis tracks a measurable parameter as it drifts toward a failure threshold rather than waiting for failures to occur. Monitoring threshold voltage shift, capacitance loss, contact resistance rise, or leakage current growth produces a trend from which time to failure can be projected, often long before any unit fails. The method shortens tests dramatically and yields information from small samples, and it requires that the relationship between the measured parameter and functional failure be established rather than assumed. It is also the foundation for condition monitoring and prognostics in fielded equipment.
Conclusion
Component reliability is fundamental to electronic system dependability. Understanding component failure mechanisms, from semiconductor wear-out through passive component aging to connector and solder degradation, is what allows engineers to make defensible decisions about selection, application, and qualification rather than hopeful ones.
The diversity of these mechanisms demands a correspondingly diverse response. Semiconductors are limited by electromigration, dielectric breakdown, bias temperature instability, and package integrity. Passive components age through material degradation and environmental attack, and their behavior differs so sharply between technologies that a single derating rule cannot serve them all. Connectors, boards, and solder joints fail through fatigue, fretting, corrosion, and migration. Each mechanism carries a characteristic stress-life relationship, and that relationship dictates the appropriate test, the appropriate derating, and the appropriate design response.
Effective component reliability management combines deliberate selection, appropriate derating, thorough qualification, and continuing monitoring through production and field service. Manufacturer data, handbook methods, and field experience each inform this work, and each must be interpreted with its limitations in view. Critical components warrant additional controls, counterfeit avoidance measures, and lifecycle planning that anticipates obsolescence.
Two habits distinguish the engineers who achieve reliability from those who merely predict it. The first is to reason from mechanism: to ask what physical process would have to occur for a component to fail in this application, and then to control that process. The second is to close the loop, feeding field and analysis results back into selection standards, derating rules, and qualification plans. Neither habit requires exotic tools, and together they account for most of the difference between products that meet their reliability requirements and products that do not.