Electronics Guide

Soft Errors and Single-Event Upsets

A soft error is a change of stored or propagating logic state caused by ionizing radiation, in a device that suffers no physical damage. Write the correct value back and the circuit behaves exactly as it did before. That property distinguishes a soft error from a hard failure, and it makes soft errors uniquely awkward for reliability engineering: the part passes every test after the event, the failure leaves no signature in a failure analysis laboratory, and the only defensible way to characterize the phenomenon is statistical. The industry name for the event itself is the single-event upset, and the family of related phenomena is collectively the single-event effects.

Radiation effects are often assumed to be a problem for satellites alone. They are not. Equipment at sea level, inside buildings, accumulates upsets at a measurable rate, and two mechanisms account for nearly all of it: alpha particles emitted by trace uranium, thorium, and their decay products in the package and interconnect materials, and neutrons produced when galactic cosmic rays collide with nuclei in the upper atmosphere. Neither source can be designed away. The alpha emitters live in the solder, the mold compound, and the underfill; the neutrons pass through a concrete roof with only modest attenuation. This article covers that terrestrial problem, a different engineering discipline from spacecraft hardening even though the physics overlaps.

The commercial stakes rose as devices shrank. A memory array holding a few kilobits and switching hundreds of femtocoulombs per cell was effectively immune; a server holding terabytes of dynamic memory, or an automotive processor holding tens of megabits of static memory inside a braking controller, is not. Soft errors now consume a substantial share of the failure rate budget for large digital systems, they are explicitly in scope for automotive and aviation safety standards, and the countermeasures against them, from ultra-low-alpha solder to lockstep processor cores, cost real silicon area and real money. Understanding where the upsets come from is the only way to spend that budget sensibly.

What Makes an Error Soft

Reliability practice separates faults by persistence. A permanent fault, such as an open bond wire or a fatigued solder joint, remains once it appears. An intermittent fault appears and disappears with temperature, vibration, or voltage, but originates in a physical defect that is always present. A transient fault, of which the radiation-induced upset is the archetype, is caused by an external event that leaves the hardware unchanged. Only the third category is properly called a soft error.

The distinction matters for the mathematics as much as for the diagnosis. A hard failure removes a unit from the population, so it belongs in the framework of survival functions and hazard rates described under the bathtub curve. A soft error removes nothing. It is a stationary arrival process, well approximated as Poisson, whose rate depends on the device, the supply voltage, the altitude, and the materials, and not at all on how long the equipment has run. A soft-error rate has no infant mortality region and no wear-out region.

The second classification concerns what the system does about the event. Following terminology Intel researchers popularized in the early 2000s, an upset that changes program output unnoticed is a silent data corruption, and one that some mechanism detects but cannot repair is a detected unrecoverable error. The consequences differ sharply: a detected unrecoverable error crashes a process or resets a controller, which is expensive but honest, while a silent data corruption produces a wrong answer that propagates. Most soft-error engineering converts the first category into the second, then the second into a corrected, invisible event.

Alpha Particles from Package and Interconnect Materials

An alpha particle is a helium nucleus, two protons and two neutrons, emitted when a heavy nucleus decays. The alpha emitters that matter in electronics come from the uranium-238 and thorium-232 decay chains, whose members are present at trace levels in almost every mineral-derived material. Emitted alphas carry between roughly four and nine megaelectronvolts. In silicon they travel a short distance, on the order of twenty to one hundred micrometers depending on energy, losing energy continuously and depositing most of it near the end of the track at the Bragg peak. Because the range is short, only emitters within a few tens of micrometers of an active device can cause an upset. A contaminated heat sink is irrelevant; a contaminated solder bump sitting directly above the transistors in a flip-chip assembly is not.

The charge arithmetic is worth committing to memory. Creating one electron-hole pair in silicon costs about 3.6 electronvolts, so one megaelectronvolt of deposited energy liberates roughly forty-four femtocoulombs of each sign, and a five-megaelectronvolt alpha stopping within the silicon generates something over two hundred femtocoulombs. Only a fraction reaches a sensitive node, but even modest collection efficiency puts the collected charge far above the switching threshold of a modern storage cell.

The 1978 Discovery and What It Established

Timothy May and Murray Woods at Intel identified alpha particles as the cause of otherwise unexplained soft errors in dynamic memories, reporting the work at the International Reliability Physics Symposium in 1978 and in the IEEE Transactions on Electron Devices in January 1979. The contamination was traced to the water used in manufacturing the ceramic packages, drawn from a source with elevated radioactivity. The episode established two principles that still govern practice: alpha emission is a materials problem solved in the supply chain rather than in the circuit, and extremely small quantities matter, because a decay rate negligible by any health-physics standard is not negligible when the target is a capacitor holding thirty femtocoulombs.

Where the Emitters Actually Live

Modern packaging concentrates the risk. In a flip-chip assembly the solder interconnect sits micrometers from the active surface, so the isotopic purity of the tin dominates. Tin refined from certain ores carries polonium-210, a strong alpha emitter with a half-life of about 138 days, and lead used in older alloys carried lead-210, which feeds polonium-210 continuously. A material can therefore be clean when measured and become dirty as a parent decays, which is why qualification specifies both emissivity and isotopic history. Other contributors include the silica filler in mold compounds and underfill, die-attach materials, ceramic package bodies and lids, and thick-film metallizations.

Atmospheric Neutrons and the Cosmic-Ray Cascade

The second mechanism begins far outside the atmosphere. Galactic cosmic rays, mostly protons with a small fraction of heavier nuclei, arrive at the top of the atmosphere with energies extending to enormous values. A primary that strikes a nitrogen or oxygen nucleus produces a shower of secondary particles, which themselves interact, producing a cascade. By the time the cascade reaches sea level it consists largely of muons, with a smaller flux of neutrons, protons, and pions. Muons are weakly ionizing and were long considered harmless, although at the smallest technology nodes their contribution has become measurable. Neutrons are the dominant terrestrial upset source.

A neutron carries no charge and deposits no energy by direct ionization. It upsets a circuit indirectly, through a nuclear reaction with a nucleus in or near the sensitive volume: elastic and inelastic scattering off silicon, spallation reactions that break a silicon nucleus into lighter fragments, and reactions with elements in the interconnect stack. The recoiling nucleus or the fragments are heavily ionizing and short-ranged, so the neutron effectively creates a miniature heavy-ion event inside the device. Because the reaction cross sections rise with energy, practice characterizes the environment by the integral flux above ten megaelectronvolts.

The reference value used throughout the industry is the flux at New York City at sea level, approximately thirteen neutrons per square centimeter per hour above ten megaelectronvolts, adopted as the normalizing condition in the JEDEC soft-error standards. Every other location is quoted as a multiplier on that reference.

Altitude, Latitude, and the Solar Cycle

Altitude is the strongest variable. The cascade is still building as it descends, so the flux increases sharply with height, roughly doubling for each thousand meters of altitude gain in the lower troposphere. Denver, at about sixteen hundred meters, sees several times the sea-level rate. Commercial aircraft cruising near eleven kilometers operate in a flux more than two orders of magnitude above sea level; figures above three hundred times the sea-level rate are commonly quoted. The flux peaks in the vicinity of eighteen to twenty kilometers, the Pfotzer maximum, and falls again above it as the primary flux has not yet cascaded.

Latitude matters because the geomagnetic field deflects charged primaries. The cutoff rigidity is highest near the magnetic equator, where the field screens most effectively, and lowest near the poles. The result is roughly a factor of two between equatorial and high-latitude flux at a given altitude, enough that avionics rate calculations specify route latitude as well as cruise altitude.

Solar activity modulates the flux inversely. A strong solar wind during solar maximum sweeps some galactic cosmic rays out of the inner solar system, so the terrestrial neutron flux is lowest when the Sun is most active, varying by a few tens of percent over the roughly eleven-year cycle. Solar particle events add short, intense excursions, a significant concern for polar flight routes and a minor one at sea level. Shielding is impractical throughout: the mean free path is long and intervening material generates its own secondaries, so only very deep overburden removes the flux, which is why underground laboratories are the control condition in soft-error experiments.

The Boron-10 Reaction and the Disappearance of BPSG

A third mechanism dominated some product families for a decade and then vanished. Borophosphosilicate glass, used as a planarizing dielectric, was formulated with several percent boron by weight. Natural boron contains about twenty percent boron-10, which has an enormous capture cross section for thermal neutrons, on the order of thousands of barns. Capture yields an alpha particle and a recoiling lithium-7 nucleus with a combined energy of about 2.8 megaelectronvolts, released within the dielectric stack, micrometers from the transistors. Boron-11, the remaining eighty percent, has a capture cross section smaller by many orders of magnitude and contributes nothing.

The reaction turned the otherwise harmless thermal neutron population into a significant upset source, and because the glass sat directly above the devices, the emitted particles reached sensitive nodes efficiently. The industry removed borophosphosilicate glass from the process rather than hardening against it, and by the early 2000s most advanced processes had done so. The episode remains instructive: a materials choice made for an unrelated reason created a radiation problem, and the fix was again a materials fix. Where boron-rich dielectrics or packaging materials remain in use, thermal neutron testing is a separate qualification step.

Charge Collection and the Critical Charge

Every soft-error mechanism ends the same way: a track of electron-hole pairs is created in the silicon, some of that charge is collected at a circuit node, and the resulting current pulse either does or does not overcome the node's ability to hold its state. The quantity that separates the two outcomes is the critical charge, written Qcrit, defined as the minimum collected charge that produces a state change or a propagating pulse of sufficient amplitude and duration to be latched.

Charge collection proceeds by two processes with different time constants. Charge generated inside or near a reverse-biased junction depletion region is swept out by drift within picoseconds, and the high carrier density along the track distorts the local field, temporarily extending the depletion region down the track in what is called funneling. Charge generated deeper in the substrate arrives more slowly by diffusion, producing a longer current tail. In bulk devices both components contribute; in silicon-on-insulator devices the buried oxide truncates the collection volume, which is one reason such devices upset less often, although the parasitic bipolar transistor formed in the floating body can amplify the small collected charge and partly offset the advantage.

To a first approximation the critical charge of a storage node is the product of node capacitance and supply voltage, modified by the drive strength of the restoring transistors. Scaling attacks that product from both directions: capacitance falls as dimensions shrink, and supply voltages have fallen from five volts to well under one. Critical charges measured in tens of femtocoulombs in the 1990s are now measured in single femtocoulombs or fractions of one. Against that, the sensitive area of each node also shrinks, so fewer particles strike it and each deposits less collectible charge. The two effects oppose each other, which is why the per-bit soft-error rate did not simply climb with each generation.

One consequence collides directly with power management. Lowering the supply voltage lowers the critical charge roughly in proportion, so dynamic voltage scaling and near-threshold operation increase the soft-error rate, sometimes steeply. Where that rate is part of a safety argument, the argument must be made at the lowest voltage the system actually uses, not at the nominal one.

How Different Circuit Elements Respond

Static Memory

A six-transistor static memory cell is a pair of cross-coupled inverters. An upset requires the collected charge to overwhelm the restoring transistor long enough for the feedback loop to latch the opposite state. Because static cells are the densest regular structure on most logic dies and are replicated in enormous numbers, they dominate the raw upset count for processors. Per-bit upset rates climbed through the 1990s, peaked near the 130-nanometer generation, then declined as the shrinking collection volume outpaced the falling critical charge. Multi-gate and fin-based devices continued that decline, because the fin presents a small collection volume raised above the substrate.

The per-bit improvement did not translate into a per-chip improvement, because bit counts grew faster than per-bit rates fell. A processor with tens of megabits of cache and a low per-bit rate can still carry a larger array failure rate than an earlier design with one megabit of much more sensitive cache, so design teams must reason in failures in time per device rather than per megabit. The organization of these arrays is treated under static memory technologies.

Dynamic Memory

Dynamic memory took a different path. Its storage cell is a capacitor, and the industry has held the stored charge roughly constant across generations, near twenty-five to thirty femtocoulombs, by building the capacitor vertically as a deep trench or stacked structure rather than letting it shrink with the lithography. That decision, made for signal-margin and refresh reasons, incidentally kept the critical charge high while the collection volume of the three-dimensional cell became very small. Per-bit cell upset rates fell by orders of magnitude, and the dominant sensitivity migrated from the array to the periphery: sense amplifiers, data path latches, address registers, and mode registers, which are ordinary logic and behave like it.

Two dynamic-memory phenomena are frequently and wrongly grouped with soft errors. Variable retention time, in which a cell's leakage changes between states over time, is a charge-trapping phenomenon. Row hammer, in which repeated activation of one row disturbs neighboring rows, is a disturbance caused by the access pattern. Both produce bit flips indistinguishable from upsets in a memory log, and the same codes correct them, but neither is a soft error and neither responds to low-alpha materials. Further background appears under dynamic memory technologies.

Latches and Flip-Flops

Sequential elements behave much like static memory cells, with two differences. They are far less numerous, so their aggregate contribution was long treated as secondary, and they are individually more vulnerable in many designs, because they are laid out for speed rather than density. An upset in a state machine register is also more consequential than one in a cache line, since it may put a controller into an undefined state rather than corrupting one datum.

Combinational Logic and Single-Event Transients

A particle striking a combinational gate produces a voltage glitch, the single-event transient, which is not an error unless a sequential element captures it. Three filters stand in the way. Logical masking occurs when the affected path does not influence the output for the current input values, as at an AND gate whose other input is low. Electrical masking occurs when the pulse attenuates as it propagates. Latching-window masking occurs when the pulse fails to arrive within the setup-and-hold window of a capturing element.

The last of the three weakens as clock frequency rises, because the vulnerable window occupies a larger fraction of every cycle. The combinational contribution therefore grows roughly in proportion to frequency, while the sequential contribution does not, and at multi-gigahertz clock rates the logic contribution becomes comparable to the flip-flop contribution in many designs. Electrical masking also weakens with scaling, because a given collected charge produces a longer pulse relative to a shorter gate delay.

Clock and reset distribution deserves separate attention. A transient injected into a clock buffer can produce a runt pulse or an extra edge that reaches thousands of flip-flops at once, converting one particle strike into widespread correlated corruption. The same applies to reset trees and global enables. These networks are small in area and large in consequence, which makes them good candidates for local hardening even in designs that do not harden the logic generally.

Multi-Cell Upsets and Bit Interleaving

Early soft-error analysis assumed one particle produced at most one flipped bit. That assumption failed as cell pitch shrank below the lateral extent of the charge cloud. A single track, or the several fragments of one spallation reaction, now deposits collectible charge in several adjacent cells, and a track entering at a shallow angle can traverse a whole row. The result is the multi-cell upset.

The terminology repays care. A multi-cell upset describes what happened physically, several neighboring cells flipping. A multi-bit upset describes what the system sees, several bits flipping within one logical word, the condition that defeats a single-error-correcting code. The mitigation strategy consists entirely of preventing the first from becoming the second.

That strategy is physical bit interleaving. The bits of one logical word are scattered so that physically neighboring cells belong to different words. With an interleaving distance of eight, a cluster of eight adjacent flipped cells becomes eight single-bit errors in eight independent words, each repaired by a single-error-correcting code. Interleaving costs area and routing complexity in the column multiplexing, and the required distance grows with the multi-cell fraction, which at advanced nodes exceeds half of all events. It is largely ineffective for register files and small distributed memories, where too few words exist to scatter, so those need different protection.

Beyond the Bit Flip: Functional Interrupts and Latch-Up

Two failure classes sit outside the simple bit-flip model and require separate treatment in a safety analysis.

A single-event functional interrupt is an upset in control state that halts or corrupts a whole device rather than a single datum: a flipped bit in a dynamic memory mode register, which changes timing for every subsequent access; a flipped configuration cell in a static-memory-based field-programmable gate array, which rewires the user logic; a corrupted serial link training state machine, which drops the link. The event is still soft, since power cycling or reconfiguration restores full function, but no data-path code detects it, and recovery takes milliseconds or seconds rather than nanoseconds. Devices for high-reliability use publish functional interrupt cross sections separately from bit cross sections, and system designers add watchdogs, periodic register readback and comparison, and configuration scrubbing.

Single-event latch-up is a different animal, because it can be destructive. Complementary metal-oxide-semiconductor structures contain a parasitic four-layer thyristor formed by the n-channel and p-channel devices and their wells. A sufficiently energetic charge injection triggers it into regenerative conduction, creating a low-resistance path from supply to ground that persists after the initiating particle has gone. The current is limited only by the supply and the metal, so the outcome ranges from an anomalous current draw cleared by a power cycle to melted interconnect and a dead die. Terrestrial latch-up rates are far lower than upset rates but are not zero, and the consequence may be permanent, which puts latch-up in the hard-failure column of a reliability budget even though a single particle caused it. Mitigations are structural: lightly doped epitaxial layers over heavily doped substrates, abundant well and substrate ties, guard rings, generous spacing between complementary devices, triple-well isolation, and silicon-on-insulator processes, which remove the parasitic path altogether. At the board level, current-limited supplies with fast trip thresholds convert a potentially destructive event into a recoverable one.

FIT Budgets and System-Level Derating

Soft-error engineering becomes tractable when expressed in the same units as every other failure mode. One failure in time is one failure per billion device-hours, so a component rated at one thousand failures in time upsets about once in one hundred and fourteen thousand hours, roughly once in thirteen years. Rates are additive across independent contributors, which makes budgeting possible and large systems vulnerable: ten thousand such components in a rack aggregate to ten million failures in time, about one event every one hundred hours. The unit and its arithmetic are treated more fully under key reliability metrics, and the series and parallel structures that combine them under system reliability calculations.

A system budget is built by allocation. The programme sets a target for the whole product, typically a maximum rate of undetected corruption plus a separate, larger allowance for detected events that force a reset. That target is divided among subsystems, then among devices, then within a device among memory arrays, register files, sequential logic, combinational logic, and configuration state. Each allocation is met by choosing a device, adding a protection mechanism, or negotiating the allocation upward at another block's expense. The discipline resembles the thermal or power budget process, and it fails the same way when nobody tracks the total.

Between the raw circuit rate and the observed system rate lie several stages of derating, and each stage must be justified rather than assumed:

  • Timing derating. The fraction of the clock cycle during which a transient can be captured. This factor grows toward one as frequency rises.
  • Logical derating. The fraction of upsets that propagate to an observable point rather than being masked by the surrounding logic values.
  • Architectural derating. The fraction of propagated errors that affect architecturally correct execution, excluding bits that are dead, speculative, or performance-only. This is the architectural vulnerability factor, introduced by Mukherjee and colleagues in 2003.
  • Application derating. The fraction of architectural errors that change the outcome the user cares about, which depends on the workload.

Derating is legitimate engineering, and the aggregate factor between raw and observed rates is often five or ten. It is also where safety arguments most often go wrong, because each factor is estimated from simulation or fault injection rather than measured, and the errors compound. Safety standards accordingly constrain it. Automotive functional safety practice discourages reducing the base failure rate used for soft errors on the strength of vulnerability factors, or of the very safety mechanisms whose coverage is claimed separately, since that credits the same mitigation twice.

Measuring Soft-Error Rate: The JESD89 Family

The industry standard for terrestrial soft-error measurement is JEDEC JESD89, "Measurement and Reporting of Alpha Particle and Terrestrial Cosmic Ray-Induced Soft Errors in Semiconductor Devices," whose revision B was published in 2021, together with a set of companion test methods. The standard defines the reference environment, the reporting units, and three complementary experimental approaches. No single one of them is sufficient, because each isolates a different part of the problem.

Alpha Source Testing

The alpha method places a calibrated thin-film source above a decapsulated or thinned die and counts errors against a known incident flux. Americium-241 is the usual choice, emitting alphas near 5.5 megaelectronvolts with a half-life of about four hundred and thirty-two years, so the source activity is stable over the life of the laboratory. Because the source flux exceeds any real package emissivity by many orders of magnitude, useful statistics accumulate in hours. The measured sensitivity, expressed as errors per incident alpha, is multiplied by the emissivity specified for the package materials to obtain the alpha contribution. The method thus separates the device's intrinsic sensitivity, fixed by the silicon, from the cleanliness of its packaging, which is a purchasing decision.

Accelerated Neutron Beam Testing

Neutron sensitivity is measured in a spallation beam whose spectrum approximates the atmospheric one. The principal facilities are the Los Alamos Neutron Science Center in New Mexico, the ChipIr instrument at the ISIS facility of the Rutherford Appleton Laboratory in the United Kingdom, TRIUMF in Vancouver, and the Research Center for Nuclear Physics at Osaka University. These beams exceed the sea-level flux by seven or eight orders of magnitude, compressing years of exposure into hours.

Three cautions apply. The beam spectrum is not identical to the atmospheric one, so the result depends on how integration and normalization are performed, and comparing facilities requires attention to the convention used. High instantaneous flux can produce more than one event per read cycle, so the test must avoid saturating detection. And the beam measures the neutron contribution only: the alpha contribution comes from the source test, and thermal neutron sensitivity, where boron is present, requires a separate thermalized beam.

Real-Time and Underground Testing

The third method exposes a very large population to the natural environment and waits. Real-time testing has no spectrum-matching uncertainty and captures every mechanism at once, including any the accelerated tests failed to anticipate. Its weakness is statistics: at natural flux a useful event count requires many gigabits of memory running continuously for months.

The elegant refinement is the paired underground control. Identical hardware runs at the surface and in a deep underground laboratory, such as the Laboratoire Souterrain de Modane in the French Alps, where the rock overburden removes essentially the whole cosmic-ray neutron flux. The underground population still experiences alpha particles from its own materials, so the underground count measures the alpha contribution directly and the difference between sites measures the neutron contribution. Mountain-altitude installations provide the complementary experiment, raising the neutron flux by a known factor while the alpha contribution stays constant. Together these give an assumption-free decomposition of the rate and serve as the reference against which the accelerated methods are validated. The broader logic of accelerating a stress and then validating the acceleration is treated under accelerated testing methods.

Qualifying Low-Alpha Materials

Because the alpha contribution scales linearly with the emissivity of nearby materials, controlling emissivity is the cheapest lever available for that half of the problem. Emissivity is quoted as counts per unit area per unit time, conventionally alphas per square centimeter per thousand hours. The industry recognizes graded levels: low-alpha material falls between two and fifty alphas per square centimeter per thousand hours, and ultra-low-alpha material below two, figures used consistently across the alpha-counting community. Commodity materials that have received no attention can measure well above the low-alpha range.

Measuring at these levels is difficult, and for years it was unreliable. Gas-flow proportional counters, the traditional instrument, place the sample outside the counting volume behind a thin window, and a round-robin study circulating identical samples among counting centers found wide disagreement. Later work with ultra-low-background ionization counters that place the sample inside the counting volume, notably the XIA UltraLo-1800, removed that variable: a multi-center study published in Nuclear Instruments and Methods in Physics Research A in 2014 reported that for low-alpha samples the remaining site-to-site variation was consistent with counting statistics alone, and that variation on ultra-low-alpha samples fell threefold, the residue attributed to differences in cosmogenic background. The lesson for a purchasing specification is that the method must be named alongside the limit, because a number without a method is not a specification. JEDEC document JESD221 addresses alpha radiation measurement in electronic materials.

A complete qualification covers every material within alpha range of the active silicon: solder alloys and bump metallurgy, underfill, mold compound and its filler, die attach, substrate metallization and solder resist, lid and stiffener materials, and conformal coatings. It also covers time, since a material containing lead-210 grows more active as that parent decays toward polonium-210, so certificates attach to lots and production dates rather than to a supplier in general. Refining tin to ultra-low-alpha grade is expensive, and the usual practice is to specify it only where geometry justifies it, chiefly for bump metallurgy sitting directly over the die.

Two boundaries are worth stating plainly. Low-alpha materials do nothing whatever about neutrons, and they do nothing about emitters inside the die itself, such as contaminated metallization, which is why wafer-level materials receive the same scrutiny.

Error Detection and Correction in Memory

Coding is the workhorse mitigation, because memory arrays hold most of the vulnerable bits and their regular structure makes coding cheap per bit.

The simplest scheme is parity, one redundant bit per word, which detects any odd number of errors and corrects none. It suffices where a clean copy exists elsewhere, which is why first-level instruction caches and write-through data caches use it: on a parity error the line is invalidated and refetched. The next step is single-error correction with double-error detection, a Hamming code with an added overall parity bit, most familiarly in the seventy-two-bit codeword protecting sixty-four data bits used in main memory for decades.

Single-error correction is defeated when one failure affects many bits of a word at once, notably the loss of an entire memory device on a rank. Symbol-oriented codes address this. The approach International Business Machines named chipkill, sold elsewhere as single-device data correction, arranges the code so that all bits contributed by any one device fall within a correctable symbol, so a whole device may fail without data loss. The cost is a wider access granularity and additional check bits.

Two more recent developments changed the landscape. On-die error correction became part of the DDR5 generation, correcting single-bit errors within the memory die before data leave it; it improves yield and masks weak cells, but protects nothing on the bus, does not substitute for system-level correction, and can hide errors from host diagnostics. Separately, link-level cyclic redundancy checking with retry now guards the interface itself, since at multi-gigabit signaling rates the channel is a plausible error source in its own right. The coding theory behind all of these is developed under error correction codes.

Scrubbing and the Accumulation Problem

A single-error-correcting code protects a word only until a second error arrives in it. Because soft errors arrive independently, the probability that a word already holding one corrected error acquires a second grows with the time between accesses. A frequently read word is corrected often and never accumulates; a word written once and read a year later is exposed for a year.

Memory scrubbing removes that exposure. A background engine walks the whole memory, reads each location, and writes back the corrected value, repairing a latent single-bit error before a second lands on it. Server memory controllers implement patrol scrubbing with a configurable period, commonly set so that a full pass completes in about a day, and demand scrubbing, which writes back the corrected value whenever a normal read finds a correctable error. The period is a design parameter rather than a default to accept: the residual uncorrectable rate falls roughly in proportion to the scrub interval while the bandwidth consumed rises inversely. Scrubbing matters equally for configuration memory in reprogrammable logic, where a readback-and-repair engine restores configuration cells before enough accumulate to change the user circuit.

Hardening the Logic: DICE, Redundancy, and Lockstep

Coding suits arrays. Sequential logic and control state need different techniques, applied at three levels.

Hardened Cells

The circuit-level answer is a storage cell that no single node can flip. The dual interlocked storage cell, described by Calin, Nicolaidis, and Velazco in the IEEE Transactions on Nuclear Science in 1996, stores each bit redundantly across four nodes interlocked so that a disturbance at any one is restored by the others. A single-node strike cannot upset the cell, which is a categorical improvement rather than a numerical one, and the price is roughly twice the transistor count with corresponding area and power.

The important caveat is charge sharing. The cell is immune to single-node upset by construction, but one particle track that deposits charge on two interlocked nodes at once defeats the interlock, and at advanced nodes those nodes may be close enough for exactly that to happen. Recovering the intended immunity requires layout discipline that separates the critical node pairs physically, which is why hardened-by-design cells are characterized by measurement rather than trusted on their schematic. Layout-aware hardening of conventional flip-flops, increasing the spacing and orientation diversity of sensitive nodes, buys a useful improvement at modest cost.

Spatial and Temporal Redundancy

Triple modular redundancy triplicates a function and votes on the three outputs, so any single corrupted copy is outvoted. It is simple, general, and expensive: three times the logic plus voters, which must themselves be protected or they become the single point of failure. Triplication is therefore applied selectively, to state machines, configuration registers, and critical control paths rather than to whole data paths. In reprogrammable logic it is combined with configuration scrubbing, since triplication protects the user state while scrubbing protects the configuration that defines the circuit; either alone leaves a hole, because accumulated configuration upsets eventually corrupt all three copies.

Temporal redundancy exploits the brevity of a transient: sampling the same signal at two or three instants separated by more than the expected pulse width, then voting, detects transients without triplicating the logic. It costs timing margin rather than area.

Lockstep Processors

The system-level answer for safety-critical controllers is dual-core lockstep. Two identical cores execute the same instruction stream and a comparator checks their outputs every cycle, raising a fault before a corrupted result reaches memory or an actuator. Practical implementations offset the second core by one or two clock cycles and mirror its inputs and outputs through delay stages, so that a disturbance affecting both at once does not produce matching wrong answers; physical separation and orientation diversity serve the same purpose. Lockstep costs essentially one hundred percent of the core area and provides detection rather than correction, so the system must still act on the fault signal, usually by a rapid reset into a safe state.

Lockstep pairs are standard in automotive microcontrollers and in processor cores marketed for functional safety, where suppliers publish diagnostic coverage figures for the comparator and its supporting logic. Those figures feed the safety metrics discussed below, and the integrator is expected to verify rather than accept them. Broader treatment appears under redundancy and fault tolerance and fault-tolerant design.

What Large-Scale Field Studies Show

Predicted soft-error rates and observed field rates have a complicated relationship, and the large-scale studies published since 2009 have repeatedly surprised the field.

The first, by Schroeder, Pinheiro, and Weber, examined memory errors across Google's fleet and appeared at SIGMETRICS in 2009. It reported correctable error rates far above the values then commonly assumed, with a large fraction of machines seeing at least one correctable error per year. More importantly, it found the errors strongly concentrated: a small number of modules produced most events, and a module that had produced an error was much more likely to produce another. That pattern is the signature of hard defects rather than random particle strikes. Hwang, Stefanovici, and Schroeder made the point explicitly in a 2012 study whose title observed that cosmic rays do not strike twice, and later work on high-performance computing installations and on a large social-media fleet reached compatible conclusions.

The correct reading is careful. This literature does not show that soft errors are unimportant. It shows that in large populations of commodity memory, permanent and intermittent defects contribute a larger share of observed events than the radiation-induced component, so a model attributing every memory error to particles mispredicts both the rate and its distribution. It also shows why controlled experiments remain necessary: a field log cannot distinguish a particle strike from a marginal cell.

A parallel caution applies to the silent data corruption reports published by major fleet operators beginning in 2021, which described processor cores producing wrong results on specific operations. Those investigations attributed the behavior to manufacturing defects and marginal devices rather than to radiation, and screening and workload-level checking are the corresponding mitigations. Distinguishing the categories is the job of physics-of-failure analysis.

Documented single-event incidents outside the laboratory are rarer than folklore suggests, because attribution after the fact is very hard. One well-attested case is the 2003 municipal election in Schaerbeek, Belgium, where an electronic tally credited a candidate with exactly four thousand and ninety-six extra votes, a value corresponding to one flipped bit, and no other explanation was found. Aviation investigators have likewise considered single-event effects among the candidate mechanisms for in-flight computer anomalies without confirming one, which is the characteristic difficulty: a soft error leaves nothing behind to find.

Standards That Make Soft-Error Rate Contractual

In several industries the soft-error rate is a term of the safety case rather than a private engineering concern, and the analysis must be documented to a defined standard.

IEC 61508, the general functional safety standard for electrical, electronic, and programmable electronic safety-related systems, appeared in its second edition in 2010. It requires dangerous undetected failure rates to be quantified against safety integrity level targets and diagnostic coverage claims to be justified. Transient faults enter as a contributor to the dangerous failure rate and through the diagnostics credited against them.

ISO 26262 adapts that framework for road vehicles, and its 2018 second edition added Part 11, devoted to applying the standard to semiconductors, which addresses transient faults, soft-error rate estimation, and the treatment of memories and processing elements explicitly. The quantitative targets are three metrics: the single-point fault metric, the latent fault metric, and the probabilistic metric for random hardware failures. For the highest automotive integrity level, ASIL D, the standard sets the single-point fault metric at ninety-nine percent or better, the latent fault metric at ninety percent or better, and the probabilistic metric below ten failures in time. Meeting a ten-failure-in-time target on a complex processor is demanding, and it is the direct reason automotive microcontrollers carry lockstep cores, error-correcting memory throughout, and extensive built-in self-test. The wider landscape is surveyed under functional safety standards.

Aviation approaches the subject from two directions. RTCA DO-254, recognized by certification authorities as guidance for the design assurance of complex airborne electronic hardware, governs how such hardware is developed and verified, and single-event effects enter as a safety consideration whose mitigation must be substantiated. The quantitative environment comes from a dedicated standard: IEC 62396-1, "Process management for avionics: Atmospheric radiation effects," whose second edition was published in 2016, provides guidance for avionics operating at altitudes up to sixty thousand feet. It characterizes the environment, defines how to compute single-event effect rates at altitude, and sets out design considerations for accommodating them. Because the neutron flux at cruise altitude exceeds the sea-level value by more than two orders of magnitude, an avionics box needs mitigation that identical ground equipment would not.

Underlying all of these, JEDEC JESD89 supplies the measurement methods and reporting conventions, and JESD221 the materials measurement. A supplier claim of a soft-error rate is meaningful only when it names the method, the reference environment, the supply voltage, and the operating mode, because all four change the answer.

Conclusion

Terrestrial soft errors are a permanent feature of digital electronics, not a defect to be eliminated. Two mechanisms produce nearly all of them. Alpha particles from trace uranium, thorium, and polonium in package and interconnect materials act only over tens of micrometers, which makes them a supply-chain problem solved by specifying and verifying material emissivity. Atmospheric neutrons, secondaries of the cosmic-ray cascade, act everywhere and cannot be shielded, which makes them an environment to be quantified by altitude, latitude, and solar phase, and then designed against.

What happens after a strike is governed by the critical charge, and scaling pushed that quantity down while shrinking the collection volume that feeds it. Per-bit rates fell while per-device rates did not; multi-cell upsets rose from a curiosity to more than half of all events; combinational logic joined sequential logic as a significant contributor as clock rates climbed; and functional interrupts and latch-up remain distinct failure classes. The response is layered. Low-alpha materials remove one source at the root. Error-correcting codes with adequate bit interleaving and a sound scrub interval handle the arrays. Hardened cells, selective triplication, and lockstep comparison handle sequential and control logic. Failure-in-time budgeting ties the layers together, the JESD89 methods supply the numbers that make the budget real, and automotive, industrial, and aviation standards make the exercise contractual.

The recurring lesson from field data is to keep the mechanisms distinct. Radiation-induced upsets, marginal cells, access-pattern disturbance, and manufacturing defects all appear in a memory log as flipped bits, and they respond to entirely different countermeasures. An engineer who can tell them apart, and who can state which environment, which voltage, and which measurement method a quoted rate refers to, is equipped to spend the soft-error budget where it actually buys reliability.

Related Topics