Risk Management
Risk management in electronics is the systematic process of identifying, analyzing, evaluating, and controlling the risks associated with electronic systems across their entire lifecycle. From concept and design through manufacturing, operation, maintenance, and eventual decommissioning, electronics professionals must find the ways a system can cause harm and then reduce that harm to a level the applicable standards, regulators, and society accept. The underlying definition is consistent across the field. ISO/IEC Guide 51, the document from which most product safety standards draw their vocabulary, defines risk as the combination of the probability that harm occurs and the severity of that harm, where harm means injury or damage to the health of people, or damage to property or the environment.
The discipline has grown in importance as electronic systems have become more complex and more deeply embedded in safety-critical applications. Modern automobiles, medical devices, industrial control systems, and aerospace platforms all depend on electronics whose failure could cause serious injury or death. That reality drove the development of international standards that codify how risk must be managed. IEC 61508, whose second edition was published in 2010, establishes the general framework for functional safety of electrical, electronic, and programmable electronic safety-related systems and defines four safety integrity levels, SIL 1 through SIL 4. Sector standards adapt that framework rather than replace it: ISO 26262 governs road vehicles through automotive safety integrity levels ASIL A through ASIL D, ISO 14971 defines the risk management process for medical devices while IEC 62304 covers medical device software, ISO 12100 covers machinery, IEC 61511 covers process industry safety instrumented systems, and DO-178C governs airborne software.
Effective risk management requires both technical expertise and disciplined process. Engineers must understand the failure modes of components and systems, estimate the probability and consequences of failure scenarios, and judge the effectiveness of candidate mitigations. They must also record their reasoning so that regulators, auditors, and certification bodies can confirm that residual risk has been reduced to an acceptable level. Many frameworks express that target through the ALARP principle, which calls for risk to be reduced as low as reasonably practicable, weighing the cost and effort of further reduction against the benefit gained. Others use different language for the same idea, including the requirement in European medical device law to reduce risk as far as possible.
The articles in this category follow the working sequence: analyze the hazards, by failure-based methods and by the system-theoretic methods that address what those miss, implement the safety functions that control them, and limit the legal exposure that follows when a product nonetheless causes harm.
Articles in This Category
What Risk Means in Practice
Precision about vocabulary prevents most of the confusion in this field. A hazard is a potential source of harm: stored electrical energy, a hot surface, a moving actuator, a lithium cell, a beam of laser light. A hazardous situation is the circumstance in which a person, property, or the environment is exposed to that hazard. Harm is the injury or damage that results. A hazard alone carries no risk; risk arises only when a credible sequence of events connects the hazard to someone who can be hurt. A battery pack stores energy whether or not anyone is nearby, but the risk appears when a cell fails short while the protection circuit is disabled and the pack sits in a user's pocket. Good analysis therefore enumerates sequences, not merely components.
Both dimensions of risk must be estimated, and they behave differently. Severity is usually the easier judgment, because the physics of burns, shock, crush, and loss of vehicle control are reasonably well understood. Probability is far harder, and the difficulty splits along an important line. Random hardware failures, such as a capacitor that shorts or a solder joint that opens, occur at rates that can be estimated from field data, reliability handbooks, or accelerated testing, and those rates can be combined arithmetically. Systematic failures cannot. A requirement written incorrectly, a design error, or a software defect is present from the first unit and will occur every time the triggering condition arises, so a failure rate for it has no physical meaning. This is why every functional safety standard has two halves: quantitative targets that address random hardware failure, and prescribed process rigor, review, and verification that address systematic failure. Techniques for estimating the random half are covered under reliability engineering and failure analysis.
No system reaches zero risk, so every framework needs a stopping rule. ALARP, formalized in United Kingdom safety practice, asks whether the cost, time, and trouble of further reduction are grossly disproportionate to the reduction achieved, and it works within tolerability boundaries: an individual risk of death of roughly one in a thousand per year for workers marks the region regarded as unacceptable, while roughly one in a million per year is treated as broadly acceptable, with the region between the two governed by the disproportion test. Other regimes phrase the rule differently. European medical device law requires risks to be reduced as far as possible without adversely affecting the benefit-risk ratio. Whatever the wording, the obligation is the same: reduce risk by design first, then justify and document what remains.
The Risk Management Process
ISO 31000:2018 supplies the generic management framework, and IEC 31010:2019 catalogs the assessment techniques that populate it. Product and functional safety standards impose a more specific sequence, but the steps are recognizably the same across sectors.
The process begins by defining the system, its intended use, and its reasonably foreseeable misuse. That last item is a requirement rather than a courtesy: ISO 12100, ISO 14971, and ISO 21448 all oblige the analyst to consider how people will predictably use the product incorrectly, because a hazard reached through ordinary human error is still a hazard the designer must address. Hazard identification follows, then risk estimation for each hazardous situation, then evaluation of the estimated risk against defined acceptance criteria. Risk control comes next, and its ordering is not optional.
ISO 12100 states that ordering as a three-step method, and ISO 14971 states it as a priority order for medical devices. First, eliminate or reduce the hazard through inherently safe design, by removing the energy, lowering the voltage, reducing the stored charge, or choosing a material that cannot ignite. Second, apply protective measures such as guards, barriers, interlocks, current limits, and safety functions implemented in the product itself. Third, provide information for safety: warnings, markings, instructions, and training. The order reflects declining reliability. A hazard removed by design stays removed for every unit, every user, and every year of service. A guard can be defeated. A warning label depends on someone reading it, understanding it, remembering it, and acting on it under pressure. Occupational practice uses the same logic in its hierarchy of controls, described under workplace and occupational safety.
The final steps are the ones most often shortchanged. Residual risk must be evaluated after controls are applied, including any new risk that a control introduces, since interlocks, diagnostics, and protective shutdowns can create their own failure modes. Implementation must then be verified, and the effectiveness of each control validated against the hazard it was meant to address. Finally, production and post-production information must flow back into the analysis, because field experience routinely reveals hazardous situations that no review meeting imagined. The methods behind these steps are treated in detail under hazard analysis and risk assessment.
Expressing Required Integrity
Functional safety standards convert a risk judgment into an engineering target by assigning an integrity level to each safety function. The level is derived from the risk reduction the function must deliver; it is not a badge selected for marketing. Each step upward costs architecture, diagnostics, verification effort, and schedule, so assigning a level higher than the analysis supports is as much a design error as assigning one too low.
IEC 61508 states its targets two ways, depending on how often the safety function is called upon. For low-demand operation, such as a protection system that acts perhaps once a year, the target is the average probability of a dangerous failure on demand: below one in ten for SIL 1, one in a hundred for SIL 2, one in a thousand for SIL 3, and one in ten thousand for SIL 4, each with the next decade as its lower bound. For high-demand or continuous operation, such as a motor control that must behave correctly at all times, the target is the average frequency of a dangerous failure per hour, which must fall below one in a hundred thousand per hour for SIL 1 and, in decade steps, below one in a hundred million per hour for SIL 4. Meeting the number is necessary but not sufficient: the standard also imposes hardware fault tolerance constraints, diagnostic coverage expectations, and systematic capability requirements on the development process.
ISO 26262 derives its automotive safety integrity level from three judgments made for each hazardous event: severity of harm, probability of exposure to the operational situation, and controllability by the driver or other persons at risk. The combination yields QM, meaning ordinary quality management suffices, or ASIL A through ASIL D. ASIL D, the most demanding, calls for a probabilistic metric for random hardware failures below one in a hundred million per hour, a single-point fault metric of at least 99 percent, and a latent fault metric of at least 90 percent. The standard also permits ASIL decomposition, in which a requirement is allocated to sufficiently independent elements at lower levels, an approach that is powerful and frequently misapplied, because the independence claim must itself be demonstrated.
Other sectors express the same idea in their own units. Machinery control systems use performance levels a through e under ISO 13849-1 or SILs under IEC 62061. Medical device software is assigned safety class A, B, or C under IEC 62304 according to the worst injury that a software failure could cause. Airborne software carries a design assurance level from A through E under DO-178C, derived from the severity of the failure condition it could contribute to, with the number of objectives and the number requiring independent verification falling at each level; the airworthiness target for a catastrophic failure condition on a transport-category aircraft is on the order of one in a billion per flight hour. Complex airborne electronic hardware follows DO-254. The aircraft and system development processes that surround them were revised together in December 2023, when SAE issued ARP4754B for development assurance and ARP4761A for safety assessment, moving the detailed analysis methods into the latter. United States defense programs apply MIL-STD-882E, issued in 2012 and updated by Change 1 in 2023, which pairs a severity and probability matrix with a rule that many engineers overlook: the resulting risk level determines how senior the official who may formally accept that risk must be. Implementation practice across these regimes is covered under functional safety implementation and, for software specifically, under software safety standards beyond medical.
Analytical Methods
No single technique finds every hazard, and mature programs deliberately combine methods whose blind spots differ.
Failure mode and effects analysis works from the bottom up. The analyst takes each component or function, postulates its failure modes, and traces the effects upward. FMEA is inductive, systematic, and thorough for single failures, and it maps naturally onto a bill of materials, which is why it dominates design and process work. Its weakness is combinations: an analysis that considers one failure at a time will not find the pair of faults that together defeat a protection scheme. Automotive and supply chain practice now follows the joint AIAG and VDA handbook published in 2019, which restructured the method into seven steps and replaced the risk priority number with an action priority rating of high, medium, or low. The change addressed a real defect in the older approach, since multiplying severity, occurrence, and detection allowed a low-severity, high-occurrence item to score the same as a potentially lethal one. Failure modes, effects, and diagnostic analysis extends FMEA with diagnostic coverage and safe failure fraction figures, which is the form functional safety work requires.
Fault tree analysis works from the top down. The analyst names an undesired top event and decomposes it through logic gates into the combinations of basic events that can produce it. FTA is deductive, handles combinations and common causes naturally, and yields both minimal cut sets and, when basic event probabilities are credible, a quantitative top event probability. Its weakness is scope: the tree answers only the question the analyst thought to ask. Hazard and operability studies attack that gap from a different direction, applying guide words such as no, more, less, reverse, and other than to each element of a design intent in a structured team session; the method originated in process engineering and remains strong wherever deviations in flow, sequence, or interface behavior matter.
Supporting techniques fill particular niches. Event tree analysis traces forward from an initiating event through the success or failure of each barrier. Markov models and reliability block diagrams handle redundancy, repair, and proof-test intervals that simple series arithmetic cannot represent. Bow-tie diagrams combine a fault tree and an event tree around a central event and communicate well to non-specialist audiences. Risk matrices remain the common language for evaluation, but they are coarse instruments: the severity and probability bands must be defined explicitly, and the ordinal categories must not be multiplied as though they were measurements. Systems-theoretic process analysis takes a different premise altogether, treating accidents as inadequate control rather than component failure, which makes it useful for software-intensive and autonomous systems in which every component performs exactly as specified and the interaction between them still produces harm. Investigation techniques applied after a failure has occurred are treated under failure analysis methodologies.
How Sectors Adapt the Framework
The common framework acquires a distinct shape in each industry, driven by the hazards that dominate there and by the regulator that supervises it.
Road vehicles work under ISO 26262, whose second edition, published in 2018, extended coverage beyond passenger cars to trucks, buses, and motorcycles and added guidance for semiconductors. ISO 26262 addresses hazards caused by malfunctioning behavior of electrical and electronic systems. It does not address the case in which nothing malfunctions and the system is still unsafe, which is precisely the problem posed by driver assistance and automated driving, where a sensor and its algorithms perform to specification but the specification proves inadequate for the situation. ISO 21448, published in 2022, covers that gap as the safety of the intended functionality, dealing with performance limitations, triggering conditions, and reasonably foreseeable misuse. Cybersecurity forms a third pillar, through ISO/SAE 21434 and, for type approval in the states that apply it, United Nations Regulation No. 155, which requires a certified cybersecurity management system. See automotive electronics standards for the wider regulatory picture.
Medical devices center on ISO 14971:2019, which specifies the risk management process and the risk management file that records it, with application guidance in ISO/TR 24971:2020. In the European Union the harmonized version is EN ISO 14971:2019 plus amendment A11:2021, whose annexes map the standard onto the Medical Device Regulation and the In Vitro Diagnostic Regulation. Device software adds IEC 62304. The distinguishing feature of the medical regime is that risk is always judged against clinical benefit, and that post-market surveillance is a legal obligation rather than a good practice. The regulatory detail appears under medical device regulations.
Machinery and process plant follow ISO 12100 for the risk assessment itself, then ISO 13849-1 or IEC 62061 for the safety-related control system, and IEC 61511 where safety instrumented systems protect a process. Aerospace and defense combine ARP4754B, ARP4761A, DO-178C, and DO-254 for civil aircraft with MIL-STD-882E for military programs; both regimes place unusual weight on independent review and on a documented safety argument. Consumer electronics has no single functional safety standard, relying instead on hazard-based product safety engineering under IEC 62368-1, on market surveillance, and on recall powers, which shifts more of the burden onto the manufacturer's own analysis. Sector requirements are compared under industry-specific regulations and industrial control standards.
Evidence and the Post-Market Loop
Analysis that leaves no trace has no regulatory value. Every regime therefore requires a durable record: the risk management file in the medical world, the safety case in functional safety, rail, and defense practice. A safety case is not a folder of documents but a structured argument, supported by evidence, that a system is acceptably safe in a defined operating context. It states claims, gives the argument that connects them, and cites the evidence that supports each step. Auditors read it in that order, and weak cases usually fail at the argument rather than at the evidence.
Traceability holds the record together. Each identified hazard should link forward to the safety requirement that addresses it, to the design element that implements the requirement, and to the verification result that demonstrates it works. That chain also makes change management tractable: when a component is substituted or a firmware routine rewritten, impact analysis follows the links to determine what analysis and testing must be repeated. Without traceability, every change forces either a full reanalysis or an act of faith. Configuration control and record retention practices are covered under quality management systems, and the testing that generates much of the evidence under safety testing procedures.
The loop closes in the field. Complaints, warranty claims, service reports, and incident investigations reveal the hazardous situations and the misuse patterns that no design review anticipated, and they supply the real failure rates that pre-production estimates only approximated. Regulated sectors formalize this as post-market surveillance with reporting duties and defined timelines; where a defect presents an unreasonable risk, corrective action extends to public notice and recall, described under product recall management. Field data also determines legal exposure, because a manufacturer that knew of a defect and failed to act stands in a far worse position than one that acted promptly. That connection is the subject of product liability prevention.
Common Failures of the Process
Risk management fails in recognizable ways, and most of them are organizational rather than technical.
The most common failure is timing. Analysis performed near the end of development, to satisfy an audit, arrives after the inherently safe design options have been spent; only guards and warnings remain, and those are the weakest controls available. The second is scope. An assessment that considers normal use by a trained operator, and ignores installation, servicing, transport, foreseeable misuse, and end of life, will miss whole classes of hazard. The third is false precision: a probability of dangerous failure computed to three significant figures from an assumed component failure rate conveys a confidence the underlying data does not justify, and it invites reviewers to argue about arithmetic instead of assumptions.
Two technical errors recur often enough to name. Independence is assumed rather than demonstrated: redundant channels that share a power supply, a clock, a connector, a calibration procedure, a compiler, or, most often, a requirements specification are not independent, and common-cause failure quietly dominates the result. And security is treated as a separate discipline: a remotely exploitable defect in a connected product can create a safety hazard directly, so the security analysis must feed the hazard analysis rather than run beside it. Regulatory expectations in that area are covered under cybersecurity regulations. A final governance failure underlies all of these: residual risk accepted by whoever is most inconvenienced by it is not an acceptance decision at all, which is why the stronger standards specify who holds that authority.
About This Category
Risk management addresses one of the most critical responsibilities of every electronics professional: ensuring that the systems we design and build operate safely. The articles here provide practical guidance on the whole cycle, from initial hazard identification through ongoing monitoring and continuous improvement. They explain the reasoning behind the standards rather than merely cataloging their clauses, so the same principles transfer across domains. The subcategories reinforce one another. Hazard analysis establishes what can go wrong and how much risk reduction is required, and system-theoretic analysis extends that to the control failures and performance limitations that component-failure methods do not reach; functional safety implementation builds and proves the safety functions that deliver the reduction; and product liability prevention preserves the evidence that the work was done competently and in good time. Related material appears under electrical safety, which covers the specific hazards most electronic products present, testing and certification, which produces the supporting evidence, and safety and protection systems, which describes the circuits that implement protection. Whether the goal is a consumer product, industrial equipment, or a safety-critical system, mastering these principles is essential for creating products that protect users and satisfy regulatory requirements.