Electronics Guide

Voltage Sag Immunity and Ride-Through

A voltage sag is a brief reduction in supply voltage, usually lasting a few cycles, that stops short of a complete interruption. Nothing about it is dramatic. The lights flicker, if anyone happens to be watching. Yet a sag of five cycles is the single most expensive power quality event in industry, because it stops production lines, scraps work in progress, and forces hours of restart on equipment that suffered no damage whatsoever. The paradox at the center of this subject is that the disturbance is trivial and the consequence is not.

The reason is a mismatch of timescales. Protective relaying on a distribution network is designed to clear a fault fast, and fast means a few cycles. During those cycles, every customer electrically near the fault sees depressed voltage. The utility considers the event a success: the fault was isolated, nothing burned, and service continued. The customer whose extruder just tripped considers it an outage. Both are correct. The sag is an unavoidable byproduct of protection working properly, and no amount of complaint to the utility will eliminate it.

That framing determines everything that follows. Because sags cannot be removed at the source, immunity must be engineered at the load. This article treats voltage sag ride-through as the plant-level discipline it actually is: first the physics and statistics that decide how often a site is hit and how deeply, then the equipment susceptibility mechanisms that decide whether a given hit causes a trip, then the tolerance curves and immunity standards that let ride-through be specified rather than hoped for, and finally the mitigation options arranged in the order that matters most in practice, which is the order of cost. Two neighboring articles carry parts of this subject in depth and are not duplicated here. Dynamic Voltage Restorers covers the series injection device itself, including its detection algorithms, control loops, and converter sizing, while Voltage Dips and Interruptions treats the same phenomenon from the electromagnetic compatibility laboratory, where the concern is qualifying a product against a standard test profile. This article sits between them, in the plant.

What a Sag Is, and Where It Comes From

Precision about terminology is worth the paragraph it costs, because the field uses two conventions that mean opposite things and confuse specifications constantly.

Definitions and the Residual Voltage Convention

IEEE 1159 defines a sag as a reduction in root-mean-square voltage to between 0.1 and 0.9 per unit of nominal, lasting from half a cycle to one minute. Below 0.1 per unit the event is an interruption. Above 0.9 per unit it is within the normal operating band. Longer than one minute and the event is reclassified as a sustained undervoltage, which is a different problem with different remedies. European practice, following IEC usage, calls the same event a voltage dip; the two words are interchangeable, and this article uses sag throughout.

The magnitude convention is the trap. A sag is normally described by its residual voltage, meaning the voltage that remains, not the voltage that was lost. A seventy percent sag is therefore a sag to seventy percent of nominal, a loss of thirty percent. Some older documents and some vendors use the opposite convention and describe the same event as a thirty percent sag. When reading a specification, an event log, or a product datasheet, establish which convention applies before drawing any conclusion. A tolerance requirement misread in this way inverts, and the resulting equipment is specified for the wrong half of the problem.

Duration is normally expressed in cycles or in milliseconds. At sixty hertz one cycle is 16.67 milliseconds; at fifty hertz it is 20 milliseconds. Standards that must serve both frequencies quote paired durations, so that a test point written as 25/30 cycles means twenty-five cycles at fifty hertz and thirty cycles at sixty hertz, both being half a second.

Faults Are the Dominant Cause

The overwhelming majority of sags that trip industrial equipment originate as short circuits somewhere on the transmission or distribution network. A tree limb contacts a conductor. Lightning flashes over an insulator. A vehicle strikes a pole. An animal bridges a bushing. A cable splice fails. In every case a low-impedance path appears between phases, or between a phase and ground, and enormous current flows through it.

That current flows through the system impedance between the source and the fault, and the resulting voltage drop depresses voltage across a wide electrical area. Protection detects the fault and opens a breaker, typically within three to ten cycles on a distribution feeder and faster still on transmission, where clearing times of two to four cycles are routine. When the breaker opens, the fault current stops and voltage recovers. The whole event lasts less than a fifth of a second.

Two consequences follow directly. First, sags are frequent, because faults are frequent across a network that spans thousands of kilometers of exposed conductor. Work by the Electric Power Research Institute's Power Electronics Applications Center, cited by Pacific Gas and Electric, concluded that a typical customer could expect on the order of twelve utility voltage sags per year. Individual sites vary enormously around that figure depending on their supply configuration and local exposure. Second, sags are short, because protection is fast. A mitigation scheme that provides half a second of support addresses nearly the whole distribution of events, which is why the economics of sag mitigation differ so sharply from the economics of outage backup.

Automatic reclosing complicates the picture. On overhead distribution, most faults are transient, so a breaker that has tripped will reclose after a short delay to test whether the fault has cleared. If it has, service resumes. If it has not, the customer experiences a second sag or interruption a fraction of a second after the first. Equipment that survived the first event may fail the second, and a reclose sequence can produce three or four disturbances in under two seconds. Any ride-through scheme sized only for a single event will be defeated by a reclose sequence.

Motor Starting and Transformer Energization

Not every sag comes from the utility. Large motors started across the line draw six to eight times their rated current until they approach speed, and that inrush depresses voltage on the shared bus. The distinguishing feature is duration: a motor-starting sag lasts as long as the acceleration, which may be several seconds for a large fan or pump, whereas a fault-induced sag is over in a fraction of a second. The distinguishing feature at the waveform level is shape: fault sags are rectangular, beginning and ending abruptly, while motor-starting sags show a sharp initial drop followed by a gradual recovery as the machine accelerates.

Transformer energization produces a similar in-plant sag through inrush current caused by asymmetric core saturation, with the added characteristic of heavy even-harmonic content that ordinary root-mean-square measurement obscures. Energization sags decay over tens of cycles as the core magnetization settles.

Distinguishing internal from external sources is the first analytical step in any investigation, and it is why serious monitoring campaigns instrument at least two points, typically the service entrance and the affected load. If both points show the sag, the source is upstream. If only the downstream point shows it, the cause is inside the facility and is usually cheaper to fix.

The Area of Vulnerability

Knowing that sags come from faults is not yet useful. The practical question is how many sags of what depth a particular site will experience per year, and that question has a rigorous answer.

The Concept

For a given piece of equipment with a given voltage tolerance, there exists a region of the surrounding power network within which a fault will depress the voltage at that equipment below its tolerance. Faults inside that region cause a trip. Faults outside it do not. The region is called the area of vulnerability, and it is measured in circuit kilometers of exposed line rather than in geographic area, because what matters is impedance from the equipment, not physical distance.

The method that formalized this analysis was IEEE Std 1346-1998, the Recommended Practice for Evaluating Electric Power System Compatibility with Electronic Process Equipment. It set out a procedure for combining the voltage sag environment of a site with the sensitivity of the process equipment installed there and reaching a financial conclusion. The standard is now classified as inactive and withdrawn, but the methodology it described remains the foundation of every commercial sag assessment, and the coordination chart it popularized is still drawn the same way. Practitioners should treat it as a source of method rather than as a document to cite in a current procurement specification.

What Sets the Size of the Area

Three factors determine how large the area of vulnerability is for a given load.

The first is the tolerance of the equipment. Equipment that trips at ninety percent residual voltage has an enormous area of vulnerability, because a fault very far away still depresses voltage by ten percent. Equipment that tolerates a sag to fifty percent has a small one. This is the factor the owner can change, and it is the reason equipment specification dominates the economics of the whole subject.

The second is the stiffness of the supply. Voltage at a load during a remote fault is set by a voltage divider between the impedance from the source to the point of common coupling and the impedance from that point to the fault. A strong supply with high available fault current holds voltage up well and shrinks the area of vulnerability; a weak rural feeder does the opposite. Sites served at transmission voltage through a dedicated substation are, other things equal, less exposed than sites served from a long radial distribution feeder shared with many other customers.

The third is the topology and exposure of the network. A site fed from a network with many kilometers of overhead line radiating from its supply substation sits inside a large sag-producing region. A site fed from an underground cable network in a dense urban area experiences fewer faults per year, although cable faults, when they occur, are permanent rather than transient and produce interruptions rather than reclose sequences.

Voltage level matters as well. Faults on the transmission system propagate over very large areas and affect many customers at once, but transmission is well protected and faults there are comparatively rare. Faults on the distribution system are far more numerous but affect a smaller region. For most industrial sites the distribution contribution dominates the count of shallow sags while the transmission contribution produces a smaller number of widely felt events.

From Vulnerability to Expected Trips per Year

The area of vulnerability becomes a number when it is combined with fault statistics. Utilities maintain fault rates per hundred circuit kilometers per year for each class of line they operate, broken out by voltage level and by overhead or underground construction. Multiplying the exposed length inside the area of vulnerability by the applicable fault rate yields an expected number of trips per year for that piece of equipment.

Performing this calculation across a range of assumed equipment tolerances produces the result that actually drives decisions: a curve showing expected annual trips against equipment sensitivity. That curve is usually steep. Moving a machine's tolerance from a trip at eighty-five percent residual voltage to a trip at seventy percent can cut expected annual trips by a large factor, because it removes from the area of vulnerability all the distant faults that produced only shallow sags, and distant faults are numerous precisely because there is so much distant line. Moving the tolerance further, from seventy percent to fifty percent, buys much less, because the remaining events are close-in faults on a short length of line. Sag mitigation therefore shows strongly diminishing returns, and recognizing where the knee of that curve lies is the whole of the engineering judgment.

Sag Type: What the Load Actually Sees

A sag is not a single number. Faults are usually unbalanced, transformers rearrange the unbalance, and the resulting waveform at the equipment terminals may differ substantially from what a phase-to-ground measurement at the service entrance suggests.

Balanced and Unbalanced Faults

Three-phase faults, in which all three conductors are shorted together, produce a balanced sag: all three phases drop by the same amount and the phase relationships are preserved. These are the deepest and most damaging sags, and they are also the rarest fault type.

The large majority of faults on distribution systems are single-line-to-ground faults, in which one conductor contacts ground. These depress one phase severely while the other two remain near normal, or in some grounding arrangements actually rise. Phase-to-phase faults, the next most common category, depress two phases and leave the third largely intact.

The practical importance is that single-phase loads on different phases of the same panel experience entirely different events from the same fault. One machine trips, its neighbor does not, and the maintenance record shows an inexplicable pattern until someone checks which phase each is connected to. Three-phase loads see a different combination again, because a three-phase rectifier draws from whichever phase pair is momentarily highest and is therefore substantially supported by the healthy phases during an unbalanced sag.

How Transformers Rearrange a Sag

Every transformer between the fault and the load changes the character of the sag. A delta-wye transformer, which is the standard arrangement for stepping down to utilization voltage, removes the zero-sequence component and redistributes the remaining unbalance. A single-phase-to-ground fault on the primary side therefore does not appear as a single-phase sag on the secondary side. It appears as a sag affecting two phases, at a different depth, with the phase angles shifted.

Math Bollen's Understanding Power Quality Problems: Voltage Sags and Interruptions, published by IEEE Press in 1999, set out the systematic classification of sag types that resulted from tracing each fault type through successive transformer connections, and that classification remains the standard analytical vocabulary. Its practical lesson for a plant engineer needs no algebra: the sag measured at the medium-voltage service entrance is not the sag the equipment sees, and the transformation depends on every winding configuration in between. A monitoring campaign that instruments only the service entrance will systematically mispredict which loads trip.

Phase-Angle Jump

When the impedance of the faulted path has a different ratio of reactance to resistance than the source impedance, the voltage during the sag is shifted in phase as well as reduced in magnitude. This shift is the phase-angle jump, and it is invisible to any instrument that records only root-mean-square magnitude.

Overhead lines are strongly inductive and their ratio of reactance to resistance is not far from that of the source, so faults on overhead networks produce modest jumps. Underground cables are relatively more resistive, and faults on cable networks can produce jumps of tens of degrees. Loads that synchronize to the supply are the ones affected: line-commutated rectifiers and thyristor converters fire their devices at angles referenced to a detected zero crossing, and an abrupt phase shift can cause a misfire, a commutation failure, or a protective trip even when the magnitude of the sag alone would have been survivable. Phase-locked loops in drive and inverter controls must reacquire, and the reacquisition transient can itself trip a supervisory function.

Phase-angle jump explains a category of trips that otherwise appears random. Two sags of identical measured depth and duration produce different outcomes because one carried a jump and the other did not, and no magnitude-only event record can distinguish them.

Point-on-Wave Effects

The instant within the cycle at which a sag begins and ends also matters, and for the same reason it is rarely captured. A sag that starts near a voltage zero crossing produces a gentle transition; one that starts near the peak produces an abrupt discontinuity that interrupts rectifier conduction at maximum current. Recovery is similarly point-dependent, and recovery at the peak can produce an inrush transient into a discharged capacitor bank that draws enough current to trip an input breaker after the supply has already returned to normal.

Electromechanical devices are strongly point-on-wave dependent. A contactor holds through a sag partly on the residual energy in its magnetic circuit, and whether it drops out can depend on where in the cycle the sag began relative to the armature's mechanical state. This is one reason repeated testing of the same contactor at the same nominal test point produces scattered results, and it is why standardized test methods require dips to be applied at multiple phase angles rather than at one.

Switch-Mode Power Supplies and Hold-Up Time

Almost every electronic load in an industrial plant is fed by a switch-mode power supply, and the behavior of that supply during a sag is the starting point for any susceptibility analysis.

The Bulk Capacitor as the Energy Reserve

A conventional offline supply rectifies the incoming alternating current and stores the result on a bulk electrolytic capacitor, from which a high-frequency converter draws to produce the regulated outputs. When input voltage falls, the rectifier stops conducting, and the converter continues to run on stored energy. Regulation is maintained until the bulk capacitor voltage falls to the converter's minimum operating point, after which the outputs collapse. The interval between the loss of input and the loss of regulation is the hold-up time, and it is the fundamental measure of a supply's inherent sag immunity.

The energy available is the difference between the energy stored at the starting voltage and the energy remaining at the minimum operating voltage, which for a capacitance C discharged from V1 to V2 is one half of C multiplied by the difference of the squares of those voltages. Dividing that energy by the load power, adjusted for converter efficiency, gives the hold-up time directly.

Three consequences fall out of that expression, and all three matter in practice. Hold-up scales linearly with capacitance, so doubling the bulk capacitor doubles the hold-up. Hold-up scales inversely with load, so a supply loaded to thirty percent of its rating holds up roughly three times as long as the same supply at full load. And hold-up depends on the square of the starting voltage, which is why the same supply behaves so differently on different nominal mains.

Why Input Voltage Headroom Dominates

A universal-input supply rated for 85 to 264 volts alternating current operated at a nominal 230 volts has a very large margin: input may fall to roughly thirty-seven percent of nominal before the supply leaves its rated range at all. The same supply operated at a nominal 120 volts has margin only to about seventy percent. The component is identical; the sag immunity differs by a factor that decides whether a plant trips.

This is the cheapest lever in the entire subject, and it costs nothing at design time. Selecting a supply whose rated input range places the site's nominal voltage near the top of that range converts a sag into a non-event. Where three-phase power is available and the supply's rating permits, connecting a single-phase supply phase to phase rather than phase to neutral gains the same kind of margin. Where a three-phase front end is an option, it is better still, because a three-phase rectifier feeding a common bus continues to be supported by the healthy phases through the unbalanced sags that constitute most events.

Power Factor Correction Front Ends

Supplies with an active power factor correction stage behave differently and generally better. The boost converter that performs the correction regulates an intermediate bus, commonly around 400 volts, above the peak of the rectified input. As input voltage falls, the boost stage increases its duty cycle and holds that bus constant, so the downstream converter sees no disturbance at all until the input falls below the point where the boost stage can no longer maintain regulation. The transition is therefore later and sharper than in a supply without correction.

The compensating behavior is that a corrected supply draws constant power. As input voltage falls, input current rises to compensate, and that rising current can trip an input protective device or overload an upstream conditioning transformer that was sized on nominal current. A mitigation device placed upstream of a bank of power-factor-corrected supplies must be rated for the current they will draw during a deep sag, not the current they draw normally.

The Undervoltage Lockout Problem

Many supplies and many downstream devices include an undervoltage lockout or brownout detector that deliberately shuts the device down when input falls below a threshold, on the reasonable theory that operating on collapsing rails risks unpredictable behavior. The threshold is frequently set conservatively by a component vendor with no knowledge of the application, and the practical result is that equipment shuts down while it still had adequate stored energy to continue. The lockout, not the physics, defines the ride-through. Where the threshold and its hysteresis are adjustable, examining them is among the highest-yield things an engineer can do, and where they are not, the specification of the next unit purchased should address them.

Contactors, Relays, and the Control Circuit

If one finding recurs across every published industrial sag investigation, it is this: the component that stops the plant is usually not the expensive one. It is a control relay or a motor contactor costing a small fraction of what it interrupts.

What the Standard Actually Permits

IEC 60947-4-1, the standard for electromechanical contactors and motor-starters, defines the operating envelope for the coil. A compliant alternating-current contactor must close satisfactorily at any control supply voltage between eighty-five and one hundred ten percent of its rated value. For drop-out, the standard specifies a window rather than a point: the measured drop-out voltage must lie between seventy-five and twenty percent of the rated control supply voltage.

Read that requirement carefully, because it is the crux of the problem. A fully compliant contactor is permitted to drop out at seventy-five percent of its coil voltage. That is a shallow sag, well inside the region that any reasonable process would otherwise have survived, and well inside the region that the site's area of vulnerability makes frequent. Two contactors from different manufacturers, both compliant, may drop out at seventy-five percent and at thirty percent respectively, and nothing in either catalog necessarily reveals which is which. Direct-current coils have a wider permitted window still, extending down to ten percent.

This is not a defect in the standard. A contactor must reliably open when commanded, and guaranteed drop-out is a safety property. But the standard is written for a device that opens when its control circuit is de-energized, not for one that must distinguish an intentional command from a four-cycle disturbance on the supply, and it makes no promise of sag ride-through whatsoever.

Why the Control Circuit Dominates

The mechanism that follows is mundane and devastating. A five-cycle sag to sixty percent arrives. The drive, the programmable controller, and the instrumentation all ride through comfortably, because their power supplies had ample hold-up. The control-circuit contactor that carries the drive's run permissive drops out. Its auxiliary contact opens. The permissive is lost. The drive, which never lost power and never had any internal reason to stop, sees its run command removed and executes a normal stop. The seal-in circuit, having opened, does not re-establish itself when voltage returns, because seal-in circuits are specifically designed not to restart machinery automatically. The line is down until an operator walks to the panel and presses a button.

Every element of that sequence performed exactly as designed. The plant stopped anyway. The economics are striking, because a contactor of higher hold-in margin, or a small ride-through module on the control transformer secondary, would have prevented it for a cost that disappears against a single hour of lost production.

Control transformers deserve specific attention because they compound the problem. A control transformer sized tightly for its burden has significant internal impedance, so the sag seen by the control circuit is deeper than the sag on the primary. Adding inrush margin to the control transformer, which is normally done for a completely different reason, incidentally improves sag performance.

Fixes at the Control Level

Several remedies exist and all are inexpensive. Contactors are available with electronically controlled coils that maintain hold-in through voltage excursions far deeper than a conventional magnet will tolerate, and swapping to one is a like-for-like replacement. A rectified and capacitor-backed direct-current coil supply gives a control circuit hold-up in the same way a bulk capacitor gives an electronic supply hold-up, and small dedicated ride-through modules that do exactly this are sold for the purpose. Time-delay relays and undervoltage relays with adjustable delay can be set so that a momentary loss of a permissive does not propagate into a stop, though this must be done with care and never in a way that defeats a safety function.

The one thing to avoid is a blanket delay on every undervoltage function in the plant. Undervoltage detection frequently protects a motor from single-phasing or a load from a genuine loss of supply, and a delay long enough to ride out a sag may be long enough to permit real damage. The correct approach discriminates: distinguish the relays that protect equipment from the relays that merely observe the supply, and adjust only the latter.

Variable Frequency Drives

Drives are simultaneously the most sag-sensitive major loads in most plants and the ones with the most sophisticated available remedies. Variable Frequency Drives covers drive architecture in full; the concern here is specifically the sag response.

The Undervoltage Trip

A drive rectifies the incoming supply onto a direct-current bus and inverts from that bus to produce variable-frequency output. When input voltage sags, the bus discharges into the load at a rate set by the output power. Drive bus capacitance is sized primarily for ripple current rather than for energy storage, so the stored energy is modest: at full load, a typical drive sustains its bus for something on the order of one to a few line cycles, roughly twenty to fifty milliseconds. When the bus falls below the undervoltage threshold, the drive trips to protect itself and the motor from operation on inadequate voltage.

That interval is shorter than most real sags. It is also shorter than the tolerance envelopes that the ITIC and SEMI F47 curves describe. A drive relying on bus capacitance alone will trip on a substantial fraction of the events a typical site experiences, and this is why drives appear so prominently in trip statistics despite being far more sophisticated than the contactors that outrank them.

Load matters greatly, and in the helpful direction. Ride-through scales inversely with output power, so a drive running a fan at forty percent speed, where the cube law makes power very low, coasts through disturbances that would trip the same drive at full torque. Trips consequently cluster at high-production periods, which is both statistically expected and commercially unfortunate.

Inertia Ride-Through and Kinetic Buffering

The elegant answer exploits energy that is already present. A spinning motor and its driven load store kinetic energy proportional to inertia and to the square of speed. During a sag, the drive can command negative torque, decelerating the load and forcing the motor to act as a generator. Regenerated power flows back through the inverter into the direct-current bus and holds it up. The control loop regulates bus voltage by modulating deceleration rate, so the bus is held at its setpoint for as long as there is kinetic energy to give.

Manufacturers market this function under several names, including kinetic buffering, inertia ride-through, flux braking control, and controlled deceleration or power-loss ride-through. The details of the control law differ, but the physics is common. When the supply returns, the drive reaccelerates to the speed setpoint.

The technique works well when the load has substantial inertia and tolerates a speed excursion. A large centrifugal fan is close to ideal: its inertia is enormous, and a momentary speed dip has no process consequence. It works poorly, or not at all, on a low-inertia load, on a load driven at very low speed where little kinetic energy exists, and on any load where speed accuracy is itself the product. A coating line whose web tension depends on exact speed ratios between sections cannot accept an uncontrolled deceleration, and a synchronized multi-drive system may hold its bus perfectly while destroying the product through loss of registration. Enabling inertia ride-through on such a line converts a clean trip into scrap.

A related and important detail is behavior on recovery. A drive that has tripped or coasted must reconnect to a motor that is still turning. Reapplying voltage at the wrong frequency to a spinning machine produces a large transient current and an overcurrent trip. Flying restart, sometimes called speed search or catch-on-the-fly, detects the actual shaft speed and matches output frequency to it before ramping. Ride-through configuration that does not also address restart behavior tends to trade an undervoltage trip for an overcurrent trip a few hundred milliseconds later.

Other Drive-Level Measures

Where inertia is unavailable, energy must be added. Common approaches include enlarging the bus capacitance, connecting a supercapacitor module to the direct-current bus through a bidirectional converter, or tying several drives to a shared bus so that regenerating machines support motoring ones. Shared-bus arrangements are particularly effective in multi-motor installations, since some axes are usually decelerating while others accelerate, and they simplify the problem from many separate ride-through requirements to one.

Active front ends, which replace the diode rectifier with a controlled converter, regulate the bus over a wider input range and hold it through sags that would defeat a diode bridge, at additional cost and complexity. Where a drive is being specified anyway for harmonic reasons, the sag benefit comes at no extra charge and should be counted in the justification.

Controllers, Instrumentation, and Hidden Dependencies

Programmable logic controllers and process instrumentation are, individually, among the better-behaved loads. Their power supplies are small, lightly loaded, and often specified with wide input range, so a controller frequently rides through events that stop everything around it. That is not entirely good news, because a controller that stays alive while its field devices reset must handle the resulting inconsistency correctly.

Controller Power and Memory

The controller power supply should be evaluated the same way as any other: rated input range, actual loading against rating, and measured hold-up. Loading a supply at forty percent rather than ninety percent is free ride-through and is good practice for thermal reasons anyway. Retentive memory backed by battery or non-volatile storage preserves program state across a genuine power loss, and its configuration deserves review, since retained data that is stale after a restart can be worse than data that is cleared.

Redundant supplies fed from different phases provide real protection against the unbalanced sags that constitute most events, because a single-phase-to-ground fault upstream will rarely depress two different phases equally. Where such a scheme is used, the diode-or arrangement that combines them must be verified to actually transfer without a gap.

Input, Output, and Field Devices

The controller is only as immune as what it reads and drives. Sensors, transmitters, analyzers, and networked remote input and output racks each have their own supplies and their own thresholds. A four-to-twenty-milliampere transmitter that resets during a sag produces a signal excursion that the controller reads as a genuine process change and may act upon. An alarm or interlock configured on that signal will fire. An operator then faces a screen full of alarms whose real cause was a five-cycle disturbance, and the diagnostic time far exceeds the disturbance.

Networked architectures introduce a further mode. A remote input and output rack that resets must re-establish its communication link, and the reconnection may take considerably longer than the electrical event: link negotiation, address resolution, and configuration verification can consume seconds. During that interval the controller sees the rack as failed and takes whatever action its fault configuration specifies. Reviewing the network fault response of every distributed node is a necessary part of any serious ride-through study, because it is entirely common for the recovery logic rather than the disturbance to determine downtime.

The Weakest Link Sets the Answer

A process is only as immune as its least immune component in the chain required for continuous operation. There is no point in specifying a drive to ride through a deep sag if the safety-related contactor upstream of it drops out at seventy-five percent, and no point in hardening the drive and the contactor if a flow transmitter three meters away resets and drives an interlock. Sag immunity is a system property, and evaluating it demands an inventory of every device in the operational chain rather than attention to the largest or most expensive ones.

This inventory is also what makes the problem tractable, because the inventory almost always reveals that a handful of cheap devices dominate. That is the recurring and encouraging finding of the field: the fix is usually small, and finding it is the hard part.

Tolerance Curves: CBEMA, ITIC, and SEMI F47

A tolerance curve plots voltage magnitude against event duration and draws a boundary inside which equipment is expected to keep operating. Three curves dominate practice, and they are routinely misapplied.

CBEMA and ITIC

The original CBEMA curve was developed by the Computer Business Equipment Manufacturers Association as a design guideline for mainframe computer power supplies. It was a smooth curve describing the voltage envelope such equipment could survive. Its successor, published by the Information Technology Industry Council in 2000 and universally called the ITI or ITIC curve, replaces the smooth profile with straight-line segments that are far easier to apply to a measured event.

The ITIC envelope applies to single-phase 120-volt equipment on 120, 120/208, and 120/240-volt sixty-hertz systems. It permits a complete dropout of up to twenty milliseconds, a residual voltage of seventy percent for up to five hundred milliseconds, and eighty percent for up to ten seconds, returning to a steady-state band of ninety to one hundred ten percent beyond that. Plotting recorded disturbances against this envelope has become the standard way of presenting a monitoring campaign, and it answers the first question any investigation asks: was this event one that the equipment should have survived?

Its limitations are systematically ignored. It was written for information technology equipment, not for motor drives, contactors, or process instrumentation. It applies to a specific voltage and frequency. It is a description of what a class of equipment was designed to tolerate, not a requirement any particular product has been tested against. Using it to argue that a given machine ought to have ridden through a given event is an argument by analogy, not by specification.

SEMI F47

SEMI F47, the Specification for Semiconductor Processing Equipment Voltage Sag Immunity, first published in 2000, is a different kind of document and the most important one in this subject. It is not a description of typical behavior. It is a requirement, written so that it can be placed in a purchase order.

It requires equipment to continue operating without interruption through sags to fifty percent of nominal for durations from fifty to two hundred milliseconds, to seventy percent for two hundred to five hundred milliseconds, and to eighty percent for five hundred to one thousand milliseconds. Behavior outside those bounds, for sags shorter than fifty milliseconds or longer than one second, is not a requirement. The specification adds recommended but not required thresholds, including tolerance of a complete dropout for one cycle, eighty percent for ten seconds, and continuous operation at ninety percent. Its scope covers semiconductor processing equipment, metrology equipment, and automated test equipment. The companion document SEMI F42 defines the test method by which conformance is demonstrated, and the 2006 revision of F47 references IEC 61000-4-34 for the test methodology, aligning the semiconductor requirement with general electromagnetic compatibility test infrastructure.

What makes F47 effective is not its technical content, which is unremarkable, but its position in a commercial transaction. Because fabs write it into equipment specifications and withhold acceptance until it is met, it propagates down the supply chain: equipment makers impose it on their module and power supply vendors, and power supply vendors publish F47 test reports so that their products can be selected. A voluntary standard became a market requirement, and ride-through became a comparable product attribute that a buyer can specify rather than discover. SEMI Standards treats the mechanism by which this happens across the whole SEMI catalog.

The specification's applicability extends well beyond semiconductors. Nothing in the envelope is fab-specific, and any purchaser of production equipment can and should cite it. Doing so at procurement is by a wide margin the cheapest available lever, because the supplier's cost of meeting it during design is a small fraction of the buyer's cost of retrofitting mitigation afterward.

What a Curve Is Not

A tolerance curve is a design target and a diagnostic aid. It is not a guarantee, and treating it as one produces two specific errors.

The first error is assuming that equipment inside the envelope will not trip. The curves describe magnitude and duration only. They say nothing about phase-angle jump, point-on-wave, sag type, or the unbalanced events that constitute the majority of real disturbances. An event plotted comfortably inside the ITIC envelope can trip equipment through a phase jump that the plot does not show.

The second error is assuming that a curve applies to a system because it applies to a component. A machine assembled from power supplies that individually meet F47 does not thereby meet F47, because the machine also contains contactors, sensors, and network devices that were never evaluated. Immunity does not aggregate; it is set by the weakest element. Only a test of the assembled machine establishes the assembled machine's tolerance.

Immunity Test Standards and Their Limits

Three families of documents define how sag immunity is measured, and they serve genuinely different purposes.

IEC 61000-4-11 and IEC 61000-4-34

IEC 61000-4-11 specifies test methods for immunity to voltage dips, short interruptions, and voltage variations for equipment with rated input current up to sixteen amperes per phase. IEC 61000-4-34 extends the same principles to equipment above sixteen amperes per phase, which covers most industrial machinery. Together they define the generator characteristics, the test levels, the number of repetitions, the phase angles at which dips are applied, and the performance criteria by which results are judged.

Those criteria are worth understanding, because they permit far more than continuous operation. Criterion A requires normal performance during the disturbance. Criterion B allows temporary degradation that recovers by itself. Criterion C permits loss of function that requires operator intervention to restore, provided the function does recover. A product that trips on every test dip and requires a manual restart may still pass its immunity test, because the applicable product standard specified Criterion C. That is a legitimate and often appropriate outcome for a consumer appliance. It is a production stoppage in a factory.

Voltage Dips and Interruptions covers these tests in detail from the compliance perspective, including the preferred test levels and the class structure that determines them.

IEEE 1668

IEEE Std 1668, the Recommended Practice for Voltage Sag and Short Interruption Ride-Through Testing for End-Use Electrical Equipment Rated Less than 1000 V, addresses a gap the IEC documents leave. Issued first as a trial-use document in 2014 and as a full recommended practice in 2017, it is not industry-specific and applies to any electrical or electronic equipment on a low-voltage system that can malfunction or shut down as a result of a supply voltage reduction lasting less than a minute.

Its distinguishing contribution is realism about sag type. It defines procedures covering single-phase, two-phase, and three-phase sags, balanced and unbalanced, which reflects what actually happens on a distribution network rather than the simplified balanced profiles that a basic immunity test applies. It also specifies certification and test reporting requirements, so that the ride-through characterization of one product can be compared with another. For a purchaser trying to specify sag immunity for equipment outside the semiconductor sector, it provides the vocabulary that F47 provides inside it.

What Compliance Does Not Prove

A compliance certificate establishes that a particular sample, in a particular operating mode, in a laboratory, met a particular set of test points to a particular performance criterion. Each of those qualifiers limits what may be inferred.

Test conditions rarely match service conditions. Equipment is normally tested at a convenient operating point, and susceptibility varies strongly with load. A machine that passes at half load may fail at full load, because hold-up scales inversely with power. Test profiles are discrete points, not a continuum, and equipment can pass every specified point while failing at a duration between two of them. Test sags are usually balanced, while service sags usually are not. Test dips carry no phase-angle jump, while service sags on cable networks carry substantial ones. Reclose sequences that deliver several events in under two seconds are not part of a standard test. And the performance criterion may permit exactly the shutdown-and-manual-restart behavior that the plant cannot tolerate.

None of this makes compliance testing worthless. It establishes a floor, it makes products comparable, and it catches gross deficiencies. It simply does not answer the question a plant engineer is asking, which is whether this machine will keep running through the sags this site actually experiences. Answering that question requires either site-specific testing of the assembled machine or the empirical evidence of a monitoring campaign correlated against trip records.

Mitigation in Order of Cost

Mitigation options span four orders of magnitude in price, and the decision is almost always made on cost. Working through them in ascending order is therefore not merely a convenient organization; it is the correct method, because each step should be exhausted before the next is considered.

First: Fix the Control Circuit

The cheapest intervention is almost always at the control level, because that is where the trips originate. Replacing a marginal contactor with one having electronic coil control, adding a small capacitor-backed direct-current supply to a control circuit, correcting an undervoltage relay setting that was never considered, or resizing an undersized control transformer are all interventions costing a small multiple of a component price. Documented case studies across many industries repeatedly find that a control-level fix eliminated the majority of a site's sag-related downtime.

The prerequisite is knowing which device tripped, which requires either instrumentation on the control circuit or careful reconstruction from event sequence records. Sites that have never done this work generally do not know, and the investigation is the first expenditure worth making.

Second: Configure What Is Already Installed

Drive firmware settings cost nothing but engineering time. Enabling inertia ride-through where the load tolerates a speed excursion, configuring flying restart so that recovery does not produce a second trip, adjusting undervoltage thresholds and their delays where they are adjustable, and setting automatic restart behavior deliberately rather than by default can eliminate a large fraction of drive trips on equipment already in place. The caution stated earlier applies throughout: each change must be evaluated against the process and against safety functions, and inertia ride-through is actively harmful on speed-critical lines.

The same applies to controller and network fault-response configuration. A remote node whose fault action is a controlled hold rather than an immediate shutdown, where the process permits, converts a link reset into a pause.

Third: Specify Immunity at Procurement

This is the single most cost-effective lever available, and it is available only before the equipment is bought. Writing SEMI F47 conformance, or an IEEE 1668 characterization, into a purchase specification transfers the problem to the supplier, who solves it during design at a small fraction of the cost of a later retrofit. The specification should require a test report on the assembled machine rather than component certificates, since immunity does not aggregate.

The corollary is organizational rather than technical. Procurement specifications are usually written by people who have never seen a sag investigation, and the requirement will not appear unless someone puts it there. A standing clause in the plant's equipment specification template costs nothing and compounds over every future purchase.

Fourth: Targeted Power Conditioning

When the load itself cannot be fixed, hardware is added, and the guiding principle is to protect the smallest possible boundary. Conditioning one control cabinet is dramatically cheaper than conditioning a service entrance, and it addresses the same trips.

Constant-voltage ferroresonant transformers regulate through magnetic saturation, respond within a cycle, and provide useful support for small, steady loads at the cost of poor light-load efficiency and sensitivity to frequency. Electronic tap-switching regulators respond within a cycle or two and handle larger loads efficiently. A dynamic voltage restorer injects a compensating voltage in series with the supply and is the standard answer for a substantial process load facing frequent moderate sags; Dynamic Voltage Restorers covers its detection, control, and sizing in depth. A static transfer switch solves a different version of the problem by exploiting source diversity, moving the load to an alternate feed within a fraction of a cycle, and it is effective only where two genuinely independent sources exist. A sag on the transmission system will depress both feeds at once, and no transfer switch can help.

Fifth: Stored Energy

Flywheel systems store kinetic energy in a rotating mass and deliver it through a motor-generator or a converter. Their characteristics suit sag mitigation unusually well: they provide seconds rather than minutes of support, which matches the duration distribution of real sags almost exactly, and they avoid the replacement cycle, temperature sensitivity, and floor loading of a large battery plant. Supercapacitor systems occupy a similar niche with no moving parts, delivering high power for short durations and tolerating very large numbers of cycles.

Both are attractive precisely because they are sized for the actual problem. A facility that experiences a dozen sags a year, each lasting under half a second, and no meaningful number of true outages, is badly served by a battery plant sized for fifteen minutes of runtime, most of which is never used and all of which must be maintained.

Sixth: Uninterruptible Supplies

A double-conversion uninterruptible power supply isolates the load from every input disturbance and rides through sags as a side effect of what it does continuously. Where the load genuinely requires protection against outages as well as sags, it is the right answer and the sag immunity is free. Backup Power Systems covers the topologies and their sizing.

Where the load requires only sag immunity, an uninterruptible supply is usually the most expensive way to obtain it, carrying conversion losses, battery replacement cycles, and maintenance burden in exchange for a runtime capability that the actual disturbance profile does not require. The honest test is whether the site's event history contains outages long enough to justify the batteries. Frequently it does not, and the correct scope is far narrower: protect the controls and the ordered-shutdown logic on an uninterruptible supply, so that a genuine outage produces a clean stop, and address the sags separately with something matched to their duration.

The Economics of the Decision

Every choice above is ultimately a comparison between the annualized cost of mitigation and the annualized cost of trips avoided. Doing that comparison honestly is harder than the engineering.

Quantifying the Cost of a Trip

The cost of a single sag-induced stoppage is far larger than lost production hours, and understating it is the most common error in these evaluations. A complete accounting includes the value of production not made during downtime and during the ramp back to rate, the material scrapped in process, the labor required to clear, clean, and restart equipment, the cost of restoring product that solidified or degraded in place, the disposal cost of that material, any quality investigation triggered by the interruption, and, where relevant, the contractual consequence of a missed delivery.

The distribution of these costs across industries is extremely wide, and that width is the point. A continuous process with a long restart, such as a glass line, a fiber draw, a plastics extrusion, or a semiconductor tool holding wafers, carries a cost per event that is orders of magnitude above a discrete assembly operation that resumes in three minutes. Any general figure quoted for the cost of a sag is meaningless; the number must be developed for the specific process, and the operations staff who perform the restart are the correct source for it.

Establishing the Event Frequency

Frequency is the other multiplicand, and it must be measured rather than assumed. Utility statistics and published averages establish an order of magnitude but not a site value, because exposure depends on the specific feeder, the specific substation, and the specific network topology serving that site.

A monitoring campaign therefore precedes almost every well-made investment decision. Instruments meeting the class A requirements of IEC 61000-4-30 measure to a defined aggregation window of ten cycles at fifty hertz or twelve at sixty, roughly two hundred milliseconds, with voltage measurement uncertainty held to one tenth of one percent. That standardization is what allows measurements at two points, or by two parties, to be compared, which matters greatly when a utility and a customer disagree about where a disturbance originated. Power Meters and Analyzers covers the instrumentation.

The campaign should run long enough to capture seasonal variation, which for sags means at least one storm season and preferably a full year, because lightning-driven fault rates vary by an order of magnitude between seasons in many regions. It should instrument at least the service entrance and the affected loads. And its output must be correlated against production records, because the finding that actually justifies expenditure is not a count of sags but a demonstrated mapping from specific recorded events to specific recorded stoppages. Sites that undertake this correlation frequently discover that a substantial fraction of their sags caused no trip at all, which sharply changes the mitigation case, and that a substantial fraction of their stoppages had causes other than power quality, which changes it further.

Building the Case

With cost per event and events per year established, the expected annual loss follows directly, and each mitigation option can be evaluated against it. The evaluation must use the residual trips each option leaves rather than assuming any of them eliminates the problem: a scheme that addresses sags to fifty percent still permits trips from deeper events, and the correct comparison is against the reduction in expected annual loss, not against the total.

The steepness of the vulnerability curve discussed earlier reappears here as the decisive economic feature. Because the first increment of tolerance removes the large population of shallow, distant-fault events, the cheapest interventions capture most of the available benefit, and each subsequent increment costs more and returns less. A rational program consequently works up the cost ladder and stops when the marginal option no longer pays, which for most sites happens well before the expensive end. The plants that spend the most on power quality hardware are frequently not the ones that suffer the most; they are the ones that skipped the investigation and bought equipment.

One further consideration belongs in the case. Semiconductor fabrication is the canonical example of a process where the cost per event is high enough to justify almost any mitigation, and the industry's development of F47 followed directly from that arithmetic; Semiconductor Fabrication Power covers the resulting facility design. Most plants are not fabs, and importing a fab's conclusions without repeating its arithmetic produces overspending as reliably as ignoring the problem produces downtime.

Conclusion

Voltage sags are a permanent feature of alternating-current power systems. They arise because protective relaying clears faults quickly, and a fast clearing time is exactly what makes the resulting disturbance short, frequent, and impossible to eliminate at the source. The engineering response is therefore immunity at the load rather than perfection at the supply.

The technical structure of that response is now well understood. The area of vulnerability turns a site's supply configuration and its equipment tolerance into an expected number of trips per year, and the relationship between the two is steep enough that modest improvements in tolerance produce large reductions in trips. Sag type, phase-angle jump, and point-on-wave explain why events of identical measured depth produce different outcomes and why magnitude-only records mislead. Equipment susceptibility follows from hold-up energy in electronic supplies, from the drop-out window that IEC 60947-4-1 permits for contactor coils, and from undervoltage thresholds in drives, with the persistent and useful finding that the cheapest component in the chain usually sets the answer.

The tolerance curves and immunity standards give that response a vocabulary. The ITIC envelope makes recorded events interpretable. SEMI F47 makes ride-through specifiable, and its real power lies in a purchase order rather than in its technical content. IEC 61000-4-11, IEC 61000-4-34, and IEEE 1668 make immunity measurable, subject to the honest limitation that a laboratory test at discrete points on a balanced supply is not the sag environment a plant inhabits.

What remains is judgment about money. Mitigation options run from a contactor replacement to a facility-wide conditioning scheme, and the correct order of consideration is ascending cost, because the low end so often suffices. The decision deserves a monitoring campaign correlated against production records first, because the two numbers it requires, cost per event and events per year, cannot be borrowed from anyone else's plant. Sites that do that work usually find the problem smaller and cheaper than they feared. Sites that skip it usually buy the wrong equipment.

Related Topics