Motion Pictures and the Coming of Sound
Between 1926 and 1930 the American film industry rebuilt itself around a piece of electronics. Studios rewrote their production methods, exhibitors rewired more than twenty thousand buildings, tens of thousands of musicians lost their livelihoods, and an art form that had developed a mature international visual language abandoned it. The proximate cause was a photoelectric cell, a vacuum-tube amplifier, and a loudspeaker large enough to fill an auditorium. Nothing about the idea was new; Thomas Edison had wanted a talking picture since the 1890s. What was new was that the parts finally existed.
This article treats that transition as an engineering problem and follows the consequences. It begins with the two obstacles that defeated every mechanical attempt at synchronized sound, amplification and synchronization, and shows why both required the vacuum tube. It then examines the two rival solutions that reached the market within eighteen months of each other: sound on disc, in the form of Warner Bros.' Vitaphone, and sound on film, in the variable-density Movietone and variable-area RCA Photophone systems. Sound on film won, and the reason is worth stating precisely, because it is a property of the medium rather than of the machinery. The article then turns to what the winning format imposed on cinema, from the frame rate to the shape of the picture, and to what the conversion cost the people and companies who paid for it.
A companion article on this site, the history of sound recording, covers the general technology of capturing and reproducing sound, from the phonautograph and the acoustic phonograph through electrical recording, magnetic tape, and the digital transition. The boundary between the two is straightforward: that article follows the recording chain as an end in itself, wherever it was used, while this one follows a single application in which sound had to be locked to a moving image, reproduced at auditorium levels, and printed on a strip of celluloid that a projectionist could splice. Those three constraints produced a distinct body of engineering. A third article, on the electronic entertainment industry between the wars, covers the same years from the side of the consumer-electronics business that supplied the equipment.
Why Synchronized Sound Waited for Electronics
The failure of pre-electronic sound film is easy to misread as a failure of imagination. It was not. Inventors attempted talking pictures continuously from the 1890s, and several systems reached commercial exhibition. They failed on physics, and they failed on the same two counts every time.
The Amplification Problem
An acoustic phonograph has no gain. Every unit of energy that moves the air in front of the horn came from the groove, and every unit in the groove came from the performer at the recording session. The system is passive throughout, and its output is set by the mechanical impedance match between a stylus, a diaphragm, and a horn of practical size. That is adequate for a parlor and marginal for a small hall. It is hopeless for a theater seating fifteen hundred people, where the loudest passage must be audible over the ventilation plant in the back row and the sound must arrive as speech rather than as a distant blur.
Escaping the constraint mechanically was tried. Léon Gaumont's Chronophone, demonstrated in Paris from 1902, used compressed air to obtain amplification without electronics, as the Auxetophone did. It produced volume at the cost of distortion and hiss, and it did not scale. Nothing before the triode could deliver clean acoustic power on the order of watts into a horn, and clean watts are what a theater requires.
The second half of the problem is the transducer. Acoustic recording captured sound through a horn, which meant the performer had to stand in front of it. In a phonograph session that is merely awkward; in a film, where the camera must see the actors and the actors must move, it is fatal. Electrical recording put a microphone in the room instead, but a microphone produces microvolts, and microvolts are useless without gain.
The Synchronization Problem
The second obstacle is subtler and, in the end, more instructive. If picture and sound live on separate carriers driven by separate mechanisms, their relative timing is an unforced error waiting to happen. Edison's Kinetophone of 1913 belted a cylinder phonograph to a projector and shipped the pair to theaters, where the belt slipped, the operator lost the start mark, and the illusion collapsed. Other systems used electrical interlocks, mechanical shafting, or a projectionist with a hand control and a good ear.
The tolerances are unforgiving. A lip-sync error of a fifth of a second is unmistakable to a general audience, and at twenty-four frames per second a fifth of a second is fewer than five frames. Any mechanism that accumulates error over a ten-minute reel must therefore hold relative speed to a few parts in ten thousand, and it must recover from every splice, every reel change, and every operator error along the way.
Note what the two problems have in common: both are problems of energy and control that a passive mechanism cannot solve, and both dissolve once an amplifier exists. Amplification supplies the missing power, and it also supplies the means to read a soundtrack optically, from a signal too weak for any mechanical linkage, which is what finally made synchronization a property of the film rather than of a coupling between machines.
The Vacuum Tube Arrives in the Theater
The technical foundation was laid outside the film industry, in the laboratories of the American telephone system. Western Electric, the manufacturing arm of AT&T, and the research organization that became Bell Telephone Laboratories in 1925 had spent a decade building the components a sound film needs, for reasons that had nothing to do with cinema.
The condenser transmitter developed by Edward C. Wente at Western Electric from 1916 gave a microphone with a flat, wide response and a predictable calibration, at a time when the alternative was a carbon button whose sensitivity drifted with packing and temperature. Vacuum-tube amplifiers built for long-distance telephone repeaters gave gain that was stable enough to cascade. Loudspeaker research aimed at public address gave moving-coil drivers coupled to exponential horns, an arrangement that trades bulk for efficiency and was the only way to obtain theater-filling output from the tens of watts a tube amplifier of the period could produce.
Those same components produced the electrical recording system that Western Electric licensed to the phonograph industry in 1925. Applying them to film was an obvious next step for the engineers, and a much less obvious one for the studios, who had watched sound pictures fail repeatedly and saw no reason to gamble a profitable silent business on another attempt.
AT&T eventually organized its film activity into a subsidiary, Electrical Research Products, Incorporated, universally called ERPI, which licensed sound equipment to studios and installed and serviced it in theaters. ERPI's position is central to the economics of the transition. The telephone company did not make films; it made and rented the apparatus, collected license fees on production, and held the patents that made both possible. When the industry converted, a large share of the money flowed out of Hollywood and into the electrical industry.
Sound on Disc: Vitaphone
Warner Bros. was a second-rank studio in the mid-1920s, ambitious and short of theaters, which is precisely the position from which a risky technical bet looks attractive. Working with Western Electric, the company formed the Vitaphone Corporation in April 1926 and put the system before the public that summer.
The Format and the Interlock
Vitaphone recorded sound electrically onto a wax master and pressed shellac discs almost sixteen inches in diameter, played at thirty-three and one-third revolutions per minute. The unusual diameter and the slow speed were chosen for one reason: playing time. A standard thousand-foot reel of 35 mm film runs about eleven minutes at twenty-four frames per second, or ninety feet per minute, and the disc had to last exactly as long as the reel. Vitaphone is therefore the origin of the thirty-three and one-third rate the long-playing record would adopt twenty years later.
The discs were cut from the inside outward, so that the groove began near a scribed synchronization arrow at the center and worked toward the rim. Inside-out cutting is the opposite of phonograph practice, and it was chosen because linear velocity under the stylus rises as the groove moves outward, putting the greatest available bandwidth at the end of the reel, where a reel typically builds toward its loudest and most detailed material.
Synchronization was mechanical. The turntable was geared to the projector so that the two ran from a common drive, which guaranteed that the ratio between disc rotation and film transport could not drift once established. Establishing it was the projectionist's job: line the start mark on the film up with the projector gate, set the needle on the arrow scribed into the disc, and start the machine. Get it right and the interlock held for the whole reel. Get it wrong and the reel played out of sync from the first frame to the last.
Don Juan, The Jazz Singer, and the Commercial Proof
Vitaphone reached the public on 5 August 1926 with Don Juan, a costume picture shot as a silent film and released with a recorded orchestral score and effects but no dialogue, preceded by short subjects in which musicians performed on camera with live-recorded sound. The strategy was deliberately conservative. The immediate commercial proposition was not talking actors; it was giving every theater in the country, including houses that could afford a pianist and no more, an orchestra on the soundtrack.
Dialogue followed. The Jazz Singer premiered on 6 October 1927, and although it is mostly a silent film with synchronized songs and a few passages of speech, it earned roughly two and a half million dollars in the United States and abroad, a figure large enough to end the industry's argument about whether audiences wanted sound. Within a year the major studios had signed with ERPI, and by mid-1930 the silent feature had effectively ceased production in Hollywood.
Why Disc Could Not Last
Vitaphone proved the market and then lost the technical argument, for reasons that follow directly from keeping the sound on a separate object.
The first failure mode is the splice. Film breaks in projection, and a projectionist repairs it by cutting out the damaged frames and cementing the ends together. Every frame removed shifts the picture permanently earlier relative to the disc, and nothing in the system can detect or correct the error. A print that has been through a busy house for a few weeks accumulates splices, and a film running acceptably in one theater could arrive at the next unwatchable.
The second is wear. Shellac played with a steel needle degrades measurably, and a Vitaphone disc was considered good for roughly twenty playings before surface noise and groove damage became objectionable. A popular film in continuous exhibition consumed discs, and exchanges had to keep replacements moving.
The third is the reel change. A feature runs to many reels, and every changeover requires the incoming projector and its turntable to be cued correctly while the outgoing one still runs: a fresh opportunity, under time pressure in a dark booth, to start the reel out of sync.
The fourth is editing, and it is the one that mattered most to the studios. A disc cannot be cut. Shortening a scene, tightening a performance, or removing a bad take meant either rerecording the disc from the beginning or dubbing through a second generation, with the noise penalty that entails. Sound on disc made the editorial flexibility that silent cinema had spent thirty years developing effectively unavailable.
Warner Bros. stopped recording directly to disc by the middle of 1931, though disc prints continued to be distributed to theaters equipped only for them until 1937. The system had done its work: it had shown that talking pictures could earn money, which was the fact the industry needed and the only fact that disc was better placed than film to establish quickly.
Sound on Film: The Case Laboratory and Movietone
The alternative approach photographs the sound. A microphone signal, amplified, modulates a light source; the modulated light exposes a narrow stripe running the length of the film alongside the picture; and in the projector a lamp shines through that stripe onto a photoelectric cell, which recovers the waveform. The idea is older than Vitaphone by decades. What it needed, again, was gain: a photoelectric cell delivers a minute current, and reading a soundtrack is impossible without an amplifier immediately behind the cell.
Three lines of work converged. In Germany, Josef Engl, Joseph Massolle, and Hans Vogt developed the Tri-Ergon system from 1919, using a glow lamp to write a variable-density track. Its patents proved commercially consequential in a way its films did not: William Fox bought rights to them in 1926, and litigation over their scope ran until 1935. In the United States, Lee de Forest promoted his Phonofilm process from the early 1920s in a small circuit of theaters. Phonofilm did not fail on principle; it failed on quality, on capital, and on de Forest's inability to interest a major studio in a system he was demonstrating with vaudeville shorts. The company was defunct by 1929.
The decisive contributions came from Theodore Case's laboratory in Auburn, New York. Case had developed the thalofide cell, a thallium-compound photodetector, between 1916 and 1918, initially for a classified United States Navy infrared signaling system tested in 1917. He then developed the AEO light, a glow lamp whose output could follow an audio-frequency signal cleanly enough to expose a soundtrack. Between 1921 and 1924 Case supplied de Forest with the components that made Phonofilm work. When de Forest presented the process publicly in 1923 without crediting the laboratory, the relationship deteriorated, and Case cut off supply in 1925.
Case then found a buyer with a studio behind him. On 23 July 1926 William Fox purchased Case's sound-on-film patents and formed the Fox-Case Corporation, and Case and his associate Earl Sponable moved to Fox to develop what became Movietone. The AEO light went into production and was used in Movietone News cameras from 1928 to 1939 and in the recording of Fox feature films from 1928 to 1931.
Fox's first strategic use of the system was not drama but news. On 20 May 1927 a Movietone crew photographed Charles Lindbergh's takeoff from Roosevelt Field with synchronized sound, and the footage was screened in a New York theater the same night. Regular Movietone News presentations began in December 1927 and the series continued in the United States until 1963. Newsreel was the ideal proving ground: the camera recorded picture and sound on one negative in one pass, so nothing could go out of register, and the material had to reach a screen within hours.
How a Variable-Density Track Works
Movietone recorded a variable-density track. The width of the exposed stripe is constant, and the audio signal modulates the intensity of the light falling on it, so the developed negative varies in optical transmission along its length. In playback, a constant light source shines through the track, and the transmitted flux, and therefore the photocell current, follows the recorded waveform.
Density recording is elegant and fragile in the same breath. Its accuracy depends on the photographic transfer function, the relationship between exposure and developed density that photographers call gamma, staying linear and consistent from the recording negative through every printing and processing stage to the release print. Gamma depends on emulsion batch, developer chemistry, temperature, and agitation. A laboratory running to a slightly different curve than the recordist assumed introduces distortion that is not obvious on inspection of the print, and densitometric control of release printing became a real discipline for exactly this reason.
Sound on Film: RCA Photophone and the Variable-Area Track
The competing optical system came out of General Electric and reached the market through RCA. Charles A. Hoxie had built a photographic recorder he called the Pallophotophone after the First World War, originally to record wireless telegraph signals, and adapted it to speech; it captured remarks by President Calvin Coolidge in 1921 for broadcast over station WGY. General Electric began developing the technique commercially for film in 1925, and in April 1928 RCA Photophone, Inc., was organized as an RCA subsidiary under David Sarnoff.
RCA needed theaters and a studio to use the system, and the corporate answer was RKO Radio Pictures, assembled from the Film Booking Offices production company, the Keith-Albee-Orpheum theater circuit, and RCA itself. RKO adopted Photophone as its house system. An FBO release, The Perfect Crime, premiered on 17 June 1928, and Syncopation followed in March 1929 as the first film recorded live with RCA Photophone.
How a Variable-Area Track Works
Photophone recorded a variable-area track. Here the exposure is nominally full or nothing, and the audio signal modulates the boundary between the clear and opaque regions of the stripe, so the developed track carries a black shape whose width traces the waveform. RCA's recorder used a fast mirror galvanometer: the audio current deflected a small mirror, the mirror swept a light beam across a triangular or V-shaped mask, and the mask converted angular deflection into a linear change in the illuminated width at the film plane. In playback, the photocell integrates the light passing through the slit, so its current follows the instantaneous track width.
The engineering advantage is that the encoding is geometric rather than photometric. What carries the signal is the position of an edge, not the greyness of an area. Photographic processing must render that edge sharply, which is a far weaker requirement than rendering a specific gamma accurately, and moderate variations in exposure and development shift the edge only slightly. Variable area is therefore substantially more tolerant of the print laboratory than variable density, and prints made under uncontrolled conditions sound closer to the original.
Both encodings share the same physical limits. High-frequency response is set by the width of the scanning slit relative to the film speed, since the slit averages over the length of track it covers, and azimuth error, meaning any tilt of the slit relative to the recorded modulation, degrades the top end sharply. The noise floor is set by film grain, by dirt and scratches, and by the dark current and shot noise of the photocell. Because grain noise scales with the illuminated area, a variable-area track can be made quieter in quiet passages by keeping most of the stripe opaque, which became the basis of the noiseless recording techniques of the early 1930s.
Variable area displaced variable density over the following two decades, and the final push came from stereophony. The Western Electric light valve, a ribbon device that varies the width of a slit under audio modulation, could record time-aligned multiple tracks, while RCA's galvanometer recorders of the day could not produce the properly registered stereo negatives that the widescreen era demanded. Once the industry required more than one channel on an optical print, stereo variable area became the standard, and it remained the analog standard for the rest of the photochemical era.
The Photoelectric Cell and the Sound Head
Reading the track is a small analog design problem with awkward constraints, and the solutions adopted in the late 1920s survived essentially unchanged for seventy years.
The reader is an assembly called the sound head, mounted below the projector head on a machine that plays analog optical tracks. An incandescent exciter lamp, run from a well-regulated direct-current supply because any ripple appears directly as hum in the audio, illuminates a mechanical slit through a small objective lens. The slit is imaged onto the moving film as a line a few thousandths of an inch high and as wide as the track. Light passing through falls on a photoelectric cell, historically a vacuum or gas-filled phototube with an alkali-metal photocathode, later a photodiode.
The cell's output is minute, of the order of microamperes into a high impedance, which makes the connection between the cell and the first stage of amplification the most delicate node in the whole theater sound system: a high-impedance, low-level circuit inside a machine containing an arc lamp, a motor, and a good deal of vibration. Preamplifiers were mounted as close to the cell as possible, the wiring was shielded, and microphonics in the first tube were a chronic complaint. This is the point at which the sound film is unambiguously an electronic technology: without gain immediately behind the photocell, there is no signal to work with at all.
The mechanical requirement is equally strict, and it is the reason the sound head is not at the picture gate. Film moves through the gate intermittently, stopping for each frame, because a projected image must be stationary while it is shown. An optical soundtrack must move past the slit at an absolutely constant speed, because any speed variation is frequency modulation of the recovered audio, heard as wow at low rates and flutter at high ones. The two motions cannot share a station, so the sound head sits separately, where the film wraps a polished sound drum carrying a heavy flywheel whose inertia filters the motion, with slack loops above and below isolating it from the gate.
The consequence is that the sound and the picture that belong together cannot occupy the same position on the film. On a 35 mm release print, the sound head is twenty-one frames after the gate, so the optical track is printed twenty-one frames ahead of the picture it accompanies. At twenty-four frames per second that is seven-eighths of a second of advance, and it is a printing convention, not an adjustment: any laboratory making a release print applies it, and every projector in the world is built to the same offset. On 16 mm prints with a mono optical track, the corresponding advance is twenty-six frames.
That offset also explains one of the medium's characteristic failures. When a print breaks and is repaired, the projectionist removes frames from a single point, but the picture and the sound belonging to that moment are twenty-one frames apart, so a splice damages the picture at one instant and the sound at a different one. The penalty is far smaller than the cumulative desynchronization that splicing inflicted on a disc system, because it is local and self-limiting, and that difference is the whole argument between the two formats in miniature.
What the Format Fixed: Frame Rate, Speed, and the Shape of the Picture
Once sound lived on the film, the film's physical parameters stopped being adjustable, and cinema acquired a set of constants it has largely kept.
Twenty-Four Frames per Second
Silent film had no fixed rate. Cameras were cranked at whatever rate suited exposure and dramatic effect, and projectionists routinely ran faster than the camera rate to shorten the running time and fit another show into the evening. Films were shot and shown at speeds that differed by twenty per cent or more, and nobody minded much.
Sound made that impossible twice over. Playback speed sets pitch, so a projectionist who runs a reel fast raises the pitch of every voice on it, an error that is immediately audible where a hurried gesture is not. More fundamentally, an optical track's high-frequency response depends on the linear speed of film past the slit, just as a phonograph groove's depends on linear velocity: at a given slit width, doubling the film speed doubles the highest frequency that can be resolved. The rate had to be fixed, and fast enough for intelligible speech.
Twenty-four frames per second was the compromise. Because 35 mm film carries sixteen frames per foot, twenty-four frames per second is exactly ninety feet per minute, or eighteen inches per second of linear track speed, which proved sufficient for the response the systems of the day could deliver. It was also close enough to prevailing silent practice that it did not waste an unreasonable amount of stock. The number has no deeper theoretical justification, and it has outlived by a century the optical constraint that produced it.
The Academy Aperture
The soundtrack also cost the picture its shape. Silent 35 mm film used the full width between the two rows of perforations, giving a frame four perforations high with an aspect ratio of about 1.33 to 1. An optical track has to go somewhere, and the only available space runs inside one row of perforations, so the sound stripe was cut out of the width of the image. Retaining the four-perforation frame height while narrowing the width produced a picture of roughly 1.19 to 1, close enough to square to look wrong on screen and to sit badly in proscenium arches designed for the older ratio.
The industry answer was to give back in height what had been lost in width. On 9 May 1932 the Society of Motion Picture Engineers adopted a projector aperture of 0.825 by 0.600 inches, roughly 21.0 by 15.2 millimeters, which the Academy of Motion Picture Arts and Sciences promoted as a standard and the trade named the Academy ratio: 1.375 to 1. Every studio film shot on 35 mm from 1932 until the widescreen upheaval of 1953 used it. The Academy frame is smaller than the silent frame in both dimensions, so sound cost the medium a meaningful fraction of its negative area, and with it some resolution and grain performance, in exchange for the stripe of emulsion that carries the sound.
The Tyranny of the Microphone
The most visible cost of the transition fell on production technique, and for about three years it made sound films look primitive beside the silent pictures they replaced.
The immediate problem is that a camera makes noise. A silent-era camera clattered, and nobody cared, because the only witness was the crew. A microphone in the same room records the clatter. The first remedy was brutal: enclose the camera and its operator in a soundproof booth, a glass-fronted cabinet the industry nicknamed the icebox, and shoot through the window. A camera in a booth cannot pan, cannot track, and cannot be moved without an hour of labor. Since a scene had to be played continuously for the recording, several such booths stood around the set, each rooted in place.
The microphone imposed its own geometry. The omnidirectional condenser microphones first used picked up the set as readily as the actor, so they had to be close, which meant hidden in a flower arrangement or behind a lamp, which meant the actor had to be close to the flower arrangement. Performers stood still and delivered lines toward the furniture. Critics who wrote about the sudden immobility of the cinema were describing an accurate consequence of the recording chain.
The recovery took most of a decade and came from four directions.
First, the camera was quieted rather than imprisoned. Sound-insulating covers built around the camera body, called blimps, and later cameras designed from the outset for quiet running, released the camera from the booth and restored the moving shot.
Second, microphones acquired directionality. RCA manufactured the Photophone Type PB-31 ribbon microphone commercially in 1931 and installed it at Radio City Music Hall in 1932; the Type 44A velocity microphone followed, with pattern and tone controls intended to manage reverberation, and unidirectional ribbon designs came out of the same laboratory during the 1930s. A microphone that rejects sound arriving from the sides and rear can be placed further from the actor for the same ratio of wanted to unwanted sound, and it can be aimed. Mounted on a boom, it could follow a moving performer, which is what finally reconciled the microphone with blocking.
Third, the industry learned to mix. Once several microphones fed a console attended by a sound engineer, coverage no longer depended on a single compromise position.
Fourth, and most consequentially, sound was separated from the moment of photography. Rerecording, which combines dialogue, music, and effects tracks in a later pass, and post-synchronization, in which an actor watches the picture and replaces the dialogue in a studio, were practical by the middle of the 1930s. Post-synchronization restored the property that had made silent cinema flexible: the ability to shoot for the picture and fix the rest afterward. It also created the modern sound department, in which the recording made on set is raw material rather than a finished product.
Wiring the Theater
On the exhibition side, the transition meant installing a small electroacoustic plant in every auditorium in the country.
The chain begins at the sound head and ends at horns behind the screen. Because the loudspeakers must be heard through the picture, the screen was replaced with perforated material, opaque enough to hold an image and open enough to pass sound. Screen perforation is a permanent feature of theatrical exhibition, and it dates from this moment.
Loudspeakers were horn-loaded from the beginning, and not for aesthetic reasons. A direct-radiating cone driver converts a small percentage of its electrical input into acoustic power, while a horn matches the high acoustic impedance at the diaphragm to the low impedance of the room and raises efficiency by an order of magnitude or more. With amplifiers delivering tens of watts, that difference decided whether a large house could be filled at all. Early Western Electric installations paired compression drivers with large exponential horns, and the arrangement set the acoustic character of the American cinema for a generation.
Single-horn systems could not cover the frequency range well, and by 1931 theater systems were being divided into separate low-, middle-, and high-frequency sections with filters distributing the signal among them. The best-known result was the two-way system developed at MGM under Douglas Shearer, head of the studio's sound department, with John Hilliard leading the engineering team and drivers supplied by James B. Lansing's Lansing Manufacturing Company. The Shearer horn used multiple fifteen-inch field-coil drivers in a folded bass horn crossed over near five hundred hertz to a multicellular high-frequency horn, the multicellular construction being a means of controlling coverage across a wide auditorium. The Academy recognized the development with a technical award in 1936, and systems built on the same plan became the industry standard.
Standardizing the loudspeakers made it possible to standardize what was fed to them. From 1938 the industry worked to a specified theater playback response, the Academy curve, which rolled off the high frequencies substantially. It was a practical accommodation to the noise of an optical track and to the horn systems then installed, and once it existed, mixers could balance a film in a studio and expect a comparable result in a distant house. It governed theatrical sound until Dolby's work in the 1970s made a wider response practical.
The building itself needed work. Auditoriums designed for live performance were generally too reverberant for reproduced speech, since reverberation smears consonants and destroys intelligibility long before it becomes objectionable on music. Absorptive treatment on walls and ceilings became a routine part of conversion, and theater acoustics shifted from supporting a live source to serving a reinforced one. Projection booths needed room for amplifier racks and power supplies, many buildings needed their electrical service upgraded, and projectionists needed training in equipment that had nothing to do with the craft they had learned.
What Conversion Cost
The bill for all this was large, and it arrived at the worst possible moment in American economic history.
Wiring a theater for sound cost on the order of fifteen thousand dollars, a sum on the order of a quarter of a million dollars in early twenty-first-century money, and there were more than twenty thousand motion picture theaters in the United States. The money went to ERPI or to RCA Photophone, in equipment purchases, installation, service contracts, and continuing license fees, which is why the transition transferred so much value from the film business to the electrical business.
Conversion was fast where capital was available and slow where it was not. Even in the United States, only about half of the theaters had been wired by 1930, which is why studios released features in both sound and silent versions well into that year. Britain had more than sixty per cent of its theaters equipped by the end of 1930. In France more than half were still projecting in silence in late 1932. In the Soviet Union, as of May 1933, fewer than one film projector in a hundred was equipped for sound.
The uneven pace had structural consequences. A small independent house that could not raise fifteen thousand dollars lost its audience to the wired theater down the street, while chains that could finance conversion across a circuit emerged larger. Studios borrowed heavily to build soundstages and re-equip, which brought investment banks into the governance of film companies to a degree they had not previously enjoyed. The consolidation of the American industry into a small group of vertically integrated majors, which later antitrust litigation would unwind, was accelerated by the capital requirements of sound.
Who Paid: Musicians, Performers, and the Language Barrier
The clearest human cost fell on musicians. Silent film was never silent in exhibition; it was accompanied, by a pianist or an organist in a small house and by a substantial orchestra in a first-run palace. Accompaniment was one of the largest sources of employment for instrumentalists in the United States, and a recorded score eliminated it. The American Federation of Musicians recorded that some twenty thousand musicians lost their theater work within two years of the arrival of talking pictures, and in 1930 the union founded the Music Defense League to campaign against what it called canned music. The campaign lost. It is one of the earliest instances of organized labor confronting the substitution of a recording for a live performance.
Performers were affected unevenly, and the popular story exaggerates. Careers did end because a voice did not suit the microphone, but the larger disruption was that sound imported a different professional standard. Studios recruited from the theater, where actors had trained to speak, and the skills that had made a silent screen performer valuable were not the same skills.
The most far-reaching commercial consequence was linguistic. A silent film crossed borders by having its intertitles retranslated and reprinted, an operation cheap enough to be an afterthought. A talking picture is locked to the language of its dialogue. Hollywood had built an export business on a product that was effectively language-neutral, and sound broke it.
The first response was to shoot the film again. From 1929 studios produced multiple-language versions, remaking a picture on the same sets with the same crew and costumes but a different cast speaking a different language, most often English, Spanish, French, or German. Paramount ran a dedicated facility at Joinville outside Paris for the purpose, and MGM, Universal, Warner Bros., and the German studio UFA all made them. The practice peaked in the early 1930s and then collapsed, partly because shooting a film three or four times was indefensibly expensive and partly because the alternatives improved. Dubbing, which replaces the dialogue track, and subtitling, which requires only a laboratory, were both far cheaper, and by the mid-1930s the market had divided into dubbing territories and subtitling territories along lines that persist today.
What Came After, in Outline
The transition described here settled the format. Everything since has been an improvement within it, and a brief sketch is enough to place the later waves.
Fidelity improved through the 1930s by attacking the optical track's weaknesses directly: noiseless recording that reduced grain noise in quiet passages, push-pull tracks that cancelled even-order distortion by recording two mirror-image modulations read differentially, and better transducers at both ends.
Multichannel sound arrived experimentally with Disney's Fantasound, developed with RCA for Fantasia in 1940, which carried its audio on a separate synchronized film and steered it among loudspeakers in the auditorium. It was too expensive to install widely, and the war ended the experiment.
Magnetic sound and widescreen arrived together in the 1950s, driven by competition with television. Magnetic stripes laminated onto the release print carried discrete channels with a signal-to-noise ratio and a frequency response that optical tracks could not approach: CinemaScope introduced four-track magnetic stereo in 1953, and Todd-AO brought six-track magnetic sound on 70 mm prints in 1955. Magnetic prints were expensive to strike and the oxide wore, so the format stayed at the premium end of exhibition.
Dolby Laboratories changed that in 1976 by making the optical track good enough. Dolby Stereo applied A-type noise reduction to a stereo variable-area track, lifting the medium's dynamic range enough to abandon the Academy curve, and used a four-to-two-to-four matrix to fold left, center, right, and a surround channel into the two optical tracks a print could carry. Center information was distributed equally between the tracks and surround information encoded out of phase between them, so a decoder could recover four channels by summing and differencing. Because a matrixed print plays acceptably as ordinary stereo or mono on unequipped machines, exhibitors could convert at their own pace. Dolby SR replaced A-type on release prints from the late 1980s.
Digital soundtracks arrived in the early 1990s in three incompatible forms, each of which found a different unused part of the print. Dolby Digital, introduced in 1992 and first widely released on Batman Returns, printed its data between the perforations on the soundtrack side of the film, carrying 5.1 channels at three hundred and twenty kilobits per second. Sony Dynamic Digital Sound, introduced the same year, used both outer edges of the print. DTS, founded in 1990, took the opposite approach: it printed only a timecode track on the film and played the audio from separate CD-ROM discs locked to it, an arrangement that debuted on Jurassic Park in 1993 and bought bandwidth at the cost of another object to keep with the print. Prints of the period commonly carried all three digital formats plus a Dolby SR analog track, so that any theater could play any print. That redundancy is a direct descendant of the format war of 1927, and its lesson had evidently been learned.
Digital distribution ended the argument by ending the print. When a feature travels as an encrypted digital cinema package and plays from a server, the soundtrack is a set of files with a timeline, and synchronization is guaranteed by the clock that drives the image. The problem that defeated Edison for twenty years, and that Vitaphone solved so precariously, no longer exists in a form that can fail.
Conclusion
The coming of sound to motion pictures is usually told as a story about The Jazz Singer and about actors with unsuitable voices. The engineering account is more useful. Synchronized sound waited for the vacuum tube because both of its obstacles were, at bottom, the same obstacle. Without gain there was no way to fill a theater from a recording, and without gain there was no way to read a signal weak enough to be carried on the film itself. Once the amplifier existed, both problems became solvable, and they were solved in two different ways within eighteen months.
Sound on disc won the market and lost the medium. It was quicker to deploy because a disc recording chain already existed and could be applied to film with little modification, and it proved to a skeptical industry that audiences would pay. But it kept the sound in a separate object, and every failure mode that followed from that separation, the splice, the reel change, the worn disc, the uneditable master, was structural. Sound on film moved the audio onto the same strip as the image, which made synchronization a property of the material rather than a service performed by a mechanism, and no amount of refinement to the disc system could have matched that.
The winning format then imposed its constraints on everything downstream. It fixed the frame rate at twenty-four per second because the optical track needed a known and adequate linear speed. It put the sound twenty-one frames ahead of its picture because film cannot move intermittently and continuously at the same point. It took a stripe out of the width of the frame, which forced the Academy aperture and gave classical cinema its shape. For a few years it rooted the camera and the actors in place, until directional microphones, blimped cameras, and post-synchronization gave the medium its mobility back.
And it was expensive. It moved a great deal of money from the film industry to the electrical industry, accelerated the consolidation of exhibition, ended a major category of musical employment, and broke the language-neutral export business on which Hollywood's international position had been built, replacing it with the dubbing and subtitling economy the world still uses. Sound was not an addition to the motion picture. It was a rebuilding of it, and the electronics were the reason it could be done at all.