Sound Fundamentals Recap with Extras
1 - Sound Homework Recap with Extras
Sound Homework Recap with Extras (Sound Fundamentals)
ARTD 2380 Video Basics, David Tamés
Updated September 26, 2024
See also: Location Sound Recording
This slide deck is updated periodically; if you have any comments or suggestions for improvement, please contact me: https://davidtames.com/contact/
© 2024 by David Tamés, some rights reserved, shared under a Creative Commons Attribution-NonCommercial-ShareAlike License, https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode.en, some materials included herein may be copyright and are being used under the terms of the fair use provisions of U.S. Copyright Law. Third-party works are attributed whenever possible.
2 - Essential reading and viewing
Essential Reading:
Chapter 6. Sound in Making Media: Foundations of Sound and Image Production by Jan Roberts-Breslin, Fourth edition is available from the Snell Library, currently available in a Firth Edition, https://amzn.to/42UzRtC
Essential Viewing: Justin Boyd: Sound and Time (Mark Lee Walley and Angela Guerra Walley, 2013, short documentary), https://vimeo.com/78213028 The Foley Artist (Oliver Holms, 2015, short film), https://vimeo.com/124053378 Recording Sound on Location (Lizi Hesling, CADARN Learning), https://www.youtube.com/watch?v=TKBzjSSaKXU The Basics of Recording Audio for Digital Video (Filmmaker IQ), https://www.youtube.com/watch?v=S9cP1WHL0Zo UCLA Post Production: How To Wrap A Cable (David McKenna), https://youtu.be/uy3axdxDdKs
3 - Summary of essential terms and concepts
Terms and concepts covered in Chapter 6. Sound in Making Media: ambience, amplitude, attenuate, balanced audio, bidirectional, binaural hearing, bit depth, cardioid, compression, condenser, decibel (dB), dynamic, frequency, frequency response, handheld microphone, harmonics, Hertz (Hz), impedance, lavaliere, level meter, lossless compression, lossy compression, monitoring, mono, omnidirectional, on-axis, overmodulation, peak, phantom power, pickup pattern, pitch, quantized, radio frequency, reverberation, ride the gain, room tone, sampled, sampling rate, shotgun microphone, signal-to-noise ratio (S/N), sound envelope, sound perspective, sound presence, stereo, streaming, timbre, unbalanced audio, unidirectional, wavelength, waves, XLR.
Additional terms that will be covered in the workshop: 16-bit, 24-bit, 48KHz, 96KHz, boom pole, dual mono, Electro-Voice RE50, K-Tek Avalon boom pole, reporter s mic, Rycote Pistol Grip, Rycote Softie, SD card, Sennheiser MKE 600, Stereo WAV, Tram TR-50, WAV, windjammer.
4 - What is sound?
What is sound? Sound is a form of energy created by physical vibrations that set molecules in motion, creating sound waves that travel through the air. These pressure waves of compression and rarefactions travel through a medium such as air, water, or solids. When these sound waves reach our ears, they cause our eardrums to vibrate, which our brain interprets as sound.
5 - Wavelength and frequency
Sound is vibrations that travel through a medium and are detected by our sense of hearing. Characteristics of sound include the frequency that determines the pitch of the sound, which is measured in hertz (Hz). Higher frequencies produce higher-pitched sounds. Amplitude determines the loudness of the sound; larger amplitudes mean louder sounds. The shape of the waveform affects the timbre or quality of the sound, distinguishing different sources like a piano and an electric guitar playing the same note. Sound waves rarely look like the perfect sine wave pictured here, which is something you would expect out of a tone generator but not natural sounds. Sound travels at different speeds depending on the medium. In air, it travels at approximately 343 meters per second at 21 C or 1,130 feet per second at sea level and at 70 F.
6 - The auditory system
The auditory system processes sound waves into electrical impulses and then interprets these impulses as recognizable sounds. The journey of a sound begins as pressure variations in a medium. The outer ear, including the pinna and ear canal, collects and funnels sound waves toward the eardrum (tympanic membrane). When sound waves hit the eardrum, it vibrates. These vibrations are passed to three tiny bones of the middle ear: malleus (hammer), incus (anvil), and stapes (stirrup). These bones amplify the vibrations and transmit them to the inner ear’s cochlea, a fluid-filled, spiral-shaped organ. Inside the cochlea, the vibrations are converted into fluid waves. This fluid movement activates specialized cells equipped with stereocilia (hair-like projections) that bend in response to the waves in the fluid. When the stereocilia bend, they create an electrical signal. The movement of these cells in different parts of the cochlea corresponds to varying sound frequencies, with high frequencies near the base and low frequencies near the apex. The electrical signals are then transmitted to the brain by the auditory nerve. The auditory nerve sends these impulses to the brainstem and ultimately to the auditory cortex. The auditory cortex, located in the brain’s temporal lobe, processes and interprets these electrical impulses. It analyzes aspects such as pitch, loudness, timbre, and direction. Our ability to interpret patterns of sounds enables us to understand speech, music, sounds in the environment, and even subtle variations in tone or rhythm. Memory and prior experience play a key role in interpreting meaning and context.
For more about the auditory system, told in an engaging and innovative manner, listen to the Musical Language episode of Radiolab is recommended, https://radiolab.org/podcast/91512-musical-language. Editor and sound designer Walter Murch tells a story of his experience with how context and experience can change our perception of music in the Making Radiolab episode of Radiolab https://radiolab.org/podcast/91746-making-radio-lab.
7 - What are the terms we use to describe a sound?
Pitch A perceptual quality that makes it possible to judge sounds as “higher” and “lower” and a major attribute along with duration, loudness, timbre and perspective. May be quantified as a frequency, but it is actually a subjective psycho-acoustical attribute.
Loudness The subjective perception of sound pressure. Perceived loudness consists of physical, physiological and psychological components. Phon is a logarithmic loudness unit of level for tones and complex sounds the accounts for variable sensitivity across different frequencies. Phon is a unit that measures how loud a sound feels to us, similar to how decibels (dB) work, but it focuses on what we actually hear rather than just the physical sound level. The phon scale helps us understand loudness in a way that matches how we experience sound in real life, not just the physical measurement of sound pressure. The phon unit is based on how loud a sound feels compared to a standard sound with a frequency of 1 kHz (which is often used in hearing tests).
Timbre The subjective perception of sound that makes it possible to judge sounds with the same loudness and pitch as qualitatively different. A harmonic is a wave with a frequency that is a multiple of the fundamental frequency, also called the 1st harmonic, other harmonics are known as higher harmonics. Harmonic Content, along with Attack and Decay (Envelope), and Vibrato/Tremolo, contribute to Timbre. See slide 12 for more details.
Presence, reverberation, direction, and movement contribute to the perspective of a sound. See slide 13 for more details.
8 - Pitch is a perceptual quality related to frequency
Pitch is a perceptual quality that allows us to judge sounds as “higher” and “lower.” Along with duration, loudness, timbre, and perspective, pitch is a significant attribute of sound. Although pitch can be quantified as the frequency of sound waves, it is actually a subjective psycho-acoustical attribute. Frequency is an objective measurement; in contrast, pitch is more about how we experience the sound. Sounds at very low or high frequencies are not perceived as strongly as mid-range sounds, even though their frequencies can be precisely measured. The perception of pitch can change based on context. When listening to complex sounds, the pitch of each individual sound might be heard differently depending on the surrounding sounds. Cultural and psychological factors can also influence pitch perception.
9 - Loudness
Loudness The subjective perception of sound pressure. Perceived loudness consists of physical, physiological and psychological components. Phon is a logarithmic loudness unit of level for tones and complex sounds the accounts for variable sensitivity across different frequencies. Phon is a unit that measures how loud a sound feels to us, similar to how decibels (dB) work, but it focuses on what we actually hear rather than just the physical sound level. The phon scale helps us understand loudness in a way that matches how we experience sound in real life, not just the physical measurement of sound pressure. The phon unit is based on how loud a sound feels compared to a standard sound with a frequency of 1 kHz (which is often used in hearing tests).
10 - Loudness
Loudness consists of physical, physiological, and psychological components.
Sound pressure is objectively measured with a logarithmic decibel scale related to the threshold of hearing.
Phon is a unit that measures how loud a sound feels to us, similar to how decibels (dB) work, but it focuses on what we actually hear rather than just the physical sound level. The phon scale helps us understand loudness in a way that matches how we experience sound in real life, not just the physical measurement of sound pressure. The phon unit is based on how loud a sound feels compared to a standard sound with a frequency of 1 kHz (which is often used in hearing tests).
Maximum sound up to 8 hour (OSHA criteria - hearing conservation program) is 80 dB
Some common sounds and their typical sound pressure levels (SPLs):
Maximum Theoretical Sound (194 dB) Jet Aircraft (during takeoff) (133 dB) Threshold of Pain (125 dB) Thunderclap (near) (120 dB) Loud Rock Concert (115 dB) Sonic Boom (110 dB) Shouting in Ear (110 dB) Chainsaw (104 dB) Night Club (3’ from speaker) (100 dB) Motorcycle (98 dB) Lawn Mower (90 dB) Blender (82 dB) Road with busy traffic (80 dB) Vacuum Cleaner (75 dB) Washing Machine (70 dB) City Traffic (inside a car) (70 dB) Normal Conversation (62 dB) Rainfall (50 dB) Quiet street (50 dB) Quiet home (40 dB) Bird Call (40 dB) Soft Whisper (30 dB) Rustling Leaves (20 dB) Whispering (at 5) (20 dB) Normal Breathing (10 dB) Threshold of Hearing (0 dB)
11 - decibel (db)
The decibel is a logarithmic units of measurement that expresses the magnitude of acoustic energy or an electrical signal relative to a specified reference level. It expresses a ratio of two quantities. The decibel is used for measurements in acoustics, electronics, and a range of other fields. The logarithmic scaling corresponds to the human perception of sound and the provides the ability to carry out multiplication of ratios by simple addition and subtraction. The definition of the decibel use base-10 logarithms. For example, a 10dB change is perceived as a doubling of sound level, thus a sound that measures 56dB would be perceived as twice as loud as one that measures 46dB.
The decibel is always relative to some 0 dB reference. In acoustics the reference level is typically set at the threshold of perception of an average human and there are common comparisons used to illustrate different levels of sound pressure. When we re using decibel to refer to audio signal levels when we re recording, the 0 dB reference level is the full scale signal that can be recorded, referred to as 0 dBfs (full scale).
One reason we use the decibel is that the human ear is capable of detecting a very large range of sound levels. Because the power in a sound wave is proportional to the square of the pressure, the ratio of the maximum power to the minimum power is above a trillion. To deal with such a wide, unwieldy range, logarithmic units are used. For example, the log of a trillion is 12, so this ratio represents a difference of 120 dB. For SPL measurements, it s easier to work with numbers in a range of 0 to 120 than a trillion, just remember that each plus 6dB change is perceived as approximately twice as loud, each minus 6 dB change is perceived as approximately half as loud. So when looking at the polar chart or frequency response diagram, you can see how the audio level falls off in terms of relative dB. Typically you ll be making changes when editing and mixing in 3 dB increments.
12 - Timbre
Timbre makes different sounds unique, even if they have the same pitch and loudness. For example, a piano and an electric guitar playing the same note at the same volume still sound different because of the timbre. The primary contributors to the quality or timbre of a sound are harmonic content, envelope, vibrato, and tremolo.
Timbre - Harmonic Content: Harmonic content is the most important of these, it is the number and relative intensity of the upper harmonics present in the sound. A sound wave consists of a fundamental frequency and a series of harmonics (an integral multiple of the frequency of the same signal). The auditory system can recognize these patterns and we use it to differentiate sounds. While sounds may have the same pitch and loudness, they will sound different based on their harmonic content which determines the waveform of the sound signal when displayed as a function of time. When an instrument or voice produces a sound, it’s not just making a single pure note (called the “fundamental”). It’s also creating a series of higher-pitched notes called harmonics or overtones. These harmonics are quieter than the fundamental note but they shape the overall sound. The combination of the fundamental frequency and unique harmonic patterns gives each sound its distinctive character. The more complex the harmonic content, the richer or more “colorful” the sound.
Timbre - Envelope (Attack and Decay): In addition to harmonic content, timbre is shaped by how a sound starts and fades. These characteristics are called attack and decay, and they play a significant role in recognizing different sounds. The attack is the very beginning of a sound, or how quickly it reaches its peak loudness after it’s played. Sounds can vary in terms of attack times. For example, a piano note has a quick, sharp attack because the sound is made instantly when a hammer strikes the string. On the other hand, a bowed violin note has a slower attack because the sound builds up more gradually as the bow contacts and then moves across the string. The decay is how a sound fades after the initial attack. Some sounds fade quickly, while others fade slowly. For example, a drum hit has a fast decay, meaning the sound dies out quickly after the strike. However, the sound of strumming strings on an acoustic guitar or a sustained vocal note will have a slower decay, meaning it holds the tone for a longer period of time before fading.
Timbre - Vibrato/Tremolo: Vibrato and tremolo add variations to a sound. Vibrato is a slight, rapid variation in pitch and can make a sound feel warmer or more alive. Singers, violinists, and guitarists often use vibrato to give their notes a richer, more emotional quality. Tremolo is a rapid change in loudness rather than pitch. It involves quickly increasing and decreasing the loudness of a sound, giving it a trembling, adding a sense of movement and texture to the sound.
13 - Perspective
Presence, reverberation, direction, and movement contribute to the perspective of a sound
Perspective - Movement, Doppler Effect
Perspective - Presence: Sound presence defines the subjective closeness the listener feels to the sound. A recorded voice that sounds close to your ear with little background noise will evoke a sense of intimacy and presence. A recorded voice with little presence sounds distant, with a good deal of air and background sound mixed in. It has little to do with loudness but can be manipulated by your choice of microphone and the placement of the microphone in relation to the sound source. A lavaliere clipped to an actor s or subject s clothing or a handheld microphone with an excellent frequency range held close to the speaker s mouth will deliver a strong sensation of sound presence. A shotgun at a distance of five feet will have less presence, even when properly aimed and with good levels. You usually want to match sound presence with shot size, if shooting video. A close-up matches a strong presence. A long shot should sound farther away. This effect can be created in post-production, so the rule of thumb in production is to record the best quality dialogue, sound effects, and ambiences you can, as close as practical to the source, and then create the appropriate sound presence in post-production. It s easier to manipulate a clean track than to fix a poorly recorded track. Most ambiences and some sound effects should be recorded in stereo to increase the sense of presence. A busy marketplace recorded in stereo will give the listener a feeling of being present in a three dimensional environment, while a mono recording of a busy marketplace will sound more like noise without any spatialization.
Perspective - Reverberation: Refers to how sound reflects off a space’s surface, creating a sense of depth and atmosphere. Reverberation shapes our perception of where a sound is coming from and the characteristics of the environment it is in. When sound waves bounce around a room or environment, we hear these reflections along with the original sound. This mixture is what we call reverberation or reverb for short. Reverberation gives us clues about the size and type of space we are in. For example, in a small room, the sound waves reflect quickly, and the reverb fades out faster. In a large gymnasium, the reverb lasts longer because the sound waves travel farther before bouncing back. The amount and type of reverberation help our brain understand whether the sound is close or far away, in a small room, in an open space, in the country, in an urban environment, and so on and so forth. We often will add reverberation to sound effects and sound mixes to establish the perspective of a sound or establish the location of the audience in the story world. Reverberation changes the timbre of a sound by adding extra reflections that blend with the original sound. This can make the sound feel fuller, richer, or more distant, depending on the space. By adding different amounts of reverb with different room models, you can create a more immersive experience for the viewer, giving them a feeling of where sounds are coming from in a three dimensional space. Dialogue with little to no reverb sounds close and intimate, as if the person is right next to you. On the other hand, a voice with much reverb will sound as if they are in a large gallery, cathedral, or a similar large space and more distant from us.
Perspective - Direction, Binaural Perception
14 - The audio signal chain from source to reproduction
The vibrating air molecules create a wave through which the sound energy moves through the air; the wave of sound energy leads to the air molecules surrounding the diaphragm in the microphone to move; these movements are converted into an analog electrical signal by the microphone, the analog electrical signal (an audio signal) is fed through a microphone cable into a recorder or a camera, the recorder or the camera s audio input circuits converts the analog audio signal into a stream of digital bits through a process of analog to digital conversion, the stream of bits are stored on a storage medium (usually an SD card) as uncompressed digital samples of the original analog data. Playback: reverses the process of recording, sending the signal through a media player to an amplifier and then to speakers or headphones or earbuds which are re-creating the original vibrations and eventually translated to the perception of sound by our nervous system. We need not go deep into the physics or neuroscience to appreciate this. What is important to understand is the essence of sound, think of sound as one of the materials you work with, you shape and mold it to create your work. It is a medium just as paint is a medium. Note: Some newer recorders and cameras support 32-bit float encoding, a format increasing in popularity since it eliminates gain adjustments during production since it can record the entire dynamic range the microphone is capable of capturing. However, it requires more management in post-production.
15 - Microphone transducer technologies
Dynamic and condenser designs are widely used in video production.
Dynamic microphones ⇐> Condenser microphones:
Dynamic: Moving-coil transducer design ⇐> Condenser: Charged-plate transducer design;
Dynamic: Less complicated and inexpensive ⇐> Condenser More complicated and expensive;
Dynamic: Usually rugged ⇐> Condenser: Delicate;
Dynamic: No external power required ⇐> Condenser: Requires power source (battery or phantom power) ;
Dynamic: Low sensitivity (requires close proximity, less sensitive to wind noise) ⇐> Condenser: High sensitivity (provides greater working distance, more sensitive to wind noise); and
Dynamic: Typically limited to hand-held microphone designs ⇐> Condenser: Available in a variety of form factors including hand-held, lavaliere, and directional microphone designs.
16 - Microphones have directional characteristics
A pickup pattern chart or a polar chart shows you the response in dB depending on position of source relative to the microphone, response also depends on frequency (shown as multiple lines each with different hatching). Take a look at the Short Shotgun polar pattern, notice how sounds on-axis are not attenuated, while sounds arriving at the microphone off-axis are attenuated, each of the rings corresponds to -5dB increments of attenuation.
Note how there s a lobe directly behind the microphone, while the null-spots are at around 125 degrees off-axis. Also note that directionality varies based on the frequency of the source. Microphone specs will also show a frequency response diagrams showing you the response in dB relative to frequency of the source being recorded with a microphone, the flatter the response the more accurate the microphone. Some frequency response diagrams will also show you what the frequency response looks like off axis.
Pickup patterns define how microphones respond to sound arriving at them from different positions in space. Some microphones accept sound from all angles, others will attenuate off-axis sound, but they not only reduced intensity, of off-axis sounds but also exhibit some coloration (change in the characteristics of the sound).
Omnidirectional: Omnidirectional microphones record sounds arriving from all positions relatively equally. Generally flatter in terms of frequency response compared to more directional microphones. They works well in close proximity. Good for recording hand-held stand up interviews, dialogue in controlled environments, ambient sound, music sources. Not very good in high ambient noise areas.
Cardioid: Cardioid microphones have a heart-shaped pick-up pattern with some rejection of sounds from behind. Generally flatter than more directional microphones, Works well in close proximity. Good for recording dialogue in controlled environments, ambient sound, music sources. Not the best choice for high ambient noise areas.
Shotgun: Shotgun microphones and super-cardioids with a lobar patters and thus more directional. Generally not as flat as cardioid or omnidirectional microphones, off-axis sound exhibits distinctive coloration. These have varying degrees of directionality, varies based on frequency of source. Good for recording when it is desirable to focus on a specific sound source and where isolation from unwanted sounds or noise is needed and when a greater working distance is required.
There are other pick up patterns and types of microphones, but the omnidirectional, cardioid, and shotgun (a.k.a. lobar) are the most common in video production.
17 - Microphones commonly used in video production
There are three microphone form factors commonly used in video production: handheld, lavalier, and boom-mounted microphones (which may also be used hand-held in a pistol grip or mounted directly on the camera). Handheld microphones (typically dynamic but some condenser models exist) are good for use in high-noise environments when close placement is possible, typically omnidirectional (e.g. Electro-Voice RE50) or cardioid (e.g. Sennheiser MD46) in terms of their pick-up pattern. Some are designed to reduce the effect of handling noise. Lavalier microphones (typically condenser) are good for use on subjects (or actors) when boom, handheld, or camera microphones will not do the trick in terms of proper placement. These are most typically omnidirectional in terms of pickup pattern, however, cardioid lavs are available, but placement much tricker, only for specialized uses.
Wireless lavalier microphones are lavaliers paired with a transmitter that is attached to the subject (or actor) which transmits the audio signal to a receiver connected to the camera or audio recording device. They are widely used in video production to give subjects (or actors) mobility without the hassle of wires. Plug-on transmitters are also available that can be plugged into a handheld or boom-mounted microphone to turn them into wireless mics as well. While trickier to use than wired microphones, wireless microphones are invaluable when subjects (or actors) are on the move or need to wander far from where they are easy to mic with a hand-held, boom, or on-camera microphone.
18 - Balanced vs. unbalanced microphone cables
Balanced microphone cables with XLR connectors are widely used for professional audio. Balanced wiring has two conductors and one ground plus special circuitry on each end in order to reduce susceptibility to electromagnetic interference (EMI) and allow for longer cable runs. The best cables use twisted pairs for each conductor ( star-quad ) for even better EMI rejection. Consumer gear uses wiring with only one conductor and one ground and is therefore more susceptible to interference. For a technical explanation of how this works, see What s the Difference Between Balanced and Unbalanced?, Aviom Blog, https://www.aviom.com/blog/balanced-vs-unbalanced/
Make sure you have the right cable lengths you ll need for your shoot. It s important to wrap cables properly. This is done for two reasons: 1. maintaining the life of the cable, and 2. it enables you to unwrap it without knotting or tangling. When done properly, you can toss the coiled cable, holding on to one end, and it will unwrap completely straight. You will find that making things neat on a production makes the work go more smoothly. See UCLA Post Production: How To Wrap A Cable (David McKenna), https://youtu.be/uy3axdxDdKs
19 - Analog to digital conversion
Analog audio signals are digitized, a process of quantizing a continuous analog value taking samples at a particular frequency (sampling rate) converting the analog values into a discrete values using a 16-bit or 24-bit fixed point binary number (bit-depth). 48 kHz / 24-bits is the video production audio recording standard. 96 kHz / 24-bits is often used for sound effects recording; the higher sampling rate allows for manipulation of the audio in post with less artifacts. WAV is an uncompressed audio file format commonly used in professional audio; polyWAV (multi-channel WAV) and BWAV (broadcast WAV) variations add multiple channels and additional metadata, respectively. Most current cameras record audio 24-bit/48KHz or 16-bit/48KHz, digital audio recorders can be set to record to many different formats. With digital audio recorders, I suggest choosing uncompressed WAV ( Wave ) with 16-bit or 24-bit at 48KHz for dialogue and 24-bit at 96KHz for sound effects that might undergo a lot of processing in post. You want to avoid compressed formats like MP3 for original recording and editing in post. Some newer recorders and cameras support 32-bit float encoding, a format increasing in popularity since it eliminates gain adjustments during production since it can record the entire dynamic range the microphone is capable of capturing. However, it requires more management in post-production.
20 - Review: Recording Sound on Location (video)
Recording Sound on Location (Lizi Hesling, CADARN Learning), https://www.youtube.com/watch?v=TKBzjSSaKXU provides an excellent introduction to the topic.
21 - Review: Recording Sound on Location (video)
What are the three Cs of good location audio?
22 - Review: Recording Sound on Location (video)
What are the three Cs of good location audio?
Good audio contributes to your audience immersing themselves in your video; you want your audio to be:
- Clear: you can clearly hear the sound that you re supposed to be hearing;
- Clean: there aren t any unwanted or distracting noises affecting the quality of the audio; and
- Consistent: the volume and quality don t keep changing unless the story or visuals are called for.
23 - Review: Recording Sound on Location (video)
What are the five principles of good location sound recording practice?
24 - Review: Recording Sound on Location
What are the five principles of good location sound recording practice?
- Assess the environment to avoid noise as much as possible
- Always monitor your audio with good headphones
- Know thy microphones
- Get close to the source
- Get your levels right