Solent University SCHOOL OF MEDIA, ARTS AND TECHNOLOGIE BSc (Hons) Popular Music Production 2023 / 2024 Sergio Quiroga Ortego “Delivering High Fidelity Audio for Streaming Platforms” Supervisor : Neil Kennedy Date of submission : May 2024 1 Acknowledgements First of all, I would like to extend my deepest gratitude to my family, who have consistently supported me throughout my academic journey. Their unwavering encouragement has been a cornerstone of my progress. I am also profoundly thankful to my initial mentors in the field of music production. These individuals were not merely educators but also guides who shepherded me through both the technical intricacies and the ethical dimensions of the profession. Their invaluable advice and relentless support have been instrumental in my continuous personal and professional development. Lastly, but certainly not least, I owe a significant debt of gratitude to the wider community of audio engineers, producers, and related professionals. This vibrant collective has always been a rich source of education, fostering an environment ripe with debate and collaboration. Their collective wisdom and camaraderie have greatly enriched my educational experience. 2 Contents Acknowledgements .................................................................................................... 2 Contents ...................................................................................................................... 3 List of Tables and Figures .......................................................................................... 5 Literature Review ........................................................................................................ 6 1. Introduction.......................................................................................................... 8 1.1 Streaming Growth in Music Distribution ......................................................... 8 1.2 Audio quality and listener experience ........................................................... 10 1.3 Goals ................................................................................................................ 12 2. Streaming Platforms and Loudness ................................................................. 14 2.1 Loudness Standards ....................................................................................... 14 2.2 The Loudness Normalization Process ........................................................... 15 2.3 Album and Track Normalization of Music for Distribution............................ 16 2.4 Normalization Workflow .................................................................................. 18 2.5 Formats and Codecs ....................................................................................... 19 3. Human Hearing .................................................................................................. 22 3.1 Our Ears ........................................................................................................... 22 3.2 Audible Spectrum ............................................................................................ 24 3.3 Dynamic Range ............................................................................................... 27 4. Pursuing Optimal Audio Quality for Playback ................................................. 29 4.1 Sampling Theorem .......................................................................................... 29 4.2 Aliasing ............................................................................................................ 32 4.3 High-Resolution Audio or Higher Sample Rates for Playback ..................... 34 4.4 Higher Sample Rates Being Harmful for Playback........................................ 36 4.5 Oversampling .................................................................................................. 39 4.6 Bit Depth .......................................................................................................... 40 4.7 How to Actually Improve Audio Quality for the End User ............................. 43 5. Modern Production Workflows ......................................................................... 44 5.1 Production Frameworks ................................................................................. 44 5.1.1 Recording Sample Rates and Bit Depths ................................................ 45 5.1.2 Production Sample Rates and Bit Depths ............................................... 46 5.2 Loudness ......................................................................................................... 49 5.3 Good Practices ................................................................................................ 53 3 6. Conclusion ............................................................................................................ 56 7. Bibliography .......................................................................................................... 60 8. Appendices ............................................................................................................. A 8.1 Appendix A: Lossless Distribution .................................................................. A 8.2 Appendix B: High-Frequency Auditory Perception and its Impact on HighEnd Audio Content Sampling ................................................................................. B 8.3 Appendix C: Considerations for Antialiasing Linear Phase Filters............... C 8.4 Appendix D: Higher Sample Rates association with Professional Audio..... D 8.5 Appendix E: Nuances of Loudness Perception............................................... E 4 List of Tables and Figures Figure 1: Stacked Bar Graph of Global Recorded Music Industry Revenues 1999-2023 (US$ Billions). .................................................................................. 9 Figure 2: Pie Chart of Global Recorded Music Revenues by Segment 2023. . 10 Figure 3: Multi-Line Graph of Listener’s Behavior After a Change in Loudness Occurred IF the Change Was Frequent.............................................................11 Figure 4: Table of Max Integrated Loudness and TP Requirements of Streaming Platforms. ........................................................................................ 15 Figure 5: Illustration of the Normalization Process. ......................................... 16 Figure 6: Illustration of the Album and Track Normalization Techniques.......... 18 Figure 7: Flow Chart of Content Delivery Workflows. ...................................... 18 Figure 8: Schematic illustration of the auditory anatomy. ................................ 23 Figure 9: Multi-Line Graph of Fletcher-Munson Loudness level Contours. ...... 25 Figure 10: Multi-Line Graph of Equal Loudness contours (2003 revision). ...... 26 Figure 11: Graphically representation as a rough stairstep pattern (depicted in red) of a sampled signal. .................................................................................. 30 Figure 12: Graphically representation as a rough stairstep pattern (depicted in red) of high frequencies sampled signal. .......................................................... 30 Figure 13: Illustration of distortion products resulting from intermodulation of a 30kHz and a 33kHz tone. ................................................................................. 36 Figure 14: Diagram illustrating the transition band width available for a 48kHz ADC/DAC (left) and a 96kHz ADC/DAC (right). ................................................ 40 Figure 15: Illustration of How Inter-Sample Peaks Occurred in the Analog Domain. ............................................................................................................ 53 5 Literature Review • Aliasing distortion: Audio distortion that occurs when a signal containing frequencies above the Nyquist frequency is improperly sampled or represented in digital form. It results in the appearance of spurious or false frequencies in the digital signal, which are not present in the original analog signal. • Audio quality: perceived fidelity, clarity, and overall fidelity of sound reproduction in an audio recording or playback. • A-weighted decibels (dBA): Unit of measurement used to quantify sound levels, taking into account the sensitivity of the human ear to different frequencies. • Bit Depth: Number of bits used to represent the amplitude of each sample in a digital audio signal. • Clipping: Distortion that occurs when an audio signal exceeds the maximum amplitude that can be represented without distortion, typically in digital audio systems. • Decibels Sound Pressure Level (dBS SPL): Unit of measurement used to quantify the intensity or loudness of a sound relative to the threshold of human hearing. • Formats: Specific encoding methods used to represent and store audio data digitally. 6 • Intermodulation distortion: Type of nonlinear distortion that occurs in audio systems when two or more different frequencies interact within the system and produce additional frequencies that were not present in the original input signals. • Loudness Target: Predetermined level or range of perceived loudness that audio content is intended to achieve during recording, mixing, mastering, or playback. It is a reference point used in audio production to ensure consistency and uniformity in the perceived loudness of audio material across different tracks, albums, or media formats. • Loudness: Subjective perception of the intensity or volume of sound. • Music streaming services: Digital platforms that enable users to access and listen to music over the internet without the need for downloading. • Normalization: Process of adjusting the amplitude of an audio signal to bring its peak or average level to a predetermined target value. • Nyquist frequency: The Nyquist frequency is defined as half of the sampling rate of a digital system. In other words, it is the highest frequency that can be accurately represented in a digital signal sampled at a given rate. • Peak (levels): Maximum instantaneous amplitude of an audio signal over a specified time period, typically measured in decibels (dB) relative to a reference level. • Sample Rate: Number of samples of audio taken per second, typically measured in Hertz (Hz). 7 • True Peak: Maximum instantaneous amplitude of an audio signal when considering inter-sample peaks, which are peaks that occur between the discrete samples of a digital audio signal. 1. Introduction 1.1 Streaming Growth in Music Distribution It is well known that music streaming services have experienced significant growth in recent years as a means of music consumption. This trend has been extensively discussed by the International Federation of the Phonographic Industry (IFPI) (2024). In 2024, as part of its annual report on the phonographic industry, the IFPI released the Global Music Report of 2023. This report sheds light on various observations regarding consumption formats and trends, as well as the markets in which they are prevalent. In 2023, the global recorded music market demonstrated robust growth, increasing by 10.2%, marking the second-highest growth rate ever recorded. With a total value of US$28.6 billion, this growth marked the ninth consecutive year of expansion. Streaming formats, particularly subscription streaming, emerged as the primary drivers of revenue growth, constituting the majority of both revenue growth and market share. Streaming revenues globally increased by 10.4% to reach US$19.3 billion, serving as the primary driver of overall global 8 growth. Streaming accounted for over two-thirds (67.3%) of the total global market, underscoring streaming platforms as the dominant method of music consumption. Figure 1: Stacked Bar Graph of Global Recorded Music Industry Revenues 1999-2023 (US$ Billions). 9 Figure 2: Pie Chart of Global Recorded Music Revenues by Segment 2023. 1.2 Audio quality and listener experience The emergence of streaming platforms as primary conduits for music dissemination has undeniably revolutionized the landscape of audio recording, mixing, post-production, and delivery workflows. However, this paradigm shift has brought to the forefront a myriad of challenges concerning audio quality and user experience. Intrinsic to the realm of sound are variations in loudness, a phenomenon that has been subject to extensive perceptual studies. These studies have elucidated that listeners generally exhibit a degree of tolerance towards moderate fluctuations in audio content. Nevertheless, there exists a discernible aversion towards 10 pronounced shifts in loudness, particularly when transitioning between unrelated audio segments, such as the interruption of a television or radio program by commercial announcements or the switch between channels or audio sources. The concept of maintaining a consistent loudness level, often termed the "comfort zone," emerges as paramount in mitigating listener discomfort stemming from abrupt changes in audio dynamics. It is imperative to recognize that the excessive employment of loud masters can precipitate audio distortion during transmission in lossy formats, thereby compromising overall audio fidelity. Figure 3: Multi-Line Graph of Listener’s Behavior After a Change in Loudness Occurred IF the Change Was Frequent. Thus, the imperative arises for the development and implementation of a loudness-based leveling solution within streaming platforms, complemented by the establishment of an appropriate loudness target. This strategic approach not only serves to preserve optimal audio quality but also engenders heightened user satisfaction by fostering a seamless and immersive listening experience across diverse audio contexts. Consequently, the integration of such measures 11 represents a pivotal step towards harmonizing audio quality and listener experience within the dynamic landscape of streaming media. 1.3 Goals In response to the aforementioned challenges, the primary objective of this document is to furnish the requisite knowledge essential for comprehending the nuances therein, thereby guiding stakeholders through a series of recommendations and guidelines. These directives are meticulously crafted with the aim of delineating and implementing an optimal Distribution Loudness framework tailored specifically for streaming and on-demand audio file playback. By adhering to these recommendations, stakeholders can elevate the overall listening experience, fostering consistency and fidelity across a spectrum of audio platforms and content categories. The overarching goals of this endeavour are delineated as follows: • Comprehensive Understanding of Streaming Platforms: Delve into the intricacies of streaming platforms to elucidate their operational mechanisms and underlying principles, thereby facilitating informed decision-making and strategic planning. • Profound Insight understanding of into Digital Audio: the fundamental Acquire a comprehensive concepts and technologies underpinning digital audio. • Establishment of Optimal Audio Formats: Identify and define the optimal format for high-fidelity audio playback. 12 • Enhancement of Listener Quality and Experience: Prioritize the enhancement of listener quality and experience by implementing measures aimed at mitigating audio artifacts, optimizing dynamic range, and ensuring consistency in loudness levels across diverse audio contexts. • Appreciation of Sample Rate and Bit Depth Dynamics: Recognize the pivotal role played by sample rate and bit depth in shaping modern production workflows, thereby enabling stakeholders to make informed decisions regarding audio quality and fidelity. • Provision of Recommendations and Best Practices: Offer a comprehensive set of recommendations and best practices aimed at empowering stakeholders to achieve exemplary standards of audio quality in their productions. 13 2. Streaming Platforms and Loudness 2.1 Loudness Standards In contemporary audio mastering, strict adherence to loudness standards is imperative. However, navigating these standards can be intricate due to the nuanced differences among various streaming platforms. Therefore, one of the primary objectives of this study is to provide a comprehensive overview of the requirements outlined by leading streaming services, elucidating their methodologies concerning format compatibility and loudness management. At the core of this investigation lies the technical document AESTD1008.1.21-9, which supersedes AESTD1004 and serves as the foundation for crafting recommendations for loudness of internet audio streaming and on-demand distribution. Through meticulous examination, the aim of this research is to elucidate the intricacies of loudness normalization processes and the methodologies embraced by platforms such as Spotify and Apple Music. To commence, it is imperative to grasp the requirements set by streaming platforms regarding the loudness and peak levels of the audio content intended for submission. To facilitate this understanding, a table containing the relevant values pertinent to this research has been compiled. 14 Figure 4: Table of Max Integrated Loudness and TP Requirements of Streaming Platforms. 2.2 The Loudness Normalization Process In the implementation of normalization, streaming services conduct an analysis of the original audio's loudness to ascertain whether it surpasses or falls short of the desired distribution loudness. Should the original audio exceed this specified threshold, normalization is employed to attenuate its level, ensuring adherence to the desired standard, a process commonly referred to as downward normalization. Conversely, if the loudness falls below the desired level, normalization introduces positive gain, termed upward normalization, which may necessitate the utilization of peak limiting techniques to prevent clipping.1 In situations where upward normalization poses a risk of clipping, the incorporation of peak limiting becomes essential to maintain the integrity of the 1 While many platforms provide the option to toggle normalization on or off, not all platforms have this feature enabled by default. 15 audio signal, albeit potentially altering its sonic characteristics. Distributors may opt for partial normalization to preserve more of the audio's dynamics, as illustrated in the accompanying figure labeled "Partially Normalized". Figure 5: Illustration of the Normalization Process. 2.3 Album and Track Normalization of Music for Distribution In the pursuit of normalization for individual tracks or entire albums, streaming platforms employ various Music Normalization Techniques. One such technique, album normalization, assumes a pivotal role in preserving the original loudness intent of a mastered album and, consequently, its artistic balance or intent. Through album normalization, streaming services adjust the loudness of all tracks within an album by applying the same gain value, thereby maintaining the relative "album loudness" of the tracks. As per the recommendations delineated in the TD1008 report by the AES, it is advisable that the album loudness value be derived from the loudest track present on the album, rather than from a comprehensive loudness measurement of the 16 entire album. This approach ensures consistency in loudness across all tracks within the album, thereby aligning with the intended sonic experience envisioned by the artist or mastering engineer. The recommendations outlined in the AES 1008 report emphasize the significant role of album normalization as a standard practice, even in situations where tracks are played outside the album context, such as in shuffle play mode. This practice stems from the deliberate mastering of quieter tracks within albums to maintain a correct relative balance with louder tracks. Consequently, the loudness of quieter tracks is intentionally attenuated. Research indicates that listeners exhibit a preference for experiencing tracks at their intended relative levels, even when encountered outside the album sequence. It is advisable for distributors to normalize the loudest track of an album to -14 LUFS, typically resulting in an integrated album loudness of -16 LUFS, considering the tendency for the loudest tracks to be approximately 2 LU louder than the album's average loudness. Adhering to this procedure ensures that the majority of quieter tracks maintain a loudness level above -20 LUFS. While Album Normalization serves as the preferred default approach, in cases where implementation proves impractical, it is recommended to normalize all tracks to 16 LUFS to maintain consistency across the album's playback experience. 17 Figure 6: Illustration of the Album and Track Normalization Techniques. Figure 6: Illustration of the Album and Track Normalization Techniques. 2.4 Normalization Workflow In the domain of distributing file-based audio to consumers, loudness normalization is often automated. Within this framework, content is categorized into two main groups: music-only or assorted, which encompasses a mixture of music, speech, interstitials, and similar elements. It is noteworthy that music-only content undergoes a distinct normalization process. Figure 7: Flow Chart of Content Delivery Workflows. 18 2.5 Formats and Codecs Presently, the predominant method employed by streaming platforms involves distributing their catalogues in lossy formats such as MP3, AAC, or OGG. These formats are favoured for their compact file size and efficient streaming capabilities. Lossy formats compress audio data, thereby reducing file size, facilitating storage and seamless music streaming, ultimately contributing to a smoother listening experience. Additionally, they boast extensive compatibility with a wide array of devices and platforms, presenting a cost-effective solution for production and distribution compared to lossless formats. However, it is noteworthy that music streaming platforms have expanded their offerings to include a variety of formats to cater to diverse user preferences and device compatibility. In the realm of digital data, compression entails the process of reducing file size by eliminating redundant or unnecessary information. Lossy and lossless audio compression represent distinct techniques for reducing the size of audio files. The primary differentiation lies in their approach to compressing audio data. Lossy Compression: Lossy compression entails the discarding of certain audio data to reduce the file size. This approach is adopted when a significant reduction in file size is required, and a certain degree of quality loss is acceptable. Algorithms utilized for lossy compression analyse the audio signal and eliminate components considered less perceptible to the human ear. By removing specific audio 19 information and simplifying the data, lossy compressed audio formats achieve further reduction in file size, albeit at the expense of audio quality. However, various technologies are employed to selectively remove portions of the sound that have minimal impact on perceived quality. These techniques also strive to minimize the introduction of audible noise during the compression process. Lossy Formats: • MP3: A lossy compression format, widely used for music streaming, with a bit rate of 128 kbps to 320 kbps. • AAC (Advanced Audio Coding): A lossy compression format, used by Apple Music, with a bit rate of 128 kbps to 320 kbps. • OGG (Ogg Vorbis): A lossy compression format, with a bit rate of 128 kbps to 320 kbps. Lossless Compression: In contrast to lossy compression, lossless compression reduces file size by eliminating extraneous metadata while preserving all audio data intact. Algorithms utilized for lossless compression condense audio data without any loss of information, resulting in a smaller file size without compromising quality. Lossless compressed audio formats effectively store data in reduced space while enabling the recreation of the original and uncompressed data from the compressed version. Unlike uncompressed audio formats, which allocate the same number of bits per unit of time to encode both sound and silence, lossless compressed 20 formats exhibit a notable distinction. Encoding a minute of absolute silence in a lossless compressed format yields a file occupying minimal space, unlike uncompressed formats where silence consumes the same space as sound. Consequently, a lossless audio file is smaller than its uncompressed counterpart. Therefore, lossless music formats aim to minimize processing time while maintaining high audio fidelity. Lossless Formats: • FLAC (Free Lossless Audio Codec): A lossless compression format, supports sample rates up to 192kHz, bit depth up to 24-bit. • ALAC (Apple Lossless Audio Codec): A lossless compression format, used by Apple Music, supports sample rates up to 192kHz, bit depth up to 24-bit. Uncompressed Formats: • WAV (Waveform Audio File Format): An uncompressed format, supports sample rates up to 192kHz, bit depth up to 32-bit. • AIFF (Audio Interchange File Format): An uncompressed format, supports sample rates up to 192kHz, bit depth up to 32-bit. It is important to note that lossless formats are not directly supported by the Bluetooth standard. Bluetooth audio codecs, such as AAC, aptX, and aptX HD, employ lossy compression techniques to reduce the bitrate of the audio signal for improved compression and transmission efficiency. Although these codecs aim to balance quality and compression ratio, they do not support lossless audio. 21 In summary, lossy compression reduces file size by discarding some audio data, while lossless compression eliminates unnecessary metadata without discarding any audio data. 3. Human Hearing 3.1 Our Ears In order to ascertain the optimal format for audio playback, it is imperative to delve into the intricacies of our auditory system, which fundamentally dictates our perception of sound. The journey of a sound wave begins with its reception by the external ear, acting as the initial gateway to our auditory experience. As these waves propagate through the ear canal, they encounter the eardrum, initiating vibrations. These mechanical vibrations set off a sequence of events within the middle ear, orchestrated by the Malleus, Incus, and Stapes, tiny yet indispensable bones that facilitate the transmission of sound energy. They modulate the sound's volume and pitch characteristics before conveying it to the cochlea, our auditory sensory organ. Within the cochlea lies the remarkable basilar membrane, akin to a resilient sheet housing specialized hair cells. These cells function as miniature sensors, each attuned to a specific range of sounds, distributed along the basilar membrane in 22 a gradient fashion. This arrangement enables us to perceive a wide spectrum of sounds, ranging from low-frequency rumbles to high-pitched tones. Figure 8: Schematic illustration of the auditory anatomy. The phenomenon of frequency selectivity demonstrated by cochlear hair cells is of paramount importance in understanding our perception of sound. This selectivity dictates that certain regions of the basilar membrane respond most vigorously to particular frequencies, creating distinct peaks of sensitivity within the auditory spectrum. Moreover, the overlapping nature of these frequency bands ensures a seamless transition between different auditory stimuli, allowing for the integration of complex sound scenes with remarkable fidelity. However, it is imperative to acknowledge the inherent limitations of our auditory apparatus. Despite its remarkable capabilities, the auditory system is bound by constraints that impact our perception of sound. One such constraint is the presence of a critical threshold beyond which sounds become imperceptible to 23 the human ear. This threshold is determined by the absence of hair cells specialized to detect specific frequencies, rendering certain sounds effectively inaudible. 3.2 Audible Spectrum Exploring the range of sounds perceptible to humans involves identifying both the highest and lowest frequencies within our auditory spectrum. This determination hinges on understanding the threshold at which sounds transition from barely audible to uncomfortably loud. In the early 1930s, Harvey Fletcher and Wilden A. Munson addressed this challenge in their seminal paper titled "Loudness, its definition, measurement, and calculation," published in 1933. Their work culminated in the creation of the influential "Fig. 4 Loudness level contours," which remain foundational in auditory research. Drawing upon the Fletcher-Munson curves and contemporary data, the consensus on human hearing range typically spans from 20 Hz to 20 kHz. This range, spanning nearly a century of empirical research, comprehensively covers the audible spectrum. It serves as a fundamental benchmark for audio reproduction, ensuring fidelity and accuracy in capturing and reproducing sounds within the human auditory range. 24 Figure 9: Multi-Line Graph of Fletcher-Munson Loudness level Contours. Recent revisions (ISO, 2003) to auditory standards have brought attention to our sensitivity to high-frequency sounds, particularly suggesting that our perception may be less acute in the upper range, specifically around 16 kHz, than previously assumed. 25 Figure 10: Multi-Line Graph of Equal Loudness contours (2003 revision). It is crucial to recognize that our perception extends beyond the conventional auditory range, even when sounds fall outside our hearing capabilities due to the absence of hair cells sensitive to their frequencies. An illustrative example of this phenomenon is evident in subsonic frequencies generated by seismic activity, such as earthquakes. These seismic waves often possess frequencies below the threshold of human hearing, typically below 20 Hz, rendering them inaudible to our ears. However, despite their subsonic nature, we can still perceive their effects through other sensory modalities. For instance, we may feel the vibrations transmitted through the ground, demonstrating the multifaceted nature of our perception beyond traditional auditory means. However, for the purpose of this study, our focus will remain exclusively on the audible spectrum of human hearing. 26 3.3 Dynamic Range Now, let's explore the range between the quietest and loudest sounds perceivable to us. One approach to defining this absolute range involves examining the curves representing the absolute threshold of hearing and the threshold of pain. The difference in decibel levels between the threshold of pain and the absolute threshold of hearing is approximately 140 decibels for a young, healthy listener. However, it is crucial to note that exposure to sound levels exceeding 130 decibels can lead to permanent hearing damage within seconds to minutes, emphasizing the critical importance of adhering to safe listening practices. Nevertheless, the expansive dynamic range of 140/130 dB is somewhat idealistic, given the practical limitations encountered in everyday environments. Such ranges are often observed in studies conducted within anechoic chambers, where ambient noise is virtually absent, allowing for the detection of even the faintest sounds, such as one's heartbeat or blood circulation. However, achieving this level of precision requires specialized equipment, meticulous calibration, and controlled conditions that diverge significantly from the typical urban environment characterized by continuous noise pollution. To provide context, a quiet broadcasting or recording studio, or a sound isolation room, typically registers a sound pressure level of around 20 dBSPL, which is approximately 28 dB louder than the quietest audible sound. For example, a study conducted by the National Institutes of Health (NIH) in 2016 investigated ambient noise level variations using an A-weighted sound level meter across four zones 27 in Vadodara city. Over a span of three hours, measurements were taken at 15minute intervals for three consecutive days, consistently at the same time. The findings revealed the highest equivalent noise level of 93.7 dBA in the commercial zone, followed by 85.5 dBA in the industrial zone, 73.2 dBA in the silence zone, and 70.2 dBA in the residential zone.2 In summary, the human auditory system perceives sound within the audible spectrum, ranging from 20 Hz to 20 kHz, with a dynamic range typically spanning between the threshold of hearing and the threshold of pain, totaling approximately 130 to 140 dB. This delineation holds paramount significance as it profoundly influences technical considerations in audio engineering, particularly regarding digital audio systems. These considerations include determining the appropriate sample rate and bit depth to accurately capture and reproduce analog sound information. Now, let us transition into the exploration of the process of encoding analog information into the digital domain, commonly referred to as sampling theory. 2 The World Health Organization (WHO) defines noise above 65 decibels (dBA) as noise pollution. 28 4. Pursuing Optimal Audio Quality for Playback 4.1 Sampling Theorem Analog waves, such as sinusoidal, represent continuous functions over time. Therefore, to sample them, we need to capture a sufficient number of unique and distinct signal values that a sinusoidal function can produce. This is achieved by sampling the signal at discrete points in time, converting it into a discrete signal. The process of converting a continuous signal into a discrete signal is governed by the Nyquist-Shannon Sampling Theorem (1928). This theorem states that any band-limited continuous-time signal can be accurately converted to and from digital signals when sampled at a rate at least twice as high as the highest frequency of the signal to be reproduced. With human auditory perception typically capped at around 20kHz, a sampling rate exceeding 40kHz becomes imperative for faithfully capturing the audible frequency spectrum. In visual representations, sampled signals often exhibit jagged stair-step patterns, which, while appearing as coarse approximations, maintain mathematical precision. Upon conversion back to analog form, these signals faithfully restore their original smooth waveform. 29 Common understanding holds that sampled signals resemble rough stair-step approximations of the continuous waveform. It is often assumed that increasing the sampling rate, along with the number of bits per sample, yields finer representations and brings digital signals closer to their analog counterparts. However, it is crucial to clarify that while signals below the Nyquist frequency are captured comprehensively through sampling alone, an infinite sampling rate is unnecessary for lossless reconstruction of the analog signal with precise timing and fidelity. Figure 11: Graphically representation as a rough stairstep pattern (depicted in red) of a sampled signal. Figure 12: Graphically representation as a rough stairstep pattern (depicted in red) of high frequencies sampled signal. 30 To correct misconceptions about depicting digital waveforms, it is crucial to recognize that representing a digital waveform as a stair-step is inaccurate. A stair-step function is continuous in time, jagged, and piecewise, with a defined value at every point in time. However, a sampled signal operates on a discrete time scale, possessing a value only at each instantaneous sample point and remaining undefined elsewhere. Consequently, there is no value assigned to the signal between sample points. Therefore, a more appropriate visualization of a discrete time signal is a lollipop graph, which accurately reflects its discrete nature and emphasizes the discrete sample points rather than attempting to interpolate values between them. This misconception may primarily stem from engineers employing zeroorder hold techniques when representing audio signals for the sake of convenience and practicality. Additionally, digital audio workstations (DAWs) often connect discrete points of a waveform graph with a straight line between points for convenience, creating a more user-friendly waveform for interpretation and manipulation. All signals with content entirely below the Nyquist frequency (half the sampling rate) are perfectly and completely captured by sampling, without the need for an infinite sampling rate. Sampling preserves frequency response and phase, allowing for lossless reconstruction of the analog signal with precise timing. 31 Real-world complications arise in audio sampling, with one of the most significant challenges being the necessity for band-limiting to prevent aliasing distortion. Signals containing frequencies above the Nyquist frequency must undergo lowpass filtering before sampling, typically achieved through an analog anti-aliasing filter. However, these filters do not immediately cut off all frequencies at the specified cutoff. They require some space to ramp down and progressively attenuate them. To accommodate this, a buffer or "guard band" between the highest frequency content we want to represent, and the Nyquist frequency is necessary. This buffer is why we need sampling rates higher than twice the Nyquist frequency, as established by the Nyquist-Shannon Sampling Theorem. For these reasons and others, sampling rates of 44.1 kHz and 48 kHz are commonly used as standards. 4.2 Aliasing Aliasing manifests within discrete signal representations, dictated by sampling parameters, wherein fewer than two data points encapsulate a single cycle of the signal. Consequently, the sampled signal forfeits the capability to faithfully reproduce frequencies beyond the Nyquist limit. Despite the loss of fidelity, the underlying data persists within the sampled signal, albeit in a distorted form. 32 Consider, for example, the introduction of distortion, a phenomenon that engenders harmonics harmonically aligned with the original signal, perceptually enhancing its timbral richness. However, should these additional harmonics surpass the Nyquist frequency threshold, they undergo folding back into the spectrum of the sampled signal, precipitating foldback distortion, commonly referred to as aliasing distortion. It is noteworthy that foldback distortion begets extraneous harmonics devoid of harmonic alignment with the original signal, thereby eliciting auditory artifacts perceived as deleterious to the overall fidelity and quality of the audio signal in the majority of instances. This divergence from harmonic coherence underscores the adverse implications of aliasing distortion, accentuating the imperative for meticulous management of sampling parameters to mitigate its deleterious effects on audio fidelity and perceptual quality. 33 4.3 High-Resolution Audio or Higher Sample Rates for Playback High-Resolution Audio encompasses a variety of digital methodologies and formats aimed at enhancing the encoding and playback of music by utilizing sampling rates and bit depths that surpass the conventional 16-bit 44.1kHz resolution commonly found in CDs, often referred to as "16/44". There is a prevailing belief among audio experts that employing higher sampling rates and bit depths during production results in audio of superior quality, ultimately enhancing the listening experience for consumers. Consequently, sampling rates such as 48kHz, 88.2kHz, 96kHz, and even 192kHz have become common standards in the industry. Additionally, the adoption of 24-bit recording has become widespread, offering increased workable headroom and a broader total dynamic range in the digital domain. However, higher sample rates do not inherently provide greater accuracy or fidelity. In fact, the only distinction between standard sample rates lies in the available bandwidth of our sampled signal. Sampling at 192 kHz undeniably results in larger file sizes, which can present challenges in terms of storage space and transmission speed. Additionally, the computational processing speed requirements associated with sampling at this rate are considerable, prompting inquiry into the rationale behind its adoption. 34 Some proponents have sought to justify the use of 192 kHz sampling by citing supposed benefits such as a narrower impulse response, suggesting enhanced spatial localization of sonic impulses or a more analog-like behaviour. However, such claims reveal a fundamental misunderstanding of signal theory principles. Bandwidth primarily concerns frequency content, while impulse response relates to the time domain. Despite their distinct definitions, they are intrinsically linked. Advocating for microsecond impulse response effectively promotes a megahertz audio system, a proposition lacking practical necessity. Human auditory capabilities do not significantly extend beyond 20 kHz, rendering frequencies above 40 kHz largely irrelevant in terms of perceptual impact. Therefore, arguments for excessively high sampling rates lack empirical support and fail to align with fundamental principles of signal theory. And remember, there is no need for higher sampling rates. There is no additional information gained from them. Moreover, instruments produce minimal sound at 96 kHz, microphones do not respond to it, speakers do not reproduce it, and the human ear cannot perceive it. In addition to the absence of benefits, operating at excessive speeds presents further disadvantages. There is an increased data size due to higher sampling rates impacting data storage and transmission speed requirements. Operating at 192 kHz significantly raises the demand for processing power, resulting in costly equipment and potential compromises in audio quality. 35 4.4 Higher Sample Rates Being Harmful for Playback The theoretical capability of 192kHz audio, extending beyond 400% of the audible limit, may initially seem promising. However, in practice, it offers negligible benefits and can even introduce subtle compromises in fidelity. Specifically, 192kHz digital music files fail to provide discernible advantages, primarily due to the inclusion of ultrasonic frequencies. These ultrasonic frequencies pose specific challenges during playback, as audio transducers and power amplifiers inherently exhibit distortion, particularly at the extremes of the frequency spectrum. When the same transducer reproduces ultrasonic frequencies alongside audible content, nonlinearities can lead to the creation of intermodulation distortion products that manifest across the entire audible spectrum. Similarly, nonlinearities in power amplifiers can yield comparable effects. While these effects may be subtle, empirical listening tests have confirmed their potential audibility, highlighting the complexities associated with high-resolution audio formats such as 192kHz. Figure 13: Illustration of distortion products resulting from intermodulation of a 30kHz and a 33kHz tone. 36 Systems not specifically engineered to reproduce ultrasonic frequencies typically exhibit elevated levels of distortion beyond the 20kHz threshold, exacerbating intermodulation distortion phenomena. Expanding a system's frequency range to accommodate ultrasonics necessitates compromises that inevitably degrade noise and distortion performance within the audible spectrum. Consequently, whether through increased distortion or compromised performance in the audible range, the unnecessary reproduction of ultrasonic content invariably diminishes overall system performance. Thus, prudent design considerations are essential to strike a balance between extending frequency response and preserving optimal performance across the audible spectrum. Monty Montgomery (2012) presents four distinct solutions to mitigate the additional distortion caused by ultrasonic content: 1 - Employing a dedicated ultrasonic-only speaker, amplifier, and crossover stage to isolate and independently reproduce ultrasonic frequencies that are inaudible to the human ear, thereby preventing interference with audible sounds. 2 - Utilizing amplifiers and transducers specifically designed for wider frequency reproduction to mitigate audible intermodulation caused by ultrasonics. However, achieving this extended frequency range entails a trade-off, potentially compromising performance in the audible portion of the spectrum. 37 3 - Implementing speakers and amplifiers meticulously engineered to avoid reproducing ultrasonic frequencies altogether, thus circumventing the issue of intermodulation distortion in the audible band. 4 - Opting not to encode such a wide frequency range from the outset. By excluding ultrasonic content from the audio encoding process, the possibility of ultrasonic intermodulation distortion within the audible band is effectively eliminated. While each solution offers a distinct approach, the fourth option appears to be the most pragmatic choice, particularly concerning the goal of creating a scalable and globally applicable solution without imposing significant resource costs. Therefore, careful consideration must be given to the trade-offs associated with handling ultrasonic frequencies, weighing the potential benefits against the costs and complexities involved. Ultimately, prioritizing improvements in the audible range is likely to yield more tangible enhancements in overall audio performance. In summary, the potential audibility of intermodulation distortion resulting from ultrasonic frequencies on a particular audio system remains subject to variability. While in some cases, the added distortion may be imperceptible, it could also be discernible in others, with the inclusion of ultrasonic content offering no discernible benefits and, in many instances, visibly degrading fidelity. Moreover, even on systems where ultrasonic content does not compromise 38 fidelity, the resources allocated to managing ultrasonics could have been redirected towards enhancing performance within the audible range. 4.5 Oversampling While sampling rates over 48kHz may appear irrelevant to high-fidelity audio data, they play an essential role in several modern digital audio techniques, with oversampling being a prime example. Unlike analog filters, digital filters have few practical limitations, allowing for greater efficiency and precision in completing the anti-aliasing process digitally. In oversampling, the raw digital signal passes through a digital anti-aliasing filter at a very high rate, ensuring effective removal of high-frequency content beyond the Nyquist frequency. This process allows for the use of low-rate audio (e.g., 44.1 kHz or 48 kHz) with all the fidelity benefits of higher sampling rates, such as smooth frequency response and low aliasing, without the drawbacks of ultrasonics that can cause intermodulation distortion and wasted space. Moreover, it even enables us to mitigate the negative effects of linear filters like pre-ringing by shifting them into inaudible content, ensuring a more accurate representation of the audio signal in the audible range. The ubiquitous integration of oversampling at remarkably high rates within contemporary analog-to-digital converters (ADCs) and digital-to-analog converters (DACs) underscores its pervasive adoption and critical role in attaining superior audio reproduction quality. 39 However, oversampling, and oversampling techniques come with drawbacks. They can introduce latency and increase CPU usage, which can be determinants to consider, especially in live scenarios. Figure 14: Diagram illustrating the transition band width available for a 48kHz ADC/DAC (left) and a 96kHz ADC/DAC (right). 4.6 Bit Depth In digital audio, the concept of bit depth refers to the number of bits allocated to represent each sample. Essentially, each bit can represent one of two possible values, typically interpreted as 'on' or 'off.' Therefore, higher bit depths permit a finer resolution in representing the amplitude of the signal. Upon conversion of a digital signal back to analog form, regardless of the bit depth employed during digitization, the resultant waveform appears smooth. This phenomenon underscores the efficacy of the digital-to-analog conversion process in faithfully reconstructing the original analog signal. However, the journey from analog to digital introduces a crucial step: quantization. Quantization occurs during digitization and introduces a level of imprecision, known as quantization noise. The magnitude of this noise is 40 contingent upon the bit depth utilized during digitization. Increasing the bit depth serves to mitigate this noise, thereby lowering the noise floor and augmenting the dynamic range of the digital signal. A fundamental aspect of quantization involves selecting the digital amplitude value nearest to the original analog amplitude. While this simplistic approach suffices in theory, its practical application often leads to undesirable outcomes. The resulting noise from this rudimentary quantization process varies depending on the input signal, leading to inconsistent noise characteristics or, worse, distortion. To address the inherent noise introduced during quantization, techniques such as dithering are employed. Dithering involves the addition of low-level noise to the signal, effectively masking quantization artifacts and improving the accuracy of the digital-to-analog conversion process. Dithering is a technique employed to mitigate the adverse effects of quantization noise in digital audio. By adding carefully crafted low-level noise to the signal, dithering effectively masks the undesirable artifacts associated with simple quantization. Moreover, dithering provides producers with a degree of control over the perceptual qualities of the introduced noise, allowing them to shape the resultant sound according to their preferences. Properly executed dithering serves to shape the quantization noise spectrum, redistributing it to frequencies that are less perceptible to the human ear. This 41 process results in a pristine, artifact-free audio file conducive to optimal playback fidelity. For instance, a standard bit depth of 16 bits yields a theoretical dynamic range of approximately 96 dB SPL (Sound Pressure Level), with each additional bit contributing an incremental 6 dB to the dynamic range. However, through the judicious application of dithering techniques, the effective dynamic range achievable with 16-bit depth can extend to approximately 120 dB SPL. In practical music playback scenarios, the dynamic range of the audio is typically fixed, determined by the bit depth and quality of the recording. Thus, optimizing the bit depth and employing appropriate dithering techniques are critical considerations in ensuring optimal audio fidelity and dynamic range in digital audio production and playback. In music production scenarios, the manipulation of gain and dynamic range is a common practice aimed at achieving desired sonic characteristics. Here, the adoption of a 24-bit depth offers distinct advantages, primarily due to its ability to provide a lower noise floor and expanded dynamic range, thus affording greater headroom for signal processing. The prevalence of modern workflows emphasizes the importance of these considerations. Contemporary production techniques facilitate extensive signal processing, wherein numerous operations are applied to individual signals or tracks within a composition. However, the accumulation of noise inherent in these 42 processes can pose challenges, particularly when mixing multiple tracks into a cohesive final product. It is also worth mentioning that increasing the bit depth of the audio representation from 16 to 24 bits does not increase the perceptible resolution or 'fineness' of the audio. It only increases the dynamic range, the range between the softest possible and the loudest possible sound, by lowering the noise floor. However, a 16-bit noise floor is already below what we can hear. In summary, 16-bit depth is and will always be sufficient to capture the dynamic range of human hearing. Therefore, the dynamic range should be increased based on the number of processes occurring before and after delivery, as well as the number of channels involved. 4.7 How to Actually Improve Audio Quality for the End User One of the most straightforward ways to enhance the quality of audio for end users is to upgrade the equipment they use to reproduce music content. By doing so, we improve the fidelity capabilities of an equipment to fullfully represent the fidelity of a signal. Another critical aspect that could significantly enhance end-user audio quality is the widespread implementation of lossless audio formats as standards in music streaming platforms. Lossless formats ensure that audio fidelity remains intact, avoiding any potential degradation. Despite some companies already offering 43 lossless formats for music consumption, it is not yet the predominant format. This issue is primarily due to technical limitations, notably the insufficient bandwidth to deliver lossless formats globally while maintaining transmission speed and costeffectiveness for streaming platforms. Finally, the most effective yet challenging approach to enhancing audio quality remains through the production of meticulously crafted mixes and masters. While seemingly straightforward, this endeavour requires a deep understanding of the intricacies of music production and mastering. Mastery of this craft comes with experience and a commitment to continuous learning. It is essential to first grasp the theoretical underpinnings as a technician before transcending boundaries as an artist, guided by a clear artistic vision. 5. Modern Production Workflows 5.1 Production Frameworks In light of the preceding discussion, it is pertinent to examine how this knowledge may be applied to refine contemporary production workflows. It is important to note that while certain engineers, such as Lavry (2004) and Bob Stuart (1997), propose that the optimal sample rates are 52 kHz, 58 kHz, or even 60 kHz, these will not be considered for the purpose of this study as they do not conform to industry standards. Although the validity of their conclusions regarding optimal sample rates is not disputed, the focus of this research is on solutions that are reproducible and widely adaptable without compatibility issues or the need for non-standard systems. Unless there is a shift in industry paradigm 44 recognizing the advantages of such sample rates and deeming them optimal, this study will adhere to the sample rates traditionally recognized as standards within the industry. 5.1.1 Recording Sample Rates and Bit Depths In the domain of audio recording, a bit depth of 24 is recommended for various reasons. Primarily, a 16-bit depth, although operational, provides limited flexibility in managing the signal relative to the noise floor. By contrast, a 24-bit depth offers an expansive dynamic range that not only maintains a considerable distance from the noise floor at the lower spectrum but also provides a broad safety margin at the upper end to prevent potential signal clipping. It is crucial to recognize that signal clipping leads to distortion, which is generally irreversible. While de-clipping processes exist, they are intended for use in exceptional circumstances and are not reliable for routine dependence. Therefore, establishing a robust signal level from the beginning is advisable. Additionally, a 32-bit depth may be advantageous in particular contexts where the audio material exhibits extensive dynamic variations. Nevertheless, the selection of bit depth must take into account the number of tracks being recorded, as higher bit depths yield larger file sizes. These considerations are vital in the decisionmaking process tailored to each specific recording scenario. Such meticulous deliberation ensures the achievement of optimal recording quality while efficiently managing data and storage demands. The dynamic range should be expanded based on the number of processes to be applied subsequently. 45 On the other hand, at this same stage, the sample rate will be determined by the type of instrument being recorded, the number of tracks, the duration of the recording, and its intended use afterward. All of this is to ensure maintaining a good cost-benefit balance. 5.1.2 Production Sample Rates and Bit Depths Modern production and mixing sessions frequently incorporate an array of plugins, necessitating the careful selection of appropriate sample rates and bit depths to meet technical goals. In the realm of production, most current Digital Audio Workstations (DAWs) operate internally at a 32-bit floating point bit depth or even higher, as exemplified by Reaper with his 64-bit floating point of internal processing. However, if this is not the case, a minimum of 24 bits is recommended for modern production involving a substantial number of tracks and processes. While 32-bit remains a viable option, the cost-benefit ratio should be carefully considered given the requirements of the project. When determining the suitable sample rate for production, various factors must be considered. Primarily, the intended final destination of the production dictates the sample rate choice; for example, if the output is destined for a CD, adherence to the standard 44.1 kHz, commonly denoted as '1644', is essential. Conversely, if alternative formats or mediums are intended, such as higher resolution audio or specialized content, a judicious selection of sample rates is imperative to prevent inefficiencies in content utilization and storage capacity. 46 Non-linear processing, such as distortion and heavy compression, may introduce undesirable aliasing distortion. To mitigate this issue, two feasible approaches are available. Firstly, one may opt to operate the project at a higher sample rate. Alternatively, a more efficient strategy involves maintaining the project at a maximum of 48kHz while selectively applying oversampling to plugins featuring oversampling capabilities or utilizing tools like the 'oversampling chain' available in Reaper. This method is generally more economical in terms of CPU usage compared to running the entire session at a higher sample rate. Moreover, it permits the targeted application of oversampling where necessary, allowing for subjective evaluation of whether the aliasing distortion detracts from or enhances the sound quality, potentially adding interest or sonic appeal. It is important to note that aliasing distortion does not inherently signify poor sound quality. When selecting a sample rate for production, it is imperative to consider the nature of the material being used. Many contemporary producers utilize samples from libraries, each with its own assigned sample rate and bit depth. It is crucial to maintain consistency within a session by avoiding the use of samples with disparate sample rates. Additionally, certain Digital Audio Workstations (DAWs) may require sample rate conversion (SRC) to align samples with the project's sample rate. Some DAWs perform SRC 'on the fly' during playback, necessitating examination of how various sample rates are managed within a session and the implications for exporting and CPU usage. However, if the need arises to integrate a sample with a different sample rate into a session (e.g., using a 44.1 kHz sample in a 48 kHz session), it is essential to 47 recognize that the upsampling process does not introduce additional highfrequency content that was not present in the original recording. Nonetheless, there may be benefits derived from the slight increase in bandwidth, particularly when employing equalizers, distortions, or other non-linear processing techniques. Another area where the consideration of different sample rates becomes pertinent is in sound design. In this capacity, the utilization of higher sample rates can offer distinct advantages for sound design techniques. Extreme pitch shifting or stretching is a common practice within the realm of sound design. Operating at higher sample rates provides increased bandwidth, which can be leveraged for these techniques. Furthermore, when manipulating bandlimited signals through pitch shifting, higher sample rates facilitate better preservation of high frequencies and mitigate the risk of anti-aliasing distortion. This preservation is attributed to the expanded frequency spectrum available at higher sample rates, allowing for greater latitude in shifting frequencies without encountering the antialiasing filter. Conversely, some individuals contend that recordings made at higher sample rates, such as 96kHz, enable the retention of more high-frequency energy when subjected to extreme stretching. This assertion suggests that higher sample rates offer advantages in retaining the integrity of high-frequency content during intensive manipulation of the signal. 48 In summary, the recommended sample rate for general production is 48 kHz. This sample rate offers a slightly increased headroom for anti-aliasing filtering, effectively minimizing undesirable artifacts within the audible spectrum, such as pre-ringing, smearing, or anti-aliasing distortions. Importantly, this advantage is achieved without imposing the heightened CPU demands or larger file sizes associated with higher sample rates like 96 kHz or 88.1 kHz, thereby optimizing the cost-benefit ratio. However, in the event of opting for a sample rate of 96 kHz, it is crucial to ensure the linearity of the system to prevent artifacts stemming from the intermodulation of ultrasonic content. Therefore, vigilance in system calibration and assessment is indispensable to uphold fidelity and integrity throughout the production process. One method to assess system linearity is to generate a pure tone at an ultrasonic frequency, such as 33 kHz, configure the audio interface to the desired sample rate (96 kHz), and listen for any artifacts. The absence of audible artifacts indicates system linearity, while the presence of noise or artifacts signifies non-linearity, potentially causing intermodulation distortion. 5.2 Loudness Another critical aspect to consider in music production is loudness. While normalization serves as a preemptive measure to enhance user experience and mitigate the need for individual distribution platforms to establish their own target levels, it is important to recognize that these volume standards are subject to change. Consequently, a song may become 'outdated' if it does not align with 49 contemporary loudness goals. In such instances, remastering the song may be necessary to ensure conformity with current norms. While established artists within the industry often possess the resources to undertake such endeavors, independent artists may encounter significant challenges in addressing this issue. Therefore, it is advisable to mix and master a song to achieve its optimal cleanest loudness, while also remaining true to any predetermined artistic intentions. By prioritizing clean and balanced loudness during the mixing and mastering process, artists can mitigate the need for frequent remastering and ensure their music remains relevant within evolving industry standards. To establish a mix anchor and framework for effective loudness management, it is essential to identify the key elements within the mix that will serve as the loudest components of the song. Typically, these elements include the kick and snare (or their equivalents), as they constitute the loudest rhythmic impulses and play a pivotal role in driving the rhythm and groove of the song, rendering it danceable. Consequently, the aim is to ensure that the master limiter is primarily triggered by these loudest rhythmic impulses, as exceeding this threshold may lead to undesirable distortion. By prioritizing the management of these key elements within the mix, producers can maintain control over the overall loudness while preserving the integrity of the audio signal. This approach facilitates the achievement of optimal loudness 50 levels without compromising sound quality, ensuring a polished and professional end result. However, it is undeniable that certain music genres inherently exhibit loudness, either as an intentional artistic choice to complement or amplify the message of a song, or perhaps due to a lack of understanding by the artists who pioneered the genre, resulting in distorted or excessively loud compositions from its inception. This latter scenario may be exemplified in certain subgenres of dubstep. Whether driven by artistic intent or an adherence to specific loudness targets, the following advice applies to producers navigating these challenges. In such productions, there is often an unconscious tendency to exploit the 32floating point internal processing capability of Digital Audio Workstations (DAWs) to push all tracks into the red before limiting them at the master limiter. To achieve cleaner mixes in loud genres, professionals like Luca Pretolesi utilize techniques such as multi-staging limiting or clipping. This method involves clipping or limiting the mixes by small increments at different stages, providing greater control over the signal and maximizing headroom at each stage while mitigating the distortion caused by excessive signal pushing. By adopting this approach, producers and engineers gain enhanced control over the signal, allowing for the application of oversampling techniques at various points in the mix and preventing distortion from occurring solely at the master 51 limiter. Moreover, structuring the session in this manner facilitates consistent gain staging, optimizing the use of emulation plugins, which often have a 'sweet spot.' Furthermore, the use of Voltage-Controlled Amplifiers (VCAs) enables precise control over the overall limited signal of the entire project or facilitates the attainment of different loudness targets. In contemporary music production, True Peak (TP) values are influenced by technological limitations, particularly the bandwidth constraints inherent in streaming platforms, which often require the use of lossy formats. It is prudent to establish a TP level that aligns with the content and loudness of the song, ensuring it never exceeds 0 dB to avoid distortion, unless intentional distortion is desired and understood. In accordance with the AESTD1008 standard, a -1 dBTP is recommended to prevent distortion during conversion by lossy codecs. Additionally, a minimum default value of -0.3 dBTP is advisable, as exceeding this threshold may lead to distortion on approximately 80% of consumer-grade playback systems. Moreover, establishing a consistent TP level helps prevent the need for individual target levels on each distribution platform. In such cases, tools such as Izotope's codec preview or Sonnox's equivalent can be invaluable for assessing how the file will translate to consumer playback environments, providing insight into potential distortion issues and allowing for adjustments as necessary. 52 Figure 15: Illustration of How Inter-Sample Peaks Occurred in the Analog Domain. 5.3 Good Practices To ensure consistent audio quality regardless of the artistic or production approach employed, adhering to a set of good practices is essential. Firstly, it is advisable to deliver uncompressed or lossless audio files, bearing in mind any specific requirements of the streaming service utilized. APPLE also advises against upsampling files beyond their original format, as this process does not enhance or restore information within the audio file. Additionally, refrain from "bit-padding" or converting 16-bit files into 24-bit format, as this does not contribute to improved audio quality. 53 While some mastering engineers may choose to manage the Sample Rate Conversion (SRC) process independently, it is recommended to deliver the highest native sample rate available. Maintaining access to the highest-resolution masters within streaming service systems ensures preparedness to leverage future technological advancements and enhancements in music quality. Utilizing True Peak Level meters is essential for accurately monitoring audio levels, as they indicate the maximum (positive or negative) value of the signal waveform in the continuous time domain (analog domain). Unlike quasi-peak meters or sample-peak meters, which may miss true peaks lying between samples, True Peak Level meters ensure that all peaks are detected. By employing an oversampling meter compliant with BS.1770 standards, true peaks can be accurately detected, allowing for precise monitoring of peak levels. This enables producers and engineers to assess whether any peaks exceed the desired true peak value, providing valuable insight into potential clipping or distortion issues. When considering the use of multistage limiting/clipping processes, it is important to carefully evaluate the implications of oversampling. While oversampling may be beneficial for limiters to ensure true peak limitation of the signal, it can have counterproductive effects when applied to clippers. Oversampling in limiters can help mitigate foldback aliasing in the audible range caused by additional harmonics introduced during nonlinear processing. 54 However, oversampled clippers may produce peaks above their designated ceiling, contradicting the intended goal of the clipping process. Therefore, it is crucial to exercise caution and discernment when deciding whether to implement oversampling in multistage limiting/clipping processes. Evaluating the specific requirements and objectives of each stage of the audio processing chain will facilitate informed decision-making and optimal results. In addition to audio quality improvements, metadata plays a crucial role in enhancing the user experience of audio content. Metadata refers to encoded information that provides descriptions of digital content streams and media files, encompassing audio, video, and photographic images. Within the realm of audio, there exist multiple types of metadata serving different purposes. Audio-related metadata serves to describe the content itself, providing essential information about the audio file. On the other hand, control-related metadata specifies how the audio should be played on a device, ensuring optimal playback settings. Examples of control-related metadata include dynamic range control (DRC) and loudness metadata, which enable compliant players to manage loudness levels for various playback scenarios. By leveraging metadata effectively, content creators can not only enhance the organization and accessibility of their audio files but also optimize the playback experience for users across different devices and environments. 55 Loudness metadata serves a crucial role in audio playback by enabling compliant players to normalize content loudness before it reaches the user's volume control. This normalization process grants playback devices control over loudness levels, ensuring a consistent listening experience across various content sources and genres. Dynamic Range Control metadata, on the other hand, is integrated into the content during encoding without altering the original audio. It allows playback devices to optionally adjust the dynamic range of the content based on factors such as the capabilities of the playback device, ambient noise levels, or listener preferences. For instance, listeners in noisy environments can opt for a limited dynamic range to improve audibility, while those in quiet settings may prefer to experience the content's full dynamic range for enhanced immersion. 6. Conclusion Navigating the uncertainties and technical intricacies inherent in audio production necessitates a thoughtful approach, considering a multitude of actions and factors. With this in mind, we will succinctly outline meticulously crafted directives aimed at defining and executing an optimal high-fidelity audio file playback and user experience. Firstly, determining the sample rate and bit depth of our projects is pivotal. The recording stage is shaped by the content to be recorded, the intended use of the material (such as CD recording), and the number of signals to be captured. 56 As an initial recommendation or starting point, a sample rate of 48 kHz and a bit depth of 24 bits are advised. However, for materials intended for extreme pitch shifting or stretching techniques, a sample rate of 96 kHz may be preferable, particularly if the number of signals is manageable or if the recording duration is not extensive. For production, mixing, and mastering, it is recommended to use a sample rate of 48 kHz and a bit depth of at least 24 bits (though your DAW likely operates on 32-bit floating point). However, it is advisable to employ oversampling for nonlinear processes like distortion, limiting, or heavy compression. Additionally, always try to work with samples at the same sample rate as your session, unless you are well acquainted with how your DAW handles these files. To proceed, another relevant aspect is loudness. In this regard, it is recommended to master each track to the cleanest loudness possible, rather than strictly adhering to a target level of -14 LUFS. Loudness should be adjusted for each song based on its composition, structure, and intent. Some songs may be above the target level, while others may be below, but they should all be normalized. However, since these target levels may change over time, avoid adjusting to a temporary value. It is important to clarify that at no point are we encouraging participation in the loudness war. Another important aspect is true peak. Current true peak values are due to technical limitations. With a true peak higher than -1 dBTP, you will experience distortion due to conversion in lossy formats. However, over time and with the 57 implementation of lossless as a standard, this headroom can be increased to 0.3 dBTP (a critical point where if exceeded, 80% of home equipment will distort the signal). If your material can adhere to the current standard of -1 dBTP, follow it. On the other hand, if your music sounds great at high volumes, at least follow the recommendation of -0.3 dBTP and use tools such as Izotope's codec preview or Sonnox's equivalent for assessing how the file will translate to consumer playback environments. Once you have your master perfectly mastered and follow all the advice to ensure audio quality, you will need to consider the following. Send uncompressed or lossless formats to the highest-quality streaming platforms without upsampling, bit-padding, or SRC. However, do not forget proper dithering with noise shaping if necessary. Add relevant loudness and DRC metadata to enhance the user experience. Since you are adding metadata, include information like the ISRC that allows you to later collect royalties from your performing rights organization if your music is played on the radio. Finally, it is important to highlight the following: the increase in quality provided by using an optimal sample rate by itself, without the synergy of encompassing processes, does not have a significant audible impact. However, as engineers striving for better sounding results, it's important to understand all the nuances of the signal we work with. Perhaps on its own, the choice of sample rate does not have a significant impact, but through all the stages, process synergies, and accumulation of errors along the way, they add up and result in a larger error that is significant. And, finally, I would like to finish with a reflection 58 from Bob Katz (2007): 'analog or digital, keep it simple or you're just compounding various forms of errors the more you do!' 59 7. Bibliography • AES, 2015. Technical Document AES TD1004.1.15-10. New York: AES [viewed 22 01 24]. Available from: https://www.aes.org/technical/documents/AESTD1004_1_15_10.pdf. • AES, 2021. Technical Document AESTD1008.1.21-9. New York: AES [viewed 28 01 24]. Available from: https://www.aes.org/technical/documentDownloads.cfm?docID=731. • AMAZON, 2024. Normalizing the Loudness of Audio Content. In: Amazon’s development documentation. 8 February 2024 [viewed 05 03 24]. Available from: https://developer.amazon.com/esES/docs/alexa/flashbriefing/normalizing-the-loudness-of-audiocontent.html. • APPLE, 2021. Apple Digital Masters: Music as the Artist and Sound Engineer Intended. [viewed 18 02 24]. Available from: https://www.apple.com/apple-music/apple-digital-masters/. • APPLE, 2024. About lossless audio in Apple Music. 7 March 2024 [viewed 21 01 24]. Available from: https://support.apple.com/enus/118295. • APPLE, n.d. Music Audio Source Profile. [viewed 18 01 24]. Available from: https://help.apple.com/itc/videoaudioassetguide/#/itc5a739206b. • Audio Masterclass, 2023. 24 bits or 96 kHz? Which makes most difference? [viewed 11 02 24]. Available from: https://www.youtube.com/watch?v=UBKJCx6UJoM. • AUDIO UNIVERSITY, 2023. Debunking the Digital Audio Myth: The Truth About the 'Stair-Step' Effect [viewed 05 01 24]. Available from: 60 https://www.youtube.com/watch?v=cD7YFUYLpDc&list=PLASEfdYtiDp0iEkeq80u0QgoFrteGFhU&index=5. • AUDIO UNIVERSITY, 2023. Why Higher Bit Depth and Sample Rates Matter in Music Production [viewed 05 01 24]. Available from: https://www.youtube.com/watch?v=VSm_7q3Ol04&. • AUDIOPHILES, 2022. What is a True Peak Limiter?. In: Audiophiles’ Blog. 1 August 2022 [viewed 10 03 24]. Available from: https://audiophiles.co/true-peak-limiter/. • BAPHOMETRIX, n.d. The Clip-To-Zero Production Strategy Gainstaging and Mixing by Loudness, not by Peak. [viewed 03 01 24] Available from: https://docs.google.com/document/d/1Ogxa5X_QdbtfLLQ_2mDEgPgHxNRLebQ7pps3rXewPM/edit#heading=h.hamz 1ram1y35. • BERG, R.E., n.d. The ear as spectrum analyzer. In: Britannica’s Blog. n.d. [viewed 13 01 24]. Available from: https://www.britannica.com/science/sound-physics/The-ear-as-spectrumanalyzer. • BRODKEY, F.D. and MADISON, WI., 2022. Hearing and the cochlea. In: MedilinePlus’ Blog. 21 July 2022 [viewed 18 01 24]. Available from: https://medlineplus.gov/ency/anatomyvideos/000063.htm. • BUTTERWORTH, B., 2020. What You Really Need to Know About Bluetooth Audio. 15 January 2020 [viewed 15 03 24]. Available from: https://www.nytimes.com/wirecutter/blog/what-you-need-to-know-aboutbluetooth-audio/. 61 • CRAVE DSP, 2017. Linear Phase EQ Explained. In: Crave DSP’s Blog. 30 May 2017 [viewed 10 03 24] Available from: https://cravedsp.com/blog/linear-phase-eq-explained. • EARTHQUAKE HAZARDS PROGRAM, n.d. Cool Earthquake Facts. In: USGS’ Blog [viewed 06 03 24]. Available from: https://www.usgs.gov/programs/earthquake-hazards/cool-earthquakefacts. • FABFILTER, 2020. Samplerates: the higher the better, right? [viewed 23 01 24]. Available from: https://www.youtube.com/watch?v=-jCwIsT0X8M. • FABFILTER, n.d. Oversampling. In: FabFilter Pro-L 2 online help. [viewed 24 02 24]. Available from: https://www.fabfilter.com/help/prol/using/oversampling. • FL STUDIO, 2013. D/A and A/D | Digital Show and Tell (Monty Montgomery @ xiph.org) [viewed 07 01 24]. Available from: https://www.youtube.com/watch?v=cIQ9IXSUzuM&. • FLETCHER, H. and Munson W.A., 1933. Loudness, Its Definition, Measurement and Calculation. Bell Telephone Laboratories. 28 August 1933 [viewed 15 01 24]. Available from: https://www.audiosciencereview.com/forum/index.php?attachments/loud ness-its-definition-measurement-and-calculation-fletcher-and-munsonpdf.86762/. • FLUX. 2022., How to Use a Limiter, Part 1 – True Peak limiting and Loudness processing. In: FLUX’s Blog. 22 November 2022 [viewed 01 03 24]. Available from: https://www.flux.audio/2022/11/22/how-to-use-alimiter-part-1-true-peak-limiting-and-loudness-processing/. 62 • FLUX. 2022., How to Use a Limiter, Part 2 – Limiter Theory – Knowing your tools. In: FLUX’s Blog. 23 November 2022 [viewed 01 03 24]. Available from: https://www.flux.audio/2022/11/23/how-to-use-a-limiterpart-2-limiter-theory-knowing-your-tools/. • HAHN, M., 2022. Lossless Audio: What It Is and How to Listen To It. In: Landr’s Blog. 29 November 2022 [viewed 13 03 24]. Available from: https://blog.landr.com/lossless-audio-streaming/. • HEARING LINK, 2024. How the ear works. In: Hearing Link’s Blog. [viewed 16 01 24] Available from: https://www.hearinglink.org/yourhearing/about-hearing/how-the-ear-works/. • HEARINGLINK, 2024. How the ear works. In: HearingLink’s Blog. April 2024 [viewed 18 01 24]. Available from: https://www.hearinglink.org/yourhearing/about-hearing/how-the-ear-works/. • HUSSAIN ATHER, S., 2020. How Does a Digital to Analog Converter Work?. In: Sciencing’s Blog. 27 December 2020 [viewed 03 03 24]. Available from: https://sciencing.com/analog-digital-converter-work4968312.html. • IBERDROLA, n.d. Noise pollution: how to reduce the impact of an invisible threat?. In: R&D Iberdrola’s Blog. [viewed 08 03 24]. Available from: https://www.iberdrola.com/sustainability/what-is-noise-pollutioncauses-effects-solutions. • IFPI, 2024. GLOBAL MUSIC REPORT 2024. Available from: https://ifpiwebsite-cms.s3.eu-west2.amazonaws.com/GMR_2023_State_of_the_Industry_ee2ea600e2.pdf. 63 • ITU, 2023. Reco01endation ITU-R BS.1770-5. Geneva: ITU [viewed 28 01 24]. Available from: https://www.itu.int/dms_pubrec/itu-r/rec/bs/RREC-BS.1770-5-202311-I!!PDF-E.pdf. • Izotope, Inc., 2020. Loudness in Mastering | Are You Listening? | S2 Ep5 [viewed 06 01 24]. Available from: https://www.youtube.com/watch?v=irgdcYD5hFE. • KAGAN, A., 2021. Should I Be Oversampling?. In: Sonarworks’ Blog. 22 february 2021 [viewed 25 02 24]. Available from: https://www.sonarworks.com/blog/learn/should-i-be-oversampling. • KATZ, B., 2007. Mastering audio : the art and the science. New York: Focal Press. • KATZ, B., 2013. ITunes music : mastering high resolution audio delivery : produce great sounding music with Mastered for iTunes. New York and London: Focal Press. • KURAKATA, K., 2016. Hearing threshold for pure tones above 20 kHz. Tokyo: Waseda University [viewed 03 03 24]. Available from: https://www.researchgate.net/publication/245524994_Hearing_threshold _for_pure_tones_above_20_kHz. • LAVRY, D., 2004 Sampling Theory For Digital Audio. Lavry Engineering [viewed 07 01 24]. Available from: https://lavryengineering.com/pdfs/lavry-sampling-theory.pdf. • LAVRY, D., 2012. The Optimal Sample Rate for Quality Audio. Lavry Engineering. 3 May 2012 [viewed 14 01 24]. Available from: https://www.lavryengineering.com/pdfs/lavry-white-paperthe_optimal_sample_rate_for_quality_audio.pdf. 64 • MACDONALD, D., 2021. Sample Rate, Bit Depth, Bit Rate, and You(r Ears), Explained [viewed 15 02 24]. Available from: https://www.youtube.com/watch?v=--VRdiFb0rk. • MASTERING THE MIX, 2020. Mastering Audio for Soundcloud, iTunes, Spotify, Amazon Music and Youtube. In: Mastering the Mix’s Blog. 26 May 2020 [viewed 08 02 24]. Available from: https://www.masteringthemix.com/blogs/learn/76296773-masteringaudio-for-soundcloud-itunes-spotify-and-youtube. • MATHIAS, K., n.d. Lossy vs Lossless Audio [Apple Music vs Spotify For Sound Quality]. In: Audio University’s Blog. [viewed 23 02 24]. Available from: https://audiouniversityonline.com/lossy-vs-lossless-audio/. • MELCHIOR, V.R., 2019. High-Resolution Audio: A History and Perspective. In: Journal of the Audio Engineering Society. My 2019 [viewed 12 02 24]. Available from: https://www.researchgate.net/publication/332954874_HighResolution_Audio_A_History_and_Perspective. • METERPLUGS, 2019. Loudness Penalty Now Supports Amazon Music, Deezer. In: MeterPlugs’ Blog. 15 October 2019 [viewed 15 02 24]. Available from: https://www.meterplugs.com/blog/2019/10/15/loudnesspenalty-amazon-deezer.html. • MIDLANDS, 2023. From a Hum to a Screech: The Bounds of Human Hearing. In: Hearing Centre Midlands’ Blog. 10 November 2023 [viewed 18 02 24]. Available from: https://www.hearingcentremidlands.co.uk/from-a-hum-to-a-screech-thebounds-of-human-hearing/. 65 • MIDLANDS, 2023. The Auditory Spectrum: Understanding the Human Hearing Range. In: Hearing Centre Midlands’ Blog. 10 November 2023 [viewed 18 01 24]. Available from: https://www.hearingcentremidlands.co.uk/the-auditory-spectrumunderstanding-the-human-hearing-range/. • MIDLANDS, 2023. The Auditory Spectrum: Understanding the Human Hearing Range. In: Hearing Centre Midlands’ Blog. 10 November 2023 [viewed 17 01 24]. Available from: https://www.hearingcentremidlands.co.uk/the-auditory-spectrumunderstanding-the-human-hearing-range/. • MOLLER, H. and PEDERSEN, C.S., 2004.. Hearing at Low and Infrasonic Frequencies. Noise and Health. Volume 6. April - June 2004 pp.37-57. Available from: https://journals.lww.com/nohe/fulltext/2004/06230/hearing_at_low_and_i nfrasonic_frequencies.5.aspx. • MONTGOMERY, M., 2012. 24/192 Music Downloads...and why they make no sense. In: Xiph’s Blog. 1 March 2012 [viewed 08 01 24]. Available from: https://people.xiph.org/~xiphmont/demo/neil-young.html. • MURTHY, A., 2020. 10. Pulse Code Modulation - Digital Audio Fundamentals [viewed 11 02 24] Available from: https://www.youtube.com/watch?v=wn71QBApCRg. • MURTHY, A., 2020. 2. Sampling Theorem - Digital Audio Fundamentals [viewed 20 01 24] Available from: https://www.youtube.com/watch?v=vrXGaFV1AmE&. 66 • MURTHY, A., 2020. 3. Common Audio Sample Rates - Digital Audio Fundamentals [viewed 20 01 24] Available from: https://www.youtube.com/watch?v=Z0EMObqS90U&. • MURTHY, A., 2020. 4. Understanding Aliasing - Digital Audio Fundamentals [viewed 21 02 24] Available from: https://www.youtube.com/watch?v=91PKZllbgds&list=PLbqhANKGP6B6V_AiS-jbvSzdd7nbwwCw&index=4. • MURTHY, A., 2020. 5. Quantization - Digital Audio Fundamentals [viewed 17 02 24] Available from: https://www.youtube.com/watch?v=1KBLguIXL30. • MURTHY, A., 2020. 6. Bit Depth - Digital Audio Fundamentals [viewed 16 02 24] Available from: https://www.youtube.com/watch?v=X4JEMCQMwOM&list=PLbqhANKGP6B6V_AiS-jbvSzdd7nbwwCw&index=6. • MURTHY, A., 2020. 7. Dithering Explained - Digital Audio Fundamentals [viewed 16 02 24] Available from: https://www.youtube.com/watch?v=48DAvO7j3zQ. • MURTHY, A., 2020. 8. Dither Types - Digital Audio Fundamentals [viewed 16 02 24] Available from: https://www.youtube.com/watch?v=t1X6DI-9_eU. • MURTHY, A., 2021. 12. Containers and File Formats - Digital Audio Fundamentals [viewed 15 02 24] Available from: https://www.youtube.com/watch?v=mfb7tuUTiZ8. 67 • MURTHY, A., 2021. 9. Noise Shaping - Digital Audio Fundamentals [viewed 11 02 24] Available from: https://www.youtube.com/watch?v=1cMae5i1Eec. • MURTHY, A., 2023. 9. Understanding Linear Phase - Digital Filter Basics [viewed 09 02 24] Available from: https://www.youtube.com/watch?v=zCdV9IUCSy8. NIH, 2015. How Do We Hear?. In: NIH’s Blog. May 2015 [viewed 03 02 24]. Available from: https://www.nidcd.nih.gov/health/how-do-we-hear. • NIH, 2022. Journey of Sound to the Brain. In: NIH’s Blog. 9 May 2022 [viewed 18 01 24]. Available from: https://www.nidcd.nih.gov/news/multimedia/journey-of-sound-video. • NIKOLIC, J., 2019. Loudness Standards – Full Comparison Table (music, film, podcast). In: Youlean’s Blog. 30 June 2019 [viewed 25 01 24]. Available from: https://youlean.co/loudness-standards-fullcomparison-table/. • OPERATING EUROVISION AND EURORADIO, 2016. EBU Tech 3344. Geneva: EBU [viewed 25 01 24]. Available from: https://tech.ebu.ch/docs/tech/tech3344.pdf. • OPERATING EUROVISION AND EURORADIO, 2023. EBU r128. Geneva: EBU [viewed 26 01 24]. Available from: https://tech.ebu.ch/docs/r/r128.pdf. • OPERATING EUROVISION AND EURORADIO, 2023. EBU r128s2. Geneva: EBU [viewed 26 01 24]. Available from: https://tech.ebu.ch/docs/r/r128s2.pdf. 68 • OPERATING EUROVISION AND EURORADIO, 2023. EBU Tech 3341. Geneva: EBU [viewed 25 01 24]. Available from: https://tech.ebu.ch/docs/tech/tech3341.pdf. • OPERATING EUROVISION AND EURORADIO, 2023. EBU Tech 3342. Geneva: EBU [viewed 25 01 24]. Available from: https://tech.ebu.ch/docs/tech/tech3342.pdf. • OPERATING EUROVISION AND EURORADIO, 2023. EBU Tech 3343. Geneva: EBU [viewed 25 01 24]. Available from: https://tech.ebu.ch/docs/tech/tech3343.pdf. • OWSINSKI, B., 2017. The mastering engineer’s handbook. Burbank, Ca: Bobby Owsinski Media Group. • PALMER, A. 2003. How the Ear Works and Why Loud Sounds Cause Hearing Loss. Nottingham: MRC Institute of Hearing Research [viewed 13 01 24] Available from: https://www.aes.org/elib/browse.cfm?elib=12261. • PRAS, A. and GUASTAVINO, C., 2010. Sampling Rate Discrimination: 44.1 kHz vs. 88.2 kHz. London: AES [viewed 21 02 24] Available from: https://www.researchgate.net/publication/257068631_Sampling_Rate_Di scrimination_441_kHz_vs_882_kHz. • PURVES D. et al., 2001. The Audible Spectrum. Available from: https://www.ncbi.nlm.nih.gov/books/NBK10924/. REISS, J., 2016. A Meta-Analysis of High Resolution Audio Perceptual Evaluation. London: University of London [viewed 02 03 24]. Available from: https://www.researchgate.net/publication/304572591_A_MetaAnalysis_of_High_Resolution_Audio_Perceptual_Evaluation. 69 • SCHORAH, J. and INGLIS, S., 2017. Mastering For Streaming Services. In: Sound On Sound’s Blog. June 2017 [viewed 10 02 24]. Available from: https://www.soundonsound.com/techniques/mastering-streamingservices. • SINGH N. et al., 2016. High ambient noise levels in Vadodara City, India, affected by urbanization. 1 December 2016 [viewed 09 03 24]. Available from: https://pubmed.ncbi.nlm.nih.gov/27902452/. • SONIBLE, 2019. Mastering Loudness. In: Sonible’s Blog. 27 May 2019 [viewed 13 02 24]. Available from: https://www.sonible.com/blog/mastering-loudness/. • SONIBLE, 2021. Normalization and streaming services. In: Sonible’s Blog. 11 December 2021 [viewed 12 02 24]. Available from: https://www.sonible.com/blog/normalization-and-streaming-services/. • SONNY ENTERTAINMENT, 2013. ASWG-R001. Audio Standards Working Group. [viewed 23 02 24]. Available from: http://gameaudiopodcast.com/ASWG-R001.pdf. • SONY, n.d. High-Resolution Audio. [viewed 07 02 24]. Available from: https://www.sony.co.uk/electronics/hi-res-audio. • SPOTIFY, n.d. Loudness normalization. [viewed 15 02 24]. Available from: https://support.spotify.com/uk/artists/article/loudnessnormalization/. • SPOTIFY, n.d. Loudness normalization. In: Spotify for Artists Articles. [viewed 05 02 24]. Available from: https://support.spotify.com/uk/artists/article/loudness-normalization/. 70 • STEWART, I., 2022. How to Master for Streaming Platforms: Normalization, LUFS, and Loudness. In: Izotope’s Blog. 20 May 2022 [viewed 10 02 24]. Available from: https://www.izotope.com/en/learn/mastering-for-streamingplatforms.html. • STEWART, I., 2022. What is the Fletcher Munson Curve? Using Equal Loudness Curves in Mixing and Mastering. In: Izotope’s Blog. 18 August 2022 [viewed 17 01 24]. Available from: https://www.izotope.com/en/learn/what-is-fletcher-munson-curve-equalloudness-curves.html. • STEWART, I., 2023. What Is a True Peak Limiter?. In: Izotope’s Blog. 21 February 2023 [viewed 01 03 24]. Available from: https://www.izotope.com/en/learn/true-peak-limiter.html. • STUART, J.R., 1997. Coding High Quality Digital Audio. [viewed 07 01 24]. Available from: https://www.researchgate.net/publication/286962221_Coding_for_HighResolution_Audio_Systems. • SWEETWATER, 2022. What Are Sample Rate and Bit Depth? [viewed 26 01 24]. Available from: https://www.youtube.com/watch?v=pJttTqdlpXo. • WATANABE, K., 2008. Objective perceptual audio quality measurement methods. In: NHK’s publications. [viewed 23 01 24]. Available from: https://www.nhk.or.jp/strl/english/publica/bt/35/2.html. 71 • WHITE SEA STUDIO, 2023. There are problems with oversampling… [viewed 08 01 24] Available from: https://www.youtube.com/watch?v=qhybrG0qIg. • WILLIAMSON, V. J., SOUTH, M. and MÜLLENSIEFEN, D., n.d. SOUND QUALITY ENHANCES THE MUSIC LISTENING EXPERIENCE. UK: University of Sheffield and University of London [viewed 20 01 24]. Available from: https://www.doc.gold.ac.uk/~mas03dm/papers/SoundQuality_Williamson SouthMullensiefen_ICMPC2014.pdf. • WORRAL, D., 2021. Linear, Logarithmic, Exponential & Perspective [viewed 23 01 24]. Available from: https://www.youtube.com/watch?v=rw7fkEDmDw. • WORRAL, D., 2022. Oversample Everything! Reaper FX and FX Chain Oversampling [viewed 24 01 24]. Available from: https://www.youtube.com/watch?v=GjtEIYXrqa8. 72 8. Appendices 8.1 Appendix A: Lossless Distribution The prevalence of lossless distribution formats is steadily increasing in contemporary audio distribution paradigms. Lossless audio formats offer a substantial reduction in the necessity for high maximum true peak buffers, as they circumvent the introduction of artifacts or distortions inherent in the conversion processes of lossy codecs. Embracing lossless distribution methodologies serves as a strategic imperative in averting generational loss, wherein each subsequent re-encode or transcode precipitates a progressive degradation in audio fidelity. Even if the initial encoding process proves transparent, subsequent iterations are susceptible to the emergence of audible artifacts, thereby jeopardizing the integrity and fidelity of the audio over time. Moreover, the adoption of lossless distribution formats engenders the preservation of high fidelity and quality over extended temporal horizons. This preservationist ethos not only safeguards the intrinsic value of audio content but also contributes to the enrichment of streaming platforms' catalogue, which serve as pivotal repositories of global historical and cultural heritage. A 8.2 Appendix B: High-Frequency Auditory Perception and its Impact on High-End Audio Content Sampling As previously discussed, the typical human auditory range spans from 20 Hz to 20 kHz for those with normal hearing, although this upper limit tends to decrease with age due to physiological changes. Beyond this conventional auditory spectrum, humans can perceive phenomena such as subsonic frequencies produced by earthquakes, which are not typically audible. Interestingly, there have been documented instances of individuals capable of perceiving frequencies as high as 24 kHz. This capability was demonstrated in studies such as 'Hearing Thresholds in Free-Field for Pure Tones above 20 kHz' by Ashihara K. et al. (2006). In this research, auditory thresholds for pure tones ranging from 2 kHz to 28 kHz were assessed. While no auditory thresholds were detected for tones above 26 kHz, thresholds at 24 kHz were observed in 4 out of 15 participants, with thresholds exceeding 88 dB SPL. Moreover, thresholds were evaluated under the masking conditions of a 20 kHz low-pass filtered noise to ascertain whether participants were perceiving subharmonics in the lower frequency range rather than the actual high-frequency signals. Considering these findings, using a sampling rate of 44.1 kHz could potentially result in the loss of extreme high-frequency content (between 20 kHz and 22 kHz). Increasing the sampling rate to 48 kHz could significantly mitigate this loss. Despite overlaps in channel frequency responses and auditory thresholds, typically these intersections occur at sound pressure levels over 100 dB SPL, a significant increase from the 88 dB SPL recorded in the study. While typical B program materials, such as music, almost never features content beyond 20 kHz and 100 dB SPL, the importance of these findings warrants further investigation, although challenges remain due to the nuanced distinctions between high-end frequency responses and auditory thresholds. B 8.3 Appendix C: Considerations for Antialiasing Linear Phase Filters In the realm of digital audio, there exists a multitude of filter options for the Antialiasing Filter of Your Signal. Among these choices is the linear phase type filter. Phase alterations inherently introduce coloration into the audio signal, but linear phase EQ stands out as a transparent option. As its name suggests, linear phase EQ ensures consistency in phase across the frequency spectrum, thereby minimizing additional coloration. It is important to note that linear phase mode does not inherently enhance audio quality or fidelity; rather, it pertains solely to how the filters manage phase. However, it is essential to consider that linear phase EQ may introduce noticeable latency, rendering it unsuitable for applications where real-time processing, such as live performances, is crucial. Additionally, the utilization of linear phase filters may lead to pre-ringing, which manifests as a peculiar sucking sound, particularly noticeable on transients. Technically, achieving a perfect linear phase necessitates the application of EQ twice—once forward in time and once backward in time—to nullify any phase changes. The latency incurred by linear phase EQ arises from the backward in time component nullifying phase alterations. While all EQs exhibit post-ringing to some extent, our auditory perception tends to be less sensitive to it. Linear phase EQ effectively mitigates post-ringing but introduces an equal and opposite amount of pre-ringing. C Pre-ringing manifests as an audible pre-echo preceding the onset of the signal's transient or wave. The steep linear-phase filters commonly employed in digital audio can exacerbate pre-ringing, potentially impacting arrival-time detection, and stereo imaging. Mitigating these effects can be achieved by reducing the steepness of the filter. Therefore, it is advisable to utilize higher sampling rates, such as 48kHz, if the filter is appropriately designed. However, if the filter lacks proper design, increasing the sample rate—thus implementing oversampling—may prove beneficial by shifting the downsides of this filter type into the audible range. C 8.4 Appendix D: Higher Sample Rates association with Professional Audio Thirty years ago, ADCs and DACs did not universally incorporate transparent oversampling techniques. During that time, certain recording consoles operated at elevated sampling rates by solely employing analog filters. Subsequent production and mastering processes capitalized on this high-rate signal to mitigate any adverse effects stemming from analog filters within the audible spectrum. Digital anti-aliasing and decimation steps, involving resampling to lower rates for formats like CDs or DAT, were typically executed during the final stages of mastering. This historical practice likely contributed to the association of 96 kHz and 192 kHz with professional music production, as they were among the sampling rates commonly utilized in this context. D 8.5 Appendix E: Nuances of Loudness Perception The human ear possesses the remarkable ability to consciously discern amplitude differences of approximately 1 dB, while studies indicate subconscious sensitivity to amplitude variances as minute as 0.2 dB. Given this remarkable sensitivity, it is noteworthy that humans tend to universally perceive louder audio as sounding better. Notably, a mere 0.2 dB difference is adequate to establish this preference. Consequently, any comparison lacking careful amplitude matching between options will likely result in the louder choice being preferred, irrespective of whether the amplitude distinction is perceptible at a conscious level. E
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )