Spotify deletes most of a song's data. Your ears let it get away with this

Spotify deletes most of a song's data. Your ears let it get away with this

Take any carriage of the last Metro out of Rajiv Chowk. It is nearly midnight, the doors sigh, the rails hum a low, steady note beneath everything. A woman by the window has closed her eyes, one earphone in, a song she has loved since college folded into the rattle of the train. The song reaching her is mostly absent.Somewhere between a studio and that earphone, a machine went through the song moment by moment and removed most of the information a CD would have carried, so carefully that she would never miss it.In this instalment of Science of Sound, we find out what was removed, and why her own ears let the streamed song get away with it.THE FILE THAT SHRINKSSound is a river, and a CD cannot carry a river. It carries snapshots of it, taken so fast that the ear hears a flow. Think of a flipbook, where each page is a still picture and the pages turn so quickly that the picture seems to move. A CD's flipbook turns 44,100 pages or snapshots every second. Why 44,100 pages? The human ear hears sounds up to about 20,000 vibrations a second or hertz. To capture any vibration faithfully, you must take at least two snapshots for every one of its cycles or frequencies. Engineers call this the Nyquist sampling rule. Twice 20,000 is 40,000. Adding a small margin of safety gave the CD its figure of 44,100 vibrations per second.Here is how much information those pages hold. Step 1: One snapshot. Each snapshot records how loud the sound is at that instant, and the machine writes that loudness as a code of 16 switches, each either on or off. Each switch is called a bit. Together, 16 switches can be set in 65,536 different ways, so a snapshot can pick from 65,536 loudness levels.Step 2: One second, one ear. Take 44,100 snapshots and give each one 16 bits.44,100 multiplied by 16 equals 705,600 bits.Step 3: One second, two ears. Music on a CD has a left channel and a right channel, so everything doubles.705,600 multiplied by 2 equals 1,411,200 bits every second.That is where the figure of 1,411 kilobits per second comes from. Kilo means a thousand, so 1,411 kilobits is about 1,411,000 bits.Step 4: What Spotify Premium sends. At its best setting, Spotify sends 320 kilobits, or 320,000 bits, every second. Set that against the CD.320,000 divided by 1,411,200 equals 0.227, or about 23 per cent.Step 5: What the free tier sends. The free tier sends about 160,000 bits every second.160,000 divided by 1,411,200 equals 0.113, or about 11 per cent.Step 6: What is left behind. Take each answer away from 100 per cent.100 minus 23 equals 77 per cent gone on Premium.100 minus 11 equals 89 per cent gone on the free tier.Step 7: What it means for a phone. One hour has 3,600 seconds. Eight bits make one byte, the unit your phone uses to count storage.On a CD: 1,411,200 multiplied by 3,600, divided by 8, equals 635,040,000 bytes, or about 635 megabytes.On Premium: 320,000 multiplied by 3,600, divided by 8, equals 144,000,000 bytes, or 144 megabytes.On the free tier: 160,000 multiplied by 3,600, divided by 8, equals 72,000,000 bytes, or 72 megabytes. Spotify streams to free listeners at about 160 kilobits per second, barely 11 per cent of the 1,411 kilobits a CD carries. (Photo: Unsplash) Spotify is not alone. Google's support page says YouTube Music Premium members can stream at up to 256 kilobits per second, using codecs called AAC and Opus, while the default setting tops out at 128.Between 77 and 89 per cent of a CD's information never leaves Spotify's servers, and the only witness is the listener who did not miss it.An hour of CD-quality music takes about 635 megabytes. At 320, it takes about 144. For a service with hundreds of millions of listeners, that difference is the whole business.This is lossy compression. Compression means the file is made smaller, and lossy means something is lost along the way. Its opposite, lossless compression, is like a zipped folder, which shrinks and then restores everything exactly.Amrit Srivastava, a musicologist based in Rome, puts it plainly. You make the bitrate smaller to get a smaller file, he told India Today Digital, and what you are doing is squeezing out some parts of the sound.The astonishing part is that the deletion is chosen, so that it falls only on what the ear would have missed anyway. How can a machine know that?WHEN ONE SOUND BLINDS ANOTHERThe answer is masking, and everyone has experienced it. Stand beside a roadside drill and try to hear a friend whisper. The whisper still travels through the air. The ear simply cannot register it under the roar.Hearing science measures this. A loud sound raises the level a quieter sound must reach to be heard, and the effect is strongest for sounds close to it in pitch. Pitch is how high or low a sound is, counted in vibrations per second, or hertz. This is frequency masking. The ear sorts pitch into about 24 bands, called critical bands, a division first set out by Eberhard Zwicker in a 1961 paper in the Journal of the Acoustical Society of America. Spotify Premium streams at up to 320 kilobits per second, so an hour of music takes about 144 megabytes against 635 for a CD. (Photo: Unsplash) Masking also works across time. After a loud burst, the ear stays partly deaf to quieter sounds for a fraction of a second. This is temporal masking. A technical report on the theory behind MP3 describes how both effects are used to decide what a compressed file can drop.So, when a cymbal crashes, a soft violin note at a similar pitch in that instant is still in the air, yet effectively gone from hearing.THE EAR'S OWN BLIND SPOTSMasking is not the only gap. The ear is also unevenly sensitive. It hears best between about 2,000 and 5,000 hertz, and grows steadily deafer towards deep bass and very high pitches, a pattern first mapped by Harvey Fletcher and Wilden Munson in 1933 and known as the equal-loudness contours.The software that does this is called a codec, short for coder and decoder.It slices the song into frames, each lasting a tiny fraction of a second. For every frame, it splits the sound into frequency bands that mirror the ear's own. It finds the loudest components and calculates how much each one masks the sounds beside it. The result is an invisible line called the masking threshold. Spotify's apps use Ogg Vorbis, an open, patent-free format that discards sound the ear is unlikely to miss. (Photo: Unsplash) Everything above the line is kept. Everything below it is judged inaudible and is not stored.This is the idea behind MP3. The audio standard it belongs to, MPEG-1, was published in 1993, and its engineers Karlheinz Brandenburg and Gerhard Stoll described the scheme in a 1994 paper in the Journal of the Audio Engineering Society. Spotify's apps use a relative called Ogg Vorbis, an open, patent-free format built on the same principle.Brandenburg has recalled how he tuned his encoder on an unaccompanied recording of Suzanne Vega singing Tom's Diner. He knew it would be nearly impossible to compress that warm a cappella voice. A lone voice leaves nothing to hide behind, which made it the hardest possible test.WHAT YOU LOSE, AND WHO NOTICESWhat disappears is subtlety. Srivastava compares the effect to hanging four carpets on your wall when the neighbours are having a party. The party is still there, but its finer textures are muffled.Whether you notice, he says, depends on how you listen. The difference shows only if you listen to a lot of good, well-mixed music on great headphones or a hi-fi. A casual listener has long been used to it. Spotify added a lossless tier for Premium listeners in September 2025, streaming FLAC at 24-bit and 44.1 kilohertz with nothing deleted. (Photo: Unsplash) Apple makes a similar point. Its support page says the difference between AAC, its compressed format, and lossless is virtually indistinguishable, though it still offers lossless as an option for those who want it.Srivastava's advice for anyone who wants the subtleties back is simple: a good pair of speakers at home, a soundbar for a television, and better headphones on the move.WHERE THE TRICK STRUGGLESMasking works beautifully inside a dense rock chorus, where plenty of loud sound is available to hide the deletions. It has less cover when a single voice or a single flute holds a note.Indian music asks a particular question here. A gamakam is an ornament, a continuous glide or oscillation around a note, and it is the soul of a raga. Professor M. Ramanathan of IIT Madras, who works on teaching machines to recognise ragas, says the music resists easy rules. Each performer can render a raga in a different manner, he told India Today Digital, so it is difficult to formulate rules, and only broad guidelines are available.We asked him two questions. Could compression make gamakams harder to catch? And might compression built mainly for simpler, Western-style music handle such variations badly? He said he would have to test both before answering properly. Even on Spotify's lossless tier, a Bluetooth earphone compresses the sound again before it reaches the ear. (Photo: Unsplash) His possible short answers were these. To the first: "No, we don't usually compress." To the second: "Maybe yes."On a related point, he was more reassuring. To identify a raga, he said, "it is probably sufficient to have a global feature recognition system without going into nuances like gamakam." In plain terms, a machine can name a raga from its overall shape, the notes it favours and the way it moves, without tracking every glide.Recognising a raga is a different task from savouring one, and the first may survive compression better than the second. The question for the next experiment is whether a method tuned to hide deletions in dense music treats a gliding note with the same care.THE LOSSLESS TWISTThe industry has an answer of its own. After announcing changes to its Premium plans in August 2025, Spotify added a lossless tier in September, streaming Free Lossless Audio Codec (FLAC) at 24-bit and 44.1 kilohertz, with nothing deleted.It arrived late. Apple's support page says Apple Music offers lossless audio in a format called ALAC (Apple Lossless Audio Codec), from CD quality up to 24-bit and 192 kilohertz, and Apple announced it in 2021. In the same month, Amazon Music said its HD tier would stream lossless audio at CD quality, with Ultra HD going up to 24-bit and 192 kilohertz, at no extra cost to Unlimited subscribers.There is one more catch, and Apple states it itself. AirPods and Beats headphones use Apple's AAC Bluetooth codec, and Bluetooth connections are not lossless. So, even on a lossless stream, the last few metres to a wireless earphone are compressed again.Every song on a phone is a quiet collaboration between the recording and the limits of the person listening. The ear has blind spots, and a machine learnt to hide inside them.Music has always been finished by the listener. Streaming has simply learnt where the listener stops.- Ends

Original Source

Read the full article at Indiatoday →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.