Your Devices Are Lying to You: Why Flat Headphones and Speakers Matter
Part 1 — The Science of Hearing What Was Actually Recorded
There is a strange problem with listening to music.
You can spend thousands of dollars on a recording studio. You can hire an experienced engineer. You can record through great microphones, mix on professional monitors, master the final track carefully—and then play that exact same recording through a phone speaker, a pair of bass-heavy headphones, and a studio monitor and hear three noticeably different versions of the same song.
The digital file didn’t change.
The playback system did.
And that distinction matters more than most people realize.
To understand why, we need to talk about frequency response, equalization, digital audio formats, smartphone speakers, headphones—and what audio engineers actually mean when they describe something as flat.
What Does “Flat” Actually Mean?
Human hearing is conventionally described as spanning roughly 20 Hz to 20,000 Hz, although sensitivity varies substantially with age, listening level and the individual.
Those frequencies correspond roughly to:
20–60 Hz: sub-bass
60–250 Hz: bass
250 Hz–2 kHz: lower and central midrange
2–6 kHz: upper mids and presence
6–20 kHz: treble and air
A theoretically perfect playback system would reproduce the signal presented to it without arbitrarily making one part of that spectrum louder or quieter than another.
That concept is called frequency response.
When engineers describe a loudspeaker as “flat,” they generally mean its measured amplitude response is relatively even across its intended operating range rather than containing large peaks and valleys.
It does not mean the speaker sounds boring.
It means the speaker is trying not to rewrite the recording.
Research by Floyd Toole, Sean Olive and colleagues at Canada’s National Research Council and later Harman found a strong relationship between listeners’ preference and loudspeakers having a smooth, relatively flat on-axis response combined with smooth off-axis behavior. Modern standards such as CTA-2034’s collection of loudspeaker measurements grew from this body of work.
A loudspeaker with a 6 dB boost around 100 Hz, for example, isn’t merely “giving you more detail.”
It’s making bass in the recording significantly louder relative to other frequencies.
A dip around 3 kHz may pull vocals backward.
An elevated treble response may make cymbals, guitar attacks and vocal consonants sound more prominent.
The recording hasn’t changed.
The playback device has effectively applied an equalizer to it.
That’s Essentially What an Equalizer Does
An equalizer is a collection of filters that changes the amplitude of selected frequency regions.
A parametric EQ commonly gives the engineer three major controls:
Frequency — where the change occurs.
Gain — how much that frequency region is boosted or reduced.
Q or bandwidth — how wide or narrow the affected area is.
Modern parametric equalizers commonly implement combinations of bell filters, shelving filters, high-pass filters and low-pass filters.
Imagine applying this EQ to a recording:
100 Hz: +5 dB
1 kHz: 0 dB
4 kHz: -3 dB
10 kHz: +4 dB
You would immediately recognize a different tonal balance.
But here’s the important part:
Headphones and speakers can effectively do the same thing acoustically even when your EQ is switched off.
Every transducer has a frequency response.
Some manufacturers intentionally tune consumer headphones toward more bass or treble because they think customers will enjoy it.
Other manufacturers deliberately attempt to minimize coloration.
That’s one of the fundamental differences between consumer audio products and professional reference monitors.
Why Producers Don’t Mix Music on Random Speakers
Imagine you’re mixing a song using speakers that exaggerate bass by 6 dB.
The kick drum sounds enormous.
The bass guitar sounds fantastic.
So you turn both of them down.
Then somebody listens to the finished song through a neutral system.
Suddenly the bass is weak.
You didn’t intentionally make a bass-light record.
Your speakers lied to you while you were making it.
Now reverse the situation.
If your speakers are bass deficient, you may keep adding bass until the mix sounds correct.
On another playback system it becomes overwhelming.
This is the entire reason monitoring accuracy matters.
The goal of a studio monitor is not necessarily to make everything sound impressive.
Its job is to expose what’s actually there so the engineer can make decisions that translate to as many playback systems as possible.
Genelec, for example, specifies the 8030C studio monitor at ±2 dB from 54 Hz to 20 kHz and explicitly describes its objective as uncolored reference reproduction.
Neumann goes even further with the KH 120 II, specifying a passband response of 48 Hz–20.5 kHz within ±1.25 dB and a deviation of only ±0.7 dB between 100 Hz and 10 kHz.
Those aren’t meaningless audiophile adjectives.
They are measurable engineering specifications.
But Headphones Are More Complicated
Here’s where “flat” becomes confusing.
A headphone that measured as a mathematically straight 20 Hz–20 kHz line at the eardrum would not necessarily sound like a flat loudspeaker in a room.
Why?
Because our ears, head and torso normally alter sound before it reaches the eardrum.
Headphones bypass parts of that acoustic interaction.
This is why serious headphone research uses target curves, artificial ears and head-and-torso simulators rather than simply chasing a ruler-straight graph.
Sean Olive and colleagues at Harman performed controlled listening experiments investigating preferred headphone response curves. Their research found that listeners tended to prefer a headphone response corresponding more closely to the sound of a good loudspeaker operating in a room than traditional “free-field” or “diffuse-field” headphone equalization.
And the science continues to evolve.
A 2025 AES paper by Sean Olive and Dan Clark examined how headphone target curves need to change when measurements are made using newer B&K 5128 and GRAS test fixtures because different artificial ears have different acoustic impedances.
So when someone online says:
“This headphone is perfectly flat.”
Be cautious.
With speakers, “flat” can be defined reasonably intuitively.
With headphones and IEMs, the question must also be:
Flat relative to what acoustic target and what measurement fixture?
What About WAV, FLAC and MP3?
This is another area filled with misinformation.
Let’s clear it up.
WAV
WAV is a container format.
It can contain several kinds of audio, although uncompressed linear PCM is extremely common.
A typical CD-quality PCM WAV file contains:
44,100 samples per second
16 bits per sample
two channels
A 44.1 kHz sampling rate can theoretically represent frequencies up to just below 22.05 kHz according to the Nyquist theorem.
FLAC
FLAC stands for Free Lossless Audio Codec.
And this is extremely important:
FLAC does NOT inherently have a smaller frequency range than WAV.
FLAC compresses PCM audio without discarding its information.
Xiph’s official FLAC documentation states that decoded FLAC audio is bit-for-bit identical to the PCM presented to the encoder.
Think of it like ZIP compression specifically optimized for audio.
If you convert a 24-bit/96 kHz WAV recording to FLAC correctly and decode it again, you haven’t converted it into a lower-bandwidth recording.
You have compressed it losslessly.
This is why FLAC files are usually smaller than uncompressed WAV files without suffering MP3-style information loss.
MP3 Is Fundamentally Different
MP3 is a lossy perceptual codec.
Rather than retaining every original sample, an MP3 encoder uses psychoacoustic models to determine which information can potentially be removed while having the smallest perceptual consequence.
Fraunhofer—the organization intimately involved in MP3’s development—describes perceptual audio coding as relying partly on “irrelevancy removal”: exploiting properties of human perception to reduce the amount of information required.
This is why a 320 kbps MP3 file can be dramatically smaller than an uncompressed WAV.
But there is another important nuance.
It is inaccurate to say every MP3 has one fixed, reduced frequency range.
The encoder, bitrate, settings and source material matter.
At lower bitrates, encoders frequently reduce high-frequency bandwidth because spending precious bits encoding frequencies near the upper limit of human hearing can produce worse perceptual results elsewhere.
At high bitrates, a good encoder can retain considerably more bandwidth.
So the scientifically accurate comparison is:
WAV/PCM: potentially uncompressed representation of the samples.
FLAC: losslessly compressed PCM; decoded signal can be identical to the source.
MP3: perceptually compressed; some information is discarded, with the amount and nature depending on encoding.
Now Let’s Put the Phone Into the Equation
This is where things get particularly interesting.
A modern smartphone contains tiny loudspeakers inside an enclosure a few millimeters thick.
Those speakers face a physical problem.
A tiny driver cannot move air like an 8-inch studio-monitor woofer.
That makes deep bass particularly difficult.
Independent measurements of phones repeatedly demonstrate dramatically reduced bass extension compared with full-size playback systems. For example, measurements of the iPhone 16 found its internal speakers relatively balanced through the mids and highs but with very limited bass output.
Engineering a modern phone speaker therefore involves much more than feeding raw audio into a tiny driver.
Designers can employ:
equalization
filtering
dynamic range compression
limiting
driver excursion protection
psychoacoustic bass enhancement
volume-dependent processing
The purpose is often to make an impossibly small speaker sound larger, louder and more balanced while preventing its driver from destroying itself.
This isn’t inherently bad engineering.
It is very clever engineering.
But it means the phone speaker is not a neutral laboratory reference.
Mobile loudspeaker engineering literature explicitly describes using complementary equalization to compensate for the uneven response of miniature speakers.
Android even provides manufacturers with a framework for automatically attaching device-specific audio effects to particular playback routes, including effects implemented in hardware DSP.
Independent Samsung testing has also measured compression and “pumping” from built-in speakers at high playback levels on some Galaxy devices—exactly the type of dynamic behavior that demonstrates why a tiny phone speaker cannot simply be treated like a transparent output device.
But Here’s an Important Correction About Headphones and Bluetooth
It is tempting to say:
“Phones manipulate the sound through the speakers, but Bluetooth, USB-C and headphone outputs bypass all processing.”
That statement is too broad.
The accurate version is:
The audio processing chain can change depending on the output route.
A phone’s built-in speaker requires substantial acoustic compensation because of its physical limitations.
A USB DAC or wired headphone output can be dramatically more electrically linear.
Older laboratory measurements, for example, found essentially flat headphone-output frequency response from both Samsung Galaxy and iPhone hardware across the audible band.
But that does not mean every external audio route is automatically untouched.
Software EQ, Dolby processing, spatial audio, accessibility settings, app-level effects, sample-rate conversion and manufacturer-specific DSP can still exist.
And Bluetooth introduces something else:
a codec.
Apple states explicitly that wireless AirPods and Beats products use AAC over conventional Bluetooth and that conventional Bluetooth playback from iPhone is not lossless. Apple Music lossless can reach 24-bit/48 kHz directly through supported wired hardware, while playback above 48 kHz requires an external DAC.
So connecting headphones does not magically guarantee that the original file reaches your ears unchanged.
The entire playback chain matters.
The Playback Chain
Think about music reproduction this way:
Master recording
↓
File/stream codec
↓
Operating system / music application
↓
DSP / EQ / volume processing
↓
DAC
↓
Amplifier
↓
Headphone or speaker
↓
Room / ear anatomy
↓
Your brain
Every stage can influence the result.
And some stages matter much more than others.
The irony is that people will sometimes debate FLAC versus 320 kbps MP3 for hours while listening through a speaker with enormous frequency-response deviations.
The codec difference may be subtle.
A 5–10 dB acoustic deviation in the playback hardware is not.
“Flat” Does Not Mean “Better for Everyone”
This distinction matters.
A neutral system is useful because it provides a reference.
That doesn’t mean everyone must prefer it for recreational listening.
You may legitimately prefer:
+4 dB of bass.
A warmer tonal balance.
Softer treble.
A V-shaped headphone.
Aggressive sub-bass.
There is nothing scientifically wrong with preference.
The problem appears when preference is confused with accuracy.
A bass-heavy speaker may be wonderful.
It simply isn’t neutral.
And once you understand that distinction, you gain something very useful:
control.
You can begin with an accurate reference and then intentionally modify it.
Instead of allowing the hardware manufacturer to decide what your music should sound like, you decide.
The Takeaway
There is no single point in the playback chain called “sound quality.” It’s a system.
Lossless files preserve information. Accurate DACs preserve the electrical signal. Flat or well-targeted transducers minimize coloration. Good room acoustics reduce the rest. And careful EQ lets you change tonal balance deliberately rather than accidentally.
Every stage either preserves what was recorded or quietly rewrites it.
Two Different Kinds of Change
Everything above has been about one axis: amplitude. How loud each frequency region is relative to the others. A bass boost, a presence dip, a phone’s DSP compensating for a 6mm driver. Same notes, different balance.
Tuning is a different axis. Retuning shifts pitch — where the notes sit. A=432 instead of A=440. The tonal balance is untouched; the reference point moves.
These are independent. You can have either without the other.
Which is exactly why the first one matters if you care about the second.
Chosen vs. Inherited
Here’s the real distinction this whole article has been building toward.
A colored playback system colors your music without asking. Some engineer decided you’d enjoy +6 dB at 100 Hz. Your phone’s DSP decided how much compression to apply at your current volume. Your Bluetooth codec decided what to discard. You weren’t consulted, you weren’t told, and you can’t switch it off.
Retuning is the opposite kind of change. A value you pick. Applied when you want it. Off in one click, so you can compare.
That’s not a small difference. It’s the difference between hearing something and knowing what you’re hearing.
If your chain is introducing 5–10 dB of coloration you can’t see or predict, you can’t cleanly evaluate anything upstream of it. Not tuning. Not a remaster. Not a new pair of headphones you just spent $400 on.
Get the reference honest first. Then change it on purpose.
That’s the whole argument for a neutral system — not that flat sounds better, but that flat gives you control.
Our Retuning Apps do the second half: system-level retuning across everything your device plays, on or off instantly, so the change you hear is the one you chose.
And that brings us to the first half — the practical question:
What speakers and headphones should you actually buy?
That will be shared in Part 2.






