Audio Codec Blind Test

Choose a simulated quality, then play the reference and compressed versions to compare.
Using generated tone
Click to play X.

Play Reference A or B to see its live spectrum. Hidden during the blind X trial so it can't give the answer away — revealed here once you guess.

0/0
Correct / Trials
What this actually is

Browsers can't perform real MP3/AAC encoding without a dedicated encoder library, which this site doesn't bundle. This tool instead approximates what low-bitrate lossy compression does perceptually — cutting high-frequency content and raising the noise floor — using real-time filtering. It's a genuine, audible simulation of those effects, not a bit-exact codec comparison. Treat a "pass" here as evidence you can hear bandwidth/noise-floor changes, not proof about any specific real-world codec.

Audio Codec Blind Test plays two versions of the same clip — one untouched, one run through a simulated low-bitrate codec — and dares you to tell which is which without knowing in advance. Upload your own audio or use the built-in tone, pick a simulated quality level, compare Reference A and Reference B, then guess X was A or X was B to build your score. Run the free vocal range test to see exactly what's working and what isn't.

An Audio Codec Blind Test strips away the packaging, the price tag, and the brand loyalty, leaving you with one honest question: can you actually hear the difference between two audio files? Whether you're comparing a 320 kbps MP3 against a lossless FLAC, or checking whether a new DAC actually resolves more detail, a properly run blind test trades guesswork for a real p-value instead of a hunch. This guide walks through the statistics behind your score, how to set up a fair listening test, and what a passing or failing result actually tells you about your ears and your gear.

What an Audio Codec Blind Test Actually Measures

At its core, an audio codec blind test is an ABX test: you get two known references, A and B, plus an unknown sample X that is secretly one of them. Your job across repeated trials is simply to say whether X matches A or matches B. Because you never know which reference the mystery sample really is, the test removes every non-auditory cue — the file name, the bitrate badge, the brand on the box — and leaves only what your ears actually pick up. Try the free free frequency identification training for a quick way to test your own listening discrimination.

A single correct guess proves nothing; flipping a coin gets you there half the time anyway. What makes the result meaningful is auditory discrimination measured across enough trials that the pattern can't plausibly be luck.

The ABX Protocol: A, B, and X

Every round follows the same shape. You get two labeled samples — A and B — that you can audition as many times as you like, plus a hidden target sample, X, randomly assigned to match one of the two audio sources.

Switch between them quickly: your auditory echoic memory holds detail for roughly three to four seconds, so long pauses between A, B, and X erase exactly the cue you're trying to compare. Commit an answer, move on, and repeat for as many rounds as your test allows.

Reading Your Blind Test Score: p-value and Whether It's Statistically Significant

Every trial you complete is a coin flip unless you can hear something real, so your score only matters relative to what pure random guessing would produce. The p-value answers that directly: it's the probability of scoring at least as well as you did if you'd been choosing at random the whole time. A result is usually called statistically significant once that probability drops below 5% — in other words, once random guessing stops being a believable explanation for your score. The check whether your impulse response room is working properly gives you a clear answer instead of guessing.

The math behind it is a straightforward binomial distribution collapsing to a coin flip:

$$P(X \geq k) = \sum_{i=k}^{n} \binom{n}{i} (0.5)^i (0.5)^{n-i}$$

With \(n\) trials and \(k\) correct answers, this tells you exactly how surprising your score really is. Conventionally, a result crosses the line at p < 0.05. A null result doesn't mean no difference exists — it means this difference wasn't audible to you, on this equipment, on this day.

Scorep-valueVerdict
5 / 100.623Pure guessing — no evidence of an audible difference.
7 / 100.172Suggestive, but not significant. Run more trials.
8 / 100.055Borderline — right on the edge of significance.
9 / 100.011Significant. Only about a 1% chance of guessing.
10 / 100.001Highly significant. You can hear it.

What Your Hit Rate Actually Proves

Your hit rate is just the raw fraction of trials you got right, and on its own it proves very little — 6 or 7 out of 10 still overlaps heavily with pure chance. The p-value converts that raw number into a claim you can actually defend, which is why two testers with the same hit rate on a different number of trials can land on opposite sides of "significant."

Lossy vs Lossless: Setting Up Your Own Listening Test

Most head-to-head comparisons come down to lossy versus lossless encoding. A lossless file (FLAC, WAV) preserves every bit of the original recording; a lossy codec — MP3, AAC — throws away data its encoder predicts you won't miss, in exchange for a smaller file.

That prediction is usually good. At high bitrates it's good enough that even trained listeners struggle to tell uncompressed audio from a well-tuned lossy encoding, which is exactly why music compression deserves a real test rather than a confident guess.

To compare fairly, feed your test the same track in both file formats — say, an AAC file ripped at a high bitrate against the original FLAC — and make sure your audio system, whether that's studio monitors, a pair of headphones, or a stack of outboard DACs and cables, isn't itself the bottleneck.

Level Matching: The Control Most Home Tests Skip

Most home blind tests skip this control entirely. Level matching is the single most important addition you can make, because a file that's even half a decibel louder tends to sound "better" regardless of its codec — any comparison that isn't level-matched is secretly testing volume, not quality. RMS normalisation removes that bias before you hear a single sample, so the audible difference you're judging under level-matched conditions is the codec's, not the mixing engineer's.

Common File Formats You Can Put Head-to-Head

Bit depth is worth testing on its own: a 16-bit file versus the same track dithered down to 8-bit makes quantization noise obvious even on modest gear. Sample rate is subtler — most high fidelity claims lean on sample rates well above what CD audio uses, though the audible payoff shrinks fast past 48kHz. And because so much listening now happens through a streaming service rather than a downloaded MP3, it's worth testing your music streaming app's own codec, not just a file sitting on your desktop.

Setting Up a Rigorous ABX Blind Test

Running your own test doesn't require special software — a few browser-based tools, and no shortage of home-brew scripts, handle the randomization for you. Here's what to have ready, and what will actually help you tell the difference between two encodes rather than just guess at it:

  • A source track you know well, encoded two different ways
  • Headphones or speakers you already trust
  • Ten uninterrupted minutes and a quiet room
  1. Pick two versions of the same audio sources — the same track, same section, encoded two different ways.
  2. Level-match the files so loudness isn't secretly doing the work for you.
  3. Loop the most revealing few seconds — a cymbal decay, sibilance, a quiet passage — rather than judging the whole track.
  4. Run at least ten trials, committing an answer before you second-guess yourself.
  5. Record your hit rate and let the p-value tell you whether it's real.

How Many Trials Do You Need?

Ten trials is the conventional floor: nine out of ten clears the standard significance bar, and ten out of ten is about as convincing as a single test gets. If you want to defend the result — say, in an audiophile forum argument — sixteen to twenty rounds pushes your confidence well past casual doubt, because a statistically significant score on more trials is harder to dismiss as a lucky streak.

What a High Fidelity Test Result Can (and Cannot) Prove

A blind test's entire value comes from removing expectation bias — the well-documented tendency to hear what you expect to hear once you know which sample is the expensive one. That's not dishonesty; it's how perception works, and it's exactly what a forced-choice design like ABX is built to strip out. As a statistical method, ABX doesn't just report whether you noticed a difference — it separates genuine auditory discrimination from a fortunate run of guesses.

Expectation Bias, Bonferroni Correction, and the Multiple Comparisons Problem

There's a subtler statistical trap worth knowing about. If you test five tracks and treat each result independently, the odds of a false positive on at least one of them climb fast — this is the multiple comparisons problem. The standard fix is a Bonferroni correction: divide your significance cut-off by the number of tracks you're testing.

$$\alpha_{adjusted} = \frac{0.05}{n}$$

Test five tracks and your real cut-off tightens from 5% to 1% per track — which is why a single lucky guess on one track out of five shouldn't convince anyone.

Why the correction gets stricter as you test more tracks

Each individual track carries its own 5% chance of a false positive. Test one track and that's your only risk.

Test five, and the chance that at least one shows a false "pass" by chance alone climbs to roughly 25% — five independent 5% risks compounding. Dividing the cut-off by the track count keeps your overall error rate pinned to 5% no matter how many comparisons you run.

Binomial
The statistical distribution behind a plain coin-flip guess — the baseline every blind test result gets compared against.
Bonferroni correction
Dividing your significance cut-off by the number of tracks you test, so multiple comparisons don't inflate your false-positive rate.
Audible difference
A gap between two samples large enough to survive a fair, level-matched, blind comparison.
Multiple comparisons problem
The statistical trap of testing many tracks and treating each pass as independently meaningful.

Beyond Your Audio Quality Quiz: Building Real-World Listening Confidence

Once you've run a formal ABX test, it's worth training the underlying skill instead of just measuring it. Audio quality perception isn't fixed — dedicated ear-training tools built on web audio APIs can sharpen your sensitivity the same way scales sharpen a musician's ear.

None of this requires expensive audio equipment. A mid-range pair of headphones, a source file, and a willingness to fail a few rounds before you succeed will tell you more about your hearing than any spec sheet. Audiophile claims are, in the end, testable claims — and a forced-choice trial is the cheapest way to find out whether you, personally, can hear the difference or just want to.

If you can consistently tell the difference between two samples across enough trials to earn a real p-value, that result is yours to keep — a shareable certificate of auditory transparency, backed by statistics instead of a confident opinion. If you can't, that's useful information too: it means your money is better spent on audiophile gear you can actually hear, not gear you're told you should.

Your Audio System, Audible Difference, and Stereo Imaging

Some tests go beyond a simple codec swap and probe your audio system itself: perfect pitch, absolute polarity (whether your system reproduces a waveform right-side up), absolute phase, and stereo imaging — the sense of instruments occupying distinct space rather than a blurred center. Each is its own forced-choice question with its own honest, statistical answer.

Frequency, Pitch, Timing, and Dynamic Range

Beyond lossy versus lossless, the same ABX approach works on almost any listening claim: the highest frequency you can reliably hear, the smallest pitch shift you can detect, the shortest timing difference that registers, or the usable dynamic range of your listening environment. They're all, at bottom, the same question — can you tell the difference, or can't you — just pointed at a different variable.