ABX Blind Test Player
You feed the ABX Blind Test Player two audio sources — generated tones or your own uploaded files, entered as Source A and Source B — then it plays an unlabeled sample and asks you to click X was Source A or X was Source B across repeated trials. You'll get a real statistical significance score at the end, since guessing alone only gets you to 50%. Watch the spectrum analyzer online update in real time as you test.
Your ears are not as reliable as you think, and that's exactly the problem a well-built ABX blind test player is designed to solve. Feed it two references — sample A and sample B — plus a hidden sample X that secretly matches one of them, and it strips away every cue except the sound itself: no labels, no price tags, no expectations. What comes back isn't a vibe, it's a number you can actually defend.
How an ABX Blind Test Player Works
Every ABX blind test player follows the same shape, whether it's comparing lossless and lossy audio, two microphones, or two mixing chains. You get two known references and one unknown, and your job is simply to say which one the unknown matches. Run that enough times and the proportion you get right separates real discrimination from chance: guessing converges on 50%, hearing a genuine difference does not. Run the sharpen your ear with absolute polarity blind for a few minutes a day to build real listening skill.
The ABX Protocol and Sample-Accurate Switching
The protocol itself is old and deliberately boring, which is the point — boring is what keeps a listening test honest. A single trial breaks down into four steps:
- Audition reference A as many times as you like.
- Audition reference B as many times as you like.
- Audition the hidden sample X, randomly assigned to match A or B each round.
- Commit an answer: is X the same as A, or the same as B?
Sample-accurate switching matters more than most first-time testers expect. If the player resets playback position every time you flip between A, B, and X, you lose the exact moment — a vocal sibilance, a cymbal crash, a bass drop — that actually revealed the difference. A good player carries your position across every switch instead.
Reference Samples, Random Guessing, and Why Order Matters
Because X is reassigned to A or B independently each round, there's no pattern to memorize and no shortcut around actually listening. That's what keeps the reference samples honest and keeps random guessing from quietly creeping into a "confident" result. If you could predict X from the order of previous rounds, the trial count would be worthless — the whole design exists to keep every round statistically independent of the last.
Reading Your p-value and Hit Rate
Once you've run a batch of trials, the player reports two numbers: your raw hit rate and the p-value behind it, the same statistical significance check that underlies psychoacoustics research generally. Hit rate is simple — Generate exactly what you need with the generate a noise, no install or sign-up required.
\( \text{Hit rate} = \dfrac{\text{correct answers}}{\text{trials}} \times 100\% \)
— but the p-value is what actually tells you whether that score means anything. It answers a narrower question than people assume: not "did you hear a difference," but "how likely is this score, or better, from pure guessing alone?"
Binomial Distribution: What Your Score Actually Proves
Because each trial is an independent coin flip when there's genuinely no audible difference, your results follow a binomial distribution. The probability of scoring at least k correct out of n trials by pure chance is:
$$ P(X \geq k) = \sum_{i=k}^{n} \binom{n}{i} \left(\frac{1}{2}\right)^{n} $$
Below is how that plays out for the standard 10-round test, plus a couple of longer runs:
| Trials correct | p-value | Verdict |
|---|---|---|
| 5 of 10 | 0.62 | Pure guessing — no evidence of an audible difference. |
| 6 of 10 | 0.38 | Inconclusive; could be a slight bias, could be chance. |
| 7 of 10 | 0.17 | Suggestive, but not statistically significant yet. |
| 8 of 10 | 0.055 | Borderline — right at the edge of significance. |
| 9 of 10 | 0.011 | Significant — roughly a 1% chance this was guessing. |
| 10 of 10 | 0.001 | Highly significant — an audible difference, confirmed. |
| 14 of 16 | 0.038 | Significant, with more confidence than the same rate on fewer rounds. |
| 18 of 20 | 0.001 | Highly significant — extremely unlikely to be chance. |
The threshold the field treats as "real" is p < 0.05. Below that line, the guessing explanation runs out; above it, you don't have enough evidence either way.
See how the p-value math scales with trial count
Fewer trials need a higher hit rate to reach significance, and more trials let a lower hit rate still qualify — a well-designed abx test lets you trade trial count for confidence deliberately, rather than stopping the moment a run looks good.
Null Result vs a Real Difference
A null result — failing to reach statistical significance — only tells you the difference wasn't audible to you, under these conditions, on this equipment. It's evidence bounding what you could detect, never proof that no difference exists. That distinction trips up more people than the statistics themselves: a blind test that comes back inconclusive is not the same as a test that proves two things sound identical.
Level Matching and Expectation Bias
Two controls do more to protect an abx test's validity than anything else: matching loudness, and hiding which sample is which.
Why Level-Matched Comparisons Matter
A difference in loudness of even half a decibel gets perceived as "better," which means any unmatched comparison quietly measures gain instead of the thing you meant to test. Level-matched tools typically normalize by RMS rather than peak:
$$ \text{RMS} = \sqrt{\frac{1}{N}\sum_{i=1}^{N} x_i^2} $$
Matching RMS rather than peak level keeps two clips with different crest factors — one heavily compressed, one dynamic — from biasing the result toward whichever one is simply louder to the ear.
Auditory Memory and Fast A/B/X Switching
Expectation bias is the second control, and it's the harder one to police yourself. Knowing which sample is the expensive cable, or which file is the "high-res" one, measurably changes what people report hearing — not through dishonesty, but because expectation is part of perception. Blinding exists specifically to remove that variable.
Common ABX Testing Mistakes
Most bad results trace back to a handful of repeatable errors rather than genuinely poor hearing:
- Deciding the trial count after seeing the score. Locking in the number of rounds beforehand is what keeps the statistics meaningful — stopping early because the run "looks good" invalidates the p-value.
- Comparing at mismatched levels. Without level matching, the louder file usually wins regardless of any real audible difference.
- Treating a null result as proof of no difference, when it only bounds what was audible under that specific test.
- Running too few trials for a forced-choice test. Three correct guesses in a row happens by coin flip about one time in eight — nowhere near significant.
- Ignoring the multiple comparisons problem when testing several tracks at once. Each extra comparison adds its own chance of a false positive, which is why a careful abx blind listening test applies something like a Bonferroni correction — tightening the significance bar for each additional track — rather than treating every test in isolation.
These aren't edge cases; they're the standard failure modes behind almost every "I can definitely hear it" claim that falls apart the moment psychometric testing methodology gets applied properly. A blind abx test isn't trying to catch you out — it's removing the one variable, expectation, that every informal listening test leaves in.
Lossless vs Lossy: FLAC, AAC, and MP3 Under the Microscope
The single most common use for a blind test player is settling the lossless-versus-lossy argument once and for all. A typical setup compares an uncompressed or FLAC reference against the same track re-encoded by a lossy codec — AAC at a mid-range bitrate, MP3 at a given kbps setting, or a specific encoder's output. Because the compression artifacts a lossy encoder introduces are usually subtle, you have to level match the two clips first; skip that step and sample-accurate switching won't save the result.
Related discrimination tests share the same abx test format: telling 16-bit from 8-bit audio, checking sensitivity to absolute polarity, or mapping the frequency response, dynamic range, and loudness thresholds your setup can actually resolve. None of these test "golden ears" in the abstract — they test one specific, falsifiable claim against one specific piece of music, on one specific playback chain, which is exactly why the result is worth trusting.
None of this requires a music streaming service's built-in test, either — a browser-based abx blind test running through the Web Audio API can decode local FLAC or MP3 files directly and apply RMS normalisation on load, without ever uploading your audio files anywhere.
Beyond File Formats: Audiophile Gear and Equipment Claims
Once you've run a listening test or two on codecs, the same approach works for any audiophile gear claim: a single dac, a pair of headphones, a set of speakers, a run of cables, even room treatment. If you can pass a blind listening test telling two cables apart on the same audio equipment, the sound quality difference is real; if you consistently can't, it probably wasn't there to begin with — no amount of confidence in your gear substitutes for a result.
Communities built around this kind of testing tend to grow into shared galleries: shootouts comparing compressors, EQs, saturation plugins, microphones, mastering chains, and even turntables, submitted as a shoutout and ranked by whoever wants to take the test. Most let you join in with no sign-up required, since a guest result is just as blind as a registered one.
One safety note worth repeating: loud comparison material can cause hearing damage with even a single exposure, so set your listening level modestly before your first trial, not after you've already been startled by one.
Playback Controls and Hotkeys for Faster Trials
Because a real abx blind listening test can mean dozens of trials, playback speed matters. Keyboard shortcuts let you switch between A, B, and X, seek within the clip, and rewind to a precise moment without ever reaching for the mouse:
A- Play or switch to reference sample A.
B- Play or switch to reference sample B.
X- Play or switch to the hidden sample X.
← / →- Seek backward or forward a few seconds within the clip.
Space- Rewind to the start of the current loop a section window.
Looping a short section — the two to five seconds around a vocal sibilance, a cymbal crash, or a bass drop — keeps every switch landing on the exact passage that matters, instead of forcing you to relocate it by ear every round.
When a session ends, most players let you export a certificate recording your score, your verdict, and whether the result reached statistically significant confidence — useful for settling an argument, or for building a track record of your own hearing profile over time.