Skip to content
Novus Examples

Audio

Audio test files usually arrive as someone's music clip. These are engineered. Silence sets document leading and trailing trim points; ASR suites pair synthetic digit utterances with transcripts and noise twins; loudness ladders, peak-clip pairs, stereo/mono, sample-rate conversion, and DTMF sequences exercise meters and harnesses. Every file is short to stay within budget and ships with documented sample rate, bit depth, and duration.

Filter audio on Browse · 396 files · 18 subcategories

Frequently asked questions

Are the ASR clips real speech recordings?

No, Wave B ASR fixtures are synthetic tone-digit utterances with JSON transcripts. They exercise ASR harnesses and loaders without shipping copyrighted speech.

What are the silence trim fixtures for?

Documented leading/trailing silence (1–5 s) on pure tones so auto-trim tools can be checked against exact timestamps. Wave E also adds mid-gap silence cases for split/trim tools.

Which formats are included?

WAV references plus compressed codecs (MP3, FLAC, OGG, and more). Specs list sample rate, bit depth, and duration.

How do I test loudness meters or peak clip detection?

Filter Browse by purpose loudness-testing for the amplitude ladder and clean↔hard-clipped peak pairs. Stereo/mono and sample-rate twins cover channel and resampler QA.

396 of 396 files

Asr

Preview of ASR: 0123 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 0123 Clean (WAV)

ASR: 0123 Clean (WAV)

Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 0123 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 0123 Noisy (WAV)

ASR: 0123 Noisy (WAV)

Noisy twin of the digit sequence 0123 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 1357 Clean (WAV)

ASR: 1357 Clean (WAV)

Synthetic clean tone sequence encoding digits 1357 (one three five seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 1357 Noisy (WAV)

ASR: 1357 Noisy (WAV)

Noisy twin of the digit sequence 1357 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Transcript (JSON)
json
244 B
Actual file preview for ASR: 1357 Transcript (JSON)

ASR: 1357 Transcript (JSON)

Ground-truth transcript for the 1357 ASR utterance pair, expected text: “one three five seven”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 24680 Clean (WAV)
wav
45.7 KB
Actual file preview for ASR: 24680 Clean (WAV)

ASR: 24680 Clean (WAV)

Synthetic clean tone sequence encoding digits 24680 (two four six eight zero). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 24680 Noisy (WAV)
wav
45.7 KB
Actual file preview for ASR: 24680 Noisy (WAV)

ASR: 24680 Noisy (WAV)

Noisy twin of the digit sequence 24680 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 24680 Transcript (JSON)
json
249 B
Actual file preview for ASR: 24680 Transcript (JSON)

ASR: 24680 Transcript (JSON)

Ground-truth transcript for the 24680 ASR utterance pair, expected text: “two four six eight zero”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 4567 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 4567 Clean (WAV)

ASR: 4567 Clean (WAV)

Synthetic clean tone sequence encoding digits 4567 (four five six seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 4567 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 4567 Noisy (WAV)

ASR: 4567 Noisy (WAV)

Noisy twin of the digit sequence 4567 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 4567 Transcript (JSON)
json
246 B
Actual file preview for ASR: 4567 Transcript (JSON)

ASR: 4567 Transcript (JSON)

Ground-truth transcript for the 4567 ASR utterance pair, expected text: “four five six seven”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 89 Clean (WAV)
wav
21.3 KB
Actual file preview for ASR: 89 Clean (WAV)

ASR: 89 Clean (WAV)

Synthetic clean tone sequence encoding digits 89 (eight nine). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 89 Noisy (WAV)
wav
21.3 KB
Actual file preview for ASR: 89 Noisy (WAV)

ASR: 89 Noisy (WAV)

Noisy twin of the digit sequence 89 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 987654 Clean (WAV)
wav
53.8 KB
Actual file preview for ASR: 987654 Clean (WAV)

ASR: 987654 Clean (WAV)

Synthetic clean tone sequence encoding digits 987654 (nine eight seven six five four). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 987654 Noisy (WAV)
wav
53.8 KB
Actual file preview for ASR: 987654 Noisy (WAV)

ASR: 987654 Noisy (WAV)

Noisy twin of the digit sequence 987654 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 987654 Transcript (JSON)
json
258 B
Actual file preview for ASR: 987654 Transcript (JSON)

ASR: 987654 Transcript (JSON)

Ground-truth transcript for the 987654 ASR utterance pair, expected text: “nine eight seven six five four”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set

Bit Depth Ladder

Preview of 440 Hz Tone @ 16-bit (WAV)
wav
172.3 KB
Actual file preview for 440 Hz Tone @ 16-bit (WAV)

440 Hz Tone @ 16-bit (WAV)

A 440 Hz sine tone stored as 16-bit PCM: part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

File
WAV · Bit Depth Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 440 Hz Tone @ 24-bit (WAV)
wav
258.4 KB
Actual file preview for 440 Hz Tone @ 24-bit (WAV)

440 Hz Tone @ 24-bit (WAV)

A 440 Hz sine tone stored as 24-bit PCM: part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

File
WAV · Bit Depth Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 440 Hz Tone @ 8-bit (WAV)
wav
86.2 KB
Actual file preview for 440 Hz Tone @ 8-bit (WAV)

440 Hz Tone @ 8-bit (WAV)

A 440 Hz sine tone stored as 8-bit PCM: part of a bit-depth ladder (8/16/24-bit) of the identical tone, for testing bit-depth conversion, dithering, and quantisation-noise handling.

File
WAV · Bit Depth Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set

Channels

Preview of 5.1 Surround Channel Identification (WAV, 6ch)
wav
3 MB
Actual file preview for 5.1 Surround Channel Identification (WAV, 6ch)

5.1 Surround Channel Identification (WAV, 6ch)

A six-channel 5.1 surround WAV that plays a 1-second tone in each channel in turn, in the standard WAV order (FL, FR, C, LFE, SL, SR). A fixture for testing multichannel decoding, downmixing, and channel-order handling.

File
WAV · Channels · 44100 Hz
Preview of Stereo Channel Map: L 440 Hz / R 880 Hz (WAV)
wav
516.8 KB
Actual file preview for Stereo Channel Map: L 440 Hz / R 880 Hz (WAV)

Stereo Channel Map: L 440 Hz / R 880 Hz (WAV)

A stereo WAV with a distinct tone in each channel (440 Hz on the left, 880 Hz on the right), so you can verify channel mapping, panning, and L/R routing unambiguously.

File
WAV · Channels · 3 s

Codec Set

Preview of AAC: ADTS
aac
71.9 KB
Actual file preview for AAC: ADTS

AAC: ADTS

The clip as raw AAC in an ADTS stream: the codec behind most streaming and mobile audio. For testing AAC decoders and remux into MP4.

File
AAC · Codec Set · AAC-LC
Use case
Conversion testingAudio analysis· Conversion set
Preview of AC-3: Dolby Digital
ac3
71 KB
Actual file preview for AC-3: Dolby Digital

AC-3: Dolby Digital

The clip as AC-3 (Dolby Digital): the multichannel codec used in DVD/broadcast. Rendered here in stereo; for testing AC-3 decoding and conversion.

File
AC3 · Codec Set · AC-3 (Dolby Digital)
Use case
Conversion testingAudio analysis· Conversion set
Preview of ALAC: Apple Lossless
m4a
184.4 KB
Actual file preview for ALAC: Apple Lossless

ALAC: Apple Lossless

The clip as Apple Lossless (ALAC) in an M4A container: lossless, unlike the AAC M4A twin. For testing ALAC decoding and lossless conversion.

File
M4A · Codec Set · ALAC
Use case
Conversion testingAudio analysis· Conversion set
Preview of AMR: Narrowband Speech
amr
4.7 KB
Actual file preview for AMR: Narrowband Speech

AMR: Narrowband Speech

The clip resampled to 8 kHz mono and encoded as AMR narrowband: the telephony/voice-note codec. For testing AMR decoding and speech-codec conversion.

File
AMR · Codec Set · AMR-NB
Use case
Conversion testingAudio analysis· Conversion set
Preview of FLAC: Lossless
flac
102.2 KB
Actual file preview for FLAC: Lossless

FLAC: Lossless

The clip as FLAC: free lossless audio compression. Byte-for-byte recoverable to the source PCM; for testing lossless decoders and conversion.

File
FLAC · Codec Set · FLAC
Use case
Conversion testingAudio analysis· Conversion set
Preview of M4A: AAC (MP4)
m4a
72.3 KB
Actual file preview for M4A: AAC (MP4)

M4A: AAC (MP4)

The clip as AAC in an MP4/M4A container (faststart): Apple's default audio container. For testing M4A parsing and MP4 audio conversion.

File
M4A · Codec Set · AAC-LC
Use case
Conversion testingAudio analysis· Conversion set
Preview of M4R: iPhone Ringtone
m4r
72.3 KB
Actual file preview for M4R: iPhone Ringtone

M4R: iPhone Ringtone

The clip as an M4R iPhone ringtone: AAC in an MP4 container with the ringtone extension. For testing that a converter maps M4R↔M4A correctly.

File
M4R · Codec Set · AAC-LC
Use case
Conversion testingAudio analysis· Conversion set
Preview of MP3: CBR 192 kbps
mp3
71.7 KB
Actual file preview for MP3: CBR 192 kbps

MP3: CBR 192 kbps

The source clip as constant-bitrate MP3 (LAME, 192 kbps): the most universally supported lossy audio format. For testing MP3 decoders, players, and conversion.

File
MP3 · Codec Set · MP3 (LAME)
Use case
Conversion testingAudio analysis· Conversion set
Preview of MP3: VBR (V2)
mp3
35.1 KB
Actual file preview for MP3: VBR (V2)

MP3: VBR (V2)

The same clip as variable-bitrate MP3 (LAME V2), for testing VBR handling, seeking, and duration estimation against the CBR twin.

File
MP3 · Codec Set · MP3 (LAME)
Use case
Conversion testingAudio analysis· Conversion set
Preview of OGG: Vorbis
ogg
16 KB
Actual file preview for OGG: Vorbis

OGG: Vorbis

The clip as Ogg Vorbis: a royalty-free lossy codec. Browser-playable; for testing Vorbis decoding and Ogg conversion.

File
OGG · Codec Set · Vorbis
Use case
Conversion testingAudio analysis· Conversion set
Preview of Opus
opus
51.9 KB
Actual file preview for Opus

Opus

The clip as Opus: the modern low-latency codec used by WebRTC and streaming. Browser-playable; for testing Opus decoding and conversion.

File
OPUS · Codec Set · Opus
Use case
Conversion testingAudio analysis· Conversion set
Preview of WAV: Codec-Set Source
wav
516.8 KB
Actual file preview for WAV: Codec-Set Source

WAV: Codec-Set Source

The lossless PCM WAV source for the audio conversion set: the same 3-second tone every other codec in this group is encoded from. Use it as the reference when diffing encoders.

File
WAV · Codec Set · PCM
Use case
Conversion testingAudio analysis· Conversion set
Preview of WMA: Windows Media Audio
wma
103.7 KB
Actual file preview for WMA: Windows Media Audio

WMA: Windows Media Audio

The clip as Windows Media Audio (WMA v2) in an ASF container: Microsoft's lossy codec. For testing WMA decoding and conversion to open formats.

File
WMA · Codec Set · WMA v2
Use case
Conversion testingAudio analysis· Conversion set

Container Formats

Preview of AIFF: 1 kHz Sine Tone
aiff
258.5 KB
Actual file preview for AIFF: 1 kHz Sine Tone

AIFF: 1 kHz Sine Tone

A pure 1 kHz sine tone stored as AIFF: Apple's big-endian 16-bit PCM container, 3 seconds, 44.1 kHz mono. A clean reference for testing AIFF decoders and WAV↔AIFF conversion.

File
AIFF · Container Formats · 3 s
Preview of AIFF: 440 Hz Sine Tone
aiff
258.5 KB
Actual file preview for AIFF: 440 Hz Sine Tone

AIFF: 440 Hz Sine Tone

A pure 440 Hz sine tone stored as AIFF: Apple's big-endian 16-bit PCM container, 3 seconds, 44.1 kHz mono. A clean reference for testing AIFF decoders and WAV↔AIFF conversion.

File
AIFF · Container Formats · 3 s
Preview of AU: 1 kHz Sine Tone
au
129.2 KB
Actual file preview for AU: 1 kHz Sine Tone

AU: 1 kHz Sine Tone

A pure 1 kHz sine tone stored as a Sun/NeXT AU file, big-endian 16-bit PCM, 3 seconds, 44.1 kHz mono. A compact reference for AU decoding and format conversion.

File
AU · Container Formats · 3 s
Preview of AU: 440 Hz Sine Tone
au
129.2 KB
Actual file preview for AU: 440 Hz Sine Tone

AU: 440 Hz Sine Tone

A pure 440 Hz sine tone stored as a Sun/NeXT AU file, big-endian 16-bit PCM, 3 seconds, 44.1 kHz mono. A compact reference for AU decoding and format conversion.

File
AU · Container Formats · 3 s

Dtmf

Preview of DTMF Dialing Tones (WAV)
wav
254.1 KB
Actual file preview for DTMF Dialing Tones (WAV)

DTMF Dialing Tones (WAV)

A DTMF (touch-tone) dialing sequence dialing a reserved 555-01xx number: each digit is the standard dual-tone pair. A fixture for testing DTMF decoders, Goertzel detectors, and tone analysis.

File
WAV · Dtmf · 44100 Hz
Preview of DTMF Sequence: *9#
wav
21.9 KB
Actual file preview for DTMF Sequence: *9#

DTMF Sequence: *9#

Synthetic DTMF tone sequence for digits “*9#”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Sequence: 042
wav
21.9 KB
Actual file preview for DTMF Sequence: 042

DTMF Sequence: 042

Synthetic DTMF tone sequence for digits “042”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Sequence: 1234
wav
28.2 KB
Actual file preview for DTMF Sequence: 1234

DTMF Sequence: 1234

Synthetic DTMF tone sequence for digits “1234”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Sequence: 13579
wav
34.4 KB
Actual file preview for DTMF Sequence: 13579

DTMF Sequence: 13579

Synthetic DTMF tone sequence for digits “13579”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Sequence: 567890
wav
40.7 KB
Actual file preview for DTMF Sequence: 567890

DTMF Sequence: 567890

Synthetic DTMF tone sequence for digits “567890”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Transcript: *9# (JSON)
json
224 B
Actual file preview for DTMF Transcript: *9# (JSON)

DTMF Transcript: *9# (JSON)

Ground-truth digit transcript for the DTMF sequence *9#.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Transcript: 042 (JSON)
json
224 B
Actual file preview for DTMF Transcript: 042 (JSON)

DTMF Transcript: 042 (JSON)

Ground-truth digit transcript for the DTMF sequence 042.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Transcript: 1234 (JSON)
json
227 B
Actual file preview for DTMF Transcript: 1234 (JSON)

DTMF Transcript: 1234 (JSON)

Ground-truth digit transcript for the DTMF sequence 1234.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Transcript: 13579 (JSON)
json
230 B
Actual file preview for DTMF Transcript: 13579 (JSON)

DTMF Transcript: 13579 (JSON)

Ground-truth digit transcript for the DTMF sequence 13579.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture

Loudness Ladder

P8 Convert

Peak Clip

Reference Tones

Preview of 1 kHz Sine Tone
wav
258.4 KB
Actual file preview for 1 kHz Sine Tone

1 kHz Sine Tone

A pure 1000 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono: a clean reference signal for level metering, spectrum analysis, and waveform rendering.

File
WAV · Reference Tones · 3 s
Preview of 220 Hz Sine Tone
wav
258.4 KB
Actual file preview for 220 Hz Sine Tone

220 Hz Sine Tone

A pure 220 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono: a clean reference signal for level metering, spectrum analysis, and waveform rendering.

File
WAV · Reference Tones · 3 s
Preview of 440 Hz Sine Tone (A4)
wav
258.4 KB
Actual file preview for 440 Hz Sine Tone (A4)

440 Hz Sine Tone (A4)

A pure 440 Hz sine tone, 3 seconds, 16-bit 44.1 kHz mono: a clean reference signal for level metering, spectrum analysis, and waveform rendering.

File
WAV · Reference Tones · 3 s
Preview of Frequency Sweep 20 Hz–20 kHz
wav
430.7 KB
Actual file preview for Frequency Sweep 20 Hz–20 kHz

Frequency Sweep 20 Hz–20 kHz

A 5-second linear sweep from 20 Hz to 20 kHz across the full audible range, for testing frequency response, spectrograms, and playback fidelity.

File
WAV · Reference Tones · 5 s
Preview of Pink Noise
wav
258.4 KB
Actual file preview for Pink Noise

Pink Noise

Three seconds of pink noise with a 1/f power spectrum: the standard reference for loudness and room-calibration testing.

File
WAV · Reference Tones · 3 s
Preview of White Noise
wav
258.4 KB
Actual file preview for White Noise

White Noise

Three seconds of white noise with a flat power spectrum: a reference for testing noise handling, gating, and spectral tools.

File
WAV · Reference Tones · 3 s

Sample Rate Ladder

Preview of 1 kHz Tone @ 16000 Hz Sample Rate (WAV)
wav
62.5 KB
Actual file preview for 1 kHz Tone @ 16000 Hz Sample Rate (WAV)

1 kHz Tone @ 16000 Hz Sample Rate (WAV)

A 1 kHz sine tone sampled at 16000 Hz: part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

File
WAV · Sample Rate Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 1 kHz Tone @ 44100 Hz Sample Rate (WAV)
wav
172.3 KB
Actual file preview for 1 kHz Tone @ 44100 Hz Sample Rate (WAV)

1 kHz Tone @ 44100 Hz Sample Rate (WAV)

A 1 kHz sine tone sampled at 44100 Hz: part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

File
WAV · Sample Rate Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 1 kHz Tone @ 48000 Hz Sample Rate (WAV)
wav
187.5 KB
Actual file preview for 1 kHz Tone @ 48000 Hz Sample Rate (WAV)

1 kHz Tone @ 48000 Hz Sample Rate (WAV)

A 1 kHz sine tone sampled at 48000 Hz: part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

File
WAV · Sample Rate Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 1 kHz Tone @ 8000 Hz Sample Rate (WAV)
wav
31.3 KB
Actual file preview for 1 kHz Tone @ 8000 Hz Sample Rate (WAV)

1 kHz Tone @ 8000 Hz Sample Rate (WAV)

A 1 kHz sine tone sampled at 8000 Hz: part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

File
WAV · Sample Rate Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set
Preview of 1 kHz Tone @ 96000 Hz Sample Rate (WAV)
wav
375 KB
Actual file preview for 1 kHz Tone @ 96000 Hz Sample Rate (WAV)

1 kHz Tone @ 96000 Hz Sample Rate (WAV)

A 1 kHz sine tone sampled at 96000 Hz: part of a sample-rate ladder (8/16/44.1/48/96 kHz) of the identical tone, for testing resampling, sample-rate detection, and aliasing.

File
WAV · Sample Rate Ladder · 2 s
Use case
Audio analysisConversion testing· Conversion set

Sample Rate Set

Signals

Preview of Clipped Tone: intentionally clipped (WAV)
wav
172.3 KB
Actual file preview for Clipped Tone: intentionally clipped (WAV)

Clipped Tone: intentionally clipped (WAV)

A 440 Hz tone driven past full scale and hard-clipped so its peaks are flat-topped: an intentionally clipped signal for testing clip and true-peak detection, distortion metering, and de-clipping tools.

File
WAV · Signals · 2 s
Preview of Digital Silence (2s WAV)
wav
172.3 KB
Actual file preview for Digital Silence (2s WAV)

Digital Silence (2s WAV)

Two seconds of true digital silence (all-zero samples): a fixture for testing silence detection, auto-trim thresholds, and noise-floor handling.

File
WAV · Signals · 2 s
Preview of Melody: C Major Scale (WAV)
wav
172.3 KB
Actual file preview for Melody: C Major Scale (WAV)

Melody: C Major Scale (WAV)

A two-second melody playing an ascending C-major scale (C4 to C5) with per-note envelopes: a simple musical signal for testing pitch detection, onset detection, and playback.

File
WAV · Signals · 2 s
Preview of Notification Chime (WAV)
wav
89.6 KB
Actual file preview for Notification Chime (WAV)

Notification Chime (WAV)

A short four-note notification chime (a C-major arpeggio with a decaying final note): a pleasant, recognisable UI sound for testing playback, short-clip handling, and notification pipelines.

File
WAV · Signals · 1 s
Preview of Quiet Audio Needing Normalization
wav
129.2 KB
Actual file preview for Quiet Audio Needing Normalization

Quiet Audio Needing Normalization

A very quiet modulated tone (~peak 0.04) for testing loudness normalization and gain staging.

File
WAV · Signals · 1.5 s
Preview of Stereo Channel Imbalance (440 Hz)
wav
172.3 KB
Actual file preview for Stereo Channel Imbalance (440 Hz)

Stereo Channel Imbalance (440 Hz)

A 440 Hz stereo tone with a loud left channel and a near-silent right channel, for balance and mono-mix testing.

File
WAV · Signals · 1 s

Silence Middle

Preview of Silence-in-Middle: Asym Lead
wav
46.9 KB
Actual file preview for Silence-in-Middle: Asym Lead

Silence-in-Middle: Asym Lead

440 Hz tone with 0.5s of digital silence in the middle (0.2s + silence + 0.8s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set
Preview of Silence-in-Middle: Asym Trail
wav
46.9 KB
Actual file preview for Silence-in-Middle: Asym Trail

Silence-in-Middle: Asym Trail

440 Hz tone with 0.5s of digital silence in the middle (0.8s + silence + 0.2s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set
Preview of Silence-in-Middle: Long Gap
wav
71.9 KB
Actual file preview for Silence-in-Middle: Long Gap

Silence-in-Middle: Long Gap

440 Hz tone with 1.5s of digital silence in the middle (0.4s + silence + 0.4s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set
Preview of Silence-in-Middle: Medium Gap
wav
54.7 KB
Actual file preview for Silence-in-Middle: Medium Gap

Silence-in-Middle: Medium Gap

440 Hz tone with 0.75s of digital silence in the middle (0.5s + silence + 0.5s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set
Preview of Silence-in-Middle: Short Gap
wav
32.9 KB
Actual file preview for Silence-in-Middle: Short Gap

Silence-in-Middle: Short Gap

440 Hz tone with 0.25s of digital silence in the middle (0.4s + silence + 0.4s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set
Preview of Silence-in-Middle: Tiny Gap
wav
24.4 KB
Actual file preview for Silence-in-Middle: Tiny Gap

Silence-in-Middle: Tiny Gap

440 Hz tone with 0.08s of digital silence in the middle (0.35s + silence + 0.35s). Fixture for mid-gap trim / split tools.

File
WAV · Silence Middle · 16000
Use case
Auto-trim testingAudio analysis· Conversion set

Silence Trim Set

Preview of Leading Silence 1s: 440 Hz Tone
wav
258.4 KB
Actual file preview for Leading Silence 1s: 440 Hz Tone

Leading Silence 1s: 440 Hz Tone

A 440 Hz tone preceded by exactly 1 second of digital silence. A direct fixture for auto-trim tools: the tone should start at 1.000s.

File
WAV · Silence Trim Set · 44100 Hz
Preview of Leading Silence 3s: 440 Hz Tone
wav
430.7 KB
Actual file preview for Leading Silence 3s: 440 Hz Tone

Leading Silence 3s: 440 Hz Tone

A 440 Hz tone preceded by exactly 3 seconds of digital silence. A direct fixture for auto-trim tools: the tone should start at 3.000s.

File
WAV · Silence Trim Set · 44100 Hz
Preview of Trailing Silence 1s: 440 Hz Tone
wav
258.4 KB
Actual file preview for Trailing Silence 1s: 440 Hz Tone

Trailing Silence 1s: 440 Hz Tone

A 440 Hz tone followed by exactly 1 second of digital silence. The tone should end at 2.0s: a direct fixture for testing trailing-silence trimming.

File
WAV · Silence Trim Set · 44100 Hz
Preview of Trailing Silence 3s: 440 Hz Tone
wav
430.7 KB
Actual file preview for Trailing Silence 3s: 440 Hz Tone

Trailing Silence 3s: 440 Hz Tone

A 440 Hz tone followed by exactly 3 seconds of digital silence. The tone should end at 2.0s: a direct fixture for testing trailing-silence trimming.

File
WAV · Silence Trim Set · 44100 Hz
Preview of Trailing Silence 5s: 440 Hz Tone
wav
603 KB
Actual file preview for Trailing Silence 5s: 440 Hz Tone

Trailing Silence 5s: 440 Hz Tone

A 440 Hz tone followed by exactly 5 seconds of digital silence. The tone should end at 2.0s: a direct fixture for testing trailing-silence trimming.

File
WAV · Silence Trim Set · 44100 Hz

Speech Ladders

Preview of Complaint call: chain telephone
wav
39.2 KB
Actual file preview for Complaint call: chain telephone

Complaint call: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: clipped
flac
145.6 KB
Actual file preview for Complaint call: clipped

Complaint call: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec g722 16k
wav
39.2 KB
Actual file preview for Complaint call: codec g722 16k

Complaint call: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec gsm 8k
gsm
8.1 KB
Actual file preview for Complaint call: codec gsm 8k

Complaint call: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mp3 128
mp3
79.5 KB
Actual file preview for Complaint call: codec mp3 128

Complaint call: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mp3 32
mp3
20 KB
Actual file preview for Complaint call: codec mp3 32

Complaint call: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mulaw 8k
wav
39.2 KB
Actual file preview for Complaint call: codec mulaw 8k

Complaint call: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec opus 24
opus
14.7 KB
Actual file preview for Complaint call: codec opus 24

Complaint call: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: dropouts
flac
135 KB
Actual file preview for Complaint call: dropouts

Complaint call: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: full spoken take
flac
382.3 KB
Actual file preview for Complaint call: full spoken take

Complaint call: full spoken take

The complete 13.5 second utterance at 24000 Hz mono FLAC, spoken by the bf_isabella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Complaint call: noise snr0
flac
333 KB
Actual file preview for Complaint call: noise snr0

Complaint call: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: noise snr10
flac
313.8 KB
Actual file preview for Complaint call: noise snr10

Complaint call: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: noise snr20
flac
297.4 KB
Actual file preview for Complaint call: noise snr20

Complaint call: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: plate
flac
145.7 KB
Actual file preview for Complaint call: plate

Complaint call: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 16000
flac
186.3 KB
Actual file preview for Complaint call: rate 16000

Complaint call: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 44100
flac
391.6 KB
Actual file preview for Complaint call: rate 44100

Complaint call: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 8000
flac
97.4 KB
Actual file preview for Complaint call: rate 8000

Complaint call: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: recogniser transcript
txt
183 B
Actual file preview for Complaint call: recogniser transcript

Complaint call: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Complaint call: reference script
txt
227 B
Actual file preview for Complaint call: reference script

Complaint call: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Complaint call: timed captions
srt
316 B
Actual file preview for Complaint call: timed captions

Complaint call: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Cost podcast: chain telephone
wav
39.2 KB
Actual file preview for Cost podcast: chain telephone

Cost podcast: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: clipped
flac
132.2 KB
Actual file preview for Cost podcast: clipped

Cost podcast: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec g722 16k
wav
39.2 KB
Actual file preview for Cost podcast: codec g722 16k

Cost podcast: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec gsm 8k
gsm
8.1 KB
Actual file preview for Cost podcast: codec gsm 8k

Cost podcast: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mp3 128
mp3
79.5 KB
Actual file preview for Cost podcast: codec mp3 128

Cost podcast: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mp3 32
mp3
20 KB
Actual file preview for Cost podcast: codec mp3 32

Cost podcast: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mulaw 8k
wav
39.2 KB
Actual file preview for Cost podcast: codec mulaw 8k

Cost podcast: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec opus 24
opus
15.3 KB
Actual file preview for Cost podcast: codec opus 24

Cost podcast: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: dropouts
flac
122 KB
Actual file preview for Cost podcast: dropouts

Cost podcast: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: full spoken take
flac
418 KB
Actual file preview for Cost podcast: full spoken take

Cost podcast: full spoken take

The complete 16.5 second utterance at 24000 Hz mono FLAC, spoken by the am_eric voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Cost podcast: noise snr0
flac
327.1 KB
Actual file preview for Cost podcast: noise snr0

Cost podcast: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: noise snr10
flac
306.2 KB
Actual file preview for Cost podcast: noise snr10

Cost podcast: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: noise snr20
flac
288.5 KB
Actual file preview for Cost podcast: noise snr20

Cost podcast: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: plate
flac
130.7 KB
Actual file preview for Cost podcast: plate

Cost podcast: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 16000
flac
186.2 KB
Actual file preview for Cost podcast: rate 16000

Cost podcast: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 44100
flac
332.9 KB
Actual file preview for Cost podcast: rate 44100

Cost podcast: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 8000
flac
98.9 KB
Actual file preview for Cost podcast: rate 8000

Cost podcast: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: recogniser transcript
txt
313 B
Actual file preview for Cost podcast: recogniser transcript

Cost podcast: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Cost podcast: reference script
txt
312 B
Actual file preview for Cost podcast: reference script

Cost podcast: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Cost podcast: timed captions
srt
403 B
Actual file preview for Cost podcast: timed captions

Cost podcast: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Delivery driver: chain telephone
wav
39.2 KB
Actual file preview for Delivery driver: chain telephone

Delivery driver: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: clipped
flac
148.8 KB
Actual file preview for Delivery driver: clipped

Delivery driver: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec g722 16k
wav
39.2 KB
Actual file preview for Delivery driver: codec g722 16k

Delivery driver: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec gsm 8k
gsm
8.1 KB
Actual file preview for Delivery driver: codec gsm 8k

Delivery driver: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mp3 128
mp3
79.5 KB
Actual file preview for Delivery driver: codec mp3 128

Delivery driver: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mp3 32
mp3
20 KB
Actual file preview for Delivery driver: codec mp3 32

Delivery driver: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mulaw 8k
wav
39.2 KB
Actual file preview for Delivery driver: codec mulaw 8k

Delivery driver: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec opus 24
opus
15.1 KB
Actual file preview for Delivery driver: codec opus 24

Delivery driver: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: dropouts
flac
138 KB
Actual file preview for Delivery driver: dropouts

Delivery driver: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: full spoken take
flac
361.6 KB
Actual file preview for Delivery driver: full spoken take

Delivery driver: full spoken take

The complete 13.0 second utterance at 24000 Hz mono FLAC, spoken by the am_adam voice at 1.05x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Delivery driver: noise snr0
flac
336 KB
Actual file preview for Delivery driver: noise snr0

Delivery driver: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: noise snr10
flac
316.4 KB
Actual file preview for Delivery driver: noise snr10

Delivery driver: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: noise snr20
flac
300.2 KB
Actual file preview for Delivery driver: noise snr20

Delivery driver: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: plate
flac
147.8 KB
Actual file preview for Delivery driver: plate

Delivery driver: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 16000
flac
190.9 KB
Actual file preview for Delivery driver: rate 16000

Delivery driver: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 44100
flac
389.9 KB
Actual file preview for Delivery driver: rate 44100

Delivery driver: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 8000
flac
100.8 KB
Actual file preview for Delivery driver: rate 8000

Delivery driver: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: recogniser transcript
txt
241 B
Actual file preview for Delivery driver: recogniser transcript

Delivery driver: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Delivery driver: reference script
txt
241 B
Actual file preview for Delivery driver: reference script

Delivery driver: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Delivery driver: timed captions
srt
438 B
Actual file preview for Delivery driver: timed captions

Delivery driver: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Haccp training: chain telephone
wav
39.2 KB
Actual file preview for Haccp training: chain telephone

Haccp training: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: clipped
flac
146.1 KB
Actual file preview for Haccp training: clipped

Haccp training: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec g722 16k
wav
39.2 KB
Actual file preview for Haccp training: codec g722 16k

Haccp training: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec gsm 8k
gsm
8.1 KB
Actual file preview for Haccp training: codec gsm 8k

Haccp training: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mp3 128
mp3
79.5 KB
Actual file preview for Haccp training: codec mp3 128

Haccp training: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mp3 32
mp3
20 KB
Actual file preview for Haccp training: codec mp3 32

Haccp training: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mulaw 8k
wav
39.2 KB
Actual file preview for Haccp training: codec mulaw 8k

Haccp training: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec opus 24
opus
15.4 KB
Actual file preview for Haccp training: codec opus 24

Haccp training: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: dropouts
flac
136.8 KB
Actual file preview for Haccp training: dropouts

Haccp training: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: full spoken take
flac
506.2 KB
Actual file preview for Haccp training: full spoken take

Haccp training: full spoken take

The complete 19.2 second utterance at 24000 Hz mono FLAC, spoken by the bm_george voice at 0.9x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Haccp training: noise snr0
flac
329.6 KB
Actual file preview for Haccp training: noise snr0

Haccp training: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: noise snr10
flac
309.4 KB
Actual file preview for Haccp training: noise snr10

Haccp training: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: noise snr20
flac
292.8 KB
Actual file preview for Haccp training: noise snr20

Haccp training: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: plate
flac
146 KB
Actual file preview for Haccp training: plate

Haccp training: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 16000
flac
192.2 KB
Actual file preview for Haccp training: rate 16000

Haccp training: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 44100
flac
359.5 KB
Actual file preview for Haccp training: rate 44100

Haccp training: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 8000
flac
99.5 KB
Actual file preview for Haccp training: rate 8000

Haccp training: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: recogniser transcript
txt
270 B
Actual file preview for Haccp training: recogniser transcript

Haccp training: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Haccp training: reference script
txt
289 B
Actual file preview for Haccp training: reference script

Haccp training: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Haccp training: timed captions
srt
401 B
Actual file preview for Haccp training: timed captions

Haccp training: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Homophone stress: chain telephone
wav
39.2 KB
Actual file preview for Homophone stress: chain telephone

Homophone stress: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: clipped
flac
129.5 KB
Actual file preview for Homophone stress: clipped

Homophone stress: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec g722 16k
wav
39.2 KB
Actual file preview for Homophone stress: codec g722 16k

Homophone stress: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec gsm 8k
gsm
8.1 KB
Actual file preview for Homophone stress: codec gsm 8k

Homophone stress: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mp3 128
mp3
79.5 KB
Actual file preview for Homophone stress: codec mp3 128

Homophone stress: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mp3 32
mp3
20 KB
Actual file preview for Homophone stress: codec mp3 32

Homophone stress: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mulaw 8k
wav
39.2 KB
Actual file preview for Homophone stress: codec mulaw 8k

Homophone stress: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec opus 24
opus
14.7 KB
Actual file preview for Homophone stress: codec opus 24

Homophone stress: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: dropouts
flac
119.5 KB
Actual file preview for Homophone stress: dropouts

Homophone stress: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: full spoken take
flac
327.4 KB
Actual file preview for Homophone stress: full spoken take

Homophone stress: full spoken take

The complete 13.3 second utterance at 24000 Hz mono FLAC, spoken by the bm_lewis voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Homophone stress: noise snr0
flac
325.2 KB
Actual file preview for Homophone stress: noise snr0

Homophone stress: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: noise snr10
flac
303.8 KB
Actual file preview for Homophone stress: noise snr10

Homophone stress: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: noise snr20
flac
286.6 KB
Actual file preview for Homophone stress: noise snr20

Homophone stress: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: plate
flac
129.4 KB
Actual file preview for Homophone stress: plate

Homophone stress: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 16000
flac
172.4 KB
Actual file preview for Homophone stress: rate 16000

Homophone stress: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 44100
flac
352.1 KB
Actual file preview for Homophone stress: rate 44100

Homophone stress: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 8000
flac
92.6 KB
Actual file preview for Homophone stress: rate 8000

Homophone stress: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: recogniser transcript
txt
215 B
Actual file preview for Homophone stress: recogniser transcript

Homophone stress: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Homophone stress: reference script
txt
216 B
Actual file preview for Homophone stress: reference script

Homophone stress: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Homophone stress: timed captions
srt
379 B
Actual file preview for Homophone stress: timed captions

Homophone stress: timed captions

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 5 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Invoice dispute: chain telephone
wav
39.2 KB
Actual file preview for Invoice dispute: chain telephone

Invoice dispute: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: clipped
flac
137.2 KB
Actual file preview for Invoice dispute: clipped

Invoice dispute: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec g722 16k
wav
39.2 KB
Actual file preview for Invoice dispute: codec g722 16k

Invoice dispute: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec gsm 8k
gsm
8.1 KB
Actual file preview for Invoice dispute: codec gsm 8k

Invoice dispute: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mp3 128
mp3
79.5 KB
Actual file preview for Invoice dispute: codec mp3 128

Invoice dispute: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mp3 32
mp3
20 KB
Actual file preview for Invoice dispute: codec mp3 32

Invoice dispute: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mulaw 8k
wav
39.2 KB
Actual file preview for Invoice dispute: codec mulaw 8k

Invoice dispute: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec opus 24
opus
15.1 KB
Actual file preview for Invoice dispute: codec opus 24

Invoice dispute: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: dropouts
flac
127.8 KB
Actual file preview for Invoice dispute: dropouts

Invoice dispute: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: full spoken take
flac
537.8 KB
Actual file preview for Invoice dispute: full spoken take

Invoice dispute: full spoken take

The complete 20.2 second utterance at 24000 Hz mono FLAC, spoken by the am_michael voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Invoice dispute: noise snr0
flac
323.5 KB
Actual file preview for Invoice dispute: noise snr0

Invoice dispute: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: noise snr10
flac
301.9 KB
Actual file preview for Invoice dispute: noise snr10

Invoice dispute: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: noise snr20
flac
284.6 KB
Actual file preview for Invoice dispute: noise snr20

Invoice dispute: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: plate
flac
137 KB
Actual file preview for Invoice dispute: plate

Invoice dispute: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 16000
flac
181.2 KB
Actual file preview for Invoice dispute: rate 16000

Invoice dispute: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 44100
flac
367.4 KB
Actual file preview for Invoice dispute: rate 44100

Invoice dispute: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 8000
flac
95.7 KB
Actual file preview for Invoice dispute: rate 8000

Invoice dispute: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: recogniser transcript
txt
233 B
Actual file preview for Invoice dispute: recogniser transcript

Invoice dispute: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Invoice dispute: reference script
txt
304 B
Actual file preview for Invoice dispute: reference script

Invoice dispute: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Invoice dispute: timed captions
srt
367 B
Actual file preview for Invoice dispute: timed captions

Invoice dispute: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Ivr menu: chain telephone
wav
39.2 KB
Actual file preview for Ivr menu: chain telephone

Ivr menu: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: clipped
flac
138.7 KB
Actual file preview for Ivr menu: clipped

Ivr menu: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec g722 16k
wav
39.2 KB
Actual file preview for Ivr menu: codec g722 16k

Ivr menu: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec gsm 8k
gsm
8.1 KB
Actual file preview for Ivr menu: codec gsm 8k

Ivr menu: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mp3 128
mp3
79.5 KB
Actual file preview for Ivr menu: codec mp3 128

Ivr menu: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mp3 32
mp3
20 KB
Actual file preview for Ivr menu: codec mp3 32

Ivr menu: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mulaw 8k
wav
39.2 KB
Actual file preview for Ivr menu: codec mulaw 8k

Ivr menu: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec opus 24
opus
15 KB
Actual file preview for Ivr menu: codec opus 24

Ivr menu: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: dropouts
flac
127.1 KB
Actual file preview for Ivr menu: dropouts

Ivr menu: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: full spoken take
flac
323.2 KB
Actual file preview for Ivr menu: full spoken take

Ivr menu: full spoken take

The complete 11.9 second utterance at 24000 Hz mono FLAC, spoken by the af_heart voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Ivr menu: noise snr0
flac
327.3 KB
Actual file preview for Ivr menu: noise snr0

Ivr menu: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: noise snr10
flac
306.5 KB
Actual file preview for Ivr menu: noise snr10

Ivr menu: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: noise snr20
flac
288.8 KB
Actual file preview for Ivr menu: noise snr20

Ivr menu: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: plate
flac
138.2 KB
Actual file preview for Ivr menu: plate

Ivr menu: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 16000
flac
181.7 KB
Actual file preview for Ivr menu: rate 16000

Ivr menu: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 44100
flac
376 KB
Actual file preview for Ivr menu: rate 44100

Ivr menu: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 8000
flac
96.7 KB
Actual file preview for Ivr menu: rate 8000

Ivr menu: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: recogniser transcript
txt
191 B
Actual file preview for Ivr menu: recogniser transcript

Ivr menu: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Ivr menu: reference script
txt
192 B
Actual file preview for Ivr menu: reference script

Ivr menu: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Ivr menu: timed captions
srt
289 B
Actual file preview for Ivr menu: timed captions

Ivr menu: timed captions

The recogniser's output with timings, 3 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 3 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Menu narration: chain telephone
wav
39.2 KB
Actual file preview for Menu narration: chain telephone

Menu narration: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: clipped
flac
148.1 KB
Actual file preview for Menu narration: clipped

Menu narration: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec g722 16k
wav
39.2 KB
Actual file preview for Menu narration: codec g722 16k

Menu narration: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec gsm 8k
gsm
8.1 KB
Actual file preview for Menu narration: codec gsm 8k

Menu narration: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mp3 128
mp3
79.5 KB
Actual file preview for Menu narration: codec mp3 128

Menu narration: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mp3 32
mp3
20 KB
Actual file preview for Menu narration: codec mp3 32

Menu narration: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mulaw 8k
wav
39.2 KB
Actual file preview for Menu narration: codec mulaw 8k

Menu narration: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec opus 24
opus
15.1 KB
Actual file preview for Menu narration: codec opus 24

Menu narration: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: dropouts
flac
139.1 KB
Actual file preview for Menu narration: dropouts

Menu narration: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: full spoken take
flac
425.9 KB
Actual file preview for Menu narration: full spoken take

Menu narration: full spoken take

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the af_bella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Menu narration: noise snr0
flac
329.4 KB
Actual file preview for Menu narration: noise snr0

Menu narration: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: noise snr10
flac
309.2 KB
Actual file preview for Menu narration: noise snr10

Menu narration: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: noise snr20
flac
292.6 KB
Actual file preview for Menu narration: noise snr20

Menu narration: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: plate
flac
148.1 KB
Actual file preview for Menu narration: plate

Menu narration: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 16000
flac
185.2 KB
Actual file preview for Menu narration: rate 16000

Menu narration: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 44100
flac
399.9 KB
Actual file preview for Menu narration: rate 44100

Menu narration: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 8000
flac
94.9 KB
Actual file preview for Menu narration: rate 8000

Menu narration: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: recogniser transcript
txt
211 B
Actual file preview for Menu narration: recogniser transcript

Menu narration: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Menu narration: reference script
txt
226 B
Actual file preview for Menu narration: reference script

Menu narration: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Menu narration: timed captions
srt
342 B
Actual file preview for Menu narration: timed captions

Menu narration: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Numbers stress: chain telephone
wav
39.2 KB
Actual file preview for Numbers stress: chain telephone

Numbers stress: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: clipped
flac
139.2 KB
Actual file preview for Numbers stress: clipped

Numbers stress: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec g722 16k
wav
39.2 KB
Actual file preview for Numbers stress: codec g722 16k

Numbers stress: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec gsm 8k
gsm
8.1 KB
Actual file preview for Numbers stress: codec gsm 8k

Numbers stress: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mp3 128
mp3
79.5 KB
Actual file preview for Numbers stress: codec mp3 128

Numbers stress: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mp3 32
mp3
20 KB
Actual file preview for Numbers stress: codec mp3 32

Numbers stress: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mulaw 8k
wav
39.2 KB
Actual file preview for Numbers stress: codec mulaw 8k

Numbers stress: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec opus 24
opus
15.2 KB
Actual file preview for Numbers stress: codec opus 24

Numbers stress: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: dropouts
flac
129 KB
Actual file preview for Numbers stress: dropouts

Numbers stress: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: full spoken take
flac
485.9 KB
Actual file preview for Numbers stress: full spoken take

Numbers stress: full spoken take

The complete 18.5 second utterance at 24000 Hz mono FLAC, spoken by the am_liam voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Numbers stress: noise snr0
flac
327.4 KB
Actual file preview for Numbers stress: noise snr0

Numbers stress: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: noise snr10
flac
307.6 KB
Actual file preview for Numbers stress: noise snr10

Numbers stress: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: noise snr20
flac
291.4 KB
Actual file preview for Numbers stress: noise snr20

Numbers stress: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: plate
flac
137.9 KB
Actual file preview for Numbers stress: plate

Numbers stress: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 16000
flac
187.3 KB
Actual file preview for Numbers stress: rate 16000

Numbers stress: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 44100
flac
346.6 KB
Actual file preview for Numbers stress: rate 44100

Numbers stress: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 8000
flac
98.1 KB
Actual file preview for Numbers stress: rate 8000

Numbers stress: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: recogniser transcript
txt
174 B
Actual file preview for Numbers stress: recogniser transcript

Numbers stress: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Numbers stress: reference script
txt
366 B
Actual file preview for Numbers stress: reference script

Numbers stress: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Numbers stress: timed captions
srt
312 B
Actual file preview for Numbers stress: timed captions

Numbers stress: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Reservation voicemail: chain telephone
wav
39.2 KB
Actual file preview for Reservation voicemail: chain telephone

Reservation voicemail: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: clipped
flac
151.7 KB
Actual file preview for Reservation voicemail: clipped

Reservation voicemail: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec g722 16k
wav
39.2 KB
Actual file preview for Reservation voicemail: codec g722 16k

Reservation voicemail: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec gsm 8k
gsm
8.1 KB
Actual file preview for Reservation voicemail: codec gsm 8k

Reservation voicemail: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mp3 128
mp3
79.5 KB
Actual file preview for Reservation voicemail: codec mp3 128

Reservation voicemail: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mp3 32
mp3
20 KB
Actual file preview for Reservation voicemail: codec mp3 32

Reservation voicemail: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mulaw 8k
wav
39.2 KB
Actual file preview for Reservation voicemail: codec mulaw 8k

Reservation voicemail: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec opus 24
opus
14.8 KB
Actual file preview for Reservation voicemail: codec opus 24

Reservation voicemail: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: dropouts
flac
141.9 KB
Actual file preview for Reservation voicemail: dropouts

Reservation voicemail: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: full spoken take
flac
455.7 KB
Actual file preview for Reservation voicemail: full spoken take

Reservation voicemail: full spoken take

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the bf_emma voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Reservation voicemail: noise snr0
flac
329.3 KB
Actual file preview for Reservation voicemail: noise snr0

Reservation voicemail: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: noise snr10
flac
310.4 KB
Actual file preview for Reservation voicemail: noise snr10

Reservation voicemail: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: noise snr20
flac
295.9 KB
Actual file preview for Reservation voicemail: noise snr20

Reservation voicemail: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: plate
flac
151.9 KB
Actual file preview for Reservation voicemail: plate

Reservation voicemail: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 16000
flac
187 KB
Actual file preview for Reservation voicemail: rate 16000

Reservation voicemail: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 44100
flac
400.8 KB
Actual file preview for Reservation voicemail: rate 44100

Reservation voicemail: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 8000
flac
96.9 KB
Actual file preview for Reservation voicemail: rate 8000

Reservation voicemail: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: recogniser transcript
txt
214 B
Actual file preview for Reservation voicemail: recogniser transcript

Reservation voicemail: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Reservation voicemail: reference script
txt
262 B
Actual file preview for Reservation voicemail: reference script

Reservation voicemail: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Reservation voicemail: timed captions
srt
379 B
Actual file preview for Reservation voicemail: timed captions

Reservation voicemail: timed captions

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 5 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Staff briefing: chain telephone
wav
39.2 KB
Actual file preview for Staff briefing: chain telephone

Staff briefing: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: clipped
flac
135.1 KB
Actual file preview for Staff briefing: clipped

Staff briefing: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec g722 16k
wav
39.2 KB
Actual file preview for Staff briefing: codec g722 16k

Staff briefing: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec gsm 8k
gsm
8.1 KB
Actual file preview for Staff briefing: codec gsm 8k

Staff briefing: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mp3 128
mp3
79.5 KB
Actual file preview for Staff briefing: codec mp3 128

Staff briefing: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mp3 32
mp3
20 KB
Actual file preview for Staff briefing: codec mp3 32

Staff briefing: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mulaw 8k
wav
39.2 KB
Actual file preview for Staff briefing: codec mulaw 8k

Staff briefing: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec opus 24
opus
14.1 KB
Actual file preview for Staff briefing: codec opus 24

Staff briefing: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: dropouts
flac
124.8 KB
Actual file preview for Staff briefing: dropouts

Staff briefing: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: full spoken take
flac
622.8 KB
Actual file preview for Staff briefing: full spoken take

Staff briefing: full spoken take

The complete 26.2 second utterance at 24000 Hz mono FLAC, spoken by the af_nicole voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Staff briefing: noise snr0
flac
326.3 KB
Actual file preview for Staff briefing: noise snr0

Staff briefing: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: noise snr10
flac
305.3 KB
Actual file preview for Staff briefing: noise snr10

Staff briefing: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: noise snr20
flac
288.7 KB
Actual file preview for Staff briefing: noise snr20

Staff briefing: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: plate
flac
135.5 KB
Actual file preview for Staff briefing: plate

Staff briefing: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 16000
flac
176.2 KB
Actual file preview for Staff briefing: rate 16000

Staff briefing: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 44100
flac
372.7 KB
Actual file preview for Staff briefing: rate 44100

Staff briefing: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 8000
flac
91.9 KB
Actual file preview for Staff briefing: rate 8000

Staff briefing: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: recogniser transcript
txt
252 B
Actual file preview for Staff briefing: recogniser transcript

Staff briefing: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Staff briefing: reference script
txt
280 B
Actual file preview for Staff briefing: reference script

Staff briefing: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Staff briefing: timed captions
srt
451 B
Actual file preview for Staff briefing: timed captions

Staff briefing: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Supplier order call: chain telephone
wav
39.2 KB
Actual file preview for Supplier order call: chain telephone

Supplier order call: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: clipped
flac
147.6 KB
Actual file preview for Supplier order call: clipped

Supplier order call: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec g722 16k
wav
39.2 KB
Actual file preview for Supplier order call: codec g722 16k

Supplier order call: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec gsm 8k
gsm
8.1 KB
Actual file preview for Supplier order call: codec gsm 8k

Supplier order call: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mp3 128
mp3
79.5 KB
Actual file preview for Supplier order call: codec mp3 128

Supplier order call: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mp3 32
mp3
20 KB
Actual file preview for Supplier order call: codec mp3 32

Supplier order call: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mulaw 8k
wav
39.2 KB
Actual file preview for Supplier order call: codec mulaw 8k

Supplier order call: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec opus 24
opus
15.3 KB
Actual file preview for Supplier order call: codec opus 24

Supplier order call: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: dropouts
flac
138 KB
Actual file preview for Supplier order call: dropouts

Supplier order call: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: full spoken take
flac
610.5 KB
Actual file preview for Supplier order call: full spoken take

Supplier order call: full spoken take

The complete 22.2 second utterance at 24000 Hz mono FLAC, spoken by the af_sarah voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Supplier order call: noise snr0
flac
328.6 KB
Actual file preview for Supplier order call: noise snr0

Supplier order call: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: noise snr10
flac
309.1 KB
Actual file preview for Supplier order call: noise snr10

Supplier order call: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: noise snr20
flac
293.4 KB
Actual file preview for Supplier order call: noise snr20

Supplier order call: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: plate
flac
147.5 KB
Actual file preview for Supplier order call: plate

Supplier order call: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 16000
flac
186.7 KB
Actual file preview for Supplier order call: rate 16000

Supplier order call: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 44100
flac
390.4 KB
Actual file preview for Supplier order call: rate 44100

Supplier order call: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 8000
flac
98.1 KB
Actual file preview for Supplier order call: rate 8000

Supplier order call: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: recogniser transcript
txt
332 B
Actual file preview for Supplier order call: recogniser transcript

Supplier order call: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Supplier order call: reference script
txt
381 B
Actual file preview for Supplier order call: reference script

Supplier order call: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Supplier order call: timed captions
srt
521 B
Actual file preview for Supplier order call: timed captions

Supplier order call: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set

Stereo Mono

Tagged

Preview of FLAC with Metadata Stripped
flac
94.2 KB
Actual file preview for FLAC with Metadata Stripped

FLAC with Metadata Stripped

The matching FLAC reduced to STREAMINFO and the exact same audio frames, so metadata-removal tests can compare compressed essence byte for byte.

File
FLAC · Tagged
Use case
Metadata testingConversion testing· Paired fixture
Preview of ID3-tagged MP3 (title / artist / album)
mp3
67.6 KB
Actual file preview for ID3-tagged MP3 (title / artist / album)

ID3-tagged MP3 (title / artist / album)

An MP3 with a full set of ID3v2.3 tags (title, artist, album, date, genre, comment) over a synthetic C-major melody. A fixture for testing tag readers, metadata editors, and music-library importers.

File
MP3 · Tagged · MP3 (LAME)