Skip to content
Novus Examples

Speech recognition fixtures

Short synthetic digit/tone utterances with transcript JSON and clean↔noise pairs, for testing ASR loaders, WER harnesses, and audio preprocessing.

315 of 315 files
Preview of ASR: 0123 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 0123 Clean (WAV)

ASR: 0123 Clean (WAV)

Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 0123 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 0123 Noisy (WAV)

ASR: 0123 Noisy (WAV)

Noisy twin of the digit sequence 0123 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 4567 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 4567 Clean (WAV)

ASR: 4567 Clean (WAV)

Synthetic clean tone sequence encoding digits 4567 (four five six seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 4567 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 4567 Noisy (WAV)

ASR: 4567 Noisy (WAV)

Noisy twin of the digit sequence 4567 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 4567 Transcript (JSON)
json
246 B
Actual file preview for ASR: 4567 Transcript (JSON)

ASR: 4567 Transcript (JSON)

Ground-truth transcript for the 4567 ASR utterance pair, expected text: “four five six seven”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 89 Clean (WAV)
wav
21.3 KB
Actual file preview for ASR: 89 Clean (WAV)

ASR: 89 Clean (WAV)

Synthetic clean tone sequence encoding digits 89 (eight nine). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 89 Noisy (WAV)
wav
21.3 KB
Actual file preview for ASR: 89 Noisy (WAV)

ASR: 89 Noisy (WAV)

Noisy twin of the digit sequence 89 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Clean (WAV)
wav
37.5 KB
Actual file preview for ASR: 1357 Clean (WAV)

ASR: 1357 Clean (WAV)

Synthetic clean tone sequence encoding digits 1357 (one three five seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Noisy (WAV)
wav
37.5 KB
Actual file preview for ASR: 1357 Noisy (WAV)

ASR: 1357 Noisy (WAV)

Noisy twin of the digit sequence 1357 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 1357 Transcript (JSON)
json
244 B
Actual file preview for ASR: 1357 Transcript (JSON)

ASR: 1357 Transcript (JSON)

Ground-truth transcript for the 1357 ASR utterance pair, expected text: “one three five seven”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 24680 Clean (WAV)
wav
45.7 KB
Actual file preview for ASR: 24680 Clean (WAV)

ASR: 24680 Clean (WAV)

Synthetic clean tone sequence encoding digits 24680 (two four six eight zero). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 24680 Noisy (WAV)
wav
45.7 KB
Actual file preview for ASR: 24680 Noisy (WAV)

ASR: 24680 Noisy (WAV)

Noisy twin of the digit sequence 24680 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 24680 Transcript (JSON)
json
249 B
Actual file preview for ASR: 24680 Transcript (JSON)

ASR: 24680 Transcript (JSON)

Ground-truth transcript for the 24680 ASR utterance pair, expected text: “two four six eight zero”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR: 987654 Clean (WAV)
wav
53.8 KB
Actual file preview for ASR: 987654 Clean (WAV)

ASR: 987654 Clean (WAV)

Synthetic clean tone sequence encoding digits 987654 (nine eight seven six five four). Pair with the noisy twin and transcript JSON for ASR evaluation.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 987654 Noisy (WAV)
wav
53.8 KB
Actual file preview for ASR: 987654 Noisy (WAV)

ASR: 987654 Noisy (WAV)

Noisy twin of the digit sequence 987654 at ~7 dB SNR. Score ASR against the shared transcript JSON.

File
WAV · Asr · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of ASR: 987654 Transcript (JSON)
json
258 B
Actual file preview for ASR: 987654 Transcript (JSON)

ASR: 987654 Transcript (JSON)

Ground-truth transcript for the 987654 ASR utterance pair, expected text: “nine eight seven six five four”.

File
JSON · Asr
Use case
ASR testingJSON parsing· Conversion set
Preview of DTMF Sequence: 1234
wav
28.2 KB
Actual file preview for DTMF Sequence: 1234

DTMF Sequence: 1234

Synthetic DTMF tone sequence for digits “1234”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Transcript: 1234 (JSON)
json
227 B
Actual file preview for DTMF Transcript: 1234 (JSON)

DTMF Transcript: 1234 (JSON)

Ground-truth digit transcript for the DTMF sequence 1234.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Sequence: 567890
wav
40.7 KB
Actual file preview for DTMF Sequence: 567890

DTMF Sequence: 567890

Synthetic DTMF tone sequence for digits “567890”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Sequence: *9#
wav
21.9 KB
Actual file preview for DTMF Sequence: *9#

DTMF Sequence: *9#

Synthetic DTMF tone sequence for digits “*9#”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Transcript: *9# (JSON)
json
224 B
Actual file preview for DTMF Transcript: *9# (JSON)

DTMF Transcript: *9# (JSON)

Ground-truth digit transcript for the DTMF sequence *9#.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Sequence: 042
wav
21.9 KB
Actual file preview for DTMF Sequence: 042

DTMF Sequence: 042

Synthetic DTMF tone sequence for digits “042”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Transcript: 042 (JSON)
json
224 B
Actual file preview for DTMF Transcript: 042 (JSON)

DTMF Transcript: 042 (JSON)

Ground-truth digit transcript for the DTMF sequence 042.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of DTMF Sequence: 13579
wav
34.4 KB
Actual file preview for DTMF Sequence: 13579

DTMF Sequence: 13579

Synthetic DTMF tone sequence for digits “13579”. Pair with the transcript JSON.

File
WAV · Dtmf · 16000
Use case
ASR testingAudio analysis· Paired fixture
Preview of DTMF Transcript: 13579 (JSON)
json
230 B
Actual file preview for DTMF Transcript: 13579 (JSON)

DTMF Transcript: 13579 (JSON)

Ground-truth digit transcript for the DTMF sequence 13579.

File
JSON · Dtmf
Use case
ASR testing· Paired fixture
Preview of Supplier order call: full spoken take
flac
610.5 KB
Actual file preview for Supplier order call: full spoken take

Supplier order call: full spoken take

The complete 22.2 second utterance at 24000 Hz mono FLAC, spoken by the af_sarah voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Supplier order call: reference script
txt
381 B
Actual file preview for Supplier order call: reference script

Supplier order call: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Supplier order call: recogniser transcript
txt
332 B
Actual file preview for Supplier order call: recogniser transcript

Supplier order call: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Supplier order call: timed captions
srt
521 B
Actual file preview for Supplier order call: timed captions

Supplier order call: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Supplier order call: plate
flac
147.5 KB
Actual file preview for Supplier order call: plate

Supplier order call: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: noise snr20
flac
293.4 KB
Actual file preview for Supplier order call: noise snr20

Supplier order call: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: noise snr10
flac
309.1 KB
Actual file preview for Supplier order call: noise snr10

Supplier order call: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: noise snr0
flac
328.6 KB
Actual file preview for Supplier order call: noise snr0

Supplier order call: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: clipped
flac
147.6 KB
Actual file preview for Supplier order call: clipped

Supplier order call: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: dropouts
flac
138 KB
Actual file preview for Supplier order call: dropouts

Supplier order call: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 44100
flac
390.4 KB
Actual file preview for Supplier order call: rate 44100

Supplier order call: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 16000
flac
186.7 KB
Actual file preview for Supplier order call: rate 16000

Supplier order call: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: rate 8000
flac
98.1 KB
Actual file preview for Supplier order call: rate 8000

Supplier order call: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mp3 128
mp3
79.5 KB
Actual file preview for Supplier order call: codec mp3 128

Supplier order call: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mp3 32
mp3
20 KB
Actual file preview for Supplier order call: codec mp3 32

Supplier order call: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec opus 24
opus
15.3 KB
Actual file preview for Supplier order call: codec opus 24

Supplier order call: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec gsm 8k
gsm
8.1 KB
Actual file preview for Supplier order call: codec gsm 8k

Supplier order call: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec mulaw 8k
wav
39.2 KB
Actual file preview for Supplier order call: codec mulaw 8k

Supplier order call: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: codec g722 16k
wav
39.2 KB
Actual file preview for Supplier order call: codec g722 16k

Supplier order call: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: chain telephone
wav
39.2 KB
Actual file preview for Supplier order call: chain telephone

Supplier order call: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: full spoken take
flac
537.8 KB
Actual file preview for Invoice dispute: full spoken take

Invoice dispute: full spoken take

The complete 20.2 second utterance at 24000 Hz mono FLAC, spoken by the am_michael voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Invoice dispute: reference script
txt
304 B
Actual file preview for Invoice dispute: reference script

Invoice dispute: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Invoice dispute: recogniser transcript
txt
233 B
Actual file preview for Invoice dispute: recogniser transcript

Invoice dispute: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Invoice dispute: timed captions
srt
367 B
Actual file preview for Invoice dispute: timed captions

Invoice dispute: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Invoice dispute: plate
flac
137 KB
Actual file preview for Invoice dispute: plate

Invoice dispute: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: noise snr20
flac
284.6 KB
Actual file preview for Invoice dispute: noise snr20

Invoice dispute: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: noise snr10
flac
301.9 KB
Actual file preview for Invoice dispute: noise snr10

Invoice dispute: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: noise snr0
flac
323.5 KB
Actual file preview for Invoice dispute: noise snr0

Invoice dispute: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: clipped
flac
137.2 KB
Actual file preview for Invoice dispute: clipped

Invoice dispute: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: dropouts
flac
127.8 KB
Actual file preview for Invoice dispute: dropouts

Invoice dispute: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 44100
flac
367.4 KB
Actual file preview for Invoice dispute: rate 44100

Invoice dispute: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 16000
flac
181.2 KB
Actual file preview for Invoice dispute: rate 16000

Invoice dispute: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: rate 8000
flac
95.7 KB
Actual file preview for Invoice dispute: rate 8000

Invoice dispute: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mp3 128
mp3
79.5 KB
Actual file preview for Invoice dispute: codec mp3 128

Invoice dispute: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mp3 32
mp3
20 KB
Actual file preview for Invoice dispute: codec mp3 32

Invoice dispute: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec opus 24
opus
15.1 KB
Actual file preview for Invoice dispute: codec opus 24

Invoice dispute: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec gsm 8k
gsm
8.1 KB
Actual file preview for Invoice dispute: codec gsm 8k

Invoice dispute: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec mulaw 8k
wav
39.2 KB
Actual file preview for Invoice dispute: codec mulaw 8k

Invoice dispute: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: codec g722 16k
wav
39.2 KB
Actual file preview for Invoice dispute: codec g722 16k

Invoice dispute: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Invoice dispute: chain telephone
wav
39.2 KB
Actual file preview for Invoice dispute: chain telephone

Invoice dispute: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: full spoken take
flac
455.7 KB
Actual file preview for Reservation voicemail: full spoken take

Reservation voicemail: full spoken take

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the bf_emma voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Reservation voicemail: reference script
txt
262 B
Actual file preview for Reservation voicemail: reference script

Reservation voicemail: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Reservation voicemail: recogniser transcript
txt
214 B
Actual file preview for Reservation voicemail: recogniser transcript

Reservation voicemail: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Reservation voicemail: timed captions
srt
379 B
Actual file preview for Reservation voicemail: timed captions

Reservation voicemail: timed captions

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 5 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Reservation voicemail: plate
flac
151.9 KB
Actual file preview for Reservation voicemail: plate

Reservation voicemail: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: noise snr20
flac
295.9 KB
Actual file preview for Reservation voicemail: noise snr20

Reservation voicemail: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: noise snr10
flac
310.4 KB
Actual file preview for Reservation voicemail: noise snr10

Reservation voicemail: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: noise snr0
flac
329.3 KB
Actual file preview for Reservation voicemail: noise snr0

Reservation voicemail: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: clipped
flac
151.7 KB
Actual file preview for Reservation voicemail: clipped

Reservation voicemail: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: dropouts
flac
141.9 KB
Actual file preview for Reservation voicemail: dropouts

Reservation voicemail: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 44100
flac
400.8 KB
Actual file preview for Reservation voicemail: rate 44100

Reservation voicemail: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 16000
flac
187 KB
Actual file preview for Reservation voicemail: rate 16000

Reservation voicemail: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: rate 8000
flac
96.9 KB
Actual file preview for Reservation voicemail: rate 8000

Reservation voicemail: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mp3 128
mp3
79.5 KB
Actual file preview for Reservation voicemail: codec mp3 128

Reservation voicemail: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mp3 32
mp3
20 KB
Actual file preview for Reservation voicemail: codec mp3 32

Reservation voicemail: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec opus 24
opus
14.8 KB
Actual file preview for Reservation voicemail: codec opus 24

Reservation voicemail: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec gsm 8k
gsm
8.1 KB
Actual file preview for Reservation voicemail: codec gsm 8k

Reservation voicemail: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec mulaw 8k
wav
39.2 KB
Actual file preview for Reservation voicemail: codec mulaw 8k

Reservation voicemail: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: codec g722 16k
wav
39.2 KB
Actual file preview for Reservation voicemail: codec g722 16k

Reservation voicemail: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Reservation voicemail: chain telephone
wav
39.2 KB
Actual file preview for Reservation voicemail: chain telephone

Reservation voicemail: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: full spoken take
flac
361.6 KB
Actual file preview for Delivery driver: full spoken take

Delivery driver: full spoken take

The complete 13.0 second utterance at 24000 Hz mono FLAC, spoken by the am_adam voice at 1.05x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Delivery driver: reference script
txt
241 B
Actual file preview for Delivery driver: reference script

Delivery driver: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Delivery driver: recogniser transcript
txt
241 B
Actual file preview for Delivery driver: recogniser transcript

Delivery driver: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Delivery driver: timed captions
srt
438 B
Actual file preview for Delivery driver: timed captions

Delivery driver: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Delivery driver: plate
flac
147.8 KB
Actual file preview for Delivery driver: plate

Delivery driver: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: noise snr20
flac
300.2 KB
Actual file preview for Delivery driver: noise snr20

Delivery driver: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: noise snr10
flac
316.4 KB
Actual file preview for Delivery driver: noise snr10

Delivery driver: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: noise snr0
flac
336 KB
Actual file preview for Delivery driver: noise snr0

Delivery driver: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: clipped
flac
148.8 KB
Actual file preview for Delivery driver: clipped

Delivery driver: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: dropouts
flac
138 KB
Actual file preview for Delivery driver: dropouts

Delivery driver: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 44100
flac
389.9 KB
Actual file preview for Delivery driver: rate 44100

Delivery driver: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 16000
flac
190.9 KB
Actual file preview for Delivery driver: rate 16000

Delivery driver: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: rate 8000
flac
100.8 KB
Actual file preview for Delivery driver: rate 8000

Delivery driver: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mp3 128
mp3
79.5 KB
Actual file preview for Delivery driver: codec mp3 128

Delivery driver: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mp3 32
mp3
20 KB
Actual file preview for Delivery driver: codec mp3 32

Delivery driver: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec opus 24
opus
15.1 KB
Actual file preview for Delivery driver: codec opus 24

Delivery driver: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec gsm 8k
gsm
8.1 KB
Actual file preview for Delivery driver: codec gsm 8k

Delivery driver: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec mulaw 8k
wav
39.2 KB
Actual file preview for Delivery driver: codec mulaw 8k

Delivery driver: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: codec g722 16k
wav
39.2 KB
Actual file preview for Delivery driver: codec g722 16k

Delivery driver: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Delivery driver: chain telephone
wav
39.2 KB
Actual file preview for Delivery driver: chain telephone

Delivery driver: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: full spoken take
flac
622.8 KB
Actual file preview for Staff briefing: full spoken take

Staff briefing: full spoken take

The complete 26.2 second utterance at 24000 Hz mono FLAC, spoken by the af_nicole voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Staff briefing: reference script
txt
280 B
Actual file preview for Staff briefing: reference script

Staff briefing: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Staff briefing: recogniser transcript
txt
252 B
Actual file preview for Staff briefing: recogniser transcript

Staff briefing: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Staff briefing: timed captions
srt
451 B
Actual file preview for Staff briefing: timed captions

Staff briefing: timed captions

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 6 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Staff briefing: plate
flac
135.5 KB
Actual file preview for Staff briefing: plate

Staff briefing: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: noise snr20
flac
288.7 KB
Actual file preview for Staff briefing: noise snr20

Staff briefing: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: noise snr10
flac
305.3 KB
Actual file preview for Staff briefing: noise snr10

Staff briefing: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: noise snr0
flac
326.3 KB
Actual file preview for Staff briefing: noise snr0

Staff briefing: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: clipped
flac
135.1 KB
Actual file preview for Staff briefing: clipped

Staff briefing: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: dropouts
flac
124.8 KB
Actual file preview for Staff briefing: dropouts

Staff briefing: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 44100
flac
372.7 KB
Actual file preview for Staff briefing: rate 44100

Staff briefing: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 16000
flac
176.2 KB
Actual file preview for Staff briefing: rate 16000

Staff briefing: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: rate 8000
flac
91.9 KB
Actual file preview for Staff briefing: rate 8000

Staff briefing: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mp3 128
mp3
79.5 KB
Actual file preview for Staff briefing: codec mp3 128

Staff briefing: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mp3 32
mp3
20 KB
Actual file preview for Staff briefing: codec mp3 32

Staff briefing: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec opus 24
opus
14.1 KB
Actual file preview for Staff briefing: codec opus 24

Staff briefing: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec gsm 8k
gsm
8.1 KB
Actual file preview for Staff briefing: codec gsm 8k

Staff briefing: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec mulaw 8k
wav
39.2 KB
Actual file preview for Staff briefing: codec mulaw 8k

Staff briefing: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: codec g722 16k
wav
39.2 KB
Actual file preview for Staff briefing: codec g722 16k

Staff briefing: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Staff briefing: chain telephone
wav
39.2 KB
Actual file preview for Staff briefing: chain telephone

Staff briefing: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: full spoken take
flac
506.2 KB
Actual file preview for Haccp training: full spoken take

Haccp training: full spoken take

The complete 19.2 second utterance at 24000 Hz mono FLAC, spoken by the bm_george voice at 0.9x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Haccp training: reference script
txt
289 B
Actual file preview for Haccp training: reference script

Haccp training: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Haccp training: recogniser transcript
txt
270 B
Actual file preview for Haccp training: recogniser transcript

Haccp training: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Haccp training: timed captions
srt
401 B
Actual file preview for Haccp training: timed captions

Haccp training: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Haccp training: plate
flac
146 KB
Actual file preview for Haccp training: plate

Haccp training: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: noise snr20
flac
292.8 KB
Actual file preview for Haccp training: noise snr20

Haccp training: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: noise snr10
flac
309.4 KB
Actual file preview for Haccp training: noise snr10

Haccp training: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: noise snr0
flac
329.6 KB
Actual file preview for Haccp training: noise snr0

Haccp training: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: clipped
flac
146.1 KB
Actual file preview for Haccp training: clipped

Haccp training: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: dropouts
flac
136.8 KB
Actual file preview for Haccp training: dropouts

Haccp training: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 44100
flac
359.5 KB
Actual file preview for Haccp training: rate 44100

Haccp training: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 16000
flac
192.2 KB
Actual file preview for Haccp training: rate 16000

Haccp training: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: rate 8000
flac
99.5 KB
Actual file preview for Haccp training: rate 8000

Haccp training: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mp3 128
mp3
79.5 KB
Actual file preview for Haccp training: codec mp3 128

Haccp training: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mp3 32
mp3
20 KB
Actual file preview for Haccp training: codec mp3 32

Haccp training: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec opus 24
opus
15.4 KB
Actual file preview for Haccp training: codec opus 24

Haccp training: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec gsm 8k
gsm
8.1 KB
Actual file preview for Haccp training: codec gsm 8k

Haccp training: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec mulaw 8k
wav
39.2 KB
Actual file preview for Haccp training: codec mulaw 8k

Haccp training: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: codec g722 16k
wav
39.2 KB
Actual file preview for Haccp training: codec g722 16k

Haccp training: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Haccp training: chain telephone
wav
39.2 KB
Actual file preview for Haccp training: chain telephone

Haccp training: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: full spoken take
flac
323.2 KB
Actual file preview for Ivr menu: full spoken take

Ivr menu: full spoken take

The complete 11.9 second utterance at 24000 Hz mono FLAC, spoken by the af_heart voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Ivr menu: reference script
txt
192 B
Actual file preview for Ivr menu: reference script

Ivr menu: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Ivr menu: recogniser transcript
txt
191 B
Actual file preview for Ivr menu: recogniser transcript

Ivr menu: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Ivr menu: timed captions
srt
289 B
Actual file preview for Ivr menu: timed captions

Ivr menu: timed captions

The recogniser's output with timings, 3 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 3 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Ivr menu: plate
flac
138.2 KB
Actual file preview for Ivr menu: plate

Ivr menu: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: noise snr20
flac
288.8 KB
Actual file preview for Ivr menu: noise snr20

Ivr menu: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: noise snr10
flac
306.5 KB
Actual file preview for Ivr menu: noise snr10

Ivr menu: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: noise snr0
flac
327.3 KB
Actual file preview for Ivr menu: noise snr0

Ivr menu: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: clipped
flac
138.7 KB
Actual file preview for Ivr menu: clipped

Ivr menu: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: dropouts
flac
127.1 KB
Actual file preview for Ivr menu: dropouts

Ivr menu: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 44100
flac
376 KB
Actual file preview for Ivr menu: rate 44100

Ivr menu: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 16000
flac
181.7 KB
Actual file preview for Ivr menu: rate 16000

Ivr menu: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: rate 8000
flac
96.7 KB
Actual file preview for Ivr menu: rate 8000

Ivr menu: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mp3 128
mp3
79.5 KB
Actual file preview for Ivr menu: codec mp3 128

Ivr menu: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mp3 32
mp3
20 KB
Actual file preview for Ivr menu: codec mp3 32

Ivr menu: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec opus 24
opus
15 KB
Actual file preview for Ivr menu: codec opus 24

Ivr menu: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec gsm 8k
gsm
8.1 KB
Actual file preview for Ivr menu: codec gsm 8k

Ivr menu: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec mulaw 8k
wav
39.2 KB
Actual file preview for Ivr menu: codec mulaw 8k

Ivr menu: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: codec g722 16k
wav
39.2 KB
Actual file preview for Ivr menu: codec g722 16k

Ivr menu: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Ivr menu: chain telephone
wav
39.2 KB
Actual file preview for Ivr menu: chain telephone

Ivr menu: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: full spoken take
flac
418 KB
Actual file preview for Cost podcast: full spoken take

Cost podcast: full spoken take

The complete 16.5 second utterance at 24000 Hz mono FLAC, spoken by the am_eric voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Cost podcast: reference script
txt
312 B
Actual file preview for Cost podcast: reference script

Cost podcast: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Cost podcast: recogniser transcript
txt
313 B
Actual file preview for Cost podcast: recogniser transcript

Cost podcast: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Cost podcast: timed captions
srt
403 B
Actual file preview for Cost podcast: timed captions

Cost podcast: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Cost podcast: plate
flac
130.7 KB
Actual file preview for Cost podcast: plate

Cost podcast: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: noise snr20
flac
288.5 KB
Actual file preview for Cost podcast: noise snr20

Cost podcast: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: noise snr10
flac
306.2 KB
Actual file preview for Cost podcast: noise snr10

Cost podcast: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: noise snr0
flac
327.1 KB
Actual file preview for Cost podcast: noise snr0

Cost podcast: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: clipped
flac
132.2 KB
Actual file preview for Cost podcast: clipped

Cost podcast: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: dropouts
flac
122 KB
Actual file preview for Cost podcast: dropouts

Cost podcast: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 44100
flac
332.9 KB
Actual file preview for Cost podcast: rate 44100

Cost podcast: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 16000
flac
186.2 KB
Actual file preview for Cost podcast: rate 16000

Cost podcast: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: rate 8000
flac
98.9 KB
Actual file preview for Cost podcast: rate 8000

Cost podcast: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mp3 128
mp3
79.5 KB
Actual file preview for Cost podcast: codec mp3 128

Cost podcast: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mp3 32
mp3
20 KB
Actual file preview for Cost podcast: codec mp3 32

Cost podcast: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec opus 24
opus
15.3 KB
Actual file preview for Cost podcast: codec opus 24

Cost podcast: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec gsm 8k
gsm
8.1 KB
Actual file preview for Cost podcast: codec gsm 8k

Cost podcast: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec mulaw 8k
wav
39.2 KB
Actual file preview for Cost podcast: codec mulaw 8k

Cost podcast: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: codec g722 16k
wav
39.2 KB
Actual file preview for Cost podcast: codec g722 16k

Cost podcast: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Cost podcast: chain telephone
wav
39.2 KB
Actual file preview for Cost podcast: chain telephone

Cost podcast: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: full spoken take
flac
382.3 KB
Actual file preview for Complaint call: full spoken take

Complaint call: full spoken take

The complete 13.5 second utterance at 24000 Hz mono FLAC, spoken by the bf_isabella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Complaint call: reference script
txt
227 B
Actual file preview for Complaint call: reference script

Complaint call: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Complaint call: recogniser transcript
txt
183 B
Actual file preview for Complaint call: recogniser transcript

Complaint call: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Complaint call: timed captions
srt
316 B
Actual file preview for Complaint call: timed captions

Complaint call: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Complaint call: plate
flac
145.7 KB
Actual file preview for Complaint call: plate

Complaint call: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: noise snr20
flac
297.4 KB
Actual file preview for Complaint call: noise snr20

Complaint call: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: noise snr10
flac
313.8 KB
Actual file preview for Complaint call: noise snr10

Complaint call: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: noise snr0
flac
333 KB
Actual file preview for Complaint call: noise snr0

Complaint call: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: clipped
flac
145.6 KB
Actual file preview for Complaint call: clipped

Complaint call: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: dropouts
flac
135 KB
Actual file preview for Complaint call: dropouts

Complaint call: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 44100
flac
391.6 KB
Actual file preview for Complaint call: rate 44100

Complaint call: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 16000
flac
186.3 KB
Actual file preview for Complaint call: rate 16000

Complaint call: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: rate 8000
flac
97.4 KB
Actual file preview for Complaint call: rate 8000

Complaint call: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mp3 128
mp3
79.5 KB
Actual file preview for Complaint call: codec mp3 128

Complaint call: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mp3 32
mp3
20 KB
Actual file preview for Complaint call: codec mp3 32

Complaint call: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec opus 24
opus
14.7 KB
Actual file preview for Complaint call: codec opus 24

Complaint call: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec gsm 8k
gsm
8.1 KB
Actual file preview for Complaint call: codec gsm 8k

Complaint call: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec mulaw 8k
wav
39.2 KB
Actual file preview for Complaint call: codec mulaw 8k

Complaint call: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: codec g722 16k
wav
39.2 KB
Actual file preview for Complaint call: codec g722 16k

Complaint call: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Complaint call: chain telephone
wav
39.2 KB
Actual file preview for Complaint call: chain telephone

Complaint call: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: full spoken take
flac
485.9 KB
Actual file preview for Numbers stress: full spoken take

Numbers stress: full spoken take

The complete 18.5 second utterance at 24000 Hz mono FLAC, spoken by the am_liam voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Numbers stress: reference script
txt
366 B
Actual file preview for Numbers stress: reference script

Numbers stress: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Numbers stress: recogniser transcript
txt
174 B
Actual file preview for Numbers stress: recogniser transcript

Numbers stress: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Numbers stress: timed captions
srt
312 B
Actual file preview for Numbers stress: timed captions

Numbers stress: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Numbers stress: plate
flac
137.9 KB
Actual file preview for Numbers stress: plate

Numbers stress: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: noise snr20
flac
291.4 KB
Actual file preview for Numbers stress: noise snr20

Numbers stress: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: noise snr10
flac
307.6 KB
Actual file preview for Numbers stress: noise snr10

Numbers stress: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: noise snr0
flac
327.4 KB
Actual file preview for Numbers stress: noise snr0

Numbers stress: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: clipped
flac
139.2 KB
Actual file preview for Numbers stress: clipped

Numbers stress: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: dropouts
flac
129 KB
Actual file preview for Numbers stress: dropouts

Numbers stress: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 44100
flac
346.6 KB
Actual file preview for Numbers stress: rate 44100

Numbers stress: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 16000
flac
187.3 KB
Actual file preview for Numbers stress: rate 16000

Numbers stress: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: rate 8000
flac
98.1 KB
Actual file preview for Numbers stress: rate 8000

Numbers stress: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mp3 128
mp3
79.5 KB
Actual file preview for Numbers stress: codec mp3 128

Numbers stress: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mp3 32
mp3
20 KB
Actual file preview for Numbers stress: codec mp3 32

Numbers stress: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec opus 24
opus
15.2 KB
Actual file preview for Numbers stress: codec opus 24

Numbers stress: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec gsm 8k
gsm
8.1 KB
Actual file preview for Numbers stress: codec gsm 8k

Numbers stress: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec mulaw 8k
wav
39.2 KB
Actual file preview for Numbers stress: codec mulaw 8k

Numbers stress: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: codec g722 16k
wav
39.2 KB
Actual file preview for Numbers stress: codec g722 16k

Numbers stress: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Numbers stress: chain telephone
wav
39.2 KB
Actual file preview for Numbers stress: chain telephone

Numbers stress: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: full spoken take
flac
327.4 KB
Actual file preview for Homophone stress: full spoken take

Homophone stress: full spoken take

The complete 13.3 second utterance at 24000 Hz mono FLAC, spoken by the bm_lewis voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Homophone stress: reference script
txt
216 B
Actual file preview for Homophone stress: reference script

Homophone stress: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Homophone stress: recogniser transcript
txt
215 B
Actual file preview for Homophone stress: recogniser transcript

Homophone stress: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Homophone stress: timed captions
srt
379 B
Actual file preview for Homophone stress: timed captions

Homophone stress: timed captions

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 5 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Homophone stress: plate
flac
129.4 KB
Actual file preview for Homophone stress: plate

Homophone stress: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: noise snr20
flac
286.6 KB
Actual file preview for Homophone stress: noise snr20

Homophone stress: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: noise snr10
flac
303.8 KB
Actual file preview for Homophone stress: noise snr10

Homophone stress: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: noise snr0
flac
325.2 KB
Actual file preview for Homophone stress: noise snr0

Homophone stress: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: clipped
flac
129.5 KB
Actual file preview for Homophone stress: clipped

Homophone stress: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: dropouts
flac
119.5 KB
Actual file preview for Homophone stress: dropouts

Homophone stress: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 44100
flac
352.1 KB
Actual file preview for Homophone stress: rate 44100

Homophone stress: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 16000
flac
172.4 KB
Actual file preview for Homophone stress: rate 16000

Homophone stress: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: rate 8000
flac
92.6 KB
Actual file preview for Homophone stress: rate 8000

Homophone stress: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mp3 128
mp3
79.5 KB
Actual file preview for Homophone stress: codec mp3 128

Homophone stress: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mp3 32
mp3
20 KB
Actual file preview for Homophone stress: codec mp3 32

Homophone stress: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec opus 24
opus
14.7 KB
Actual file preview for Homophone stress: codec opus 24

Homophone stress: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec gsm 8k
gsm
8.1 KB
Actual file preview for Homophone stress: codec gsm 8k

Homophone stress: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec mulaw 8k
wav
39.2 KB
Actual file preview for Homophone stress: codec mulaw 8k

Homophone stress: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: codec g722 16k
wav
39.2 KB
Actual file preview for Homophone stress: codec g722 16k

Homophone stress: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Homophone stress: chain telephone
wav
39.2 KB
Actual file preview for Homophone stress: chain telephone

Homophone stress: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: full spoken take
flac
425.9 KB
Actual file preview for Menu narration: full spoken take

Menu narration: full spoken take

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the af_bella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis· Conversion set
Preview of Menu narration: reference script
txt
226 B
Actual file preview for Menu narration: reference script

Menu narration: reference script

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Menu narration: recogniser transcript
txt
211 B
Actual file preview for Menu narration: recogniser transcript

Menu narration: recogniser transcript

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

File
TXT · Speech Ladders
Use case
ASR testing· Paired fixture
Preview of Menu narration: timed captions
srt
342 B
Actual file preview for Menu narration: timed captions

Menu narration: timed captions

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

File
SRT · Speech Ladders · 4 cues
Use case
ASR testingMedia accessibility· Conversion set
Preview of Menu narration: plate
flac
148.1 KB
Actual file preview for Menu narration: plate

Menu narration: plate

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: noise snr20
flac
292.6 KB
Actual file preview for Menu narration: noise snr20

Menu narration: noise snr20

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: noise snr10
flac
309.2 KB
Actual file preview for Menu narration: noise snr10

Menu narration: noise snr10

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: noise snr0
flac
329.4 KB
Actual file preview for Menu narration: noise snr0

Menu narration: noise snr0

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: clipped
flac
148.1 KB
Actual file preview for Menu narration: clipped

Menu narration: clipped

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: dropouts
flac
139.1 KB
Actual file preview for Menu narration: dropouts

Menu narration: dropouts

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 44100
flac
399.9 KB
Actual file preview for Menu narration: rate 44100

Menu narration: rate 44100

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 16000
flac
185.2 KB
Actual file preview for Menu narration: rate 16000

Menu narration: rate 16000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: rate 8000
flac
94.9 KB
Actual file preview for Menu narration: rate 8000

Menu narration: rate 8000

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
FLAC · Speech Ladders · flac
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mp3 128
mp3
79.5 KB
Actual file preview for Menu narration: codec mp3 128

Menu narration: codec mp3 128

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mp3 32
mp3
20 KB
Actual file preview for Menu narration: codec mp3 32

Menu narration: codec mp3 32

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
MP3 · Speech Ladders · mp3
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec opus 24
opus
15.1 KB
Actual file preview for Menu narration: codec opus 24

Menu narration: codec opus 24

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
OPUS · Speech Ladders · opus
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec gsm 8k
gsm
8.1 KB
Actual file preview for Menu narration: codec gsm 8k

Menu narration: codec gsm 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
GSM · Speech Ladders · gsm
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec mulaw 8k
wav
39.2 KB
Actual file preview for Menu narration: codec mulaw 8k

Menu narration: codec mulaw 8k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: codec g722 16k
wav
39.2 KB
Actual file preview for Menu narration: codec g722 16k

Menu narration: codec g722 16k

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · adpcm_g722
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Menu narration: chain telephone
wav
39.2 KB
Actual file preview for Menu narration: chain telephone

Menu narration: chain telephone

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

File
WAV · Speech Ladders · pcm_mulaw
Use case
ASR testingAudio analysis+1· Conversion set
Preview of Supplier order call: ladder answer key
json
4.3 KB
Actual file preview for Supplier order call: ladder answer key

Supplier order call: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Invoice dispute: ladder answer key
json
4.3 KB
Actual file preview for Invoice dispute: ladder answer key

Invoice dispute: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Reservation voicemail: ladder answer key
json
4.3 KB
Actual file preview for Reservation voicemail: ladder answer key

Reservation voicemail: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Delivery driver: ladder answer key
json
4.3 KB
Actual file preview for Delivery driver: ladder answer key

Delivery driver: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Staff briefing: ladder answer key
json
4.3 KB
Actual file preview for Staff briefing: ladder answer key

Staff briefing: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Haccp training: ladder answer key
json
4.3 KB
Actual file preview for Haccp training: ladder answer key

Haccp training: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Ivr menu: ladder answer key
json
4.3 KB
Actual file preview for Ivr menu: ladder answer key

Ivr menu: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Cost podcast: ladder answer key
json
4.3 KB
Actual file preview for Cost podcast: ladder answer key

Cost podcast: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Complaint call: ladder answer key
json
4.3 KB
Actual file preview for Complaint call: ladder answer key

Complaint call: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Numbers stress: ladder answer key
json
4.3 KB
Actual file preview for Numbers stress: ladder answer key

Numbers stress: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Homophone stress: ladder answer key
json
4.3 KB
Actual file preview for Homophone stress: ladder answer key

Homophone stress: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of Menu narration: ladder answer key
json
4.3 KB
Actual file preview for Menu narration: ladder answer key

Menu narration: ladder answer key

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

File
JSON · Speech Ladders
Use case
ASR testingJSON parsing· Conversion set
Preview of ASR Digit Utterances Dataset (JSONL)
jsonl
764 B
Actual file preview for ASR Digit Utterances Dataset (JSONL)

ASR Digit Utterances Dataset (JSONL)

JSON Lines ASR training/eval set for the Wave B synthetic digit utterances: each row points at a clean WAV and carries the expected transcript.

File
JSONL · Speech · 8 records