
ASR: 0123 Clean (WAV)
Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.
- File
- WAV · Asr · 16000
- Use case
- ASR testingAudio analysis· Paired fixture
Short synthetic digit/tone utterances with transcript JSON and clean↔noise pairs, for testing ASR loaders, WER harnesses, and audio preprocessing.

Synthetic clean tone sequence encoding digits 0123 (zero one two three). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 0123 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 0123 ASR utterance pair, expected text: “zero one two three”.

Synthetic clean tone sequence encoding digits 4567 (four five six seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 4567 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 4567 ASR utterance pair, expected text: “four five six seven”.

Synthetic clean tone sequence encoding digits 89 (eight nine). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 89 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 89 ASR utterance pair, expected text: “eight nine”.

Synthetic clean tone sequence encoding digits 1357 (one three five seven). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 1357 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 1357 ASR utterance pair, expected text: “one three five seven”.

Synthetic clean tone sequence encoding digits 24680 (two four six eight zero). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 24680 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 24680 ASR utterance pair, expected text: “two four six eight zero”.

Synthetic clean tone sequence encoding digits 987654 (nine eight seven six five four). Pair with the noisy twin and transcript JSON for ASR evaluation.

Noisy twin of the digit sequence 987654 at ~7 dB SNR. Score ASR against the shared transcript JSON.

Ground-truth transcript for the 987654 ASR utterance pair, expected text: “nine eight seven six five four”.

Short synthetic cue tone for “one”: clean reference for wake-word / digit ASR smoke tests.

Noisy twin of the “one” cue tone (~5 dB SNR).

Short synthetic cue tone for “zero”: clean reference for wake-word / digit ASR smoke tests.

Noisy twin of the “zero” cue tone (~5 dB SNR).

Extra synthetic digit utterance 1111 (one one one one): clean twin.

Noisy twin of digit sequence 1111.

Ground-truth transcript for extra ASR utterance 1111.

Extra synthetic digit utterance 2222 (two two two two): clean twin.

Noisy twin of digit sequence 2222.

Ground-truth transcript for extra ASR utterance 2222.

Extra synthetic digit utterance 3333 (three three three three): clean twin.

Noisy twin of digit sequence 3333.

Ground-truth transcript for extra ASR utterance 3333.

Extra synthetic digit utterance 4444 (four four four four): clean twin.

Noisy twin of digit sequence 4444.

Ground-truth transcript for extra ASR utterance 4444.

Extra synthetic digit utterance 5555 (five five five five): clean twin.

Noisy twin of digit sequence 5555.

Ground-truth transcript for extra ASR utterance 5555.

Extra synthetic digit utterance 6666 (six six six six): clean twin.

Noisy twin of digit sequence 6666.

Ground-truth transcript for extra ASR utterance 6666.

Extra synthetic digit utterance 7777 (seven seven seven seven): clean twin.

Noisy twin of digit sequence 7777.

Ground-truth transcript for extra ASR utterance 7777.

Extra synthetic digit utterance 8888 (eight eight eight eight): clean twin.

Noisy twin of digit sequence 8888.

Ground-truth transcript for extra ASR utterance 8888.

Extra synthetic digit utterance 9999 (nine nine nine nine): clean twin.

Noisy twin of digit sequence 9999.

Ground-truth transcript for extra ASR utterance 9999.

Extra synthetic digit utterance 0000 (zero zero zero zero): clean twin.

Noisy twin of digit sequence 0000.

Ground-truth transcript for extra ASR utterance 0000.

Synthetic DTMF tone sequence for digits “1234”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence 1234.

Synthetic DTMF tone sequence for digits “567890”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence 567890.

Synthetic DTMF tone sequence for digits “*9#”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence *9#.

Synthetic DTMF tone sequence for digits “042”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence 042.

Synthetic DTMF tone sequence for digits “13579”. Pair with the transcript JSON.

Ground-truth digit transcript for the DTMF sequence 13579.

The complete 22.2 second utterance at 24000 Hz mono FLAC, spoken by the af_sarah voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 20.2 second utterance at 24000 Hz mono FLAC, spoken by the am_michael voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the bf_emma voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 13.0 second utterance at 24000 Hz mono FLAC, spoken by the am_adam voice at 1.05x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 26.2 second utterance at 24000 Hz mono FLAC, spoken by the af_nicole voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 6 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 19.2 second utterance at 24000 Hz mono FLAC, spoken by the bm_george voice at 0.9x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 11.9 second utterance at 24000 Hz mono FLAC, spoken by the af_heart voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 3 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 16.5 second utterance at 24000 Hz mono FLAC, spoken by the am_eric voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 13.5 second utterance at 24000 Hz mono FLAC, spoken by the bf_isabella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 18.5 second utterance at 24000 Hz mono FLAC, spoken by the am_liam voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 13.3 second utterance at 24000 Hz mono FLAC, spoken by the bm_lewis voice at 1.0x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 5 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

The complete 15.6 second utterance at 24000 Hz mono FLAC, spoken by the af_bella voice at 0.95x rate. The five-second excerpt in this group's ladder is cut from this take, so anything needing real duration uses this file and anything comparing damage uses the ladder. Lossless, so it is the archival reference for the whole group.

Exactly what was spoken, which is what the synthesiser was given: numbers, currency and times are spelled out as words because that is how they were said. This is ground truth by construction - it existed before the audio did - and it is deliberately NOT the same string as the recogniser transcript paired with it.

What a speech recogniser returned for this take. It normalises: spoken "forty three dollars and eighteen cents" comes back as digits and a currency symbol. Compared with the reference script without normalising first, this disagrees on every number in the utterance and a word error rate computed that way reports a fault that does not exist. That gap is the point of the pair.

The recogniser's output with timings, 4 cues over the full take. Carries the same normalisation as the plain transcript in this group, so it inherits the same scoring trap, and adds the timing dimension: a cue that starts before the word is spoken is a different defect from a cue with the wrong words in it.

A 5 second excerpt of this group's take, reference. The clean five-second excerpt every other file in this ladder is scored against. Delivered as flac at 24000 Hz mono. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, noise. Broadband noise mixed in at a measured signal-to-noise ratio, which is where a denoiser is either working or not. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, clipping. Samples driven past a fixed threshold and flattened, the damage a hot microphone does and the one no codec can undo. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, dropouts. Short spans silenced outright, as a packet-switched call drops them, so a transcriber must decide between a pause and a missing word. Delivered as flac at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 44100 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, resample. The same audio at a different sample rate, which is the step most speech pipelines get wrong silently by assuming their model's rate. Delivered as flac at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as mp3 at 24000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as opus at 48000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as gsm at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, codec. Encoded and decoded through a real speech codec, so the artefacts are the ones the format actually produces rather than a simulation of them. Delivered as adpcm_g722 at 16000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

A 5 second excerpt of this group's take, telephone chain. Band-limited, resampled and companded in one pass: the whole telephone path applied at once rather than one effect at a time. Delivered as pcm_mulaw at 8000 Hz mono. Scored against plate.flac in the same group, which is the same five seconds undamaged. The parameters in the specification are what the file MEASURED, not what was requested of it.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

Machine-readable description of all 16 variants in this group: the file, the kind of damage, the measured parameters that produced it and what each one is expected to demonstrate. Provided so a test suite can walk the ladder programmatically instead of hard-coding filenames.

JSON Lines ASR training/eval set for the Wave B synthetic digit utterances: each row points at a clean WAV and carries the expected transcript.