Skip to content
Testaroo

Explore the test library

Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.

444 results

Page 10 of 19; 24 results per page.

Show the canonical directory
Preview of Interpolation Input: pan-city, Motion-Blurred Decimation
mp4
27 KB
Actual file preview for Interpolation Input: pan-city, Motion-Blurred Decimation

Interpolation Input: pan-city, Motion-Blurred Decimation

Halved frame rate where each output frame is the AVERAGE of the two it replaces, rather than one of them: what a long shutter angle actually produces. Substantially harder than clean decimation, because the interpolator must undo motion blur as well as invent the missing instants, and no input frame matches any ground-truth frame exactly.

File
MP4 · Interp · 640x360
Use case
Frame interpolationVideo deblur+1· Conversion set
Preview of JSON: Caption Cue List
json
876 B
Actual file preview for JSON: Caption Cue List

JSON: Caption Cue List

The same five cues as a plain JSON array with float second timings: the shape most caption pipelines use internally between parsing one format and writing another. Handy as the expected intermediate when testing a converter, since it removes timestamp-formatting differences from the comparison.

File
JSON · Captions Formats · 5 cues
Use case
Subtitle parsingConversion testing+1· Conversion set
Preview of JSON: Timed-Text Suite Index
json
17.1 KB
Actual file preview for JSON: Timed-Text Suite Index

JSON: Timed-Text Suite Index

A machine-readable index of all 67 timed-text, streaming-manifest and ad-signalling fixtures in this wave, with id, format, path and byte size for each. Useful as a work-list when running a parser across the whole suite, and as a manifest to diff against after regenerating.

File
JSON · Index
Preview of JSON: Video QC Note 09
json
100 B
Actual file preview for JSON: Video QC Note 09

JSON: Video QC Note 09

Video QC companion SAMPLE (av-sync-notes.json) for Wave G suites.

File
JSON · Notes
Preview of LRC: Enhanced Word-Level Timing
lrc
309 B
Actual file preview for LRC: Enhanced Word-Level Timing

LRC: Enhanced Word-Level Timing

Enhanced LRC with per-word timings in angle brackets alongside the usual per-line timestamps: the format karaoke and lyric-sync apps consume. Includes the standard metadata tags, an offset field, and a final empty timestamp that clears the display. Simple LRC parsers read only the line timings and silently render the word markers as visible text.

File
LRC · Captions Edge
Use case
Subtitle parsingConversion testing+1· Conversion set
Preview of M4V: iTunes H.264 Clip
m4v
12.2 KB
Actual file preview for M4V: iTunes H.264 Clip

M4V: iTunes H.264 Clip

The clip as M4V: Apple's MP4 variant used by iTunes. Browser-playable; for testing M4V handling and M4V↔MP4 conversion.

File
M4V · M4v · 480x270
Preview of Market vendor sorting produce
mp4
424.2 KB
Actual file preview for Market vendor sorting produce

Market vendor sorting produce

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-p-market-sen_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0563 (active). Substantial frame-to-frame change, the end of the range where naive scene detection starts firing. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 640 × 480 px
Preview of Market walk: 24-second continuous film
mp4
2.3 MB
Actual file preview for Market walk: 24-second continuous film

Market walk: 24-second continuous film

A 24 second 512x288 H.264 film at 24 fps, 576 frames, written with the moov atom first so it can be seeked before it has finished downloading. Twelve generated segments crossfaded into one continuous take. Across 575 transitions the mean inter-frame change is 0.0217 and the strongest single transition is 0.2597, with 0 above 0.3. A deliberate hard cut between two unrelated films measures 0.311 by the same method, so the worst join here sits 0.0513 below a real cut. At eleven times the length of any other clip in this library, it is the only footage here long enough to seek through, chapter, or scrub. Stylised rather than photographic: this is the 1.3B model at 512x288, the subject drifts over 24 seconds, and a dark rendering artefact recurs in a few frames. It is a fixture for duration, seeking and stitch measurement, not showcase footage.

File
MP4 · Generated Films · 512 × 288 px
Use case
Video QAMedia handling· Conversion set
Preview of Market walk: chapter track
vtt
831 B
Actual file preview for Market walk: chapter track

Market walk: chapter track

A WebVTT chapter track for the 24 second film in this group, with 12 cues at 2.125 second intervals. The cues are not decoration: each one marks a boundary between two of the twelve generated segments, verified against the measured transitions, so seeking to a cue lands on a join. Chapter tracks over two-second clips have nowhere to seek to; this is the first in this library with somewhere to go.

File
VTT · Generated Films · 12 cues
Use case
Media accessibilitySubtitle parsing· Conversion set
Preview of Market walk: scrub-preview index
vtt
2.1 KB
Actual file preview for Market walk: scrub-preview index

Market walk: scrub-preview index

The WebVTT index for the sprite sheet in this group: 24 cues, one per second, each naming a rectangle of the sheet with an xywh media fragment. The sheet and this index are each useless alone, which is why they are paired.

File
VTT · Generated Films · 24 cues
Use case
Subtitle parsingMedia handling· Paired fixture
Preview of Matting Ground Truth: Per-Frame Alpha Matte
mp4
32.9 KB
Actual file preview for Matting Ground Truth: Per-Frame Alpha Matte

Matting Ground Truth: Per-Frame Alpha Matte

The exact per-frame alpha used to composite every clip in this group, as an 8-bit greyscale clip: white is opaque subject, black is background, and edges carry genuine intermediate values because the matte is anti-aliased rather than binary. This is what makes matting output measurable. Compare a predicted alpha against this frame by frame instead of inspecting a composite and forming an opinion.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Matting Input: Subject on Busy patterned
mp4
28.7 KB
Actual file preview for Matting Input: Subject on Busy patterned

Matting Input: Subject on Busy patterned

The same subject and the same alpha, composited over busy patterned: no chroma screen, which is the case a general background-removal model actually has to handle. Score the predicted alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Matting Input: Subject on Chroma Green
mp4
26 KB
Actual file preview for Matting Input: Subject on Chroma Green

Matting Input: Subject on Chroma Green

The canonical keying setup: the subject over a chroma-green field carrying a deliberate vertical lighting falloff, because a perfectly flat key is unrealistically easy. Pull a key, then score the resulting alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo QA· Conversion set
Preview of Matting Input: Subject on Colour close to subject
mp4
9.4 KB
Actual file preview for Matting Input: Subject on Colour close to subject

Matting Input: Subject on Colour close to subject

The same subject and the same alpha, composited over colour close to subject: no chroma screen, which is the case a general background-removal model actually has to handle. The background colour deliberately sits close to the subject's own, so colour alone cannot separate them and the model must use shape and motion. Score the predicted alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Matting Input: Subject on Dark studio
mp4
17 KB
Actual file preview for Matting Input: Subject on Dark studio

Matting Input: Subject on Dark studio

The same subject and the same alpha, composited over dark studio: no chroma screen, which is the case a general background-removal model actually has to handle. Score the predicted alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Matting Input: Subject on Office interior
mp4
33 KB
Actual file preview for Matting Input: Subject on Office interior

Matting Input: Subject on Office interior

The same subject and the same alpha, composited over office interior: no chroma screen, which is the case a general background-removal model actually has to handle. Score the predicted alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Matting Input: Subject on Outdoor daylight
mp4
15.7 KB
Actual file preview for Matting Input: Subject on Outdoor daylight

Matting Input: Subject on Outdoor daylight

The same subject and the same alpha, composited over outdoor daylight: no chroma screen, which is the case a general background-removal model actually has to handle. Score the predicted alpha against the ground-truth matte in this group.

File
MP4 · Matting · 640x360
Use case
Video mattingVideo segmentation+1· Conversion set
Preview of Mist drifting through a mountain valley
mp4
166 KB
Actual file preview for Mist drifting through a mountain valley

Mist drifting through a mountain valley

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-l-mountain_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0070 (subtle). Motion is deliberate but slight, closer to a living photograph than to action. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 640 × 480 px
Preview of Mist drifting through a mountain valley: degraded input
mp4
14 KB
Actual file preview for Mist drifting through a mountain valley: degraded input

Mist drifting through a mountain valley: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/mountain.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 34.359 dB PSNR and 0.915 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Mist drifting through a mountain valley: restored by Real-ESRGAN
mp4
134.8 KB
Actual file preview for Mist drifting through a mountain valley: restored by Real-ESRGAN

Mist drifting through a mountain valley: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/mountain.mp4. Measured with ffmpeg's own filters it scores 32.503 dB PSNR and 0.92 SSIM; the same degraded input enlarged by plain bicubic scores 34.359 dB and 0.915, so the two metrics DISAGREE: -1.86 dB of PSNR against it, +0.0050 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of MKV: Matroska H.264 Clip
mkv
11.5 KB
Actual file preview for MKV: Matroska H.264 Clip

MKV: Matroska H.264 Clip

The clip in a Matroska (MKV) container with H.264 video: the flexible open container used for rich multi-track media. For testing MKV demuxing and remux/conversion.

File
MKV · Mkv · 480x270
Preview of MOV: QuickTime H.264 Clip
mov
12.1 KB
Actual file preview for MOV: QuickTime H.264 Clip

MOV: QuickTime H.264 Clip

The clip in a QuickTime (MOV) container with H.264 video: Apple's container, common from cameras and editors. For testing MOV parsing and MOV→MP4 conversion.

File
MOV · Mov · 480x270
Preview of MP4 Layout: Faststart (moov Atom First)
mp4
23.2 KB
Actual file preview for MP4 Layout: Faststart (moov Atom First)

MP4 Layout: Faststart (moov Atom First)

The moov index is relocated to the front of the file, so a player can begin playback after the first few kilobytes. Required for progressive download to work at all. Byte-for-byte the same encode as its twin in this group; only the atom order differs, which is why comparing the two is the clean way to demonstrate the effect.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of MP4 Layout: Fragmented (fMP4 / CMAF)
mp4
23.2 KB
Actual file preview for MP4 Layout: Fragmented (fMP4 / CMAF)

MP4 Layout: Fragmented (fMP4 / CMAF)

Fragmented MP4: an empty moov followed by independent moof/mdat fragment pairs, rather than one monolithic index. This is what CMAF streaming actually delivers, and what makes a segment playable without the rest of the file. Parsers written against progressive MP4 frequently fail here, because there is no sample table to read up front.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set