Skip to content
Testaroo

Explore the test library

Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.

444 results

Page 2 of 19; 24 results per page.

Show the canonical directory
Preview of Bitstream: All-Intra (Every Frame a Keyframe)
mp4
119.8 KB
Actual file preview for Bitstream: All-Intra (Every Frame a Keyframe)

Bitstream: All-Intra (Every Frame a Keyframe)

Every frame is an I-frame, so any frame is an independent cut point. Much larger, and the shape editing and frame-accurate seeking want. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Bitstream: Heavy B-Frames (8 Consecutive)
mp4
22.8 KB
Actual file preview for Bitstream: Heavy B-Frames (8 Consecutive)

Bitstream: Heavy B-Frames (8 Consecutive)

Eight consecutive B-frames with a B-pyramid, so decode order and display order diverge sharply and DTS runs well behind PTS. The case that breaks naive timestamp handling and any code that assumes frames arrive in display order. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Bitstream: Long GOP (24 Frames, One Per Second)
mp4
23.2 KB
Actual file preview for Bitstream: Long GOP (24 Frames, One Per Second)

Bitstream: Long GOP (24 Frames, One Per Second)

One keyframe per second: the delivery default, and the setting that determines HLS/DASH segment boundaries, since a segment must start on a keyframe. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Bitstream: No B-Frames
mp4
29.5 KB
Actual file preview for Bitstream: No B-Frames

Bitstream: No B-Frames

I and P frames only. Required by some low-latency and legacy decoders, and it removes the reordering that makes decode order differ from display order. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Bitstream: Short GOP (8 Frames)
mp4
35.9 KB
Actual file preview for Bitstream: Short GOP (8 Frames)

Bitstream: Short GOP (8 Frames)

A keyframe every 8 frames with scene-cut detection disabled, so the GOP length is exactly what it says. Short GOPs cost bitrate but bound seek latency: the trade-off streaming packagers make explicitly. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Bitstream: Single GOP (One Keyframe Only)
mp4
23.2 KB
Actual file preview for Bitstream: Single GOP (One Keyframe Only)

Bitstream: Single GOP (One Keyframe Only)

Exactly one keyframe, at the start. Maximally efficient and nearly unseekable: a player must decode from frame zero to reach any position, which is what makes long-GOP archives painful to scrub. Same picture and same codec as every other clip in this group; only the GOP and frame-type structure differ, so the effect on size and seekability is directly attributable.

File
MP4 · Structure · 480x270
Use case
Video codecsConversion testing+1· Conversion set
Preview of Cat on a windowsill
mp4
207 KB
Actual file preview for Cat on a windowsill

Cat on a windowsill

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-n-cat-window_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0182 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 640 × 480 px
Preview of Cat on a windowsill: degraded input
mp4
20.1 KB
Actual file preview for Cat on a windowsill: degraded input

Cat on a windowsill: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/cat-window.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 32.598 dB PSNR and 0.917 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Cat on a windowsill: restored by Real-ESRGAN
mp4
225.7 KB
Actual file preview for Cat on a windowsill: restored by Real-ESRGAN

Cat on a windowsill: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/cat-window.mp4. Measured with ffmpeg's own filters it scores 31.398 dB PSNR and 0.932 SSIM; the same degraded input enlarged by plain bicubic scores 32.598 dB and 0.917, so the two metrics DISAGREE: -1.20 dB of PSNR against it, +0.0150 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Chapters: Embedded in Matroska (MKV)
mkv
25.8 KB
Actual file preview for Chapters: Embedded in Matroska (MKV)

Chapters: Embedded in Matroska (MKV)

Four named chapters (Cold Open, Titles, Main Segment, Credits) on two-second boundaries. The clip burns the chapter name and a running timecode into the picture, so the chapter marks can be verified by eye against what the player's chapter menu claims. Matroska stores chapters as a proper EditionEntry structure with nanosecond timestamps and per-language names, and every desktop player exposes them. This is the reference case.

File
MKV · Chapters · 8s
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: Embedded in MP4
mp4
27.7 KB
Actual file preview for Chapters: Embedded in MP4

Chapters: Embedded in MP4

Four named chapters (Cold Open, Titles, Main Segment, Credits) on two-second boundaries. The clip burns the chapter name and a running timecode into the picture, so the chapter marks can be verified by eye against what the player's chapter menu claims. MP4 has no chapter box. FFmpeg writes them as a QuickTime chapter track (a hidden text track referenced by the video track), which QuickTime, VLC and most desktop players read and which browsers ignore completely. Remux this to MKV and back and the chapters usually survive; convert it with a tool that maps streams individually and they usually do not.

File
MP4 · Chapters · 8s
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: FFMETADATA1 Chapter File
txt
339 B
Actual file preview for Chapters: FFMETADATA1 Chapter File

Chapters: FFMETADATA1 Chapter File

FFmpeg's own metadata format, and the one used to build the two embedded-chapter files in this group. START and END are integers in the declared TIMEBASE, which is the detail that trips people up: change TIMEBASE and every timestamp silently means something else. This is the exact file the MKV and MP4 chapter fixtures here were built from.

File
TXT · Chapters
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: Matroska Chapter XML
xml
1.5 KB
Actual file preview for Chapters: Matroska Chapter XML

Chapters: Matroska Chapter XML

The XML dialect `mkvmerge` and `mkvpropedit` read to write chapters into a Matroska file without re-muxing it. Timestamps are nanosecond-precision `HH:MM:SS.nnnnnnnnn`, and each ChapterAtom carries a UID that must be stable across edits: the field most hand-written generators omit, which is why re-running them renumbers every chapter.

File
XML · Chapters
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: None (MKV Control)
mkv
25.5 KB
Actual file preview for Chapters: None (MKV Control)

Chapters: None (MKV Control)

The same clip with the chapter marks explicitly removed. The control for this group, and the file that tells you whether a chapter reader returns an empty list or reports the whole-file duration as a single unnamed chapter: both are common, and they are not the same answer.

File
MKV · Chapters · 8s
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: Timestamp List in a Description
txt
110 B
Actual file preview for Chapters: Timestamp List in a Description

Chapters: Timestamp List in a Description

The plain-text convention used in video descriptions: `M:SS Title`, one per line, first entry at zero. There is no specification, so every parser guesses, at leading zeros, at `HH:MM:SS` versus `M:SS`, at whether a dash or a dot separates the timestamp from the title. Useful for testing a scraper against the format people actually paste.

File
TXT · Chapters
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chapters: WebVTT Chapter Track
vtt
267 B
Actual file preview for Chapters: WebVTT Chapter Track

Chapters: WebVTT Chapter Track

A WebVTT file intended for `<track kind="chapters">` rather than `kind="subtitles"`. Structurally it is an ordinary WebVTT file, which is the point, nothing inside it says 'chapters', so the same bytes behave as captions if the kind attribute is wrong, and the chapter titles are rendered over the video as subtitles. A common and very visible mistake.

File
VTT · Chapters
Use case
Media metadataVideo QA+1· Conversion set
Preview of Chef plating a dish
mp4
265.3 KB
Actual file preview for Chef plating a dish

Chef plating a dish

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-p-kitchen-chef_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0287 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 640 × 480 px
Preview of Chef plating a dish: degraded input
mp4
36.2 KB
Actual file preview for Chef plating a dish: degraded input

Chef plating a dish: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/kitchen-chef.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 27.165 dB PSNR and 0.861 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Chef plating a dish: restored by Real-ESRGAN
mp4
329.2 KB
Actual file preview for Chef plating a dish: restored by Real-ESRGAN

Chef plating a dish: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/kitchen-chef.mp4. Measured with ffmpeg's own filters it scores 28.023 dB PSNR and 0.906 SSIM; the same degraded input enlarged by plain bicubic scores 27.165 dB and 0.861, so it beats bicubic on both, by +0.86 dB and +0.0450 SSIM. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of City street at golden hour (text-to-video)
mp4
108.9 KB
Actual file preview for City street at golden hour (text-to-video)

City street at golden hour (text-to-video)

A 2.1 second 512x288 H.264 clip at 8 fps, 17 frames. Generated from a text prompt alone, kept as the low-resolution, non-standard-frame-rate case that decoder tests need and photographic clips do not provide. Measured mean inter-frame change 0.0174 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 512 × 288 px
Preview of Clouds drifting behind a fantasy keep
mp4
164.1 KB
Actual file preview for Clouds drifting behind a fantasy keep

Clouds drifting behind a fantasy keep

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-g-fantasy-keep_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0051 (subtle). Motion is deliberate but slight, closer to a living photograph than to action. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

File
MP4 · Generated Clips · 640 × 480 px
Preview of Clouds drifting behind a fantasy keep: degraded input
mp4
15.5 KB
Actual file preview for Clouds drifting behind a fantasy keep: degraded input

Clouds drifting behind a fantasy keep: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/game-keep.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 34.952 dB PSNR and 0.906 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Clouds drifting behind a fantasy keep: restored by Real-ESRGAN
mp4
123 KB
Actual file preview for Clouds drifting behind a fantasy keep: restored by Real-ESRGAN

Clouds drifting behind a fantasy keep: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/game-keep.mp4. Measured with ffmpeg's own filters it scores 34.455 dB PSNR and 0.923 SSIM; the same degraded input enlarged by plain bicubic scores 34.952 dB and 0.906, so the two metrics DISAGREE: -0.50 dB of PSNR against it, +0.0170 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Coast drift: 24-second continuous film
mp4
1.5 MB
Actual file preview for Coast drift: 24-second continuous film

Coast drift: 24-second continuous film

A 24 second 512x288 H.264 film at 24 fps, 576 frames, written with the moov atom first so it can be seeked before it has finished downloading. Twelve generated segments crossfaded into one continuous take. Across 575 transitions the mean inter-frame change is 0.0139 and the strongest single transition is 0.1614, with 0 above 0.3. A deliberate hard cut between two unrelated films measures 0.311 by the same method, so the worst join here sits 0.1496 below a real cut. At eleven times the length of any other clip in this library, it is the only footage here long enough to seek through, chapter, or scrub. Stylised rather than photographic: this is the 1.3B model at 512x288, the subject drifts over 24 seconds, and a dark rendering artefact recurs in a few frames. It is a fixture for duration, seeking and stitch measurement, not showcase footage.

File
MP4 · Generated Films · 512 × 288 px
Use case
Video QAMedia handling· Conversion set