Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.
A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-sp-runner_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0508 (active). Substantial frame-to-frame change, the end of the range where naive scene detection starts firing. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.
A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-owner-salon_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0082 (subtle). Motion is deliberate but slight, closer to a living photograph than to action. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.
SubViewer/SBV as YouTube exports it: no cue indices, a comma rather than an arrow between start and end, and single-digit hours. Superficially similar enough to SubRip that format detection by eye fails, which makes it a good test of sniffing logic: a parser that guesses SubRip from the timestamps then trips on the missing index.
The same interlacing with the opposite field order, the DV and SD convention. Deinterlacing with the wrong field order makes motion jitter backwards: a subtle, very common bug that this pair makes reproducible. In Matroska rather than MP4 so the catalog renders a poster instead of advertising an inline player, combing artefacts are best judged from a still anyway.
Alternate lines carry alternate instants, with the top field displayed first: the HD broadcast convention. Shown progressively it combs on every moving edge. In Matroska rather than MP4 so the catalog renders a poster instead of advertising an inline player, combing artefacts are best judged from a still anyway.
Whole frames, captured and displayed at once: the reference for this group and how essentially all modern video is shot. In Matroska rather than MP4 so the catalog renders a poster instead of advertising an inline player, combing artefacts are best judged from a still anyway.
An intentionally corrupt SCC file whose first caption is valid but whose second carries byte pairs with the wrong parity. CEA-608 requires odd parity in bit 7 of every byte, and broadcast decoders use it to detect transmission errors. Tests whether a decoder checks parity at all, and whether it drops just the bad caption or the whole file.
Broadcast closed captions as Scenarist SCC: SMPTE timecodes followed by hexadecimal CEA-608 byte pairs, with odd parity set on every byte as the standard requires. Uses pop-on mode: RCL to load, ENM to clear non-displayed memory, a preamble address code for row 15, the character pairs, then EOC to display. This is the fixture that exposes decoders which treat captions as text: the parity bits, the two-byte control codes and the frame-accurate timecodes all have to be handled.
The matching return cue for the break-start splice_insert in this group: the same spliceEventId with outOfNetworkIndicator false, signalling the feed is coming back from the ad break. It carries no BreakDuration because the return is explicit rather than automatic. Pairing start and end cues by event id is exactly what an ad-insertion implementation has to get right, and this is the pair to test it on.
An SCTE-35 splice_insert as XML: the cue that tells a downstream packager a linear ad break starts. outOfNetworkIndicator marks entry into the break, ptsTime is 8 seconds on the 90 kHz PTS clock, and the 30-second break auto-returns. This is the signal a manifest manipulator converts into an HLS EXT-X-CUE-OUT or a new DASH period.
The modern SCTE-35 form: a time_signal command carrying a segmentation descriptor rather than a bare splice_insert. Type 34 marks a provider placement opportunity start, and the delivery restrictions declare that this break may not be filled for web delivery: the field that ad-insertion logic is supposed to honour and frequently ignores. Includes a base64 UPID identifying the content.
Per-frame instance masks for the orbit plate, colour-coded so the three objects are separable: red is the disc, green the square, blue the centre marker, black the background. Instance identity is stable across frames, which is what makes this usable for video object segmentation and tracking rather than only per-frame segmentation. Note the mask is encoded lossily like every other clip here, so threshold rather than testing for exact equality.
Two hard-edged objects orbiting in antiphase over a grid, plus a static centre marker. The objects cross in front of the grid and reach the frame edges, so occlusion and boundary handling both get exercised. Every frame has an exact mask in this group.
A three-level trimap derived from the instance masks: definite foreground, definite background, and a nine-pixel unknown band around every boundary. This is the input interactive matting and rotoscoping tools actually take, and the band is where all the difficulty lives: a matte is only as good as its handling of that region.
A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-owner-shop_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0107 (subtle). Motion is deliberate but slight, closer to a living photograph than to action. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.
Standard SMPTE-style colour bars with a 1 kHz reference tone: the classic broadcast test pattern. A fixture for calibrating colour, testing decoders, and verifying audio alongside a known video signal.
A clip with a soft (muxed, toggleable) SRT subtitle track in a Matroska container: the subtitles are a separate stream, not burned in. A fixture for testing subtitle extraction, rendering, and toggling.
The same clip with a soft subtitle track as MP4 mov_text (the MP4-native subtitle format): the format twin of the MKV/SRT file, for testing subtitle handling across containers.
An intentionally corrupt SubRip file whose first cue ends three seconds before it begins. Nothing is malformed at the text level, so parsers accept it and the problem only surfaces downstream, as a negative duration, a caption that never displays, or a sort that puts the timeline out of order.
An intentionally corrupt SubRip file with a lone UTF-8 continuation byte appended: a byte that can never legally start a character. Both cues are otherwise fine. Exercises the difference between strict decoding, which raises, and lenient decoding, which substitutes a replacement character and continues.
File
SRT · Captions Corrupt · UTF-8 with one invalid sequence
An intentionally corrupt SubRip file: the first cue's timing line is missing the --> separator, so it cannot be parsed as a time range. The second cue is well-formed, which is the point: a robust parser should report the bad cue and still return the good one rather than failing the whole file.
An intentionally corrupt SubRip file with two separate index problems: the first cue is numbered with a word, and index 2 is then used twice. SubRip indices are advisory rather than load-bearing, so this distinguishes parsers that key cues by index (and lose one to the collision) from those that do not.
An intentionally corrupt SubRip file that stops in the middle of the third cue's timestamp, as a truncated download or an interrupted write would. There is no trailing newline. Tests whether a parser returns the two complete cues or discards everything because the tail is unparseable.
Five hundred short sequential cues over four minutes. Small in bytes but long enough to expose quadratic parsing, per-cue DOM churn, and UI that re-renders the whole caption list on every cue change.