
3GP: Mobile H.264 Clip
The clip as 3GP: the 3GPP mobile container from the feature-phone era. For testing 3GP demuxing and conversion.
- File
- 3GP · 3gp · 480x270
- Use case
- Conversion testing
Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.
Page 1 of 19; 24 results per page.

The clip as 3GP: the 3GPP mobile container from the feature-phone era. For testing 3GP demuxing and conversion.

A 60 frames-per-second clip with a smoothly sweeping marker: double the usual frame rate. A fixture for testing high-frame-rate playback, frame-rate detection, and fps conversion (60→30 decimation).

A five-second clip that flashes white and plays a 1 kHz beep on every whole second: the first clip in the library WITH an audio track. A direct fixture for measuring and correcting audio/video sync (lip-sync) offset.

The same codec and container with NO alpha plane, as the control for the alpha clip in this group. Useful for checking that alpha detection reads the pixel format rather than assuming every WebM is transparent, and for confirming a compositing bug is in the alpha handling rather than in the player.

Video with a real per-pixel alpha channel: VP9 yuva420p in WebM, the one alpha path browsers decode natively. Composite it over a page background and the transparency is genuine, not a chroma key. Roughly 90% of each frame is fully transparent. Two traps this file exists to expose. First, auto-alt-ref must be disabled at encode time or libvpx silently drops the alpha plane, producing a valid file with no transparency and no error. Second, WebM stores VP9 alpha in BlockAdditional and signals it with AlphaMode=1, so FFmpeg's NATIVE vp9 decoder reports pix_fmt yuv420p and decodes fully opaque: you must force `-c:v libvpx-vp9` to see the alpha at all. Probing this file with default settings and concluding it has no alpha is the expected mistake.

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-v-truck-artic_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0269 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/truck-depot.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 24.913 dB PSNR and 0.793 SSIM against that original - the number any model has to beat to be worth running.

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/truck-depot.mp4. Measured with ffmpeg's own filters it scores 25.299 dB PSNR and 0.842 SSIM; the same degraded input enlarged by plain bicubic scores 24.913 dB and 0.793, so it beats bicubic on both, by +0.39 dB and +0.0490 SSIM. Both figures are measurements against the same original, which is what makes them comparable at all.

Advanced SubStation Alpha using the features that distinguish it from SubRip: per-syllable \k karaoke timings in centiseconds, a full V4+ style definition with primary and secondary colours, and an \an override that repositions a line. Converting this to SRT necessarily loses all of it, which makes it a good test of whether a converter warns about that or drops it silently.

The clip as DivX-style MPEG-4 ASP in an AVI container: the codec that defined early desktop video. For testing MPEG-4 Part 2 decoding and AVI conversion.

The clip as Motion-JPEG in a classic AVI (RIFF) container: every frame an independent JPEG. For testing legacy AVI readers, MJPEG decoding, and AVI→modern-codec conversion.

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-owner-cafe_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0175 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/cafe-barista.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 30.791 dB PSNR and 0.911 SSIM against that original - the number any model has to beat to be worth running.

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/cafe-barista.mp4. Measured with ffmpeg's own filters it scores 30.639 dB PSNR and 0.941 SSIM; the same degraded input enlarged by plain bicubic scores 30.791 dB and 0.911, so the two metrics DISAGREE: -0.15 dB of PSNR against it, +0.0300 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-owner-cafe_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0261 (moderate). Clear movement without a scene change, which is the ordinary case for short footage. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.

A 24-patch chart drifting slowly so it is genuinely moving footage rather than a still. Patch values are fixed and documented, so colourisation and colour-management output can be measured per patch. The values are fictional and not a reproduction of any licensed reference chart. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

Far, mid and near layers translating at 0.6, 2.4 and 6.0 pixels per frame. Relative depth is defined by the motion ratio rather than guessed from cues, which gives depth-from-video and optical-flow output something objective to be scored against. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

A 36-spoke Siemens star under a slow zoom, plus bar-pair wedges from 16 pixels down to 2. Detail runs right down to the Nyquist limit, which is exactly where super-resolution and denoise either recover structure or invent it. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

A smooth vertical gradient whose hue drifts across the clip, with one soft glow for structure. Almost no high-frequency detail, so it provokes banding and blocking in exactly the way flat skies do in real footage: the hardest case for a low-bitrate encoder. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

A green disc and a red square orbiting the centre in antiphase over a grid. Hard edges and flat fills make the boundary unambiguous, so segmentation and tracking output can be scored against exact geometry rather than a judgement call. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

A skyline scrolling at a constant 3.5 pixels per frame: pure horizontal translation with no rotation or scale change. The lit windows give sparse high-contrast features to track, and the constant velocity means the correct answer for interpolation and optical flow is known exactly. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

An anti-aliased figure moving over a chroma-green field with a deliberate lighting falloff, so the key is not perfectly flat. Ships with a per-frame alpha ground truth, which is what makes matting output measurable instead of merely inspectable. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

Scrolling monospaced terminal output with a blinking cursor. Thin high-contrast glyph edges are what chroma subsampling and low bitrates destroy first, and legibility after processing is a pass/fail signal that needs no metric. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

A 2.0 second 640x480 H.264 clip at 24 fps, 49 frames. Animated from the still nss-v-bike-city_00001_.png, so the first frame is a known image and frame extraction can be checked against it. Measured mean inter-frame change 0.0586 (active). Substantial frame-to-frame change, the end of the range where naive scene detection starts firing. Synthetic footage: two seconds at this size is a decoder and pipeline fixture rather than showcase material, hands degrade in later frames, and no text in shot is legible.