Skip to content
Testaroo

Downscaled clips with their full-resolution ground truth

Low-resolution clips produced from a documented high-quality source by a recorded filter, so super-resolution output can be measured against the original rather than judged by eye.

43 of 43 files
Preview of Base Plate: Detail, Siemens Star and Frequency Wedges
mp4
68.3 KB
Actual file preview for Base Plate: Detail, Siemens Star and Frequency Wedges

Base Plate: Detail, Siemens Star and Frequency Wedges

A 36-spoke Siemens star under a slow zoom, plus bar-pair wedges from 16 pixels down to 2. Detail runs right down to the Nyquist limit, which is exactly where super-resolution and denoise either recover structure or invent it. One of eight shared base plates: every AI-video suite in this library degrades one of these rather than inventing its own footage, so results across suites are comparable. Encoded at CRF 14 (well above the house CRF 30), because a reference compressed as hard as the material under test puts the measurement floor above the effect being measured.

File
MP4 · Base Plates · 640x360
Use case
Video upscalingVideo denoise+1· Conversion set
Preview of Super-Resolution Ground Truth: detail-chart
mp4
68.3 KB
Actual file preview for Super-Resolution Ground Truth: detail-chart

Super-Resolution Ground Truth: detail-chart

The full-resolution reference for the detail-chart super-resolution set at 640x360. Every low-resolution input in this group was produced by downscaling these exact pixels with a recorded filter, so an upscaler's output can be compared against the true original instead of against another upscale.

File
MP4 · Superres · 640x360
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷2 Bicubic
mp4
32.8 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷2 Bicubic

Super-Resolution Input: detail-chart, ÷2 Bicubic

The detail-chart plate downscaled 2× to 320x180 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷2 Area / box
mp4
34.6 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷2 Area / box

Super-Resolution Input: detail-chart, ÷2 Area / box

The detail-chart plate downscaled 2× to 320x180 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷2 Nearest neighbour
mp4
28.6 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷2 Nearest neighbour

Super-Resolution Input: detail-chart, ÷2 Nearest neighbour

The detail-chart plate downscaled 2× to 320x180 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷3 Bicubic
mp4
20.8 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷3 Bicubic

Super-Resolution Input: detail-chart, ÷3 Bicubic

The detail-chart plate downscaled 3× to 212x120 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷3 Area / box
mp4
22.1 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷3 Area / box

Super-Resolution Input: detail-chart, ÷3 Area / box

The detail-chart plate downscaled 3× to 212x120 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷3 Nearest neighbour
mp4
17.4 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷3 Nearest neighbour

Super-Resolution Input: detail-chart, ÷3 Nearest neighbour

The detail-chart plate downscaled 3× to 212x120 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷4 Bicubic
mp4
14.6 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷4 Bicubic

Super-Resolution Input: detail-chart, ÷4 Bicubic

The detail-chart plate downscaled 4× to 160x90 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷4 Area / box
mp4
15.5 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷4 Area / box

Super-Resolution Input: detail-chart, ÷4 Area / box

The detail-chart plate downscaled 4× to 160x90 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: detail-chart, ÷4 Nearest neighbour
mp4
12.4 KB
Actual file preview for Super-Resolution Input: detail-chart, ÷4 Nearest neighbour

Super-Resolution Input: detail-chart, ÷4 Nearest neighbour

The detail-chart plate downscaled 4× to 160x90 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Ground Truth: text-motion
mp4
37.7 KB
Actual file preview for Super-Resolution Ground Truth: text-motion

Super-Resolution Ground Truth: text-motion

The full-resolution reference for the text-motion super-resolution set at 640x360. Every low-resolution input in this group was produced by downscaling these exact pixels with a recorded filter, so an upscaler's output can be compared against the true original instead of against another upscale.

File
MP4 · Superres · 640x360
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷2 Bicubic
mp4
18.5 KB
Actual file preview for Super-Resolution Input: text-motion, ÷2 Bicubic

Super-Resolution Input: text-motion, ÷2 Bicubic

The text-motion plate downscaled 2× to 320x180 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷2 Area / box
mp4
26.9 KB
Actual file preview for Super-Resolution Input: text-motion, ÷2 Area / box

Super-Resolution Input: text-motion, ÷2 Area / box

The text-motion plate downscaled 2× to 320x180 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷2 Nearest neighbour
mp4
27 KB
Actual file preview for Super-Resolution Input: text-motion, ÷2 Nearest neighbour

Super-Resolution Input: text-motion, ÷2 Nearest neighbour

The text-motion plate downscaled 2× to 320x180 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷3 Bicubic
mp4
9.4 KB
Actual file preview for Super-Resolution Input: text-motion, ÷3 Bicubic

Super-Resolution Input: text-motion, ÷3 Bicubic

The text-motion plate downscaled 3× to 212x120 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷3 Area / box
mp4
15.3 KB
Actual file preview for Super-Resolution Input: text-motion, ÷3 Area / box

Super-Resolution Input: text-motion, ÷3 Area / box

The text-motion plate downscaled 3× to 212x120 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷3 Nearest neighbour
mp4
31.2 KB
Actual file preview for Super-Resolution Input: text-motion, ÷3 Nearest neighbour

Super-Resolution Input: text-motion, ÷3 Nearest neighbour

The text-motion plate downscaled 3× to 212x120 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷4 Bicubic
mp4
6.5 KB
Actual file preview for Super-Resolution Input: text-motion, ÷4 Bicubic

Super-Resolution Input: text-motion, ÷4 Bicubic

The text-motion plate downscaled 4× to 160x90 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷4 Area / box
mp4
8.8 KB
Actual file preview for Super-Resolution Input: text-motion, ÷4 Area / box

Super-Resolution Input: text-motion, ÷4 Area / box

The text-motion plate downscaled 4× to 160x90 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: text-motion, ÷4 Nearest neighbour
mp4
22.9 KB
Actual file preview for Super-Resolution Input: text-motion, ÷4 Nearest neighbour

Super-Resolution Input: text-motion, ÷4 Nearest neighbour

The text-motion plate downscaled 4× to 160x90 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Ground Truth: pan-city
mp4
37.9 KB
Actual file preview for Super-Resolution Ground Truth: pan-city

Super-Resolution Ground Truth: pan-city

The full-resolution reference for the pan-city super-resolution set at 640x360. Every low-resolution input in this group was produced by downscaling these exact pixels with a recorded filter, so an upscaler's output can be compared against the true original instead of against another upscale.

File
MP4 · Superres · 640x360
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷2 Bicubic
mp4
21 KB
Actual file preview for Super-Resolution Input: pan-city, ÷2 Bicubic

Super-Resolution Input: pan-city, ÷2 Bicubic

The pan-city plate downscaled 2× to 320x180 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷2 Area / box
mp4
25.2 KB
Actual file preview for Super-Resolution Input: pan-city, ÷2 Area / box

Super-Resolution Input: pan-city, ÷2 Area / box

The pan-city plate downscaled 2× to 320x180 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷2 Nearest neighbour
mp4
19.3 KB
Actual file preview for Super-Resolution Input: pan-city, ÷2 Nearest neighbour

Super-Resolution Input: pan-city, ÷2 Nearest neighbour

The pan-city plate downscaled 2× to 320x180 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 320x180
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷3 Bicubic
mp4
25.2 KB
Actual file preview for Super-Resolution Input: pan-city, ÷3 Bicubic

Super-Resolution Input: pan-city, ÷3 Bicubic

The pan-city plate downscaled 3× to 212x120 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷3 Area / box
mp4
26 KB
Actual file preview for Super-Resolution Input: pan-city, ÷3 Area / box

Super-Resolution Input: pan-city, ÷3 Area / box

The pan-city plate downscaled 3× to 212x120 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷3 Nearest neighbour
mp4
20.8 KB
Actual file preview for Super-Resolution Input: pan-city, ÷3 Nearest neighbour

Super-Resolution Input: pan-city, ÷3 Nearest neighbour

The pan-city plate downscaled 3× to 212x120 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 212x120
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷4 Bicubic
mp4
14.8 KB
Actual file preview for Super-Resolution Input: pan-city, ÷4 Bicubic

Super-Resolution Input: pan-city, ÷4 Bicubic

The pan-city plate downscaled 4× to 160x90 using a bicubic filter. The standard downscale in most benchmarks. Mild ringing at edges, and the filter most super-resolution models are trained to invert, so it flatters them. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷4 Area / box
mp4
16.2 KB
Actual file preview for Super-Resolution Input: pan-city, ÷4 Area / box

Super-Resolution Input: pan-city, ÷4 Area / box

The pan-city plate downscaled 4× to 160x90 using a area / box filter. Simple pixel averaging, what a camera's binning path actually does. Softer than bicubic and not what most models saw in training, which makes it a fairer test. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Super-Resolution Input: pan-city, ÷4 Nearest neighbour
mp4
13.9 KB
Actual file preview for Super-Resolution Input: pan-city, ÷4 Nearest neighbour

Super-Resolution Input: pan-city, ÷4 Nearest neighbour

The pan-city plate downscaled 4× to 160x90 using a nearest neighbour filter. Point sampling with no filtering at all, so downscaling aliases hard. The Siemens star folds into moiré, and no amount of upscaling can recover what aliasing destroyed. Upscale it back to 640x360 and score against the ground truth in this group: the filter is recorded because which one was used changes the difficulty far more than the scale factor does.

File
MP4 · Superres · 160x90
Use case
Video upscalingVideo QA· Conversion set
Preview of Barista at the espresso machine: degraded input
mp4
26.7 KB
Actual file preview for Barista at the espresso machine: degraded input

Barista at the espresso machine: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/cafe-barista.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 30.791 dB PSNR and 0.911 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Barista at the espresso machine: restored by Real-ESRGAN
mp4
245.9 KB
Actual file preview for Barista at the espresso machine: restored by Real-ESRGAN

Barista at the espresso machine: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/cafe-barista.mp4. Measured with ffmpeg's own filters it scores 30.639 dB PSNR and 0.941 SSIM; the same degraded input enlarged by plain bicubic scores 30.791 dB and 0.911, so the two metrics DISAGREE: -0.15 dB of PSNR against it, +0.0300 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Cat on a windowsill: degraded input
mp4
20.1 KB
Actual file preview for Cat on a windowsill: degraded input

Cat on a windowsill: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/cat-window.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 32.598 dB PSNR and 0.917 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Cat on a windowsill: restored by Real-ESRGAN
mp4
225.7 KB
Actual file preview for Cat on a windowsill: restored by Real-ESRGAN

Cat on a windowsill: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/cat-window.mp4. Measured with ffmpeg's own filters it scores 31.398 dB PSNR and 0.932 SSIM; the same degraded input enlarged by plain bicubic scores 32.598 dB and 0.917, so the two metrics DISAGREE: -1.20 dB of PSNR against it, +0.0150 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Articulated lorry at a loading bay: degraded input
mp4
27 KB
Actual file preview for Articulated lorry at a loading bay: degraded input

Articulated lorry at a loading bay: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/truck-depot.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 24.913 dB PSNR and 0.793 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Articulated lorry at a loading bay: restored by Real-ESRGAN
mp4
300.4 KB
Actual file preview for Articulated lorry at a loading bay: restored by Real-ESRGAN

Articulated lorry at a loading bay: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/truck-depot.mp4. Measured with ffmpeg's own filters it scores 25.299 dB PSNR and 0.842 SSIM; the same degraded input enlarged by plain bicubic scores 24.913 dB and 0.793, so it beats bicubic on both, by +0.39 dB and +0.0490 SSIM. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Chef plating a dish: degraded input
mp4
36.2 KB
Actual file preview for Chef plating a dish: degraded input

Chef plating a dish: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/kitchen-chef.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 27.165 dB PSNR and 0.861 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Chef plating a dish: restored by Real-ESRGAN
mp4
329.2 KB
Actual file preview for Chef plating a dish: restored by Real-ESRGAN

Chef plating a dish: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/kitchen-chef.mp4. Measured with ffmpeg's own filters it scores 28.023 dB PSNR and 0.906 SSIM; the same degraded input enlarged by plain bicubic scores 27.165 dB and 0.861, so it beats bicubic on both, by +0.86 dB and +0.0450 SSIM. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Clouds drifting behind a fantasy keep: degraded input
mp4
15.5 KB
Actual file preview for Clouds drifting behind a fantasy keep: degraded input

Clouds drifting behind a fantasy keep: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/game-keep.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 34.952 dB PSNR and 0.906 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Clouds drifting behind a fantasy keep: restored by Real-ESRGAN
mp4
123 KB
Actual file preview for Clouds drifting behind a fantasy keep: restored by Real-ESRGAN

Clouds drifting behind a fantasy keep: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/game-keep.mp4. Measured with ffmpeg's own filters it scores 34.455 dB PSNR and 0.923 SSIM; the same degraded input enlarged by plain bicubic scores 34.952 dB and 0.906, so the two metrics DISAGREE: -0.50 dB of PSNR against it, +0.0170 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Mist drifting through a mountain valley: degraded input
mp4
14 KB
Actual file preview for Mist drifting through a mountain valley: degraded input

Mist drifting through a mountain valley: degraded input

The input a super-resolution tool is given: the published clip /files/video/comfy-realism-v1/mountain.mp4 reduced from 640x480 to 160x120 by an exact 4x area downscale, over all 49 frames at 24.0 fps. Because the reduction is an exact integer factor of a clip that is already in this catalogue, anything a tool produces from this file can be MEASURED against the original rather than judged by eye. For reference, a plain bicubic enlargement back to 640x480 scores 34.359 dB PSNR and 0.915 SSIM against that original - the number any model has to beat to be worth running.

File
MP4 · Superres · 160 × 120 px
Use case
Video upscalingSuper-resolution· Paired fixture
Preview of Mist drifting through a mountain valley: restored by Real-ESRGAN
mp4
134.8 KB
Actual file preview for Mist drifting through a mountain valley: restored by Real-ESRGAN

Mist drifting through a mountain valley: restored by Real-ESRGAN

The 160x120 input in this group restored to 640x480 by Real-ESRGAN x4plus (BSD-3-Clause), frame for frame with all 49 frames intact, so it compares directly against the published ground truth /files/video/comfy-realism-v1/mountain.mp4. Measured with ffmpeg's own filters it scores 32.503 dB PSNR and 0.92 SSIM; the same degraded input enlarged by plain bicubic scores 34.359 dB and 0.915, so the two metrics DISAGREE: -1.86 dB of PSNR against it, +0.0050 of SSIM for it. That is the signature of a perceptual upscaler - it invents texture, which restores structure while moving individual pixels further from the original, and it is why a super-resolution result reported as one number is not reportable. Both figures are measurements against the same original, which is what makes them comparable at all.

File
MP4 · Superres · 640 × 480 px
Use case
Video upscalingSuper-resolution· Paired fixture