Find files, editable templates and browser test targets by what you need to make or test. The directory below is cut by format; the two collections under it cut the same library by subject and by workflow.
The same view of city bicycle leaning against a brick wall, soft daylight, with a foreign object composited over 7.55% of the frame as two disconnected regions - the case a reader that keeps only the largest connected component, or takes the bounding box of both, gets wrong. This is the file an object-removal tool is given. Outside the mask it is byte-identical to the clean plate beside it, so any difference a tool leaves there is damage it did rather than content it was handed.
City bicycle leaning against a brick wall, soft daylight, with the object taken back out by Stable Diffusion 1.5 inpainting and the gap reconstructed from the surrounding context alone - the masked latents are erased before sampling, so the model never saw what it was painting over. Inside the mask it differs from the source by 53.072/255 and from the ground-truth plate by 26.044/255; the second number is NOT expected to be small, because an inpainter invents plausible content rather than recovering what was there. Beyond a 16-pixel ring around the mask the frame changes by only 7.882/255, which is the full-frame VAE round trip and not an edit.
The exact footprint of the object sitting over city bicycle leaning against a brick wall, soft daylight, as an 8-bit mask covering 7.55% of the frame as two disconnected regions - the case a reader that keeps only the largest connected component, or takes the bounding box of both, gets wrong. It holds only the values 0 and 255. The footprint is what DREW the object, so it is ground truth by construction rather than a segmentation of it. Hard-edged on purpose: a feathered edge has no exact footprint, and the exactness is the point of shipping it.
A tangent-space surface normal map for the published plate v-bike-city.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.996, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-boat-harbour.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 8.0% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A clean 512x512 crop of small fishing boat moored in a harbour, morning light, taken from the published plate v-boat-harbour.png before anything was added to it. This is the ANSWER KEY for its group: the object in the source file was composited onto this image, so this is exactly what was behind it. Nothing else in the group came from a second tool's guess.
The same view of small fishing boat moored in a harbour, morning light, with a foreign object composited over 5.61% of the frame as a single organic region, neither a rectangle nor an ellipse. This is the file an object-removal tool is given. Outside the mask it is byte-identical to the clean plate beside it, so any difference a tool leaves there is damage it did rather than content it was handed.
Small fishing boat moored in a harbour, morning light, with the object taken back out by Stable Diffusion 1.5 inpainting and the gap reconstructed from the surrounding context alone - the masked latents are erased before sampling, so the model never saw what it was painting over. Inside the mask it differs from the source by 30.711/255 and from the ground-truth plate by 25.736/255; the second number is NOT expected to be small, because an inpainter invents plausible content rather than recovering what was there. Beyond a 16-pixel ring around the mask the frame changes by only 7.026/255, which is the full-frame VAE round trip and not an edit.
The exact footprint of the object sitting over small fishing boat moored in a harbour, morning light, as an 8-bit mask covering 5.61% of the frame as a single organic region, neither a rectangle nor an ellipse. It holds only the values 0 and 255. The footprint is what DREW the object, so it is ground truth by construction rather than a segmentation of it. Hard-edged on purpose: a feathered edge has no exact footprint, and the exactness is the point of shipping it.
A tangent-space surface normal map for the published plate v-boat-harbour.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9956, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-bus-city.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 12.6% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A tangent-space surface normal map for the published plate v-bus-city.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9956, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-car-classic.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 6.2% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A tangent-space surface normal map for the published plate v-car-classic.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9954, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-car-ev.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 7.7% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A clean 512x512 crop of white electric hatchback at a charging point, urban car park, taken from the published plate v-car-ev.png before anything was added to it. This is the ANSWER KEY for its group: the object in the source file was composited onto this image, so this is exactly what was behind it. Nothing else in the group came from a second tool's guess.
The same view of white electric hatchback at a charging point, urban car park, with a foreign object composited over 4.31% of the frame as two disconnected regions - the case a reader that keeps only the largest connected component, or takes the bounding box of both, gets wrong. This is the file an object-removal tool is given. Outside the mask it is byte-identical to the clean plate beside it, so any difference a tool leaves there is damage it did rather than content it was handed.
White electric hatchback at a charging point, urban car park, with the object taken back out by Stable Diffusion 1.5 inpainting and the gap reconstructed from the surrounding context alone - the masked latents are erased before sampling, so the model never saw what it was painting over. Inside the mask it differs from the source by 89.979/255 and from the ground-truth plate by 38.547/255; the second number is NOT expected to be small, because an inpainter invents plausible content rather than recovering what was there. Beyond a 16-pixel ring around the mask the frame changes by only 4.784/255, which is the full-frame VAE round trip and not an edit.
The exact footprint of the object sitting over white electric hatchback at a charging point, urban car park, as an 8-bit mask covering 4.31% of the frame as two disconnected regions - the case a reader that keeps only the largest connected component, or takes the bounding box of both, gets wrong. It holds only the values 0 and 255. The footprint is what DREW the object, so it is ground truth by construction rather than a segmentation of it. Hard-edged on purpose: a feathered edge has no exact footprint, and the exactness is the point of shipping it.
A tangent-space surface normal map for the published plate v-car-ev.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9964, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-car-night.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 8.8% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A tangent-space surface normal map for the published plate v-car-night.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9952, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.
A canny edge map for the published plate v-car-sedan.png, 1024x1024. Canny edge detection at low threshold 0.1 and high 0.3, run at the plate's own resolution so the edges land on the same pixels as the photograph they came from. Lit coverage measures 7.8% of the frame. The plate and this map are the same scene at the same size, so they can be compared pixel for pixel rather than by eye.
A tangent-space surface normal map for the published plate v-car-sedan.png, 1024x1024 - the same size as the plate, so the two compare pixel for pixel with no resample in between. RGB encodes the XYZ surface direction remapped from -1..1 into 0..255. Decoded back to vectors this file measures a mean length of 0.9961, which is the number to check your own decode against: a reader that transposes the channels or inverts the remap still produces a plausible-looking image, and lights the surface the wrong way.