Data
Data files are where parsers quietly break, so this category leans into the edge cases. The CSV set runs from clean to deliberately messy, quoted commas, embedded newlines, ragged rows, semicolon and tab delimiters, a headerless file, a ten-thousand-row file, and a Latin-1 encoded file. JSON comes flat, deeply nested, as JSON Lines, and intentionally invalid for error-handling tests. OpenAPI and JSON Schema twins, paginated API and webhook fixtures, and irregular time-series edges cover modern API QA. There's XML, YAML, and a real SQLite database with related tables, plus faker-generated user and order datasets with documented schemas. Each file spells out its exact quirks in the spec, so when your importer chokes you know precisely which case did it.
Filter data on Browse · 1255 files · 73 subcategories
Frequently asked questions
What CSV edge cases do you cover?+
Encoding matrix (UTF-8 BOM, UTF-16, CP1252, Shift-JIS, Latin-1), delimiter variants, fixed-width, ragged rows, quoted newlines, and domain mini-datasets (SaaS, POS, lab, marketing).
Are datasets realistic or random?+
Schemas are documented and values are deterministic synthetic SAMPLE data, shaped like production exports without real customer PII.
How do I test encoding detection?+
Filter Browse by purpose encoding-detection or open the encodings subcategory under Data.
Do you include OpenAPI, JSON Schema, or time-series edges?+
Yes, later waves add valid/invalid schema twins, pagination and webhook API fixtures, and irregular timestamp / DST-gap series. Filter by schema-testing or timeseries-testing.