Skip to content
Novus Examples

Documentation

Contracts, fixtures, editors, and local creation, from a stable URL to a verified export.

Updated Applies to Novus Examples 2026.08

What Novus Examples is

Novus Examples is a library of real, downloadable example and test files. Every file is generated by a deterministic script, not collected from the web. That single design choice is what makes the library useful: because a script produced each file, every file ships with an exact specification, a fixed random seed where randomness is involved, and, where it matters, a paired counterpart for before/after testing. Nothing here is copyrighted stock content, so everything is free to use.

Reading a file page

Each file has its own page with four things:

  • A preview: an inline image, an audio player, or the first lines of a text file.
  • A spec table, the file's exact properties: dimensions, colour space, noise sigma and seed, silence timestamps, form field names, encoding, formulas, and so on.
  • Purpose tags: what the file is good for testing, each linking to a purpose page.
  • One download button: a direct link to the static file. There is exactly one per page.

The spec table is the point. You are testing against a documented input, not a mystery file.

Testing contracts

New P6 fixtures add a visible testing contract to the ordinary preview and spec sheet. It names the scenario, the exact result a test should assert, and the artifact's role: valid, recoverable, intentionally invalid, or a reference control. The contract is structured catalog data, so the same expectation appears consistently on file, form, and document-template pages.

Use the contract as the acceptance criterion in a parser, converter, importer, or model-runtime test. The grouped calendar, vCard, metadata, schema, ONNX, and graph suites publish cross-file relationships and expected outcomes instead of asking you to infer behavior from a filename.

Pairs

A pair is a 1:1 relationship between two files that belong together: a clean image and its noisy version, a colour image and its greyscale conversion, a text document and its image-only scanned twin. Pairs are what serious testing needs: one file is the input, the other is the expected result or the ground truth you score against. When a file has a pair, its page links straight to it.

Groups and conversion sets

A group ties together a family of files that share the same content: one test card exported across PNG, WebP, AVIF, and more, or a resolution ladder from 16px to 4K, or a template offered as both DOCX and PDF. Groups are how you test converters. Convert one member and diff against the expected twin.

Purpose tags

Every file is tagged with what it helps you test. Browse them under Browse by purpose: denoise testing, CSV parsing, OCR, auto-trim, form parsing, encoding detection, and more. Purpose pages are the fastest way to find the exact fixtures for a task.

The libraries

Every fixture belongs to exactly one library, and every library is a page listing its own files with the same preview, spec sheet, contract, pair and group furniture. The split is by what a file is, not by what you are testing with it - that is what purpose tags are for, and one file usually carries several.

  • Images - enhancement, segmentation, OCR, visual-diff and vision fixtures, with degraded inputs paired to clean references.
  • Audio - reference tones, ASR pairs, loudness ladders and silence-trim sets.
  • Documents - PDFs, Office files, annotations, markup and text encodings.
  • Data - CSV, JSON, schemas, API payloads and time series, clean and deliberately messy.
  • Templates - visual scenes and document starting points, each with a local Studio.
  • Forms - live on-site demos plus HTML and AcroForm fixtures for real-world form cases.
  • Archives - ZIP, TAR, gzip, 7z and ISO, including nested, unicode-named, empty and password-protected members.
  • E-books - EPUB, FB2 and CBZ, valid and intentionally invalid.
  • Models - 3D and CAD interchange: STL, OBJ, 3MF, DXF, STEP, IGES and EPS.
  • Fonts - one demonstration typeface exported as TTF, OTF, WOFF and WOFF2.
  • Video - one documented clip across MP4, WebM, MKV, MOV, AVI and OGV.
  • Email - RFC 822 .eml messages and .mbox mailboxes.
  • AI and ML - training data, embeddings, annotations and evaluation sets.
  • Prompts - provider-neutral plans, message fixtures, tool calls and evaluation contracts.
  • Creator audio - foley, ambience, voice, stems, cue sheets and controlled editing cases.
  • Creator video - b-roll, social variants, overlays, timelines, captions and sync references.
  • Source code - idiomatic snippets across many languages, with intentionally invalid lint fixtures labelled as such.
  • Web assets - favicons, PWA manifests, service workers, robots and sitemap files, social cards and .well-known files.
  • Localization - PO/POT, XLIFF, .strings, ARB, Android XML, RESX and i18next, with RTL and CJK variants.
  • Security - PKI, JWT and JWKS, OAuth token responses, CORS and CSP header text, and SSH samples.
  • Testing and QA - JUnit, TAP, Gherkin, coverage and HAR: the reports CI actually has to parse.
  • Version control - unified diffs, patches, conflict markers and repository config files.
  • Observability - OTLP traces, Prometheus metrics, structured logs and alert payloads.
  • Supply chain - SBOMs, lockfiles, provenance attestations and vulnerability reports.
  • Geospatial - GeoJSON, shapefiles, KML and KMZ, GeoPackage, WKT and GPS traces.
  • Scientific - chemistry, bioinformatics, gridded data and citation formats.
  • Pipelines - CI/CD workflows, infrastructure-as-code configs, build files and orchestration DAGs.

Nothing in the security library is a real identity, and none of it should ever be used as one. Every key, token and certificate there is generated for parser and harness tests and is marked SAMPLE on its page.

Finding a fixture

Five surfaces, each answering a different question.

  • Browse - faceted by category, format, purpose, validity role and which workspace can open the file. Use it when you know roughly what you want and need to narrow.
  • Search - one box across files, visual templates, live targets, pages and articles. Use it when you know a word.
  • Format coverage - the matrix of every format the library provides, what is planned, and what is deliberately absent with the reason. Use it to settle "do you have a .parquet" without guessing at a URL.
  • Glossary - one page per format in plain language: what the format is, and what a fixture of it is good for. Use it when the extension is the unknown.
  • Tool map - the complete route index for the site, including the machine-readable feeds. Use it when you want to see everything at once.

Business examples has a different shape again. A kit is not one fixture but a set of files that refer to one another - records, templates, schemas and the expected result - so you can exercise a whole import or reconciliation workflow instead of a single parse. Each kit names the platform profile it was checked against and marks its invalid members as invalid.

Where the catalogue is thin

Coverage is published rather than implied, because a count of files is not a count of cases.

Every fixture carries a validity role: valid, recoverable, intentionally invalid, or a reference control. The useful question about a format is therefore not how many files it has but whether it ships anything that fails. Measured across the whole catalogue on 2026-09-11, 202 of 288 extensions ship no failing fixture at all, and png alone holds 416 fixtures with no negative path. That is a real hole rather than a rounding error: if you are testing error handling for one of those formats, this library currently hands you the happy path and nothing else.

The same measurement separated cells that are volume from cells that are coverage - several hundred fixtures can collapse to a handful of distinct contract shapes when one generator loops over a parameter. Both lists are named, not summarised, in docs/COVERAGE_CUBE.md in the source repository.

Format coverage is the browsable half of this: what exists, what is planned, and what is deliberately absent with the reason why.

Live targets

Targets are the one part of the site that is not a download. Each is a page built to behave in one specific awkward way, so a scraper, a screen reader, an accessibility audit or a browser-automation script can be pointed at difficulty that is documented instead of accidental. They are permanent URLs like everything else, so a failing test can link to the exact page that broke it.

Text that is not where a DOM query expects it: accessible name only, canvas-rendered text, SVG text, CSS generated content, eight ways to hide text, bidirectional text, icon glyphs and their labels.

Content that arrives late, or in pieces: slow load, delayed hydration, lazy loading, skeleton swap, infinite scroll, streaming rows, virtualised list, client-side pagination.

Boundaries a selector does not cross: shadow DOM, shadow DOM form, web components, iframe nest, deep iframe nest, iframe sandbox.

Widgets with their own keyboard and ARIA contract: accordion, ARIA tabs, custom listbox, keyboard menu, typeahead, ARIA live regions, ARIA landmarks, toast queue, slider precision, table sort and filter.

Things that get between you and a click: blocking modal, nested modals, sticky overlay, pointer-events overlay, scroll lock, hover-only content, cookie banner.

Forms that resist automation: live validation, client-side stepper, sign-in shaped form, paste-only field, disabled and readonly states, contenteditable region, file drop zone.

Pointer and drag: drag-and-drop reorder, pointer drag on a canvas.

State that does not live in the element tree: hash-fragment state, scroll spy, forced dark, forced light, print stylesheet.

Licensing

Every file in the library is free for any use, with no attribution required, including commercial use. See the license for the plain-language terms. The intentionally corrupt files are provided as-is for error-handling tests; they are not valid files by design and are clearly labelled as such.

Stable URLs

File URLs, ids, and download paths are permanent. Once a file ships, its address never changes. That means you can link to a fixture from a README, embed its download URL in a CI pipeline, or bookmark it, and it will keep working.

Using fixtures in CI

Because every download is a static URL, you can fetch fixtures directly in a test pipeline:

curl -O https://examples.novusstreamsolutions.com/files/data/csv/messy-csv-quoted-commas-newlines.csv

Point your parser, converter, or filter at the downloaded file and assert against the documented spec. For paired fixtures, download both the input and its reference and score one against the other. Nothing requires an account or an API key.

Reading the site from a script

Fetching a known file needs nothing but its URL. When a script needs to know what is here, there is a small read-only JSON API under /api/v1. It is the same shape on every Novus site, so one client can walk all of them from any entry point.

  • /api/v1 - the index: name, description, version, and the address of every other route, so a client that knows only the origin can discover the rest in one request.
  • /api/v1/health - status, version and the current time. Never cached, because a cached health check reports the past.
  • /api/v1/site - what this site is, and the address of its MCP endpoint.
  • /api/v1/capabilities - what the site can do for a machine caller, projected from the same tool definitions /mcp dispatches on rather than retyped alongside them.
  • /api/v1/siblings - the other Novus properties.

The rules, so you do not have to discover them. GET and HEAD only; anything else answers 405 with a JSON body saying the API is read-only. No key, no account, no cookies. Access-Control-Allow-Origin: * and no credentials, so a browser client works and cannot be tricked into attaching a session. Reads are shared-cacheable for five minutes at the browser and ten at the edge, with a day of stale-while-revalidate behind that, so the CDN answers most of them. Rate limiting is per route and per caller at 60 requests a minute - a person reading cannot reach it, a tight loop will, and exhausting one route cannot exhaust another.

This API describes the site; it does not enumerate the catalogue. To find fixtures, use the MCP tools below or search; for one link-rich summary of the whole site, generated from the live catalogue, fetch /llms.txt. The files themselves are always at their permanent /files/… URLs.

Other endpoints exist and are deliberately not part of that contract. /api/auth/*, /api/saved, /api/saves, /api/account/* and /api/notifications serve the optional account. /api/chat/* serves the live chat on the contact page. /api/status and /api/health report the site's own health for the status page. /api/demo-submit is the endpoint behind the live form demos: same-origin only, POST only, form content types only, honeypot-guarded, rate-limited, byte-capped while it reads the body, and it stores and forwards nothing.

Connecting an AI client

The site runs a Model Context Protocol server at /mcp, and the MCP server page is its documentation: the address, the client configuration to paste, every tool it exposes, and what it deliberately will not do. The endpoint is POST-only JSON-RPC, so opening it in a browser answers 405 on purpose - that page is the half a person can read, and its tool list is derived from the same object the endpoint dispatches on rather than retyped.

The six tools locate things: the categories, the purposes, files matching a description, one file's full record, the documentation, and the sibling Novus sites. They do not open, convert, or check a file for you. That limit is repeated inside the tool descriptions themselves, so an assistant reading them cannot promise more than the server can do.

Choose a creation workspace

Create is the shared starting point for the site's six local-first creation models:

  • Text and data: start with the blank buffer or a format starter in the editor, then preview, convert, and download.
  • Forms: open any of the 243 catalog fixtures in Form Studio. HTML and PDF use layout editors; JSON, CSV, TXT, FDF, and XFDF use honest field-value editors with original-format export and re-import.
  • Documents: start with a blank block shell or one of 36 editable curated families in Document Studio, then export HTML, Markdown, DOCX, XLSX, or print to PDF. A definition with exactly one table can also export CSV without inventing a multi-table flattening rule.
  • Visuals: start from the blank canvas or choose a validated scene in Visual Template Studio, edit its text, tables, shapes, colours, logos, and images, then export locally.
  • Audio: synthesize a tone, sweep, seeded noise signal, or silence in the audio workspace, then export a PCM or encoded test fixture with recipe, source, and output SHA-256 evidence.
  • Video: synthesize colour bars, a countdown, or a motion grid in the video workspace, using only codecs reported by this browser session.

The workspaces share an entry point and local-first privacy model, but keep separate editors because their contracts differ. A CSV cell, an AcroForm widget, a document block, and a positioned canvas node should not silently coerce into one another.

What "open in editor" actually promises

"Open in editor" is not one promise, and the difference matters before you rely on it. Six capability levels are published, each mapped from the code that runs rather than from intent:

  • Open and read - the file is shown in a format-aware way: an inline player or viewer, a decoded text view, an archive's member list, or the read-only hex-and-ASCII dump with its base64 copy. No write path at all. Hex and base64 viewing belong here and nowhere else.
  • Structured edit - the file is parsed into something you can change (document text, typed columns, named form cells, members inside a container) and written back in the same format. It does not promise byte fidelity for anything the representation did not model.
  • Round-trip preserve - open the original, change one part, and get the format back with everything you did not touch intact.
  • Recomposition - a new document is built in that format from the site's own model. Your file is not read; you get a clean document that resembles the template, not your document back.
  • Native-app edit - a deliberate hand-off: download it and edit it in the named application, which keeps the native layout and formulas.
  • Unsupported - nothing beyond the download.

Producing a format never raises that format's level. Visual Template Studio exports PNG and the audio workspace encodes FLAC, but neither reads one you already have, so both formats stay at open and read.

What genuinely round-trips. Nine ZIP document containers - docx, xlsx, pptx, odt, ods, odp, epub, jar and 3mf - survive an edit through the archive explorer. Every member you did not touch is re-extracted and re-written carrying its original storage flag and entry order, which is the property an EPUB's mimetype entry depends on. That was proven rather than asserted: every ZIP-family fixture in the catalogue was rebuilt and each member's SHA-256 compared against the original. What is still withheld is a document model - you are hand-editing raw OOXML or ODF, and nothing repairs a broken relationship or content-type part for you.

Two limits worth knowing before you type. The text editor normalises CRLF to LF on the first keystroke, because its document model splits on both and rejoins with one; an untouched file is safe, an edited one is not. And only UTF-8, UTF-8 with BOM and UTF-16LE decode - everything else fails the read, deliberately.

That second limit used to be invisible at exactly the wrong moment. Until 2026-09-11 the "Open in editor" link was decided by the extension alone, so 22 fixtures whose own spec sheets declare them undecodable - Latin-1 and Shift-JIS CSVs, UTF-16BE captions, binary STL and PLY meshes, a log full of control characters - offered an editor and delivered the read-only hex inspector instead. The link is now decided by what the fixture says about itself, so those pages no longer make the offer. The fixtures are unchanged: being non-UTF-8 is their entire value.

The full mapping, format by format with the evidence for each, is in docs/EDITOR_CAPABILITY.md in the source repository.

Browser media encoders

The audio and video creation workspaces generate synthetic fixtures only. They never request a camera or microphone, accept an upload, create an account, or send source media to a server. Each job runs in a dedicated module worker that is created only after Generate is pressed and is terminated on completion or cancellation. Recipe limits cap memory and output size.

The workspaces pin Mediabunny and its official MP3, AAC, FLAC, and AC-3 extensions at version 1.51.0. Mediabunny and its extensions use the Mozilla Public License 2.0. The MP3 extension embeds LAME 3.100 under the LGPL; the AAC and AC-3 extensions use size-limited FFmpeg encoder builds; the FLAC extension uses libFLAC. The site does not bundle stock ffmpeg.wasm. Full notices are recorded in docs/MEDIA_ENCODER_NOTICES.md in the source repository.

PCM and pinned encoder-extension paths use the byte reproducibility tier for a fixed recipe and package version. WebCodecs video and browser-provided Ogg encoders use the content tier: recipe, source fingerprint, duration, dimensions, and codec intent are stable, while encoder bytes can vary by browser, operating system, and hardware. The output SHA-256 records the exact artifact produced by the current session; it is not a claim that every conforming browser emits identical bytes.

Local collections and saved drafts

The faceted browser can collect up to 100 files and 200 MiB on the current device. Export produces a deterministic ZIP with a sorted manifest and SHA-256 values. The selection is stored only in browser storage; files are fetched from their existing static download URLs when you export the bundle. No collection, manifest, or fixture is uploaded to a Novus server.

Form and Document Studio can retain a local draft. Visual Studio can retain user-selected image assets in the browser's IndexedDB. These are convenience copies on the current device, not accounts or cloud saves. Clear the draft in its workspace or clear this site's browser data to remove them. Generated media remains in memory only long enough to preview or download it, and object URLs and worker resources are released on completion, cancellation, retry, or page close.

Templates and Visual Template Studio

Templates covers two product surfaces:

  • Document templates: invoices, CVs, spreadsheets, decks, and HTML email starting points (same permanent /templates/... download URLs as before).
  • Visual templates: product studio, lifestyle, marketplace, social, and marketing scenes under /visual-templates. Published families only appear after real previews, size exports, variants, and an editor scene pass validation.

Visual Template Studio opens either a published scene or an empty canvas so you can add or replace images and backgrounds, edit text, tables, shapes, and palette colours, and export PNG/JPEG/WebP/SVG locally. Uploads stay in the browser. Start blank when you want to build every layer yourself, or use a validated template when its tested layout is the better starting point.

The catalog includes 840 scene-backed families, including dedicated Software/SaaS Product Marketing and Travel, Tourism & Hospitality collections. Each detail page explains what the scene is best for, which fields are editable, how to customize it, available exports, and related templates; every visible preview element must have a corresponding editable scene node.

Model inference fixtures

The model inference testing suite groups each validated ONNX model with JSON input and expected-output artifacts. The page contract records tensor shapes, dtypes, opset, numeric tolerance, and the expected operation, while the model page exposes a graph preview. That makes each trio usable as a small runtime compatibility test rather than a generic binary download.

Image transformation suites

Beyond denoise and format ladders, the Images library includes paired suites for photo restoration, inpainting, deblur, super-resolution, and background removal (including hard-edge cases). Use them as inputs and ground truth when measuring filters and models.

Accessibility, motion and languages

The target is WCAG 2.2 Level AA, and the accessibility page states it with its limits attached: that is a target and a self-assessment, not a certification, and the page lists the contrast ratios it was measured against rather than asking you to take the claim on trust. Public routes are scanned with axe in a real browser as part of the release checks, and the same checks run at a phone viewport for the things that only exist once a browser has done layout - the download control sitting above the fold on a file page, the bottom bar staying reachable, and no page scrolling sideways.

Motion is a choice with four settings, not a switch, and the control is on that same page. Match my system is the default and follows your operating system. Full motion runs everything even if your system asks for less. Reduced motion removes decorative movement but keeps the short transitions that tell you a control responded, because losing those loses information rather than motion. No motion animates nothing and silences vibration feedback too. The preference is applied before the first paint, so a viewer who asked for less never sees a burst of animation while a script loads, and a viewer with JavaScript off still has prefers-reduced-motion honoured in plain CSS.

Languages. The interface is served in English only. Languages lists what is supported and what is merely planned, and it derives that from the routes that actually exist, so it cannot advertise a locale this site does not serve. Fixture content is a separate matter and is not English-only: the localization library ships Arabic, Hebrew, Japanese and Chinese variants precisely so bidirectional text and font fallback can be tested.

The catalogue is static and the workspaces run in your browser, so most of what you do here never reaches a server: editing, generating, exporting and collecting are local, and files are fetched from the same public URLs everyone else gets.

Consent is asked differently depending on where you are, and the region comes from your browser's timezone rather than from a lookup of your address - an IP lookup would be exactly the third-party request the consent model exists to avoid. European timezones, plus the EEA territories that do not sit under Europe/, get the opt-in regime: a blocking banner, and nothing non-essential until you choose. Everywhere else gets opt-out: analytics and advertising start granted with a dismissible notice and a one-click refusal. An unreadable timezone falls back to opt-in, because the safe failure is asking someone who did not need to be asked. Global Privacy Control, where your browser sends it, forces the strict default everywhere.

Advertising. The site carries display and native ad units and an export link, all gated on the advertising choice above. Six route groups carry no advertising under any consent state: privacy, cookies, terms, license, accessibility and security policy. Signing in changes none of this in either direction, and an automated check compares the ad slots on a page signed-out against the same page signed-in and requires them to match.

Unsolicited navigation is blocked. Ad networks occasionally try to take a page over with no click behind it. A guard watches navigation events and lets through everything a person actually started - ordinary links, clicked ads, history moves, same-origin and Novus destinations - while cancelling a scripted navigation to a third party that no gesture asked for. The ad-free routes above are exempt, because nothing there could have started one.

The binding terms are on cookies and privacy.

Accounts and saved fixtures

An account is optional and changes nothing about access. Every fixture, every specification and every download is public and stays public; there is no members-only tier and no paywalled file. What an account adds is a saved list that follows you between devices, and a display name you can change.

What is stored. A saved entry is the catalogue id of a fixture and the date you saved it. The file itself is never copied into your account: the catalogue is generated, byte-pinned and served from stable URLs, so the id plus the manifest is the whole truth and a copy could only drift from it. The list is capped at 500 entries, which is a shortlist rather than storage.

If a fixture is retired between saves, its row stays and the account page says the fixture is no longer in the catalogue rather than showing a dead link or quietly dropping the entry.

Signing in. Email and password, or Google where this deployment has an OAuth client configured. A Google sign-in for an address that already has a password account links to that same account rather than creating a second one, because Google verifies the address before asserting it. A forgotten password is reset with one of the recovery codes shown when the account was created, or with a reset link our team sends after checking it is you; an emailed reset link is offered only where outbound mail is actually configured, so the site never tells you to check an inbox that will receive nothing.

Chat. The contact page has a live chat. A help assistant answers first from these guides, quoting the page it found and linking to it, and a person answers when you ask for one. Signed in, your conversations are listed in your account.

Advertising is unchanged when you sign in. No account state removes or reduces advertising anywhere on this site, and an automated test compares the ad slots present signed-out against signed-in on the same page and requires them to match. Cookie and advertising choices remain per browser rather than per account, because they are a property of the device you are reading on; they live on the cookie page, which can also reopen Mediavine's own consent prompt wherever Mediavine asks for one, and signing in does not alter them.

Deleting. The account page deletes the account and its saved list outright, behind a typed confirmation. Nothing about the public catalogue is affected.

How the library is kept honest

Most of the promises above are enforced by the build rather than by anyone remembering them.

  • The manifest and the disk must agree exactly. Every shipped file has a manifest entry and every manifest entry has a file. The build fails on either half being untrue, and it fails on the totals moving at all unless the expected totals move with them - a minimum would let a quiet loss pass, so these are equalities.
  • Bytes are pinned. A committed ledger records the SHA-256 of every file in the catalogue, and a test diffs the disk against it. A regeneration that silently rewrote a fixture fails that test instead of replacing a file someone has already wired into a pipeline.
  • Counts quoted in prose are checked against the catalogue, not recalled. A stale sentence fails the build the same way a stale constant does.
  • Downloads are downloads. Every /files/… response is sent as an attachment. The one exception is streaming media under /files/stream/…, which has to be fetchable inline and cross-origin or an HLS or DASH player cannot read a manifest and its segments at all.
  • An intentionally corrupt file does not get a broken player. A fixture whose extension previews inline but whose bytes deliberately cannot decode renders an information card describing the damage, rather than a video element that can never load and reads as a defect of the site.
  • What machines may read is written down. Robots and AI access names the fetchers allowed by name, what they may reach, and what stays closed; /llms.txt is the machine-readable summary of the site.
  • The brand's accounts are listed, not implied. Social is the one list of every account Novus Stream Solutions runs, and the site's structured data claims that same list rather than a second copy of it.
  • When something breaks, status says what can and cannot go wrong on a site shaped like this one and where to check, and an error page shows neither a stack trace nor an internal message.

Where to start

Browse the Images, Audio, Documents, Data, and Templates libraries, explore visual templates, search from any page, or open a text file directly in the in-browser editor.

Was this page helpful?

Found an error? Send a correction.