CSV with Mixed-Type Columns
A CSV whose 'value' column mixes integers, floats, scientific notation, dates, hex, thousands separators, and whitespace, and whose 'flag' column mixes a dozen boolean spellings. A torture test for type inference and schema detection.
id,value,flag
1,42,true
2,3.14159,false
3,hello,TRUE
4,2026-01-15,0
5,-7,1
6,1e6,yes
7,true,no
8,0x1F,Y
9,,N
10,NaN,t
11, 12 ,f
12,"1,234.56",True
Specifications
- Rows
- 12
- Note
- value mixes int, float, string, date, hex, thousands-sep, whitespace; flag mixes many boolean spellings
Testing contract
Expected to pass- Scenario
- Exercise CSV with Mixed-Type Columns in its edge cases workflow. A CSV whose 'value' column mixes integers, floats, scientific notation, dates, hex, thousands separators, and whitespace, and whose 'flag' column mixes a dozen boolean spellings.
- Expected result
- 12 data records using ',' delimiters and UTF-8; header fields are id, value, flag; data-record widths (columns:count) are {"3":12}.
What is a .csv file?
CSV (Comma-Separated Values) is a plain-text tabular format where rows are lines and fields are separated by commas, with quoting rules for values that contain delimiters, quotes, or newlines. It has no formal type system and depends on encoding and dialect conventions. It is the most portable format for tabular data exchange.
How to use this file
Use an example CSV to test parsers against quoting and embedded-delimiter edge cases, header handling, encoding detection, and import pipelines into databases or spreadsheets.
How to use this file for testing
“CSV with Mixed-Type Columns” is a deterministic Testaroo fixture for CSV parsing, Error handling, Data import. Clean and deliberately messy CSVs, quoted commas, embedded newlines, ragged rows, odd delimiters, and encodings.
Documented properties for this file: 12 rows. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks (delimiters, encodings, null handling, schema, and row counts) in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Feed the file to your parser and assert it handles the documented quirks, quoted delimiters, embedded newlines, ragged rows, or invalid syntax; the valid↔invalid distinction is labelled in the title.
Code examples
import pandas as pd
df = pd.read_csv("mixed-types.csv")
print(df.head())
print(df.dtypes)Generated by generation/data_edgecases.py. Free for any use, no attribution required, license.
Related files
- csvCSV with Mixed Null RepresentationsA CSV where missing values are written seven different ways (empty string, NULL, NA, N/A, null, None, and a dash) across text and numeric columns. A fixture for testing null-detection and coercion in CSV importers.

- csvCSV with Mixed Timestamp FormatsA CSV listing timestamps in eleven formats: ISO 8601 with Z and numeric offsets, millisecond precision, naive local, date-only, Unix epoch in seconds and milliseconds, US AM/PM, and RFC 1123. A fixture for testing date parsing and timezone normalisation.

- csvAlder Table Bistro: Allergen matrix missing a menu itemAllergen matrix missing a menu item for Alder Table Bistro. Intentionally incomplete. 3 rows where the recipes cost 4 menu items. MENU-04 (Potato gratin) has no row, although it has costed recipe lines in plate-costs.csv and sales in sales-mix-pmix.csv. Its correct row would declare: milk: Cream.

- csvAlder Table Bistro: Allergen menu missing a dish on the menuAllergen menu missing a dish on the menu for Alder Table Bistro. Intentionally incomplete. 3 rows where allergen-menu.csv has 4. MENU-02 Tomato soup is printed on printed-menu.pdf at 11.50 CAD and declares 1 allergen in the costing folder, and it has no row here. The remaining rows are correct.

- csvAlder Table Bistro: Cost of goods sold that contradicts the count sheetCost of goods sold that contradicts the count sheet for Alder Table Bistro. Intentionally inconsistent. Row 3 reports a closing count of 22.20 kg for ING-03 (Potatoes), the counted figure with the two decimal digits transposed, where inventory-count-sheet.csv, inventory-variance.xlsx and cogs-report.json all count 22.02 kg. Every derived value follows the wrong count, so the file is internally consistent and its total of 1574.49 CAD differs from the correct 1574.81 CAD by -0.32 CAD.

- csvAlder Table Bistro: Inventory count sheet as transcribedInventory count sheet as transcribed for Alder Table Bistro. The same closing counts as inventory-count-sheet.csv, in the shape a paper count actually arrives in. CRLF line endings, a UTF-8 byte order mark, a semicolon delimiter and a decimal comma. Two blank lines, one section banner that is not a record, one ingredient split across two locations, one counted as whole packs plus a remainder, one superseded mid-shift row at 14:20, trailing whitespace in three fields, a leading-zero bin code, an empty bin code, and one record carrying ten fields where the header declares 9. Every quirk is named in README.md.
