Convert v2 ORC Employee Table Source
Binary orc source for the five-row P8 employee conversion table, preserving ids, names, departments, booleans, and scores. Stable P8 artifact p8-convert-orc-source.
| id | name | department | active | score | joined | |
|---|---|---|---|---|---|---|
| 1001 | Ada Lovelace | ada.lovelace@example.com | Engineering | true | 98.5 | 2021-03-01 |
| 1002 | Alan Turing | alan.turing@example.com | Research | true | 95 | 2020-06-15 |
| 1003 | Grace Hopper | grace.hopper@example.com | Engineering | false | 91.2 | 2019-11-20 |
| 1004 | Katherine Johnson | katherine.johnson@example.com | Operations | true | 96.8 | 2022-01-10 |
| 1005 | Edsger Dijkstra | edsger.dijkstra@example.com | Research | false | 89.4 | 2018-09-05 |
Specifications
- Rows
- 5
- Columns
- 7
- Source Format
- orc
- Delivery Mode
- download-only
- Provider
- converter-v2
- Provenance
- Synthetic deterministic P8 fixture generated by generation/p8_content.py; seed namespace 2026082300
- Fixture Reserve
- convert-v2
Testing contract
Expected to pass- Scenario
- Read the ORC artifact and project the five contract columns in row order.
- Expected result
- Five rows and all seven columns match ids 1001 through 1005; Ada's email and 2021-03-01 join date survive, Grace is inactive, and scores remain exact.
What is a .orc file?
Apache ORC (Optimized Row Columnar, .orc) is a binary columnar format from the Hadoop ecosystem. It stores data in stripes with lightweight indexes, per-column compression, and embedded statistics, and is common in Hive and big-data pipelines.
How to use this file
Use an example .orc file to test ORC readers, stripe and index handling, and ORC-to-Parquet/CSV conversion.
How to use this file for testing
“Convert v2 ORC Employee Table Source” is a deterministic Testaroo fixture for Conversion testing, Data engineering, Serialization testing. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.
Documented properties for this file: 5 rows · 7 columns. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks (delimiters, encodings, null handling, schema, and row counts) in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Generated by generation/p8_content.py. Free for any use, no attribution required, license.
Related files
- pbConvert v2 Protobuf Creator RecordValid Protobuf wire record with string, uint32, and string fields for schema-guided decoding. Stable P8 artifact p8-convert-protobuf-source.

- jsonConvert v2 Protobuf Expected JSONExpected JSON semantic result for the schema-guided Protobuf decode. Stable P8 artifact p8-convert-protobuf-expected.

- pbConvert v2 Protobuf Missing-Schema Controlled FailureValid wire bytes intentionally supplied without a matching schema so tools must preserve unknown field 15. Stable P8 artifact p8-convert-protobuf-missing-schema.

- protoConvert v2 Protobuf Schema CompanionProto3 schema companion declaring the three fields used by the binary creator record. Stable P8 artifact p8-convert-protobuf-schema.

- avroAvro: Row Binary + SchemaThe same records as Apache Avro: a compact row-based binary format that embeds its own schema, widely used in Kafka pipelines. For testing Avro decoders and schema evolution.

- bsonBSON: Binary JSON (MongoDB)The records as BSON: the binary-JSON encoding MongoDB stores documents in. For testing BSON decoders and JSON↔BSON conversion.
