Scientific
Research software reads formats that predate most of the web and rarely come with a clean sample. This category ships small, valid, fully synthetic ones. Chemistry covers MDL molfiles and SDfiles, XYZ coordinates, SMILES strings, crystallographic CIF, and PDB structures for an invented molecule. Bioinformatics covers FASTA and FASTQ sequences with documented quality encodings, SAM alignments, BED intervals, GFF3 annotations, and Newick phylogenetic trees. Numerical and gridded data arrive as NetCDF, HDF5, FITS, and Matrix Market alongside CSV twins so a loader can be checked against readable values. Citation formats (BibTeX, RIS) complete the set. Nothing here describes a real sample, organism, patient, or observation, the values are generated from fixed seeds and documented in each file's spec.
Filter scientific on Browse · 128 files · 11 subcategories
Frequently asked questions
What chemistry and bioinformatics formats are included?+
MDL molfiles and SDfiles, XYZ, SMILES, CIF, and PDB on the chemistry side; FASTA, FASTQ, SAM, BED, GFF3, and Newick trees on the bioinformatics side.
Is any of this real experimental data?+
No. Every molecule, sequence, structure, and measurement is generated from a fixed seed and documented in the spec table, no real organism, patient, sample, or observation is described.
Can I read the binary formats without special tooling?+
NetCDF, HDF5, and FITS fixtures ship with a documented variable/header listing and, where useful, a CSV twin carrying the same values in readable form.