Sample files for testing parsers and error handling
A parser test has two halves: files that must load, and load correctly, and files that must
not load or must fail cleanly. The valid files below carry props in their manifest entries
that say what a correct read finds. The broken ones carry an outcome, shown beside each
link: must-fail, may-recover or varies.
Proving a file loaded correctly
https://loremfile.dev/pdf/a4-3pages.pdf · 4.4 KB (4,351 bytes) · application/pdf
Its manifest entry lists contains_text, strings taken from the text it was generated from.
Assert that each one appears in what your extractor returns, after collapsing whitespace,
and the test proves the pages were read rather than that the call returned.
https://loremfile.dev/docx/with-tracked-changes.docx · 37.4 KB (37,393 bytes) · application/vnd.openxmlformats-officedocument.wordprocessingml.document · 3 paragraphs · 12 words
One deletion, one insertion and one comment. Joining all the XML text returns the deleted word as if it were current; the entry's props record the deleted, inserted and current text to compare with.
https://loremfile.dev/pdf/scanned-1page.pdf · 107 KB (107,030 bytes) · application/pdf
A page that looks readable and has no text layer, so extraction returns an empty string. A pipeline should notice and send it to OCR rather than store an empty document.
https://loremfile.dev/json/all-types.json · 708 bytes · application/json; charset=utf-8 · utf-8 · LF · 23 lines
Integers beyond 2^53 and 2^63, negative zero and a surrogate pair: the values that change on a round trip through a double or a careless string decoder.
Text that splits in the wrong place
https://loremfile.dev/csv/people-10-quoted-newlines.csv · 2.7 KB (2,697 bytes) · text/csv; charset=utf-8 · 10 rows · 14 columns · utf-8 · LF · 20 lines
A quoted field holding a real newline. Splitting the file into lines before parsing it as CSV cuts that record in two.
https://loremfile.dev/csv/people-10-utf8-bom.csv · 2.7 KB (2,682 bytes) · text/csv; charset=utf-8 · 10 rows · 14 columns · utf-8 · LF · 11 lines
A UTF-8 byte-order mark, which many readers leave attached to the first column name, so a lookup by that name fails.
https://loremfile.dev/txt/utf16be-no-bom.txt · 3.6 KB (3,614 bytes) · text/plain; charset=utf-16be · utf-16be · LF · 5 lines
UTF-16 big-endian with no byte-order mark for a detector to find. Code that decides the encoding from a BOM reads it as the wrong one.
Files that must fail
https://loremfile.dev/edge/json-trailing-comma.json · 64 bytes · application/json; charset=utf-8 · must-fail
Commas before closing brackets: valid JavaScript, invalid JSON. A conforming JSON parser refuses it, so a reader that accepts it is not parsing JSON.
https://loremfile.dev/edge/xml-unclosed-tag.xml · 122 bytes · application/xml; charset=utf-8 · must-fail
An element that is never closed. XML 1.0 requires a parser to stop rather than guess where the element ends.
https://loremfile.dev/edge/zero-byte.json · 0 bytes · application/json; charset=utf-8 · must-fail
No bytes at all, where JSON requires a value. Check that the error says the file is empty, not something about an unexpected end of input.
Files that may be recovered
https://loremfile.dev/edge/zip-truncated-50pct.zip · 911 bytes · application/zip · may-recover
The first half of a zip: the local headers survive and the central directory does not. A reader that walks local headers recovers the surviving entries; one that starts from the central directory finds nothing.
https://loremfile.dev/edge/pdf-truncated-60pct.pdf · 2.6 KB (2,610 bytes) · application/pdf · may-recover
The first part of a PDF, cut before its cross-reference table and trailer. Readers that rebuild the table by scanning the file recover the pages that survive.
https://loremfile.dev/edge/utf8-invalid-bytes.txt · 143 bytes · text/plain; charset=utf-8 · may-recover
Text served as UTF-8 with byte sequences UTF-8 does not allow. Strict decoding fails, and replacing the bad bytes recovers the rest. Either is right; silently dropping lines is not.
Files readers disagree on
https://loremfile.dev/edge/json-bom.json · 80 bytes · application/json; charset=utf-8 · varies
Valid JSON behind a byte-order mark that RFC 8259 does not allow. Some parsers skip the mark and some refuse the file, so pin down which behaviour your code relies on.
https://loremfile.dev/edge/csv-ragged-rows.csv · 65 bytes · text/csv; charset=utf-8 · varies
Rows with fewer and more fields than the header. RFC 4180 says rows should match, not what a reader does when they do not: pad, trim or refuse, and test that your code does the one you chose.
Limits
https://loremfile.dev/json/nested-100-levels.json · 1.4 KB (1,380 bytes) · application/json; charset=utf-8 · utf-8 · LF · 1 line
An object nested 100 levels deep. A recursive parser with a lower depth limit should refuse it with an error rather than crash.
https://loremfile.dev/txt/very-long-line-1mb.txt · 1 MB (1,000,000 bytes) · text/plain; charset=utf-8 · utf-8 · LF · 1 line
A single line filling the whole file, for line-based readers that hold a line in a fixed buffer.
Every broken file, with the exact damage and the valid file to compare it with, is on the edge cases page.