Sample files for testing parsers and error handling

A parser test has two halves: files that must load, and load correctly, and files that must not load or must fail cleanly. The valid files below carry props in their manifest entries that say what a correct read finds. The broken ones carry an outcome, shown beside each link: must-fail, may-recover or varies.

Proving a file loaded correctly

https://loremfile.dev/pdf/a4-3pages.pdf · 4.4 KB (4,351 bytes) · application/pdf

Its manifest entry lists contains_text, strings taken from the text it was generated from. Assert that each one appears in what your extractor returns, after collapsing whitespace, and the test proves the pages were read rather than that the call returned.

https://loremfile.dev/docx/with-tracked-changes.docx · 37.4 KB (37,393 bytes) · application/vnd.openxmlformats-officedocument.wordprocessingml.document · 3 paragraphs · 12 words

One deletion, one insertion and one comment. Joining all the XML text returns the deleted word as if it were current; the entry's props record the deleted, inserted and current text to compare with.

https://loremfile.dev/pdf/scanned-1page.pdf · 107 KB (107,030 bytes) · application/pdf

A page that looks readable and has no text layer, so extraction returns an empty string. A pipeline should notice and send it to OCR rather than store an empty document.

https://loremfile.dev/json/all-types.json · 708 bytes · application/json; charset=utf-8 · utf-8 · LF · 23 lines

Integers beyond 2^53 and 2^63, negative zero and a surrogate pair: the values that change on a round trip through a double or a careless string decoder.

Text that splits in the wrong place

https://loremfile.dev/csv/people-10-quoted-newlines.csv · 2.7 KB (2,697 bytes) · text/csv; charset=utf-8 · 10 rows · 14 columns · utf-8 · LF · 20 lines

A quoted field holding a real newline. Splitting the file into lines before parsing it as CSV cuts that record in two.

https://loremfile.dev/csv/people-10-utf8-bom.csv · 2.7 KB (2,682 bytes) · text/csv; charset=utf-8 · 10 rows · 14 columns · utf-8 · LF · 11 lines

A UTF-8 byte-order mark, which many readers leave attached to the first column name, so a lookup by that name fails.

https://loremfile.dev/txt/utf16be-no-bom.txt · 3.6 KB (3,614 bytes) · text/plain; charset=utf-16be · utf-16be · LF · 5 lines

UTF-16 big-endian with no byte-order mark for a detector to find. Code that decides the encoding from a BOM reads it as the wrong one.

Files that must fail

https://loremfile.dev/edge/json-trailing-comma.json · 64 bytes · application/json; charset=utf-8 · must-fail

Commas before closing brackets: valid JavaScript, invalid JSON. A conforming JSON parser refuses it, so a reader that accepts it is not parsing JSON.

https://loremfile.dev/edge/xml-unclosed-tag.xml · 122 bytes · application/xml; charset=utf-8 · must-fail

An element that is never closed. XML 1.0 requires a parser to stop rather than guess where the element ends.

https://loremfile.dev/edge/zero-byte.json · 0 bytes · application/json; charset=utf-8 · must-fail

No bytes at all, where JSON requires a value. Check that the error says the file is empty, not something about an unexpected end of input.

Files that may be recovered

https://loremfile.dev/edge/zip-truncated-50pct.zip · 911 bytes · application/zip · may-recover

The first half of a zip: the local headers survive and the central directory does not. A reader that walks local headers recovers the surviving entries; one that starts from the central directory finds nothing.

https://loremfile.dev/edge/pdf-truncated-60pct.pdf · 2.6 KB (2,610 bytes) · application/pdf · may-recover

The first part of a PDF, cut before its cross-reference table and trailer. Readers that rebuild the table by scanning the file recover the pages that survive.

https://loremfile.dev/edge/utf8-invalid-bytes.txt · 143 bytes · text/plain; charset=utf-8 · may-recover

Text served as UTF-8 with byte sequences UTF-8 does not allow. Strict decoding fails, and replacing the bad bytes recovers the rest. Either is right; silently dropping lines is not.

Files readers disagree on

https://loremfile.dev/edge/json-bom.json · 80 bytes · application/json; charset=utf-8 · varies

Valid JSON behind a byte-order mark that RFC 8259 does not allow. Some parsers skip the mark and some refuse the file, so pin down which behaviour your code relies on.

https://loremfile.dev/edge/csv-ragged-rows.csv · 65 bytes · text/csv; charset=utf-8 · varies

Rows with fewer and more fields than the header. RFC 4180 says rows should match, not what a reader does when they do not: pad, trim or refuse, and test that your code does the one you chose.

Limits

https://loremfile.dev/json/nested-100-levels.json · 1.4 KB (1,380 bytes) · application/json; charset=utf-8 · utf-8 · LF · 1 line

An object nested 100 levels deep. A recursive parser with a lower depth limit should refuse it with an error rather than crash.

https://loremfile.dev/txt/very-long-line-1mb.txt · 1 MB (1,000,000 bytes) · text/plain; charset=utf-8 · utf-8 · LF · 1 line

A single line filling the whole file, for line-based readers that hold a line in a fixed buffer.

Every broken file, with the exact damage and the valid file to compare it with, is on the edge cases page.