Steven Gonsalvez

Software Engineer

allenai/olmocr: Toolkit for linearizing PDFs for LLM datasets/training

Why CEREBRO kept it

PDF linearization for LLM training datasets.

The text below is an automated extraction of the article at https://github.com/allenai/olmocr, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (github.com).

A toolkit for converting PDFs and other image-based document formats into clean, readable, plain text format. Try the online demo: https://olmocr.allenai.org/ Features: - Convert PDF, PNG, and JPEG based documents into clean Markdown - Support for equations, tables, handwriting, and complex formatting - Automatically removes headers and footers - Convert into text with a natural reading order, even in the presence of figures, multi-column layouts, and insets - Efficient, less than $200 USD per million pages converted - (Based on a 7B parameter VLM, so it requires a GPU) - October 21, 2025 - v0

Backlinks

Appeared in 1 briefing

Related

Shares tags: ai/llm-mechanics

Also from github.com