
Baidu's Unlimited OCR Processes Long Documents in One Pass, Mimicking Human Parsing Memory
Baidu released Unlimited OCR on 23 June 2026, a model designed to parse multi-page documents in a single inference pass by maintaining semantic state across pages—solving a core limitation of chunking-based OCR. The weights and code are immediately available on GitHub and Hugging Face. Conventional systems fragment documents, losing context between pages; this architecture targets long-horizon inputs where information from page three disambiguates content on page thirty.
Published