DOE OSTI · code-171746
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
Abstract
Code for ‘Lost in OCR Translation?’: robust document retrieval under degradation. Compares OCR-based, vision-only, and hybrid pipelines; includes SambaNova LLaMA Vision OCR, Nougat, and ViDoRe baselines. Provides QA data generation, RAG evaluation, and metrics (Levenshtein, nDCG@k, Recall@k, EM/F1) with reproducible scripts. Includes dataset guides
Keep this discovery
Explore connections, maps & timelines
Bhattarai, Manish [Los Alamos National Labs], Most, Alexander. 2025-08-27. Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval. https://doi.org/10.11578/dc.20251212.7
Cite the original work for its findings. Save a collection to share your selection of sources.