Engineering PapersSearch

DOE OSTI · code-164482

AdaParse

Abstract

SF-25-127AdaParse (Adaptive Parallel PDF Parsing and Resource Scaling Engine) enables scalable, high-accuracy PDF parsing. AdaParse is a data-driven strategy that assigns an appropriate parser to each document, offering high accuracy for any computational budget. Moreover, it offers a workflow of various PDF parsing software that includes extraction tools: PyMuPDF, pypdf traditional OCR: Tesseract, modern OCR (e.g., Vision Transformers): Nougat and Marker

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Siebenschuh, Carlo, Hippe, Kyle, Gokdemir, Ozan, Brace, Alexander, Khan, Arham, Hosssain, MD Khalid, Babuji, YaduNand, Chia, Nicholas, Vishwanath, Venkatram, Stevens, RickL, Foster, IanT, Underwood, Robert. 2025-09-19. AdaParse. https://doi.org/10.11578/dc.20250919.2

Cite the original work for its findings. Save a collection to share your selection of sources.