The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word.
I'm also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF.
The project is open source and I'd love feedback on the architecture, extraction quality, and useful use cases.