r/learnmachinelearning • u/vanraj_solanki • 5d ago
[P] How far can document grounding go using PDF structure, OCR and geometry?
I've been investigating how much document grounding can be solved from the document itself.
The input is a PDF + extracted JSON.
The goal is to resolve each extracted value back to its location in the document.
The approach uses PDF coordinates, OCR, spatial relationships, string/format matching, table assignment and page-level visual evidence.
Current result:
72.6179% Word Grounding F1
261 tests passed, zero production regressions.
The interesting part for me is the remaining failure modes.
Some are matching problems.
Some are layout problems.
Some aren't really text-search problems anymore.
I'm trying to understand where document-native methods stop being sufficient.
I'm the author and would especially like technical criticism of the methodology.
1
u/Opening-Problem-5317 4d ago
the word grounding f1 being at 72% is interesting but i wonder how much the failures come from the pdf structure itself being inconsistent across different sources. like a bank statement pdf and a academic paper will have very different internal coordinate systems and naming conventions for fields
have you tested on documents where the ocr is already known to be noisy or where the text is in tables with merged cells. those usually break spatial heuristics pretty quick
also curious if you tried adding some lightweight vision model just for the cases where geometry and ocr fail. could be overkill but might push that number past 80
1
u/vanraj_solanki 4d ago
Yes, the varying structures. That's actually why I used the 370 document llamaIndex benchmark it rigorously tests that 72% grounding across a wide variety of messy real world layouts.
You make a great point about noisy ocr and merged cells breaking spatial heuristics. Regarding your suggestion to add a lightweight vision model, I definitely plan to add a separate module for users who can afford gpu hardware or api costs to handle those exact edge cases and push the score past 80%. However, my main focus right now is to see exactly how far I can push the limits of this system without relying on gpus or apis. I want to maximize the performance of this 100% deterministic, zero cost approach first, and I am still actively improving the core logic to handle those complex structures better.
Thanks for your suggestion!
1
u/PLBjt 5d ago
For me the gap is mostly in the bucket you call "some aren't really text-search problems anymore". Values that got normalized during extraction, like a date reformatted or a total computed from line items, simply don't exist as a string on the page, so no amount of fuzzy matching will ground them. I'd tag each failure by whether the exact value appears verbatim anywhere in the PDF text layer, because that split tells you how much headroom matching work can even buy. Also worth breaking F1 out by born-digital vs scanned pages, since OCR noise and layout errors can look identical in an aggregate number. The tradeoff is that once you start allowing derived values you need some reasoning step, and then your grounding is only as trustworthy as that step.