fig3
Figure 3. Two document-parsing paradigms for materials-science papers. (A) Modular OCR- and rule-based pipelines separate layout, table, chart, and chemical-structure extraction; (B) Unified end-to-end multimodal LLM parsing directly maps page images or parsed blocks to machine-readable structured output. The two routes differ in controllability, error propagation, cost, and traceability. OCR: Optical character recognition; LLM: large language model.






