| 1 |
An, Z., Ding, X., Fu, Y.-C., Chu, C.-C., Li, Y., and Du, W. (2024). Golden-retriever:
High-fidelity agentic retrieval augmented generation for industrial knowledge base.
arXiv preprint arXiv:2408.00798.
|
|
| 2 |
DeepMind, G. (2025a). Gemini 3 flash: Frontier intelligence built for speed. https:
//blog.google/products/gemini/gemini-3-flash.
|
|
| 3 |
DeepMind, G. (2025b). Gemini 3 pro: State-of-the-art multimodal ai model. https:
//blog.google/products/gemini/gemini-3.
|
|
| 4 |
Gao, S., Zhao, S., Jiang, X., Duan, L., Chng, Y. X., Chen, Q.-G., Luo, W., Zhang, K.,
Bian, J.-W., and Gong, M. (2025). Scaling beyond context: A survey of multimo-
dal retrieval-augmented generation for document understanding. arXiv preprint ar-
Xiv:2510.15253.
|
|
| 5 |
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang,
H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv
preprint arXiv:2312.10997.
|
|
| 6 |
Gu, J., Kuen, J., Morariu, V. I., Zhao, H., Barmpalios, N., Jain, R., Nenkova, A., and Sun,
T. (2021). Unified pretraining framework for document understanding. In Advances in
Neural Information Processing Systems (NeurIPS 2021).
|
|
| 7 |
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and
Park, S. (2022). Ocr-free document understanding transformer. In Proceedings of the
17th European Conference on Computer Vision (ECCV 2022), pages 498–517, Berlin,
Heidelberg. Springer-Verlag.
|
|
| 8 |
Mandal, S., Talewar, A., Ahuja, P., and Juvatkar, P. (2025). Nanonets-ocr-s: A mo-
del for transforming documents into structured markdown with intelligent content
recognition and semantic tagging. https://huggingface.co/nanonets/
Nanonets-OCR-s
|
|
| 9 |
McKie, J. X. (2024). PyMuPDF: Python bindings for the MuPDF library. Version 1.24.x
|
|
| 10 |
Mei, L., Mo, S., Yang, Z., and Chen, C. (2025). A survey of multimodal retrieval-
augmented generation. arXiv preprint arXiv:2504.08748.
|
|
| 11 |
Qwen (2025). Qwen2.5-vl-3b: Instruction-tuned vision-language model. https://
huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct
|
|
| 12 |
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., and Zhou, M. (2020). Layoutlm: Pre-training
of text and layout for document image understanding. In Proceedings of the 26th ACM
SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD
’20, pages 1192–1200, New York, NY, USA. Association for Computing Machinery
|
|