LLM-integrated architecture for document intelligence and semantic retrieval
Tóm tắt: 5
|
PDF: 7
##plugins.themes.academic_pro.article.main##
Author
-
Pham Minh TuanThe University of Danang - University of Science and Technology, VietnamNguyen Cong CuongThe University of Danang - University of Science and Technology, VietnamQuoc Khanh TrinhThe University of Danang - University of Science and Technology, VietnamPhuc Hao DoDanang Architecture University, VietnamNguyen Nang Hung VanThe University of Danang - University of Science and Technology, Vietnam
Từ khóa:
Tóm tắt
This study presents an Artificial Intelligence-enabled digital transformation system for intelligent enterprise document management. The proposed platform combines Optical Character Recognition, semantic embeddings, vector search, and Retrieval-Augmented Generation to improve document retrieval and knowledge access in organizational environments. Unlike traditional keyword-based systems, it supports context-aware search and evidence-grounded question answering over internal documents. The system is implemented with a practical full-stack architecture using React, FastAPI, and PostgreSQL with pgvector. To improve retrieval quality, we introduce a K-Means-based filtering step that separates highly relevant results from noisy ones before answer generation. Experimental observations show that this hybrid approach provides more focused outputs and better practical precision while preserving strong ranking quality. The results demonstrate strong potential for scalable enterprise knowledge intelligence applications.
Tài liệu tham khảo
-
[1] C. M. Bishop, Pattern Recognition and Machine Learning. New York: Springer, 2006.
[2] A. K. Jain, “Data clustering: 50 years beyond K-means,” Pattern Recognition Letters, vol. 31, no. 8, pp. 651–666, 2010.
[3] S. Russell and P. Norvig, Artificial Intelligence: A Modern Approach, 4th edition. Harlow, UK: Pearson, 2021.
[4] D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models, 3rd ed., Stanford University, Stanford, CA, USA, 2026. [Online]. Available: https://web.stanford.edu/~jurafsky/slp3/ [Accessed March 12, 2026].
[5] W. X. Zhao et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023.
[6] scikit-learn Developers, “Clustering: K-Means,” scikit-learn documentation, 2025. [Online]. Available: https://scikit-learn.org/stable/modules/clustering.html#k-means [Accessed March 12, 2026].
[7] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781, 2013.
[8] W. Fan et al., “A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models,” in Proc. ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, Barcelona, Spain, 2024, pp. 6491–6501. https://doi.org/10.1145/3637528.3671470
[9] S. V. Rice, G. Nagy and T. A. Nartker, Optical Character Recognition: An Illustrated Guide to the Frontier. Boston, MA: Springer, 1999.
[10] pgvector Developers, “pgvector: Open-source vector similarity search for Postgres,” GitHub repository, 2026. [Online]. Available: https://github.com/pgvector/pgvector [Accessed March 12, 2026].
[11] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” arXiv preprint arXiv:1908.10084, 2019.
[12] Google DeepMind, “Gemma: Open models for responsible AI,” 2024.
[13] OpenAI, “Text embeddings guide,” OpenAI Platform Documentation, 2025. [Online]. Available: https://platform.openai.com/docs/guides/embeddings [Accessed March 12, 2026].
[14] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems, 2020.
[15] R. T. Fielding, Architectural Styles and the Design of Network-based Software Architectures, Doctoral dissertation, University of California, Irvine, Irvine, CA, USA, 2000.
[16] M. Richards, Software Architecture Patterns. Sebastopol, CA: O’Reilly Media, 2015.
[17] S. Ramírez, “FastAPI framework, high performance APIs with Python,” FastAPI Docs, 2025. [Online]. Available: https://fastapi.tiangolo.com [Accessed March 12, 2026].
[18] PostgreSQL Global Development Group, “PostgreSQL 16 Documentation,” PostgreSQL.org, 2025. [Online]. Available: https://www.postgresql.org/docs/16 [Accessed March 12, 2026].
[19] Jaided AI, “EasyOCR: Ready-to-use OCR with 80+ languages,” GitHub, 2025. [Online]. Available: https://github.com/JaidedAI/EasyOCR [Accessed March 12, 2026].
[20] Mistral AI, “OCR,” Mistral Docs, March 6, 2025. [Online]. Available: https://docs.mistral.ai/models/ocr-25-03 [Accessed March 12, 2026].
[21] A. Singhal, “Modern information retrieval: A brief overview,” IEEE Data Engineering Bulletin, vol. 24, no. 4, pp. 35–43, 2001.
[22] O. Khattab and M. Zaharia, “ColBERT: Efficient and effective passage search via contextualized late interaction,” in Proc. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2020, pp. 39–48.
[23] N. N. H. Van, P. H. Do, V. N. Hoang, T. T. K. Nguyen and M. T. Pham, “AI-Powered University Admission Counseling: A Use Case of Large Language Models in Student Guidance,” IEEE Transactions on Learning Technologies, vol. 18, pp. 856-868, 2025, doi: 10.1109/TLT.2025.3604096.

