KnowledgeTrust-RAG: Reliability-Aware Retrieval for Enterprise Knowledge Management Using Amazon BedrockA Systematic Literature Review

Authors

  • Swapna Pulti Salesforce Solution Manager, Pearson Education Inc Author

DOI:

https://doi.org/10.71238/snnst.v2i02.184

Abstract

Large Language Models are now commonly used for discovering knowledge held within unstructured repositories, but the ability to directly use a generative system to access a proprietary knowledge base suffers from issues with factual hallucination, poor provenance and lack of auditability. To build a conceptual framework, KnowledgeTrust-RAG, for enterprise knowledge management reliability-aware retrieval-augmented generation, this review combines 25 peer-reviewed and preprint sources published up to 2024, using the architectural affordances of a managed generative artificial-intelligence cloud platform like Amazon Bedrock as an illustration. The synthesis connects various previously unexplored components of self-attention-based encoders and dense passage retrieval, the original retrieval-augmented generation formulation, and the self-critiquing and graph-structured models, and summarizes the parallel literatures on hallucination taxonomy, automated evaluation, and enterprise deployment barriers. The results show that reliability of enterprise retrieval-augmented generation is shaped by four interacting processes: dense semantic retrieval, retrieval-necessity gating, global context aggregation via a graph-based approach, and reference-free evaluation metrics like faithfulness and answer relevance. The review also reveals that the variety of evaluation frameworks extends beyond just the number of dimensions, with one framework examining five reliability dimensions while a comparator examines three, and that enterprise-specific challenges, such as data security, access governance and integration complexity, are not sufficiently addressed by generic pipelines. Embedded ingestion, retrieval, orchestration, evaluation and governance layers are proposed, each of which is auditable and could be deployed in the managed cloud separately. The review provides a common classification of the mechanisms for retrieval reliability, a comparison of different architectural variations, and a deployment baseline in a governance perspective which provides a technological basis for designing trustworthy enterprise knowledge systems and can serve as a baseline for future empirical testing on production workloads.

Downloads

Download data is not yet available.

References

M. Alavi and D. E. Leidner, “Review: Knowledge Management and Knowledge Management Systems: Conceptual Foundations and Research Issues,” MIS Q., vol. 25, no. 1, pp. 107–136, 2001, doi: 10.2307/3250961.

Ş. E. Şeker, “Editorial: Large Language Models in Work and Business,” Front. Artif. Intell., vol. 7, 2024, doi: 10.3389/frai.2024.1516832.

S. Zhao, Y. Yang, Z. Wang, Z. He, L. K. Qiu, and L. Qiu, “Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make Your LLMs Use External Data More Wisely,” 2024. doi: 10.48550/arXiv.2409.14924.

S. Gupta, R. Ranjan, and S. N. Singh, “A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions,” 2024. doi: 10.48550/arXiv.2410.12837.

Y. Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” 2023. doi: 10.48550/arXiv.2312.10997.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics, 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.

V. Karpukhin et al., “Dense Passage Retrieval for Open-Domain Question Answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, 2020, pp. 6769–6781. doi: 10.18653/v1/2020.emnlp-main.550.

K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang, “REALM: Retrieval-Augmented Language Model Pre-Training,” in Proceedings of the 37th International Conference on Machine Learning, 2020, pp. 3929–3938. doi: 10.48550/arXiv.2002.08909.

P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems 33, 2020. doi: 10.48550/arXiv.2005.11401.

A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection,” 2023. doi: 10.48550/arXiv.2310.11511.

Z. Guo, L. Xia, Y. Yu, T. Ao, and C. Huang, “LightRAG: Simple and Fast Retrieval-Augmented Generation,” in Findings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp. 10746–10761. doi: 10.18653/v1/2025.findings-emnlp.568.

A. Vaswani et al., “Attention Is All You Need,” in Advances in Neural Information Processing Systems 30, 2017, pp. 5998–6008. doi: 10.48550/arXiv.1706.03762.

N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks,” in Proceedings of EMNLP-IJCNLP 2019, 2019, pp. 3982–3992. doi: 10.18653/v1/D19-1410.

D. Edge et al., “From Local to Global: A Graph RAG Approach to Query-Focused Summarization,” 2024. doi: 10.48550/arXiv.2404.16130.

S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu, “Unifying Large Language Models and Knowledge Graphs: A Roadmap,” IEEE Trans. Knowl. Data Eng., vol. 36, no. 7, pp. 3580–3599, 2024, doi: 10.1109/TKDE.2024.3352100.

T. Bruckhaus, “RAG Does Not Work for Enterprises,” 2024. doi: 10.48550/arXiv.2406.04369.

S. Es, J. James, L. Espinosa Anke, and S. Schockaert, “RAGAs: Automated Evaluation of Retrieval Augmented Generation,” in Proceedings of the 18th EACL: System Demonstrations, 2024, pp. 150–158. doi: 10.18653/v1/2024.eacl-demo.16.

J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, “ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems,” in Proceedings of NAACL 2024, 2024, pp. 338–354. doi: 10.18653/v1/2024.naacl-long.20.

V. Rawte, A. Sheth, and A. Das, “A Survey of Hallucination in Large Foundation Models,” 2023. doi: 10.48550/arXiv.2309.05922.

L. Huang et al., “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” ACM Trans. Inf. Syst., vol. 43, no. 2, pp. 1–55, 2025, doi: 10.1145/3703155.

S. M. T. I. Tonmoy et al., “A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models,” 2024. doi: 10.48550/arXiv.2401.01313.

Z. Ji et al., “Survey of Hallucination in Natural Language Generation,” ACM Comput. Surv., vol. 55, no. 12, pp. 1–38, 2023, doi: 10.1145/3571730.

Y. Liu et al., “Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment,” 2024. doi: 10.48550/arXiv.2308.05374.

X. Huang et al., “A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation,” Artif. Intell. Rev., vol. 57, 2024, doi: 10.1007/s10462-024-10824-0.

Downloads

Published

2025-12-31

How to Cite

KnowledgeTrust-RAG: Reliability-Aware Retrieval for Enterprise Knowledge Management Using Amazon BedrockA Systematic Literature Review. (2025). Sciences Du Nord Nature Science and Technology, 2(02), 72-86. https://doi.org/10.71238/snnst.v2i02.184