Large Language Model-Based Program Comprehension for Enterprise Software Repositories: A Context-Aware Framework for Code Summarization, Dependency Analysis, and Impact Prediction
DOI:
https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I4P127Keywords:
Large Language Models, Program Comprehension, Software Repository Mining, Code Summarization, Dependency Analysis, Change Impact Prediction, Retrieval-Augmented Generation, Enterprise Software Engineering, Software Maintenance, Technical DebtAbstract
Enterprise software repositories contain complex, long-lived, and continuously evolving code assets whose behavior is distributed across services, frameworks, configuration files, database schemas, event streams, and deployment scripts. Conventional program comprehension tools provide useful static analysis, search, and dependency visualization capabilities, but they often fail to explain why a code element exists, how a change propagates across enterprise boundaries, and which downstream artifacts are likely to be affected by a modification. Large language models have recently demonstrated strong ability in source code understanding, documentation generation, and natural-language reasoning over program artifacts. However, direct application of generic LLMs to enterprise repositories is insufficient because enterprise comprehension requires contextual grounding, dependency-aware retrieval, architectural constraints, security awareness, and traceable impact prediction. This paper proposes RepoComprehend-LLM, a context-aware framework for LLM-based program comprehension in enterprise software repositories. The framework integrates repository mining, static dependency extraction, semantic code representation, retrieval-augmented generation, graph-based impact propagation, and confidence-calibrated explanation generation. Unlike isolated method-level summarization, the proposed framework generates hierarchical summaries at method, class, service, API, and business-capability levels; constructs multi-layer dependency graphs across code, configuration, database, event, and deployment artifacts; and predicts change impact using structural, semantic, historical, and developer-activity signals. The paper presents the architecture, methodology, scoring functions, evaluation protocol, implementation considerations, and enterprise governance controls for deploying such a framework in regulated software engineering environments. The central contribution is a research-oriented, implementation-ready approach that treats program comprehension as a context-aware, evidence-grounded, and continuously updated intelligence layer for software maintenance, modernization, and release risk assessment.
References
1. Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A Pre-Trained Model for Programming and Natural Languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, Online, 2020, pp. 1536–1547, https://doi.org/10.18653/v1/2020.findings-emnlp.139.
2. Sivva, S. D., Thalakanti, R. R., Bandari, S. S. G., & Yettapu, S. D. R. (2023). AI-Driven Decision Intelligence for Agile Software Lifecycle Governance: An Architecture-Centered Framework Integrating Machine Learning Defect Prediction and Automated Testing. International Journal of Emerging Trends in Computer Science and Information Technology, 4(4), 167-172. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I4P118
3. D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. K. Deng, C. Fu, and M. Zhou, “GraphCodeBERT: Pre-training Code Representations with Data Flow,” in International Conference on Learning Representations, 2021, https://openreview.net/forum?id=jLoC4ez43PZ.
4. Thalakanti, R. R., & Goud Bandari, S. S. (2024). Intelligent Continuous Integration and Delivery for Banking Systems using Machine Learning Driven Risk Detection with Real World Deployment Evaluation. International Journal of AI, BigData, Computational and Management Studies, 5(4), 168-175. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I4P118
5. Y. Wang, W. Wang, S. Joty, and S. C. H. Hoi, “CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 8696–8708, https://doi.org/10.18653/v1/2021.emnlp-main.685.
6. Gudi, S. R. (2024). AI-Driven Fax-to-Digital Prescription Automation: A Cloud-Native Framework Using OCR, Machine Learning, and Microservices for Pharmacy Operations. International Journal of Emerging Research in Engineering and Technology, 5(1), 111-116. https://doi.org/10.63282/3050-922X.IJERET-V5I1P113
7. D. Guo, S. Lu, N. Duan, Y. Wang, M. Zhou, and J. Yin, “UniXcoder: Unified Cross-Modal Pre-training for Code Representation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland, 2022, pp. 7212–7225, https://doi.org/10.18653/v1/2022.acl-long.499.
8. Mutyam, N. (2024). Graph-based modeling of service dependencies for predicting failure propagation in distributed systems. International Journal of Multidisciplinary Evolutionary Research, 5(1), 113–116. https://doi.org/10.54660/IJMER.2024.5.1.113-116
9. S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang, G. Li, L. Zhou, L. Shou, L. Zhou, M. Tufano, M. Gong, M. Zhou, N. Duan, N. Sundaresan, S. K. Deng, S. Fu, and S. Liu, “CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” Proceedings of Machine Learning and Systems Datasets and Benchmarks Track, 2021, https://arxiv.org/abs/2102.04664.
10. Gunda SK, Yettapu SDR, Bodakunti S, Bikki SB. Decision Intelligence Methodology for AI-Driven Agile Software Lifecycle Governance and Architecture-Centered Project Management, 2023 Mar. 30;4(1):102-8. https://doi.org/10.63282/3050-9262.IJAIDSML-V4I1P112
11. H. Husain, H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt, “CodeSearchNet Challenge: Evaluating the State of Semantic Code Search,” arXiv preprint arXiv:1909.09436, 2019, https://arxiv.org/abs/1909.09436.
12. Gudi, S. R. (2023). Enhancing Reliability in Java Enterprise Systems through Comparative Analysis of Automated Testing Frameworks. International Journal of Emerging Trends in Computer Science and Information Technology, 4(2), 151-160. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I2P115
13. T. Liu, C. Xu, and J. McAuley, “RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems,” arXiv preprint arXiv:2306.03091, 2023, https://arxiv.org/abs/2306.03091.
14. Thalakanti, R. R., Goud Bandari, S. S., & Sivva, S. D. . (2024). Federated Learning for Privacy Preserving Fraud Detection across Financial Institutions: Architecture Protocols and Operational Governance. International Journal of Emerging Research in Engineering and Technology, 5(2), 108-114. https://doi.org/10.63282/3050-922X.IJERET-V5I2P111
15. R. Bairi, A. Sonwane, A. Kanade, V. D. C., A. Iyer, S. Parthasarathy, S. Rajamani, B. Ashok, and S. Shet, “CodePlan: Repository-level Coding using LLMs and Planning,” arXiv preprint arXiv:2309.12499, 2023, https://arxiv.org/abs/2309.12499.
16. Gudi, S. R. (2024). Design and Evaluation of Secure Microservices Architecture for HIPAA-Compliant Prescription Processing on AWS and OpenShift. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(2), 144-149. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I2P116
17. S. Lehnert, “A Taxonomy for Software Change Impact Analysis,” in Proceedings of the 12th International Workshop on Principles of Software Evolution and the 7th Annual ERCIM Workshop on Software Evolution, 2011, pp. 41–50, https://doi.org/10.1145/2024445.2024454.
18. Bandari, S. S. G., Sivva, S. D., & Thalakanti, R. R. (2024). Regulatory Grade Fraud Detection using Explainable Artificial Intelligence with Auditable Decision Pathways and Empirical Validation on Banking Data. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(3), 139-147. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I3P115
19. Y. Wang, H. Le, A. D. Gotmare, N. D. Q. Bui, J. Li, and S. C. H. Hoi, “CodeT5+: Open Code Large Language Models for Code Understanding and Generation,” arXiv preprint arXiv:2305.07922, 2023, https://arxiv.org/abs/2305.07922.
20. Gunda, S. K. G. (2023). The Future of Software Development and the Expanding Role of ML Models. International Journal of Emerging Research in Engineering and Technology, 4(2), 126-129. https://doi.org/10.63282/3050-922X.IJERET-V4I2P113
21. U. Alon, M. Zilberstein, O. Levy, and E. Yahav, “code2vec: Learning Distributed Representations of Code,” Proceedings of the ACM on Programming Languages, 3(POPL), 2019, pp. 1–29, https://doi.org/10.1145/3290353.
22. Gudi, S. R. (2024). Leveraging Predictive Analytics and Redis-Backed Caching to Optimize Specialty Medication Fulfillment and Pharmacy Inventory Management. International Journal of AI, BigData, Computational and Management Studies, 5(3), 155-160. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I3P116
23. M. Allamanis, E. T. Barr, P. Devanbu, and C. Sutton, “A Survey of Machine Learning for Big Code and Naturalness,” ACM Computing Surveys, 51(4), 2018, pp. 1–37, https://doi.org/10.1145/3212695.
24. Sivva, S. D. (2023). An end-to-end AI-based systems engineering paradigm for lifecycle governance, predictive quality assurance, automation economics, and cybersecurity intelligence. Journal of Frontiers in Multidisciplinary Research, 4(1), 600–604. https://doi.org/10.54660/.JFMR.2023.4.1.600-604
25. G. Sridhara, E. Hill, D. Muppaneni, L. Pollock, and K. Vijay-Shanker, “Towards Automatically Generating Summary Comments for Java Methods,” in Proceedings of the IEEE/ACM International Conference on Automated Software Engineering, 2010, pp. 43–52, https://doi.org/10.1145/1858996.1859006.
26. S. K. Gunda, “Comparative Analysis of Machine Learning Models for Software Defect Prediction,” 2024 International Conference on Power, Energy, Control and Transmission Systems (ICPECTS), Chennai, India, 2024, pp. 1-6, https://doi.org/10.1109/ICPECTS62210.2024.10780167.
27. A. Marcus, A. Sergeyev, V. Rajlich, and J. I. Maletic, “An Information Retrieval Approach to Concept Location in Source Code,” in Proceedings of the 11th Working Conference on Reverse Engineering, 2004, pp. 214–223, https://doi.org/10.1109/WCRE.2004.10.
28. Balerao, M. (2023). A converged artificial intelligence architecture for innovation, software lifecycle optimization, and cybersecurity risk mitigation. International Journal of Multidisciplinary Futuristic Development, 4(1), 117–120. https://doi.org/10.54660/IJMFD.2023.4.1.117-120
29. S. Haque, T. Ahmed, and P. Devanbu, “Context-aware Code Summary Generation,” arXiv preprint arXiv:2408.09006, 2024, https://arxiv.org/abs/2408.09006.
30. Manga I, Sivva SD, Manga VK. The Adaptive Intelligence in Cloud Systems: A Unified Architecture for AI Enhanced Observability and Automated Root Cause Analysis. 2024 Mar. 30 5(1):160-6. https://ijaidsml.org/index.php/ijaidsml/article/view/366
31. S. K. Gunda, “Fault Prediction Unveiled: Analyzing the Effectiveness of Random Forest, Logistic Regression, and KNeighbors,” 2024 2nd International Conference on Self Sustainable Artificial Intelligence Systems (ICSSAS), Erode, India, 2024, pp. 107-113, https://doi.org/10.1109/ICSSAS64001.2024.10760620.
32. “Enhancing Code Understanding for Impact Analysis by Combining Dependence Graph Information and Conceptual Coupling,” ACM Transactions on Software Engineering and Methodology, 2024, https://doi.org/10.1145/3643770.