Intelligent Test Generation Using Large Language Models in Microservices-Based Enterprise Applications
DOI:
https://doi.org/10.63282/3050-9416.IJAIBDCMS-V4I1P118Keywords:
Large Language Models, Test Generation, Microservices, Spring Boot, AI-Augmented Testing, Code Coverage, Contract Testing, CI/CD, Software Quality, Prompt EngineeringAbstract
Automated software testing has become a fundamental requirement for maintaining quality, reliability, and scalability in modern enterprise applications. The rapid adoption of microservices architecture has significantly improved system modularity and deployment flexibility; however, it has introduced complex testing challenges due to distributed service interactions, dynamic API contracts, independent deployment cycles, and heterogeneous technology stacks. Traditional automated testing approaches often fail to generate comprehensive test cases for microservices because they lack contextual understanding of service behavior, business logic, and inter-service dependencies. This research presents an intelligent test generation framework using Large Language Models (LLMs) to automate the creation of unit tests, integration tests, and contract tests for microservices-based enterprise applications. The proposed approach leverages LLM-based code generation capabilities combined with service metadata, API specifications, dependency graphs, historical defect information, and application source code to produce context-aware test suites. The framework focuses on Spring Boot-based microservices commonly used in enterprise environments and integrates AI-driven test generation into continuous integration and continuous deployment (CI/CD) workflows. Unlike conventional test automation methods that depend heavily on manually written test scripts, the proposed model analyzes application semantics, identifies potential failure scenarios, and generates optimized test cases with higher coverage. The research methodology includes extraction of microservice architecture information, preprocessing of service artifacts, prompt engineering strategies, LLM-based test generation, automated validation, and feedback-based refinement. The generated test cases are evaluated using software quality metrics including code coverage, defect detection rate, execution success rate, and developer productivity improvement. Experimental analysis on enterprise retail microservices demonstrates that LLM-assisted testing improves overall testing efficiency by reducing manual effort and increasing coverage of complex service interactions. The study further investigates the role of LLMs in handling API evolution, contract validation, and distributed transaction testing. Results indicate that intelligent test generation provides significant benefits in accelerating software delivery while maintaining reliability. The proposed framework establishes a practical foundation for integrating generative artificial intelligence into enterprise software quality engineering. The findings suggest that LLM-based testing can become an essential component of future DevOps ecosystems by enabling adaptive, scalable, and intelligent software validation.
References
1. Myers, G. J., Badgett, T., Thomas, T. M., & Sandler, C. (2004). The art of software testing (Vol. 2). Chichester: John Wiley & Sons.
2. Ammann, P., & Offutt, J. (2017). Introduction to software testing. Cambridge University Press.
3. Tillmann, N., & De Halleux, J. (2008, April). Pex–white box test generation for. net. In International conference on tests and proofs (pp. 134-153). Berlin, Heidelberg: Springer Berlin Heidelberg.
4. Fraser, G., & Arcuri, A. (2011, September). Evosuite: automatic test suite generation for object-oriented software. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering (pp. 416-419).
5. McMinn, P. (2004). Search‐based software test data generation: a survey. Software testing, Verification and reliability, 14(2), 105-156.
6. Kumar, M. S. (2022). An AI-Driven Framework for Data Governance, Quality Management, and Metadata Integration in Enterprise Systems. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 3(2), 165-175.
7. Aluri, Y. S. (2021). Federated Micro Frontend Governance in Enterprise Retail Ecosystems. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 2(2), 114-125.
8. Yuvaraj, N. (2022). LLM-Augmented Conversational Intelligence for Customer Workflow Continuity. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 3(4), 171-183.
9. Cherukuri, R., & Putchakayala, R. (2021). Frontend-Driven Metadata Governance: A Full-Stack Architecture for High-Quality Analytics and Privacy Assurance. International Journal of Emerging Research in Engineering and Technology, 2(3), 95-108.
10. Yallavula, R., & Putchakayala, R. (2022). A Data Governance and Analytics-Enhanced Approach to Mitigating Cyber Threats in NoSQL Database Systems. International Journal of Emerging Trends in Computer Science and Information Technology, 3(3), 90-100.
11. Harman, M., Jia, Y., & Zhang, Y. (2015, April). Achievements, open problems and challenges for search based software testing. In 2015 IEEE 8th international conference on software testing, verification and validation (ICST) (pp. 1-12). IEEE.
12. Catal, C. (2011). Software fault prediction: A literature review and current trends. Expert systems with applications, 38(4), 4626-4636.
13. Kamei, Y., Shihab, E., Adams, B., Hassan, A. E., Mockus, A., Sinha, A., & Ubayashi, N. (2012). A large-scale empirical study of just-in-time quality assurance. IEEE Transactions on Software Engineering, 39(6), 757-773.
14. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. nature, 521(7553), 436-444.
15. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., ... & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
16. Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., ... & Amodei, D. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33, 1877-1901.
17. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., ... & Zaremba, W. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
18. Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., ... & Zhou, M. (2020, November). Codebert: A pre-trained model for programming and natural languages. In Findings of the association for computational linguistics: EMNLP 2020 (pp. 1536-1547).
19. Riccio, V., Jahangirova, G., Stocco, A., Humbatova, N., Weiss, M., & Tonella, P. (2020). Testing machine learning based systems: a systematic mapping. Empirical Software Engineering, 25(6), 5193-5254.
20. Newman, Sam. Building microservices: designing fine-grained systems. " O'Reilly Media, Inc.", 2021.
21. Sneha, K., & Malle, G. M. (2017, August). Research on software testing techniques and software automation testing tools. In 2017 international conference on energy, communication, data analytics and soft computing (ICECDS) (pp. 77-81). IEEE.
22. Hourani, H., Hammad, A., & Lafi, M. (2019, April). The impact of artificial intelligence on software testing. In 2019 IEEE Jordan International Joint Conference on Electrical Engineering and Information Technology (JEEIT) (pp. 565-570).