Open Access

A Robust Framework for Reliable and Secure Large Language Model-Assisted Software Development

4 Faculty of Computing and Software Engineering, Malaysian Institute of Digital Technology, Kuala Lumpur, Malaysia
4 Faculty of Information Technology, Malaysian Institute of Digital Technology, Kuala Lumpur, Malaysia

Abstract

Large Language Models (LLMs) are increasingly being incorporated into software-development workflows for code generation, test creation, debugging, refactoring, documentation, and maintenance. Although LLM-assisted development can improve developer productivity, generated code cannot automatically be regarded as reliable or secure. In particular, nondeterministic behavior in software tests, unreliable validation procedures, inadequate historical testing information, and insufficient root-cause analysis can weaken confidence in AI-generated software. This paper proposes a robust framework for reliable and secure LLM-assisted software development by integrating LLM-based code generation with systematic validation, flaky-test detection, historical test analysis, static prediction, and automated root-cause assessment. The framework synthesizes findings from existing research on flaky-test detection and prediction and extends their implications to AI-assisted development environments. The methodology defines a multi-stage pipeline covering requirement interpretation, constrained code generation, static analysis, test generation, flakiness assessment, reliability-oriented test prioritization, security validation, and human-controlled release decisions. The analysis indicates that reliability should be treated as a continuous property rather than a final-stage testing outcome. The proposed framework also highlights the importance of separating genuine software defects from nondeterministic test failures, since unreliable tests can distort the evaluation of generated code. The paper contributes a conceptual architecture for combining trustworthy LLM-assisted development with established software-quality mechanisms and identifies directions for empirical validation, adaptive risk scoring, and enterprise-scale deployment.

How to Cite

Amirul Rahman, & Nur Aisyah Hassan. (2026). A Robust Framework for Reliable and Secure Large Language Model-Assisted Software Development. Frontiers in Emerging Computer Science and Information Technology, 3(09), 07–13. https://doi.org/10.64917/fecsit/Volume03Issue09-02

References

REFERENCES
A. Ahmad, F. G. Oliveira Neto, Z. Shi, K. Sandahl, and O. Leifler, “A multi-factor approach for flaky test detection and automated root cause analysis,” in Proc. 28th Asia-Pacific Softw. Eng. Conf. (APSEC), Piscataway, NJ, USA : IEEE Press, 2021, pp. 338–348.
A. Alshammari, C. Morris, M. Hilton, and J. Bell, “FlakeFlagger: Predicting flakiness without rerunning tests,” in Proc. IEEE/ACM 43rd Int. Conf. Softw. Eng. (ICSE), Piscataway, NJ, USA : IEEE Press, 2021, pp. 1572–1584.
J. Bell, O. Legunsen, M. Hilton, L. Eloussi, T. Yung, and D. Marinov, “DeFlaker: Automatically detecting flaky tests,” in Proc. 40th Int. Conf. Softw. Eng., 2018, pp. 433–444.
M. Cordy, R. Rwemalika, M. Papadakis, and M. Harman, “FlakiMe: Laboratory-controlled test flakiness impact assessment. A case study on mutation testing and program repair,” 2019, arXiv:1912.03197.
E. Fallahzadeh and P. C. Rigby, “The impact of flaky tests on historical test prioritization on chrome,” in Proc. 44th Int. Conf. Softw. Eng.: Softw. Eng. Pract., 2022, pp. 273–282.
S. Fatima, T. A. Ghaleb, and L. Briand, “Flakify: A black-box, language model-based predictor for flaky tests,” IEEE Trans. Softw. Eng., vol. 49, no. 4, pp. 1912–1927, Apr. 2023.
M. Gruber, M. Heine, N. Oster, M. Philippsen, and G. Fraser, “Practical flaky test prediction using common code evolution and test history data,” in Proc. IEEE Conf. Softw. Testing, Verification Validation (ICST), Piscataway, NJ, USA : IEEE Press, 2023, pp. 210–221.
S. Habchi, G. Haben, M. Papadakis, M. Cordy, and Y. L. Traon, “A qualitative study on the sources, impacts, and mitigation strategies of flaky tests,” in Proc. IEEE Conf. Softw. Testing, Verification Validation (ICST), Piscataway, NJ, USA : IEEE Press, 2022, pp. 244–255.
G. Haben, S. Habchi, M. Papadakis, M. Cordy, and Y. L. Traon, “A replication study on the usability of code vocabulary in predicting flaky tests,” in Proc. IEEE/ACM 18th Int. Conf. Mining Softw. Repositories (MSR), Piscataway, NJ, USA : IEEE Press, 2021, pp. 219–229.
B. Harry, “How we approach testing VSTS to enable continuous delivery,” 2019. [Online]. Available: https://devblogs.microsoft.com/bharry/testing-in-a-cloud-delivery-cadence/
W. Lam, S. Winter, A. Wei, T. Xie, D. Marinov, and J. Bell, “A large-scale longitudinal study of flaky tests,” Proc. ACM Program. Lang., vol. 4, no. OOPSLA, pp. 1–29, 2020.
Q. Luo, F. Hariri, L. Eloussi, and D. Marinov, “An empirical analysis of flaky tests,” in Proc. 22nd ACM SIGSOFT Int. Symp. Found. Softw. Eng., 2014, pp. 643–653.
G. Pinto, B. Miranda, S. Dissanayake, M. d’Amorim, C. Treude, and A. Bertolino, “What is the vocabulary of flaky tests? ” in Proc. 17th Int. Conf. Mining Softw. Repositories, 2020, pp. 492–502.
V. Pontillo, F. Palomba, and F. Ferrucci, “Toward static test flakiness prediction: A feasibility study,” in Proc. 5th Int. Workshop Mach. Learn. Techn. Softw. Qual. Evol., 2021, pp. 19–24.
Y. Qin, “Peeler: Learning to effectively predict flakiness without running tests,” in Proc. IEEE Int. Conf. Softw. Maintenance Evol. (ICSME), Piscataway, NJ, USA : IEEE Press, 2022, pp. 257–268.
R. Verdecchia, E. Cruciani, B. Miranda, and A. Bertolino, “Know you neighbor: Fast static prediction of test flakiness,” IEEE Access, vol. 9, pp. 76119–76134, 2021.
Kongari, S. S. R., Kumar, S. K.., & Kumar, A. (2026). Trustworthy and Secure LLM-Assisted Code Generation for Enterprise Software Development. International Journal of Data Science and Machine Learning, 6(01), 222-237. https://doi.org/10.55640/ijdsml-06-01-04