Evaluating the Quality, Maintainability, and Security of AI-Generated Source Code in Enterprise Systems

Digital Transformation and Technology Dynamics

Rashadul Islam Samrat, , Kanita Haider

Department of Marketing, University of Barishal, Barishal, Bangladesh; Department of Computer and Information Sciences, University of the Cumberlands, Kentucky, USA; Department of Computer Science and Engineering, International Islamic University Chittagong, Chittagong, Bangladesh

Digital Transformation and Technology DynamicsVol. 2, Issue 2November 29, 2022

View Full PDF

Abstract

Large language model (LLM)–based code-generation tools such as GitHub Copilot, ChatGPT, and Amazon CodeWhisperer are increasingly embedded in enterprise software development lifecycles, with studies reporting notable gains in developer task-completion speed. Despite rapid adoption, systematic evaluation of the code these tools produce has lagged. Existing empirical work often isolates functional correctness, security, or maintainability, relying on inconsistent corpora, static-analysis toolchains, and reporting conventions, which complicates enterprise leaders’ ability to assess aggregate quality risks. This paper addresses that gap by synthesizing the empirical literature on AI-generated code quality, maintainability, and security, and by proposing an integrated evaluation framework, the Enterprise Code Quality Index (ECQI). A thematic review of peer-reviewed and archival studies published between 2021 and 2026 was conducted, encompassing correctness benchmarks, static-analysis security studies, controlled maintainability experiments, and productivity trials. Building on this evidence, a topic-specific research design is outlined, combining repository-grounded prompt sampling, multi-tool code generation, static security analysis with CWE mapping, maintainability metrics, functional testing, and structured expert review. The reviewed evidence highlights a consistent rate of security weaknesses in AI-generated code, ranging from 12 to 40 percent across independent studies, mixed but cautious signals on maintainability and technical debt, and robust productivity gains when human oversight is retained. Because this study did not execute new empirical experiments, illustrative results are presented separately from literature-derived findings. The contribution lies in offering a thematic synthesis, a reproducible evaluation methodology, a composite quality index for governance, and actionable recommendations for secure, maintainable enterprise adoption of AI coding assistants.

Keywords

AI-generated code; Software security; Code maintainability; Enterprise software engineering; Static analysis; Technical debt; Large language models; Software quality assurance

Article Information

Published
November 29, 2022
Journal
Digital Transformation and Technology Dynamics
Volume / Issue
2 / 2
Article No.
DTTD-2022003
Year
2022

Browse

All published articles · Journal archive