Analisis Komparatif Kode GitHub Copilot: Tinjauan Metrik Clean Code dan Efisiensi Algoritma

Authors

  • Muhammad Arvind Alaric
  • Aji Gautama Putrada

Abstract

Pemanfaatan GitHub Copilot sebagai asisten pemrograman menjanjikan peningkatan produktivitas, namun memunculkan pertanyaan mengenai validitas fungsional dan kualitas struktural kode yang dihasilkan. Penelitian ini bertujuan mengevaluasi kualitas kode GitHub Copilot dibandingkan dengan programmer manusia (mahasiswa tingkat akhir) pada masalah pemrograman berorientasi objek kompleksitas menengah. Metode penelitian menggunakan eksperimen terkontrol dengan pendekatan mixed-method: evaluasi fungsional menggunakan automated test suite (Jest), analisis kualitas statis (Clean Code) menggunakan SonarQube, dan analisis efisiensi algoritma (Big O) secara kualitatif. Hasil penelitian menunjukkan dikotomi kompetensi yang signifikan; AI unggul dalam validitas struktural dengan skor Cognitive Complexity jauh lebih rendah (p < 0.05 pada uji Wilcoxon) dan konsistensi tinggi. Sebaliknya, programmer manusia unggul dalam validitas logika dan penanganan kasus tepi (edge cases). Ditemukan fenomena "Clean Garbage" pada AI, di mana kode terlihat rapi secara struktur namun gagal secara fungsional. Kesimpulannya, peran manusia tetap krusial sebagai arsitek solusi dan validator logika dalam ekosistem pengembangan perangkat lunak berbasis AI.

Kata kunci—GitHub Copilot, Kualitas Kode, Clean Code, Kompleksitas Algoritmik, Eksperimen Terkontrol.

References

A. Moradi Dakhel, V. Majdinasab, A. Nikanjam, F. Khomh, M. C. Desmarais, and Z. M. Jiang. “GitHub Copilot AI pair programmer: Asset or Liability?”. J. Syst. Softw., vol. 203, 2023.

B. Yetistiren, I. Ozsoy, and E. Tuzun. “Assessing the quality of GitHub copilot’s code generation”. Proc. 18th Int. Conf. Predict. Model. Data Anal. Softw. Eng. (PROMISE), pp. 62–71, 2022.

N. Nguyen and S. Nadi. “An Empirical Evaluation of GitHub Copilot’s Code Suggestions”. Proc. Mining Softw. Repos. (MSR), vol. 1, no. 1, 2022.

S. Peng, E. Kalliamvakou, P. Cihon, and M. Demirer. (2023). “The Impact of AI on Developer Productivity: Evidence from GitHub Copilot”. arXiv preprint arXiv:2302.06590. [Online]. Available: http://arxiv.org/abs/2302.06590

M. M. Barón, M. Wyrich, and S. Wagner. “An empirical validation of cognitive complexity as a measure of source code understandability”. Int. Symp. Empir. Softw. Eng. Meas., 2020.

M. A. Florez Muñoz, J. C. Jaramillo De La Torre, S. Pareja López, S. Herrera, and C. A. Candela Uribe. “Comparative Study of AI Code Generation Tools: Quality Assessment and Performance Analysis”. LatIA, vol. 2, p. 104, 2024.

D. Alawad, M. Panta, M. Zibran, and M. R. Islam. “An empirical study of the relationships between code readability and software complexity”. 27th Int. Conf. Softw. Eng. Data Eng. (SEDE), pp. 122–127, 2018.

W. Takerngsaksiri, M. Fu, C. Tantithamthavorn, J. Pasuksmit, K. Chen, and M. Wu. (2025, June). “Code Readability in the Age of Large Language Models: An Industrial Case Study from Atlassian”. Companion Proc. 33rd ACM Symp. Found. Softw. Eng. (FSE ’25). [Online]. Vol. 1. Available: https://arxiv.org/abs/2501.11264v1

[9] D. Tosi. “Studying the Quality of Source Code Generated by Different AI Generative Engines: An Empirical Evaluation”. Future Internet, vol. 16, no. 6, 2024.

D. Smit, H. Smuts, P. Louw, J. Pielmeier, and C. Eidelloth. “The impact of GitHub Copilot on developer productivity from a software engineering body of knowledge perspective”. Proc. 30th Americas Conf. Inf. Syst. (AMCIS), 2024.

R. P. Buse and W. Weimer. “Learning a Metric for Code Readability”. IEEE Trans. Softw. Eng., vol. 36, no. 4, pp. 546–558, 2010.

T. Heričko and B. Šumak. “Exploring Maintainability Index Variants for Software Maintainability Measurement in Object-Oriented Systems”. Appl. Sci., vol. 13, no. 5, p. 2972, 2023.

D. G. Paul, H. Zhu, and I. Bayley. (2024). “Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review”. arXiv preprint arXiv:2406.12655. [Online].

B. Yetiştiren, I. Özsoy, M. Ayerdem, and E. Tüzün. (2023). “Evaluating the Code Quality of AI-Assisted Code Generation Tools: An Empirical Study on GitHub Copilot, Amazon CodeWhisperer, and ChatGPT”. arXiv preprint arXiv:2304.10778. [Online].

J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou. “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”. Advances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022.

P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig. “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing”. ACM Comput. Surv., vol. 55, no. 9, art. 195, 2023.

Published

2026-07-01

Issue

Section

Prodi S1 Teknologi Informasi