Multiphase code clone detection using NSGA-II
Neha Sainiprofnehasaini@gmail.comDepartment of Computer ScienceGovernment College, Chhachhrauli (Yamuna Nagar)Chhachhrauli, Yamuna Nagar, Haryana, 135103, IndiaView full profile → , *ShaluCorresponding authorsingshalu2609@gmail.comDepartment of Computer Science & TechnologyManav Rachna UniversityFaridabad, Haryana, 121004, India0000-0002-5516-0357View full profile → , Upasna Joshiupasnajoshi19@gmail.comSchool of Computer Science Engineering & TechnologyBennett UniversityGreater Noida, Uttar Pradesh, 201310, IndiaView full profile → , Ankit Gambhirer.ankit.gambhir@gmail.comDepartment of Computer Science & EngineeringTrinity Institute of Professional StudiesGuru Gobind Singh Indraprastha UniversityGreater Noida, Uttar Pradesh, 201310, IndiaView full profile → , Aniket Singhaniketsingh@mru.edu.inDepartment of Computer Science & TechnologyManav Rachna UniversityFaridabad, Haryana, 121004, IndiaView full profile → , Mohd Anas Khananas.cse786@gmail.comDepartment of Computer EngineeringJamia Millia IslamiaJamia Nagar, New Delhi, 110025, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 01 Mar 2026
- Published Online:
- 31 Jul 2026
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-2326
- Pages:
- 2713–2722
Abstract
This duplication of source code is termed as code cloning. This is the most widely used source code reusing technique in software development. When one bug is discovered in a single section of a code, there is need to test all the other sections that are the same code to test the bug. Consequently, the given cloning technique can lead to the spread of some issues, which significantly affects maintenance expenses. Code clone detection (CCD) seems to be an active field of research. CCD method based on semantic similarity is applicable in software engineering (e.g., software evolution, software reuse, etc). The classical code clone detection systems emphasize semantic similarity to a lesser degree and more on the syntax similarity. This leads to ignoring of candidate code that has similar meanings in semantics. To address this issue semantic similarity-based code clone detection algorithm named as Multiphase Code Clone Detection using non-dominated sorting algorithm (NSGA- II) has been put forward which has given a type-1 clone detection of 97.37% and a type-2, type-3, and type-4 clones detection of 96.70%, 95% and 92.50% respectively.
Keywords
Subject Classifications
References
[1] A. Sheneamer, S. Roy, and J. Kalita, “A detection framework for semantic code clones and obfuscated code,” Expert Systems with Applications, vol. 97, pp. 129–143 (2018), doi: 10.1016/j.eswa.2017.12.040.
[2] D. Rattan, R. Bhatia, and M. Singh, “Software clone detection: A systematic review,” Information and Software Technology, vol. 55, no. 7, pp. 1165–1199 (2013), doi: 10.1016/j.infsof.2013.01.008.
[3] S. K. Panda, S. Mohapatra, and S. Das, “Artificial neural network-based virtual machine allocation in cloud computing,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 24, no. 6, pp. 1709–1722 (2021), doi: 10.1080/09720529.2021.1878626.
[4] H. Sajnani, V. Saini, J. Svajlenko, C. K. Roy, and C. V. Lopes, “SourcererCC and SourcererCC-I: Tools to detect clones in batch mode and during software development,” in Proc. 38th Int. Conf. Software Engineering Companion (ICSE-C), pp. 441–444 (2016), doi: 10.1145/2889160.2889165.
[5] P. Vijai and P. Bagavathi Sivakumar, “A hybrid multi-objective optimization approach with NSGA-II for feature selection,” Decision Analytics Journal, vol. 14, art. no. 100550 (2025), doi: 10.1016/j.dajour.2025.100550.
[6] E. Elabd, W. M. Ead, and Shalu, “An optimized differential private stochastic gradient descent (DP-SGD) approach to combat membership inference attacks in neural networks,” Int. J. Inf. Technol., vol. 17, no. 4, pp. 2369–2374 (2025), doi: 10.1007/s41870-025-02442-y.
[7] M. Karnalim and Simon, “Current trends in source code analysis, plagiarism detection and issues of analysis big datasets,” Procedia Computer Science, vol. 124, pp. 329–338 (2017), doi: 10.1016/j.procs.2017.12.162.
[8] S. Singh and D. Singh, “A bio-inspired VM migration using re-initialization and decomposition based-whale optimization,” ICT Express, vol. 9, no. 1, pp. 92–99 (2023), doi: 10.1016/j.icte.2022.02.003.
[9] S. K. Saini, H. Sajnani, and C. V. Lopes, “Semantic code clone detection for Internet of Things applications using reaching definition and liveness analysis,” The Journal of Supercomputing, vol. 73, no. 11, pp. 4997–5026 (2017), doi: 10.1007/s11227-016-1832-6.
[10] B. Wan, S. Dong, J. Zhou, and Y. Qian, “SJBCD: A Java code clone detection method based on bytecode using Siamese neural network,” Applied Sciences, vol. 13, no. 17, art. no. 9580 (2023), doi: 10.3390/app13179580.
[11] G. A. Kumar, D. C. R. K. Reddy, and D. A. Govardhan, “Software code clone detection using AST,” International Journal of P2P Network Trends and Technology, vol. 9, no. 4, pp. 15–19 (2014), doi: 10.14445/22492615/IJPTT-V9P407.
[12] A. Yasmin, R. Venkatesan, and K. Gaverchand, “Enhancing the randomness of a symmetric cryptographic technique based on linear 1-D cellular automata,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 28, no. 1, pp. 185–204 (2025), doi: 10.47974/JDMSC-2133.
[13] T. Kamiya, S. Kusumoto, and K. Inoue, “CCFinder: A multilinguistic token-based code clone detection system for large scale source code,” IEEE Transactions on Software Engineering, vol. 28, no. 7, pp. 654–670 (2002), doi: 10.1109/TSE.2002.1019480.
[14] H. Liu and Z. Ma, “Detecting duplications in sequence diagrams based on suffix trees,” in Proc. Asia-Pacific Software Engineering Conf., pp. 269–276 (2008).
[15] M. Gabel, L. Jiang, and Z. Su, “Challenging cloning related problems with GPU-based algorithms,” in Proc. 4th Int. Workshop on Software Clones, pp. 1–7 (2010), doi: 10.1145/1808901.1808905.
[16] W. Rahman, F. Pu, and X. Jia, “Clone detection on large Scala codebases,” in 2020 IEEE 14th International Workshop on Software Clones (IWSC), pp. 38–44 (2020), doi: 10.1109/IWSC50091.2020.9047640.
[17] M. Ďuračík, E. Kršák, and P. Hrkút, “Current trends in source code analysis, plagiarism detection and issues of analysis big datasets,” Procedia Engineering, vol. 192, pp. 136–141 (2017), doi: 10.1016/j.proeng.2017.06.024.
[18] R. Tekchandani, R. Bhatia, and M. Singh, “Semantic code clone detection for Internet of Things applications using reaching definition and liveness analysis,” The Journal of Supercomputing, vol. 74, no. 9, pp. 4199–4226 (Sep. 2018), doi: 10.1007/s11227-016-1832-6.




