TARU PUBLICATIONS
 Journal of Statistics and Management Systems cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0014·ISSN (Print): 0972-0510

Monthly Journal: Publishes peer-reviewed aticles on theoretical and applied statistics and management systems, expoloring industrial statistics, actuarial and decision sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Comparison of machine learning classification algorithms under different data distributions

* ,

* Corresponding author · click or hover a name for details

pp. 553–572Vol. 29Issue 5May 2026DOI: 10.47974/JSMS-1607XML
Received:
01 Sep 2025
Published Online:
20 Mar 2026
Article type:
Research Article
Language:
EN
Article no.:
JSMS-1607
Pages:
553–572

Abstract

In this study, we compare the performance of various machine learning algorithms through simulations conducted on datasets with different distributional characteristics. Widely used models such as Logistic Regression, Naive Bayes, k-Nearest Neighbors, Decision Trees, Support Vector Machines, Random Forest, AdaBoost, Gradient Boosting, and XGBoost are examined. The results indicate that the performance of classification algorithms is strongly influenced by the underlying data distribution and class balance. Specifically, Logistic Regression, Naive Bayes, and Support Vector Machines achieved high accuracy on normally distributed datasets, whereas tree-based methods such as Random Forest, Gradient Boosting, and XGBoost demonstrated greater robustness and superior outcomes on non-normal distributions and imbalanced class structures.

Keywords

Subject Classifications

62H3068T01

References

[1] S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms. Cambridge, U.K.: Cambridge University Press (2014).
[2] M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim, “Do we need hundreds of classifiers to solve real world classification problems?” J. Mach. Learn. Res., vol. 15, no. 1, pp. 3133–3181 (2014).
[3] J. Charoenpong, B. Pimpunchat, S. Amornsamankul, W. Triampo, and N. Nuttavut, “A comparison of machine learning algorithms and their applications,” Int. J. Simulation: Systems, Science & Technology, vol. 20, no. 4 (2019).
[4] N. Macià and E. Bernadó-Mansilla, “Towards UCI+: A mindful repository design,” Information Sciences, vol. 261, pp. 237–262 (2014).
[5] A. V. Aglarci and C. Bal, “Classification performance of machine learning methods in different data structures,” Communications in Statistics – Simulation and Computation, vol. 53, no. 12, pp. 6471–6489 (2024).
[6] Y. S. Li and C. Y. Guo, “Random logistic machine (RLM): Transforming statistical models into machine learning approach,” Communications in Statistics – Theory and Methods, vol. 53, no. 21, pp. 7517–7525 (2024).
[7] S. Ray, “A quick review of machine learning algorithms,” in Proc. Int. Conf. on Machine Learning, Big Data, Cloud and Parallel Computing (COMITCon), Bhubaneswar, India, pp. 35–39 (Feb. 2019).
[8] S. Zhang, M. Zong, K. Sun, Y. Liu, and D. Cheng, “Efficient kNN algorithm based on graph sparse reconstruction,” in Proc. 10th Int. Conf. on Advanced Data Mining and Applications (ADMA), Guilin, China, pp. 356–369 (Dec. 2014).
[9] A. Niculescu-Mizil and R. Caruana, “Predicting good probabilities with supervised learning,” in Proc. 22nd Int. Conf. on Machine Learning (ICML), Bonn, Germany, pp. 625–632 (Aug. 2005).
[10] M. Denil, D. Matheson, and N. de Freitas, “Narrowing the gap: Random forests in theory and in practice,” in Proc. Int. Conf. on Machine Learning (ICML), Atlanta, GA, USA, pp. 665–673 (Jan. 2014).
[11] Y. Freund and R. E. Schapire, “Experiments with a new boosting algorithm,” in Proc. 13th Int. Conf. on Machine Learning (ICML), pp. 148–156 (July 1996).
[12] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Annals of Statistics, vol. 29, no. 5, pp. 1189–1232 (2001).
[13] J. Yoon, “Forecasting of real GDP growth using machine learning models: Gradient boosting and random forest approach,” Computational Economics, vol. 57, no. 1, pp. 247–265 (2021).

Views: 72Downloads: 22Citations: 0