TARU PUBLICATIONS
 Journal of Statistics and Management Systems cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0014·ISSN (Print): 0972-0510

Monthly Journal: Publishes peer-reviewed aticles on theoretical and applied statistics and management systems, expoloring industrial statistics, actuarial and decision sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Developing a mixture decision tree (MDT) algorithm with k-means clustering to improve classification accuracy

* , ,

* Corresponding author · click or hover a name for details

pp. 733–749Vol. 27Issue 4May 2024DOI: 10.47974/JSMS-863XML
Received:
03 Dec 2020
Accepted:
07 Sep 2021
Published Online:
27 May 2024
Article type:
Research Article
Language:
EN
Article no.:
JSMS-863
Pages:
733–749

Abstract

Statistical research based on data is one of the primary keys to the development of the world. But the size of the data is getting increased very rapidly. For this reason, statistical data analysis is getting complicated, and sometimes the existing data mining techniques are not providing significant and satisfactory results. On the other hand, traditional classification techniques provide results that are not statistically significant every time. In this research, a Mixture Decision Tree algorithm is being proposed that is a more powerful classification technique, and that may provide better results than the existing methods. In this Mixture Decision Tree, the Classification and Regression Tree (CART) algorithm is combined with k-means clustering, by which the primary data is filtered in two stages. Applying this method to some predefined data, significant enhancement is found in the classification precision for the current classification techniques. The Mixture Decision Tree algorithm is slightly more complex than the existing methods, but computer programming techniques can be done quickly.

Keywords

Subject Classifications

62J0503C45

References

[1] Aitkenhead, M. J., A co-evolving decision tree classification method. Expert System with Application, 34(1), 18-25 (2008). doi:https://doi.org/10.1016/j.eswa.2006.08.008
[2] Biswas, A., Farid, D. M., & Rahman, C. M., A New Decision Tree Learning Approach forNovel Class Detection in Concept Drifting Data Stream Classification. Journal of Computer Science and Engineering, 14(1), 1-8 (2012).
[3] Chandra, B., & Gupta, M., Robust approach for estimating probabilities in Naïve–Bayes Classifier for gene expression data. Expert Systems with Applications, 38(3), 1293-1298 (2011). doi:https://doi.org/10.1016/j.eswa.2010.06.076.
[4] Chandra, B., & Varghese, P. P., Fuzzifying Gini Index based decision trees. Expert Systems with Applications, 36(4), 8549–8559 (2009). doi:https://doi.org/10.1016/j.eswa.2008.10.053.
[5] Chena, J., Huang, H., Tian, S., & Qu, Y., Feature selection for text classification with Naïve Bayes. Expert Systems with Applications, 36(3), 5432-5435 (2009). doi:https://doi.org/10.1016/j.eswa.2008.06.054.
[6] Farid, A. M., Zhang, L., Rahman, C. M., Hossain, M. A., & Strachan, R., Hybrid decision tree and naïve Bayes classifiers for multi-class classification tasks. Expert Systems with Applications, 41(4), 1937-1946 (2014). doi:https://doi.org/10.1016/j.eswa.2013.08.089.
[7] Farid, D. M., & Rahman, M. Z., Anomaly Network Intrusion Detection Based on Improved Self Adaptive Bayesian Algorithm. Journal of Computers, 5(1), 23-31 (2010). doi:DOI:10.4304/jcp.5.1.23-31.
[8] Farid, D. M., Maruf, G. M., & Rahman, C. M., A new approach of Boosting using decision tree classifier for classifying noisy data. International Conference on Informatics, Electronics & Vision (ICIEV) (2013). doi:DOI: 10.1109/ICIEV.2013.6572718.
[9] Frank, A., & Asuncion, A., UCI ((University College London) Machine Learning Repository (2010). Retrieved from https://archive.ics.uci.edu/ml/datasets/Iris.
[10] Hechmi, J. M., Khlaifi, H., Amine, B., & Zrelli, A., Intrusion Detection Using Data Fusion and Machine Learning. 26th International Conference on Software, Telecommunications and Computer Networks (SoftCOM) (2018). doi:10.23919/SOFTCOM.2018.8555800.
[11] Hsu, C.-C., Huang, Y.-P., & Chang, K.-W., Extended Naive Bayes classifier for mixed data. Expert Systems with Applications, 35(3), 1080–1083 (2008). doi:https://doi.org/10.1016/j.eswa.2007.08.031.
[12] Kaklamanis, M. M., & Filippakis, M. E., A comparative survey of machine learning classification algorithms for breast cancer detection. PCI ‘19: Proceedings of the 23rd Pan-Hellenic Conference on Informatics, pp. 93-103 (2019). doi:https://doi.org/10.1145/3368640.3368642.
[13] Kemal, P., & Salih, G., A novel hybrid intelligent method based on C4.5 decision tree classifier and one-against-all approach for multi-class classification problems. Expert Systems with Applications, 36(2), 1587-1592 (2009). doi:https://doi.org/10.1016/j.eswa.2007.11.051.
[14] Koc, L., Mazzuchi, T. A., & Sarkani, S., A network intrusion detection system based on a Hidden Naïve Bayes multiclass classifier. Expert Systems with Applications, 39(18), 13492-13500 (2012). doi:https://doi.org/10.1016/j.eswa.2012.07.009.
[15] Kwok, J. T.-Y., Moderating the Outputs of Support Vector Machine Classifiers. IEEE Transactions On Neural Networks, 10(5), 1018-1031 (1999).
[16] Liberman, N., Decision Trees and Random Forests, Towards Data Science (2017). 
[17] Mingo, L., Aslanyan, L., Castellanos, J., Díaz, M., & Riazanov, V., FOURIER NEURAL NETWORKS: AN APPROACH WITH SINUSOIDAL ACTIVATION. International Journal “Information Theories & Applications, 11, 52-55 (2006).
[18] Mugambi, E. M., Hunter, A., Oatleyd, G., & Kennedy, L., Polynomial-fuzzy decision tree structures for classifying medical data. Knowledge-Based Systems, 17(2-4), 81-87 (2004). doi:https://doi.org/10.1016/j.knosys.2004.03.003.
Nagarajan, P. R., Sandhiya, K., Vinod, V. M., & Mekala, V., Offline Navigation:GPS based assisting system in Sathuragiri forests 
using Machine Learning. International Conference on Intelligent Computing and Communication for Smart World (I2C2SW) (2018). doi:10.1109/I2C2SW45816.2018.8997523.
[19] Polat, K., & Güneş, S., The effect to diagnostic accuracy of decision tree classifier of fuzzy and k-NN based weighted pre-processing methods to diagnosis of erythemato-squamous diseases. Digital Signal Processing, 16(6), 922-930 (2006). doi:https://doi.org/10.1016/j.dsp.2006.04.007.
[20] Ram, N. P., Sandhiya, K., Vinod, V. M., & Mekala, V., Offline Navigation:GPS based assisting system in Sathuragiri forests using Machine Learning. International Conference on Intelligent Computing and Communication for Smart World (I2C2SW), pp. 326-331 (2018). doi:10.1109/I2C2SW45816.2018.8997523.
[21] Sabbirantor. K means clustering | K Means ++ (Slideshow) (2019). 
[22] Siswantoro, J., Prabuwono, A. S., Abdullah, A., & Idrus, B. A linear model based on Kalman filter for improving neural network classification performance. Expert Systems with Applications, 112-122 (2016). doi:https://doi.org/10.1016/j.eswa.2015.12.012.
[23] Tolson, E., Machine Learning in the Area of Image. Advanced Undergraduate Project Spring 2001, 1-35 (2001).
[24] Tuovinen, T. S., Real-time classification of SMEs credit and risk ratings and the impact of financial indicators and payment behaviour. Bachelor’s Thesis, University of Applied Sciences (2020).
[25] Jyothi Varanasi & M. M. Tripathi. K-means clustering based photo voltaic power forecasting using artificial neural network, particle swarm optimization and support vector regression, Journal of Information and Optimization Sciences, 40:2, 309-328 (2019). DOI: 10.1080/02522667.2019.1578091.
[26] Satvik Vats & B. B. Sagar, Performance evaluation of K-means clustering on Hadoop infrastructure, Journal of Discrete Mathematical Sciences and Cryptography, 22:8, 1349-1363 (2019), DOI: 10.1080/09720529.2019.1692444.

Views: 286Downloads: 86Citations: 0