TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

A conditionally positive definite kernel function for clustering of incomplete data

* ,

* Corresponding author · click or hover a name for details

pp. 403–412Vol. 45Issue 2March 2024DOI: 10.47974/JIOS-1557XML
Published Online:
17 Mar 2024
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1557
Pages:
403–412

Abstract

Clustering of incomplete data sets that contains missing features is one of the most widely studied problems in the literature, and several imputation and non-imputation techniques are used to solve this problem. A weighted sum of the Euclidean distance from the datum to the corresponding clusters is used in  Fuzzy c-means clustering. It has been observed that the kernel-based clustering techniques outperform the conventional algorithms in terms of accuracy. This is due to their ability to handle non-linear data and map it to higher dimensional space while preserving its internal structure. Kernel functions are really important when it comes to the performance of kernel-based clustering methods. Choosing the right kernel function isn’t simple. Among the various clustering algorithms that have been examined in the literature, the, Gaussian kernel function has been found to be more useful.. This paper suggests a conditionally positive definite kernel function that can be used in the unsupervised clustering of incomplete data. Numerical analysis shows that the conditionally positive definite kernel function also performs well on datasets with incomplete features.

Keywords

Subject Classifications

58-XX

References

[1] Jerez, J. M., Molina, I., García-Laencina, P. J., Alba, E., Ribelles, N., Martín, M., & Franco, L. Missing data imputation using statistical and machine learning methods in a real breast cancer problem. Artificial Intelligence in Medicine, 50(2), 105-115 (2010).
[2] Hathaway, R. J., & Bezdek, J. C., Fuzzy c-means clustering of incomplete data. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 31(5), 735-744 (2001).
[3] García-Laencina, P. J., Sancho-Gómez, J. L., Figueiras-Vidal, A. R., & Verleysen, M., K nearest neighbours with mutual information for simultaneous classification and missing data imputation. Neurocomputing, 72(7-9), 1483-1493 (2009).
[4] Little, R.J. and Rubin, D.B., Statistical Analysis with Missing Data. 793, John Wiley & Sons, Hoboken (2019).
[5] Enders,C. K.,& Baraldi, A. N., Missing data handling methods. The Wiley handbook of psychometric testing: A multidisciplinary reference on survey, scale and test development, 139-185 (2018).
[6] Farhangfar,A.,Kurgan, L.A.,& Pedrycz, W., A novel framework for imputation of missing values in databases. IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans, 37(5), 692-709 (2007).
[7] Weed, L., Lok, R., Chawra, D.,& Zeitzer, J., The impact of missing data and imputation methods on the analysis of 24-hour activity patterns. Clocks & Sleep, 4(4), 497-507 (2022).
[8] Zhang, S., Nearest neighbor selection for iteratively kNN imputation. Journal of Systems and Software, 85(11), 2541-2552 (2012).
[9] Noor, N. M., Al Bakri Abdullah, M. M., Yahaya, A. S., & Ramli, N. A., Comparison of linear interpolation method and mean method to replace the missing values in environmental data set. In Materials Science Forum, 803, 278-281 (2015, January) Trans Tech Publications Ltd.
[10] Goel, S., & Tushir, M., Different approaches for missing data handling in fuzzy clustering: a review. Recent Advances in Electrical & Electronic Engineering, 13(6), 833-846 (2020).
[11] Hwang, S., Oh, J., Cox, J., Tang, S. J., & Tibbals, H. F., Blood detection in wireless capsule endoscopy using expectation maximization clustering. In Medical Imaging 2006: Image Processing, 6144, p. 61441P. International Society for Optics and Photonics (2016).
[12] K. R. Muller, S. Mika, G. Ratsch and K. Tsuda, An introduction to kernel-based learning algorithms, IEEE transactions on neural networks, 12(2), 181-201 (2001).
[13] M. Girolami, “Mercer kernel-based clustering in feature space,” IEEE transactions on Neural Networks,13(3), 780-784 (2002).
[14] M. Tushir and S. Srivastava, A new kernel based hybrid c-means clustering model, In 2007 IEEE International Fuzzy Systems Conference, 1-5 (2007).
[15] Nigam, J., Tushir, M., & Rai, D., A conditionally positive definite kernel functions for possibilistic clustering. International Journal of Artificial Intelligence and Soft Computing, 6(1), 75-91 (2017).
[16] K. Bache and M. Lichman, UCI Machine Learning Repository [http://archive. ics. uci.edu/ml]. Irvine, CA: University of California, “School of information and computer science (2013). 

Views: 181Downloads: 86Citations: 1