TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate

The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to: • Information Sciences • Optimization Sciences • Control Theory • Operational Research • Decision Sciences • Information Theory • Information Technology • Computer Networks and Communications • Mathematical Programming • Modelling and Simulation • Database Management • Applications to Engineering Sciences • Applications to Technology

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

A comprehensive study based on MFCC and spectrogram for audio classification

, , * ,

* Corresponding author · click or hover a name for details

pp. 1057–1074Vol. 44Issue 6September 2023DOI: 10.47974/JIOS-1431XML
Published Online:
04 Sep 2023
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1431
Pages:
1057–1074

Abstract

Music Assortment is a music information retrieval (MIR) function to decide music connotation computationally. In recent years, deep neural networks have been proven to be effective in numerous classification tasks, including music genre categorisation. In this paper, we employ a comparative study between the two different music classification techniques. The first technique uses the audio’s spectrogram image and computes the music’s genre based on its spectrogram, using the CNN model trained on the spectrograms. The second approach computes the MFCC’s (Mel-Frequency Cepstral Coefficients) musical features and utilises them to classify the music using ANN. This paper aims to study the two algorithms closely against different audio signals and check the performance report of the above-mentioned techniques to see which of them is better for music genre classification.

Keywords

Subject Classifications

68T07

References

[1] R. Agarwal, S. Singh, and S. Vats, “Implementation of an improved algorithm for frequent itemset mining using Hadoop,” Proceeding - IEEE Int. Conf. Comput. Commun. Autom. ICCCA 2016, pp. 13–18 (Jan. 2017), doi: 10.1109/CCAA.2016.7813719.[2] S. Vats, S. K. Dubey, and N. K. Pandey, “Criminal Face Identification System Criminal Face Identification.”[3] S. Vats and B. B. Sagar, “Data lake: A plausible big data science for business intelligence,” Commun. Comput. Syst. Proc. 2nd Int. Conf. Commun. Comput. Syst. ICCCS 2018, pp. 442–448, Oct. 2019, doi: 10.1201/9780429444272-70/DATA-LAKE-PLAUSIBLE-BIG-DATA-SCIENCE-BUSINESS-INTELLIGENCE-SATVIK-VATS-SAGAR.[4] M. Bhatia, V. Sharma, P. Singh, and M. Masud, “Multi-Level P2P Traffic Classification Using Heuristic and Statistical-Based Techniques: A Hybrid Approach,” Symmetry 2020, Vol. 12, Page 2117, vol. 12, no. 12, p. 2117 (Dec. 2020), doi: 10.3390/SYM12122117.[5] S. Vikrant, P. R. B, and B. H. S, “Policy for planned placement of sensor nodes in large scale wireless sensor network,” KSII Trans. INTERNET Inf. Syst., vol. 10, no. 7 (2016), doi: 10.3837/tiis.2016.07.019.[6] D. Prasad, A. Kumar, V. Sharma, and D. Prasad, “SDRS: Split Detection and Reconfiguration Scheme for Quick Healing of Wireless Sensor Network,” Int. J. Adv. Comput., vol. 46, no. 3, pp. 2051–0845 (Accessed: Feb. 03, 2023). [Online]. Available: https://www.researchgate.net/publication/263661007.[7] D. Arora, A. Singh, V. Sharma, H. S. Bhaduria, and R. B. Patel, “HgsDb: Haplogroups Database to understand migration and molecular risk assessment,” Bioinformation, vol. 11, no. 6, p. 272 (Jun. 2015), doi: 10.6026/97320630011272.[8] V. Sharma, R. B. Patel, H. S. Bhaduria, and D. Prasad, “Policy for random aerial deployment in large scale Wireless Sensor Networks,” Int. Conf. Comput. Commun. Autom. ICCCA 2015, pp. 367–373 (Jul. 2015), doi: 10.1109/CCAA.2015.7148445.[9] S. Vats, B. B. Sagar, K. Singh, A. Ahmadian, and B. A. Pansera, “Performance Evaluation of an Independent Time Optimized Infrastructure for Big Data Analytics that Maintains Symmetry,” Symmetry 2020, Vol. 12, Page 1274, vol. 12, no. 8, p. 1274 (Aug. 2020), doi: 10.3390/ SYM12081274.[10] S. Vats, S. Singh, G. Kala, R. Tarar, and S. D. FRCR, “iDoc-X: An artificial intelligence model for tuberculosis diagnosis and localization,” https://doi.org/10.1080/09720529.2021.1932910, vol. 24, no. 5, pp. 1257–1272 (Jul. 2021), doi: 10.1080/09720529.2021.1932910.[11] V. Sharma, R. B. Patel, H. S. Bhadauria, and D. Prasad, “NADS: Neighbor Assisted Deployment Scheme for Optimal Placement of Sensor Nodes to Achieve Blanket Coverage in Wireless Sensor Network,” Wirel. Pers. Commun., vol. 90, no. 4, pp. 1903–1933 (Oct. 2016), doi: 10.1007/S11277-016-3430-6/METRICS. [12] D. Prasad, A. Kumar, V. Sharma, and D. Prasad, “Distributed Deployment Scheme for Homogeneous Distribution of Randomly Deployed Mobile Sensor Nodes in Wireless Sensor Network,” IJACSA) Int. J. Adv. Comput. Sci. Appl., vol. 4, no. 4 (2013), Accessed: Feb. 03, 2023. [Online]. Available: www.ijacsa.thesai.org.[13] V. Sharma, R. B. Patel, H. S. Bhadauria, and D. Prasad, “Deployment schemes in wireless sensor network to achieve blanket coverage in large-scale open area: A review,” Egypt. Informatics J., vol. 17, no. 1, pp. 45–56 (Mar. 2016), doi: 10.1016/J.EIJ.2015.08.003.[14] S. Vats and B. B. Sagar, “An independent time optimized hybrid infrastructure for big data analytics,” https://doi.org/10.1142/S021798492050311X, vol. 34, no. 28 (Jul. 2020), doi: 10.1142/S021798492050311X.[15] S. Vats and B. B. Sagar, “Performance evaluation of K-means clustering on Hadoop infrastructure,” https://doi.org/10.1080/09720529.2019.1692444, vol. 22, no. 8, pp. 1349–1363 (Nov. 2020), doi: 10.1080/09720529.2019.1692444.[16] R. Agarwal, S. Singh, and S. Vats, “Review of parallel apriori algorithm on mapreduce framework for performance enhancement,” Adv. Intell. Syst. Comput., vol. 654, pp. 403–411 (2018), doi: 10.1007/978-981-10-6620-7_38/COVER.[17] S. Vikrant, R. B. Patel, H. S. Bhadauria, and D. Prasad, “Glider assisted schemes to deploy sensor nodes in Wireless Sensor Networks,” Rob. Auton. Syst., vol. 100, pp. 1–13 (Feb. 2018), doi: 10.1016/J.ROBOT.2017.10.015.[18] V. Sharma et al., “OGAS: Omni-directional Glider Assisted Scheme for autonomous deployment of sensor nodes in open area wireless sensor network,” ISA Trans., (Aug. 2022), doi: 10.1016/J.ISATRA.2022.08.001.[19] J. P. Bhati, D. Tomar, and S. Vats, “Examining Big Data Management Techniques for Cloud-Based IoT Systems,” pp. 164–191 (2018), doi: 10.4018/978-1-5225-3445-7.CH009:[20] V. Sharma, R. B. Patel, H. S. Bhadauria, and D. Prasad, “Pneumatic Launcher Based Precise Placement Model for Large-Scale Deployment in Wireless Sensor Networks,” IJACSA) Int. J. Adv. Comput. Sci. Appl., vol. 6, no. 12 (2015), Accessed: Feb. 03, 2023. [Online]. Available: www.ijacsa.thesai.org.[21] Lee, H., Pham, P., Largman, Y., & Ng, A. : Unsupervised feature learning for audio classification using convolutional deep belief networks. Advances in neural information processing systems, 22 (2009).[22] Hamel, P., & Eck, D. : Learning features from music audio with deep belief networks. In ISMIR, Vol. 10, pp. 339-344 (2010, August).[23] Haggblade, M., Hong, Y., & Kao, K. : Music genre classification. Department of Computer Science, Stanford University (2011).[24] Costa, Y. M., Oliveira, L. S., Koerich, A. L., Gouyon, F., & Martins, J. G. : Music genre classification using LBP textural features. Signal Processing, 92(11), 2723-2737 (2012).[25] Rajanna, A. R., Aryafar, K., Shokoufandeh, A., & Ptucha, R. : Deep neural networks: A case study for music genre classification. In 2015 IEEE 14th international conference on machine learning and applications (ICMLA), pp. 655-660 (2015, December). IEEE. [26] Poria, S., Gelbukh, A., Hussain, A., Bandyopadhyay, S., & Howard, N. : Music genre classification: A semi-supervised approach. In Pattern Recognition: 5th Mexican Conference, MCPR 2013, Querétaro, Mexico, June 26-29, 2013. Proceedings 5, pp. 254-263 (2013). Springer Berlin Heidelberg.[27] Baniya, B. K., Ghimire, D., & Lee, J. : Automatic music genre classification using timbral texture and rhythmic content features. In 2015 17th International Conference on Advanced Communication Technology (ICACT), pp. 434-443 (2015, July). IEEE.[28] Nanni, L., Costa, Y. M., Lumini, A., Kim, M. Y., & Baek, S. R. : Combining visual and acoustic features for music genre classification. Expert Systems with Applications, 45, 108-117 (2016).[29] Kumar, A., Singh, K., & Khan, T. : L-RTAM: Logarithm based reliable trust assessment model for WBSNs. Journal of Discrete Mathematical Sciences and Cryptography, 24(6), 1701-1716 (2021).[30] Khan, T., & Singh, K. : Resource management based secure trust model for WSN. Journal of Discrete Mathematical Sciences and Cryptography, 22(8), 1453-1462 (2019).[31] S. Vishnupriya and K. Meenakshi, “Automatic Music Genre Classification using Convolutional Neural Network,” 2018 International Conference on Computer Communication and Informatics (ICCCI), Coimbatore, India, pp. 1-4 (2018), doi: 10.1109/ICCCI.2018.8441340.[32] A. Elbir, H. Bilal Çam, M. Emre Iyican, B. Öztürk and N. Aydin, “Music Genre Classification and Recommendation by Using Machine Learning Techniques,” 2018 Innovations in Intelligent Systems and Applications Conference (ASYU), Adana, Turkey, pp. 1-5 (2018), doi: 10.1109/ASYU.2018.8554016.[33] Sugianto, S., & Suyanto, S. : Voting-based music genre classification using mel spectrogram and convolutional neural network. In 2019 International Seminar on Research of Information Technology and Intelligent Systems (ISRITI), pp. 330-333 (2019, December). IEEE.[34] Pelchat, N., & Gelowitz, C. M. : Neural network music genre classification. Canadian Journal of Electrical and Computer Engineering, 43(3), 170-173 (2020).[35] Singh, Y., & Biswas, A. : Robustness of musical features on deep learning models for music genre classification. Expert Systems with Applications, 199, 116879 (2022).[36] N. Ndou, R. Ajoodha and A. Jadhav, “Music Genre Classification: A Review of Deep-Learning and Traditional Machine-Learning Approaches,” 2021 IEEE International IOT, Electronics and Mechatronics Conference (IEMTRONICS), Toronto, ON, Canada,  pp. 1-6 (2022), doi: 10.1109/IEMTRONICS52119.2021.9422487.[37] Hore, P., & Sharma, A. : Code-switched end-to-end Marathi speech recognition for especially abled people. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 771-784 (2022).[38] Umamaheswaran, S., John, R., Deepthi, S. S., & Dharmarajlu, S. M. : Caption positioning structure for hard of hearing people using deep learning method. Journal of Discrete Mathematical Sciences and Cryptography, 25(3), 623-633 (2022).[39] Gupta, V., Juyal, S., Singh, G. P., Killa, C., & Gupta, N. : Emotion recognition of audio/speech data using deep learning approaches. Journal of Information and Optimization Sciences, 41(6), 1309-1317 (2020).[40] Jaber, H. Q., & Abdulbaqi, H. A. : Real time Arabic speech recognition based on convolution neural network. Journal of Information and Optimization Sciences, 42(7), 1657-1663 (2021).[41] Poulose, A., Reddy, C. S., Dash, S., & Sahu, B. J. R. : Music recommender system via deep learning. Journal of Information and Optimization Sciences, 43(5), 1081-1088 (2022).
Views: 368Downloads: 22Citations: 61