Review of speaker recognition : Concepts, challenges, architectures, and future directions
*Neelam NehraCorresponding authorneelamnehra@msit.inDepartment of Electronics & Communication EngineeringMaharaja Surajmal Institute of TechnologyJanakpuri, Delhi, 110058, IndiaView full profile → , Geetanjali Sharmagsharma@msit.inDepartment of Electronics & Communication EngineeringMaharaja Surajmal Institute of TechnologyJanakpuri, Delhi, 110058, IndiaView full profile → , Parveen Kumarparveen@msit.inDepartment of Information TechnologyMaharaja Surajmal Institute of TechnologyJanakpuri, Delhi, 110058, IndiaView full profile → , Dinesh Sheorandineshsheoran@msit.inDepartment of Electronics & Communication EngineeringMaharaja Surajmal Institute of TechnologyJanakpuri, Delhi, 110058, IndiaView full profile → , Amita Yadavamita.yadav@msit.inDepartment of Computer Science and EngineeringMaharaja Surajmal Institute of TechnologyJanakpuri, Delhi, 110058, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 07 Aug 2024
- Published Online:
- 19 Feb 2025
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-1857
- Pages:
- 121–132
Abstract
Keywords
Subject Classifications
References
[1] J. P. Campbell, “Speaker recognition: A tutorial,” Proceedings of the IEEE, vol. 85, no. 9, pp. 1437-1462 (1997).
[2] A. K. Jain, A. Ross, and S. Prabhakar, “An introduction to biometric recognition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 14, no. 1, pp. 4-20 (2004).
[3] D. Deshwal, P. Sangwan, and D. Kumar, “A structured approach towards robust database collection for language identification,” in 2020 21st International Arab Conference on Information Technology (ACIT), pp. 1-6 (2020).
[4] D. Deshwal, P. Sangwan, N. Dahiya, N. Nehra, and A. Dahiya, “A comprehensive approach for performance evaluation of Indian language identification systems,” Journal of Intelligent & Fuzzy Systems, (Preprint), pp. 1-17.
[5] F. Richardson, D. Reynolds, and N. Dehak, “A unified deep neural network for speaker and language recognition,” arXiv preprint arXiv:1504.00923 (2015).
[6] N. Dehak, P. A. Torres-Carrasquillo, D. Reynolds, and R. Dehak, “Language recognition via i-vectors and dimensionality reduction,” in Twelfth Annual Conference of the International Speech Communication Association (2011).
[7] K. J. Devi and K. Thongam, “Automatic speaker recognition with enhanced swallow swarm optimization and ensemble classification model from speech signals,” Journal of Ambient Intelligence and Humanized Computing, pp. 1-14 (2019).
[8] R. Chakroun and M. Frikha, “Robust features for text-independent speaker recognition with short utterances,” Neural Computing and Applications, pp. 1-21 (2020).
[9] H. Hassanzadeh, J. A. Qadir, S. M. Omer, M. H. Ahmed, and E. Khezri, “Deep learning for speaker recognition: A comparative analysis of 1D-CNN and LSTM models using diverse datasets,” in 2024 4th Interdisciplinary Conference on Electrics and Computer (INTCEC), pp. 1-8 (2024).
[10] N. Chauhan, T. Isshiki, and D. Li, “Speaker recognition using fusion of features with feedforward artificial neural network and support vector machine,” in 2020 International Conference on Intelligent Engineering and Management (ICIEM), pp. 170-176 (2020).
[11] S. Rosenthal, P. Atanasova, G. Karadzhov, M. Zampieri, and P. Nakov, “A large-scale semi-supervised dataset for offensive language identification,” arXiv preprint arXiv:2004.14454 (2020).
[12] H. Mukherjee, S. Ghosh, S. Sen, O. S. Md, K. C. Santosh, S. Phadikar, and K. Roy, “Deep learning for spoken language identification: Can we visualize speech signal patterns?,” Neural Computing and Applications, vol. 31, no. 12, pp. 8483-8501 (2019).
[13] S. Ganji, K. Dhawan, and R. Sinha, “IITG-HingCoS corpus: A Hinglish code-switching database for automatic speech recognition,” Speech Communication, vol. 110, pp. 76-89 (2019).
[14] M. K. Mustafa, T. Allen, and K. Appiah, “A comparative review of dynamic neural networks and hidden Markov model methods for mobile on-device speech recognition,” Neural Computing and Applications, vol. 31, no. 2, pp. 891-899 (2019).
[15] B. Mak, J. C. Junqua, and B. Reaves, “A robust speech/non-speech detection algorithm using time and frequency-based features,” in ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 1, pp. 269-272 (Mar. 1992).
[16] R. V. Pawar, R. M. Jalnekar, and J. S. Chitode, “Review of various stages in speaker recognition system, performance measures and recognition toolkits,” Analog Integrated Circuits and Signal Processing, vol. 94, no. 2, pp. 247-257 (2018).
[17] B. Singh, V. Rani, and N. Mahajan, “Preprocessing in ASR for computer machine interaction with humans: A review,” International Journal of Advanced Research in Computer Science and Software Engineering, vol. 2, no. 3, pp. 396-399 (2012).
[18] B. Singh, V. Rani, and N. Mahajan, “Preprocessing in ASR for computer machine interaction with humans: A review,” International Journal of Advanced Research in Computer Science and Software Engineering, vol. 2, no. 3, pp. 396-399 (2012).
[19] N. Chauhan, T. Isshiki, and D. Li, “Speaker recognition using LPC, MFCC, ZCR features with ANN and SVM classifier for large input database,” in 2019 IEEE 4th International Conference on Computer and Communication Systems (ICCCS), pp. 130-133 (2019).
[20] A. Chowdhury and A. Ross, “Fusing MFCC and LPC features using 1D triplet CNN for speaker recognition in severely degraded audio signals,” IEEE Transactions on Information Forensics and Security, vol. 15, pp. 1616-1629 (2019).
[21] P. K. Nayana, D. Mathew, and A. Thomas, “Performance comparison of speaker recognition systems using GMM and i-vector methods with PNCC and RASTA PLP features,” in 2017 International Conference on Intelligent Computing, Instrumentation and Control Technologies (ICICICT), pp. 438-443 (2017).
[22] P. K. Nayana, D. Mathew, and A. Thomas, “Comparison of text independent speaker identification systems using GMM and i-vector methods,” Procedia Computer Science, vol. 115, pp. 47-54 (2017).
[23] G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5115-5119 (2016).
[24] N. Chen, Y. Qian, and K. Yu, “Multi-task learning for text-dependent speaker verification,” in Sixteenth Annual Conference of the International Speech Communication Association (2015).
[25] S. Dey, S. Madikeri, M. Ferras, and P. Motlicek, “Deep neural network based posteriors for text-dependent speaker verification,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5050-5054 (2016).
[26] A. H. T. Mahde, “Speech recognition by improving the performance of algorithms used in discrimination,” International Journal of Computer Science & Information Technology (IJCSIT), vol. 11 (2019).
[27] A. Kanagasundaram, D. Dean, S. Sridharan, and C. Fookes, “DNN based speaker recognition on short utterances,” arXiv preprint arXiv:1610.03190 (2016).
[28] D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur, “Deep neural network-based speaker embeddings for end-to-end speaker verification,” in 2016 IEEE Spoken Language Technology Workshop (SLT), pp. 165-170 (2016).
[29] P. Kenny, T. Stafylakis, P. Ouellet, V. Gupta, and M. J. Alam, “Deep Neural Networks for extracting Baum-Welch statistics for Speaker Recognition,” in Odyssey, vol. 2014, pp. 293-298 (2014).
[30] D. Deshwal, P. Sangwan, and D. Kumar, “A language identification system using hybrid features and back-propagation neural network,” Applied Acoustics, vol. 164, p. 107289 (2020).
[31] C. C. Bhanja, M. A. Laskar, and R. H. Laskar, “A pre-classification-based language identification for Northeast Indian languages using prosody and spectral features,” Circuits, Systems, and Signal Processing, vol. 38, no. 5, pp. 2266-2296 (2019).
[32] P. Kumar, A. Biswas, A. N. Mishra, and M. Chandra, “Spoken language identification using hybrid feature extraction methods,” arXiv preprint arXiv:1003.5623, 2010.
[33] A. Hajavi and A. Etemad, “A deep neural network for short-segment speaker recognition,” arXiv preprint arXiv:1907.10420 (2019).
[34] P. Sangwan, D. Deshwal, and N. Dahiya, “Performance of a language identification system using hybrid features and ANN learning algorithms,” Applied Acoustics, vol. 175, p. 107815 (2021).




