Open Access
·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate
The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to:
• Information Sciences
• Optimization Sciences
• Control Theory
• Operational Research
• Decision Sciences
• Information Theory
• Information Technology
• Computer Networks and Communications
• Mathematical Programming
• Modelling and Simulation
• Database Management
• Applications to Engineering Sciences
• Applications to Technology
Issues up to 2022 co-published with and available at:
Exploring alternative techniques for accelerating deep learning convergence for batch normalization
Shyam Deshmukhdshyam100@yahoo.comDepartment of Information Technology Pune Institute of Computer TechnologyPune, Maharashtra, 411043, IndiaView full profile →
, *Bhavana TipleCorresponding authorbhavana.tiple@mitwpu.edu.inDepartment of Computer Engineering and Technology Dr. Vishwanath Karad MIT-World Peace UniversityDepartment of Computer Engineering and Technology Dr. Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, 411038, IndiaView full profile →
, Gurunath T. Chavangt.chavan@gmail.comDepartment of Information Technology Vishwakarma Institute of TechnologyPune, Maharashtra, 411037, IndiaView full profile →
, Madhura Phatakmadhura.phatak@mitwpu.edu.inDepartment of Computer Engineering and Technology Dr. Vishwanath Karad MIT-World Peace UniversityDepartment of Computer Engineering and Technology Dr Vishwanath Karad MIT World Peace UniversityPune, Maharashtra, 411038, IndiaView full profile →
, Gaurav Gondhalekargg.ycce@gmail.comDepartment of Electrical Engineering Yeshwatrao Chavan College of EngineeringNagpur, Maharashtra, 441110, IndiaView full profile →
, Rajesh B. Rautrautrb@rknec.eduDepartment of Electronics and Communication Engineering Ramdeobaba UniversityNagpur, Maharashtra, 440013, IndiaView full profile →
* Corresponding author · click or hover a name for details
Deep learning models are the most important part of current AI systems because they can learn complex patterns from data in a very impressive way. However, creating these models often takes a lot of time and computer power, especially when methods like Batch Normalization (BN) are used. By balancing activations, BN helps to stabilize and speed up neural network training, but it adds extra work to the computer, which slows down convergence speed. This paper looks at different ways to speed up the convergence of deep learning models that use BN. This study looks into different ways to slow down the convergence process caused by BN layers by reading a lot of current research and doing experiments. Some methods, like weight normalization, group normalization, and layer normalization, are looked at to see if they can be used instead of or in addition to BN to keep or speed up convergence. It also looked into how to change the architecture and make optimization techniques that work with these different standardization methods. Our results show that while BN is still a popular way to keep neural network training stable, other methods show promise for speeding up convergence without affecting performance. By using these other approaches, professionals can cut down on the amount of work that needs to be done on computers and the time it takes to train models. This makes it easier to build and use deep learning models in places with limited resources. This study adds to the ongoing search for deep learning methods that work well and can be used by many people. It also opens up new areas for research and improvement in model training.
[1] S. Wu, G. Li, F. Chen, and L. Shi, “L1 -Norm Batch Normalization for Efficient Training of Deep Neural Networks,” in IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 7, pp. 2043-2051 (July 2019).[2] Y. S. Ting, Y. F. Teng and T. D. Chiueh, “Batch Normalization Processor Design for Convolution Neural Network Training and Inference,” 2021 IEEE International Symposium on Circuits and Systems (ISCAS), Daegu, Korea, pp. 1-4 (2021)[3] A. Muhammad, F. Shamshad and S. H. Bae, “Adversarial Attacks and Batch Normalization: A Batch Statistics Perspective,” in IEEE Access, vol. 11, pp. 96449-96459 (2023).[4] M. Alwani, H. Chen, M. Ferdman, and P. Milder, “Fused-layer CNN accelerators,” in Proceedings of the 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pp. 1–12 (2016), doi: 10.1109/MICRO.2016.7783722.[5] Z. Yang, B. Ni, X. Huang, and D. Chen, “Bactran: A hardware batch normalization implementation for CNN training engine,” IEEE Embedded Systems Letters, vol. 12, no. 4, pp. 111–114 (Dec. 2020), doi: 10.1109/LES.2020.3002759.[6] S. Du, J. Liu and L. Zhao, “Research on Rolling Bearing Condition Monitoring Method Based on Deep Learning,” 2020 Chinese Automation Congress (CAC), Shanghai, China, pp. 74-78 (2020).[7] Y. Li, T. Zhang, and C. Shan, Deep Learning Convolutional Neural Network: From Entry to Mastery. Beijing, China: Mechanical Industry Press (2018).[8] W. F. Hidayat, M. F. Julianto, Y. Malau, A. Setiadi and Sriyadi, “Implementation of LSTM and Adam Optimization as a Cryptocurrency Polygon Price Predictor,” 2023 International Conference on Information Technology Research and Innovation (ICITRI), Jakarta, Indonesia, pp. 123-127 (2023).[9] G. Ardiyansyah, F. Ferdiansyah and U. Ependi, “Deep Learning Model Analysis and Web-Based Implementation of Cryptocurrency Prediction”, J. Inf. Syst. Informatics, vol. 4, no. 4, pp. 958-974 (2022).[10] K. Maharana, S. Mondal and B. Nemade, “A review: Data pre-processing and data augmentation techniques”, Glob. Transitions Proc., vol. 3, no. 1, pp. 91-99 (2022).[11] S. Saika, S. Shewali, M. J. Borah, and D. J. Mahanta, “Arithmetic mean of interval valued intuitionistic fuzzy soft sets and its application in diagnosing infectious diseases,” J. Interdiscip. Math., vol. 27, no. 5, pp. 1079–1090 (2024), doi: 10.47974/JIM-1934.[12] P. Benz, C. Zhang and I. S. Kweon, “Batch normalization increases adversarial vulnerability and decreases adversarial transferability: A non-robust feature perspective”, Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), pp. 7798-7807 (Oct. 2021).[13] J. Bruce and D. Erman, “A probabilistic approach to systems of parameters and noether normalization”, arXiv:1604.01704 (2016).
Views: 166Downloads: 9Citations: 0
Install Journal of Information and Optimization SciencesFaster access from your home screen