A lightweight vision-language framework for infrastructure-free indoor navigation
Arpit Jaindr.jainarpit@gmail.comDepartment of Computer Science and EngineeringKoneru Lakshmaiah Education Foundation (Deemed to be University)Vadeshawaram, Andhra Pradesh, 522302, IndiaView full profile → , Bhuvan Unhelkarbunhelkar@usf.eduDepartment of Information TechnologyMuma College of BusinessUniversity of South FloridaTampa, Florida, USAView full profile → , Prasun Chakrabartidrprasun.cse@gmail.comDepartment of Computer Science and EngineeringSir Padampat Singhania UniversityUdaipur, Rajasthan, 313601, India0000-0001-8062-4144View full profile → , Shilpa Choudharyshilpachoudhary1987@gmail.comDepartment of Computer Science and Engineering (AIML)Neil Gogte Institute of TechnologyHyderabad, Telangana, 500088, IndiaView full profile → , *Khemraj SharmaCorresponding authorkhemraj.sharma@kiit.ac.inSchool of ManagementKalinga Institute of Industrial TechnologyBhubaneswar, Odisha, 751024, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 01 Dec 2025
- Published Online:
- 31 Jul 2026
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-2327
- Pages:
- 2723–2741
Abstract
A lightweight vision-only indoor navigation framework that constructs a room connectivity graph from a simple monocular video walkthrough and uses it for real-time guidance. Our approach harnesses vision-language models and classical computer vision techniques: we use CLIP image embeddings to recognize scenes, optical flow to estimate motion (detecting turns and transitions), and blur detection to filter out unclear frames. Each distinct room in the environment is automatically identified and represented by three key reference images (capturing the start, middle, and end views of the room), which serve as nodes in a topological graph. The graph encodes how rooms connect and the direction of movement between them (e.g. forward transitions, left or right turns). During live navigation, frames from a user’s smartphone or an IP camera are continuously matched against these stored references via CLIP-based similarity to localize the user’s position. A breadth-first search (BFS) on the room graph then finds the shortest path to the desired destination, and a text-to-speech module provides step-by-step audio instructions (for example, “Turn right to enter the corridor”). We evaluate our system on custom indoor video datasets: a fine-tuned CLIP ViT-B/32 model achieved an accuracy of 0.938 with an F1-score of 0.940, markedly higher than the ResNet’s 0.852 accuracy and 0.847 F1-score. Similarly, the ViT attained superior precision (0.912 vs 0.762) and slightly higher recall(0.971 vs 0.958).
Keywords
Subject Classifications
References
[1] A. Cheddad, J. Condell, K. Curran, and P. M. Kevitt, “Digital image steganography: Survey and analysis of current methods,” Signal Processing, vol. 90, no. 3, pp. 727–752 (Mar. 2010).
[2] F. Zafari, A. Gkelias, and K. K. Leung, “A survey of indoor localization systems and technologies,” IEEE Communications Surveys & Tutorials, vol. 21, no. 3, pp. 2568–2599 (2019).
[3] R. Faragher and R. Harle, “Location fingerprinting with Bluetooth Low Energy beacons,” IEEE Journal on Selected Areas in Communications, vol. 33, no. 11, pp. 2418–2428 (Nov. 2015).
[4] A. Alarifi, A. Al-Salman, M. Alsaleh, A. Alnafessah, S. Al-Hadhrami, M. Al-Ammar, and H. Al-Khalifa, “Ultra-wideband indoor positioning technologies: Analysis and recent advances,” Sensors, vol. 16, no. 5, Art. no. 707 (2016).
[5] R. Mur-Artal and J. D. Tardós, “ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras,” IEEE Transactions on Robotics, vol. 33, no. 5, pp. 1255–1262 (Oct. 2017).
[6] J. Engel, T. Schöps, and D. Cremers, “LSD-SLAM: Large-scale direct monocular SLAM,” in Proc. European Conf. Computer Vision (ECCV), pp. 834–849 (2014).
[7] A. J. Davison, I. D. Reid, N. D. Molton, and O. Stasse, “MonoSLAM: Real-time single camera SLAM,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, no. 6, pp. 1052–1067 (Jun. 2007).
[8] G. Klein and D. Murray, “Parallel tracking and mapping for small AR workspaces,” in Proc. IEEE/ACM Int. Symp. Mixed and Augmented Reality (ISMAR), pp. 225–234 (2007).
[9] J. Engel, V. Koltun, and D. Cremers, “Direct Sparse Odometry,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 3, pp. 611–625 (Mar. 2018).
[10] J. Czarnowski, T. Laidlow, R. Clark, and A. J. Davison, “DeepFactors: Real-time probabilistic dense monocular SLAM,” IEEE Transactions on Robotics, vol. 36, no. 6, pp. 1593–1609 (Dec. 2020).
[11] M. J. Milford and G. F. Wyeth, “SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1643–1649 (2012).
[12] R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “NetVLAD: CNN architecture for weakly supervised place recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 5297–5307 (2016).
[13] A. Chakravarty, A. Jain, and A. K. Saxena, “Sugarcane disease identification using mobile deep learning solutions,” Journal of Information Technology Management, vol. 17, Special Issue: Intelligent Security and Management, pp. 198–214 (2025), doi: 10.22059/jitm.2025.104554.
[14] A. Torii, R. Arandjelović, J. Sivic, and T. Pajdla, “24/7 place recognition by view synthesis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 2, pp. 257–271 (Feb. 2018).
[15] J. Revaud, M. Douze, C. Schmid, and H. Jégou, “Learning with average precision: Training image retrieval with a listwise loss,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), pp. 5107–5116 (2019).
[16] J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” arXiv preprint arXiv:1804.02767 (Apr. 2018).
[17] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proc. European Conf. Computer Vision (ECCV), pp. 801–818 (2018).
[18] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), pp. 2961–2969 (2017).
[19] M. Cummins and P. Newman, “FAB-MAP: Probabilistic localization and mapping in the space of appearance,” The International Journal of Robotics Research, vol. 27, no. 6, pp. 647–665 (2008).
[20] A. Chakravarty, A. Jain, and A. K. Saxena, “Deep learning approach to sugarcane disease identification: From image analysis to mobile application,” in Proc. 4th Int. Conf. Technological Advancements in Computational Sciences (ICTACS), Tashkent, Uzbekistan, pp. 1696–1702 (2024), doi: 10.1109/ICTACS62700.2024.10840413.
[21] A. Chakravarty, A. Jain, and A. K. Saxena, “Disease detection of plants using deep learning approach—A review,” in Proc. 11th Int. Conf. System Modeling & Advancement in Research Trends (SMART), Moradabad, India, pp. 1285–1292 (Dec. 2022), doi: 10.1109/SMART55829.2022.10047097.




