TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

A lightweight vision-language framework for infrastructure-free indoor navigation

, , , , *

* Corresponding author · click or hover a name for details

pp. 2723–2741Vol. 47Issue 7July 2026DOI: 10.47974/JIOS-2327XML
Received:
01 Dec 2025
Published Online:
31 Jul 2026
Article type:
Research Article
Language:
EN
Article no.:
JIOS-2327
Pages:
2723–2741

Abstract

A lightweight vision-only indoor navigation framework that constructs a room connectivity graph from a simple monocular video walkthrough and uses it for real-time guidance. Our approach harnesses vision-language models and classical computer vision techniques: we use CLIP image embeddings to recognize scenes, optical flow to estimate motion (detecting turns and transitions), and blur detection to filter out unclear frames. Each distinct room in the environment is automatically identified and represented by three key reference images (capturing the start, middle, and end views of the room), which serve as nodes in a topological graph. The graph encodes how rooms connect and the direction of movement between them (e.g. forward transitions, left or right turns). During live navigation, frames from a user’s smartphone or an IP camera are continuously matched against these stored references via CLIP-based similarity to localize the user’s position. A breadth-first search (BFS) on the room graph then finds the shortest path to the desired destination, and a text-to-speech module provides step-by-step audio instructions (for example, “Turn right to enter the corridor”). We evaluate our system on custom indoor video datasets: a fine-tuned CLIP ViT-B/32 model achieved an accuracy of 0.938 with an F1-score of 0.940, markedly higher than the ResNet’s 0.852 accuracy and 0.847 F1-score. Similarly, the ViT attained superior precision (0.912 vs 0.762) and slightly higher recall(0.971 vs 0.958).

Keywords

Subject Classifications

68T4590B40

References

[1] A. Cheddad, J. Condell, K. Curran, and P. M. Kevitt, “Digital image steganography: Survey and analysis of current methods,” Signal Processing, vol. 90, no. 3, pp. 727–752 (Mar. 2010).

[2] F. Zafari, A. Gkelias, and K. K. Leung, “A survey of indoor localization systems and technologies,” IEEE Communications Surveys & Tutorials, vol. 21, no. 3, pp. 2568–2599 (2019).

[3] R. Faragher and R. Harle, “Location fingerprinting with Bluetooth Low Energy beacons,” IEEE Journal on Selected Areas in Communications, vol. 33, no. 11, pp. 2418–2428 (Nov. 2015).

[4] A. Alarifi, A. Al-Salman, M. Alsaleh, A. Alnafessah, S. Al-Hadhrami, M. Al-Ammar, and H. Al-Khalifa, “Ultra-wideband indoor positioning technologies: Analysis and recent advances,” Sensors, vol. 16, no. 5, Art. no. 707 (2016).

[5] R. Mur-Artal and J. D. Tardós, “ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras,” IEEE Transactions on Robotics, vol. 33, no. 5, pp. 1255–1262 (Oct. 2017).

[6] J. Engel, T. Schöps, and D. Cremers, “LSD-SLAM: Large-scale direct monocular SLAM,” in Proc. European Conf. Computer Vision (ECCV), pp. 834–849 (2014).

[7] A. J. Davison, I. D. Reid, N. D. Molton, and O. Stasse, “MonoSLAM: Real-time single camera SLAM,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, no. 6, pp. 1052–1067 (Jun. 2007).

[8] G. Klein and D. Murray, “Parallel tracking and mapping for small AR workspaces,” in Proc. IEEE/ACM Int. Symp. Mixed and Augmented Reality (ISMAR), pp. 225–234 (2007).

[9] J. Engel, V. Koltun, and D. Cremers, “Direct Sparse Odometry,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 3, pp. 611–625 (Mar. 2018).

[10] J. Czarnowski, T. Laidlow, R. Clark, and A. J. Davison, “DeepFactors: Real-time probabilistic dense monocular SLAM,” IEEE Transactions on Robotics, vol. 36, no. 6, pp. 1593–1609 (Dec. 2020).

[11] M. J. Milford and G. F. Wyeth, “SeqSLAM: Visual route-based navigation for sunny summer days and stormy winter nights,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), pp. 1643–1649 (2012).

[12] R. Arandjelović, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, “NetVLAD: CNN architecture for weakly supervised place recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 5297–5307 (2016).

[13] A. Chakravarty, A. Jain, and A. K. Saxena, “Sugarcane disease identification using mobile deep learning solutions,” Journal of Information Technology Management, vol. 17, Special Issue: Intelligent Security and Management, pp. 198–214 (2025), doi: 10.22059/jitm.2025.104554.

[14] A. Torii, R. Arandjelović, J. Sivic, and T. Pajdla, “24/7 place recognition by view synthesis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 2, pp. 257–271 (Feb. 2018).

[15] J. Revaud, M. Douze, C. Schmid, and H. Jégou, “Learning with average precision: Training image retrieval with a listwise loss,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), pp. 5107–5116 (2019).

[16] J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” arXiv preprint arXiv:1804.02767 (Apr. 2018).

[17] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proc. European Conf. Computer Vision (ECCV), pp. 801–818 (2018).

[18] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), pp. 2961–2969 (2017).

[19] M. Cummins and P. Newman, “FAB-MAP: Probabilistic localization and mapping in the space of appearance,” The International Journal of Robotics Research, vol. 27, no. 6, pp. 647–665 (2008).

[20] A. Chakravarty, A. Jain, and A. K. Saxena, “Deep learning approach to sugarcane disease identification: From image analysis to mobile application,” in Proc. 4th Int. Conf. Technological Advancements in Computational Sciences (ICTACS), Tashkent, Uzbekistan, pp. 1696–1702 (2024), doi: 10.1109/ICTACS62700.2024.10840413.

[21] A. Chakravarty, A. Jain, and A. K. Saxena, “Disease detection of plants using deep learning approach—A review,” in Proc. 11th Int. Conf. System Modeling & Advancement in Research Trends (SMART), Moradabad, India, pp. 1285–1292 (Dec. 2022), doi: 10.1109/SMART55829.2022.10047097.

Views: 57Downloads: 41Citations: 0