TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Enhancing information retrieval through advanced query expansion techniques and semantic analysis

* ,

* Corresponding author · click or hover a name for details

pp. 1587–1603Vol. 46Issue 5July 2025DOI: 10.47974/JIOS-1839XML
Received:
09 Apr 2024
Published Online:
29 Apr 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1839
Pages:
1587–1603

Abstract

Optimizing the speed and precision with which useful information could be retrieved is the primary goal of Information Retrieval Enhancement. The authors discuss their motivation for creating better data retrieval methods in this work. The study investigates the use of semantic analysis and complex Query Expansion (QE) techniques to enhance the quality of search engine results. This research aims to enhance current techniques for retrieving data by using cutting-edge algorithms and semantic analysis tools. This study proposed the novel Bidirectional Encoder Representations from Transformers (BERT) and conditional random fields (CRF) to extract information from the Yahoo Dataset. Contextualized word embeddings with BERT and sequence labelling with CRF improve data analysis and extraction. The experimental results presented in the research demonstrate the effectiveness of these enhanced tactics in producing more precise and relevant search results by intelligently extending query terms and understanding the context of user searches. The results indicate that the proposed method exhibits a high degree of accuracy (99.7%), a substantial rate of recall (98.0%), a considerable rate of precision (99.8%), and a noteworthy value of F1-score (98%). Future research should focus on intelligent search engines that make it easy to access large volumes of data from several sources. These tools would use state-of-the-art algorithms for machine learning (ML) and methods for natural language processing.

Keywords

Subject Classifications

Primary 68T50 Natural language processingSecondary 68P20 Information storage and retrieval of data

References

[1] T. Guo, R. Zhou, and C. Tian, “On the information leakage in private information retrieval systems,” IEEE Trans. Inf. Forensics Security, vol. 15, pp. 2999-3012 (2020).
[2] T. Barik and V. Singh, “AQtpUIR: Adaptive query term proximity based user information retrieval,” J. Inf. Optim. Sci., vol. 41, no. 6, pp. 1479–1497 (2020), doi: 10.1080/02522667.2020.1820190.
[3] V. Suma, “A novel information retrieval system for the distributed cloud using a hybrid deep fuzzy hashing algorithm,” J. Inf. Technol. Digit. World, vol. 2, no. 3, pp. 151-160 (2020).
[4] L. Afuan, A. Ashari, and Y. Suyanto, “A study: query expansion methods in information retrieval,” in J. Phys.: Conf. Ser., vol. 1367, no. 1, p. 012001 (2019).
[5] V. Suma, “A novel information retrieval system for the distributed cloud using a hybrid deep fuzzy hashing algorithm,” J. Inf. Technol. Digit. World, vol. 2, no. 3, pp. 151-160 (2020).
[6] S. Shaukat, A. Shaukat, K. Shahzad, and A. Daud, “Using TREC for developing semantic information retrieval benchmark for Urdu,” Inf. Process. Manag., vol. 59, no. 3, p. 102939 (2022).
[7] G. Pasi, G. J. F. Jones, L. Goeuriot, L. Kelly, S. Marrara, and C. Sanvitto, “Overview of the CLEF 2019 personalized information retrieval lab (PIR-CLEF 2019),” in Exp. IR Meets Multilinguality, Multimodality, Interaction: 10th Int. Conf. CLEF Assoc., Lugano, Switzerland, pp. 417-424 (2019).
[8] I. Soboroff, “Overview of TREC 2021,” in 30th Text Retrieval Conf., Gaithersburg, Maryland (2021).
[9] L. Goeuriot, P. Ruch, J. J. McAuley, D. G. Oppenheim, M. Y. Tse, M. C. G. Scherp, J. Garcia, and B. P. D. Rojas, “Overview of the CLEF eHealth evaluation lab 2020,” in Int. Conf. Cross-Language Eval. Forum Eur. Lang., Cham, Switzerland, pp. 255-271 (2020).
[10] H. K. Azad and A. Deepak, “Query expansion techniques for information retrieval: a survey,” Inf. Process. Manag., vol. 56, no. 5, pp. 1698-1735 (2019).
[11] V. Raskin and J. M. Taylor, “Fuzziness, uncertainty, vagueness, possibility, and probability in natural language,” in 2014 IEEE Conf. Norbert Wiener 21st Century (21CW), pp. 1-6 (2014).
[12] L. Zhao and J. Callan, “Automatic term mismatch diagnosis for selective query expansion,” in Proc. 35th Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 515-524 (2012).
[13] P. Gupta, K. Bali, R. E. Banchs, M. Choudhury, and P. Rosso, “Query expansion for mixed-script information retrieval,” in Proc. 37th Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 677-686 (2014).
[14] M. Fernández, J. M. Albornoz, and M. O. M. Gutiérrez, “Semantically enhanced information retrieval: An ontology-based approach,” J. Web Semantics, vol. 9, no. 4, pp. 434-452 (2011).
[15] M. W. Glier, D. A. McAdams, and J. S. Linsey, “Exploring automated text classification to improve keyword corpus search results for bioinspired design,” J. Mech. Des., vol. 136, no. 11, p. 111103 (2014).
[16] S. Lim and C. S. Tucker, “A Bayesian sampling method for product feature extraction from large-scale textual data,” J. Mech. Des., vol. 138, no. 6, p. 061403 (2016).
[17] L. Lan, Y. Liu, and W. F. Lu, “Discovering a hierarchical design process model using text mining,” in Int. Design Eng. Tech. Conf. Comput. Inf. Eng. Conf., vol. 50084 (2016), p. V01BT02A024.
[18] Y. Rezgui, S. Boddy, M. Wetherill, and G. Cooper, “Past, present and future of information and knowledge sharing in the construction industry: Towards semantic service-based e-construction?” Comput.-Aided Des., vol. 43, no. 5, pp. 502-515 (2011).
[19] X. Chang, R. Rai, and J. Terpenny, “Development and utilization of ontologies in design for manufacturing,” p. 021009 (2010).
[20] Y. Liu, S. C. J. Lim, and W. B. Lee, “Product family design through ontology-based faceted component analysis, selection, and optimization,” J. Mech. Des., vol. 135, no. 8, p. 081007 (2013).
[21] F. Shi, L. Chen, J. Han, and P. Childs, “A data-driven text mining and semantic network analysis for design information retrieval,” J. Mech. Des., vol. 139, no. 11, p. 111402 (2017).
[22] Y. Ding and S. Foo, “Ontology research and development. Part 1-a review of ontology generation,” J. Inf. Sci., vol. 28, no. 2, pp. 123-136 (2002).
[23] R. Jagerman, H. Zhuang, Z. Qin, X. Wang, and M. Bendersky, “Query expansion by prompting large language models,” arXiv preprint arXiv:2305.03653 (2023).
[24] T. Russell-Rose, P. Gooch, and U. Kruschwitz, “Interactive query expansion for professional search applications,” Bus. Inf. Rev., vol. 38, no. 3, pp. 127-137 (2021).
[25] L. Gao, F. Qi, X. Qian, Z. Yang, and J. Zhang, “Complement lexical retrieval model with semantic residual embeddings,” in Adv. Inf. Retrieval: 43rd Eur. Conf. IR Res., ECIR 2021, Virtual Event, Mar. 28–Apr. 1, 2021, Proc., Part I 43, Springer Int. Publishing, pp. 146-160 (2021).
[26] H. ALMarwi, M. Ghurab, and I. Al-Baltah, “A hybrid semantic query expansion approach for Arabic information retrieval,” J. Big Data, vol. 7, no. 1, pp. 1-19 (2020).
[27] M. U. Devi and G. M. Gandhi, “Scalable information retrieval system in the semantic web by query expansion and ontological-based LSA ranking similarity measurement,” Int. J. Adv. Intell. Paradigms, vol. 17, no. 1-2, pp. 44-66 (2020).
[28] H. Zamani, S. Dehghani, and W. B. Croft, “Generating clarifying questions for information retrieval,” in Proc. Web Conf. 2020, pp. 418-428 (2020).
[29] Y. Qu, L. Chen, and X. Gao, “RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering,” arXiv preprint arXiv:2010.08191 (2020).
[30] V. Karpukhin, L. MacAvaney, D. M. Radev, J. Miller, and R. Lewis, “Dense passage retrieval for open-domain question answering,” arXiv preprint arXiv:2004.04906 (2020).
[31] I. Safder and S. Hassan, “Bibliometric-enhanced information retrieval: a novel deep feature engineering approach for algorithm searching from full-text publications,” Scientometrics, vol. 119, pp. 257-277 (2019).
[32] J. Liu, P. Wang, and Y. Zhou, “Neural query expansion for code search,” in Proc. 3rd ACM Singplan Int. Workshop Mach. Learn. Program. Lang., pp. 29-37 (2019).
[33] H. K. Azad and A. Deepak, “A new approach for query expansion using Wikipedia and WordNet,” Inf. Sci., vol. 492, pp. 147-163 (2019).
[34] H. Khalifi, H. Zare, and M. M. Sadjadi, “Query expansion based on clustering and personalized information retrieval,” Prog. Artif. Intell., vol. 8, pp. 241-251 (2019).
[35] M. Morik, M. Kim, and P. M. M. Patel, “Controlling fairness and bias in dynamic learning-to-rank,” in Proc. 43rd Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, pp. 429-438 (2020).
[36] Q. Song, J. Zhang, and R. M. Raj, “Stock portfolio selection using learning-to-rank algorithms with news sentiment,” Neurocomputing, vol. 264, pp. 20-28 (2017).
[37] Z. Hu, Z. Wang, and D. M. Martinez, “Unbiased lambdamart: an unbiased pairwise learning-to-rank algorithm,” in World Wide Web Conf., pp. 2830-2836 (2019).
[38] R. de FSM Russo and R. Camanho, “Criteria in AHP: A systematic review of literature,” Procedia Comput. Sci., vol. 55, pp. 1123-1132 (2015).
[39] A. Darko, M. Osei, and D. D. Nyarku, “Review of the application of analytic hierarchy process (AHP) in construction,” Int. J. Constr. Manag., vol. 19, no. 5, pp. 436-452 (2019).
[40] A. Pereira, M. J. Silva, and S. Pinto, “Querying semantic catalogues of biomedical databases,” J. Biomed. Inf., vol. 137, p. 104272 (2023).
[41] J.-P. Calbimonte, D. R. Mercier, and F. Ciravegna, “Decentralized semantic provision of personal health streams,” J. Web Semantics, vol. 76, p. 100774 (2023).
[42] F. Zhao, Z. Zhu, and P. Han, “A novel model for semantic similarity measurement based on WordNet and word embedding,” J. Intell. Fuzzy Syst., vol. 40, no. 5, pp. 9831-9842 (2021).
[43] X. Pu, L. Yuan, J. Leng, T. Wu, and X. Gao, “Lexical knowledge enhanced text matching via distilled word sense disambiguation,” Knowl.-Based Syst., vol. 263, p. 110282 (2023).
[44] M. Bevilacqua, T. Pasini, A. Raganato, and R. Navigli, “Recent trends in word sense disambiguation: A survey,” in Proc. Int. Joint Conf. Artif. Intell., pp. 4330-4338 (2021).
[45] H. Husni, Y. Kustiyahningsih, F. H. Rachman, E. M. S. Rochman, and H. Yulian, “Query expansion using pseudo relevance feedback based on the Bahasa version of the Wikipedia dataset,” in AIP Conf. Proc., vol. 2679, no. 1, AIP Publishing (2023).
[46] M. J. Hussain, H. Bai, S. H. Wasti, G. Huang, and Y. Jiang, “Evaluating semantic similarity and relatedness between concepts by combining taxonomic and non-taxonomic semantic features of WordNet and Wikipedia,” Inf. Sci., vol. 625, pp. 673-699 (2023).
[47] G. A. Miller and C. Fellbaum, “WordNet then and now,” Lang. Resour. Eval., vol. 41, pp. 209-214 (2007).
[48] G. Puccetti, V. Giordano, I. Spada, F. Chiarello, and G. Fantoni, “Technology identification from patent texts: A novel named entity recognition method,” Technol. Forecast. Soc. Change, vol. 186, p. 122160, 2023.J. Mao and W. Liu, “Hadoken: a BERT-CRF Model for Medical Document Anonymization,” in IberLEF@ SEPLN, pp. 720-726 (2019).
[49] M. Liu, Z. Tu, Z. Wang, and X. Xu, “LTP: A new active learning strategy for BERT-CRF based named entity recognition,” arXiv preprint arXiv:2001.02524 (2020).
[50] X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” Advances in Neural Information Processing Systems, vol. 28 (2015).
[51] N. Moniz and L. Torgo, “Multi-source social feedback of online news feeds,” arXiv preprint arXiv:1801.07055 (2018).
[52] T. Barik and V. Singh, “AQtpUIR: Adaptive query term proximity based user information retrieval,” Journal of Information & Optimization Sciences, vol. 41, no. 6, pp. 1479–1497 (2020), doi: 10.1080/02522667.2020.1820190.
[53] A. Pawar and V. Mago, “Calculating the similarity between words and sentences using a lexical database and corpus statistics,” (2018). arXiv preprint arXiv:1802.05667.

Views: 383Downloads: 87Citations: 0