Designing a structured approach for consistency verification in AI systems for heartbeat disorder-related conversations
*Muhammad Badruddin KhanCorresponding authormbkhan@imamu.edu.saInformation Systems DepartmentCollege of Computer and Information SciencesImam Mohammad Ibn Saud Islamic University (IMSIU)Riyadh, 11432, Saudi ArabiaView full profile → , Abdul Khader Jilani Saudagaraksaudagar@imamu.edu.saInformation Systems DepartmentCollege of Computer and Information SciencesImam Mohammad Ibn Saud Islamic University (IMSIU)Riyadh, 11432, Saudi ArabiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 12 Jun 2024
- Published Online:
- 18 Dec 2024
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JSMS-1402
- Pages:
- 1733–1741
Abstract
Keywords
Subject Classifications
References
[1] W. X. Zhao, X. Y. Li, Z. H. Wang, and A. B. Chen, “A survey of large language models,” arXiv, Oct. 13 (2024), arXiv:2303.18223. doi: 10.48550/arXiv.2303.18223.
[2] I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar, “GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models,” Oct. 07 (2024), arXiv: arXiv:2410.05229. doi: 10.48550/arXiv.2410.05229.
[3] K. Cobbe, A. S. Patel, J. S. Brown, and L. D. Miller, “Training verifiers to solve math word problems,” arXiv, Nov. 18 (2021), arXiv:2110.14168. doi: 10.48550/arXiv.2110.14168.
[4] N. Dziri, S. J. Lee, P. M. Sharma, and R. T. Zhang, “Faith and fate: Limits of transformers on compositionality,” arXiv, Oct. 31 (2023), arXiv:2305.18654. doi: 10.48550/arXiv.2305.18654.
[5] M. Sallam, “ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns,” Healthc. Basel Switz., vol. 11, no. 6, p. 887, Mar. (2023), doi: 10.3390/healthcare11060887.
[6] M. Cascella, J. Montomoli, V. Bellini, and E. Bignami, “Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios,” J. Med. Syst., vol. 47, no. 1, p. 33 (2023), doi: 10.1007/s10916-023-01925-4.
[7] Y. Li, Z. Li, K. Zhang, R. Dan, S. Jiang, and Y. Zhang, “ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge,” Jun. 24 (2023), arXiv: arXiv:2303.14070. doi: 10.48550/arXiv.2303.14070.
[8] Y. Chang, Z. X. Li, M. P. Kim, and A. B. Zhao, “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 3, pp. 39:1-39:45, Mar. (2024), doi: 10.1145/3641289.
[9] Y. Liu, S. M. Zhao, X. T. Zhang, and Q. W. Huang, “Trustworthy LLMs: A survey and guideline for evaluating large language models’ alignment,” arXiv, Mar. 21 (2024), arXiv:2308.05374. doi: 10.48550/arXiv.2308.05374.
[10] Y. Sun, J. K. Wang, L. Z. Xu, and F. G. Yu, “LLM4Vuln: A unified evaluation framework for decoupling and enhancing LLMs’ vulnerability reasoning,” arXiv, Sep. 5 (2024), arXiv:2401.16185. doi: 10.48550/arXiv.2401.16185.




