CFP last date
20 October 2026
Reseach Article

A Hybrid Deep Learning Framework for Detecting and Mitigating Hallucinations in Large Language Models

by Moqbel Tawhib Salah Abdo, Kasmi Manal
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 142
Year of Publication: 2026
Authors: Moqbel Tawhib Salah Abdo, Kasmi Manal
10.5120/ijca14a09bd7b6b5

Moqbel Tawhib Salah Abdo, Kasmi Manal . A Hybrid Deep Learning Framework for Detecting and Mitigating Hallucinations in Large Language Models. International Journal of Computer Applications. 187, 142 ( Oct 2026), 71-77. DOI=10.5120/ijca14a09bd7b6b5

@article{ 10.5120/ijca14a09bd7b6b5,
author = { Moqbel Tawhib Salah Abdo, Kasmi Manal },
title = { A Hybrid Deep Learning Framework for Detecting and Mitigating Hallucinations in Large Language Models },
journal = { International Journal of Computer Applications },
issue_date = { Oct 2026 },
volume = { 187 },
number = { 142 },
month = { Oct },
year = { 2026 },
issn = { 0975-8887 },
pages = { 71-77 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number142/a-hybrid-deep-learning-framework-for-detecting-and-mitigating-hallucinations-in-large-language-models/ },
doi = { 10.5120/ijca14a09bd7b6b5 },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-10-11T01:23:33.555164+05:30
%A Moqbel Tawhib Salah Abdo
%A Kasmi Manal
%T A Hybrid Deep Learning Framework for Detecting and Mitigating Hallucinations in Large Language Models
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 142
%P 71-77
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

Large language models (LLMs) such as GPT-4, LLaMA-3, and Gemini 1.5 remain prone to generating fluent but factually incorrect content—a failure mode termed hallucination. This paper presents HalluGuard v2, a novel hybrid deep learning framework integrating a Self-Consistency Checker (SCC) and a Retrieval-Based Verifier (RBV) through a Query-Adaptive Gating Network (QA-GN). The QA-GN is a lightweight two-layer MLP that dynamically predicts per-query fusion weights, enabling context-sensitive combination of intra-model consistency signals and evidence-grounded verification. An Adversarial Probing Module activates as a post-hoc step for low-confidence factual decisions, challenging the model to expose fragile false confidence. A Triple-Signal OOD Handler activates dynamic web-search verification when retrieval confidence falls below a corpus-coverage threshold. A new expert-annotated benchmark, HalluDetect-3K, comprising 3,000 QA pairs spanning six domains and five hallucination categories is constructed. HalluGuard v2 achieves 90.3% accuracy and 88.1% macro F1-score on HalluDetect-3K, outperforming the strongest prior baseline by 11.4 and 11.7 percentage points, with AUC-ROC = 0.941 confirmed by five-fold cross-validation (88.9% ± 0.35%). The QA-GN contributes 5.9 pp over static logistic regression; adversarial probing adds a further 0.7 pp. A dedicated adversarial robustness evaluation demonstrates that a Suspicion Score mechanism reduces evasion from 31.4% to 12.5%. Corrected responses achieve 4.61/5 factual accuracy while preserving 91.2% BERTScore fluency. Latency optimisation via token-probability entropy approximation reduces inference from 3.6 s to 1.2 s at a 2.4 pp accuracy cost.

References
  1. OpenAI. (2023). GPT-4 Technical Report. arXiv:2303.08774.
  2. Meta AI. (2024). LLaMA 3: Open Foundation Large Language Models. Meta AI Research Blog.
  3. Google. (2024). Gemini 1.5: Unlocking multi-modal understanding across millions of tokens. arXiv:2403.05530.
  4. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., & Fung, P. (2023). Survey of hallucination in natural lan-guage generation. ACM Computing Surveys, 55(12), 1–38.
  5. Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On faithfulness and factuality in abstractive summarization. In Proc. ACL 2020 (pp. 1906–1919).
  6. Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., & Liu, T. (2023). A survey on hallucination in LLMs. arXiv:2311.05232.
  7. Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., & Zhou, D. (2023). Self-consistency im-proves chain of thought reasoning. In Proc. ICLR 2023.
  8. Manakul, P., Liusie, A., & Gales, M. J. (2023). Self-CheckGPT: Zero-resource black-box hallucination de-tection for generative LLMs. In Proc. EMNLP 2023 (pp. 9004–9017).
  9. Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W. T., Koh, P. W., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision. In Proc. EMNLP 2023.
  10. Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling LLMs to generate text with citations. In Proc. EMNLP 2023.
  11. Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., & Weston, J. (2023). Chain-of-verification reduces hallucination in LLMs. arXiv:2309.11495.
  12. Hu, Y., Tu, T., Koh, P. W., Huang, J., & Salakhutdinov, R. (2023). RefChecker: Reference-based fine-grained hallucination checker for LLMs. arXiv:2405.14486.
  13. Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2023). Self-RAG: Learning to retrieve, generate, and critique. arXiv:2310.11511.
  14. Kuhn, L., Gal, Y., & Farquhar, S. (2023). Semantic un-certainty: Linguistic invariances for uncertainty estima-tion in NLG. In Proc. ICLR 2023.
  15. Tang, L., Peng, P., Chen, B., Ma, S., Pan, X., & Yu, D.(2024). MiniCheck: Efficient fact-checking of LLMs on grounding documents. In Proc. EMNLP 2024.
  16. Joshi, M., Choi, E., Weld, D. S., & Zettlemoyer, L. (2017). TriviaQA: A large scale challenge dataset for reading comprehension. In Proc. ACL 2017.
  17. Jin, D., Pan, E., Oufattole, N., Weng, W. H., Fang, H., & Szolovits, P. (2021). What disease does this patient have? MedQA. Applied Sciences, 11(14), 6421.
  18. Welbl, J., Liu, N. F., & Gardner, M. (2017). Crowd-sourcing multiple choice science questions. In EMNLP Workshop 2017.
  19. Popat, K., Mukherjee, S., Yates, A., & Weikum, G. (2019). DeClarE: Debunking fake news using evidence-aware deep learning. In Proc. EMNLP 2018.
  20. Talmor, A., Herzig, J., Lourie, N., & Berant, J. (2019). CommonsenseQA: A challenge targeting com-monsense knowledge. In Proc. NAACL-HLT 2019.
  21. Guha, N., Nyarko, J., Ho, D., Re´, C., Chilton, A., Nara-hari, A., & Koreeda, Y. (2023). LegalBench: Measuring legal reasoning in LLMs. arXiv:2308.11462
Index Terms

Computer Science
Information Sciences

Keywords

Large language models (LLMs) such as GPT-4 LLaMA-3 and Gemini 1.5 remain prone to generating fluent but factually incorrect content—a failure mode termed hallucination. This paper presents HalluGuard v2 a novel hybrid deep learning framework integrating a Self-Consistency Checker (SCC) and a Retrieval-Based Verifier (RBV) through a Query-Adaptive Gating Network (QA-GN). The QA-GN is a lightweight two-layer MLP that dynamically predicts per-query fusion weights enabling context-sensitive combination of intra-model consistency signals and evidence-grounded verification. An Adversarial Probing Module activates as a post-hoc step for low-confidence factual decisions challenging the model to expose fragile false confidence. A Triple-Signal OOD Handler activates dynamic web-search verification when retrieval confidence falls below a corpus-coverage threshold. A new expert-annotated benchmark HalluDetect-3K comprising 3 000 QA pairs spanning six domains and five hallucination categories is constructed. HalluGuard v2 achieves 90.3% accuracy and 88.1% macro F1-score on HalluDetect-3K outperforming the strongest prior baseline by 11.4 and 11.7 percentage points with AUC-ROC = 0.941 confirmed by five-fold cross-validation (88.9% ± 0.35%). The QA-GN contributes 5.9 pp over static logistic regression; adversarial probing adds a further 0.7 pp. A dedicated adversarial robustness evaluation demonstrates that a Suspicion Score mechanism reduces evasion from 31.4% to 12.5%. Corrected responses achieve 4.61/5 factual accuracy while preserving 91.2% BERTScore fluency. Latency optimisation via token-probability entropy approximation reduces inference from 3.6 s to 1.2 s at a 2.4 pp accuracy cost.