حرکت به‌سوی بازیابی معنایی در علوم اسلامی: از مدل‌های سنتی تا معماری‌های مدرن

نوع مقاله : مقاله پژوهشی

نویسنده

گروه اشاعه اطلاعات و دانش، پژوهشکده مدیریت اطلاعات و مدارک اسلامی، پژوهشگاه علوم و فرهنگ اسلامی، قم، ایران

10.22081/jikm.2026.74693.1109

چکیده

موتورهای جستجوی کلیدواژه‌ای به دلیل ناتوانی ذاتی در درک چندمعنایی، هم‌معنایی و روابط مفهومی پیچیده معمولاً نتایج نامرتبط و حجیمی ارائه می‌دهند. این مسئله در حوزۀ علوم اسلامی که اصطلاحات تخصصی چندلایه، متون دوزبانه (فارسی و عربی) و ترکیب منحصر‌به‌فردی از زبان کلاسیک و مدرن دارند، به‌مراتب حادتر به نظر می‌رسد. در دهۀ اخیر، با ظهور فناوری‌های نوینی مانند وب معنایی، گراف‌های دانش، مدل‌های زبانی بزرگ (LLMs) و معماری بازیابی‌ـ‌افزوده با تولید (RAG)، فرصت‌های بی‌سابقه‌ای برای ارتقای بازیابی معنایی در حوزه‌های تخصصی فراهم شده است.
هدف پژوهش حاضر، ارائۀ تحلیل نظام‌مند و به‌روز از چالش‌های بازیابی معنایی در علوم اسلامی، بررسی انتقادی رویکردهای موجود در پرتو فناوری‌های جدید و ارائۀ چارچوبی یکپارچه مبتنی‌بر شواهد تجربی با هدف بهبود وضعیت موجود است. این پژوهش با روش مرور نظام‌مند ادبیات و تحلیل محتوای کیفی ۵۲ منبع معتبر بین‌المللی و داخلی (از ۲۰۱۴ تا ۲۰۲۵) انجام شده است. یافته‌ها نشان می‌دهد چالش‌های بازیابی معنایی را می‌توان در چهار سطح طبقه‌بندی کرد: مفهومی (چندمعنایی، هم‌معنایی، سلسله‌مراتبِ مفاهیم)، زبانی (پیچیدگی‌های صرفی عربی، تنوع رسم‌الخطی فارسی، دوزبانگی)، فنی (محدودیت رویکردهای کلیدواژه‌ای، ناتوانی در فهم پرسش‌های پیچیده، توهم مدل‌های زبانی) و ساختاری (نبود هستی‌شناسی استاندارد، پراکندگی داده‌ها، کمبود داده‌های آموزشی).
چارچوب پیشنهادی این پژوهش پنج لایۀ مکمل دارد: (۱) هستی‌شناسی ساختاریافته OWL/SKOS برای علوم اسلامی، (۲) گراف دانش یکپارچۀ چندزبانه، (۳) مدل‌های زبانی بزرگ تنظیم‌شده برای دامنه، (۴) سیستم RAG بومی‌سازی‌شده با بازیابی ترکیبی و (۵) رابط جستجوی هوشمند با بسط پرسش معنایی. شواهد تجربی از پروژه‌های SemanticHadith، FarsBase، AraBERT، مدل Fanar و نتایج کنفرانس QIAS 2025 نشان می‌دهد سیستم‌های ترکیبی بومی‌سازی‌شده می‌توانند در وظایف تخصصی اسلامی حتی از قدرتمندترین مدل‌های عمومی نیز پیشی بگیرند.
 

کلیدواژه‌ها

موضوعات


عنوان مقاله [English]

Toward Semantic Retrieval in Islamic Sciences: From Traditional Models to Modern Architectures

نویسنده [English]

  • Ali Mirarab
Information and Knowledge Dissemination, Islamic Information and Documents Management Research Center, Qom, Iran
چکیده [English]

Keyword-based search engines inherently struggle to understand polysemy, synonymy, and complex conceptual relationships, and therefore often produce large numbers of irrelevant results. This problem is particularly acute in the field of Islamic sciences, where specialized terminology is multilayered, texts are bilingual (Persian and Arabic), and classical and modern forms of language are uniquely combined. Over the past decade, the emergence of new technologies such as the Semantic Web, knowledge graphs, large language models (LLMs), and Retrieval-Augmented Generation (RAG) architectures has created unprecedented opportunities for improving semantic retrieval in specialized domains. The present study aims to provide a systematic and up-to-date analysis of the challenges of semantic retrieval in Islamic sciences, critically examine existing approaches in light of emerging technologies, and propose an integrated evidence-based framework for improving the current state of the field. The study employs a systematic literature review and qualitative content analysis of 52 authoritative international and domestic sources published between 2014 and 2025. The findings indicate that the challenges of semantic retrieval can be classified into four levels: conceptual (polysemy, synonymy, and conceptual hierarchies); linguistic (Arabic morphological complexities, variations in Persian orthography, and bilingualism); technical (limitations of keyword-based approaches, inability to understand complex queries, and hallucinations in language models); and structural (lack of standardized ontologies, data fragmentation, and insufficient training data). The proposed framework consists of five complementary layers: (1) a structured OWL/SKOS ontology for Islamic sciences; (2) an integrated multilingual knowledge graph; (3) domain-adapted large language models; (4) a localized RAG system with hybrid retrieval; and (5) an intelligent search interface incorporating semantic query expansion. Empirical evidence from the SemanticHadith and FarsBase projects, AraBERT, the Fanar model, and findings presented at the QIAS 2025 conference indicates that localized hybrid systems can outperform even the most powerful general-purpose models in specialized Islamic-domain tasks.
 

کلیدواژه‌ها [English]

  • Semantic Retrieval
  • Islamic Sciences
  • Knowledge Graph
  • Large Language Model
  • Persian Language
  • Natural Language Processing
Andago, M., Phoebe, T. P. L., & Thanoun, B. A. M. (2010). Evaluation of a semantic search engine against a keyword search engine using first 20 precision. International Journal for the Advancement of Science & Arts, 1(2), pp. 55-63.
Antoun, W., Baly, F., & Hajj, H. (2020, May). Arabert: Transformer-based model for arabic language understanding. In Proceedings of the 4th workshop on open-source arabic corpora and processing tools, with a shared task on offensive language detection (pp. 9-15).
Asgari-Bidhendi, M., Hadian, A., & Minaei-Bidgoli, B. (2019). Farsbase: The persian knowledge graph. Semantic Web, 10(6), 1169-1196.
Alothman, M., Altammami, A., Alhoshan, M., Almazrua, A., Al-Rasheed, R., Almatham, R., ... & Alfaifi, A. (2026). ARAG: An Agent-Based Hybrid Semantic-Lexical Retrieval-Augmented Generation Pipeline for Arabic Texts. Procedia Computer Science, 275, 682-691.
Bouchekif, A., Rashwani, S., Mohamed, E. S. A., Alkhatib, M., Sbahi, H., Gaben, S., ... & Ghaly, M. (2025, November). Qias 2025: Overview of the shared task on islamic inheritance reasoning and knowledge assessment. In Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks (pp. 851-860).
Darwish, K., & Magdy, W. (2014). Arabic information retrieval. Foundations and Trends in Information Retrieval, 7(4), 239-342. doi:10.1561/1500000031
Abbas, U., Ahmad, M. S., Alam, F., Altinisik, E., Asgari, E., Boshmaf, Y., Boughorbel, S., Chawla, S., Chowdhury, S., Dalvi, F., Darwish, K., Durrani, N., Elfeky, M., Elmagarmid, A., Eltabakh, M., Fatehkia, M., Fragkopoulos, A., Hasanain, M., Hawasly, M., Husaini, M., Jung, S.-G., Lucas, J. K., Magdy, W., Messaoud, S., Mohamed, A., Mohiuddin, T., Mousi, B., Mubarak, H., Musleh, A., Naeem, Z., Ouzzani, M., Popovic, D., Sadeghi, A., Sencar, H. T., Shinoy, M., Sinan, O., Zhang, Y., Ali, A., El Kheir, Y., Ma, X., & Ruan, C. (2025). Fanar: An Arabic-centric multimodal generative AI platform. arXiv. https://arxiv.org/abs/2501.13944.
Farahani, M., Gharachorloo, M., Farahani, M., & Manthouri, M. (2021). ParsBERT: Transformer-based model for Persian language understanding. Neural Processing Letters, 53, 3831-3847. https://doi.org/10.1007/s11063-021-10528-4
Scherp, A., Groener, G., Škoda, P., Hose, K., & Vidal, M. E. (2024). Semantic Web: Past, Present, and Future (with Machine Learning on Knowledge Graphs and Language Models on Knowledge Graphs). arXiv preprint arXiv:2412.17159.
Moniri, S., Schlosser, T., & Kowerko, D. (2024). Investigating the Challenges and Opportunities in Persian Language Information Retrieval through Standardized Data Collections and Deep Learning. Computers, 13(8), 212. https://doi.org/10.3390/computers13080212
Ghafouri, A., Naderi, H., Aghajani Asl, M., & Firouzmandi, M. (2023). IslamicPCQA: A dataset for Persian multi-hop complex question answering in Islamic text resources. arXiv. https://arxiv.org/abs/2304.11664
Jarrar, M. (2021). The Arabic ontology: An Arabic WordNet with ontologically clean content. Applied Ontology Journal, 16(1), 1-26. Doi:10.3233/AO-200241
Jarrar, M., Akra, D., & Hammouda, T. (2024). Alma: Fast Lemmatizer and POS Tagger for Arabic. Procedia Computer Science. 244. 378-387. 10.1016/j.procs.2024.10.212..
Hammouda, T., Jarrar, M., & Khalilia, M. (2024). SinaTools: Open source toolkit for Arabic natural language processing. arXiv. https://arxiv.org/abs/2411.01523.
Kamran, A. B., Abro, B., & Basharat, A. (2023). SemanticHadith: An ontology-driven knowledge graph for the hadith corpus. Journal of Web Semantics, 78, 100797. https://doi.org/10.1016/j.websem.2023.100797
Kamran, A. B., Butt, N. A., & Basharat, A. (2026). Semantic Enrichment of Hadith Corpus—Knowledge Graph Generation From Islamic Text. Semantic Web, 17(2), 22104968261431425.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W. T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, 33, pp 9459-9474.
Malhas, R., & Elsayed, T. (2022). Qur'an QA 2022: Overview of the first shared task on question answering over the Holy Qur'an. In Proceedings of the 5th Workshop )OSACT 2022( (pp. 79-87). Marseille, France: Association for Computational Linguistics.
Malhas, R., & Elsayed, T. (2020). AyaTEC: Building a reusable verse-based test collection for Arabic question answering on the Holy Qur’an. ACM Transactions on Asian and Low-Resource Language Information Processing, 19(6), Article 78, 1–21. https://doi.org/10.1145/3400396
Nacar, O., & Koubaa, A. (2024). Enhancing semantic similarity understanding in Arabic NLP with nested embedding learning. arXiv preprint, arXiv: 2407.21139. Doi:10.48550/arXiv.2407.21139.
O'Leary, M. (2010). Hakia gets serious with semantic search. Information Today, 27(6), pp 38-43.
Pastor, J. A., Martinez, F. J., & Rodriguez, J. V. (2009). Advantages of thesaurus representation using the Simple Knowledge Organization System (SKOS) compared with proposed alternatives. Information Research, 14(4). Retrieved from http://InformationR.net/ir/14-4/paper422.html
Pavlova, V. (2025). Multi-stage training of bilingual Islamic LLM for neural passage retrieval. In Proceedings of the New Horizons in Computational Linguistics for Religious Texts (pp. 42–52). Abu Dhabi, UAE: Association for Computational Linguistics.
Pavlova, V., & Makhlouf, M. (2024, November). Building an efficient multilingual non-profit IR system for the islamic domain leveraging multiprocessing design in rust. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 981-990).
Al-Smadi, M. (2025). QU-NLP at QIAS 2025 shared task: A two-phase LLM fine-tuning and retrieval-augmented generation approach for Islamic inheritance reasoning. In K. Darwish, A. Ali, I. Abu Farha, S. Touileb, I. Zitouni, A. Abdelali, S. Al-Ghamdi, S. Alkhereyf, W. Zaghouani, S. Khalifa, B. AlKhamissi, R. Almatham, I. Hamed, Z. Alyafeai, A. Alowisheq, G. Inoue, K. Mrini, & W. Alshammari (Eds.), Proceedings of the Third Arabic Natural Language Processing Conference: Shared Tasks (pp. 892–898). Association for Computational Linguistics. https://doi.org/ 10.18653/v1/2025.arabicnlp-sharedtasks.123.
Ahmad, M. A., Ballout, M., Ahmad, R. A., & Bruni, E. (2025). Transformer Tafsir at QIAS 2025 shared task: Hybrid retrieval-augmented generation for Islamic knowledge question answering. arXiv [Preprint]. https://doi.org/ 10.48550/arXiv.2509.23793
Tümer, D., Shah, M. A., & Bitirim, Y. (2009, May). An empirical evaluation on semantic search performance of keyword-based and semantic search engines: Google, yahoo, msn and hakia. In 2009 Fourth International Conference on Internet Monitoring and Protection (pp. 51-55). IEEE..