English
Related papers

Related papers: Hakim: Farsi Text Embedding Model

200 papers

Traditional topic models often struggle with contextual nuances and fail to adequately handle polysemy and rare words. This limitation typically results in topics that lack coherence and quality. Large Language Models (LLMs) can mitigate…

Computation and Language · Computer Science 2025-05-13 Hajar Sakai , Sarah S. Lam

Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces Persian Speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Nima Sedghiyeh , Sara Sadeghi , Reza Khodadadi , Farzin Kashani , Omid Aghdaei , Somayeh Rahimi , Mohammad Sadegh Safari

We present a system that allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. We also release trained models for English, Arabic, German, French, Spanish, Russian, Turkish, and…

Computation and Language · Computer Science 2023-06-06 John Wieting , Kevin Gimpel , Graham Neubig , Taylor Berg-Kirkpatrick

The availability of different pre-trained semantic models enabled the quick development of machine learning components for downstream applications. Despite the availability of abundant text data for low resource languages, only a few…

Computation and Language · Computer Science 2022-02-24 Seid Muhie Yimam , Abinew Ali Ayele , Gopalakrishnan Venkatesh , Ibrahim Gashaw , Chris Biemann

Persian Poetry has consistently expressed its philosophy, wisdom, speech, and rationale on the basis of its couplets, making it an enigmatic language on its own to both native and non-native speakers. Nevertheless, the notice able gap…

Computation and Language · Computer Science 2022-09-22 Reza Khanmohammadi , Mitra Sadat Mirshafiee , Yazdan Rezaee Jouryabi , Seyed Abolghasem Mirroshandel

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. Second, it…

Computation and Language · Computer Science 2020-03-04 Muhammad Nabeel Asim , Muhammad Usman Ghani , Muhammad Ali Ibrahim , Sheraz Ahmad , Waqar Mahmood , Andreas Dengel

Reading comprehension, a fundamental cognitive ability essential for knowledge acquisition, is a complex skill, with a notable number of learners lacking proficiency in this domain. This study introduces innovative tasks for Brain-Computer…

Human-Computer Interaction · Computer Science 2024-01-30 Yuhong Zhang , Shilai Yang , Gert Cauwenberghs , Tzyy-Ping Jung

This paper introduces an innovative approach using Retrieval-Augmented Generation (RAG) pipelines with Large Language Models (LLMs) to enhance information retrieval and query response systems for university-related question answering. By…

Information Retrieval · Computer Science 2024-12-03 Arshia Hemmat , Kianoosh Vadaei , Mohammad Hassan Heydari , Afsaneh Fatemi

Recent advances in language models (LMs), have demonstrated significant efficacy in tasks related to the arts and humanities. While LMs have exhibited exceptional performance across a wide range of natural language processing tasks, there…

Computation and Language · Computer Science 2023-12-07 Amir Panahandeh , Hanie Asemi , Esmaeil Nourani

While most of the knowledge bases already support the English language, there is only one knowledge base for the Persian language, known as FarsBase, which is automatically created via semi-structured web information. Unlike English…

Computation and Language · Computer Science 2020-05-06 Majid Asgari-Bidhendi , Behrooz Janfada , Behrouz Minaei-Bidgoli

Despite rapid advances in large language models (LLMs), low-resource languages remain excluded from NLP, limiting digital access for millions. We present PunGPT2, the first fully open-source Punjabi generative model suite, trained on a 35GB…

Computation and Language · Computer Science 2025-10-06 Jaskaranjeet Singh , Rakesh Thakur

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

Computation and Language · Computer Science 2026-05-26 Isun Chehreh , Ebrahim Ansari

In modern e-commerce search systems, dense retrieval has become an indispensable component. By computing similarities between query and item (product) embeddings, it efficiently selects candidate products from large-scale repositories. With…

Information Retrieval · Computer Science 2025-10-20 Jianting Tang , Dongshuai Li , Tao Wen , Fuyu Lv , Dan Ou , Linli Xu

Semantic caching enhances the efficiency of large language model (LLM) systems by identifying semantically similar queries, storing responses once, and serving them for subsequent equivalent requests. However, existing semantic caching…

Machine Learning · Computer Science 2025-07-10 Shervin Ghaffari , Zohre Bahranifard , Mohammad Akbari

Link prediction with knowledge graph embedding (KGE) is a popular method for knowledge graph completion. Furthermore, training KGEs on non-English knowledge graph promote knowledge extraction and knowledge graph reasoning in the context of…

Artificial Intelligence · Computer Science 2023-03-28 Najmeh Torabian , Behrouz Minaei-Bidgoli , Mohsen Jahanshahi

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more comprehensive evaluation, we introduce the Massive…

Deep learning based models have dominated the current landscape of production recommender systems. Furthermore, recent years have witnessed an exponential growth of the model scale--from Google's 2016 model with 1 billion parameters to the…

The growth of conversational AI services has increased demand for effective information retrieval from dialogue data. However, existing methods often face challenges in capturing semantic intent or require extensive labeling and…

Information Retrieval · Computer Science 2025-03-07 Sangyeop Kim , Hangyeul Lee , Yohan Lee

In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes three key technical contributions. First, we propose ReMixer, a new data synthesis method that…

Information Retrieval · Computer Science 2026-04-21 Jianlyu Chen , Junwei Lan , Chaofan Li , Defu Lian , Zheng Liu

Nowadays, dialogue systems are used in many fields of industry and research. There are successful instances of these systems, such as Apple Siri, Google Assistant, and IBM Watson. Task-oriented dialogue system is a category of these, that…

Computation and Language · Computer Science 2024-01-02 Keyvan Mahmoudi , Heshaam Faili