中文
相关论文

相关论文: Hakim: Farsi Text Embedding Model

200 篇论文

Effectively adapting powerful pretrained foundation models to diverse tasks remains a key challenge in AI deployment. Current approaches primarily follow two paradigms:discrete optimization of text prompts through prompt engineering, or…

计算与语言 · 计算机科学 2025-08-06 Xiaoming Hou , Jiquan Zhang , Zibin Lin , DaCheng Tao , Shengli Zhang

As a digraphic language, the Persian language utilizes two written standards: Perso-Arabic in Afghanistan and Iran, and Tajik-Cyrillic in Tajikistan. Despite the significant similarity between the dialects of each country, script…

计算与语言 · 计算机科学 2025-10-10 Rayyan Merchant , Kevin Tang

App usage prediction is important for smartphone system optimization to enhance user experience. Existing modeling approaches utilize historical app usage logs along with a wide range of semantic information to predict the app usage;…

机器学习 · 计算机科学 2021-08-27 Yonchanok Khaokaew , Mohammad Saiedur Rahaman , Ryen W. White , Flora D. Salim

This paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplicate text retrieval, clustering, and dataset deduplication…

计算与语言 · 计算机科学 2023-11-30 Marina Zhang , Owen Vallis , Aysegul Bumin , Tanay Vakharia , Elie Bursztein

Embedding models play a crucial role in representing and retrieving information across various NLP applications. Recent advances in large language models (LLMs) have further enhanced the performance of embedding models. While these models…

计算与语言 · 计算机科学 2025-09-15 Yixuan Tang , Yi Yang

This research assesses the effectiveness of state-of-the-art large language models (LLMs), including ChatGPT, Llama, Aya, Jais, and ACEGPT, in the task of Arabic automated essay scoring (AES) using the AR-AES dataset. It explores various…

计算与语言 · 计算机科学 2025-01-29 Rayed Ghazawi , Edwin Simpson

Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However,…

信息检索 · 计算机科学 2024-10-15 Weiwei Sun , Zhengliang Shi , Jiulong Wu , Lingyong Yan , Xinyu Ma , Yiding Liu , Min Cao , Dawei Yin , Zhaochun Ren

Large Language Models (LLMs) have shown remarkable capabilities, not only in generating human-like text, but also in acquiring knowledge. This highlights the need to go beyond the typical Natural Language Processing downstream benchmarks…

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creation and validation of…

计算与语言 · 计算机科学 2026-05-05 Mullosharaf K. Arabov

Smart cities need the involvement of their residents to enhance quality of life. Conversational query-answering is an emerging approach for user engagement. There is an increasing demand of an advanced conversational question-answering that…

计算与语言 · 计算机科学 2024-04-16 Pardis Moradbeiki , Nasser Ghadiri

Finding the appropriate words to convey concepts (i.e., lexical access) is essential for effective communication. Reverse dictionaries fulfill this need by helping individuals to find the word(s) which could relate to a specific concept or…

计算与语言 · 计算机科学 2021-11-02 Arman Malekzadeh , Amin Gheibi , Ali Mohades

Over the past years, Automated Essay Scoring (AES) systems have gained increasing attention as scalable and consistent solutions for assessing the proficiency of student writing. Despite recent progress, support for Arabic AES remains…

计算与语言 · 计算机科学 2026-05-20 Hoor Elbahnasawi , Marwan Sayed , Sohaila Eltanbouly , Fatima Brahamia , Tamer Elsayed

In the recent decade, with the enormous growth of digital content in internet and databases, sentiment analysis has received more and more attention between information retrieval and natural language processing researchers. Sentiment…

计算与语言 · 计算机科学 2014-12-30 Ayoub Bagheri , Mohamad Saraee

Relation extraction is the task of extracting semantic relations between entities in a sentence. It is an essential part of some natural language processing tasks such as information extraction, knowledge extraction, and knowledge base…

计算与语言 · 计算机科学 2020-05-15 Majid Asgari-Bidhendi , Mehrdad Nasser , Behrooz Janfada , Behrouz Minaei-Bidgoli

Effective feature representations play a critical role in enhancing the performance of text generation models that rely on deep neural networks. However, current approaches suffer from several drawbacks, such as the inability to capture the…

计算与语言 · 计算机科学 2024-02-27 Omama Hamad , Ali Hamdi , Khaled Shaban

New models for natural language understanding have recently made an unparalleled amount of progress, which has led some researchers to suggest that the models induce universal text representations. However, current benchmarks are…

计算与语言 · 计算机科学 2022-04-05 Damien Sileo , Tim Van-de-Cruys , Camille Pradel , Philippe Muller

Social media hold valuable, vast and unstructured information on public opinion that can be utilized to improve products and services. The automatic analysis of such data, however, requires a deep understanding of natural language. Current…

计算与语言 · 计算机科学 2019-10-01 Kia Dashtipour , Mandar Gogate , Jingpeng Li , Fengling Jiang , Bin Kong , Amir Hussain

While emerging Persian NLP benchmarks have expanded into pragmatics and politeness, they rarely distinguish between memorized cultural facts and the ability to reason about implicit social norms. We introduce DivanBench, a diagnostic…

计算与语言 · 计算机科学 2026-02-20 Alireza Sakhaeirad , Ali Ma'manpoosh , Arshia Hemmat

This paper focuses on how to extract opinions over each Persian sentence-level text. Deep learning models provided a new way to boost the quality of the output. However, these architectures need to feed on big annotated data as well as an…