English
Related papers

Related papers: DuReader_retrieval: A Large-scale Chinese Benchmar…

200 papers

We study multi-answer retrieval, an under-explored problem that requires retrieving passages to cover multiple distinct answers for a given question. This task requires joint modeling of retrieved passages, as models should not repeatedly…

Computation and Language · Computer Science 2021-09-21 Sewon Min , Kenton Lee , Ming-Wei Chang , Kristina Toutanova , Hannaneh Hajishirzi

In the realm of e-commerce search, the significance of semantic matching cannot be overstated, as it directly impacts both user experience and company revenue. Along this line, query rewriting, serving as an important technique to bridge…

Information Retrieval · Computer Science 2024-03-05 Wenjun Peng , Guiyang Li , Yue Jiang , Zilong Wang , Dan Ou , Xiaoyi Zeng , Derong Xu , Tong Xu , Enhong Chen

Traditional legal retrieval systems designed to retrieve legal documents, statutes, precedents, and other legal information are unable to give satisfactory answers due to lack of semantic understanding of specific questions. Large Language…

Computation and Language · Computer Science 2024-08-02 Nan Xie , Yuelin Bai , Hengyuan Gao , Feiteng Fang , Qixuan Zhao , Zhijian Li , Ziqiang Xue , Liang Zhu , Shiwen Ni , Min Yang

Effective disaster management requires timely access to accurate and contextually relevant information. Existing Information Retrieval (IR) benchmarks, however, focus primarily on general or specialized domains, such as medicine or finance,…

Information Retrieval · Computer Science 2025-09-23 Kai Yin , Xiangjue Dong , Chengkai Liu , Lipai Huang , Yiming Xiao , Zhewei Liu , Ali Mostafavi , James Caverlee

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

This paper contributes a new large-scale dataset for weakly supervised cross-media retrieval, named Twitter100k. Current datasets, such as Wikipedia, NUS Wide and Flickr30k, have two major limitations. First, these datasets are lacking in…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yuting Hu , Liang Zheng , Yi Yang , Yongfeng Huang

Most existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introduce a new…

Computer Vision and Pattern Recognition · Computer Science 2020-02-14 Yipeng Sun , Jiaming Liu , Wei Liu , Junyu Han , Errui Ding , Jingtuo Liu

Scene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Jianqi Ma , Zhetong Liang , Wangmeng Xiang , Xi Yang , Lei Zhang

An approximate textual retrieval algorithm for searching sources with high levels of defects is presented. It considers splitting the words in a query into two overlapping segments and subsequently building composite regular expressions…

Information Retrieval · Computer Science 2007-05-23 Pere Constans

In this work we focus on multi-turn passage retrieval as a crucial component of conversational search. One of the key challenges in multi-turn passage retrieval comes from the fact that the current turn query is often underspecified due to…

Information Retrieval · Computer Science 2020-05-26 Nikos Voskarides , Dan Li , Pengjie Ren , Evangelos Kanoulas , Maarten de Rijke

The current communication presents a simple exercise with the aim of solving a singular problem: the retrieval of extremely large amounts of items in the Web of Science interface. As it is known, Web of Science interface allows a user to…

Digital Libraries · Computer Science 2009-11-09 Ricardo Arencibia-Jorge , Loet Leydesdorff , Zaida Chinchilla-Rodriguez , Ronald Rousseau , Soren W. Paris

Searching for mathematical results remains difficult: most existing tools retrieve entire papers, while mathematicians and theorem-proving agents often seek a specific theorem, lemma, or proposition that answers a query. While semantic…

Information Retrieval · Computer Science 2026-03-10 Luke Alexander , Eric Leonen , Sophie Szeto , Artemii Remizov , Ignacio Tejeda , Jarod Alper , Giovanni Inchiostro , Vasily Ilin

Text correction, especially the semantic correction of more widely used scenes, is strongly required to improve, for the fluency and writing efficiency of the text. An adversarial multi-task learning method is proposed to enhance the…

Computation and Language · Computer Science 2023-06-29 Fanyu Wang , Zhenping Xie

This study presents the first comprehensive safety evaluation of the DeepSeek models, focusing on evaluating the safety risks associated with their generated content. Our evaluation encompasses DeepSeek's latest generation of large language…

Cryptography and Security · Computer Science 2025-03-20 Zonghao Ying , Guangyi Zheng , Yongxin Huang , Deyue Zhang , Wenxin Zhang , Quanchen Zou , Aishan Liu , Xianglong Liu , Dacheng Tao

Different from the traditional translation tasks, classical Chinese poetry translation requires both adequacy and fluency in translating culturally and historically significant content and linguistic poetic elegance. Large language models…

Computation and Language · Computer Science 2024-12-31 Andong Chen , Lianzhang Lou , Kehai Chen , Xuefeng Bai , Yang Xiang , Muyun Yang , Tiejun Zhao , Min Zhang

Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However,…

Computation and Language · Computer Science 2025-08-14 Kangwei Liu , Siyuan Cheng , Bozhong Tian , Xiaozhuan Liang , Yuyang Yin , Meng Han , Ningyu Zhang , Bryan Hooi , Xi Chen , Shumin Deng

This paper presents BiPaR, a bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support multilingual and cross-lingual reading comprehension. The biggest difference between BiPaR and existing reading…

Computation and Language · Computer Science 2019-10-14 Yimin Jing , Deyi Xiong , Yan Zhen

Retrieval-Augmented Generation (RAG) has emerged as a promising technology for legal document consultation, yet its application in Chinese legal scenarios faces two key limitations: existing benchmarks lack specialized support for joint…

Computation and Language · Computer Science 2026-03-13 Yaocong Li , Qiang Lan , Leihan Zhang , Le Zhang

A reverse dictionary takes the description of a target word as input and outputs the target word together with other words that match the description. Existing reverse dictionary methods cannot deal with highly variable input queries and…

Computation and Language · Computer Science 2019-12-20 Lei Zhang , Fanchao Qi , Zhiyuan Liu , Yasheng Wang , Qun Liu , Maosong Sun

This paper introduces a cross-lingual statutory article retrieval (SAR) dataset designed to enhance legal information retrieval in multilingual settings. Our dataset features spoken-language-style legal inquiries in English, paired with…

Computation and Language · Computer Science 2024-10-16 Yen-Hsiang Wang , Feng-Dian Su , Tzu-Yu Yeh , Yao-Chung Fan
‹ Prev 1 4 5 6 7 8 10 Next ›