中文
相关论文

相关论文: Naver Labs Europe (SPLADE) @ TREC NeuCLIR 2022

200 篇论文

This paper describes our participation to the 2022 TREC Deep Learning challenge. We submitted runs to all four tasks, with a focus on the full retrieval passage task. The strategy is almost the same as 2021, with first stage retrieval being…

信息检索 · 计算机科学 2023-02-27 Carlos Lassance , Stéphane Clinchant

This report presents our participation to the WSDM Cup 2026 shared task on multilingual document retrieval from English queries. The task provides a challenging benchmark for cross-lingual generalization. It also provides a natural testbed…

信息检索 · 计算机科学 2026-02-25 Thibault Formal , Maxime Louis , Hervé Déjean , Stéphane Clinchant

This is the first year of the TREC Neural CLIR (NeuCLIR) track, which aims to study the impact of neural approaches to cross-language information retrieval. The main task in this year's track was ad hoc ranked retrieval of Chinese, Persian,…

信息检索 · 计算机科学 2023-09-26 Dawn Lawrie , Sean MacAvaney , James Mayfield , Paul McNamee , Douglas W. Oard , Luca Soldaini , Eugene Yang

This paper describes our participation in the 2023 WSDM CUP - MIRACL challenge. Via a combination of i) document translation; ii) multilingual SPLADE and Contriever; and iii) multilingual RankT5 and many other models, we were able to get…

信息检索 · 计算机科学 2023-03-01 Carlos Lassance

The HLTCOE team applied PLAID, an mT5 reranker, and document translation to the TREC 2023 NeuCLIR track. For PLAID we included a variety of models and training techniques -- the English model released with ColBERT v2, translate-train~(TT),…

计算与语言 · 计算机科学 2024-04-15 Eugene Yang , Dawn Lawrie , James Mayfield

The principal goal of the TREC Neural Cross-Language Information Retrieval (NeuCLIR) track is to study the impact of neural approaches to cross-language information retrieval. The track has created four collections, large collections of…

信息检索 · 计算机科学 2024-04-15 Dawn Lawrie , Sean MacAvaney , James Mayfield , Paul McNamee , Douglas W. Oard , Luca Soldaini , Eugene Yang

This paper presents NAVER LABS Europe's systems for Tamasheq-French and Quechua-Spanish speech translation in the IWSLT 2023 Low-Resource track. Our work attempts to maximize translation quality in low-resource settings using multilingual…

计算与语言 · 计算机科学 2023-06-14 Edward Gow-Smith , Alexandre Berard , Marcely Zanon Boito , Ioan Calapodescu

This paper reports on a study of cross-lingual information retrieval (CLIR) using the mT5-XXL reranker on the NeuCLIR track of TREC 2022. Perhaps the biggest contribution of this study is the finding that despite the mT5 model being…

信息检索 · 计算机科学 2023-03-29 Vitor Jeronymo , Roberto Lotufo , Rodrigo Nogueira

Retrieval systems generally focus on web-style queries that are short and underspecified. However, advances in language models have facilitated the nascent rise of retrieval models that can understand more complex queries with diverse…

This paper describes the systems that we submitted to the WMT19 Machine Translation robustness task. This task aims to improve MT's robustness to noise found on social media, like informal language, spelling mistakes and other orthographic…

计算与语言 · 计算机科学 2019-07-16 Alexandre Bérard , Ioan Calapodescu , Claude Roux

To measure advances in retrieval, test collections with relevance judgments that can faithfully distinguish systems are required. This paper presents NeuCLIRBench, an evaluation collection for cross-language and multilingual retrieval. The…

The principal goal of the TREC Neural Cross-Language Information Retrieval (NeuCLIR) track is to study the effect of neural approaches on cross-language information access. The track has created test collections containing Chinese, Persian,…

信息检索 · 计算机科学 2025-09-19 Dawn Lawrie , Sean MacAvaney , James Mayfield , Paul McNamee , Douglas W. Oard , Luca Soldaini , Eugene Yang

This paper describes our system for the SemEval2022 task of matching dictionary glosses to word embeddings. We focus on the Reverse Dictionary Track of the competition, which maps multilingual glosses to reconstructed vector…

计算与语言 · 计算机科学 2022-06-09 Zhiyong Wang , Ge Zhang , Nineli Lashkarashvili

The advent of multilingual language models has generated a resurgence of interest in cross-lingual information retrieval (CLIR), which is the task of searching documents in one language with queries from another. However, the rapid pace of…

The HLTCOE team applied PLAID, an mT5 reranker, GPT-4 reranker, score fusion, and document translation to the TREC 2024 NeuCLIR track. For PLAID we included a variety of models and training techniques -- Translate Distill (TD), Generate…

信息检索 · 计算机科学 2025-10-02 Eugene Yang , Dawn Lawrie , Orion Weller , James Mayfield

Measuring advances in retrieval requires test collections with relevance judgments that can faithfully distinguish systems. This paper presents NeuCLIRTech, an evaluation collection for cross-language retrieval over technical information.…

This paper describes our system, which placed third in the Multilingual Track (subtask 11), fourth in the Code-Mixed Track (subtask 12), and seventh in the Chinese Track (subtask 9) in the SemEval 2022 Task 11: MultiCoNER Multilingual…

计算与语言 · 计算机科学 2022-04-18 Weichao Gan , Yuanping Lin , Guangbo Yu , Guimin Chen , Qian Ye

This paper describes our approach for SemEval-2023 Task 3: Detecting the category, the framing, and the persuasion techniques in online news in a multi-lingual setup. For Subtask 1 (News Genre), we propose an ensemble of fully trained and…

计算与语言 · 计算机科学 2023-11-10 Ben Wu , Olesya Razuvayevskaya , Freddy Heppell , João A. Leite , Carolina Scarton , Kalina Bontcheva , Xingyi Song

Evaluation efforts such as TREC, CLEF, NTCIR and FIRE, alongside public leaderboard such as MS MARCO, are intended to encourage research and track our progress, addressing big questions in our field. However, the goal is not simply to…

信息检索 · 计算机科学 2021-05-11 Nick Craswell , Bhaskar Mitra , Emine Yilmaz , Daniel Campos , Jimmy Lin

Recently, neural models led to significant improvements in both machine translation (MT) and natural language generation tasks (NLG). However, generation of long descriptive summaries conditioned on structured data remains an open…

计算与语言 · 计算机科学 2019-11-01 Fahimeh Saleh , Alexandre Bérard , Ioan Calapodescu , Laurent Besacier
‹ 上一页 1 2 3 10 下一页 ›