English
Related papers

Related papers: A Statutory Article Retrieval Dataset in French

200 papers

We propose Bijective Universal Scene-Specific Anomalous Relationship Detection (BUSSARD), a normalizing flow-based model for detecting anomalous relations in scene graphs, generated from images. Our work follows a multimodal approach,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Melissa Schween , Mathis Kruse , Bodo Rosenhahn

This paper introduces Fetch-A-Set (FAS), a comprehensive benchmark tailored for legislative historical document analysis systems, addressing the challenges of large-scale document retrieval in historical contexts. The benchmark comprises a…

Information Retrieval · Computer Science 2024-06-18 Adrià Molina , Oriol Ramos Terrades , Josep Lladós

Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone task, it has received little attention. We address this by…

Computation and Language · Computer Science 2026-04-28 Michelle Wastl , Jannis Vamvas , Rico Sennrich

We present preliminary results about Legistix, a tool we are developing to automatically consolidate the French and European law. Legistix is based both on regular expressions used in several compound grammars, similar to the successive…

Computation and Language · Computer Science 2023-01-18 Georges-André Silber

Generative retrieval (Wang et al., 2022; Tay et al., 2022) is a popular approach for end-to-end document retrieval that directly generates document identifiers given an input query. We introduce summarization-based document IDs, in which…

Computation and Language · Computer Science 2024-10-31 Haoxin Li , Daniel Cheng , Phillip Keung , Jungo Kasai , Noah A. Smith

Contextual spelling correction models are an alternative to shallow fusion to improve automatic speech recognition (ASR) quality given user vocabulary. To deal with large user vocabularies, most of these models include candidate retrieval…

Computation and Language · Computer Science 2023-06-06 Alexandra Antonova , Evelina Bakhturina , Boris Ginsburg

Search engine has become a fundamental component in various web and mobile applications. Retrieving relevant documents from the massive datasets is challenging for a search engine system, especially when faced with verbose or tail queries.…

Information Retrieval · Computer Science 2020-08-11 Kuan Fang , Long Zhao , Zhan Shen , RuiXing Wang , RiKang Zhour , LiWen Fan

Regulatory texts are inherently long and complex, presenting significant challenges for information retrieval systems in supporting regulatory officers with compliance tasks. This paper introduces a hybrid information retrieval system that…

Computation and Language · Computer Science 2025-02-25 Jhon Rayo , Raul de la Rosa , Mario Garrido

In legal eDiscovery, the parties are required to search through their electronically stored information to find documents that are relevant to a specific case. Negotiations over the scope of these searches are often based on a fear that…

Information Retrieval · Computer Science 2020-08-03 Herbert L. Roitblat

Retrieval-Augmented Generation (RAG) systems are showing promising potential, and are becoming increasingly relevant in AI-powered legal applications. Existing benchmarks, such as LegalBench, assess the generative capabilities of Large…

Artificial Intelligence · Computer Science 2024-08-21 Nicholas Pipitone , Ghita Houir Alami

Legal case retrieval, which aims to retrieve relevant cases to a given query case, benefits judgment justice and attracts increasing attention. Unlike generic retrieval queries, legal case queries are typically long and the definition of…

Information Retrieval · Computer Science 2023-12-07 Youchao Zhou , Heyan Huang , Zhijing Wu

Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevance is dominated by aspects of similarity in texts, leaving unobserved whether the text is…

Information Retrieval · Computer Science 2026-04-24 Tobias Schimanski , Stefanie Lewandowski , Christian Woerle , Nicola Reichenau , Yauheni Huryn , Markus Leippold

Today, full-texts of scientific articles are often stored in different locations than the used datasets. Dataset registries aim at a closer integration by making datasets citable but authors typically refer to datasets using inconsistent…

Digital Libraries · Computer Science 2016-11-08 Behnam Ghavimi , Philipp Mayr , Christoph Lange , Sahar Vahdati , Sören AUER

Large Language Models (LLMs) have demonstrated strong potential across legal tasks, yet the problem of legal citation prediction remains under-explored. At its core, this task demands fine-grained contextual understanding and precise…

Computation and Language · Computer Science 2025-05-23 Jiuzhou Han , Paul Burgess , Ehsan Shareghi

This technical report outlines the filtering approach applied to the collection of the Bielefeld Academic Search Engine (BASE) data to extract articles from the political science domain. We combined hard and soft filters to address entries…

Digital Libraries · Computer Science 2024-06-25 Nina Smirnova

Though exponentially growing health-related literature has been made available to a broad audience online, the language of scientific articles can be difficult for the general public to understand. Therefore, adapting this expert-level…

Computation and Language · Computer Science 2022-10-25 Kush Attal , Brian Ondov , Dina Demner-Fushman

Existing summarization systems mostly generate summaries purely relying on the content of the source document. However, even for humans, we usually need some references or exemplars to help us fully understand the source document and write…

Computation and Language · Computer Science 2021-12-14 Chenxin An , Ming Zhong , Zhichao Geng , Jianqiang Yang , Xipeng Qiu

We consider the problem of creating document representations in which inter-document similarity measurements correspond to semantic similarity. We first present a novel subspace-based framework for formalizing this task. Using this…

Computation and Language · Computer Science 2007-05-23 Rie Kubota Ando , Lillian Lee

Personalized content recommendations have been pivotal to the content experience in digital media from video streaming to social networks. However, several domain specific challenges have held back adoption of recommender systems in news…

We study the problem of retrieval with instructions, where users of a retrieval system explicitly describe their intent along with their queries. We aim to develop a general-purpose task-aware retrieval system using multi-task instruction…

Computation and Language · Computer Science 2022-12-21 Akari Asai , Timo Schick , Patrick Lewis , Xilun Chen , Gautier Izacard , Sebastian Riedel , Hannaneh Hajishirzi , Wen-tau Yih
‹ Prev 1 4 5 6 7 8 10 Next ›