English
Related papers

Related papers: Citation Data of Czech Apex Courts

200 papers

We present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0 (PDT-C 1.0), the purpose of which is - as it always been the case for the family of the Prague Dependency Treebanks - to…

Computation and Language · Computer Science 2020-06-09 Jan Hajič , Eduard Bejček , Jaroslava Hlaváčová , Marie Mikulová , Milan Straka , Jan Štěpánek , Barbora Štěpánková

Text summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and other high-resource…

Computation and Language · Computer Science 2025-08-15 Václav Tran , Jakub Šmíd , Jiří Martínek , Ladislav Lenc , Pavel Král

The availability of structured legal data is important for advancing Natural Language Processing (NLP) techniques for the German legal system. One of the most widely used datasets, Open Legal Data, provides a large-scale collection of…

Computation and Language · Computer Science 2026-01-06 Harshil Darji , Martin Heckelmann , Christina Kratsch , Gerard de Melo

Artificial intelligence is being utilized in many domains as of late, and the legal system is no exception. However, as it stands now, the number of well-annotated datasets pertaining to legal documents from the Supreme Court of the United…

Computation and Language · Computer Science 2021-12-08 Mohammad Alali , Shaayan Syed , Mohammed Alsayed , Smit Patel , Hemanth Bodala

In this paper, we present a method of automatic catchphrase extracting from legal case documents. We utilize deep neural networks for constructing scoring model of our extraction system. We achieve comparable performance with systems using…

Computation and Language · Computer Science 2018-09-17 Vu Tran , Minh Le Nguyen , Ken Satoh

We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can…

Computation and Language · Computer Science 2007-05-23 Anoop Sarkar , Daniel Zeman

As a pivotal task in natural language processing, element extraction has gained significance in the legal domain. Extracting legal elements from judicial documents helps enhance interpretative and analytical capacities of legal cases, and…

Computation and Language · Computer Science 2023-10-11 Xue Zongyue , Liu Huanghai , Hu Yiran , Kong Kangle , Wang Chenlu , Liu Yun , Shen Weixing

In this paper we present a new paradigm for the identification of datasets extracted from the Virtual Atomic and Molecular Data Centre (VAMDC) e-science infrastructure. Such identification includes information on the origin and version of…

Digital Libraries · Computer Science 2016-06-02 Carlo Maria Zwölf , Nicolas Moreau , Marie-Lise Dubernet

Citation analysis is one of the most frequently used methods in research evaluation. We are seeing significant growth in citation analysis through bibliometric metadata, primarily due to the availability of citation databases such as the…

Digital Libraries · Computer Science 2020-09-01 Sehrish Iqbal , Saeed-Ul Hassan , Naif Radi Aljohani , Salem Alelyani , Raheel Nawaz , Lutz Bornmann

In legal document writing, one of the key elements is properly citing the case laws and other sources to substantiate claims and arguments. Understanding the legal domain and identifying appropriate citation context or cite-worthy sentences…

Computation and Language · Computer Science 2023-05-08 Mann Khatri , Pritish Wadhwa , Gitansh Satija , Reshma Sheik , Yaman Kumar , Rajiv Ratn Shah , Ponnurangam Kumaraguru

Wikipedia is an essential component of the open science ecosystem, yet it is poorly integrated with academic open science initiatives. Wikipedia Citations is a project that focuses on extracting and releasing comprehensive datasets of…

Digital Libraries · Computer Science 2024-06-28 Natallia Kokash , Giovanni Colavizza

The escalating number of pending cases is a growing concern world-wide. Recent advancements in digitization have opened up possibilities for leveraging artificial intelligence (AI) tools in the processing of legal documents. Adopting a…

Information Retrieval · Computer Science 2023-10-19 Subinay Adhikary , Sagnik Das , Sagnik Saha , Procheta Sen , Dwaipayan Roy , Kripabandhu Ghosh

We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper…

Computation and Language · Computer Science 2024-01-03 Andreas Östling , Holli Sargeant , Huiyuan Xie , Ludwig Bull , Alexander Terenin , Leif Jonsson , Måns Magnusson , Felix Steffek

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

Computation and Language · Computer Science 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

The paper presents first results of the CitEcCyr project funded by RANEPA. The project aims to create a source of open citation data for research papers written in Russian. Compared to existing sources of citation data, CitEcCyr is working…

Digital Libraries · Computer Science 2017-10-03 Jose Manuel Barrueco , Thomas Krichel , Sergey Parinov , Victor Lyapunov , Oxana Medvedeva , Varvara Sergeeva

Legal practitioners and judicial institutions face an ever-growing volume of case-law documents characterised by formalised language, lengthy sentence structures, and highly specialised terminology, making manual triage both time-consuming…

Computation and Language · Computer Science 2026-04-21 Moinul Hossain , Sourav Rabi Das , Zikrul Shariar Ayon , Sadia Afrin Promi , Ahnaf Atef Choudhury , Shakila Rahman , Jia Uddin

Sentence-by-sentence information extraction from long documents is an exhausting and error-prone task. As the indicator of document skeleton, catalogs naturally chunk documents into segments and provide informative cascade semantics, which…

Computation and Language · Computer Science 2023-05-01 Tong Zhu , Guoliang Zhang , Zechang Li , Zijian Yu , Junfei Ren , Mengsong Wu , Zhefeng Wang , Baoxing Huai , Pingfu Chao , Wenliang Chen

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

Computation and Language · Computer Science 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

The significance and influence of US Supreme Court majority opinions derive in large part from opinions' roles as precedents for future opinions. A growing body of literature seeks to understand what drives the use of opinions as precedents…

Applications · Statistics 2021-01-19 Christian S. Schmid , Ted Hsuan Yun Chen , Bruce A. Desmarais

Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved.…

Computation and Language · Computer Science 2019-11-18 Guillaume Wenzek , Marie-Anne Lachaux , Alexis Conneau , Vishrav Chaudhary , Francisco Guzmán , Armand Joulin , Edouard Grave