English
Related papers

Related papers: Citation Data of Czech Apex Courts

200 papers

Arguments, counter-arguments, facts, and evidence obtained via documents related to previous court cases are of essential need for legal professionals. Therefore, the process of automatic information extraction from documents containing…

Computation and Language · Computer Science 2019-08-20 Gathika Ratnayaka , Thejan Rupasinghe , Nisansa de Silva , Viraj Salaka Gamage , Menuka Warushavithana , Amal Shehan Perera

This paper uses Critical Discourse Analysis (CDA) to show how Sino-judicial activism shapes Data Intellectual Property Rights (DIPR) in China. We identify two complementary judicial discourses. Local courts (exemplified by the Zhejiang High…

Computers and Society · Computer Science 2025-09-24 Chanhou Lou

This short paper presents a compact overview of the Czech approach to implementing the European Open Science Cloud and plans for developing a Czech national infrastructure for FAIR research data. Its purpose is to provide an…

Digital Libraries · Computer Science 2024-02-22 Matej Antol , Jiri Marek , Michaela Capandova , Jaroslav Juracek , Ludek Matyska

Legal citation in common-law systems depends not only on factual similarity, but also on the legal principle for which a precedent is invoked. However, existing benchmarks for legal citation retrieval use case facts, citation context, or…

Information Retrieval · Computer Science 2026-05-21 Shannon Lee Yueh Ern , Kaidong Feng , Yingpeng Du , Chloe Lee En Jia , Zhu Sun

In this paper, we introduce a novel Czech dataset for aspect-based sentiment analysis (ABSA), which consists of 3.1K manually annotated reviews from the restaurant domain. The dataset is built upon the older Czech dataset, which contained…

Computation and Language · Computer Science 2025-08-12 Jakub Šmíd , Pavel Přibáň , Ondřej Pražák , Pavel Král

This paper presents thirteen datasets for binary, multiclass and multilabel classification based on the European Court of Human Rights judgments since its creation. The interest of such datasets is explained through the prism of the…

Machine Learning · Computer Science 2019-02-05 Alexandre Quemy

Despite growing interest, accurately and reliably representing unstructured data, such as court decisions, in a structured form, remains a challenge. Recent advancements in generative AI applied to language modeling enabled the…

Information Retrieval · Computer Science 2025-01-27 Alexander V. Belikov , Sacha Raoult

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

Computation and Language · Computer Science 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

This article focuses on the development and evaluation of Small-sized Czech sentence embedding models. Small models are important components for real-time industry applications in resource-constrained environments. Given the limited…

Computation and Language · Computer Science 2023-11-27 Jiří Bednář , Jakub Náplava , Petra Barančíková , Ondřej Lisický

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often not generalizable, based on synthetic data, or not publicly…

Digital Libraries · Computer Science 2026-03-27 Parth Sarin , Juan Pablo Alperin , Adam Buttrick , Dione Mentis

Grammatical error correction in English is a long studied problem with many existing systems and datasets. However, there has been only a limited research on error correction of other languages. In this paper, we present a new dataset…

Computation and Language · Computer Science 2019-10-17 Jakub Náplava , Milan Straka

Court transcripts and judgments are rich repositories of legal knowledge, detailing the intricacies of cases and the rationale behind judicial decisions. The extraction of key information from these documents provides a concise overview of…

Computation and Language · Computer Science 2024-03-20 Joana Ribeiro de Faria , Huiyuan Xie , Felix Steffek

This paper tackles the task of legal extractive summarization using a dataset of 430K U.S. court opinions with key passages annotated. According to automated summary quality metrics, the reinforcement-learning-based MemSum model is best and…

Computation and Language · Computer Science 2023-05-16 Emmanuel Bauer , Dominik Stammbach , Nianlong Gu , Elliott Ash

Open data promotes transparency and accountability as everyone can analyse it. Law enforcement and the judiciary are increasingly making data available, to increase trust and confidence in the criminal justice system. Due to privacy…

Computers and Society · Computer Science 2021-04-12 Pieter Hartel , Rolf van Wegberg , Mark van Staalduinen

In common law jurisdictions, legal practitioners rely on precedents to construct arguments, in line with the doctrine of \emph{stare decisis}. As the number of cases grow over the years, prior case retrieval (PCR) has garnered significant…

Computation and Language · Computer Science 2024-04-02 T. Y. S. S Santosh , Rashid Gustav Haddad , Matthias Grabmair

Large language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination. In this work, our aim is to allow LLMs to generate text with citations, improving their factual…

Computation and Language · Computer Science 2023-11-01 Tianyu Gao , Howard Yen , Jiatong Yu , Danqi Chen

Retrieving case law is a time-consuming task predominantly carried out by querying databases. We provide a comparison of two models in three different settings for Czech Constitutional Court decisions: (i) a large general-purpose embedder…

Computation and Language · Computer Science 2025-12-08 Tereza Novotna , Jakub Harasta

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from…

Computation and Language · Computer Science 2016-03-23 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

A typical information extraction pipeline consists of token- or span-level classification models coupled with a series of pre- and post-processing scripts. In a production pipeline, requirements often change, with classes being added and…

Artificial Intelligence · Computer Science 2022-01-19 Ramon Pires , Fábio C. de Souza , Guilherme Rosa , Roberto A. Lotufo , Rodrigo Nogueira

As legal case law databases such as HUDOC continue to grow rapidly, it has become essential for legal researchers to find efficient methods to handle such large-scale data sets. Such case law databases usually consist of the textual content…

Information Retrieval · Computer Science 2024-01-31 M. Mohammadi , L. M. Bruijn , M. Wieling , M. Vols