English
Related papers

Related papers: The Czech Court Decisions Corpus (CzCDC): Availabi…

200 papers

In this paper, we introduce the citation data of the Czech apex courts (Supreme Court, Supreme Administrative Court and Constitutional Court). This dataset was automatically extracted from the corpus of texts of Czech court decisions -…

Computation and Language · Computer Science 2020-02-07 Jakub Harašta , Tereza Novotná , Jaromír Šavelka

We introduce the Cambridge Law Corpus (CLC), a dataset for legal AI research. It consists of over 250 000 court cases from the UK. Most cases are from the 21st century, but the corpus includes cases as old as the 16th century. This paper…

Computation and Language · Computer Science 2024-01-03 Andreas Östling , Holli Sargeant , Huiyuan Xie , Ludwig Bull , Alexander Terenin , Leif Jonsson , Måns Magnusson , Felix Steffek

We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam$.$cz. To the best of our knowledge, CWRCzech is…

Information Retrieval · Computer Science 2024-07-16 Josef Vonášek , Milan Straka , Rostislav Krč , Lenka Lasoňová , Ekaterina Egorova , Jana Straková , Jakub Náplava

We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Error Correction…

Computation and Language · Computer Science 2022-04-22 Jakub Náplava , Milan Straka , Jana Straková , Alexandr Rosen

As stakeholders' pressure on corporates for disclosing their corporate social responsibility operations grows, it is crucial to understand how efficient corporate disclosure systems are in bridging the gap between corporate social…

General Economics · Economics 2023-01-10 Xhesilda Vogli , Erion Çano

Indian Judiciary is suffering from burden of millions of cases that are lying pending in its courts at all the levels. In this paper, we analyze the data that we have collected on the pendency of 24 high courts in the Republic of India as…

Computers and Society · Computer Science 2023-07-21 Kshitiz Verma

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available…

Computation and Language · Computer Science 2018-02-01 Pavel Král , Ladislav Lenc

Courts must justify their decisions, but systematically analyzing judicial reasoning at scale remains difficult. This study tests claims about formalistic judging in Central and Eastern Europe (CEE) by developing automated methods to detect…

Computation and Language · Computer Science 2026-03-24 Tomáš Koref , Lena Held , Mahammad Namazov , Harun Kumru , Yassine Thlija , Ivan Habernal

With the rapid increase of published open datasets, it is crucial to support the open data progress in smart cities while considering the open data quality. In the Czech Republic, and its National Open Data Catalogue (NODC), the open…

Databases · Computer Science 2023-03-06 Dasa Kusnirakova , Mouzhi Ge , Leonard Walletzky , Barbora Buhnova

We present a Chinese judicial reading comprehension (CJRC) dataset which contains approximately 10K documents and almost 50K questions with answers. The documents come from judgment documents and the questions are annotated by law experts.…

Computation and Language · Computer Science 2019-12-21 Xingyi Duan , Baoxin Wang , Ziyue Wang , Wentao Ma , Yiming Cui , Dayong Wu , Shijin Wang , Ting Liu , Tianxiang Huo , Zhen Hu , Heng Wang , Zhiyuan Liu

We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the sound recordings of the Czech parliamentary speeches with…

Computation and Language · Computer Science 2025-09-09 Vladislav Stankov , Matyáš Kopp , Ondřej Bojar

An automated system that could assist a judge in predicting the outcome of a case would help expedite the judicial process. For such a system to be practically useful, predictions by the system should be explainable. To promote research in…

Computation and Language · Computer Science 2021-06-01 Vijit Malik , Rishabh Sanjay , Shubham Kumar Nigam , Kripa Ghosh , Shouvik Kumar Guha , Arnab Bhattacharya , Ashutosh Modi

Publication databases rely on accurate metadata extraction from diverse web sources, yet variations in web layouts and data formats present challenges for metadata providers. This paper introduces CRAWLDoc, a new method for contextual…

Computation and Language · Computer Science 2025-06-05 Fabian Karl , Ansgar Scherp

Standardization of data items collected in paediatric clinical trials is an important but challenging issue. The Clinical Data Interchange Standards Consortium (CDISC) data standards are well understood by the pharmaceutical industry but…

This short paper presents a compact overview of the Czech approach to implementing the European Open Science Cloud and plans for developing a Czech national infrastructure for FAIR research data. Its purpose is to provide an…

Digital Libraries · Computer Science 2024-02-22 Matej Antol , Jiri Marek , Michaela Capandova , Jaroslav Juracek , Ludek Matyska

Retrieving case law is a time-consuming task predominantly carried out by querying databases. We provide a comparison of two models in three different settings for Czech Constitutional Court decisions: (i) a large general-purpose embedder…

Computation and Language · Computer Science 2025-12-08 Tereza Novotna , Jakub Harasta

In clinical care, obtaining a correct diagnosis is the first step towards successful treatment and, ultimately, recovery. Depending on the complexity of the case, the diagnostic phase can be lengthy and ridden with errors and delays. Such…

Information Retrieval · Computer Science 2019-08-26 Carsten Eickhoff , Floran Gmehlin , Anu V. Patel , Jocelyn Boullier , Hamish Fraser

Big data in healthcare has made a positive difference in advancing analytical capabilities and lowering the costs of medical care. In addition to providing analytical capabilities on platforms supporting current and near-future AI with…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-07 Martin Štufi , Boris Bačić , Leonid Stoimenov

The availability of structured legal data is important for advancing Natural Language Processing (NLP) techniques for the German legal system. One of the most widely used datasets, Open Legal Data, provides a large-scale collection of…

Computation and Language · Computer Science 2026-01-06 Harshil Darji , Martin Heckelmann , Christina Kratsch , Gerard de Melo

Access to legal information is fundamental to access to justice. Yet accessibility refers not only to making legal documents available to the public, but also rendering legal information comprehensible to them. A vexing problem in bringing…

Computation and Language · Computer Science 2025-05-08 Mingruo Yuan , Ben Kao , Tien-Hsuan Wu , Michael M. K. Cheung , Henry W. H. Chan , Anne S. Y. Cheung , Felix W. H. Chan , Yongxi Chen
‹ Prev 1 2 3 10 Next ›