English
Related papers

Related papers: HLDC: Hindi Legal Documents Corpus

200 papers

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

Computation and Language · Computer Science 2026-05-18 Nuwan I. Senaratna

The widespread adoption of Large Language Models (LLMs) and awareness around multilingual LLMs have raised concerns regarding the potential risks and repercussions linked to the misapplication of AI-generated text, necessitating increased…

Computation and Language · Computer Science 2024-10-08 Ishan Kavathekar , Anku Rani , Ashmit Chamoli , Ponnurangam Kumaraguru , Amit Sheth , Amitava Das

Leveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Chaohu Liu , Kun Yin , Haoyu Cao , Xinghua Jiang , Xin Li , Yinsong Liu , Deqiang Jiang , Xing Sun , Linli Xu

The development of robust language models for low-resource languages is frequently bottlenecked by the scarcity of high-quality, coherent, and domain-appropriate training corpora. In this paper, we introduce the Multilingual TinyStories…

Computation and Language · Computer Science 2026-03-17 Deepon Halder , Angira Mukherjee

Current advancements in Natural Language Processing (NLP) have largely favored resource-rich languages, leaving a significant gap in high-quality datasets for low-resource languages like Hindi. This scarcity is particularly evident in text…

Computation and Language · Computer Science 2026-01-06 Praveenkumar Katwe , RakeshChandra Balabantaray , Kaliprasad Vittala

While large language models (LLMs) have showcased impressive capabilities, they struggle with addressing legal queries due to the intricate complexities and specialized expertise required in the legal field. In this paper, we introduce…

Computation and Language · Computer Science 2024-06-24 Zhiwei Fei , Songyang Zhang , Xiaoyu Shen , Dawei Zhu , Xiao Wang , Maosong Cao , Fengzhe Zhou , Yining Li , Wenwei Zhang , Dahua Lin , Kai Chen , Jidong Ge

Introduction: Clinical text classification using natural language processing (NLP) models requires adequate training data to achieve optimal performance. For that, 200-500 documents are typically annotated. The number is constrained by time…

Computation and Language · Computer Science 2026-01-23 Jaya Chaturvedi , Saniya Deshpande , Chenkai Ma , Robert Cobb , Angus Roberts , Robert Stewart , Daniel Stahl , Diana Shamsutdinova

One type of machine learning, text classification, is now regularly applied in the legal matters involving voluminous document populations because it can reduce the time and expense associated with the review of those documents. One form of…

Information Retrieval · Computer Science 2019-04-04 Rishi Chhatwal , Nathaniel Huber-Fliflet , Robert Keeling , Jianping Zhang , Haozhen Zhao

The impressive progress in NLP techniques has been driven by the development of multi-task benchmarks such as GLUE and SuperGLUE. While these benchmarks focus on tasks for one or two input sentences, there has been exciting work in…

Computation and Language · Computer Science 2025-10-20 G Thomas Hudson , Noura Al Moubayed

The availability of large on-line text corpora provides a natural and promising bridge between the worlds of natural language processing (NLP) and machine learning (ML). In recent years, the NLP community has been aggressively investigating…

cmp-lg · Computer Science 2008-02-03 Stephen Soderland , Wendy Lehnert

The advent of large language models (LLMs) has led to significant achievements in various domains, including legal text processing. Leveraging LLMs for legal tasks is a natural evolution and an increasingly compelling choice. However, their…

Computation and Language · Computer Science 2025-07-29 Tan-Minh Nguyen , Hoang-Trung Nguyen , Trong-Khoi Dao , Xuan-Hieu Phan , Ha-Thanh Nguyen , Thi-Hai-Yen Vuong

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering total of 1.17 billion speakers and 121 languages have more…

Document layout analysis is essential for downstream tasks such as information retrieval, extraction, OCR, and digitization. However, existing large-scale datasets like PubLayNet and DocBank lack fine-grained region labels and multilingual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Oikantik Nath , Sahithi Kukkala , Mitesh Khapra , Ravi Kiran Sarvadevabhatla

Knowledge-intensive language tasks (KILTs) typically require retrieving relevant documents from trustworthy corpora, e.g., Wikipedia, to produce specific answers. Very recently, a pre-trained generative retrieval model for KILTs, named…

Information Retrieval · Computer Science 2024-02-27 Jiafeng Guo , Changjiang Zhou , Ruqing Zhang , Jiangui Chen , Maarten de Rijke , Yixing Fan , Xueqi Cheng

The legal landscape encompasses a wide array of lawsuit types, presenting lawyers with challenges in delivering timely and accurate information to clients, particularly concerning critical aspects like potential imprisonment duration or…

Artificial Intelligence · Computer Science 2024-07-30 Jia-Hong Huang , Chao-Chun Yang , Yixian Shen , Alessio M. Pacces , Evangelos Kanoulas

Legal charge prediction, an essential task in legal AI, seeks to assign accurate charge labels to case descriptions, attracting significant recent interest. Existing methods primarily employ diverse neural network structures for modeling…

Computation and Language · Computer Science 2024-08-06 Jingyun Sun , Chi Wei , Yang Li

We address the task of hierarchical multi-label classification (HMC) of scientific documents at an industrial scale, where hundreds of thousands of documents must be classified across thousands of dynamic labels. The rapid growth of…

Artificial Intelligence · Computer Science 2024-12-09 Seyed Amin Tabatabaei , Sarah Fancher , Michael Parsons , Arian Askari

The judiciary, as one of democracy's three pillars, is dealing with a rising amount of legal issues, needing careful use of judicial resources. This research presents a complex framework that leverages Data Science methodologies, notably…

Information Retrieval · Computer Science 2025-07-03 Puspendu Banerjee , Aritra Mazumdar , Wazib Ansar , Saptarsi Goswami , Amlan Chakrabarti

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present in any other LLM…

Computation and Language · Computer Science 2026-04-02 Jesse van Oort , Frank Brinkkemper , Erik de Graaf , Bram Vanroy , Saskia Lensink

In today's legal environment, lawsuits and regulatory investigations require companies to embark upon increasingly intensive data-focused engagements to identify, collect and analyze large quantities of data. When documents are staged for…

Information Retrieval · Computer Science 2019-04-04 Rishi Chhatwal , Peter Gronvall , Nathaniel Huber-Fliflet , Robert Keeling , Jianping Zhang , Haozhen Zhao