中文
相关论文

相关论文: HLDC: Hindi Legal Documents Corpus

200 篇论文

Large Language Models (LLMs), trained on extensive datasets from the web, exhibit remarkable general reasoning skills. Despite this, they often struggle in specialized areas like law, mainly because they lack domain-specific pretraining.…

计算与语言 · 计算机科学 2025-11-27 Mann Khatri , Mirza Yusuf , Rajiv Ratn Shah , Ponnurangam Kumaraguru

In the Indian court system, pending cases have long been a problem. There are more than 4 crore cases outstanding. Manually summarising hundreds of documents is a time-consuming and tedious task for legal stakeholders. Many state-of-the-art…

计算与语言 · 计算机科学 2023-02-20 Satyajit Ghosh , Mousumi Dutta , Tanaya Das

The data article presents the large bilingual parallel corpus of low-resourced language pair Sanskrit-Hindi, named SAHAAYAK 2023. The corpus contains total of 1.5M sentence pairs between Sanskrit and Hindi. To make the universal usability…

计算与语言 · 计算机科学 2023-07-04 Vishvajitsinh Bakrola , Jitendra Nasariwala

The Digital Corpus of Sanskrit records around 650,000 sentences along with their morphological and lexical tagging. But inconsistencies in morphological analysis, and in providing crucial information like the segmented word, urges the need…

计算与语言 · 计算机科学 2020-05-15 Sriram Krishnan , Amba Kulkarni , Gérard Huet

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel…

计算与语言 · 计算机科学 2025-04-23 Rahul Raja , Arpita Vats

Large, high-quality datasets are crucial for training Large Language Models (LLMs). However, so far, there are few datasets available for specialized critical domains such as law and the available ones are often only for the English…

计算与语言 · 计算机科学 2024-05-21 Joel Niklaus , Veton Matoshi , Matthias Stürmer , Ilias Chalkidis , Daniel E. Ho

Mining parallel document pairs for document-level machine translation (MT) remains challenging due to the limitations of existing Cross-Lingual Document Alignment (CLDA) techniques. Existing methods often rely on metadata such as URLs,…

计算与语言 · 计算机科学 2025-11-11 Sanjay Suryanarayanan , Haiyue Song , Mohammed Safi Ur Rahman Khan , Anoop Kunchukuttan , Raj Dabre

The recent advances of deep learning have dramatically changed how machine learning, especially in the domain of natural language processing, can be applied to legal domain. However, this shift to the data-driven approaches calls for larger…

计算与语言 · 计算机科学 2022-10-06 Wonseok Hwang , Dongjun Lee , Kyoungyeon Cho , Hanuhl Lee , Minjoon Seo

Language Identification in textual documents is the process of automatically detecting the language contained in a document based on its content. The present Language Identification techniques presume that a document contains text in one of…

计算与语言 · 计算机科学 2021-06-30 Mohd Zeeshan Ansari , Tanvir Ahmad , Noaima Bari

Automatic topic classification has been studied extensively to assist managing and indexing scientific documents in a digital collection. With the large number of topics being available in recent years, it has become necessary to arrange…

计算与语言 · 计算机科学 2022-11-08 Mobashir Sadat , Cornelia Caragea

Legal NLP remains underdeveloped in regions like India due to the scarcity of structured datasets. We introduce IndianBailJudgments-1200, a new benchmark dataset comprising 1200 Indian court judgments on bail decisions, annotated across 20+…

计算与语言 · 计算机科学 2025-07-04 Sneha Deshmukh , Prathmesh Kamble

In this paper we describe some ways to utilize various lexical resources to improve the quality of statistical machine translation system. We have augmented the training corpus with various lexical resources such as IndoWordnet semantic…

计算与语言 · 计算机科学 2017-03-07 Sreelekha S , Pushpak Bhattacharyya

One of the principal tasks of machine learning with major applications is text classification. This paper focuses on the legal domain and, in particular, on the classification of lengthy legal documents. The main challenge that this study…

计算与语言 · 计算机科学 2019-12-17 Lulu Wan , George Papageorgiou , Michael Seddon , Mirko Bernardoni

Here we search for the best automated classification approach for a set of complex legal documents. Our classification task is not trivial: our aim is to classify ca 30,000 public courthouse records from 12 states and 267 counties at two…

计算与语言 · 计算机科学 2023-12-13 Glen Hopkins , Kristjan Kalm

In this paper, we conduct one of the very first studies for cross-corpora performance evaluation in the spoken language identification (LID) problem. Cross-corpora evaluation was not explored much in LID research, especially for the Indian…

音频与语音处理 · 电气工程与系统科学 2021-05-13 Spandan Dey , Goutam Saha , Md Sahidullah

Automated requirements assessment traditionally relies on universal patterns as proxies for defectiveness, implemented through rule-based heuristics or machine learning classifiers trained on large annotated datasets. However, what…

软件工程 · 计算机科学 2026-01-06 Max Unterbusch , Andreas Vogelsang

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises…

Given the large size and volumes of contracts and their underlying inherent complexity, manual reviews become inefficient and prone to errors, creating a clear need for automation. Automatic Legal Contract Classification (LCC)…

计算与语言 · 计算机科学 2025-07-30 Amrita Singh , Aditya Joshi , Jiaojiao Jiang , Hye-young Paik

Large language models (LLMs) are increasingly used to access legal information. Yet, their deployment in multilingual legal settings is constrained by unreliable retrieval and the lack of domain-adapted, open-embedding models. In…

计算与语言 · 计算机科学 2026-02-11 Narges Baba Ahmadi , Jan Strich , Martin Semmann , Chris Biemann