English
Related papers

Related papers: Linguistic Legal Concept Extraction in Portuguese

200 papers

The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically designed for the Portuguese legal system. In this work, we…

Computation and Language · Computer Science 2025-02-25 Beatriz Canaverde , Telmo Pessoa Pires , Leonor Melo Ribeiro , André F. T. Martins

We present and make available pre-trained language models (Phraser, Word2Vec, Doc2Vec, FastText, and BERT) for the Brazilian legal language, a Python package with functions to facilitate their use, and a set of demonstrations/tutorials…

Being able to clearly interpret legal texts and fully understanding our rights, obligations and other legal norms has become progressively more important in the digital society. However, simply giving citizens access to the laws is not…

Computation and Language · Computer Science 2022-03-11 Maria Duarte , Pedro A. Santos , João Dias , Jorge Baptista

In Brazil, all legal professionals must demonstrate their knowledge of the law and its application by passing the OAB exams, the national bar exams. The OAB exams therefore provide an excellent benchmark for the performance of legal…

Computation and Language · Computer Science 2017-12-15 Pedro Delfino , Bruno Cuconato , Edward Hermann Haeusler , Alexandre Rademaker

The objective of this paper is to develop predictive models to classify Brazilian legal proceedings in three possible classes of status: (i) archived proceedings, (ii) active proceedings, and (iii) suspended proceedings. This problem's…

Computation and Language · Computer Science 2021-06-24 Felipe Maia Polo , Itamar Ciochetti , Emerson Bertolo

Event extraction is an Information Retrieval task that commonly consists of identifying the central word for the event (trigger) and the event's arguments. This task has been extensively studied for English but lags behind for Portuguese,…

Computation and Language · Computer Science 2024-09-02 Luís Filipe Cunha , Ricardo Campos , Alípio Jorge

Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of…

Computation and Language · Computer Science 2025-01-22 Marlo Souza , Bruno Cabral , Daniela Claro , Lais Salvador

Automatic reasoning about textual information is a challenging task in modern Natural Language Processing (NLP) systems. In this work we describe our proposal for representing and reasoning about Portuguese documents by means of Linked Data…

Computation and Language · Computer Science 2019-09-13 Paulo Quaresma , Vitor Beires Nogueira , Kashyap Raiyani , Roy Bayot , Teresa Gonçalves

The Brazilian judiciary has a large workload, resulting in a long time to finish legal proceedings. Brazilian National Council of Justice has established in Resolution 469/2022 formal guidance for document and process digitalization opening…

Computation and Language · Computer Science 2023-06-02 Daniel da Silva Junior , Paulo Roberto dos S. Corval , Aline Paes , Daniel de Oliveira

This work addresses the challenge of capturing the complexities of legal knowledge by proposing a multi-layered embedding-based retrieval method for legal and legislative texts. Creating embeddings not only for individual articles but also…

Artificial Intelligence · Computer Science 2025-03-13 João Alberto de Oliveira Lima

This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leaderboard comes as a way to address a gap in the evaluation of LLM for…

Computation and Language · Computer Science 2026-03-16 João Silva , Luís Gomes , António Branco

Legal Judgment Prediction is one of the most acclaimed fields for the combined area of NLP, AI, and Law. By legal prediction we mean an intelligent systems capable to predict specific judicial characteristics, such as judicial outcome, a…

Machine Learning · Computer Science 2022-12-29 Vithor Gomes Ferreira Bertalan , Evandro Eduardo Seron Ruiz

Despite the recent advances in Large Language Models, benchmarks for evaluating legal writing remain scarce due to the inherent complexity of assessing open-ended responses in this domain. One of the key challenges in evaluating language…

Computation and Language · Computer Science 2025-05-01 Ramon Pires , Roseval Malaquias Junior , Rodrigo Nogueira

Since a lexicon-based approach is more elegant scientifically, explaining the solution components and being easier to generalize to other applications, this paper provides a new approach for offensive language and hate speech detection on…

Recent advances in deep learning have promoted the advent of many computational systems capable of performing intelligent actions that, until then, were restricted to the human intellect. In the particular case of human languages, these…

Computation and Language · Computer Science 2024-05-02 Rúben Almeida , Evelin Amorim

This paper introduces JurisTCU, a Brazilian Portuguese dataset for legal information retrieval (LIR). The dataset is freely available and consists of 16,045 jurisprudential documents from the Brazilian Federal Court of Accounts, along with…

Data derived from the realm of the social sciences is often produced in digital text form, which motivates its use as a source for natural language processing methods. Researchers and practitioners have developed and relied on artificial…

Computation and Language · Computer Science 2024-06-04 Raysa Benatti , Fabiana Severi , Sandra Avila , Esther Luna Colombini

Much recent effort has been devoted to creating large-scale language models. Nowadays, the most prominent approaches are based on deep neural networks, such as BERT. However, they lack transparency and interpretability, and are often seen…

Computation and Language · Computer Science 2023-06-08 Rodrigo Wilkens , Leonardo Zilio , Aline Villavicencio

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese…

Information Retrieval · Computer Science 2024-04-11 Mirelle Bueno , Eduardo Seiti de Oliveira , Rodrigo Nogueira , Roberto A. Lotufo , Jayr Alencar Pereira

Ambiguity is a characteristic of natural language, which makes expression ideas flexible. However, in a domain that requires accurate statements, it becomes a barrier. Specifically, a single word can have many meanings and multiple words…

Computation and Language · Computer Science 2021-06-28 Ha-Thanh Nguyen , Vu Tran , Phuong Minh Nguyen , Thi-Hai-Yen Vuong , Quan Minh Bui , Chau Minh Nguyen , Binh Tran Dang , Minh Le Nguyen , Ken Satoh
‹ Prev 1 2 3 10 Next ›