English
Related papers

Related papers: Romanian Multiword Expression Detection Using Mult…

200 papers

Correctly identifying multiword expressions (MWEs) is an important task for most natural language processing systems since their misidentification can result in ambiguity and misunderstanding of the underlying text. In this work, we…

Computation and Language · Computer Science 2023-06-21 Andrei-Marius Avram , Verginica Barbu Mititelu , Vasile Păiş , Dumitru-Clementin Cercel , Ştefan Trăuşan-Matu

With the rise of bidirectional encoder representations from Transformer models in natural language processing, the speech community has adopted some of their development methodologies. Therefore, the Wav2Vec models were introduced to reduce…

Computation and Language · Computer Science 2023-07-03 Andrei-Marius Avram , Răzvan-Alexandru Smădu , Vasile Păiş , Dumitru-Clementin Cercel , Radu Ion , Dan Tufiş

We propose a multilingual adversarial training model for determining whether a sentence contains an idiomatic expression. Given that a key challenge with this task is the limited size of annotated data, our model relies on pre-trained…

Computation and Language · Computer Science 2022-06-08 Lis Kanashiro Pereira , Ichiro Kobayashi

In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English…

We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141…

Offensive language detection is a crucial task in today's digital landscape, where online platforms grapple with maintaining a respectful and inclusive environment. However, building robust offensive language detection models requires large…

Computation and Language · Computer Science 2024-07-31 Elena-Beatrice Nicola , Dumitru-Clementin Cercel , Florin Pop

Multilingual BERT (mBERT), XLM-RoBERTa (XLMR) and other unsupervised multilingual encoders can effectively learn cross-lingual representation. Explicit alignment objectives based on bitexts like Europarl or MultiUN have been shown to…

Computation and Language · Computer Science 2020-10-07 Shijie Wu , Mark Dredze

This paper describes our approach to the task of identifying offensive languages in a multilingual setting. We investigate two data augmentation strategies: using additional semi-supervised labels with different thresholds and cross-lingual…

Computation and Language · Computer Science 2020-08-05 Hwijeen Ahn , Jimin Sun , Chan Young Park , Jungyun Seo

This paper presents a language-independent deep learning architecture adapted to the task of multiword expression (MWE) identification. We employ a neural architecture comprising of convolutional and recurrent layers with the addition of an…

Computation and Language · Computer Science 2018-09-11 Shiva Taslimipoor , Omid Rohanian

Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30k dataset into Romanian…

Computation and Language · Computer Science 2025-12-18 George-Andrei Dima , Dumitru-Clementin Cercel

Pre-trained language models (PLMs) have consistently demonstrated outstanding performance across a diverse spectrum of natural language processing tasks. Nevertheless, despite their success with unseen data, current PLM-based…

Computation and Language · Computer Science 2024-03-19 Javad Rafiei Asl , Prajwal Panzade , Eduardo Blanco , Daniel Takabi , Zhipeng Cai

The remarkable achievements obtained by open-source large language models (LLMs) in recent years have predominantly been concentrated on tasks involving the English language. In this paper, we aim to advance the performance of Llama2 models…

Computation and Language · Computer Science 2024-10-08 George-Andrei Dima , Andrei-Marius Avram , Cristian-George Crăciun , Dumitru-Clementin Cercel

Ensuring that both new and experienced drivers master current traffic rules is critical to road safety. This paper evaluates Large Language Models (LLMs) on Romanian driving-law QA with explanation generation. We release a 1{,}208-question…

Computation and Language · Computer Science 2025-09-30 Eduard Barbu , Adrian Marius Dumitran

This paper presents a novel approach for multi-label emotion detection, where Llama-3 is used to generate explanatory content that clarifies ambiguous emotional expressions, thereby enhancing RoBERTa's emotion classification performance. By…

Machine Learning · Computer Science 2025-04-17 Niloofar Ranjbar , Hamed Baghbani

This paper highlights the significance of natural language processing (NLP) within artificial intelligence, underscoring its pivotal role in comprehending and modeling human language. Recent advancements in NLP, particularly in…

Computation and Language · Computer Science 2024-10-01 Vlad-Cristian Matei , Iulian-Marius Tăiatu , Răzvan-Alexandru Smădu , Dumitru-Clementin Cercel

Sarcasm detection identifies natural language expressions whose intended meaning is different from what is implied by its surface meaning. It finds applications in many NLP tasks such as opinion mining, sentiment analysis, etc. Today,…

Multimedia · Computer Science 2021-10-04 Sundesh Gupta , Aditya Shah , Miten Shah , Laribok Syiemlieh , Chandresh Maurya

Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the availability and curation of high-quality data. Data quality is…

Computation and Language · Computer Science 2025-11-04 Vlad Negoita , Mihai Masala , Traian Rebedea

We present a comprehensive approach for multiword expression (MWE) identification that combines binary token-level classification, linguistic feature integration, and data augmentation. Our DeBERTa-v3-large model achieves 69.8% F1 on the…

Computation and Language · Computer Science 2026-01-28 Diego Rossini , Lonneke van der Plas

This paper presents the PALI team's winning system for SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation. We fine-tune XLM-RoBERTa model to solve the task of word in context disambiguation, i.e., to…

Artificial Intelligence · Computer Science 2021-06-08 Shuyi Xie , Jian Ma , Haiqin Yang , Lianxin Jiang , Yang Mo , Jianping Shen

Lip reading or visual speech recognition has gained significant attention in recent years, particularly because of hardware development and innovations in computer vision. While considerable progress has been obtained, most models have only…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Emilian-Claudiu Mănescu , Răzvan-Alexandru Smădu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop
‹ Prev 1 2 3 10 Next ›