English
Related papers

Related papers: Including Dialects and Language Varieties in Autho…

200 papers

Human personality is significantly represented by those words which he/she uses in his/her speech or writing. As a consequence of spreading the information infrastructures (specifically the Internet and social media), human communications…

Generative large language models (LLMs) have become central to everyday life, producing human-like text across diverse domains. A growing body of research investigates whether these models also exhibit personality- and demographic-like…

Computation and Language · Computer Science 2025-10-14 Dana Sotto Porat , Ella Rabinovich

The paper describes a transformer-based system designed for SemEval-2023 Task 9: Multilingual Tweet Intimacy Analysis. The purpose of the task was to predict the intimacy of tweets in a range from 1 (not intimate at all) to 5 (very…

Computation and Language · Computer Science 2023-12-19 Anna Glazkova

The widespread adoption of large language models (LLMs) has made it difficult to distinguish human writing from machine-produced text in many real applications. Detectors that were effective for one generation of models tend to degrade when…

Computation and Language · Computer Science 2025-12-09 Sepyan Purnama Kristanto , Lutfi Hakim , Dianni Yusuf

Authorship representation (AR) learning, which models an author's unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual settings-mostly in…

Computation and Language · Computer Science 2025-09-23 Junghwan Kim , Haotian Zhang , David Jurgens

The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. While recent research has…

Computation and Language · Computer Science 2025-10-01 Sergio E. Zanotto , Segun Aroyehun

his paper describes our techniques to detect hate speech against women and immigrants on Twitter in multilingual contexts, particularly in English and Spanish. The challenge was designed by SemEval-2019 Task 5, where the participants need…

Computation and Language · Computer Science 2020-11-30 Alvi Md Ishmam

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

Machine Learning · Statistics 2017-02-07 Bruno Gonçalves , David Sánchez

This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech. Demographic feature prediction plays a crucial role in…

Computation and Language · Computer Science 2025-02-18 Yuchen Yang , Thomas Thebaud , Najim Dehak

Forensic authorship profiling uses linguistic markers to infer characteristics about an author of a text. This task is paralleled in dialect classification, where a prediction is made about the linguistic variety of a text based on the text…

Computation and Language · Computer Science 2024-07-02 Dana Roemling , Yves Scherrer , Aleksandra Miletic

Content has historically been the primary lens used to study language in online communities. This paper instead focuses on the linguistic style of communities. While we know that individuals have distinguishable styles, here we ask whether…

Computation and Language · Computer Science 2022-09-28 Osama Khalid , Padmini Srinivasan

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

Computation and Language · Computer Science 2019-04-08 Shikha Bordia , Samuel R. Bowman

We report our models for detecting age, language variety, and gender from social media data in the context of the Arabic author profiling and deception detection shared task (APDA). We build simple models based on pre-trained bidirectional…

Computation and Language · Computer Science 2019-11-01 Chiyu Zhang , Muhammad Abdul-Mageed

Semantics, morphology and syntax are strongly interdependent. However, the majority of computational methods for semantic change detection use distributional word representations which encode mostly semantics. We investigate an alternative…

Computation and Language · Computer Science 2021-09-23 Mario Giulianelli , Andrey Kutuzov , Lidia Pivovarova

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and abusive language.…

Computation and Language · Computer Science 2019-05-30 Thomas Davidson , Debasmita Bhattacharya , Ingmar Weber

With the growth of social medias, such as Twitter, plenty of user-generated data emerge daily. The short texts published on Twitter -- the tweets -- have earned significant attention as a rich source of information to guide many…

Artificial Intelligence · Computer Science 2021-06-01 Sérgio Barreto , Ricardo Moura , Jonnathan Carvalho , Aline Paes , Alexandre Plastino

AI models have become extremely popular and accessible to the general public. However, they are continuously under the scanner due to their demonstrable biases toward various sections of the society like people of color and non-binary…

Computers and Society · Computer Science 2023-10-11 Siddharth D Jaiswal , Ankit Kumar Verma , Animesh Mukherjee

The evolution of the internet has created an abundance of unstructured data on the web, a significant part of which is textual. The task of author profiling seeks to find the demographics of people solely from their linguistic and…

Machine Learning · Computer Science 2019-06-24 Joey Hong , Chris Mattmann , Paul Ramirez

Predicting gender by the first name is not a simple task. In many applications, especially in the natural language processing (NLP) field, this task may be necessary, mainly when considering foreign names. In this paper, we examined and…

Machine Learning · Computer Science 2021-10-05 Rosana C. B. Rego , Verônica M. L. Silva , Victor M. Fernandes

This paper presents a software allowing to describe voices using a continuous Voice Femininity Percentage (VFP). This system is intended for transgender speakers during their voice transition and for voice therapists supporting them in this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-24 David Doukhan , Simon Devauchelle , Lucile Girard-Monneron , Mía Chávez Ruz , V. Chaddouk , Isabelle Wagner , Albert Rilliard
‹ Prev 1 4 5 6 7 8 10 Next ›