中文
相关论文

相关论文: Multilingual textual data: an approach through mul…

200 篇论文

We compare the outcomes of multilingual and crosslingual training for related and unrelated Australian languages with similar phonological inventories. We use the Montreal Forced Aligner to train acoustic models from scratch and adapt a…

计算与语言 · 计算机科学 2025-04-25 Alessio Tosolini , Claire Bowern

Exploring and understanding language data is a fundamental stage in all areas dealing with human language. It allows NLP practitioners to uncover quality concerns and harmful biases in data before training, and helps linguists and social…

计算与语言 · 计算机科学 2024-06-26 Alan Ramponi , Camilla Casula , Stefano Menini

Large Language Models (LLMs) are becoming increasingly capable across global languages. However, the ability to communicate across languages does not necessarily translate to appropriate cultural representations. A key concern is US-centric…

计算与语言 · 计算机科学 2025-09-03 Jonathan Rystrøm , Hannah Rose Kirk , Scott Hale

Text data is inherently temporal. The meaning of words and phrases changes over time, and the context in which they are used is constantly evolving. This is not just true for social media data, where the language used is rapidly influenced…

计算与语言 · 计算机科学 2025-03-05 Kai-Robin Lange , Niklas Benner , Lars Grönberg , Aymane Hachcham , Imene Kolli , Jonas Rieger , Carsten Jentsch

In-context learning is a popular inference strategy where large language models solve a task using only a few labeled demonstrations without needing any parameter updates. Although there have been extensive studies on English in-context…

计算与语言 · 计算机科学 2024-06-10 Miaoran Zhang , Vagrant Gautam , Mingyang Wang , Jesujoba O. Alabi , Xiaoyu Shen , Dietrich Klakow , Marius Mosbach

Linguistic accommodation is the process in which speakers adjust their accent, diction, vocabulary, and other aspects of language according to the communication style of one another. Previous research has shown how linguistic accommodation…

计算与语言 · 计算机科学 2021-01-26 Aviv Ben Haim , Oren Tsur

People communicate in more than 7,000 languages around the world, with around 780 languages spoken in India alone. Despite this linguistic diversity, research on Sentiment Analysis has predominantly focused on English text data, resulting…

计算与语言 · 计算机科学 2024-09-04 Aekansh Kathunia , Mohammad Kaif , Nalin Arora , N Narotam

Computer-based interactive items have become prevalent in recent educational assessments. In such items, the entire human-computer interactive process is recorded in a log file and is known as the response process. This paper aims at…

应用统计 · 统计学 2025-01-08 Xueying Tang , Zhi Wang , Qiwei He , Jingchen Liu , Zhiliang Ying

Probing techniques for large language models (LLMs) have primarily focused on English, overlooking the vast majority of the world's languages. In this paper, we extend these probing methods to a multilingual context, investigating the…

计算与语言 · 计算机科学 2025-02-03 Daoyang Li , Haiyan Zhao , Qingcheng Zeng , Mengnan Du

This paper generalises dynamic factor models for multidimensional dependent data. In doing so, it develops an interpretable technique to study complex information sources ranging from repeated surveys with a varying number of respondents to…

计量经济学 · 经济学 2023-01-31 Matteo Barigozzi , Filippo Pellegrino

When humans read a text, their eye movements are influenced by the structural complexity of the input sentences. This cognitive phenomenon holds across languages and recent studies indicate that multilingual language models utilize…

计算与语言 · 计算机科学 2023-02-28 Charlotte Pouw , Nora Hollenstein , Lisa Beinborn

Large language models (LLMs) provide detailed and impressive responses to queries in English. However, are they really consistent at responding to the same query in other languages? The popular way of evaluating for multilingual performance…

计算与语言 · 计算机科学 2025-05-29 Ashim Gupta , Maitrey Mehta , Zhichao Xu , Vivek Srikumar

Learning interpretable representations of data generative latent factors is an important topic for the development of artificial intelligence. With the rise of the large multimodal model, it can align images with text to generate answers.…

机器学习 · 计算机科学 2024-04-19 Mengdan Zhu , Zhenke Liu , Bo Pan , Abhinav Angirekula , Liang Zhao

Structural probing work has found evidence for latent syntactic information in pre-trained language models. However, much of this analysis has focused on monolingual models, and analyses of multilingual models have employed correlational…

计算与语言 · 计算机科学 2022-10-27 Aaron Mueller , Yu Xia , Tal Linzen

This paper evaluates global-scale dialect identification for 14 national varieties of English as a means for studying syntactic variation. The paper makes three main contributions: (i) introducing data-driven language mapping as a method…

计算与语言 · 计算机科学 2019-04-12 Jonathan Dunn

High-dimensional multivariate longitudinal data, which arise when many outcome variables are measured repeatedly over time, are becoming increasingly common in social, behavioral and health sciences. We propose a latent variable model for…

统计方法学 · 统计学 2025-12-09 Sze Ming Lee , Yunxiao Chen , Tony Sit

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

We propose a Bayesian modeling framework for jointly analyzing multiple functional responses of different types (e.g. binary and continuous data). Our approach is based on a multivariate latent Gaussian process and models the dependence…

统计方法学 · 统计学 2016-01-12 Beth A. Tidemann-Miller , Brian J. Reich , Ana-Maria Staicu

We are proposing a simple, but efficient basic approach for a number of multilingual and cross-lingual language technology applications that are not limited to the usual two or three languages, but that can be applied with relatively little…

计算与语言 · 计算机科学 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Camelia Ignat

Factor analysis is a flexible technique for assessment of multivariate dependence and codependence. Besides being an exploratory tool used to reduce the dimensionality of multivariate data, it allows estimation of common factors that often…

应用统计 · 统计学 2020-05-08 Vitor G. C. da Silva , Kelly C. M. Gonçalves , João B. M. Pereira