English
Related papers

Related papers: Bridging the Data Provenance Gap Across Text, Spee…

200 papers

AI-powered search systems are emerging as new information gatekeepers, fundamentally transforming how users access news and information. Despite their growing influence, the citation patterns of these systems remain poorly understood. We…

Information Retrieval · Computer Science 2025-07-09 Kai-Cheng Yang

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-resource languages,…

Computation and Language · Computer Science 2025-04-01 Xinyu Wang , Wenbo Zhang , Sarah Rajtmajer

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and online platforms give rise to many scientific problems…

Multimedia · Computer Science 2025-02-11 Yu Shang , Chen Gao , Nian Li , Yong Li

The increasing prevalence of mental disorders globally highlights the urgent need for effective digital screening methods that can be used in multilingual contexts. Most existing studies, however, focus on English data, overlooking critical…

Computation and Language · Computer Science 2026-01-27 Ana-Maria Bucur , Marcos Zampieri , Tharindu Ranasinghe , Fabio Crestani

In 2022, with the release of ChatGPT, large-scale language models gained widespread attention. ChatGPT not only surpassed previous models in terms of parameters and the scale of its pretraining corpus but also achieved revolutionary…

Artificial Intelligence · Computer Science 2024-11-13 Yiming Ju , Huanhuan Ma

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

Machine Learning · Computer Science 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

Artificial Intelligence (AI) is now firmly at the center of evidence-based medicine. Despite many success stories that edge the path of AI's rise in healthcare, there are comparably many reports of significant shortcomings and unexpected…

Text representation models are prone to exhibit a range of societal biases, reflecting the non-controlled and biased nature of the underlying pretraining data, which consequently leads to severe ethical issues and even bias amplification.…

Computation and Language · Computer Science 2021-06-08 Soumya Barikeri , Anne Lauscher , Ivan Vulić , Goran Glavaš

The increasing tendency to collect large and uncurated datasets to train vision-and-language models has raised concerns about fair representations. It is known that even small but manually annotated datasets, such as MSCOCO, are affected by…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Noa Garcia , Yusuke Hirota , Yankun Wu , Yuta Nakashima

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

Human-Computer Interaction · Computer Science 2026-04-07 Alexis Carrillo , Enrique Taietta , Ali Aghazadeh Ardebili , Giuseppe Alessandro Veltri , Massimo Stella

Machine learning-based behavioral models rely on features extracted from audio-visual recordings. The recordings are processed using open-source tools to extract speech features for classification models. These tools often lack validation…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Tahiya Chowdhury , Veronica Romero

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under…

People spend a substantial portion of their lives engaged in conversation, and yet our scientific understanding of conversation is still in its infancy. In this report we advance an interdisciplinary science of conversation, with findings…

Computation and Language · Computer Science 2022-03-02 Andrew Reece , Gus Cooney , Peter Bull , Christine Chung , Bryn Dawson , Casey Fitzpatrick , Tamara Glazer , Dean Knox , Alex Liebscher , Sebastian Marin

In the age of information overload, professionals across various fields face the challenge of navigating vast amounts of documentation and ever-evolving standards. Ensuring compliance with standards, regulations, and contractual obligations…

Cryptography and Security · Computer Science 2024-07-22 Shohreh Deldari , Mohammad Goudarzi , Aditya Joshi , Arash Shaghaghi , Simon Finn , Flora D. Salim , Sanjay Jha

People increasingly use videos on the Web as a source for learning. To support this way of learning, researchers and developers are continuously developing tools, proposing guidelines, analyzing data, and conducting experiments. However, it…

Multimedia · Computer Science 2023-08-15 Evelyn Navarrete , Andreas Nehring , Sascha Schanze , Ralph Ewerth , Anett Hoppe

The rise of Multimodal Large Language Models (MLLMs) has become a transformative force in the field of artificial intelligence, enabling machines to process and generate content across multiple modalities, such as text, images, audio, and…

Computation and Language · Computer Science 2025-12-09 Ming Li , Keyu Chen , Ziqian Bi , Ming Liu , Xinyuan Song , Zekun Jiang , Tianyang Wang , Benji Peng , Qian Niu , Junyu Liu , Jinlang Wang , Sen Zhang , Xuanhe Pan , Jiawei Xu , Pohsun Feng

Corporate AI-washing-the strategic misrepresentation of AI capabilities via exaggerated or fabricated cross-channel disclosures-has emerged as a systemic threat to capital market information integrity with the widespread adoption of…

Computers and Society · Computer Science 2026-04-14 Zhanjie Wen , Jingqiao Guo

Conversational search is a relatively young area of research that aims at automating an information-seeking dialogue. In this paper we help to position it with respect to other research areas within conversational Artificial Intelligence…

Information Retrieval · Computer Science 2021-06-09 Svitlana Vakulenko , Evangelos Kanoulas , Maarten de Rijke