English
Related papers

Related papers: SpeechMirror: A Multimodal Visual Analytics System…

200 papers

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Alexander H. Liu , Sung-Lin Yeh , James Glass

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in human-machine…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Tan-Hanh Pham , Hoang-Nam Le , Phu-Vinh Nguyen , Chris Ngo , Truong-Son Hy

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

Multimedia · Computer Science 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

We present Speakerly, a new real-time voice-based writing assistance system that helps users with text composition across various use cases such as emails, instant messages, and notes. The user can interact with the system through…

Computation and Language · Computer Science 2023-10-26 Dhruv Kumar , Vipul Raheja , Alice Kaiser-Schatzlein , Robyn Perry , Apurva Joshi , Justin Hugues-Nuger , Samuel Lou , Navid Chowdhury

Public speaking and presentation competence plays an essential role in many areas of social interaction in our educational, professional, and everyday life. Since our intention during a speech can differ from what is actually understood by…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Ömer Sümer , Cigdem Beyan , Fabian Ruth , Olaf Kramer , Ulrich Trautwein , Enkelejda Kasneci

Advances in AI have enabled ESL learners to practice speaking through conversational systems. However, most tools rely on explicit correction, which can interrupt the conversation and undermine confidence. Grounded in second language…

Human-Computer Interaction · Computer Science 2026-02-05 Minju Park , Seunghyun Lee , Juhwan Ma , Dongwook Yoon

We present VOICE, a novel approach to science communication that connects large language models' (LLM) conversational capabilities with interactive exploratory visualization. VOICE introduces several innovative technical contributions that…

Human-Computer Interaction · Computer Science 2026-05-19 Donggang Jia , Alexandra Irger , Lonni Besancon , Ondrej Strnad , Deng Luo , Johanna Bjorklund , Anders Ynnerman , Ivan Viola

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages…

Multimedia · Computer Science 2023-10-24 Joanna Hong , Se Jin Park , Yong Man Ro

In remote video meetings, visual non-verbal cues, such as facial expressions or head movements, are seen continuously but often only partially. This increases ambiguity compared to in-person settings and can cause misinterpretation or…

Human-Computer Interaction · Computer Science 2026-05-04 Gun Woo Warren Park , Anthony Tang , Fanny Chevalier

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

Sound · Computer Science 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

Most popular goal-oriented dialogue agents are capable of understanding the conversational context. However, with the surge of virtual assistants with screen, the next generation of agents are required to also understand screen context in…

Machine Learning · Computer Science 2021-11-26 Sanchit Agarwal , Jan Jezabek , Arijit Biswas , Emre Barut , Shuyang Gao , Tagyoung Chung

How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural…

Computer Vision and Pattern Recognition · Computer Science 2019-05-24 Tae-Hyun Oh , Tali Dekel , Changil Kim , Inbar Mosseri , William T. Freeman , Michael Rubinstein , Wojciech Matusik

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

Computation and Language · Computer Science 2025-08-19 Shumin Que , Anton Ragni

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

The think aloud method is an important and commonly used tool for usability optimization. However, analyzing think aloud data could be time consuming. In this paper, we put forth an automatic analysis of verbal protocols and test the link…

Human-Computer Interaction · Computer Science 2023-07-12 Supriya Murali , Tina Walber , Christoph Schaefer , Sezen Lim

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

Image and Video Processing · Electrical Eng. & Systems 2022-09-27 Rahul Sharma , Shrikanth Narayanan