中文
相关论文

相关论文: V(is)owel: An Interactive Vowel Chart to Understan…

200 篇论文

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Shang-Yi Chuang , Hsin-Min Wang , Yu Tsao

Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream. Advances in deep learning and the availability of large audio-visual datasets have led to the development…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Pingchuan Ma , Stavros Petridis , Maja Pantic

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Sylvain Guy , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, a recurring challenge is ensuring these models utilize visual information as effectively as…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Estelle Aflalo , Gabriela Ben Melech Stan , Tiep Le , Man Luo , Shachar Rosenman , Sayak Paul , Shao-Yen Tseng , Vasudev Lal

Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information…

声音 · 计算机科学 2021-07-02 Bochen Li , Yuxuan Wang , Zhiyao Duan

This paper unravels the design tricks adopted by us, the champion team MReaL-BDAI, for Visual Dialog Challenge 2019: two causal principles for improving Visual Dialog (VisDial). By "improving", we mean that they can promote almost every…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Jiaxin Qi , Yulei Niu , Jianqiang Huang , Hanwang Zhang

With the recent developments in speech synthesis via machine learning, this study explores incorporating linguistics knowledge to visualise and evaluate synthetic speech model training. If changes to the first and second formant (in turn,…

音频与语音处理 · 电气工程与系统科学 2022-08-23 Binu Abeysinghe , Jesin James , Catherine I. Watson , Felix Marattukalam

Visual question answering (VQA) is the task of answering questions about an image. The task assumes an understanding of both the image and the question to provide a natural language answer. VQA has gained popularity in recent years due to…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Deepanway Ghosal , Navonil Majumder , Roy Ka-Wei Lee , Rada Mihalcea , Soujanya Poria

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Ziyi Ni , Minglun Han , Feilong Chen , Linghui Meng , Jing Shi , Pin Lv , Bo Xu

Speech perception is key to verbal communication. For people with hearing loss, the capability to recognize speech is restricted, particularly in a noisy environment or the situations without visual cues, such as lip-reading unavailable via…

声音 · 计算机科学 2020-12-21 Rung-Yu Tseng , Tao-Wei Wang , Szu-Wei Fu , Chia-Ying Lee , Yu Tsao

A preliminary experimental study is presented, that aims at eliciting the contribution of oral messages to facilitating visual search tasks on crowded visual displays. Results of quantitative and qualitative analyses suggest that…

人机交互 · 计算机科学 2007-09-05 Noëlle Carbonell , Suzanne Kieffer

Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the ability to interpret and use tactile signals, limiting their…

机器人学 · 计算机科学 2025-07-30 Jianxin Bi , Kevin Yuchen Ma , Ce Hao , Mike Zheng Shou , Harold Soh

As communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack…

人机交互 · 计算机科学 2023-09-12 Zeyuan Huang , Qiang He , Kevin Maher , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Sheng-feng Qin , Yong-Jin Liu , Hongan Wang

Vocabulary acquisition in early education often relies on rote memorization and passive screen-based tools, which can fail to engage students kinesthetically and collaboratively. This paper introduces Vocabuild, an augmented tangible…

人机交互 · 计算机科学 2025-09-16 Siying Hu , Zhenhao Zhang

Image-based visual-language (I-VL) pre-training has shown great success for learning joint visual-textual representations from large-scale web data, revealing remarkable ability for zero-shot generalisation. This paper presents a simple but…

计算机视觉与模式识别 · 计算机科学 2022-07-18 Chen Ju , Tengda Han , Kunhao Zheng , Ya Zhang , Weidi Xie

Rubrics and oral feedback are approaches to help students improve performance and meet learning outcomes. However, their effect on the actual improvement achieved is inconclusive. This paper evaluates the effect of rubrics and oral feedback…

软件工程 · 计算机科学 2023-07-25 Sebastian Barney , Mahvish Khurum , Kai Petersen , Michael Unterkalmsteiner , Ronald Jabangwe

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

The ability to read, understand, and comprehend visual information representations is subsumed under the term visualization literacy (VL). One possibility to improve the use of information visualizations is to introduce adaptations.…

人机交互 · 计算机科学 2022-07-18 Marc Satkowski , Franziska Kessler , Susanne Narciss , Raimund Dachselt

Optical see-through augmented reality (OST-AR) overlays digital targets and annotations on the physical world, offering promising guidance for hands-on tasks such as medical needle insertion or assembly. Recent work on OST-AR depth…

图形学 · 计算机科学 2025-10-03 Hu Guo , Lily Patel , Rohan Gupt

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha