中文
相关论文

相关论文: V(is)owel: An Interactive Vowel Chart to Understan…

200 篇论文

With the rise of the gig economy, online language tutoring platforms are becoming increasingly popular. These platforms provide temporary and flexible jobs for native speakers as tutors and allow language learners to have one-on-one…

人机交互 · 计算机科学 2022-04-19 Meng Xia , Yankun Zhao , Jihyeong Hong , Mehmet Hamza Erol , Taewook Kim , Juho Kim

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such…

计算与语言 · 计算机科学 2022-03-18 Woojeong Jin , Dong-Ho Lee , Chenguang Zhu , Jay Pujara , Xiang Ren

Objectives: Several studies have focused on the role of working memory (WM) in predicting mathematical and reading literacy. Alternative models of WM have been proposed and a modality-dependent model of WM, distinguishing between verbal and…

历史与综述 · 数学 2022-01-04 David Giofrè , Enrica Donolato , Irene C. Mammarella

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected speaker not…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Tomoya Yoshinaga , Keitaro Tanaka , Shigeo Morishima

The 'keyword method' is an effective technique for learning vocabulary of a foreign language. It involves creating a memorable visual link between what a word means and what its pronunciation in a foreign language sounds like in the…

In automotive domain, operation of secondary tasks like accessing infotainment system, adjusting air conditioning vents, and side mirrors distract drivers from driving. Though existing modalities like gesture and speech recognition systems…

人机交互 · 计算机科学 2021-06-11 Gowdham Prabhakar , Priyam Rajkhowa , Pradipta Biswas

This innovative practice article reports on the piloting of vibe coding (using natural language to create software applications with AI) for English as a Foreign Language (EFL) education. We developed a human-AI meta-languaging framework…

计算机与社会 · 计算机科学 2025-09-12 David James Woo , Kai Guo , Yangyang Yu

Input multimodality combining speech and hand gestures has motivated numerous usability studies. Contrastingly, issues relating to the design and ergonomic evaluation of multimodal output messages combining speech with visual modalities…

人机交互 · 计算机科学 2007-09-05 Suzanne Kieffer , Noëlle Carbonell

Sign language datasets are often not representative in terms of vocabulary, underscoring the need for models that generalize to unseen signs. Vector quantization is a promising approach for learning discrete, token-like representations, but…

计算与语言 · 计算机科学 2025-09-08 Lee Kezar , Zed Sehyr , Jesse Thomason

The rise of wearable smart devices raises unprecedented opportunities for self-improvement through ubiquitous behavior tracking and guidance. However, the design of effective wearable behavior intervention systems remains relatively…

人机交互 · 计算机科学 2025-07-08 Zhang Youpeng , Nuwan Janaka , Ashwin Ram , Yin Peilin , Tian Yang , Shengdong Zhao , Pierre Dragicevic

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as `visemes'. In this…

图像与视频处理 · 电气工程与系统科学 2019-09-17 Helen Bear , Richard Harvey

What makes speeches effective has long been a subject for debate, and until today there is broad controversy among public speaking experts about what factors make a speech effective as well as the roles of these factors in speeches.…

人机交互 · 计算机科学 2021-11-01 Kevin Maher , Zeyuan Huang , Jiancheng Song , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Hao Wang , Yong-Jin Liu , Hongan Wang

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to…

音频与语音处理 · 电气工程与系统科学 2025-04-29 Zhaofeng Lin , Naomi Harte

We present the results of an exploratory study on how pairs interact with speech commands and touch gestures on a wall-sized display during a collaborative sensemaking task. Previous work has shown that speech commands, alone or in…

人机交互 · 计算机科学 2025-09-18 Gabriela Molina León , Anastasia Bezerianos , Olivier Gladin , Petra Isenberg

A key challenge to visualization authoring is the process of getting familiar with the complex user interfaces of authoring tools. Natural Language Interface (NLI) presents promising benefits due to its learnability and usability. However,…

人机交互 · 计算机科学 2022-10-25 Yun Wang , Zhitao Hou , Leixian Shen , Tongshuang Wu , Jiaqi Wang , He Huang , Haidong Zhang , Dongmei Zhang

Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval. To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential. However, frame-level…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuanbo Hou , Frank K. Soong , Jian Luan , Shengchen Li

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This…

A main challenge of Visual-Language Tracking (VLT) is the misalignment between visual inputs and language descriptions caused by target movement. Previous trackers have explored many effective feature modification methods to preserve more…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Yihao Zhen , Qiang Wang , Yu Qiao , Liangqiong Qu , Huijie Fan

English speech rhythm, the temporal patterns of stressed syllables, is essential for English as a second language (ESL) learners to produce natural-sounding and comprehensible speech. Rhythm training is generally based on imitation of…

人机交互 · 计算机科学 2025-07-28 Chang Chen , Sicheng Song , Shuchang Xu , Zhicheng Li , Huamin Qu , Yanna Lin

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden