中文
相关论文

相关论文: RhythmTA: A Visual-Aided Interactive System for ES…

200 篇论文

Speech rhythms have been dealt with in three main ways: from the introspective analyses of rhythm as a correlate of syllable and foot timing in linguistics and applied linguistics, through analyses of durations of segments of utterances…

神经元与认知 · 定量生物学 2019-03-14 Dafydd Gibbon , Xuewei Lin

In the last ongoing years, there has been a significant ascending on the field of Natural Language Processing (NLP) for performing multiple tasks including English Language Teaching (ELT). An effective strategy to favor the learning process…

计算与语言 · 计算机科学 2023-09-21 Carlos Morales-Torres , Mario Campos-Soberanis , Diego Campos-Sobrino

It is proposed that the theory of dynamical systems offers appropriate tools to model many phonological aspects of both speech production and perception. A dynamic account of speech rhythm is shown to be useful for description of both…

cmp-lg · 计算机科学 2008-02-03 Robert Port , Fred Cummins , Michael Gasser

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Ziqi Dai , Yiting Chen , Jiacheng Xu , Liufei Xie , Yuchen Wang , Zhenchuan Yang , Bingsong Bai , Yangsheng Gao , Wenjiang Zhou , Weifeng Zhao , Ruohua Zhou

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to…

Speech evaluation is an essential component in computer-assisted language learning (CALL). While speech evaluation on English has been popular, automatic speech scoring on low resource languages remains challenging. Work in this area has…

计算与语言 · 计算机科学 2021-07-09 Huayun Zhang , Ke Shi , Nancy F. Chen

Despite renewed awareness of the importance of articulation, it remains a challenge for instructors to handle the pronunciation needs of language learners. There are relatively scarce pedagogical tools for pronunciation teaching and…

计算机视觉与模式识别 · 计算机科学 2020-05-15 M. Hamed Mozaffari , Won-Sook Lee

Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to…

声音 · 计算机科学 2023-05-05 Tenglong Ao , Qingzhe Gao , Yuke Lou , Baoquan Chen , Libin Liu

This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of…

计算与语言 · 计算机科学 2025-12-30 Jeongsoo Choi , Jaehun Kim , Joon Son Chung

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that…

音频与语音处理 · 电气工程与系统科学 2025-06-10 Nick Mehlman , Thomas Thebaud , Dani Byrd , Shri Narayanan

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

音频与语音处理 · 电气工程与系统科学 2019-07-08 Noé Tits

Expressive text-to-speech (TTS) aims to synthesize different speaking style speech according to human's demands. Nowadays, there are two common ways to control speaking styles: (1) Pre-defining a group of speaking style and using…

声音 · 计算机科学 2023-06-27 Dongchao Yang , Songxiang Liu , Rongjie Huang , Chao Weng , Helen Meng

With the rise of the gig economy, online language tutoring platforms are becoming increasingly popular. These platforms provide temporary and flexible jobs for native speakers as tutors and allow language learners to have one-on-one…

人机交互 · 计算机科学 2022-04-19 Meng Xia , Yankun Zhao , Jihyeong Hong , Mehmet Hamza Erol , Taewook Kim , Juho Kim

Among various methods to learn a second language (L2), such as listening and shadowing, Extensive Viewing involves learning L2 by watching many videos. However, it is difficult for many L2 learners to smoothly and effortlessly comprehend…

人机交互 · 计算机科学 2025-04-15 Naoto Nishida , Hinako Nozaki , Buntarou Shizuki

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplified workflow where…

计算与语言 · 计算机科学 2025-11-19 Rui Liu , Yuan Zhao , Zhenqi Jia

This innovative practice article reports on the piloting of vibe coding (using natural language to create software applications with AI) for English as a Foreign Language (EFL) education. We developed a human-AI meta-languaging framework…

计算机与社会 · 计算机科学 2025-09-12 David James Woo , Kai Guo , Yangyang Yu

Cross-lingual emotional text-to-speech (TTS) aims to produce speech in one language that captures the emotion of a speaker from another language while maintaining the target voice's timbre. This process of cross-lingual emotional speech…

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber, the first video…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Ngoc-Son Nguyen , Thanh V. T. Tran , Jeongsoo Choi , Hieu-Nghia Huynh-Nguyen , Truong-Son Hy , Van Nguyen

Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Neha Sahipjohn , Ashishkumar Gudmalwar , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Cross-lingual dubbing of lecture videos requires the transcription of the original audio, correction and removal of disfluencies, domain term discovery, text-to-text translation into the target language, chunking of text using target…

‹ 上一页 1 2 3 10 下一页 ›