English
Related papers

Related papers: FoleyDirector: Fine-Grained Temporal Steering for …

200 papers

Developing text-driven symbolic music generation models remains challenging due to the scarcity of aligned text-music datasets and the unreliability of automated captioning pipelines. While most efforts have focused on MIDI, sheet music…

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Ilpo Viertola , Vladimir Iashin , Esa Rahtu

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yihao Zhi , Xiaodong Cun , Xuelin Chen , Xi Shen , Wen Guo , Shaoli Huang , Shenghua Gao

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the…

Sound · Computer Science 2024-03-14 Shentong Mo , Jing Shi , Yapeng Tian

In movie productions, the Foley Artist is responsible for creating an overlay soundtrack that helps the movie come alive for the audience. This requires the artist to first identify the sounds that will enhance the experience for the…

Sound · Computer Science 2020-06-29 Sanchita Ghose , John J. Prevost

Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Maksim Siniukov , Di Chang , Minh Tran , Hongkun Gong , Ashutosh Chaubey , Mohammad Soleymani

Recent advancements in text-to-image (T2I) generation using diffusion models have enabled cost-effective video-editing applications by leveraging pre-trained models, eliminating the need for resource-intensive training. However, the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yangfan He , Sida Li , Jianhui Wang , Kun Li , Xinyuan Song , Xinhang Yuan , Keqin Li , Kuan Lu , Menghao Huo , Jingqun Tang , Yi Xin , Jiaqi Chen , Miao Zhang , Xueqian Wang

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Fine-grained emotion recognition (FER) plays a vital role in various fields, such as disease diagnosis, personalized recommendations, and multimedia mining. However, existing FER methods face three key challenges in real-world applications:…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Jingyao Wang , Wenwen Qiang , Changwen Zheng , Fuchun Sun

Existing works have made strides in video generation, but the lack of sound effects (SFX) and background music (BGM) hinders a complete and immersive viewer experience. We introduce a novel semantically consistent v ideo-to-audio generation…

Multimedia · Computer Science 2024-04-29 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Kaixin Shen , Ruijie Quan , Linchao Zhu , Jun Xiao , Yi Yang

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline -…

Sound · Computer Science 2025-07-24 Tobias Morocutti , Jonathan Greif , Paul Primus , Florian Schmid , Gerhard Widmer

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model…

Visual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes. Existing autoregressive methods, which rely on previous frame-sentence pairs, struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Sixiao Zheng , Yanwei Fu

Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. While diffusion models have shown promise in this text-to-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Wenshuo Chen , Haozhe Jia , Songning Lai , Lei Wang , Yuqi Lin , Hongru Xiao , Lijie Hu , Yutao Yue

Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Guangzhao Li , Yanming Yang , Chenxi Song , Chi Zhang

Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yuji Wang , Moran Li , Xiaobin Hu , Ran Yi , Jiangning Zhang , Han Feng , Weijian Cao , Yabiao Wang , Chengjie Wang , Lizhuang Ma

Text-to-Time Series generation holds significant potential to address challenges such as data sparsity, imbalance, and limited availability of multimodal time series datasets across domains. While diffusion models have achieved remarkable…

Machine Learning · Computer Science 2025-05-09 Yunfeng Ge , Jiawei Li , Yiji Zhao , Haomin Wen , Zhao Li , Meikang Qiu , Hongyan Li , Ming Jin , Shirui Pan

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zijun Ding , Mingdie Xiong , Congcong Zhu , Jingrun Chen

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational…

Sound · Computer Science 2023-05-04 Jinlong Xue , Yayue Deng , Fengping Wang , Ya Li , Yingming Gao , Jianhua Tao , Jianqing Sun , Jiaen Liang
‹ Prev 1 4 5 6 7 8 10 Next ›