中文
相关论文

相关论文: TalkCuts: A Large-Scale Dataset for Multi-Shot Hum…

200 篇论文

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric, overlooking the…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Peiran Wu , Yunze Liu , Zhengdong Zhu , Enmin Zhou , Junxiao Shen

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirable to design a…

计算与语言 · 计算机科学 2024-08-27 Chien-yu Huang , Min-Han Shih , Ke-Han Lu , Chi-Yuan Hsiao , Hung-yi Lee

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

计算与语言 · 计算机科学 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic…

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

计算与语言 · 计算机科学 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

In recent research on dialogue systems and corpora, there has been a significant focus on two distinct categories: task-oriented (TOD) and open-domain (chit-chat) dialogues. TOD systems aim to satisfy specific user goals, such as finding a…

计算与语言 · 计算机科学 2023-08-29 Wen-Yu Chang , Yun-Nung Chen

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models during dataset…

The one-shot talking-head synthesis task aims to animate a source image to another pose and expression, which is dictated by a driving frame. Recent methods rely on warping the appearance feature extracted from the source, by using motion…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Kangning Liu , Yu-Chuan Su , Wei , Hong , Ruijin Cang , Xuhui Jia

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-language datasets, which…

音频与语音处理 · 电气工程与系统科学 2024-07-22 Xinhao Mei , Chutong Meng , Haohe Liu , Qiuqiang Kong , Tom Ko , Chengqi Zhao , Mark D. Plumbley , Yuexian Zou , Wenwu Wang

Conversational assistants are increasingly popular across diverse real-world applications, highlighting the need for advanced multimodal speech modeling. Speech, as a natural mode of communication, encodes rich user-specific characteristics…

计算与语言 · 计算机科学 2024-12-23 Maximillian Chen , Ruoxi Sun , Sercan Ö. Arık

Recently, multi-modal large language models have made significant progress. However, visual information lacking of guidance from the user's intention may lead to redundant computation and involve unnecessary visual noise, especially in…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Zheng Cheng , Rendong Wang , Zhicheng Wang

Multi-modal large language models are regarded as a crucial step towards Artificial General Intelligence (AGI) and have garnered significant interest with the emergence of ChatGPT. However, current speech-language models typically adopt the…

计算与语言 · 计算机科学 2023-05-22 Dong Zhang , Shimin Li , Xin Zhang , Jun Zhan , Pengyu Wang , Yaqian Zhou , Xipeng Qiu

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

声音 · 计算机科学 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

We consider the challenging problem of audio to animated video generation. We propose a novel method OneShotAu2AV to generate an animated video of arbitrary length using an audio clip and a single unseen image of a person as an input. The…

计算机视觉与模式识别 · 计算机科学 2021-02-22 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall , Mujtaba Hasan , Pranshu Agarwal , Dipankar Sarkar

The objective of this work is person-clustering in videos -- grouping characters according to their identity. Previous methods focus on the narrower task of face-clustering, and for the most part ignore other cues such as the person's…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Andrew Brown , Vicky Kalogeiton , Andrew Zisserman

With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple…

多媒体 · 计算机科学 2022-09-19 Davide Salvi , Brian Hosler , Paolo Bestagini , Matthew C. Stamm , Stefano Tubaro

Multimedia compression allows us to watch videos, see pictures and hear sounds within a limited bandwidth, which helps the flourish of the internet. During the past decades, multimedia compression has achieved great success using hand-craft…

多媒体 · 计算机科学 2023-08-21 Yuhao Cheng , Siru Zhang , Yiqiang Yan , Rong Chen , Yun Zhang

We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer). Trained on 147M conversation-like exchanges extracted from Reddit comment chains over a period spanning…

计算与语言 · 计算机科学 2020-05-05 Yizhe Zhang , Siqi Sun , Michel Galley , Yen-Chun Chen , Chris Brockett , Xiang Gao , Jianfeng Gao , Jingjing Liu , Bill Dolan