中文
相关论文

相关论文: TalkCuts: A Large-Scale Dataset for Multi-Shot Hum…

200 篇论文

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tuning on videos.…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Yan Wang , Yawen Zeng , Jingsheng Zheng , Xiaofen Xing , Jin Xu , Xiangmin Xu

Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken…

计算与语言 · 计算机科学 2025-06-25 Shuzheng Si , Wentao Ma , Haoyu Gao , Yuchuan Wu , Ting-En Lin , Yinpei Dai , Hangyu Li , Rui Yan , Fei Huang , Yongbin Li

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where,…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Mathew Monfort , SouYoung Jin , Alexander Liu , David Harwath , Rogerio Feris , James Glass , Aude Oliva

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously,…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yingjie Zhou , Xilei Zhu , Siyu Ren , Ziyi Zhao , Ziwen Wang , Farong Wen , Yu Zhou , Jiezhang Cao , Xiongkuo Min , Fengjiao Chen , Xiaoyu Li , Xuezhi Cao , Guangtao Zhai , Xiaohong Liu

Achieving realistic, vivid, and human-like synthesized conversational gestures conditioned on multi-modal data is still an unsolved problem due to the lack of available datasets, models and standard evaluation metrics. To address this, we…

计算机视觉与模式识别 · 计算机科学 2022-09-21 Haiyang Liu , Zihao Zhu , Naoya Iwamoto , Yichen Peng , Zhengqing Li , You Zhou , Elif Bozkurt , Bo Zheng

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is urgently needed to…

多媒体 · 计算机科学 2024-08-28 Zeyu Jin , Jia Jia , Qixin Wang , Kehan Li , Shuoyi Zhou , Songtao Zhou , Xiaoyu Qin , Zhiyong Wu

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

Transcripts of teaching episodes can be effective tools to understand discourse patterns in classroom instruction. According to most educational experts, sustained classroom discourse is a critical component of equitable, engaging, and rich…

计算与语言 · 计算机科学 2022-04-21 Abhijit Suresh , Jennifer Jacobs , Charis Harty , Margaret Perkoff , James H. Martin , Tamara Sumner

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Youxin Pang , Xiaochen Zhao , Chao Xu , Lizhen Wang , Hongjiang Xiao , Shi Yan , Hongwen Zhang , Yebin Liu

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

People spend a substantial portion of their lives engaged in conversation, and yet our scientific understanding of conversation is still in its infancy. In this report we advance an interdisciplinary science of conversation, with findings…

Spoken dialogue is essential for human-AI interactions, providing expressive capabilities beyond text. Developing effective spoken dialogue systems (SDSs) requires large-scale, high-quality, and diverse spoken dialogue corpora. However,…

计算与语言 · 计算机科学 2026-04-03 Wataru Nakata , Kentaro Seki , Hitomi Yanaka , Yuki Saito , Shinnosuke Takamichi , Hiroshi Saruwatari

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Yifeng Ma , Suzhen Wang , Yu Ding , Bowen Ma , Tangjie Lv , Changjie Fan , Zhipeng Hu , Zhidong Deng , Xin Yu

Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this paper, we open-source a…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Babak Naderi , Ross Cutler

Robust task-oriented spoken dialogue agents require exposure to the full diversity of how people interact through speech. Building spoken user simulators that address this requires large-scale spoken task-oriented dialogue (TOD) data…

计算与语言 · 计算机科学 2026-03-18 Jonggeun Lee , Junseong Pyo , Jeongmin Park , Yohan Jo

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

人工智能 · 计算机科学 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency