中文
相关论文

相关论文: ISExplore:Informative Segment Selection for Effici…

200 篇论文

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

计算机视觉与模式识别 · 计算机科学 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

Recent large vision-language models have achieved strong performance on short- and medium-length video understanding, yet they remain inadequate for ultra-long or even infinite video reasoning, where models must preserve coherent memory…

人工智能 · 计算机科学 2026-05-08 Peizheng Yan , Yu Zhao , Liang Xie , Juntong Qi , Mingming Wang , Erwei Yin

State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as a generalist. Personalized speech enhancement could be a…

音频与语音处理 · 电气工程与系统科学 2025-08-06 Tsun-An Hsieh , Minje Kim

Flow-based generative models have greatly improved text-to-speech (TTS) synthesis quality, but inference speed remains limited by the iterative sampling process and multiple function evaluations (NFE). The recent MeanFlow model accelerates…

声音 · 计算机科学 2025-10-10 Wei Wang , Rong Cao , Yi Guo , Zhengyang Chen , Kuan Chen , Yuanyuan Huo

Retrieval-augmented generation from videos requires systems to retrieve relevant audiovisual evidence from large corpora and synthesize it into coherent, attributed text. Current approaches struggle at both ends: retrieval methods fail on…

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

多媒体 · 计算机科学 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Information-seeking conversation systems are increasingly popular in real-world applications, especially for e-commerce companies. To retrieve appropriate responses for users, it is necessary to compute the matching degrees between…

计算与语言 · 计算机科学 2022-11-03 Haojie Pan , Cen Chen , Chengyu Wang , Minghui Qiu , Liu Yang , Feng Ji , Jun Huang

Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listening states…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Lei Zhu , Lijian Lin , Ye Zhu , Jiahao Wu , Xuehan Hou , Yu Li , Yunfei Liu , Jie Chen

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Jin Liu , Xi Wang , Xiaomeng Fu , Yesheng Chai , Cai Yu , Jiao Dai , Jizhong Han

We present a text-based tool for editing talking-head video that enables an iterative editing workflow. On each iteration users can edit the wording of the speech, further refine mouth motions if necessary to reduce artifacts and manipulate…

计算机视觉与模式识别 · 计算机科学 2020-11-24 Xinwei Yao , Ohad Fried , Kayvon Fatahalian , Maneesh Agrawala

Unsupervised keyphrase prediction has gained growing interest in recent years. However, existing methods typically rely on heuristically defined importance scores, which may lead to inaccurate informativeness estimation. In addition, they…

计算与语言 · 计算机科学 2025-06-02 Lam Thanh Do , Aaditya Bodke , Pritom Saha Akash , Kevin Chen-Chuan Chang

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Changpeng Cai , Guinan Guo , Jiao Li , Junhao Su , Fei Shen , Chenghao He , Jing Xiao , Yuanxu Chen , Lei Dai , Feiyu Zhu

Providing accurate predictions is challenging for machine learning algorithms when the number of features is larger than the number of samples in the data. Prior knowledge can improve machine learning models by indicating relevant variables…

Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast scale and high complexity of video data. Current RAG approaches…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Nianbo Zeng , Haowen Hou , Fei Richard Yu , Si Shi , Ying Tiffany He

Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and average facial shape…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Ziyu Yao , Xuxin Cheng , Zhiqi Huang

\Ac{RAG} has emerged as a crucial technique for enhancing large models with real-time and domain-specific knowledge. While numerous improvements and open-source tools have been proposed to refine the \ac{RAG} framework for accuracy,…

信息检索 · 计算机科学 2025-02-20 Yixing Fan , Qiang Yan , Wenshan Wang , Jiafeng Guo , Ruqing Zhang , Xueqi Cheng

Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with predefined speaking…

图形学 · 计算机科学 2024-08-20 Xukun Zhou , Fengxin Li , Ziqiao Peng , Kejian Wu , Jun He , Biao Qin , Zhaoxin Fan , Hongyan Liu

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling on the input…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Weihan Xu , Paul Pu Liang , Haven Kim , Julian McAuley , Taylor Berg-Kirkpatrick , Hao-Wen Dong

Existing deep facial animation coding techniques efficiently compress talking head videos by applying deep generative models. Instead of compressing the entire video sequence, these methods focus on compressing only the keyframe and the…

图像与视频处理 · 电气工程与系统科学 2025-03-14 Riku Takahashi , Ryugo Morita , Fuma Kimishima , Kosuke Iwama , Jinjia Zhou

Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zirui Zhu , Hailun Xu , Yang Luo , Yong Liu , Kanchan Sarkar , Zhenheng Yang , Yang You