中文
相关论文

相关论文: Self-Paced and Self-Corrective Masked Prediction f…

200 篇论文

Movie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To streamline this process, we propose an automatic trailer generation…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Dawit Mureja Argaw , Mattia Soldan , Alejandro Pardo , Chen Zhao , Fabian Caba Heilbron , Joon Son Chung , Bernard Ghanem

Movie trailers perform multiple functions: they introduce viewers to the story, convey the mood and artistic style of the film, and encourage audiences to see the movie. These diverse functions make trailer creation a challenging endeavor.…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Pinelopi Papalampidi , Frank Keller , Mirella Lapata

Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which may lead to error…

Accurate temporal prediction is the bridge between comprehensive scene understanding and embodied artificial intelligence. However, predicting multiple fine-grained states of a scene at multiple temporal scales is difficult for…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Zhitao Zeng , Guojian Yuan , Junyuan Mao , Yuxuan Wang , Xiaoshuang Jia , Yueming Jin

Trailers are short promotional videos designed to provide audiences with a glimpse of a movie. The process of creating a trailer typically involves selecting key scenes, dialogues and action sequences from the main content and editing them…

多媒体 · 计算机科学 2026-02-02 Roberto Balestri , Pasquale Cascarano , Mirko Degli Esposti , Guglielmo Pescatore

As the pretraining technique is growing in popularity, little work has been done on pretrained learning-based motion prediction methods in autonomous driving. In this paper, we propose a framework to formalize the pretraining task for…

机器人学 · 计算机科学 2023-09-19 Yi Yang , Qingwen Zhang , Thomas Gilles , Nazre Batool , John Folkesson

We introduce a novel method for movie genre classification, capitalizing on a diverse set of readily accessible pretrained models. These models extract high-level features related to visual scenery, objects, characters, text, speech, music,…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Serkan Sulun , Paula Viana , Matthew E. P. Davies

Current video generation models usually convert signals indicating appearance and motion received from inputs (e.g., image, text) or latent spaces (e.g., noise vectors) into consecutive frames, fulfilling a stochastic generation process for…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Xue Song , Jingjing Chen , Bin Zhu , Yu-Gang Jiang

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP's pretraining on…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhe Li , Weihao Yuan , Yisheng He , Lingteng Qiu , Shenhao Zhu , Xiaodong Gu , Weichao Shen , Yuan Dong , Zilong Dong , Laurence T. Yang

Human motion prediction model has applications in various fields of computer vision. Without taking into account the inherent stochasticity in the prediction of future pose dynamics, such methods often converges to a deterministic undesired…

计算机视觉与模式识别 · 计算机科学 2018-12-07 Jogendra Nath Kundu , Maharshi Gor , R. Venkatesh Babu

The domain of automatic video trailer generation is currently undergoing a profound paradigm shift, transitioning from heuristic-based extraction methods to deep generative synthesis. While early methodologies relied heavily on low-level…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Abhishek Dharmaratnakar , Srivaths Ranganathan , Debanshu Das , Anushree Sinha

The performance of pre-trained masked diffusion models is often constrained by their sampling procedure, which makes decisions irreversible and struggles in low-step generation regimes. We introduce a novel sampling algorithm that works…

Generative models are increasingly being explored in click-through rate (CTR) prediction field to overcome the limitations of the conventional discriminative paradigm, which rely on a simple binary classification objective. However,…

信息检索 · 计算机科学 2025-11-19 Moyu Zhang , Yujun Jin , Yun Chen , Jinxin Hu , Yu Zhang , Xiaoyi Zeng

This paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important…

计算与语言 · 计算机科学 2021-02-17 Mana Ihori , Naoki Makishima , Tomohiro Tanaka , Akihiko Takashima , Shota Orihashi , Ryo Masumura

Recently, text-to-image (T2I) synthesis has undergone significant advancements, particularly with the emergence of Large Language Models (LLM) and their enhancement in Large Vision Models (LVM), greatly enhancing the instruction-following…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Weijin Cheng , Jianzhi Liu , Jiawen Deng , Fuji Ren

Trajectory prediction is an important task that involves modeling the indeterminate nature of traffic actors to forecast future trajectories given the observed trajectory sequences. However, current methods confine themselves to presumed…

机器人学 · 计算机科学 2023-12-18 Pranav Singh Chib , Pravendra Singh

Masked autoregressive models (MAR) have emerged as a powerful paradigm for image and video generation, combining the flexibility of masked modeling with the expressiveness of continuous tokenizers. However, when sampling individual frames,…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Zian Li , Muhan Zhang

This paper addresses the problem of self-supervised video representation learning from a new perspective -- by video pace prediction. It stems from the observation that human visual system is sensitive to video pace, e.g., slow motion, a…

计算机视觉与模式识别 · 计算机科学 2020-09-07 Jiangliu Wang , Jianbo Jiao , Yun-Hui Liu

Successful video analysis relies on accurate recognition of pixels across frames, and frame reconstruction methods based on video correspondence learning are popular due to their efficiency. Existing frame reconstruction methods, while…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Zihan Zhou , Changrui Dai , Aibo Song , Xiaolin Fang

Video-based movie genre classification has garnered considerable attention due to its various applications in recommendation systems. Prior work has typically addressed this task by adapting models from traditional video classification…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Zhongping Zhang , Yiwen Gu , Bryan A. Plummer , Xin Miao , Jiayi Liu , Huayan Wang
‹ 上一页 1 2 3 10 下一页 ›