中文
相关论文

相关论文: Beyond FVD: Enhanced Evaluation Metrics for Video …

200 篇论文

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Chunyang Cheng , Tianyang Xu , Xiao-Jun Wu , Tao Zhou , Hui Li , Zhangyong Tang , Josef Kittler

We present Stable Video 3D (SV3D) -- a latent video diffusion model for high-resolution, image-to-multi-view generation of orbital videos around a 3D object. Recent work on 3D generation propose techniques to adapt 2D generative models for…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Vikram Voleti , Chun-Han Yao , Mark Boss , Adam Letts , David Pankratz , Dmitry Tochilkin , Christian Laforte , Robin Rombach , Varun Jampani

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Xuyang Shen , Dong Li , Jinxing Zhou , Zhen Qin , Bowen He , Xiaodong Han , Aixuan Li , Yuchao Dai , Lingpeng Kong , Meng Wang , Yu Qiao , Yiran Zhong

Since its introduction, the Virtual Element Method (VEM) was shown to be able to deal with a large variety of polygons, while achieving good convergence rates. The regularity assumptions proposed in the VEM literature to guarantee the…

数值分析 · 数学 2021-02-15 Tommaso Sorgente , Silvia Biasotti , Gianmarco Manzini , Michela Spagnuolo

The past few years have seen impressive progress in the development of deep generative models capable of producing high-dimensional, complex, and photo-realistic data. However, current methods for evaluating such models remain incomplete:…

机器学习 · 计算机科学 2024-03-14 Marco Jiralerspong , Avishek Joey Bose , Ian Gemp , Chongli Qin , Yoram Bachrach , Gauthier Gidel

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xinyue Liang , Zhinyuan Ma , Lingchen Sun , Yanjun Guo , Lei Zhang

Improving the diversity of generated results while maintaining high visual quality remains a significant challenge in image generation tasks. Fractal Generative Models (FGMs) are efficient in generating high-quality images, but their…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Haowei Zhang , Yuanpei Zhao , Ji-Zhe Zhou , Mao Li

Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches operate purely in the 2D pixel space and lack explicit…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Yunpeng Bai , Shaoheng Fang , Chaohui Yu , Fan Wang , Qixing Huang

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Zhengfei Kuang , Shengqu Cai , Hao He , Yinghao Xu , Hongsheng Li , Leonidas Guibas , Gordon Wetzstein

Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-trained Image-to-Video Diffusion Models (I2V-DMs) by…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Liuhan Chen , Xiaodong Cun , Xiaoyu Li , Xianyi He , Shenghai Yuan , Jie Chen , Ying Shan , Li Yuan

Video anomaly detection (VAD) plays a critical role in public safety applications such as intelligent surveillance. However, the rarity, unpredictability, and high annotation cost of real-world anomalies make it difficult to scale VAD…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Suhang Cai , Xiaohao Peng , Chong Wang , Xiaojie Cai , Jiangbo Qian

As video generation models advance rapidly, assessing the quality of generated videos has become increasingly critical. Existing metrics, such as Fr\'echet Video Distance (FVD), Inception Score (IS), and ClipSim, measure quality primarily…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Zihan Wang , Songlin Li , Lingyan Hao , Xinyu Hu , Bowen Song

We present the design of a mesh quality indicator that can predict the behavior of the Virtual Element Method (VEM) on a given mesh family or finite sequence of polyhedral meshes (dataset). The mesh quality indicator is designed to measure…

数值分析 · 数学 2021-12-22 Tommaso Sorgente , Silvia Biasotti , Gianmarco Manzini , Michela Spagnuolo

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Yufan Deng , Zilin Pan , Hongyu Zhang , Xiaojie Li , Ruoqing Hu , Yufei Ding , Yiming Zou , Yan Zeng , Daquan Zhou

In this paper, we propose a new full-reference quality metric for mobile 3D content. Our method is modeled around the Human Visual System, fusing the information of both left and right channels, considering color components, the cyclopean…

图像与视频处理 · 电气工程与系统科学 2018-03-19 Amin Banitalebi-Dehkordi , Mahsa T. Pourazad , Panos Nasiopoulos

Left ventricular ejection fraction (LVEF) is a key indicator of cardiac function and plays a central role in the diagnosis and management of cardiovascular disease. Echocardiography, as a readily accessible and non-invasive imaging…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Shravan Saranyan , Pramit Saha

Diffusion models have demonstrated remarkable performance in image and video synthesis. However, scaling them to high-resolution inputs is challenging and requires restructuring the diffusion pipeline into multiple independent components,…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Ivan Skorokhodov , Willi Menapace , Aliaksandr Siarohin , Sergey Tulyakov

Stereoscopic video technologies have been introduced to the consumer market in the past few years. A key factor in designing a 3D system is to understand how different visual cues and distortions affect the perceptual quality of…

图像与视频处理 · 电气工程与系统科学 2018-03-14 Amin Banitalebi-Dehkordi , Mahsa T. Pourazad , Panos Nasiopoulos

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparameters. Fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Minghan Li , Chenxi Xie , Yichen Wu , Lei Zhang , Mengyu Wang

Video diffusion models have achieved impressive realism and controllability but are limited by high computational demands, restricting their use on mobile devices. This paper introduces the first mobile-optimized video diffusion model.…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Haitam Ben Yahia , Denis Korzhenkov , Ioannis Lelekas , Amir Ghodrati , Amirhossein Habibian