中文
相关论文

相关论文: ColoDiff: Integrating Dynamic Consistency With Con…

200 篇论文

Reliable co-speech motion generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically operate on local joint rotations, which are defined hierarchically based…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Xiangyue Zhang , Jianfang Li , Jianqiang Ren , Jiaxu Zhang

Spatial computer vision techniques have the potential to improve the diagnostic performance of colonoscopy. However, the lack of 3D colonoscopy datasets for training and validation hinders their development. This paper introduces C3VDv2,…

图像与视频处理 · 电气工程与系统科学 2025-09-12 Mayank V. Golhar , Lucas Sebastian Galeano Fretes , Loren Ayers , Venkata S. Akshintala , Taylor L. Bobrow , Nicholas J. Durr

Collaborative 3D object detection holds significant importance in the field of autonomous driving, as it greatly enhances the perception capabilities of each individual agent by facilitating information exchange among multiple agents.…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Zhe Huang , Shuo Wang , Yongcai Wang , Lei Wang

Detection and diagnosis of colon polyps are key to preventing colorectal cancer. Recent evidence suggests that AI-based computer-aided detection (CADe) and computer-aided diagnosis (CADx) systems can enhance endoscopists' performance and…

图像与视频处理 · 电气工程与系统科学 2024-03-05 Carlo Biffi , Giulio Antonelli , Sebastian Bernhofer , Cesare Hassan , Daizen Hirata , Mineo Iwatate , Andreas Maieron , Pietro Salvagnini , Andrea Cherubini

Generative models have made remarkable advancements and are capable of producing high-quality content. However, performing controllable editing with generative models remains challenging, due to their inherent uncertainty in outputs. This…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Yikun Ma , Yiqing Li , Jiawei Wu , Xing Luo , Zhi Jin

Low-dose computed tomography (CT) denoising is crucial for reduced radiation exposure while ensuring diagnostically acceptable image quality. Despite significant advancements driven by deep learning (DL) in recent years, existing DL-based…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Zhihao Chen , Qi Gao , Zilong Li , Junping Zhang , Yi Zhang , Jun Zhao , Hongming Shan

Colonoscopy is the choice procedure to diagnose colon and rectum cancer, from early detection of small precancerous lesions (polyps), to confirmation of malign masses. However, the high variability of the organ appearance and the complex…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Josué Ruano , Martín Gómez , Eduardo Romero , Antoine Manzanera

This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation aims to synthesize diverse samples with similar visual content to the source image by…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Yinxi He , Kang Liao , Chunyu Lin , Tianyi Wei , Yao Zhao

Colonoscopy analysis, particularly automatic polyp segmentation and detection, is essential for assisting clinical diagnosis and treatment. However, as medical image annotation is labour- and resource-intensive, the scarcity of annotated…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Yuhao Du , Yuncheng Jiang , Shuangyi Tan , Xusheng Wu , Qi Dou , Zhen Li , Guanbin Li , Xiang Wan

We present a novel task called online video editing, which is designed to edit \textbf{streaming} frames while maintaining temporal consistency. Unlike existing offline video editing assuming all frames are pre-established and accessible,…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Feng Chen , Zhen Yang , Bohan Zhuang , Qi Wu

Medical image segmentation has been significantly advanced with the rapid development of deep learning (DL) techniques. Existing DL-based segmentation models are typically discriminative; i.e., they aim to learn a mapping from the input…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Tao Chen , Chenhui Wang , Zhihao Chen , Yiming Lei , Hongming Shan

The field of generative models has recently witnessed significant progress, with diffusion models showing remarkable performance in image generation. In light of this success, there is a growing interest in exploring the application of…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Ariel Lapid , Idan Achituve , Lior Bracha , Ethan Fetaya

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Yupeng Zhou , Daquan Zhou , Ming-Ming Cheng , Jiashi Feng , Qibin Hou

Diffusion Models have shown remarkable proficiency in image and video synthesis. As model size and latency increase limit user experience, hybrid edge-cloud collaborative framework was recently proposed to realize fast inference and…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Jiajian Xie , Shengyu Zhang , Zhou Zhao , Fan Wu , Fei Wu

Videos can be an effective way to deliver contextualized, just-in-time medical information for patient education. However, video analysis, from topic identification and retrieval to extraction and analysis of medical information and…

图像与视频处理 · 电气工程与系统科学 2024-10-07 Yawen Guo , Xiao Liu , Anjana Susarla , Padman Rema

Video diffusion models have recently made great progress in generation quality, but are still limited by the high memory and computational requirements. This is because current video diffusion models often attempt to process…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Sihyun Yu , Weili Nie , De-An Huang , Boyi Li , Jinwoo Shin , Anima Anandkumar

We introduce the Cross Human Motion Diffusion Model (CrossDiff), a novel approach for generating high-quality human motion based on textual descriptions. Our method integrates 3D and 2D information using a shared transformer network within…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Zeping Ren , Shaoli Huang , Xiu Li

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video,…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Shentong Mo , Zehua Chen , Fan Bao , Jun Zhu

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synthesized images or…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yuanhao Zhai , Kevin Lin , Linjie Li , Chung-Ching Lin , Jianfeng Wang , Zhengyuan Yang , David Doermann , Junsong Yuan , Zicheng Liu , Lijuan Wang

Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like memory and cognition. Among these top-down influences, language…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Yolo Yunlong Tang , Gen Zhan , Li Yang , Yiting Liao , Chenliang Xu