中文
相关论文

相关论文: DAGE: Dual-Stream Architecture for Efficient and F…

200 篇论文

360-degree visual content is widely shared on platforms such as YouTube and plays a central role in virtual reality, robotics, and autonomous navigation. However, consumer-grade dual-fisheye systems consistently yield imperfect panoramas…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Changha Shin , Woong Oh Cho , Seon Joo Kim

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zihui Gao , Ke Liu , Donny Y. Chen , Duochao Shi , Guosheng Lin , Hao Chen , Chunhua Shen

The proliferation of sophisticated AI-generated deepfakes poses critical challenges for digital media authentication and societal security. While existing detection methods perform well within specific generative domains, they exhibit…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Naseem Khan , Tuan Nguyen , Amine Bermak , Issa Khalil

Diffusion models are known for generating high-quality images, causing serious security concerns. To combat this, most efforts rely on deep neural networks (e.g., CNNs and Transformers), while largely overlooking the potential of…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Mengxin Fu , Yuezun Li

Student engagement is crucial for improving learning outcomes in group activities. Highly engaged students perform better both individually and contribute to overall group success. However, most existing automated engagement recognition…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Saniah Kayenat Chowdhury , Muhammad E. H. Chowdhury

A dramatic influx of diffusion-generated images has marked recent years, posing unique challenges to current detection technologies. While the task of identifying these images falls under binary classification, a seemingly straightforward…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Yewon Lim , Changyeon Lee , Aerin Kim , Oren Etzioni

We present FLARE, a feed-forward model designed to infer high-quality camera poses and 3D geometry from uncalibrated sparse-view images (i.e., as few as 2-8 inputs), which is a challenging yet practical setting in real-world applications.…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Shangzhan Zhang , Jianyuan Wang , Yinghao Xu , Nan Xue , Christian Rupprecht , Xiaowei Zhou , Yujun Shen , Gordon Wetzstein

Segmentation-oriented Industrial Anomaly Synthesis (SIAS) plays a pivotal role in enhancing the performance of downstream anomaly segmentation, as it provides an effective means of expanding abnormal data. However, existing SIAS methods…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Xichen Xu , Yanshu Wang , Jinbao Wang , Qunyi Zhang , Xiaoning Lei , Guoyang Xie , Guannan Jiang , Zhichao Lu

Training modern neural networks on large datasets is computationally and energy intensive. We present SAGE, a streaming data-subset selection method that maintains a compact Frequent Directions (FD) sketch of gradient geometry in $O(\ell…

机器学习 · 计算机科学 2025-10-10 Ashish Jha , Salman Ahmadi-Asl

Despite significant progress has been made in image deraining, we note that most existing methods are often developed for only specific types of rain degradation and fail to generalize across diverse real-world rainy scenes. How to…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Qianfeng Yang , Qiyuan Guan , Xiang Chen , Jiyu Jin , Guiyue Jin , Jiangxin Dong

In recent years, consumer-level depth cameras have been adopted for various applications. However, they often produce depth maps at only a moderately high frame rate (approximately 30 frames per second), preventing them from being used for…

图形学 · 计算机科学 2018-11-06 Ming-Ze Yuan , Lin Gao , Hongbo Fu , Shihong Xia

We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Shuo Sun , Unal Artan , Malcolm Mielle , Achim J. Lilienthaland , Martin Magnusson

We introduce RAGE, an image compression framework that achieves four generally conflicting objectives: 1) good compression for a wide variety of color images, 2) computationally efficient, fast decompression, 3) fast random access of images…

图像与视频处理 · 电气工程与系统科学 2024-02-12 Christian D. Rask , Daniel E. Lucani

Transformer-based object detectors often struggle with occlusions, fine-grained localization, and computational inefficiency caused by fixed queries and dense attention. We propose DAMM, Dual-stream Attention with Multi-Modal queries, a…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Noreen Anwar , Guillaume-Alexandre Bilodeau , Wassim Bouachir

Convolutional Neural Networks have demonstrated superior performance on single image depth estimation in recent years. These works usually use stacked spatial pooling or strided convolution to get high-level information which are common…

计算机视觉与模式识别 · 计算机科学 2018-09-05 Zhixiang Hao , Yu Li , Shaodi You , Feng Lu

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Ying Zang , Yidong Han , Chaotao Ding , Yuanqi Hu , Deyi Ji , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu

Unsupervised image-to-image translation aims at learning a mapping between two visual domains. However, learning a translation across large geometry variations always ends up with failure. In this work, we present a novel…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Wayne Wu , Kaidi Cao , Cheng Li , Chen Qian , Chen Change Loy

Medical image understanding requires meticulous examination of fine visual details, with particular regions requiring additional attention. While radiologists build such expertise over years of experience, it is challenging for AI models to…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Ying Jin , Zhuoran Zhou , Haoquan Fang , Jenq-Neng Hwang

Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Yamei Chen , Yan Di , Guangyao Zhai , Fabian Manhardt , Chenyangguang Zhang , Ruida Zhang , Federico Tombari , Nassir Navab , Benjamin Busam

Image animation is the task of transferring the motion of a driving video to a given object in a source image. While great progress has recently been made in unsupervised motion transfer, requiring no labeled data or domain priors, many…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Peirong Liu , Rui Wang , Xuefei Cao , Yipin Zhou , Ashish Shah , Ser-Nam Lim