中文
相关论文

相关论文: Fusion Complexity Inversion: Why Simpler Cross Vie…

200 篇论文

We study rotation-robust learning for image inputs using Convolutional Model Trees (CMTs) [1], whose split and leaf coefficients can be structured on the image grid and transformed geometrically at deployment time. In a controlled MNIST…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Hongyi Li , William Ward Armstrong , Jun Xu

The amount of available Earth observation data has increased dramatically in the recent years. Efficiently making use of the entire body information is a current challenge in remote sensing and demands for light-weight problem-agnostic…

机器学习 · 计算机科学 2020-10-26 Marc Rußwurm , Marco Körner

Text-guided image generation and editing using diffusion models have achieved remarkable advancements. Among these, tuning-free methods have gained attention for their ability to perform edits without extensive model adjustments, offering…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Wenyi Mo , Tianyu Zhang , Yalong Bai , Bing Su , Ji-Rong Wen

Cancer survival prediction from whole slide images (WSIs) is a challenging task in computational pathology due to the large size, irregular shape, and high granularity of the WSIs. These characteristics make it difficult to capture the full…

图像与视频处理 · 电气工程与系统科学 2025-03-05 Rustin Soraki , Huayu Wang , Joann G. Elmore , Linda Shapiro

Early and accurate interpretation of screening mammograms is essential for effective breast cancer detection, yet it remains a complex challenge due to subtle imaging findings and diagnostic ambiguity. Many existing AI approaches fall short…

图像与视频处理 · 电气工程与系统科学 2025-07-24 Yalda Zafari , Roaa Elalfy , Mohamed Mabrok , Somaya Al-Maadeed , Tamer Khattab , Essam A. Rashed

Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tuning incurs prohibitive computational and storage overhead. Parameter-Efficient…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Lingyu Xiong , Jinjin Shi , Xuran Xu , Cong Luo , Runyu Shi , Ying Huang

Monocular depth estimation from a single RGB image remains a fundamental challenge in computer vision due to inherent scale ambiguity and the absence of explicit geometric cues. Existing approaches typically rely on increasingly complex…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Wuqi Su , Huilun Song , Chen Zhao , Chi Xu

Current few-shot learning models capture visual object relations in the so-called meta-learning setting under a fixed-resolution input. However, such models have a limited generalization ability under the scale and location mismatch between…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Hongguang Zhang , Philip H. S. Torr , Piotr Koniusz

Operational phase unwrapping is the primary computational bottleneck in InSAR-based volcanic and seismic monitoring. We challenge the industry trend of adopting high-complexity computer vision architectures, such as attention mechanisms,…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Prabhjot Singh , Manmeet Singh

Unsupervised pre-training has emerged as a transformative paradigm, displaying remarkable advancements in various domains. However, the susceptibility to domain shift, where pre-training data distribution differs from fine-tuning, poses a…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Abhiroop Talasila , Maitreya Maity , U. Deva Priyakumar

Machine-learning algorithms have gained popularity in recent years in the field of ecological modeling due to their promising results in predictive performance of classification problems. While the application of such algorithms has been…

Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic…

机器学习 · 计算机科学 2026-03-20 Youjin Wang , Jiaqiao Zhao , Rong Fu , Run Zhou , Ruizhe Zhang , Jiani Liang , Suisuai Cao , Feng Zhou

This study introduces RicEns-Net, a novel Deep Ensemble model designed to predict crop yields by integrating diverse data sources through multimodal data fusion techniques. The research focuses specifically on the use of synthetic aperture…

图像与视频处理 · 电气工程与系统科学 2025-02-11 Akshay Dagadu Yewle , Laman Mirzayeva , Oktay Karakuş

Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain. We investigate the failure of 3D point cloud encoders to achieve similar generalization…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Chun-Peng Chang , Shaoxiang Wang , Alain Pagani , Dariu Gavrila , Holger Caesar

Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervised segmentation. To bridge foundational computer vision with practical metallurgical…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Abdul Mueez , Shruti Vyas

Window-based transformers have demonstrated outstanding performance in super-resolution tasks due to their adaptive modeling capabilities through local self-attention (SA). However, they exhibit higher computational complexity and inference…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Zhenyu Hu , Wanjie Sun

Few-shot semantic segmentation has attracted growing interest for its ability to generalize to novel object categories using only a few annotated samples. To address data scarcity, recent methods incorporate multiple foundation models to…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Wei Zhuo , Zhiyue Tang , Wufeng Xue , Hao Ding , Junkai Ji , Linlin Shen

Transformers have become one of the dominant architectures in deep learning, particularly as a powerful alternative to convolutional neural networks (CNNs) in computer vision. However, Transformer training and inference in previous works…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Zizheng Pan , Bohan Zhuang , Haoyu He , Jing Liu , Jianfei Cai

Assessment of forest biodiversity is crucial for ecosystem management and conservation. While traditional field surveys provide high-quality assessments, they are labor-intensive and spatially limited. This study investigates whether deep…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Simon B. Jensen , Stefan Oehmcke , Andreas Møgelmose , Meysam Madadi , Christian Igel , Sergio Escalera , Thomas B. Moeslund

Camera-LiDAR fusion models significantly enhance perception performance in autonomous driving. The fusion mechanism leverages the strengths of each modality while minimizing their weaknesses. Moreover, in practice, camera-LiDAR fusion…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Shiqi Sun , Yantao Lu , Ning Liu , Bo Jiang , JinChao Chen , Ying Zhang