English
Related papers

Related papers: Solving Spatial Supersensing Without Spatial Super…

200 papers

Semantic Scene Completion (SSC) transforms an image of single-view depth and/or RGB 2D pixels into 3D voxels, each of whose semantic labels are predicted. SSC is a well-known ill-posed problem as the prediction model has to "imagine" what…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Fengyun Wang , Dong Zhang , Hanwang Zhang , Jinhui Tang , Qianru Sun

Ensuring reliable confidence scores from deep networks is of pivotal importance in critical decision-making systems, notably in the medical domain. While recent literature on calibrating deep segmentation networks has led to significant…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Balamurali Murugesan , Sukesh Adiga , Bingyuan Liu , Hervé Lombaert , Ismail Ben Ayed , Jose Dolz

Despite the large progress in supervised learning with neural networks, there are significant challenges in obtaining high-quality, large-scale and accurately labelled datasets. In such a context, how to learn in the presence of noisy…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Chen Feng , Georgios Tzimiropoulos , Ioannis Patras

We investigate the problem of recovering a structured sparse signal from a linear observation model with an uncertain dynamic grid in the sensing matrix. The state-of-the-art expectation maximization based compressed sensing (EM-CS)…

Signal Processing · Electrical Eng. & Systems 2024-07-25 An Liu , Yufan Zhou , Wenkang Xu

We address the problem of video representation learning without human-annotated labels. While previous efforts address the problem by designing novel self-supervised tasks using video data, the learned features are merely on a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Yunhui Liu , Wei Liu

Single Image Super-Resolution (SISR) reconstructs high-resolution images from low-resolution inputs, enhancing image details. While Vision Transformer (ViT)-based models improve SISR by capturing long-range dependencies, they suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Junyoung Kim , Youngrok Kim , Siyeol Jung , Donghyun Min

Causal reasoning is often challenging with spatial data, particularly when handling high-dimensional inputs. To address this, we propose a neural network (NN) based framework integrated with an approximate Gaussian process to manage spatial…

Machine Learning · Computer Science 2024-12-06 Ziyang Jiang , Zach Calhoun , Yiling Liu , Lei Duan , David Carlson

Recent single-image super-resolution (SISR) networks, which can adapt their network parameters to specific input images, have shown promising results by exploiting the information available within the input data as well as large external…

Computer Vision and Pattern Recognition · Computer Science 2021-03-19 Jinsu Yoo , Tae Hyun Kim

Supervised learning techniques are at the center of many tasks in remote sensing. Unfortunately, these methods, especially recent deep learning methods, often require large amounts of labeled data for training. Even though satellites…

Machine Learning · Computer Science 2021-08-03 Pablo Gómez , Gabriele Meoni

Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness (recognizing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Pascal Benschop , Justin Dauwels , Jan van Gemert

Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpected artifacts, but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiongze Yu , Xiangbo Gao , Pooja Verlani , Akshay Gadde , Yilin Wang , Balu Adsumilli , Zhengzhong Tu

Labeled data is a critical resource for training and evaluating machine learning models. However, many real-life datasets are only partially labeled. We propose a semi-supervised machine learning training strategy to improve event detection…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Florian Dubost , Erin Hong , Nandita Bhaskhar , Siyi Tang , Daniel Rubin , Christopher Lee-Messer

Scarcity of pixel-level labels is a significant challenge in practical scenarios. In specific domains like industrial smoke, acquiring such detailed annotations is particularly difficult and often requires expert knowledge. To alleviate…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Zheyuan Zhang , Yen-chia Hsu

Traditional regression models assume stationary relationships between predictors and responses, failing to capture the spatial heterogeneity present in many environmental, epidemiological, and ecological processes. To address this…

Methodology · Statistics 2025-05-27 Justice Akuoko-Frimpong , Edward Shao , Jonathan Ta

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

Machine Learning · Computer Science 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,''…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shaoxiong Zhan , Yanlin Lai , Zheng Liu , Hai Lin , Shen Li , Xiaodong Cai , Zijian Lin , Wen Huang , Hai-Tao Zheng

Compared with image scene parsing, video scene parsing introduces temporal information, which can effectively improve the consistency and accuracy of prediction. In this paper, we propose a Spatial-Temporal Semantic Consistency method to…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Xingjian He , Weining Wang , Zhiyong Xu , Hao Wang , Jie Jiang , Jing Liu

3D super-resolution aims to reconstruct high-fidelity 3D models from low-resolution (LR) multi-view images. Early studies primarily focused on single-image super-resolution (SISR) models to upsample LR images into high-resolution images.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Hyun-kyu Ko , Dongheok Park , Youngin Park , Byeonghyeon Lee , Juhee Han , Eunbyung Park

Recent advancements in real-time super-resolution have enabled higher-quality video streaming, yet existing methods struggle with the unique challenges of compressed video content. Commonly used datasets do not accurately reflect the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Evgeney Bogatyrev , Khaled Abud , Ivan Molodetskikh , Nikita Alutis , Dmitriy Vatolin

While multimodal large language models (MLLMs) exhibit strong performance on single-video tasks (e.g., video question answering), their capability for spatiotemporal pattern reasoning across multiple videos remains a critical gap in pattern…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Nannan Zhu , Yonghao Dong , Teng Wang , Xueqian Li , Shengjun Deng , Yijia Wang , Zheng Hong , Tiantian Geng , Guo Niu , Hanyan Huang , Xiongfei Yao , Shuaiwei Jiao
‹ Prev 1 3 4 5 6 7 10 Next ›