English
Related papers

Related papers: PuzzleGPT: Emulating Human Puzzle-Solving Ability …

200 papers

Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human…

Human pose estimation aims to figure out the keypoints of all people in different scenes. Current approaches still face some challenges despite promising results. Existing top-down methods deal with a single person individually, without the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Shuaitao Zhao , Kun Liu , Yuhang Huang , Qian Bao , Dan Zeng , Wu Liu

Vision-language models (VLM) have demonstrated impressive performance in image recognition by leveraging self-supervised training on large datasets. Their performance can be further improved by adapting to the test sample using test-time…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ramya Hebbalaguppe , Tamoghno Kandar , Abhinav Nagpal , Chetan Arora

Visual place recognition is the task of recognizing a place depicted in an image based on its pure visual appearance without metadata. In visual place recognition, the challenges lie upon not only the changes in lighting conditions, camera…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Duc Canh Le , Chan Hyun Youn

Recent studies have highlighted the limitations of large language models in mathematical reasoning, particularly their inability to capture the underlying logic. Inspired by meta-learning, we propose that models should acquire not only…

Computation and Language · Computer Science 2024-12-19 Kejie Chen , Lin Wang , Qinghai Zhang , Renjun Xu

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Konstantin Klemmer , Esther Rolf , Caleb Robinson , Lester Mackey , Marc Rußwurm

Visual place recognition (VPR) is a highly challenging task that has a wide range of applications, including robot navigation and self-driving vehicles. VPR is particularly difficult due to the presence of duplicate regions and the lack of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Yifan Xu , Pourya Shamsolmoali , Jie Yang

Existing volumetric methods for predicting 3D human pose estimation are accurate, but computationally expensive and optimized for single time-step prediction. We present TEMPO, an efficient multi-view pose estimation model that learns a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Rohan Choudhury , Kris Kitani , Laszlo A. Jeni

This paper presents a method for improving any object tracking algorithm based on machine learning. During the training phase, important trajectory features are extracted which are then used to calculate a confidence value of trajectory.…

Computer Vision and Pattern Recognition · Computer Science 2010-07-21 Duc Phu Chau , Francois Bremond , Etienne Corvee , Monique Thonnat

The application of machine learning (ML) in a range of geospatial tasks is increasingly common but often relies on globally available covariates such as satellite imagery that can either be expensive or lack predictive power. Here we…

Computation and Language · Computer Science 2024-02-27 Rohin Manvi , Samar Khanna , Gengchen Mai , Marshall Burke , David Lobell , Stefano Ermon

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 An-Chieh Cheng , Hongxu Yin , Yang Fu , Qiushan Guo , Ruihan Yang , Jan Kautz , Xiaolong Wang , Sifei Liu

Can we better anticipate an actor's future actions (e.g. mix eggs) by knowing what commonly happens after his/her current action (e.g. crack eggs)? What if we also know the longer-term goal of the actor (e.g. making egg fried rice)? The…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Qi Zhao , Shijie Wang , Ce Zhang , Changcheng Fu , Minh Quan Do , Nakul Agarwal , Kwonjoon Lee , Chen Sun

Vision-Language Models like CLIP create aligned embedding spaces for text and images, making it possible for anyone to build a visual classifier by simply naming the classes they want to distinguish. However, a model that works well in one…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Kevin Robbins , Xiaotong Liu , Yu Wu , Le Sun , Grady McPeak , Abby Stylianou , Robert Pless

Neural models such as YOLO and HuBERT can be used to detect local properties such as objects ("car") and emotions ("angry") in individual frames of videos and audio clips respectively. The likelihood of these detections is indicated by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Avishree Khare , Hideki Okamoto , Bardh Hoxha , Georgios Fainekos , Rajeev Alur

Large language models (LLMs) can be used as accessible and intelligent chatbots by constructing natural language queries and directly inputting the prompt into the large language model. However, different prompt' constructions often lead to…

Computation and Language · Computer Science 2023-12-14 Jinta Weng , Jiarui Zhang , Yue Hu , Daidong Fa , Xiaofeng Xuand , Heyan Huang

Reasoning segmentation is a challenging vision-language task that aims to output the segmentation mask with respect to a complex, implicit, and even non-visual query text. Previous works incorporated multimodal Large Language Models (MLLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Shiu-hong Kao , Yu-Wing Tai , Chi-Keung Tang

Human visual attention is a complex phenomenon that has been studied for decades. Within it, the particular problem of scanpath prediction poses a challenge, particularly due to the inter- and intra-observer variability, among other…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Daniel Martin , Diego Gutierrez , Belen Masia

While grasps must satisfy the grasping stability criteria, good grasps depend on the specific manipulation scenario: the object, its properties and functionalities, as well as the task and grasp constraints. In this paper, we consider such…

Large-scale pre-trained models have shown promising open-world performance for both vision and language tasks. However, their transferred capacity on 3D point clouds is still limited and only constrained to the classification task. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Xiangyang Zhu , Renrui Zhang , Bowei He , Ziyu Guo , Ziyao Zeng , Zipeng Qin , Shanghang Zhang , Peng Gao

This paper introduces ClimateGPT, a model family of domain-specific large language models that synthesize interdisciplinary research on climate change. We trained two 7B models from scratch on a science-oriented dataset of 300B tokens. For…

‹ Prev 1 8 9 10 Next ›