English
Related papers

Related papers: PhysLab: A Benchmark Dataset for Multi-Granularity…

200 papers

As CLIP's global alignment limits its ability to capture fine-grained details, recent efforts have focused on enhancing its region-text alignment. However, current remote sensing (RS)-specific CLIP variants still inherit this limited…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Zhenshi Li , Weikang Yu , Dilxat Muhtar , Xueliang Zhang , Pengfeng Xiao , Pedram Ghamisi , Xiao Xiang Zhu

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

Computation and Language · Computer Science 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Wenliang Dai , Junnan Li , Dongxu Li , Anthony Meng Huat Tiong , Junqi Zhao , Weisheng Wang , Boyang Li , Pascale Fung , Steven Hoi

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xiangting Meng , Jiaqi Yang , Mingshu Chen , Chenxin Yan , Yujiao Shi , Wenchao Ding , Laurent Kneip

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and…

Robotics · Computer Science 2024-09-27 Nghia Nguyen , Minh Nhat Vu , Tung D. Ta , Baoru Huang , Thieu Vo , Ngan Le , Anh Nguyen

As a pioneering vision-language model, CLIP (Contrastive Language-Image Pre-training) has achieved significant success across various domains and a wide range of downstream vision-language tasks. However, the text encoders in popular CLIP…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Mothilal Asokan , Kebin Wu , Fatima Albreiki

Interactive and spatially aware technologies are transforming educational frameworks, particularly in K-12 settings where hands-on exploration fosters deeper conceptual understanding. However, during collaborative tasks, existing systems…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Changsoo Jung , Sheikh Mannan , Jack Fitzgerald , Nathaniel Blanchard

Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-06 Xun Long Ng , Kian Eng Ong , Qichen Zheng , Yun Ni , Si Yong Yeo , Jun Liu

Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from text inputs emerges as a promising solution that can unlock…

Graphics · Computer Science 2026-01-01 Luca Collorone , Mert Kiray , Indro Spinelli , Fabio Galasso , Benjamin Busam

Understanding objects through multiple sensory modalities is fundamental to human perception, enabling cross-sensory integration and richer comprehension. For AI and robotic systems to replicate this ability, access to diverse, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Samuel Clarke , Suzannah Wistreich , Yanjie Ze , Jiajun Wu

We introduce \textbf{LongInsightBench}, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating \textbf{visual,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 ZhaoYang Han , Qihan Lin , Hao Liang , Bowen Chen , Zhou Liu , Wentao Zhang

Large datasets are the cornerstone of recent advances in computer vision using deep learning. In contrast, existing human motion capture (mocap) datasets are small and the motions limited, hampering progress on learning models of human…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Naureen Mahmood , Nima Ghorbani , Nikolaus F. Troje , Gerard Pons-Moll , Michael J. Black

Despite advances in physics-based 3D motion synthesis, current methods face key limitations: reliance on pre-reconstructed 3D Gaussian Splatting (3DGS) built from dense multi-view images with time-consuming per-scene optimization; physics…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Chunji Lv , Zequn Chen , Donglin Di , Weinan Zhang , Hao Li , Wei Chen , Yinjie Lei , Changsheng Li

Visual representation learning hold great promise for robotics, but is severely hampered by the scarcity and homogeneity of robotics datasets. Recent works address this problem by pre-training visual representations on large-scale but…

Robotics · Computer Science 2023-10-16 Sudeep Dasari , Mohan Kumar Srirama , Unnat Jain , Abhinav Gupta

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Meiqi Sun , Zhonghan Zhao , Wenhao Chai , Hanjun Luo , Shidong Cao , Yanting Zhang , Jenq-Neng Hwang , Gaoang Wang

We propose a new long video dataset (called Track Long and Prosper - TLP) and benchmark for single object tracking. The dataset consists of 50 HD videos from real world scenarios, encompassing a duration of over 400 minutes (676K frames),…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Abhinav Moudgil , Vineet Gandhi

Cognitive science has shown that humans perceive videos in terms of events separated by the state changes of dominant subjects. State changes trigger new events and are one of the most useful among the large amount of redundant information…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Yuxuan Wang , Difei Gao , Licheng Yu , Stan Weixian Lei , Matt Feiszli , Mike Zheng Shou

We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for robotic imitation learning. Current datasets, such as Ego4D,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Hongxiang Zhao , Xingchen Liu , Mutian Xu , Yiming Hao , Weikai Chen , Xiaoguang Han

Manipulating garments and fabrics has long been a critical endeavor in the development of home-assistant robots. However, due to complex dynamics and topological structures, garment manipulations pose significant challenges. Recent…

Robotics · Computer Science 2024-12-24 Haoran Lu , Ruihai Wu , Yitong Li , Sijie Li , Ziyu Zhu , Chuanruo Ning , Yan Shen , Longzan Luo , Yuanpei Chen , Hao Dong
‹ Prev 1 8 9 10 Next ›