中文
相关论文

相关论文: SANPO: A Scene Understanding, Accessibility and Hu…

200 篇论文

Robust perception is critical for autonomous driving, especially under adverse weather and lighting conditions that commonly occur in real-world environments. In this paper, we introduce the Stereo Image Dataset (SID), a large-scale…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Zaid A. El-Shair , Abdalmalek Abu-raddaha , Aaron Cofield , Hisham Alawneh , Mohamed Aladem , Yazan Hamzeh , Samir A. Rawashdeh

Aiming at facilitating a real-world, ever-evolving and scalable autonomous driving system, we present a large-scale dataset for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw…

计算机视觉与模式识别 · 计算机科学 2021-11-09 Jianhua Han , Xiwen Liang , Hang Xu , Kai Chen , Lanqing Hong , Jiageng Mao , Chaoqiang Ye , Wei Zhang , Zhenguo Li , Xiaodan Liang , Chunjing Xu

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Kedi Ying , Ruiping Liu , Chongyan Chen , Mingzhe Tao , Hao Shi , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

With a number of marine populations in rapid decline, collecting and analyzing data about marine populations has become increasingly important to develop effective conservation policies for a wide range of marine animals, including whales.…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Akshaj Gaur , Cheng Liu , Xiaomin Lin , Nare Karapetyan , Yiannis Aloimonos

We introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jinsub Yim , Hyungtae Lee , Sungmin Eum , Yi-Ting Shen , Yan Zhang , Heesung Kwon , Shuvra S. Bhattacharyya

Road scene understanding is crucial in autonomous driving, enabling machines to perceive the visual environment. However, recent object detectors tailored for learning on datasets collected from certain geographical locations struggle to…

计算机视觉与模式识别 · 计算机科学 2024-02-13 Hasib Zunair , Shakib Khan , A. Ben Hamza

Vision-based end-to-end (E2E) driving has garnered significant interest in the research community due to its scalability and synergy with multimodal large language models (MLLMs). However, current E2E driving benchmarks primarily feature…

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has been done to…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Fadila Wendigoundi Douamba , Jianjun Song , Ling Fu , Yuliang Liu , Xiang Bai

We present a new, publicly-available image dataset generated by the NVIDIA Deep Learning Data Synthesizer intended for use in object detection, pose estimation, and tracking applications. This dataset contains 144k stereo image pairs that…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Mona Jalal , Josef Spjut , Ben Boudaoud , Margrit Betke

The accelerating development of autonomous driving technology has placed greater demands on obtaining large amounts of high-quality data. Representative, labeled, real world data serves as the fuel for training deep learning networks,…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Pengchuan Xiao , Zhenlei Shao , Steven Hao , Zishuo Zhang , Xiaolin Chai , Judy Jiao , Zesong Li , Jian Wu , Kai Sun , Kun Jiang , Yunlong Wang , Diange Yang

Exploring to what humans pay attention in dynamic panoramic scenes is useful for many fundamental applications, including augmented reality (AR) in retail, AR-powered recruitment, and visual language navigation. With this goal in mind, we…

计算机视觉与模式识别 · 计算机科学 2021-11-15 Yi Zhang

Predicting the interaction between pedestrian and vehicle is essential for autonomous driving safety in unstructured and semi-structured scenarios; however, this task is severely hindered by the scarcity of public datasets that feature…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Haoyang Peng , Qian Hu , Songan Zhang , Ming Yang

According to WHO statistics, the number of visually impaired people is increasing annually. One of the most critical necessities for visually impaired people is the ability to navigate safely. This paper proposes a navigation system based…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Mohammad Javadian Farzaneh , Hossein Mahvash Mohammadi

Maintaining situational awareness (SA) is critical in human-robot teams. Yet, under high workload and dynamic conditions, operators often experience SA gaps. Automated detection of SA gaps could provide timely assistance for operators.…

Motivated by the impact of large-scale datasets on ML systems we present the largest self-driving dataset for motion prediction to date, containing over 1,000 hours of data. This was collected by a fleet of 20 autonomous vehicles along a…

计算机视觉与模式识别 · 计算机科学 2020-11-18 John Houston , Guido Zuidhof , Luca Bergamini , Yawei Ye , Long Chen , Ashesh Jain , Sammy Omari , Vladimir Iglovikov , Peter Ondruska

The value of roadside perception, which could extend the boundaries of autonomous driving and traffic management, has gradually become more prominent and acknowledged in recent years. However, existing roadside perception approaches only…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Ruiyang Hao , Siqi Fan , Yingru Dai , Zhenlin Zhang , Chenxi Li , Yuntian Wang , Haibao Yu , Wenxian Yang , Jirui Yuan , Zaiqing Nie

It is estimated that 285 million people globally are visually impaired. A majority of these people live in developing countries and are among the elderly population. One of the most difficult tasks faced by the visually impaired is…

计算机与社会 · 计算机科学 2015-06-04 Shonal Chaudhry , Rohitash Chandra

This paper describes the COCO-Text dataset. In recent years large-scale datasets like SUN and Imagenet drove the advancement of scene understanding and object recognition. The goal of COCO-Text is to advance state-of-the-art in text…

计算机视觉与模式识别 · 计算机科学 2016-06-21 Andreas Veit , Tomas Matera , Lukas Neumann , Jiri Matas , Serge Belongie

Accurate 3D trajectory data is crucial for advancing autonomous driving. Yet, traditional datasets are usually captured by fixed sensors mounted on a car and are susceptible to occlusion. Additionally, such an approach can precisely…

We address the challenge of predicting human visual attention during real-world navigation by measuring and modeling egocentric pedestrian eye gaze in an outdoor campus setting. We introduce the EgoCampus dataset, which spans 25 unique…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Ronan John , Aditya Kesari , Vincenzo DiMatteo , Kristin Dana