中文
相关论文

相关论文: RoadText-1K: Text Detection & Recognition Dataset …

200 篇论文

Vision-based road detection is an essential functionality for supporting advanced driver assistance systems (ADAS) such as road following and vehicle and pedestrian detection. The major challenges of road detection are dealing with shadows…

计算机视觉与模式识别 · 计算机科学 2014-12-11 José M. Álvarez , Ferran Diego , Joan Serrat , Antonio M. López

Dynamic scene understanding is the ability of a computer system to interpret and make sense of the visual information present in a video of a real-world scene. In this thesis, we present a series of frameworks for dynamic scene…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Salman Khan

Satellites are capable of capturing high-resolution videos. It makes vehicle perception from satellite become possible. Compared to street surveillance, drive recorder or other equipments, satellite videos provide a much broader city-scale…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Bin Zhao , Pengfei Han , Xuelong Li

The accelerating development of autonomous driving technology has placed greater demands on obtaining large amounts of high-quality data. Representative, labeled, real world data serves as the fuel for training deep learning networks,…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Pengchuan Xiao , Zhenlei Shao , Steven Hao , Zishuo Zhang , Xiaolin Chai , Judy Jiao , Zesong Li , Jian Wu , Kai Sun , Kun Jiang , Yunlong Wang , Diange Yang

While several datasets for autonomous navigation have become available in recent years, they tend to focus on structured driving environments. This usually corresponds to well-delineated infrastructure such as lanes, a small number of…

计算机视觉与模式识别 · 计算机科学 2018-11-27 Girish Varma , Anbumani Subramanian , Anoop Namboodiri , Manmohan Chandraker , C V Jawahar

Autonomous trucking offers significant benefits, such as improved safety and reduced costs, but faces unique perception challenges due to trucks' large size and dynamic trailer movements. These challenges include extensive blind spots and…

机器人学 · 计算机科学 2025-08-11 Tenghui Xie , Zhiying Song , Fuxi Wen , Jun Li , Guangzhao Liu , Zijian Zhao

This study focuses on improving the optical character recognition (OCR) data for panels in the COMICS dataset, the largest dataset containing text and images from comic books. To do this, we developed a pipeline for OCR processing and…

计算与语言 · 计算机科学 2023-01-02 Gürkan Soykan , Deniz Yuret , Tevfik Metin Sezgin

Recently, video scene text detection has received increasing attention due to its comprehensive applications. However, the lack of annotated scene text video datasets has become one of the most important problems, which hinders the…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Jiajun Zhu , Xiufeng Jiang , Zhiwei Jia , Shugong Xu , Shan Cao

We present a new traffic dataset, METEOR, which captures traffic patterns and multi-agent driving behaviors in unstructured scenarios. METEOR consists of more than 1000 one-minute videos, over 2 million annotated frames with bounding boxes…

计算机视觉与模式识别 · 计算机科学 2022-03-21 Rohan Chandra , Xijun Wang , Mridul Mahajan , Rahul Kala , Rishitha Palugulla , Chandrababu Naidu , Alok Jain , Dinesh Manocha

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Thanadol Singkhornart , Olarik Surinta

In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets.…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Piera Riccio , Francesco Galati , Kajetan Schweighofer , Noa Garcia , Nuria Oliver

Recently, video text detection, tracking, and recognition in natural scenes are becoming very popular in the computer vision community. However, most existing algorithms and benchmarks focus on common text cases (e.g., normal size, density)…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Mike Zheng Shou , Umapada Pal , Dimosthenis Karatzas , Xiang Bai

Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Haruki Sakajo , Hiroshi Takato , Hiroshi Tsutsui , Komei Soda , Hidetaka Kamigaito , Taro Watanabe

With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either…

计算机视觉与模式识别 · 计算机科学 2019-09-18 Quang-Hieu Pham , Pierre Sevestre , Ramanpreet Singh Pahwa , Huijing Zhan , Chun Ho Pang , Yuda Chen , Armin Mustafa , Vijay Chandrasekhar , Jie Lin

Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve…

机器学习 · 计算机科学 2019-06-04 Zhengping Che , Guangyu Li , Tracy Li , Bo Jiang , Xuefeng Shi , Xinsheng Zhang , Ying Lu , Guobin Wu , Yan Liu , Jieping Ye

Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel dataset, SUTD-TrafficQA…

计算机视觉与模式识别 · 计算机科学 2021-07-07 Li Xu , He Huang , Jun Liu

Driving Scene understanding is a key ingredient for intelligent transportation systems. To achieve systems that can operate in a complex physical and social environment, they need to understand and learn how humans drive and interact with…

计算机视觉与模式识别 · 计算机科学 2018-11-07 Vasili Ramanishka , Yi-Ting Chen , Teruhisa Misu , Kate Saenko

Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Ridouane Ghermi , Xi Wang , Vicky Kalogeiton , Ivan Laptev

Video-Text Retrieval (VTR) aims to search for the most relevant video related to the semantics in a given sentence, and vice versa. In general, this retrieval task is composed of four successive steps: video and textual feature…

计算机视觉与模式识别 · 计算机科学 2023-02-27 Cunjuan Zhu , Qi Jia , Wei Chen , Yanming Guo , Yu Liu

Traffic light perception is an essential component of the camera-based perception system for autonomous vehicles, enabling accurate detection and interpretation of traffic lights to ensure safe navigation through complex urban environments.…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Rupert Polley , Nikolai Polley , Dominik Heid , Marc Heinrich , Sven Ochs , J. Marius Zöllner