中文
相关论文

相关论文: OmniHD-Scenes: A Next-Generation Multimodal Datase…

200 篇论文

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

声音 · 计算机科学 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Tai Wang , Xiaohan Mao , Chenming Zhu , Runsen Xu , Ruiyuan Lyu , Peisen Li , Xiao Chen , Wenwei Zhang , Kai Chen , Tianfan Xue , Xihui Liu , Cewu Lu , Dahua Lin , Jiangmiao Pang

3D semantic segmentation is one of the key tasks for autonomous driving system. Recently, deep learning models for 3D semantic segmentation task have been widely researched, but they usually require large amounts of training data. However,…

机器人学 · 计算机科学 2020-02-24 Yancheng Pan , Biao Gao , Jilin Mei , Sibo Geng , Chengkun Li , Huijing Zhao

Most existing mobile robotic datasets primarily capture static scenes, limiting their utility for evaluating robotic performance in dynamic environments. To address this, we present a mobile robot oriented large-scale indoor dataset,…

机器人学 · 计算机科学 2024-12-12 Zeshun Li , Fuhao Li , Wanting Zhang , Zijie Zheng , Xueping Liu , Yongjin Liu , Long Zeng

We present UrbanScene3D, a large-scale data platform for research of urban scene perception and reconstruction. UrbanScene3D contains over 128k high-resolution images covering 16 scenes including large-scale real urban regions and synthetic…

计算机视觉与模式识别 · 计算机科学 2022-07-20 Liqiang Lin , Yilin Liu , Yue Hu , Xingguang Yan , Ke Xie , Hui Huang

Robust perception is critical for autonomous driving, especially under adverse weather and lighting conditions that commonly occur in real-world environments. In this paper, we introduce the Stereo Image Dataset (SID), a large-scale…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Zaid A. El-Shair , Abdalmalek Abu-raddaha , Aaron Cofield , Hisham Alawneh , Mohamed Aladem , Yazan Hamzeh , Samir A. Rawashdeh

Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-based perception approaches often struggle with sensor noise,…

机器人学 · 计算机科学 2025-08-29 Zifan Wang , Teli Ma , Yufei Jia , Xun Yang , Jiaming Zhou , Wenlong Ouyang , Qiang Zhang , Junwei Liang

High-definition (HD) semantic mapping of complex intersections poses significant challenges for traditional vehicle-based approaches due to occlusions and limited perspectives. This paper introduces a novel camera-LiDAR fusion framework…

机器人学 · 计算机科学 2025-07-15 Zhongzhang Chen , Miao Fan , Shengtong Xu , Mengmeng Yang , Kun Jiang , Xiangzeng Liu , Haoyi Xiong

Unmanned surface vehicles can encounter a number of varied visual circumstances during operation, some of which can be very difficult to interpret. While most cases can be solved only using color camera images, some weather and lighting…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Jon Muhovič , Janez Perš

4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this work, we contribute…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Zhongang Cai , Daxuan Ren , Ailing Zeng , Zhengyu Lin , Tao Yu , Wenjia Wang , Xiangyu Fan , Yang Gao , Yifan Yu , Liang Pan , Fangzhou Hong , Mingyuan Zhang , Chen Change Loy , Lei Yang , Ziwei Liu

Multimodal multitask learning has attracted an increasing interest in recent years. Singlemodal models have been advancing rapidly and have achieved astonishing results on various tasks across multiple domains. Multimodal learning offers…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Ye Xue , Diego Klabjan , Jean Utke

Conventional frame camera is the mainstream sensor of the autonomous driving scene perception, while it is limited in adverse conditions, such as low light. Event camera with high dynamic range has been applied in assisting frame camera for…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Shihan Peng , Hanyu Zhou , Hao Dong , Zhiwei Shi , Haoyue Liu , Yuxing Duan , Yi Chang , Luxin Yan

Person re-identification (ReID) suffers from a lack of large-scale high-quality training data due to challenges in data privacy and annotation costs. While previous approaches have explored pedestrian generation for data augmentation, they…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Changxiao Ma , Chao Yuan , Xincheng Shi , Yuzhuo Ma , Yongfei Zhang , Longkun Zhou , Yujia Zhang , Shangze Li , Yifan Xu

Camera-based 3D semantic scene completion (SSC) plays a crucial role in autonomous driving, enabling voxelized 3D scene understanding for effective scene perception and decision-making. Existing SSC methods have shown efficacy in improving…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Zhiwen Yang , Yuxin Peng

Optical flow is the motion of a pixel between at least two consecutive video frames and can be estimated through an end-to-end trainable convolutional neural network. To this end, large training datasets are required to improve the accuracy…

计算机视觉与模式识别 · 计算机科学 2021-04-19 Roman Seidel , André Apitzsch , Gangolf Hirtz

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Haoyi Tao , Chaozheng Huang , Nan Wang , Han Lyu , Linfeng Zhang , Guolin Ke , Xi Fang

Omnidirectional image and video super-resolution is a crucial research topic in low-level vision, playing an essential role in virtual reality and augmented reality applications. Its goal is to reconstruct high-resolution images or video…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Qianqian Zhao , Chunle Guo , Tianyi Zhang , Junpei Zhang , Peiyang Jia , Tan Su , Wenjie Jiang , Chongyi Li

Robust 3D occupancy prediction is essential for autonomous driving, particularly under adverse weather conditions where traditional vision-only systems struggle. While the fusion of surround-view 4D radar and cameras offers a promising…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Long Yang , Lianqing Zheng , Wenjin Ai , Minghao Liu , Sen Li , Qunshu Lin , Shengyu Yan , Jie Bai , Zhixiong Ma , Tao Huang , Xichan Zhu

Understanding the complex urban infrastructure with centimeter-level accuracy is essential for many applications from autonomous driving to mapping, infrastructure monitoring, and urban management. Aerial images provide valuable information…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Seyed Majid Azimi , Corentin Henry , Lars Sommer , Arne Schumann , Eleonora Vig

This work presents a novel video dataset recorded from overlapping highway traffic cameras along an urban interstate, enabling multi-camera 3D object tracking in a traffic monitoring context. Data is released from 3 scenes containing video…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Derek Gloudemans , Yanbing Wang , Gracie Gumm , William Barbour , Daniel B. Work