中文
相关论文

相关论文: Joint Audio-Visual Idling Vehicle Detection with S…

200 篇论文

Leveraging the computing and sensing capabilities of vehicles, vehicular federated learning (VFL) has been applied to edge training for connected vehicles. The dynamic and interconnected nature of vehicular networks presents unique…

机器学习 · 计算机科学 2025-06-10 Jintao Yan , Tan Chen , Yuxuan Sun , Zhaojun Nan , Sheng Zhou , Zhisheng Niu

One of the critical pieces of the self-driving puzzle is understanding the surroundings of a self-driving vehicle (SDV) and predicting how these surroundings will change in the near future. To address this task we propose MultiXNet, an…

Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance, they are all restricted to pre-defined relation categories,…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Kaifeng Gao , Siqi Chen , Hanwang Zhang , Jun Xiao , Yueting Zhuang , Qianru Sun

In recent years, unmanned aerial vehicle (UAV) imaging is a suitable solution for real-time monitoring different vehicles on the urban scale. Real-time vehicle detection with the use of uncertainty estimation in deep meta-learning for the…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Mehdi Khoshboresh-Masouleh , Reza Shah-Hosseini

Vehicle weaving on highways contributes to traffic congestion, raises safety issues, and underscores the need for sophisticated traffic management systems. Current tools are inadequate in offering precise and comprehensive data on…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Mei Qiu , Wei Lin , Stanley Chien , Lauren Christopher , Yaobin Chen , Shu Hu

Detection of violence and weaponized violence in closed-circuit television (CCTV) footage requires a comprehensive approach. In this work, we introduce the \emph{Smart-City CCTV Violence Detection (SCVD)} dataset, specifically designed to…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Toluwani Aremu , Li Zhiyuan , Reem Alameeri , Mustaqeem Khan , Abdulmotaleb El Saddik

Latest advances have achieved realistic virtual try-on (VTON) through localized garment inpainting using latent diffusion models, significantly enhancing consumers' online shopping experience. However, existing VTON technologies neglect the…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Fei Shen , Xin Jiang , Xin He , Hu Ye , Cong Wang , Xiaoyu Du , Zechao Li , Jinhui Tang

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV representation learning frameworks, the importance of AV data…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Shentong Mo , Yibing Song

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and…

人工智能 · 计算机科学 2025-09-23 Jia Li , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Vehicle re-identification (ReID) endeavors to associate vehicle images collected from a distributed network of cameras spanning diverse traffic environments. This task assumes paramount importance within the spectrum of vehicle-centric…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Ali Amiri , Aydin Kaya , Ali Seydi Keceli

On-board 3D object detection in autonomous vehicles often relies on geometry information captured by LiDAR devices. Albeit image features are typically preferred for detection, numerous approaches take only spatial data as input. Exploiting…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Alejandro Barrera , Carlos Guindel , Jorge Beltrán , Fernando García

We present Hybrid Voxel Network (HVNet), a novel one-stage unified network for point cloud based 3D object detection for autonomous driving. Recent studies show that 2D voxelization with per voxel PointNet style feature extractor leads to…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Maosheng Ye , Shuangjie Xu , Tongyi Cao

In recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve…

计算机视觉与模式识别 · 计算机科学 2022-07-18 Jiaqi Gu , Zhiyu Xiang , Pan Zhao , Tingming Bai , Lingxuan Wang , Xijun Zhao , Zhiyuan Zhang

Computer Vision has played a major role in Intelligent Transportation Systems (ITS) and traffic surveillance. Along with the rapidly growing automated vehicles and crowded cities, the automated and advanced traffic management systems (ATMS)…

计算机视觉与模式识别 · 计算机科学 2022-07-05 Mahdi Rezaei , Mohsen Azarmi , Farzam Mohammad Pour Mir

Traffic problems have increased in modern life due to a huge number of vehicles, big cities, and ignoring the traffic rules. Vehicular ad hoc network (VANET) has improved the traffic system in previous some and plays a vital role in the…

网络与互联网体系结构 · 计算机科学 2023-06-02 Muhammad Shoaib Farooq , Sawera Kanwal

Traffic congestion and violations pose significant challenges for urban mobility and road safety. Traditional traffic monitoring systems, such as fixed cameras and sensor-based methods, are often constrained by limited coverage, low…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Ali Khanpour , Tianyi Wang , Afra Vahidi-Shams , Wim Ectors , Farzam Nakhaie , Amirhossein Taheri , Christian Claudel

Accurate and reliable object detection is critical for ensuring the safety and efficiency of Connected Autonomous Vehicles (CAVs). Traditional on-board perception systems have limited accuracy due to occlusions and blind spots, while…

机器人学 · 计算机科学 2025-09-25 Everett Richards , Bipul Thapa , Lena Mashayekhy

With the incoming introduction of 5G networks and the advancement in technologies, such as Network Function Virtualization and Software Defined Networking, new and emerging networking technologies and use cases are taking shape. One such…

网络与互联网体系结构 · 计算机科学 2021-02-23 Dimitrios Michael Manias , Abdallah Shami

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal…

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Okan Köpüklü , Maja Taseska , Gerhard Rigoll