中文
相关论文

相关论文: Vi-SAFE: A Spatial-Temporal Framework for Efficien…

200 篇论文

Temporal sentence grounding (TSG) aims to localize the temporal segment which is semantically aligned with a natural language query in an untrimmed video.Most existing methods extract frame-grained features or object-grained features by 3D…

计算机视觉与模式识别 · 计算机科学 2023-02-22 Zeyu Xiong , Daizong Liu , Pan Zhou , Jiahao Zhu

Spatio-temporal action detection (STAD) is an important fine-grained video understanding task. Current methods require box and label supervision for all action classes in advance. However, in real-world applications, it is very likely to…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Tao Wu , Shuqiu Ge , Jie Qin , Gangshan Wu , Limin Wang

In this work, we propose a motion robust and high-speed detection pipeline which better leverages the event data. First, we design an event stream representation called temporal active focus (TAF), which efficiently utilizes the…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Bingde Liu , Chang Xu , Wen Yang , Huai Yu , Lei Yu

Lane segment topology reasoning provides comprehensive bird's-eye view (BEV) road scene understanding, which can serve as a key perception module in planning-oriented end-to-end autonomous driving systems. Existing lane topology reasoning…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yiming Yang , Hongbin Lin , Yueru Luo , Suzhong Fu , Chao Zheng , Xinrui Yan , Shuqi Mei , Kun Tang , Shuguang Cui , Zhen Li

This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through spatial localization on time-frequency spectrograms. Vibration signals are transformed into continuous wavelet transform (CWT) spectrograms…

信号处理 · 电气工程与系统科学 2026-04-21 Po-Heng Chou , Wei-Lung Mao , Ru-Ping Lin , Jen-Yu Chiu , Chun-Yu Yeh

Nowadays, navigation and ride-sharing apps have collected numerous images with spatio-temporal data. A core technology for analyzing such images, associated with spatiotemporal information, is Traffic Scene Understanding (TSU), which aims…

多媒体 · 计算机科学 2025-11-13 Jingtian Ma , Jingyuan Wang , Wayne Xin Zhao , Guoping Liu , Xiang Wen

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrates modalities such as…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Wenping Jin , Li Zhu , Jing Sun

Early detection of anxiety is crucial for reducing the suffering of individuals with mental disorders and improving treatment outcomes. Utilizing an mHealth platform for anxiety screening can be particularly practical in improving screening…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Haimiao Mo , Yuchen Li , Shanlin Yang , Wei Zhang , Shuai Ding

Convolutional Neural Networks (CNN) are commonly used for the problem of object detection thanks to their increased accuracy. Nevertheless, the performance of CNN-based detection models is ambiguous when detection speed is considered. To…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Ioanna Gogou , Dimitrios Koutsomitropoulos

In this paper, we present a unified, end-to-end trainable spatiotemporal CNN model for VOS, which consists of two branches, i.e., the temporal coherence branch and the spatial segmentation branch. Specifically, the temporal coherence branch…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Kai Xu , Longyin Wen , Guorong Li , Liefeng Bo , Qingming Huang

Tram-human interaction safety is an important challenge, given that trams frequently operate in densely populated areas, where collisions can range from minor injuries to fatal outcomes. This paper addresses the issue from the perspective…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Ondřej Valach , Ivan Gruber

Combining Simultaneous Localisation and Mapping (SLAM) estimation and dynamic scene modelling can highly benefit robot autonomy in dynamic environments. Robot path planning and obstacle avoidance tasks rely on accurate estimations of the…

机器人学 · 计算机科学 2021-12-16 Jun Zhang , Mina Henein , Robert Mahony , Viorela Ila

Cooperative perception via communication among intelligent traffic agents has great potential to improve the safety of autonomous driving. However, limited communication bandwidth, localization errors and asynchronized capturing time of…

计算机视觉与模式识别 · 计算机科学 2024-08-23 Yunshuang Yuan , Monika Sester

Video super-resolution (VSR) is a task that aims to reconstruct high-resolution (HR) frames from the low-resolution (LR) reference frame and multiple neighboring frames. The vital operation is to utilize the relative misaligned frames for…

计算机视觉与模式识别 · 计算机科学 2022-11-04 Meiqin Liu , Shuo Jin , Chao Yao , Chunyu Lin , Yao Zhao

Detection of pedestrians on embedded devices, such as those on-board of robots and drones, has many applications including road intersection monitoring, security, crowd monitoring and surveillance, to name a few. However, the problem can be…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Mohamed Afifi , Yara Ali , Karim Amer , Mahmoud Shaker , Mohamed Elhelw

We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close to real-time performance, supporting applications including…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Gergely Dinya , Péter Halász , András Lőrincz , Kristóf Karacs , Anna Gelencsér-Horváth

object detection framework plays crucial role in autonomous driving. In this paper, we introduce the real-time object detection framework called You Only Look Once (YOLOv1) and the related improvements of YOLOv2. We further explore the…

计算机视觉与模式识别 · 计算机科学 2019-05-14 Shouyu Wang , Weitao Tang

Identifying drones and birds correctly is essential for keeping the skies safe and improving security systems. Using the VIP CUP 2025 dataset, which provides both RGB and infrared (IR) images, this study presents EGD-YOLOv8n, a new…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Sudipto Sarkar , Mohammad Asif Hasan , Khondokar Ashik Shahriar , Fablia Labiba , Nahian Tasnim , Sheikh Anawarul Haq Fattah

While single image shadow detection has been improving rapidly in recent years, video shadow detection remains a challenging task due to data scarcity and the difficulty in modelling temporal consistency. The current video shadow detection…

计算机视觉与模式识别 · 计算机科学 2021-08-02 Shilin Hu , Hieu Le , Dimitris Samaras

Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhaoxing Zhang , Junda Cheng , Gangwei Xu , Xiaoxiang Wang , Can Zhang , Xin Yang
‹ 上一页 1 8 9 10 下一页 ›