中文
相关论文

相关论文: The Third Place Solution for CVPR2022 AVA Accessib…

200 篇论文

We present a vision-action policy that won 1st place in the 2025 BEHAVIOR Challenge - a large-scale benchmark featuring 50 diverse long-horizon household tasks in photo-realistic simulation, requiring bimanual manipulation, navigation, and…

机器人学 · 计算机科学 2025-12-23 Ilia Larchenko , Gleb Zarin , Akash Karnatak

Visual Question Answering (VQA) is a challenging task of natural language processing (NLP) and computer vision (CV), attracting significant attention from researchers. English is a resource-rich language that has witnessed various…

计算与语言 · 计算机科学 2024-04-18 Ngan Luu-Thuy Nguyen , Nghia Hieu Nguyen , Duong T. D Vo , Khanh Quoc Tran , Kiet Van Nguyen

The Vision Challenge Track 1 for Data-Effificient Defect Detection requires competitors to instance segment 14 industrial inspection datasets in a data-defificient setting. This report introduces the technical details of the team…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Xian Tao , Zhen Qu , Hengliang Luo , Jianwen Han , Yonghao He , Danfeng Liu , Chengkan Lv , Fei Shen , Zhengtao Zhang

Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak…

机器人学 · 计算机科学 2026-05-29 Zhongyu Xia , Yousen Tang , Bingqing Wei , Yongtao Wang

Road segmentation in challenging domains, such as night, snow or rain, is a difficult task. Most current approaches boost performance using fine-tuning, domain adaptation, style transfer, or by referencing previously acquired imagery. These…

计算机视觉与模式识别 · 计算机科学 2022-05-30 Connor Malone , Sourav Garg , Ming Xu , Thierry Peynot , Michael Milford

This report presents a brief description of our winning solution to the AVA Active Speaker Detection (ASD) task at ActivityNet Challenge 2022. Our underlying model UniCon+ continues to build on our previous work, the Unified Context Network…

计算机视觉与模式识别 · 计算机科学 2022-06-23 Yuanhang Zhang , Susan Liang , Shuang Yang , Shiguang Shan

Video panoptic segmentation is an advanced task that extends panoptic segmentation by applying its concept to video sequences. In the hope of addressing the challenge of video panoptic segmentation in diverse conditions, We utilize DVIS++…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Ruipu Wu , Jifei Che , Han Li , Chengjing Wu , Ting Liu , Luoqi Liu

Stand-alone Visual Place Recognition (VPR) systems have little defence against a well-designed adversarial attack, which can lead to disastrous consequences when deployed for robot navigation. This paper extensively analyzes the effect of…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Connor Malone , Owen Claxton , Iman Shames , Michael Milford

Predicting vehicle trajectories, angle and speed is important for safe and comfortable driving. We demonstrate the best predicted angle, speed, and best performance overall winning the top three places of the ICCV 2019 Learning to Drive…

计算机视觉与模式识别 · 计算机科学 2019-11-21 Michael Diodato , Yu Li , Antonia Lovjer , Minsu Yeom , Albert Song , Yiyang Zeng , Abhay Khosla , Benedikt Schifferer , Manik Goyal , Iddo Drori

OOD-CV challenge is an out-of-distribution generalization task. To solve this problem in object detection track, we propose a simple yet effective Generalize-then-Adapt (G&A) framework, which is composed of a two-stage domain generalization…

计算机视觉与模式识别 · 计算机科学 2023-01-13 Wei Zhao , Binbin Chen , Weijie Chen , Shicai Yang , Di Xie , Shiliang Pu , Yueting Zhuang

Autonomous driving technology is developing rapidly and nowadays first autonomous rides are being provided in city areas. This requires the highest standards for the safety and reliability of the technology. Motion prediction part of the…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Stepan Konev

Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present…

机器人学 · 计算机科学 2025-09-29 Yushan Liu , Shilong Mu , Xintao Chao , Zizhen Li , Yao Mu , Tianxing Chen , Shoujie Li , Chuqiao Lyu , Xiao-Ping Zhang , Wenbo Ding

Self-driving vehicles and autonomous ground robots require a reliable and accurate method to analyze the traversability of the surrounding environment for safe navigation. This paper proposes and evaluates a real-time machine learning-based…

In order to deal with the task of video panoptic segmentation in the wild, we propose a robust integrated video panoptic segmentation solution. In our solution, we regard the video panoptic segmentation task as a segmentation target…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Jinming Su , Wangwang Yang , Junfeng Luo , Xiaolin Wei

Multi-class product counting and recognition identifies product items from images or videos for automated retail checkout. The task is challenging due to the real-world scenario of occlusions where product items overlap, fast movement in…

计算机视觉与模式识别 · 计算机科学 2022-04-26 Md. Istiak Hossain Shihab , Nazia Tasnim , Hasib Zunair , Labiba Kanij Rupty , Nabeel Mohammed

Motivated by the rapid development of autonomous vehicle technology, this work focuses on the challenges of introducing them in ride-hailing platforms with conventional strategic human drivers. We consider a ride-hailing platform that…

计算机科学与博弈论 · 计算机科学 2024-06-28 Shuqin Gao , Xinyuan Wu , Antonis Dimakis , Costas Courcoubetis

This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable of perceiving and reasoning about specific AVR problems,…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Ke Zhu , Yu Wang , Jiangjiang Liu , Qunyi Xie , Shanshan Liu , Gang Zhang

Visual place recognition (VPR) is an important component technology for camera-based mapping and navigation applications. This is a challenging problem because images of the same place may appear quite different for reasons including…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Nick Trinh , Damian Lyons

The self driving challenge in 2021 is this century's technological equivalent of the space race, and is now entering the second major decade of development. Solving the technology will create social change which parallels the invention of…

机器学习 · 计算机科学 2021-08-13 Jeffrey Hawke , Haibo E , Vijay Badrinarayanan , Alex Kendall

This technical report describes the methods we employed for the Driving with Language track of the CVPR 2024 Autonomous Grand Challenge. We utilized a powerful open-source multimodal model, InternVL-1.5, and conducted a full-parameter…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Jiahan Li , Zhiqi Li , Tong Lu