中文
相关论文

相关论文: MABe22: A Multi-Species Multi-Task Benchmark for L…

200 篇论文

We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, Gemini-2.5-Flash and…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Thomas Heap , Laurence Aitchison , Emma Cahill , Adriana Casado Rodriguez

Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation…

计算机视觉与模式识别 · 计算机科学 2018-03-22 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

Fetching, which includes approaching, grasping, and retrieving, is a critical challenge for robot manipulation tasks. Existing methods primarily focus on table-top scenarios, which do not adequately capture the complexities of environments…

机器人学 · 计算机科学 2024-10-21 Beining Han , Meenal Parakh , Derek Geng , Jack A Defay , Gan Luyang , Jia Deng

Rather than simply recognizing the action of a person individually, collective activity recognition aims to find out what a group of people is acting in a collective scene. Previ- ous state-of-the-art methods using hand-crafted potentials…

计算机视觉与模式识别 · 计算机科学 2017-09-21 Yongyi Tang , Peizhen Zhang , Jian-Fang Hu , Wei-Shi Zheng

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Quantification of behavior is critical in applications ranging from neuroscience, veterinary medicine and animal conservation efforts. A common key step for behavioral analysis is first extracting relevant keypoints on animals, known as…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Shaokai Ye , Anastasiia Filippova , Jessy Lauer , Steffen Schneider , Maxime Vidal , Tian Qiu , Alexander Mathis , Mackenzie Weygandt Mathis

Video data and algorithms have been driving advances in multi-object tracking (MOT). While existing MOT datasets focus on occlusion and appearance similarity, complex motion patterns are widespread yet overlooked. To address this issue, we…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Xiaoyan Cao , Yiyao Zheng , Yao Yao , Huapeng Qin , Xiaoyu Cao , Shihui Guo

Investigating children's embodied learning in mixed-reality environments, where they collaboratively simulate scientific processes, requires analyzing complex multimodal data to interpret their learning and coordination behaviors. Learning…

Micro-gesture recognition and behavior-based emotion prediction are both highly challenging tasks that require modeling subtle, fine-grained human behaviors, primarily leveraging video and skeletal pose data. In this work, we present two…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Arman Martirosyan , Shahane Tigranyan , Maria Razzhivina , Artak Aslanyan , Nazgul Salikhova , Ilya Makarov , Andrey Savchenko , Aram Avetisyan

Numerous fields, such as ecology, biology, and neuroscience, use animal recordings to track and measure animal behaviour. Over time, a significant volume of such data has been produced, but some computer vision techniques cannot explore it…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Jose Sosa , Sharn Perry , Jane Alty , David Hogg

The study of mouse social behaviours has been increasingly undertaken in neuroscience research. However, automated quantification of mouse behaviours from the videos of interacting mice is still a challenging problem, where object tracking…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Zheheng Jiang , Zhihua Liu , Long Chen , Lei Tong , Xiangrong Zhang , Xiangyuan Lan , Danny Crookes , Ming-Hsuan Yang , Huiyu Zhou

Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these…

计算机视觉与模式识别 · 计算机科学 2021-06-10 Grant Van Horn , Elijah Cole , Sara Beery , Kimberly Wilber , Serge Belongie , Oisin Mac Aodha

We introduce PACE (Pose Annotations in Cluttered Environments), a large-scale benchmark designed to advance the development and evaluation of pose estimation methods in cluttered scenarios. PACE provides a large-scale real-world benchmark…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Yang You , Kai Xiong , Zhening Yang , Zhengxiang Huang , Junwei Zhou , Ruoxi Shi , Zhou Fang , Adam W. Harley , Leonidas Guibas , Cewu Lu

This technical report introduces our solution, MEEV, proposed to the EgoBody Challenge at ECCV 2022. Captured from head-mounted devices, the dataset consists of human body shape and motion of interacting people. The EgoBody dataset has…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Nicolas Monet , Dongyoon Wee

Deep learning models have achieved excellent recognition results on large-scale video benchmarks. However, they perform poorly when applied to videos with rare scenes or objects, primarily due to the bias of existing video datasets. We…

计算机视觉与模式识别 · 计算机科学 2022-09-21 Haodong Duan , Yue Zhao , Kai Chen , Yuanjun Xiong , Dahua Lin

Panoptic tracking enables pixel-level scene interpretation of videos by integrating instance tracking in panoptic segmentation. This provides robots with a spatio-temporal understanding of the environment, an essential attribute for their…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Juana Valeria Hurtado , Sajad Marvi , Rohit Mohan , Abhinav Valada

Sensor data streams from wearable devices and smart environments are widely studied in areas like human activity recognition (HAR), person identification, or health monitoring. However, most of the previous works in activity and sensor…

机器学习 · 计算机科学 2023-08-09 Taoran Sheng , Manfred Huber

We present DogMo, a large-scale multi-view RGB-D video dataset capturing diverse canine movements for the task of motion recovery from images. DogMo comprises 1.2k motion sequences collected from 10 unique dogs, offering rich variation in…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Zan Wang , Siyu Chen , Luya Mo , Xinfeng Gao , Yuxin Shen , Lebin Ding , Wei Liang

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEBE) framework to…

多智能体系统 · 计算机科学 2025-07-11 Sinem Erisken , Timothy Gothard , Martin Leitgab , Ram Potham

Instruction data is crucial for improving the capability of Large Language Models (LLMs) to align with human-level performance. Recent research LIMA demonstrates that alignment is essentially a process where the model adapts instructions'…

计算与语言 · 计算机科学 2024-10-01 Yiwei Li , Jiayi Shi , Shaoxiong Feng , Peiwen Yuan , Xinglin Wang , Boyuan Pan , Heda Wang , Yao Hu , Kan Li