中文
相关论文

相关论文: MABe22: A Multi-Species Multi-Task Benchmark for L…

200 篇论文

Large language models (LLMs) excel at explicit reasoning, but their implicit computational strategies remain underexplored. Decades of psychophysics research show that humans intuitively process and integrate noisy signals using…

计算与语言 · 计算机科学 2025-12-03 Julian Ma , Jun Wang , Zafeirios Fountas

Holistic methods based on dense trajectories are currently the de facto standard for recognition of human activities in video. Whether holistic representations will sustain or will be superseded by higher level video encoding in terms of…

计算机视觉与模式识别 · 计算机科学 2014-07-29 Leonid Pishchulin , Mykhaylo Andriluka , Bernt Schiele

Engagement detection in online learning environments is vital for improving student outcomes and personalizing instruction. We present ViBED-Net (Video-Based Engagement Detection Network), a novel deep learning framework designed to assess…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Prateek Gothwal , Deeptimaan Banerjee , Ashis Kumer Biswas

Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person trajectories are then recovered by a procedure of data…

计算机视觉与模式识别 · 计算机科学 2019-05-30 Hefeng Wu , Yafei Hu , Keze Wang , Hanhui Li , Lin Nie , Hui Cheng

3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Hongyu Ke , Jack Morris , Kentaro Oguchi , Xiaofei Cao , Yongkang Liu , Haoxin Wang , Yi Ding

Human movement analysis is a key area of research in robotics, biomechanics, and data science. It encompasses tracking, posture estimation, and movement synthesis. While numerous methodologies have evolved over time, a systematic and…

机器人学 · 计算机科学 2023-05-11 Brenda Elizabeth Olivas-Padilla , Alina Glushkova , Sotiris Manitsaris

In the image classification task, the most common approach is to resize all images in a dataset to a unique shape, while reducing their precision to a size which facilitates experimentation at scale. This practice has benefits from a…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Ferran Parés , Anna Arias-Duart , Dario Garcia-Gasulla , Gema Campo-Francés , Nina Viladrich , Eduard Ayguadé , Jesús Labarta

Automated capture of animal pose is transforming how we study neuroscience and social behavior. Movements carry important social cues, but current methods are not able to robustly estimate pose and shape of animals, particularly for social…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Marc Badger , Yufu Wang , Adarsh Modh , Ammon Perkes , Nikos Kolotouros , Bernd G. Pfrommer , Marc F. Schmidt , Kostas Daniilidis

Animals perceive the world to plan their actions and interact with other agents to accomplish complex tasks, demonstrating capabilities that are still unmatched by AI systems. To advance our understanding and reduce the gap between the…

Objective monitoring and assessment of human motor behavior can improve the diagnosis and management of several medical conditions. Over the past decade, significant advances have been made in the use of wearable technology for continuously…

计算机视觉与模式识别 · 计算机科学 2019-09-23 Behnaz Rezaei , Yiorgos Christakis , Bryan Ho , Kevin Thomas , Kelley Erb , Sarah Ostadabbas , Shyamal Patel

Interactive applications demand believable characters that respond naturally to dynamic environments. Traditional character animation techniques often struggle to handle arbitrary situations, leading to a growing trend of dynamically…

图形学 · 计算机科学 2025-10-28 Jose Luis Ponton , Sheldon Andrews , Carlos Andujar , Nuria Pelechano

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models face challenges when it comes to…

As robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of…

机器人学 · 计算机科学 2025-12-09 Md Mofijul Islam , Alexi Gladstone , Sujan Sarker , Ganesh Nanduru , Md Fahim , Keyan Du , Aman Chadha , Tariq Iqbal

Estimating the pose of animals can facilitate the understanding of animal motion which is fundamental in disciplines such as biomechanics, neuroscience, ethology, robotics and the entertainment industry. Human pose estimation models have…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Moira Shooter , Charles Malleson , Adrian Hilton

Facial valence/arousal, expression and action unit are related tasks in facial affective analysis. However, the tasks only have limited performance in the wild due to the various collected conditions. The 4th competition on affective…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Yanan Chang , Yi Wu , Xiangyu Miao , Jiahe Wang , Shangfei Wang

Applied Behavior Analysis (ABA) is a clinical discipline whose documentation, teaching programs and multi-session behavioral logs, is formulaic and high-volume, yet real session data is HIPAA-protected and bound by professional…

计算与语言 · 计算机科学 2026-05-26 Festus Kahunla

Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and motor demands. But current benchmarks are often limited by small scale, unimodal sensing or…

密码学与安全 · 计算机科学 2026-05-18 Ishpuneet Singh , Gursmeep Kaur , Uday Pratap Singh Atwal , Guramrit Singh , Gurjot Singh , Maninder Singh

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing MLLMs and benchmarks primarily focus on single-image input…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Haowei Liu , Xi Zhang , Haiyang Xu , Yaya Shi , Chaoya Jiang , Ming Yan , Ji Zhang , Fei Huang , Chunfeng Yuan , Bing Li , Weiming Hu

While a great variety of 3D cameras have been introduced in recent years, most publicly available datasets for object recognition and pose estimation focus on one single camera. In this work, we present a dataset of 32 scenes that have been…

机器人学 · 计算机科学 2020-09-30 Till Grenzdörffer , Martin Günther , Joachim Hertzberg

Generative models for audio-conditioned dance motion synthesis map music features to dance movements. Models are trained to associate motion patterns to audio patterns, usually without an explicit knowledge of the human body. This approach…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Davide Moltisanti , Jinyi Wu , Bo Dai , Chen Change Loy