中文
相关论文

相关论文: The Multiscenario Multienvironment BioSecure Multi…

200 篇论文

We present a demonstration of a web-based system called M2LADS ("System for Generating Multimodal Learning Analytics Dashboards"), designed to integrate, synchronize, visualize, and analyze multimodal data recorded during computer-based…

人机交互 · 计算机科学 2025-03-17 Alvaro Becerra , Roberto Daza , Ruth Cobos , Aythami Morales , Julian Fierrez

We explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a…

Accurate motion understanding of the dynamic objects within the scene in bird's-eye-view (BEV) is critical to ensure a reliable obstacle avoidance system and smooth path planning for autonomous vehicles. However, this task has received…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Hiep Truong Cong , Ajay Kumar Sigatapu , Arindam Das , Yashwanth Sharma , Venkatesh Satagopan , Ganesh Sistu , Ciaran Eising

Breast Magnetic Resonance Imaging (MRI) demonstrates the highest sensitivity for breast cancer detection among imaging modalities and is standard practice for high-risk women. Interpreting the multi-sequence MRI is time-consuming and prone…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Luyang Luo , Mingxiang Wu , Mei Li , Yi Xin , Qiong Wang , Varut Vardhanabhuti , Winnie CW Chu , Zhenhui Li , Juan Zhou , Pranav Rajpurkar , Hao Chen

Developmental dysgraphia is a neurological disorder that hinders children's writing skills. In recent years, researchers have increasingly explored machine learning methods to support the diagnosis of dysgraphia based on offline and online…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Jayakanth Kunhoth , Somaya Al-Maadeed , Moutaz Saleh , Younes Akbari

In this paper we list the sensors commonly available in modern smartphones and provide a general outlook of the different ways these sensors can be used for modeling the interaction between human and smartphones. We then provide a taxonomy…

密码学与安全 · 计算机科学 2020-06-02 Alejandro Acien , Aythami Morales , Ruben Vera-Rodriguez , Julian Fierrez

This paper presents a database of human faces for persons wearing spectacles. The database consists of images of faces having significant variations with respect to illumination, head pose, skin color, facial expressions and sizes, and…

计算机视觉与模式识别 · 计算机科学 2016-08-12 Anirban Dasgupta , Shubhobrata Bhattacharya , Aurobinda Routray

Gait recognition has emerged as a powerful biometric technique for identifying individuals at a distance without requiring user cooperation. Most existing methods focus primarily on RGB-derived modalities, which fall short in real-world…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Chenye Wang , Qingyuan Cai , Saihui Hou , Aoqi Li , Yongzhen Huang

This paper describes a pipeline for collecting acoustic scene data by using crowdsourcing. The detailed process of crowdsourcing is explained, including planning, validation criteria, and actual user interfaces. As a result of data…

音频与语音处理 · 电气工程与系统科学 2022-11-07 Il-Young Jeong , Jeongsoo Park

Multimodal tactile sensing could potentially enable robots to improve their performance at manipulation tasks by rapidly discriminating between task-relevant objects. Data-driven approaches to this tactile perception problem show promise,…

机器人学 · 计算机科学 2015-11-13 Joshua Wade , Tapomayukh Bhattacharjee , Charles C. Kemp

This paper presents a novel multimodal human activity recognition system. It uses a two-stream decision level fusion of vision and inertial sensors. In the first stream, raw RGB frames are passed to a part affinity field-based pose…

计算机视觉与模式识别 · 计算机科学 2023-06-29 Santosh Kumar Yadav , Muhtashim Rafiqi , Egna Praneeth Gummana , Kamlesh Tiwari , Hari Mohan Pandey , Shaik Ali Akbara

The Real Face Dataset is a pedestrian face detection benchmark dataset in the wild, comprising over 11,000 images and over 55,000 detected faces in various ambient conditions. The dataset aims to provide a comprehensive and diverse…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Leonardo Ramos Thomas

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging…

计算机视觉与模式识别 · 计算机科学 2022-04-15 Bharat Lal Bhatnagar , Xianghui Xie , Ilya A. Petrov , Cristian Sminchisescu , Christian Theobalt , Gerard Pons-Moll

The rapid evolution of Large Multimodal Models (LMMs) has enabled agents to perform complex digital and physical tasks, yet their deployment as autonomous decision-makers introduces substantial unintentional behavioral safety risks.…

人工智能 · 计算机科学 2026-03-30 Yuxuan Li , Yi Lin , Peng Wang , Shiming Liu , Xuetao Wei

In this study, we present a comprehensive public dataset for driver drowsiness detection, integrating multimodal signals of facial, behavioral, and biometric indicators. Our dataset includes 3D facial video using a depth camera, IR camera…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Morteza Bodaghi , Majid Hosseini , Raju Gottumukkala , Ravi Teja Bhupatiraju , Iftikhar Ahmad , Moncef Gabbouj

Current mobile user authentication systems based on PIN codes, fingerprint, and face recognition have several shortcomings. Such limitations have been addressed in the literature by exploring the feasibility of passive authentication on…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Giuseppe Stragapede , Ruben Vera-Rodriguez , Ruben Tolosana , Aythami Morales , Alejandro Acien , Gael Le Lan

Multimodal Foundation Models (FMs) offer a path to learn general-purpose representations from heterogeneous ecological data, easily transferable to downstream tasks. However, practical biodiversity modelling remains fragmented; separate…

Can generative agents be trusted in multimodal environments? Despite advances in large language and vision-language models that enable agents to act autonomously and pursue goals in rich settings, their ability to reason about safety,…

人工智能 · 计算机科学 2025-10-10 Alhim Vera , Karen Sanchez , Carlos Hinojosa , Haidar Bin Hamid , Donghoon Kim , Bernard Ghanem

Predicting stroke risk is a complex challenge that can be enhanced by integrating diverse clinically available data modalities. This study introduces a self-supervised multimodal framework that combines 3D brain imaging, clinical data, and…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Camille Delgrange , Olga Demler , Samia Mora , Bjoern Menze , Ezequiel de la Rosa , Neda Davoudi

Synthesizing information from multiple data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly…