中文
相关论文

相关论文: CALM: Cognitive Assessment using Light-insensitive…

200 篇论文

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Zhiqiu Lin , Samuel Yu , Zhiyi Kuang , Deepak Pathak , Deva Ramanan

Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-reliance on a dominant modality. We introduce a context-aligned reasoning framework that…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Sumra Khan , Sagar Chhabriya , Aizan Zafar , Sheeraz Arif , Amgad Muneer , Anas Zafar , Shaina Raza , Rizwan Qureshi

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions…

计算与语言 · 计算机科学 2025-01-15 Hui Lee , Singh Suniljit , Yong Siang Ong

Accurately recognizing health-related conditions from wearable data is crucial for improved healthcare outcomes. To improve the recognition accuracy, various approaches have focused on how to effectively fuse information from multiple…

机器学习 · 计算机科学 2022-02-18 Huiyuan Yang , Han Yu , Kusha Sridhar , Thomas Vaessen , Inez Myin-Germeys , Akane Sano

Accurate classification of medical device risk levels is essential for regulatory oversight and clinical safety. We present a Transformer-based multimodal framework that integrates textual descriptions and visual information to predict…

机器学习 · 计算机科学 2025-05-02 Yu Han , Aaron Ceross , Jeroen H. M. Bergmann

Existing Large Vision-Language Models (LVLMs) primarily align image features of vision encoder with Large Language Models (LLMs) to leverage their superior text generation capabilities. However, the scale disparity between vision encoder…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Shi Liu , Kecheng Zheng , Wei Chen

Recent advancements in Large Vision-Language Models (LVLMs) have significantly expanded their utility in tasks like image captioning and visual question answering. However, they still struggle with object hallucination, where models…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Yeongjae Cho , Keonwoo Kim , Taebaek Hwang , Sungzoon Cho

Multimodal sensors provide complementary information to develop accurate machine-learning methods for human activity recognition (HAR), but introduce significantly higher computational load, which reduces efficiency. This paper proposes an…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Ziqi Gao , Yuntao Wang , Jianguo Chen , Junliang Xing , Shwetak Patel , Xin Liu , Yuanchun Shi

In today's fast-paced world, accurately monitoring stress levels is crucial. Sensor-based stress monitoring systems often need large datasets for training effective models. However, individual-specific models are necessary for personalized…

The emergence of mobile eye trackers embedded in next generation smartphones or VR displays will make it possible to trace not only what objects we look at but also the level of attention in a given situation. Exploring whether we can…

人机交互 · 计算机科学 2016-02-17 Per Bækgaard , Michael Kai Petersen , Jakob Eg Larsen

Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assessments, and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Yuheng Li , Yue Zhang , Abdoul Aziz Amadou , Yuxiang Lai , Jike Zhong , Tiziano Passerini , Dorin Comaniciu , Puneet Sharma

Long-term monitoring of patients with epilepsy presents a challenging problem from the engineering perspective of real-time detection and wearable devices design. It requires new solutions that allow continuous unobstructed monitoring and…

机器学习 · 计算机科学 2022-04-11 Una Pale , Tomas Teijeiro , David Atienza

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

声音 · 计算机科学 2021-03-03 Michael Neumann , Ngoc Thang Vu

Emergency Medical Services (EMS) responders often operate under time-sensitive conditions, facing cognitive overload and inherent risks, requiring essential skills in critical thinking and rapid decision-making. This paper presents…

人工智能 · 计算机科学 2024-10-27 Keshara Weerasinghe , Saahith Janapati , Xueren Ge , Sion Kim , Sneha Iyer , John A. Stankovic , Homa Alemzadeh

Cultural heritage sites face accelerating degradation due to climate change, yet tradi- tional monitoring relies on unimodal analysis (visual inspection or environmental sen- sors alone) that fails to capture the complex interplay between…

人工智能 · 计算机科学 2025-10-17 David Roqui , Adèle Cormier , nistor Grozavu , Ann Bourges

Attitude determination using the smartphone's inertial sensors poses a major challenge due to the sensor low-performance grade and variate nature of the walking pedestrian. In this paper, data-driven techniques are employed to address that…

信号处理 · 电气工程与系统科学 2022-09-13 Eran Vertzberger , Itzik Klein

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated significant progress in tackling complex multimodal tasks. Among these cutting-edge developments, Google's Bard stands out for its remarkable multimodal…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Wenqi Shao , Meng Lei , Yutao Hu , Peng Gao , Kaipeng Zhang , Fanqing Meng , Peng Xu , Siyuan Huang , Hongsheng Li , Yu Qiao , Ping Luo

Retrieving visual and textual information from medical literature and hospital records can enhance diagnostic accuracy for clinical image interpretation. However, multimodal retrieval-augmented diagnosis is highly challenging. We explore a…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Nir Mazor , Tom Hope

Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degradations, such as fog, blur, or low-light conditions. Infrared…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Abrar Majeedi , Zhiyuan Ruan , Ziyi Zhao , Hongcheng Wang , Jianglin Lu , Yin Li

Timely implementation of interventions to slow cognitive decline among older adults requires accurate monitoring to detect changes in cognitive function. Data gathered using wearable devices that can continuously monitor factors known to be…

信号处理 · 电气工程与系统科学 2024-03-26 Collin Sakal , Tingyou Li , Juan Li , Xinyue Li