中文
相关论文

相关论文: Attention-Driven Multimodal Alignment for Long-ter…

200 篇论文

Action Quality Assessment (AQA) evaluates diverse skills but models struggle with non-stationary data. We propose Continual AQA (CAQA) to refine models using sparse new data. Feature replay preserves memory without storing raw inputs.…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Kanglei Zhou , Liyuan Wang , Xingxing Zhang , Hubert P. H. Shum , Frederick W. B. Li , Jianguo Li , Xiaohui Liang

While abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models~(LMMs), their visual quality assessment~(IQA) ability has been relatively under-explored. Here we…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Hanwei Zhu , Xiangjie Sui , Baoliang Chen , Xuelin Liu , Peilin Chen , Yuming Fang , Shiqi Wang

As long-context language modeling becomes increasingly important, the cost of maintaining and attending to large Key/Value (KV) caches grows rapidly, becoming a major bottleneck in both training and inference. While prior works such as…

机器学习 · 计算机科学 2026-03-25 Dong Liu , Yanxuan Yu , Ben Lengerich , Ying Nian Wu

Real-world fine manipulation, particularly in bimanual manipulation, typically requires low-latency control and stable visual localization, while collecting large-scale data is costly and limited demonstrations may lead to localization…

机器人学 · 计算机科学 2026-05-04 Xianbo Cai , Hideyuki Ichiwara , Masaki Yoshikawa , Tetsuya Ogata

Long-term Action Quality Assessment (AQA) aims to evaluate the quantitative performance of actions in long videos. However, existing methods face challenges due to domain shifts between the pre-trained large-scale action recognition…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Kanglei Zhou , Hubert P. H. Shum , Frederick W. B. Li , Xingxing Zhang , Xiaohui Liang

Audio question answering (AQA), acting as a widely used proxy task to explore scene understanding, has got more attention. The AQA is challenging for it requires comprehensive temporal reasoning from different scales' events of an audio…

声音 · 计算机科学 2023-05-30 Guangyao Li , Yixin Xu , Di Hu

Human pose serves as a cornerstone of action quality assessment (AQA), where subtle spatial-temporal variations in pose often distinguish excellence from mediocrity. In high-level competitions, these nuanced differences become decisive…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Shuaikang Zhu , Yang Yang , Chen Sun

Weakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods…

计算机视觉与模式识别 · 计算机科学 2021-03-26 Ashraful Islam , Chengjiang Long , Richard Radke

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Shaoning Xiao , Long Chen , Kaifeng Gao , Zhao Wang , Yi Yang , Zhimeng Zhang , Jun Xiao

As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual…

图像与视频处理 · 电气工程与系统科学 2024-07-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Weisi Lin , Guangtao Zhai

This work presents a new multimodal system for remote attention level estimation based on multimodal face analysis. Our multimodal approach uses different parameters and signals obtained from the behavior and physiological processes that…

计算机视觉与模式识别 · 计算机科学 2023-01-24 Roberto Daza , Luis F. Gomez , Aythami Morales , Julian Fierrez , Ruben Tolosana , Ruth Cobos , Javier Ortega-Garcia

Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention…

人机交互 · 计算机科学 2025-05-22 Lu Li , Cunhang Fan , Hongyu Zhang , Jingjing Zhang , Xiaoke Yang , Jian Zhou , Zhao Lv

Multi-modal large language models (MLLMs) have demonstrated considerable potential across various downstream tasks that require cross-domain knowledge. MLLMs capable of processing videos, known as Video-MLLMs, have attracted broad interest…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Jiajun Fei , Dian Li , Zhidong Deng , Zekun Wang , Gang Liu , Hui Wang

Physical rehabilitation programs frequently begin with a brief stay in the hospital and continue with home-based rehabilitation. Lack of feedback on exercise correctness is a significant issue in home-based rehabilitation. Automated…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Aditya Kanade , Mansi Sharma , Manivannan Muniyandi

The attention mechanism is blooming in computer vision nowadays. However, its application to video quality assessment (VQA) has not been reported. Evaluating the quality of in-the-wild videos is challenging due to the unknown of pristine…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Fengchuang Xing , Yuan-Gen Wang , Hanpin Wang , Leida Li , Guopu Zhu

Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely through semantic perception. However, this paradigm diverges from human aesthetic cognition,…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Liwen Yu , Chi Liu , Xiaotong Han , Congcong Zhu , Minghao Wang , Sheng Shen

Long-context modeling has drawn more and more attention in the area of Large Language Models (LLMs). Continual training with long-context data becomes the de-facto method to equip LLMs with the ability to process long inputs. However, it…

计算与语言 · 计算机科学 2025-10-14 Jianghao Chen , Junhong Wu , Yangyifan Xu , Jiajun Zhang

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato

Multimodal Learning Analytics (MMLA) leverages advanced sensing technologies and artificial intelligence to capture complex learning processes, but integrating diverse data sources into cohesive insights remains challenging. This study…

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Najmeh Sadoughi , Xinyu Li , Avijit Vajpayee , David Fan , Bing Shuai , Hector Santos-Villalobos , Vimal Bhat , Rohith MV