中文
相关论文

相关论文: MTGA: Multi-View Temporal Granularity Aligned Aggr…

200 篇论文

Recent approaches have shown that large-scale vision-language models such as CLIP can improve semantic segmentation performance. These methods typically aim for pixel-level vision-language alignment, but often rely on low resolution image…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Anurag Das , Xinting Hu , Li Jiang , Bernt Schiele

Skeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Lianyu Hu , Liqing Gao , Zekang Liu , Wei Feng

Lipreading, the technology of decoding spoken content from silent videos of lip movements, holds significant application value in fields such as public security. However, due to the subtle nature of articulatory gestures, existing…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Matteo Rossi

Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignment, a process that obscures local patch information and fine-grained spatio-temporal…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yerim So , Jiyeong Kim , Jiwon Yoon , Dongbo Min

We propose a new information aggregation method which called Localized Feature Aggregation Module based on the similarity between the feature maps of an encoder and a decoder. The proposed method recovers positional information by…

图像与视频处理 · 电气工程与系统科学 2021-12-06 Ryouichi Furukawa , Kazuhiro Hotta

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for spatial grounding…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Yuanyuan Jiang , Jianqin Yin

Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion…

计算机视觉与模式识别 · 计算机科学 2020-04-06 Yan Li , Bin Ji , Xintian Shi , Jianguo Zhang , Bin Kang , Limin Wang

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Najmeh Sadoughi , Xinyu Li , Avijit Vajpayee , David Fan , Bing Shuai , Hector Santos-Villalobos , Vimal Bhat , Rohith MV

This paper presents a sensory fusion neuromorphic dataset collected with precise temporal synchronization using a set of Address-Event-Representation sensors and tools. The target application is the lip reading of several keywords for…

Streaming applications from health care analytics to algorithmic trading deploy Kleene queries to detect and aggregate event trends. Rich event matching semantics determine how to compose events into trends. The expressive power of…

数据库 · 计算机科学 2020-10-08 Olga Poppe , Chuan Lei , Elke A. Rundensteiner , David Maier

Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain challenging due to the difficulty of aligning dense spatio-temporal video representations…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Kedi Sun , Chaohui Dang , Yue Feng , James Glasbey , Theodoros N. Arvanitis , Le Zhang

Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query to moments results in insufficient discriminability and…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Ran Ran , Jiwei Wei , Shuchang Zhou , Yitong Qin , Shiyuan He , Zeyu Ma , Yuyang Zhou , Yang Yang

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Wenjun Yu , Yinchen Zhou , Jia-Xuan Jiang , Shubin Zeng , Yuee Li , Zhong Wang

Audiovisual data is everywhere in this digital age, which raises higher requirements for the deep learning models developed on them. To well handle the information of the multi-modal data is the key to a better audiovisual modal. We observe…

声音 · 计算机科学 2023-09-27 Meng Liu , Ke Liang , Dayu Hu , Hao Yu , Yue Liu , Lingyuan Meng , Wenxuan Tu , Sihang Zhou , Xinwang Liu

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Bowen Hao , Dongliang Zhou , Xiaojie Li , Xingyu Zhang , Liang Xie , Jianlong Wu , Erwei Yin

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Yuchen Hu , Yu Gu , Chenxing Li , Rilin Chen , Dong Yu

Several training strategies and temporal models have been recently proposed for isolated word lip-reading in a series of independent works. However, the potential of combining the best strategies and investigating the impact of each of them…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Pingchuan Ma , Yujiang Wang , Stavros Petridis , Jie Shen , Maja Pantic

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion and challenging illumination conditions. However, existing Event-based Object Detection…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Meisen Wang , Hao Deng , Wei Bao , Ma Yuanxiao , Chengjie Wang , Zhiqiang Tian , Shaoyi Du , Siqi Li

The facts and time in the document are intricately intertwined, making temporal reasoning over documents challenging. Previous work models time implicitly, making it difficult to handle such complex relationships. To address this issue, we…

计算与语言 · 计算机科学 2023-11-09 Zheng Chu , Zekun Wang , Jiafeng Liang , Ming Liu , Bing Qin