English
Related papers

Related papers: A Multimodal Transformer for Live Streaming Highli…

200 papers

Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition. However, compared to LSTM models, the heavy computational cost of the Transformer during inference is a key issue to…

Computation and Language · Computer Science 2021-03-02 Xie Chen , Yu Wu , Zhenghao Wang , Shujie Liu , Jinyu Li

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Mehdi Fatan , Emanuele Mincato , Dimitra Pintzou , Mariella Dimiccoli

In dynamic traffic environments, motion forecasting models must be able to accurately estimate future trajectories continuously. Streaming-based methods are a promising solution, but despite recent advances, their performance often degrades…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Alexander Prutsch , Christian Fruhwirth-Reisinger , David Schinagl , Horst Possegger

Streaming video recognition reasons about objects and their actions in every frame of a video. A good streaming recognition model captures both long-term dynamics and short-term changes of video. Unfortunately, in most existing methods, the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Yue Zhao , Philipp Krähenbühl

Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mechanism that…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Akash Gupta , Rohun Tripathi , Wondong Jang

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processing various…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Yiyuan Zhang , Kaixiong Gong , Kaipeng Zhang , Hongsheng Li , Yu Qiao , Wanli Ouyang , Xiangyu Yue

Voice conversion models have developed for decades, and current mainstream research focuses on non-streaming voice conversion. However, streaming voice conversion is more suitable for practical application scenarios than non-streaming voice…

Sound · Computer Science 2022-06-16 Ziyi Chen , Haoran Miao , Pengyuan Zhang

Many learning tasks involve multi-modal data streams, where continuous data from different modes convey a comprehensive description about objects. A major challenge in this context is how to efficiently interpret multi-modal information in…

Machine Learning · Computer Science 2020-07-24 Amila Silva , Shanika Karunasekera , Christopher Leckie , Ling Luo

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Mengmeng Wang , Jiazheng Xing , Boyuan Jiang , Jun Chen , Jianbiao Mei , Xingxing Zuo , Guang Dai , Jingdong Wang , Yong Liu

Recently, the Transformer module has been transplanted from natural language processing to computer vision. This paper applies the Transformer to video-based person re-identification, where the key issue is to extract the discriminative…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Tianyu Zhang , Longhui Wei , Lingxi Xie , Zijie Zhuang , Yongfei Zhang , Bo Li , Qi Tian

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and datasets. However, such approaches that target generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Ashutosh Anshul , Shreyas Gopal , Deepu Rajan , Eng Siong Chng

Although Transformers excel in natural language processing, their extension to time series forecasting remains challenging due to insufficient consideration of the differences between textual and temporal modalities. In this paper, we…

Machine Learning · Computer Science 2025-10-09 Zhipeng Liu , Peibo Duan , Xuan Tang , Baixin Li , Yongsheng Huang , Mingyang Geng , Changsheng Zhang , Bin Zhang , Binwu Wang

Streaming neural network models for fast frame-wise responses to various speech and sensory signals are widely adopted on resource-constrained platforms. Hence, increasing the learning capacity of such streaming models (i.e., by adding more…

Deadline-aware transmission scheduling in immersive video streaming is crucial. The objective is to guarantee that at least a certain block in multi-links is fully delivered within their deadlines, which is referred to as delivery ratio.…

Networking and Internet Architecture · Computer Science 2024-09-02 Tongtong Feng , Qi Qi , Bo He , Jingyu Wang

Extracting real-time insights from multi-modal data streams from various domains such as healthcare, intelligent transportation, and satellite remote sensing remains a challenge. High computational demands and limited knowledge scope…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Murugan Sankaradas , Ravi K. Rajendran , Srimat T. Chakradhar

Streaming perception is a critical task in autonomous driving that requires balancing the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they only rely on the current and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Chenyang Li , Zhi-Qi Cheng , Jun-Yan He , Pengyu Li , Bin Luo , Hanyuan Chen , Yifeng Geng , Jin-Peng Lan , Xuansong Xie

Multimodal Transformers are emerging artificial intelligence (AI) models designed to process a mixture of signals from diverse modalities. Digital computing-in-memory (CIM) architectures are considered promising for achieving high…

Hardware Architecture · Computer Science 2025-02-11 Shantian Qin , Ziqing Qiang , Zhihua Fan , Wenming Li , Xuejun An , Xiaochun Ye , Dongrui Fan

Survival prediction plays a crucial role in assisting clinicians with the development of cancer treatment protocols. Recent evidence shows that multimodal data can help in the diagnosis of cancer disease and improve survival prediction.…

Image and Video Processing · Electrical Eng. & Systems 2023-11-14 Ruiquan Ge , Xiangyang Hu , Rungen Huang , Gangyong Jia , Yaqi Wang , Renshu Gu , Changmiao Wang , Elazab Ahmed , Linyan Wang , Juan Ye , Ye Li

Human state recognition is a critical topic with pervasive and important applications in human-machine systems. Multi-modal fusion, the combination of metrics from multiple data sources, has been shown as a sound method for improving the…

Human-Computer Interaction · Computer Science 2023-04-12 Ruiqi Wang , Wonse Jo , Dezhong Zhao , Weizheng Wang , Baijian Yang , Guohua Chen , Byung-Cheol Min