English
Related papers

Related papers: MPAI-EEV: Standardization Efforts of Artificial In…

200 papers

Video understanding has been considered as one critical step towards world modeling, which is an important long-term problem in AI research. Recently, multimodal foundation models have shown such potential via large-scale pretraining. These…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Boyu Chen , Siran Chen , Kunchang Li , Qinglin Xu , Yu Qiao , Yali Wang

The visual signal compression is a long-standing problem. Fueled by the recent advances of deep learning, exciting progress has been made. Despite better compression performance, existing end-to-end compression algorithms are still designed…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Shurun Wang , Zhao Wang , Shiqi Wang , Yan Ye

Token Communication (TokenCom) is a new paradigm, motivated by the recent success of Large AI Models (LAMs) and Multimodal Large Language Models (MLLMs), where tokens serve as unified units of communication and computation, enabling…

Information Theory · Computer Science 2026-03-04 Jingxuan Men , Mahdi Boloursaz Mashhadi , Ning Wang , Yi Ma , Mike Nilsson , Rahim Tafazolli

Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Yihan Wu , Yifan Peng , Yichen Lu , Xuankai Chang , Ruihua Song , Shinji Watanabe

Compressed video quality enhancement (CVQE) is crucial for improving user experience with lossy video codecs like H.264/AVC, H.265/HEVC, and H.266/VVC. While deep learning based CVQE has driven significant progress, existing surveys still…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Xiem HoangVan , Dang BuiDinh , Sang NguyenQuang , Wen-Hsiao Peng

This paper describes the development of a novel algorithm to tackle the problem of real-time video stabilization for unmanned aerial vehicles (UAVs). There are two main components in the algorithm: (1) By designing a suitable model for the…

Computer Vision and Pattern Recognition · Computer Science 2017-01-16 Anli Lim , Bharath Ramesh , Yue Yang , Cheng Xiang , Zhi Gao , Feng Lin

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Models (MLLMs). However, video understanding remains a…

Analyzing video for traffic categorization is an important pillar of Intelligent Transport Systems. However, it is difficult to analyze and predict traffic based on image frames because the representation of each frame may vary…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Somdip Dey , Amit K. Singh , Dilip K. Prasad , Klaus D. McDonald-Maier

Semantic segmentation of aerial videos has been extensively used for decision making in monitoring environmental changes, urban planning, and disaster management. The reliability of these decision support systems is dependent on the…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Girisha S , Ujjwal Verma , Manohara Pai M M , Radhika Pai

Unmanned Aerial Vehicles (UAVs) based video text spotting has been extensively used in civil and military domains. UAV's limited battery capacity motivates us to develop an energy-efficient video text spotting solution. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Zhenyu Hu , Zhenyu Wu , Pengcheng Pi , Yunhe Xue , Jiayi Shen , Jianchao Tan , Xiangru Lian , Zhangyang Wang , Ji Liu

Recent advances in video compression have seen significant coding performance improvements with the development of new standards and learning-based video codecs. However, most of these works focus on application scenarios that allow a…

Multimedia · Computer Science 2025-02-18 Siyue Teng , Yuxuan Jiang , Ge Gao , Fan Zhang , Thomas Davis , Zoe Liu , David Bull

Recent advancements in video anomaly understanding (VAU) have opened the door to groundbreaking applications in various fields, such as traffic monitoring and industrial automation. While the current benchmarks in VAU predominantly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Hang Du , Guoshun Nan , Jiawen Qian , Wangchenhui Wu , Wendi Deng , Hanqing Mu , Zhenyan Chen , Pengxuan Mao , Xiaofeng Tao , Jun Liu

Data encoding is a common and central operation in most data analysis tasks. The performance of other models downstream in the computational process highly depends on the quality of data encoding. One of the most powerful ways to encode…

Machine Learning · Computer Science 2025-09-03 Teddy Lazebnik , Liron Simon-Keren

Existing approaches to drone visual geo-localization predominantly adopt the image-based setting, where a single drone-view snapshot is matched with images from other platforms. Such task formulation, however, underutilizes the inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Hao Ju , Shaofei Huang , Si Liu , Zhedong Zheng

Electroencephalography (EEG) is an invaluable tool in neuroscience, offering insights into brain activity with high temporal resolution. Recent advancements in machine learning and generative modeling have catalyzed the application of EEG…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yashvir Sabharwal , Balaji Rama

This paper explores the integration of model-based and data-driven approaches within the realm of neural speech and audio coding systems. It highlights the challenges posed by the subjective evaluation processes of speech and audio codecs…

Sound · Computer Science 2025-01-08 Minje Kim , Jan Skoglund

Storage and transport of six degrees of freedom (6DoF) dynamic volumetric visual content for immersive applications requires efficient compression. ISO/IEC MPEG has recently been working on a standard that aims to efficiently code and…

Image and Video Processing · Electrical Eng. & Systems 2022-06-07 Maria Santamaria , Vinod Kumar Malamal Vadakital , Lukasz Kondrad , Antti Hallapuro , Miska M. Hannuksela

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

Video captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence. The encoder-decoder framework is the most popular paradigm for this task in recent years. However, there exist some…

Computer Vision and Pattern Recognition · Computer Science 2021-02-15 Haoran Chen , Jianmin Li , Xiaolin Hu

In 2021, a new track has been initiated in the Challenge for Learned Image Compression~: the video track. This category proposes to explore technologies for the compression of short video clips at 1 Mbit/s. This paper proposes to generate…

Image and Video Processing · Electrical Eng. & Systems 2021-05-21 Théo Ladune , Pierrick Philippe
‹ Prev 1 8 9 10 Next ›