English
Related papers

Related papers: Multimedia Communication Quality Assessment Testbe…

200 papers

Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g., Mermaid). In production, these systems process arbitrary inputs for which no ground-truth code…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Giang Son Nguyen , Zi Pong Lim , Sarthak Ketanbhai Modi , Yon Shin Teo , Wenya Wang

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zhenyu Yang , Yuhang Hu , Zemin Du , Dizhan Xue , Shengsheng Qian , Jiahong Wu , Fan Yang , Weiming Dong , Changsheng Xu

Recent years have witnessed an ever-expandingvolume of user-generated content (UGC) videos available on the Internet. Nevertheless, progress on perceptual quality assessmentof UGC videos still remains quite limited. There are many…

Multimedia · Computer Science 2019-09-13 Yang Li , Shengbin Meng , Xinfeng Zhang , Shiqi Wang , Yue Wang , Siwei Ma

Ultra-high-resolution streaming and emerging immersive services are driving rapidly increasing wireless video traffic. However, perceptually pleasing video transmission over bandwidth-limited and latency-constrained wireless links remains…

Image and Video Processing · Electrical Eng. & Systems 2026-05-20 Yinhuan Huang , Zhijin Qin

Multimedia compression allows us to watch videos, see pictures and hear sounds within a limited bandwidth, which helps the flourish of the internet. During the past decades, multimedia compression has achieved great success using hand-craft…

Multimedia · Computer Science 2023-08-21 Yuhao Cheng , Siru Zhang , Yiqiang Yan , Rong Chen , Yun Zhang

In recent years, user-generated content (UGC) has become one of the major video types consumed via streaming networks. Numerous research contributions have focused on assessing its visual quality through subjective tests and objective…

Image and Video Processing · Electrical Eng. & Systems 2024-08-15 Zihao Qi , Chen Feng , Fan Zhang , Xiaozhong Xu , Shan Liu , David Bull

Video streaming, in various forms of video on demand (VOD), live, and 360 degree streaming, has grown dramatically during the past few years. In comparison to traditional cable broadcasters whose contents can only be watched on TVs, video…

Multimedia · Computer Science 2020-12-01 Xiangbo Li , Mahmoud Darwich , Magdy Bayoumi , Mohsen Amini Salehi

Immersive video offers the freedom to navigate inside virtualized environment. Instead of streaming the bulky immersive videos entirely, a viewport (also referred to as field of view, FoV) adaptive streaming is preferred. We often stream…

Multimedia · Computer Science 2018-02-19 Shaowei Xie , Qiu Shen , Yiling Xu , Qiaojian Qian , Shaowei Wang , Zhan Ma , Wenjun Zhang

With neural video codecs (NVCs) emerging as promising alternatives for traditional compression methods, it is increasingly important to determine whether existing quality metrics remain valid for evaluating their performance. However, few…

Image and Video Processing · Electrical Eng. & Systems 2026-05-19 Benjamin Herb , Rakesh Rao Ramachandra Rao , Steve Göring , Alexander Raake

The growing popularity of virtual and augmented reality communications and 360{\deg} video streaming is moving video communication systems into much more dynamic and resource-limited operating settings. The enormous data volume of 360{\deg}…

Multimedia · Computer Science 2018-03-23 Jacob Chakareski , Ridvan Aksu , Xavier Corbillon , Gwendal Simon , Viswanathan Swaminathan

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Yang Yang , Yury Kartynnik , Yunpeng Li , Jiuqiang Tang , Xing Li , George Sung , Matthias Grundmann

Rapid advancements in video diffusion models have enabled the creation of realistic videos, raising concerns about unauthorized use and driving the demand for techniques to protect model ownership. Existing watermarking methods, while…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 MinHyuk Jang , Youngdong Jang , JaeHyeok Lee , Feng Yang , Gyeongrok Oh , Jongheon Jeong , Sangpil Kim

To open up new possibilities to assess the multimodal perceptual quality of omnidirectional media formats, we proposed a novel open source 360 audiovisual (AV) quality dataset. The dataset consists of high-quality 360 video clips in…

Multimedia · Computer Science 2022-05-18 Randy F Fela , Andréas Pastor , Patrick Le Callet , Nick Zacharov , Toinon Vigier , Søren Forchhammer

Contemporary deep learning based video captioning follows encoder-decoder framework. In encoder, visual features are extracted with 2D/3D Convolutional Neural Networks (CNNs) and a transformed version of those features is passed to the…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Nayyer Aafaq , Naveed Akhtar , Wei Liu , Ajmal Mian

Conventional recommendation systems frequently fail to fully exploit the high-dimensional semantic signals inherent in multimedia content, thereby limiting the fidelity of user preference modeling. While Multimodal Large Language Models…

Information Retrieval · Computer Science 2026-05-12 Yiming Zhu , Xu Liu , Ziyun Xu , Zheng Wu , Joena Zhang , Sirius Chen , Chenheli Hua , Silvester Yao , Qichao Que , Wentao Shi , Junfeng Pan , Linhong Zhu

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehension, necessitating…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Junming Lin , Zheng Fang , Chi Chen , Zihao Wan , Fuwen Luo , Peng Li , Yang Liu , Maosong Sun

Thanks to the abundance of Web platforms and broadband connections, HTTP Adaptive Streaming has become the de facto choice for multimedia delivery nowadays. However, the visual quality of adaptive video streaming may fluctuate strongly…

Multimedia · Computer Science 2020-02-26 Huyen T. T. Tran , Nam Pham Ngoc , Tobias Hoßfeld , Michael Seufert , Truong Cong Thang

We present an AI-based framework for semantic transmission of multimedia data over band-limited, time-varying channels. The method targets scenarios where large content is split into multiple packets, with an unknown number potentially…

Multimedia · Computer Science 2026-01-29 Homa Esfahanizadeh , Nargis Fayaz , Jinfeng Du , Harish Viswanathan

In this paper, we propose a deep learning based video quality assessment (VQA) framework to evaluate the quality of the compressed user's generated content (UGC) videos. The proposed VQA framework consists of three modules, the feature…

Image and Video Processing · Electrical Eng. & Systems 2021-06-03 Wei Sun , Tao Wang , Xiongkuo Min , Fuwang Yi , Guangtao Zhai

Neural audio codecs are a fundamental component of modern generative audio pipelines. Although recent codecs achieve strong low-bitrate reconstruction and provide powerful representations for downstream tasks, most are non-streamable,…

Sound · Computer Science 2025-09-22 Luca Della Libera , Cem Subakan , Mirco Ravanelli
‹ Prev 1 4 5 6 7 8 10 Next ›