English
Related papers

Related papers: Perceptual Evaluation on Audio-visual Dataset of 3…

200 papers

While there are several widely used object detection datasets, current computer vision algorithms are still limited in conventional images. Such images narrow our vision in a restricted region. On the other hand, 360{\deg} images provide a…

Computer Vision and Pattern Recognition · Computer Science 2019-10-07 Shih-Han Chou , Cheng Sun , Wen-Yen Chang , Wan-Ting Hsu , Min Sun , Jianlong Fu

Based on the Just-Noticeable-Difference (JND) criterion, a subjective video quality assessment (VQA) dataset, called the VideoSet, was constructed recently. In this work, we propose a JND-based VQA model using a probabilistic framework to…

Multimedia · Computer Science 2018-07-04 Haiqiang Wang , Xinfeng Zhang , Chao Yang , C. -C. Jay Kuo

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Xinyu Chen , Yunxin Li , Haoyuan Shi , Baotian Hu , Wenhan Luo , Yaowei Wang , Min Zhang

We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We…

Image and Video Processing · Electrical Eng. & Systems 2024-10-04 Yu-Chih Chen , Avinab Saha , Alexandre Chapiro , Christian Häne , Jean-Charles Bazin , Bo Qiu , Stefano Zanetti , Ioannis Katsavounidis , Alan C. Bovik

This paper focuses on numeric data, with emphasis on distinct characteristics like varying significance, unstructured format, mass volume and real-time processing. We propose a novel, context-dependent valuation framework specifically…

Databases · Computer Science 2018-10-23 Milen S. Marev , Ernesto Compatangelo , Wamberto Vasconcelos

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

We present the outcomes of a recent large-scale subjective study of Mobile Cloud Gaming Video Quality Assessment (MCG-VQA) on a diverse set of gaming videos. Rapid advancements in cloud services, faster video encoding technologies, and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Avinab Saha , Yu-Chih Chen , Chase Davis , Bo Qiu , Xiaoming Wang , Rahul Gowda , Ioannis Katsavounidis , Alan C. Bovik

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Jaden Pieper , Stephen D. Voran

In this paper, we consider the problem of multimodal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

With rapid advancements in virtual reality (VR) headsets, effectively measuring stereoscopic quality of experience (SQoE) has become essential for delivering immersive and comfortable 3D experiences. However, most existing stereo metrics…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Netanel Y. Tamir , Shir Amir , Ranel Itzhaky , Noam Atia , Shobhita Sundaram , Stephanie Fu , Ron Sokolovsky , Phillip Isola , Tali Dekel , Richard Zhang , Miriam Farber

Recent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand…

Multimedia · Computer Science 2023-04-18 Kaiyuan Hu , Yili Jin , Haowen Yang , Junhua Liu , Fangxin Wang

The wide popularity of digital photography and social networks has generated a rapidly growing volume of multimedia data (i.e., image, music, and video), resulting in a great demand for managing, retrieving, and understanding these data.…

Multimedia · Computer Science 2019-11-14 Sicheng Zhao , Shangfei Wang , Mohammad Soleymani , Dhiraj Joshi , Qiang Ji

Despite substantial efforts dedicated to the design of heuristic models for omnidirectional (i.e., 360$^\circ$) image quality assessment (OIQA), a conspicuous gap remains due to the lack of consideration for the diversity of viewing…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Xiangjie Sui , Hanwei Zhu , Xuelin Liu , Yuming Fang , Shiqi Wang , Zhou Wang

In this article, we introduce the ContentWise Impressions dataset, a collection of implicit interactions and impressions of movies and TV series from an Over-The-Top media service, which delivers its media contents over the Internet. The…

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Zhucun Xue , Jiangning Zhang , Teng Hu , Haoyang He , Yinan Chen , Yuxuan Cai , Yabiao Wang , Chengjie Wang , Yong Liu , Xiangtai Li , Dacheng Tao

Nowadays, 360{\deg} video/image has been increasingly popular and drawn great attention. The spherical viewing range of 360{\deg} video/image accounts for huge data, which pose the challenges to 360{\deg} video/image processing in solving…

Image and Video Processing · Electrical Eng. & Systems 2019-10-29 Chen Li , Mai Xu , Shanyi Zhang , Patrick Le Callet

In this paper we explore audiovisual emotion recognition under noisy acoustic conditions with a focus on speech features. We attempt to answer the following research questions: (i) How does speech emotion recognition perform on noisy data?…

Sound · Computer Science 2021-03-03 Michael Neumann , Ngoc Thang Vu

This paper proposes a novel framework to evaluate fluid simulation methods based on crowd-sourced user studies in order to robustly gather large numbers of opinions. The key idea for a robust and reliable evaluation is to use a reference…

Graphics · Computer Science 2020-11-23 Kiwon Um , Xiangyu Hu , Nils Thuerey

In this paper we examine the ability of low-level multimodal features to extract movie similarity, in the context of a content-based movie recommendation approach. In particular, we demonstrate the extraction of multimodal representation…

Information Retrieval · Computer Science 2019-12-19 Konstantinos Bougiatiotis , Theodore Giannakopoulos

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

Computer Vision and Pattern Recognition · Computer Science 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu
‹ Prev 1 8 9 10 Next ›