English
Related papers

Related papers: THQA: A Perceptual Quality Assessment Database for…

200 papers

Speech quality assessment (SQA) refers to the evaluation of speech quality, and developing an accurate automatic SQA method that reflects human perception has become increasingly important, in order to keep up with the generative AI boom.…

Sound · Computer Science 2025-08-29 Wen-Chin Huang

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Juntong Wang , Jiarui Wang , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

Measurement of interaction quality is a critical task for the improvement of spoken dialog systems. Existing approaches to dialog quality estimation either focus on evaluating the quality of individual turns, or collect dialog-level quality…

The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xinqi Xiong , Prakrut Patel , Qingyuan Fan , Amisha Wadhwa , Sarathy Selvam , Xiao Guo , Luchao Qi , Xiaoming Liu , Roni Sengupta

Recent years have witnessed an exponential increase in the demand for face video compression, and the success of artificial intelligence has expanded the boundaries beyond traditional hybrid video coding. Generative coding approaches have…

Image and Video Processing · Electrical Eng. & Systems 2023-10-31 Yixuan Li , Bolin Chen , Baoliang Chen , Meng Wang , Shiqi Wang , Weisi Lin

Video live streaming is gaining prevalence among video streaming services, especially for the delivery of popular sporting events. Many objective Video Quality Assessment (VQA) models have been developed to predict the perceptual quality of…

Image and Video Processing · Electrical Eng. & Systems 2021-06-17 Zaixi Shang , Joshua P. Ebenezer , Alan C. Bovik , Yongjun Wu , Hai Wei , Sriram Sethuraman

An accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Shaolin Su , Hanhe Lin , Vlad Hosu , Oliver Wiedemann , Jinqiu Sun , Yu Zhu , Hantao Liu , Yanning Zhang , Dietmar Saupe

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Jiayi Lyu , Leigang Qu , Wenjing Zhang , Hanyu Jiang , Kai Liu , Zhenglin Zhou , Xiaobo Xia , Jian Xue , Tat-Seng Chua

As multimedia services such as video streaming, video conferencing, virtual reality (VR), and online gaming continue to expand, ensuring high perceptual visual quality becomes a priority to maintain user satisfaction and competitiveness.…

Multimedia · Computer Science 2025-03-04 Wei Zhou , Hadi Amirpour , Christian Timmerer , Guangtao Zhai , Patrick Le Callet , Alan C. Bovik

Humans gather information by engaging in conversations involving a series of interconnected questions and answers. For machines to assist in information gathering, it is therefore essential to enable them to answer conversational questions.…

Computation and Language · Computer Science 2019-04-02 Siva Reddy , Danqi Chen , Christopher D. Manning

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

In recent years, image generation technology has rapidly advanced, resulting in the creation of a vast array of AI-generated images (AIGIs). However, the quality of these AIGIs is highly inconsistent, with low-quality AIGIs severely…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Jiquan Yuan , Fanyi Yang , Jihe Li , Xinyan Cao , Jinming Che , Jinlong Lin , Xixin Cao

The advent of AI has influenced many aspects of human life, from self-driving cars and intelligent chatbots to text-based image and video generation models capable of creating realistic images and videos based on user prompts…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Abhijay Ghildyal , Yuanhan Chen , Saman Zadtootaghaj , Nabajeet Barman , Alan C. Bovik

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu

The rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Yunhao Li , Sijing Wu , Wei Sun , Zhichao Zhang , Yucheng Zhu , Zicheng Zhang , Huiyu Duan , Xiongkuo Min , Guangtao Zhai

Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Kirill Koltsov , Aleksandr Gushchin , Anastasia Antsiferova , Dmitriy Vatolin

In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Bingyang Cui , Qi Yang , Kaifa Yang , Yiling Xu , Xiaozhong Xu , Shan Liu

We introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets. We have developed a strong and robust question engine that leverages scene…

Computation and Language · Computer Science 2019-07-12 Drew A. Hudson , Christopher D. Manning

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang