English
Related papers

Related papers: LMME3DHF: Benchmarking and Evaluating Multimodal 3…

200 papers

Although recent large multimodal models (LMMs) demonstrate impressive progress on vision language tasks, their alignment with human centered (HC) principles, such as fairness, ethics, inclusivity, empathy, and robustness; remains poorly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Shaina Raza , Aravind Narayanan , Vahid Reza Khazaie , Ashmal Vayani , Ahmed Y. Radwan , Mukund S. Chettiar , Amandeep Singh , Mubarak Shah , Deval Pandya

Conventional, classification-based AI-generated image detection methods cannot explain why an image is considered real or AI-generated in a way a human expert would, which reduces the trustworthiness and persuasiveness of these detection…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Michael Yang , Shijian Deng , William T. Doan , Kai Wang , Tianyu Yang , Harsh Singh , Yapeng Tian

Evaluating Generative 3D models remains challenging due to misalignment between automated metrics and human perception of quality. Current benchmarks rely on image-based metrics that ignore 3D structure or geometric measures that fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Dylan Ebert

In recent years, AI has demonstrated remarkable capabilities in simulating human behaviors, particularly those implemented with large language models (LLMs). However, due to the lack of systematic evaluation of LLMs' simulated behaviors,…

Computation and Language · Computer Science 2024-06-18 Yang Xiao , Yi Cheng , Jinlan Fu , Jiashuo Wang , Wenjie Li , Pengfei Liu

Collecting and labeling training data is one important step for learning-based methods because the process is time-consuming and biased. For face analysis tasks, although some generative models can be used to generate face data, they can…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Dingyun Zhang , Chenglai Zhong , Yudong Guo , Yang Hong , Juyong Zhang

Talking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs. This paper provides a comprehensive review…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Vineet Kumar Rakesh , Soumya Mazumdar , Research Pratim Maity , Sarbajit Pal , Amitabha Das , Tapas Samanta

Face video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Sijing Wu , Yunhao Li , Ziwen Xu , Yixuan Gao , Huiyu Duan , Wei Sun , Guangtao Zhai

Recent advances in deep learning methods have increased the performance of face detection and recognition systems. The accuracy of these models relies on the range of variation provided in the training data. Creating a dataset that…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Shubhajit Basak , Hossein Javidnia , Faisal Khan , Rachel McDonnell , Michael Schukat

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

Sound · Computer Science 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

The growing interest in automatic survey generation (ASG), a task that traditionally required considerable time and effort, has been spurred by recent advances in large language models (LLMs). With advancements in retrieval-augmented…

Computation and Language · Computer Science 2025-08-18 Beichen Guo , Zhiyuan Wen , Yu Yang , Peng Gao , Ruosong Yang , Jiaxing Shen

Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However, many benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Jinsheng Huang , Liang Chen , Taian Guo , Fu Zeng , Yusheng Zhao , Bohan Wu , Ye Yuan , Haozhe Zhao , Zhihui Guo , Yichi Zhang , Jingyang Yuan , Wei Ju , Luchen Liu , Tianyu Liu , Baobao Chang , Ming Zhang

Evaluating text-to-image generation models requires alignment with human perception, yet existing human-centric metrics are constrained by limited data coverage, suboptimal feature extraction, and inefficient loss functions. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Yuhang Ma , Yunhao Shui , Xiaoshi Wu , Keqiang Sun , Hongsheng Li

Advances in generative modeling have made it increasingly easy to fabricate realistic portrayals of individuals, creating serious risks for security, communication, and public trust. Detecting such person-driven manipulations requires…

Human face generation and editing represent an essential task in the era of computer vision and the digital world. Recent studies have shown remarkable progress in multi-modal face generation and editing, for instance, using face…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Mohammadreza Mofayezi , Reza Alipour , Mohammad Ali Kakavand , Ehsaneddin Asgari

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Shunsuke Saito , Tomas Simon , Jason Saragih , Hanbyul Joo

Multimodal Large Language Models (MLLMs) have recently been explored as face verification systems that determine whether two face images are of the same person. Unlike dedicated face recognition systems, MLLMs approach this task through…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Ünsal Öztürk , Hatef Otroshi Shahreza , Sébastien Marcel

3D face reconstruction (3DFR) algorithms are based on specific assumptions tailored to the limits and characteristics of the different application scenarios. In this study, we investigate how multiple state-of-the-art 3DFR algorithms can be…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Simone Maurizio La Cava , Roberto Casula , Sara Concas , Giulia Orrù , Ruben Tolosana , Martin Drahansky , Julian Fierrez , Gian Luca Marcialis

The Audio-to-3D-Gesture (A2G) task has enormous potential for various applications in virtual reality and computer graphics, etc. However, current evaluation metrics, such as Fr\'echet Gesture Distance or Beat Constancy, fail at reflecting…

Multimedia · Computer Science 2026-03-27 Zhilin Gao , Yunhao Li , Sijing Wu , Yuqin Cao , Huiyu Duan , Guangtao Zhai

Photo-real digital human avatars are of enormous importance in graphics, as they enable immersive communication over the globe, improve gaming and entertainment experiences, and can be particularly beneficial for AR and VR settings.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Marc Habermann , Lingjie Liu , Weipeng Xu , Gerard Pons-Moll , Michael Zollhoefer , Christian Theobalt

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang