English
Related papers

Related papers: Enhancing Lie Detection Accuracy: A Comparative St…

200 papers

We present a multi-purpose algorithm for simultaneous face detection, face alignment, pose estimation, gender recognition, smile detection, age estimation and face recognition using a single deep convolutional neural network (CNN). The…

Computer Vision and Pattern Recognition · Computer Science 2016-11-04 Rajeev Ranjan , Swami Sankaranarayanan , Carlos D. Castillo , Rama Chellappa

Handwritten Mathematical Expression Recognition (HMER) methods have made remarkable progress, with most existing HMER approaches based on either a hybrid CNN/RNN-based with GRU architecture or Transformer architectures. Each of these has…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Kehua Chen , Haoyang Shen , Lifan Zhong , Mingyi Chen

Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hugo Bohy , Minh Tran , Kevin El Haddad , Thierry Dutoit , Mohammad Soleymani

Model observers are computational tools to evaluate and optimize task-based medical image quality. Linear model observers, such as the Channelized Hotelling Observer (CHO), predict human accuracy in detection tasks with a few possible…

Image and Video Processing · Electrical Eng. & Systems 2024-10-24 Aditya Jonnalagadda , Bruno B. Barufaldi , Andrew D. A. Maidment , Susan P. Weinstein , Craig K. Abbey , Miguel P. Eckstein

This paper proposes one of the first clinical applications of multimodal large language models (LLMs) as an assistant for radiologists to check errors in their reports. We created an evaluation dataset from real-world radiology datasets…

Computation and Language · Computer Science 2024-03-05 Jinge Wu , Yunsoo Kim , Eva C. Keller , Jamie Chow , Adam P. Levine , Nikolas Pontikos , Zina Ibrahim , Paul Taylor , Michelle C. Williams , Honghan Wu

In this paper, we aim to address the problem of heterogeneous or cross-spectral face recognition using machine learning to synthesize visual spectrum face from infrared images. The synthesis of visual-band face images allows for more…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Kenneth Lai , Svetlana N. Yanushkevich

Multimodal out-of-context (OOC) misinformation is misinformation that repurposes real images with unrelated or misleading captions. Detecting such misinformation is challenging because it requires resolving the context of the claim before…

Machine Learning · Computer Science 2025-05-27 Sharad Duwal , Mir Nafis Sharear Shopnil , Abhishek Tyagi , Adiba Mahbub Proma

This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve the first problem, we…

Machine Learning · Statistics 2017-09-12 Szu-Wei Fu , Ting-yao Hu , Yu Tsao , Xugang Lu

Deepfakes are synthetically generated images, videos or audios, which fraudsters use to manipulate legitimate information. Current deepfake detection systems struggle against unseen data. To address this, we employ three different deep…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Sohail Ahmed Khan , Alessandro Artusi , Hang Dai

We describe a deep learning based method for estimating 3D facial expression coefficients. Unlike previous work, our process does not relay on facial landmark detection methods as a proxy step. Recent methods have shown that a CNN can be…

Computer Vision and Pattern Recognition · Computer Science 2018-02-05 Feng-Ju Chang , Anh Tuan Tran , Tal Hassner , Iacopo Masi , Ram Nevatia , Gerard Medioni

In this work, we address the problem of improvement of robustness of feature representations learned using convolutional neural networks (CNNs) to image deformation. We argue that higher moment statistics of feature distributions could be…

Computer Vision and Pattern Recognition · Computer Science 2017-07-26 Zhun Sun , Mete Ozay , Takayuki Okatani

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems is the Audio Visual Scene-Aware Dialog (AVSD). Most…

Computation and Language · Computer Science 2022-02-22 Yoshihiro Yamazaki , Shota Orihashi , Ryo Masumura , Mihiro Uchida , Akihiko Takashima

Recent studies on deepfake detection have achieved promising results when training and testing faces are from the same dataset. However, their results severely degrade when confronted with forged samples that the model has not yet seen…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Tiewen Chen , Shanmin Yang , Shu Hu , Zhenghan Fang , Ying Fu , Xi Wu , Xin Wang

In this paper, the multi-task learning of lightweight convolutional neural networks is studied for face identification and classification of facial attributes (age, gender, ethnicity) trained on cropped faces without margins. The necessity…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Andrey V. Savchenko

We present a statistical model for $3$D human faces in varying expression, which decomposes the surface of the face using a wavelet transform, and learns many localized, decorrelated multilinear models on the resulting coefficients. Using…

Computer Vision and Pattern Recognition · Computer Science 2014-07-02 Alan Brunton , Timo Bolkart , Stefanie Wuhrer

Multimodal learning involves integrating information from various modalities to enhance learning and comprehension. We compare three modality fusion strategies in person identification and verification by processing two modalities: voice…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Aref Farhadipour , Masoumeh Chapariniya , Teodora Vukovic , Volker Dellwo

This paper presents several novel findings on the explainability of vision reflection in large multimodal models (LMMs). First, we show that prompting an LMM to verify the prediction of a specialized vision model can improve recognition…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Guoyuan An , JaeYoon Kim , SungEui Yoon

Recent works on convolutional neural networks (CNNs) for facial alignment have demonstrated unprecedented accuracy on a variety of large, publicly available datasets. However, the developed models are often both cumbersome and…

Computer Vision and Pattern Recognition · Computer Science 2019-06-12 TianXing Li , Zhi Yu , Edmund Phung , Brendan Duke , Irina Kezele , Parham Aarabi

This project investigates the capabilities of large language models (LLMs) to determine the difficulty of data visualization literacy test items. We explore whether features derived from item text (question and answer options), the…

Artificial Intelligence · Computer Science 2026-03-06 Samin Khan

Feature extraction with convolutional neural networks (CNNs) is a popular method to represent images for machine learning tasks. These representations seek to capture global image content, and ideally should be independent of geometric…

Machine Learning · Computer Science 2022-03-03 Jake Lee , Junfeng Yang , Zhangyang Wang