English
Related papers

Related papers: EmoArt: A Multidimensional Dataset for Emotion-Awa…

200 papers

Text-to-Image (T2I) diffusion models (DM) have garnered widespread adoption due to their capability in generating high-fidelity outputs and accessibility to anyone able to put imagination into words. However, DMs are often predisposed to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Zhe Jin , Tat-Seng Chua

A hallmark of human intelligence is the ability to create complex artifacts through structured multi-step processes. Generating procedural tutorials with AI is a longstanding but challenging goal, facing three key obstacles: (1) scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Yiren Song , Cheng Liu , Mike Zheng Shou

This study investigates the key characteristics and suitability of widely used Facial Expression Recognition (FER) datasets for training deep learning models. In the field of affective computing, FER is essential for interpreting human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 F. Xavier Gaya-Morey , Cristina Manresa-Yee , Célia Martinie , Jose M. Buades-Rubio

With the explosive growth of social media, opinionated postings with emojis have increased explosively. Many emojis are used to express emotions, attitudes, and opinions. Emoji representation learning can be helpful to improve the…

Computation and Language · Computer Science 2022-05-24 Xiaowei Yuan , Jingyuan Hu , Xiaodan Zhang , Honglei Lv

The integration of generative AI in visual art has revolutionized not only how visual content is created but also how AI interacts with and reflects the underlying domain knowledge. This survey explores the emerging realm of diffusion-based…

Artificial Intelligence · Computer Science 2024-08-23 Bingyuan Wang , Qifeng Chen , Zeyu Wang

Diffusion models have demonstrated exceptional capabilities in generating a broad spectrum of visual content, yet their proficiency in rendering text is still limited: they often generate inaccurate characters or words that fail to blend…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jianyi Zhang , Yufan Zhou , Jiuxiang Gu , Curtis Wigington , Tong Yu , Yiran Chen , Tong Sun , Ruiyi Zhang

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yijie Zhu , Lingsen Zhang , Zitong Yu , Rui Shao , Tao Tan , Liqiang Nie

We present Follow-Your-Emoji-Faster, an efficient diffusion-based framework for freestyle portrait animation driven by facial landmarks. The main challenges in this task are preserving the identity of the reference portrait, accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yue Ma , Zexuan Yan , Hongyu Liu , Hongfa Wang , Heng Pan , Yingqing He , Junkun Yuan , Ailing Zeng , Chengfei Cai , Heung-Yeung Shum , Zhifeng Li , Wei Liu , Linfeng Zhang , Qifeng Chen

Text-to-image diffusion models have demonstrated remarkable capabilities in generating artistic content by learning from billions of images, including popular artworks. However, the fundamental question of how these models internally…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Alfio Ferrara , Sergio Picascia , Elisabetta Rocchetti

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is urgently needed to…

Multimedia · Computer Science 2024-08-28 Zeyu Jin , Jia Jia , Qixin Wang , Kehan Li , Shuoyi Zhou , Songtao Zhou , Xiaoyu Qin , Zhiyong Wu

Automated affective computing in the wild setting is a challenging problem in computer vision. Existing annotated databases of facial expressions in the wild are small and mostly cover discrete emotions (aka the categorical model). There…

Computer Vision and Pattern Recognition · Computer Science 2018-02-06 Ali Mollahosseini , Behzad Hasani , Mohammad H. Mahoor

Automated human emotion recognition from facial expressions is a well-studied problem and still remains a very challenging task. Some efficient or accurate deep learning models have been presented in the literature. However, it is quite…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Monu Verma , Murari Mandal , Satish Kumar Reddy , Yashwanth Reddy Meedimale , Santosh Kumar Vipparthi

Recent years have witnessed increasing attention in cartoon media, powered by the strong demands of industrial applications. As the first step to understand this media, cartoon face recognition is a crucial but less-explored task with few…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Yi Zheng , Yifan Zhao , Mengyuan Ren , He Yan , Xiangju Lu , Junhui Liu , Jia Li

Text-to-image generation has recently witnessed remarkable achievements. We introduce a text-conditional image diffusion model, termed RAPHAEL, to generate highly artistic images, which accurately portray the text prompts, encompassing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Zeyue Xue , Guanglu Song , Qiushan Guo , Boxiao Liu , Zhuofan Zong , Yu Liu , Ping Luo

Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-conditioned systems suffer from two fundamental limitations: (i)…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Ivan Rinaldi , Matteo Mendula , Nicola Fanelli , Florence Levé , Matteo Testi , Giovanna Castellano , Gennaro Vessio

Due to the complex nature of human emotions and the diversity of emotion representation methods in humans, emotion recognition is a challenging field. In this research, three input modalities, namely text, audio (speech), and video, are…

Artificial Intelligence · Computer Science 2024-02-13 Minoo Shayaninasab , Bagher Babaali

Diffusion models have demonstrated great success in the field of text-to-image generation. However, alleviating the misalignment between the text prompts and images is still challenging. The root reason behind the misalignment has not been…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Dongzhi Jiang , Guanglu Song , Xiaoshi Wu , Renrui Zhang , Dazhong Shen , Zhuofan Zong , Yu Liu , Hongsheng Li

Emotions widely affect human decision-making. This fact is taken into account by affective computing with the goal of tailoring decision support to the emotional states of individuals. However, the accurate recognition of emotions within…

Computation and Language · Computer Science 2018-11-14 Bernhard Kratzwald , Suzana Ilic , Mathias Kraus , Stefan Feuerriegel , Helmut Prendinger

Generating high-quality, multi-layer transparent images from text prompts can unlock a new level of creative control, allowing users to edit each layer as effortlessly as editing text outputs from LLMs. However, the development of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Junwen Chen , Heyang Jiang , Yanbin Wang , Keming Wu , Ji Li , Chao Zhang , Keiji Yanai , Dong Chen , Yuhui Yuan

Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Omri Avrahami , Thomas Hayes , Oran Gafni , Sonal Gupta , Yaniv Taigman , Devi Parikh , Dani Lischinski , Ohad Fried , Xi Yin