中文
相关论文

相关论文: Towards Disentangled Representations for Human Ret…

200 篇论文

Multimodal data are prevalent across various domains, and learning robust representations of such data is paramount to enhancing generation quality and downstream task performance. To handle heterogeneity and interconnections among…

机器学习 · 计算机科学 2025-09-30 Yijie Zhang , Yiyang Shen , Weiran Wang

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to demonstrate the…

音频与语音处理 · 电气工程与系统科学 2022-10-26 Hui Lu , Disong Wang , Xixin Wu , Zhiyong Wu , Xunying Liu , Helen Meng

Variational Autoencoders for multimodal data hold promise for many tasks in data analysis, such as representation learning, conditional generation, and imputation. Current architectures either share the encoder output, decoder input, or…

Although Generative Adversarial Networks (GANs) have made significant progress in face synthesis, there lacks enough understanding of what GANs have learned in the latent representation to map a random code to a photo-realistic image. In…

计算机视觉与模式识别 · 计算机科学 2020-10-30 Yujun Shen , Ceyuan Yang , Xiaoou Tang , Bolei Zhou

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for this task: 1) lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this work, we…

计算机视觉与模式识别 · 计算机科学 2019-12-19 Hsin-Ying Lee , Hung-Yu Tseng , Qi Mao , Jia-Bin Huang , Yu-Ding Lu , Maneesh Singh , Ming-Hsuan Yang

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Zihui Xue , Kristen Grauman

Deep generative modelling for human body analysis is an emerging problem with many interesting applications. However, the latent space learned by such approaches is typically not interpretable, resulting in less flexibility. In this work,…

计算机视觉与模式识别 · 计算机科学 2020-02-18 Rodrigo de Bem , Arnab Ghosh , Thalaiyasingam Ajanthan , Ondrej Miksik , Adnane Boukhayma , N. Siddharth , Philip Torr

Representation disentanglement may help AI fundamentally understand the real world and thus benefit both discrimination and generation tasks. It currently has at least three unresolved core issues: (i) heavy reliance on label annotation and…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Xin Jin , Bohan Li , BAAO Xie , Wenyao Zhang , Jinming Liu , Ziqiang Li , Tao Yang , Wenjun Zeng

Deep learning (DL) methods where interpretability is intrinsically considered as part of the model are required to better understand the relationship of clinical and imaging-based attributes with DL outcomes, thus facilitating their use in…

图像与视频处理 · 电气工程与系统科学 2022-12-13 Irem Cetin , Maialen Stephens , Oscar Camara , Miguel Angel Gonzalez Ballester

We present a new model DrNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each frame into a…

机器学习 · 计算机科学 2024-03-15 Remi Denton , Vighnesh Birodkar

We address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intra-class…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Hyunjong Park , Sanghoon Lee , Junghyup Lee , Bumsub Ham

Existing methods for multi-modal time series representation learning aim to disentangle the modality-shared and modality-specific latent variables. Although achieving notable performances on downstream tasks, they usually assume an…

机器学习 · 计算机科学 2024-05-28 Ruichu Cai , Zhifang Jiang , Zijian Li , Weilin Chen , Xuexin Chen , Zhifeng Hao , Yifan Shen , Guangyi Chen , Kun Zhang

In Emotion Recognition in Conversations (ERC), the emotions of target utterances are closely dependent on their context. Therefore, existing works train the model to generate the response of the target utterance, which aims to recognise…

计算与语言 · 计算机科学 2023-05-30 Kailai Yang , Tianlin Zhang , Sophia Ananiadou

Skeleton-based human action recognition has been drawing more interest recently due to its low sensitivity to appearance changes and the accessibility of more skeleton data. However, even the 3D skeletons captured in practice are still…

计算机视觉与模式识别 · 计算机科学 2022-09-26 Cunling Bian , Wei Feng , Fanbo Meng , Song Wang

Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, purely unsupervised approaches have proven successful on fully disentangled synthetic data, but fail to…

机器学习 · 计算机科学 2026-01-30 Alexandre Myara , Nicolas Bourriez , Thomas Boyer , Thomas Lemercier , Ihab Bendidi , Auguste Genovesio

We propose a Convolutional Neural Network (CNN)-based model "RotationNet," which takes multi-view images of an object as input and jointly estimates its pose and object category. Unlike previous approaches that use known viewpoint labels…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Asako Kanezaki , Yasuyuki Matsushita , Yoshifumi Nishida

Clustering multi-view data has been a fundamental research topic in the computer vision community. It has been shown that a better accuracy can be achieved by integrating information of all the views than just using one view individually.…

计算机视觉与模式识别 · 计算机科学 2019-07-24 Ming Yin , Weitian Huang , Junbin Gao

Learning to reconstruct 3D shapes using 2D images is an active research topic, with benefits of not requiring expensive 3D data. However, most work in this direction requires multi-view images for each object instance as training…

计算机视觉与模式识别 · 计算机科学 2021-07-14 Bo Peng , Wei Wang , Jing Dong , Tieniu Tan

Recent works have shown that a rich set of semantic directions exist in the latent space of Generative Adversarial Networks (GANs), which enables various facial attribute editing applications. However, existing methods may suffer poor…

计算机视觉与模式识别 · 计算机科学 2021-05-28 Yuxuan Han , Jiaolong Yang , Ying Fu

Most recent view-invariant action recognition and performance assessment approaches rely on a large amount of annotated 3D skeleton data to extract view-invariant features. However, acquiring 3D skeleton data can be cumbersome, if not…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Faegheh Sardari , Björn Ommer , Majid Mirmehdi