English
Related papers

Related papers: Language-free Compositional Action Generation via …

200 papers

Linking neural representations to linguistic factors is crucial in order to build and analyze NLP models interpretable by humans. Among these factors, syntactic roles (e.g. subjects, direct objects,$\dots$) and their realizations are…

Computation and Language · Computer Science 2022-06-23 Ghazi Felhi , Joseph Le Roux , Djamé Seddah

In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Louis Airale , Xavier Alameda-Pineda , Stéphane Lathuilière , Dominique Vaufreydaz

Generative models have thrived in computer vision, enabling unprecedented image processes. Yet the results in audio remain less advanced. Our project targets real-time sound synthesis from a reduced set of high-level parameters, including…

Sound · Computer Science 2019-06-25 Adrien Bitton , Philippe Esling , Antoine Caillon , Martin Fouilleul

3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the second drawer of the cabinet near the bed''). Progress has…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Jaime Corsetti , Francesco Giuliari , Davide Boscaini , Pedro Hermosilla , Andrea Pilzer , Guofeng Mei , Alexandros Delitzas , Francis Engelmann , Fabio Poiesi

Even with strong sequence models like Transformers, generating expressive piano performances with long-range musical structures remains challenging. Meanwhile, methods to compose well-structured melodies or lead sheets (melody + chords),…

Sound · Computer Science 2023-03-08 Shih-Lun Wu , Yi-Hsuan Yang

We present a novel method to generate human motion to populate 3D indoor scenes. It can be controlled with various combinations of conditioning signals such as a path in a scene, target poses, past motions, and scenes represented as 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Nicolas Ugrinovic , Thomas Lucas , Fabien Baradel , Philippe Weinzaepfel , Gregory Rogez , Francesc Moreno-Noguer

We introduce Autodecompose, a novel self-supervised generative model that decomposes data into two semantically independent properties: the desired property, which captures a specific aspect of the data (e.g. the voice in an audio signal),…

Machine Learning · Computer Science 2023-02-14 Mohammad Reza Bonyadi

Autoregressive generative transformers are key in music generation, producing coherent compositions but facing challenges in human-machine collaboration. We propose RefinPaint, an iterative technique that improves the sampling process. It…

Sound · Computer Science 2024-11-12 Pedro Ramoneda , Martin Rocamora , Taketo Akama

We introduce a novel resampling criterion using lift scores, for improving compositional generation in diffusion models. By leveraging the lift scores, we evaluate whether generated samples align with each single condition and then compose…

Machine Learning · Computer Science 2025-05-27 Chenning Yu , Sicun Gao

The focus of the action understanding literature has predominately been classification, how- ever, there are many applications demanding richer action understanding such as mobile robotics and video search, with solutions to classification,…

Computer Vision and Pattern Recognition · Computer Science 2014-10-23 Ran Xu , Gang Chen , Caiming Xiong , Wei Chen , Jason J. Corso

We study domain-adaptive image synthesis, the problem of teaching pretrained image generative models a new style or concept from as few as one image to synthesize novel images, to better understand the compositional image synthesis. We…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Kihyuk Sohn , Albert Shaw , Yuan Hao , Han Zhang , Luisa Polania , Huiwen Chang , Lu Jiang , Irfan Essa

We present PartComposer: a framework for part-level concept learning from single-image examples that enables text-to-image diffusion models to compose novel objects from meaningful components. Existing methods either struggle with…

Graphics · Computer Science 2025-09-16 Junyu Liu , R. Kenny Jones , Daniel Ritchie

We propose a framework that learns to execute natural language instructions in an environment consisting of goal-reaching tasks that share components of their task descriptions. Our approach leverages the compositionality of both value…

Machine Learning · Computer Science 2021-10-12 Vanya Cohen , Geraud Nangue Tasse , Nakul Gopalan , Steven James , Matthew Gombolay , Benjamin Rosman

We target the problem of sparse 3D reconstruction of dynamic objects observed by multiple unsynchronized video cameras with unknown temporal overlap. To this end, we develop a framework to recover the unknown structure without sequencing…

Computer Vision and Pattern Recognition · Computer Science 2016-05-24 Enliang Zheng , Dinghuang Ji , Enrique Dunn , Jan-Michael Frahm

Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving such a problem requires a model to recognize unseen actions…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Rongchang Li , Zhenhua Feng , Tianyang Xu , Linze Li , Xiao-Jun Wu , Muhammad Awais , Sara Atito , Josef Kittler

Active learning improves annotation efficiency by selecting the most informative samples for annotation and model training. While most prior work has focused on selecting informative images for classification tasks, we investigate the more…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jingna Qiu , Frauke Wilm , Mathias Öttl , Jonas Utz , Maja Schlereth , Moritz Schillinger , Marc Aubreville , Katharina Breininger

Image paragraph generation is the task of producing a coherent story (usually a paragraph) that describes the visual content of an image. The problem nevertheless is not trivial especially when there are multiple descriptive and diverse…

Computer Vision and Pattern Recognition · Computer Science 2019-08-02 Jing Wang , Yingwei Pan , Ting Yao , Jinhui Tang , Tao Mei

Image compositing is a task of combining regions from different images to compose a new image. A common use case is background replacement of portrait images. To obtain high quality composites, professionals typically manually perform…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 He Zhang , Jianming Zhang , Federico Perazzi , Zhe Lin , Vishal M. Patel

Autoregressive generation is a powerful approach for high-fidelity image synthesis, but it remains computationally demanding and slow even on the most advanced accelerators. While speculative decoding has been explored to mitigate this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Selin Yildirim , Subhajit Dutta Chowdhury , Mohammad Mahdi Kamani , Vikram Appia , Deming Chen

In designing generative models, it is commonly believed that in order to learn useful latent structure, we face a fundamental tension between expressivity and structure. In this paper we challenge this view by proposing a new approach to…

Machine Learning · Statistics 2026-04-03 Alex Markham , Isaac Hirsch , Jeri A. Chang , Liam Solus , Bryon Aragam