English
Related papers

Related papers: AIpparel: A Multimodal Foundation Model for Digita…

200 papers

General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to-image models suffer from fine-grained semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Shiyue Zhang , Zheng Chong , Xujie Zhang , Hanhui Li , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

We present a data-driven method for learning to generate animations of 3D garments using a 2D image diffusion model. In contrast to existing methods, typically based on fully connected networks, graph neural networks, or generative…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Raquel Vidaurre , Elena Garces , Dan Casas

The process of fashion design usually involves sketching, refining, and coloring, with designers drawing inspiration from various images to fuel their creative endeavors. However, conventional image search methods often yield irrelevant…

Human-Computer Interaction · Computer Science 2024-10-01 Jianan Jiang , Di Wu , Hanhui Deng , Yidan Long , Wenyi Tang , Xiang Li , Can Liu , Zhanpeng Jin , Wenlei Zhang , Tangquan Qi

We present Any-Modality Augmented Language Model (AnyMAL), a unified model that reasons over diverse input modality signals (i.e. text, image, video, audio, IMU motion sensor), and generates textual responses. AnyMAL inherits the powerful…

We have recently seen great progress in building photorealistic animatable full-body codec avatars, but generating high-fidelity animation of clothing is still difficult. To address these difficulties, we propose a method to build an…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Donglai Xiang , Fabian Prada , Timur Bagautdinov , Weipeng Xu , Yuan Dong , He Wen , Jessica Hodgins , Chenglei Wu

Time-series foundation models excel at tasks like forecasting across diverse data types by leveraging informative waveform representations. Wearable sensing data, however, pose unique challenges due to their variability in patterns and…

Machine Learning · Computer Science 2025-05-19 Yunfei Luo , Yuliang Chen , Asif Salekin , Tauhidur Rahman

Eye-tracking data reveals valuable insights into users' cognitive states but is difficult to analyze due to its structured, non-linguistic nature. While large language models (LLMs) excel at reasoning over text, they struggle with temporal…

Human-Computer Interaction · Computer Science 2025-07-25 Dongyang Guo , Yasmeen Abdrabou , Enkeleda Thaqi , Enkelejda Kasneci

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

Machine Learning · Computer Science 2024-05-01 Paul Pu Liang

Clothing plays a fundamental role in digital humans. Current approaches to animate 3D garments are mostly based on realistic physics simulation, however, they typically suffer from two main issues: high computational run-time cost, which…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Andrés Casado-Elvira , Marc Comino Trinidad , Dan Casas

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

Robotics · Computer Science 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

The realism of digital avatars is crucial in enabling telepresence applications with self-expression and customization. While physical simulations can produce realistic motions for clothed humans, they require high-quality garment assets…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yifei Li , Hsiao-yu Chen , Egor Larionov , Nikolaos Sarafianos , Wojciech Matusik , Tuur Stuyck

Large multimodal models (LMMs) have revolutionized text-to-image generation, but they risk perpetuating the harmful social biases in their training data. Prior work has identified gender bias in these models, but methodological limitations…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Juan Manuel Contreras

Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal coherence, their dynamics are often constrained by the continuous nature of their training data.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zhoujie Fu , Xianfang Zeng , Jinghong Lan , Xinyao Liao , Cheng Chen , Junyi Chen , Jiacheng Wei , Wei Cheng , Shiyu Liu , Yunuo Chen , Gang Yu , Guosheng Lin

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

The standardized sizes used in the garment industry do not cover the range of individual differences in body shape for most people, leading to ill-fitting clothes, high return rates and overproduction. Recent research efforts in both…

Graphics · Computer Science 2022-09-21 Katja Wolff , Philipp Herholz , Verena Ziegler , Frauke Link , Nico Brügel , Olga Sorkine-Hornung

In-context learning (ICL) facilitates Large Language Models (LLMs) exhibiting emergent ability on downstream tasks without updating billions of parameters. However, in the area of multi-modal Large Language Models (MLLMs), two problems…

Multimedia · Computer Science 2024-07-02 Jun Gao , Qian Qiao , Ziqiang Cao , Zili Wang , Wenjie Li

Garments are important to humans. A visual system that can estimate and track the complete garment pose can be useful for many downstream tasks and real-world applications. In this work, we present a complete package to address the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Han Xue , Wenqiang Xu , Jieyi Zhang , Tutian Tang , Yutong Li , Wenxin Du , Ruolin Ye , Cewu Lu

Per-garment virtual try-on methods collect garment-specific datasets and train networks tailored to each garment to achieve superior results. However, these approaches often struggle with loose-fitting garments due to two key limitations:…

Graphics · Computer Science 2025-09-05 Zaiqiang Wu , I-Chao Shen , Takeo Igarashi

Multimodal Entity Linking (MEL) is a crucial task that aims at linking ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, such as Wikipedia. Existing methods focus heavily on using complex…

Artificial Intelligence · Computer Science 2024-08-22 Liu Qi , He Yongyi , Lian Defu , Zheng Zhi , Xu Tong , Liu Che , Chen Enhong

The capability to generate simulation-ready garment models from 3D shapes of clothed humans will significantly enhance the interpretability of captured geometry of real garments, as well as their faithful reproduction in the virtual world.…

Graphics · Computer Science 2024-10-22 Boyang Yu , Frederic Cordier , Hyewon Seo