中文
相关论文

相关论文: AIpparel: A Multimodal Foundation Model for Digita…

200 篇论文

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

We present FashionEngine, an interactive 3D human generation and editing system that creates 3D digital humans via user-friendly multimodal controls such as natural languages, visual perceptions, and hand-drawing sketches. FashionEngine…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Tao Hu , Fangzhou Hong , Zhaoxi Chen , Ziwei Liu

We propose a new type of full-body human avatars, which combines parametric mesh-based body model with a neural texture. We show that with the help of neural textures, such avatars can successfully model clothing and hair, which usually…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Artur Grigorev , Karim Iskakov , Anastasia Ianina , Renat Bashirov , Ilya Zakharkin , Alexander Vakhitov , Victor Lempitsky

A virtual try-on method takes a product image and an image of a model and produces an image of the model wearing the product. Most methods essentially compute warps from the product image to the model image and combine using image…

计算机视觉与模式识别 · 计算机科学 2020-03-30 Kedan Li , Min Jin Chong , Jingen Liu , David Forsyth

Multi-Instance Multi-Label learning (MIML) models complex objects (bags), each of which is associated with a set of interrelated labels and composed with a set of instances. Current MIML solutions still focus on a single-type of objects and…

机器学习 · 计算机科学 2021-11-09 Yuanlin Yang , Guoxian Yu , Jun Wang , Lei Liu , Carlotta Domeniconi , Maozu Guo

With the growth of online shopping for fashion products, accurate fashion recommendation has become a critical problem. Meanwhile, social networks provide an open and new data source for personalized fashion analysis. In this work, we study…

计算机视觉与模式识别 · 计算机科学 2020-05-27 Haitian Zheng , Kefei Wu , Jong-Hwi Park , Wei Zhu , Jiebo Luo

Information workers' productivity is significantly influenced by their cognitive states and physiological responses. AI assistants such as ChatGPT, Copilot, and others have become integral components of knowledge-intensive workplaces. These…

人机交互 · 计算机科学 2026-05-12 Amog Rao , Utkarsh Agarwal , Amol Harsh , Siddharth Siddharth

We present BootComp, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments. Here, the main bottleneck is data acquisition for training: collecting a large-scale…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Yisol Choi , Sangkyung Kwak , Sihyun Yu , Hyungwon Choi , Jinwoo Shin

How to recommend outfits has gained considerable attention in both academia and industry in recent years. Many studies have been carried out regarding fashion compatibility learning, to determine whether the fashion items in an outfit are…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Dongliang Zhou , Haijun Zhang , Qun Li , Jianghong Ma , Xiaofei Xu

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-garment model that…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Jan Ackermann , Kiyohiro Nakayama , Guandao Yang , Tong Wu , Gordon Wetzstein

Multimodal retrieval systems are becoming increasingly vital for cutting-edge AI technologies, such as embodied AI and AI-driven digital content industries. However, current multimodal retrieval tasks lack sufficient complexity and…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Bangwei Liu , Yicheng Bao , Shaohui Lin , Xuhong Wang , Xin Tan , Yingchun Wang , Yuan Xie , Chaochao Lu

Analyzing fashion trends is essential in the fashion industry. Current fashion forecasting firms, such as WGSN, utilize the visual information from around the world to analyze and predict fashion trends. However, analyzing fashion trends is…

计算机视觉与模式识别 · 计算机科学 2020-05-05 Mengyun Shi , Van Dyk Lewis

Manipulating clothes is challenging due to their complex dynamics, high deformability, and frequent self-occlusions. Garments exhibit a nearly infinite number of configurations, making explicit state representations difficult to define. In…

机器人学 · 计算机科学 2025-05-13 Oriol Barbany , Adrià Colomé , Carme Torras

High-quality 3D garment reconstruction plays a crucial role in mitigating the sim-to-real gap in applications such as digital avatars, virtual try-on and robotic manipulation. However, existing garment reconstruction methods typically rely…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Ming Li , Hui Shan , Kai Zheng , Chentao Shen , Siyu Liu , Yanwei Fu , Zhen Chen , Xiangru Huang

Multimodal models are expected to be a critical component to future advances in artificial intelligence. This field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural…

Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored,…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Yu He , Ting Zhu , Yichun Liu , Lichen Ma , Xinyuan Shan , Jingling Fu , Yu Shi , Junshi Huang , Yan Li

We introduce a multi-modal discriminative and generative frame-work capable of assisting humans in producing visual content re-lated to a given theme, starting from a collection of documents(textual, visual, or both). This framework can be…

图像与视频处理 · 电气工程与系统科学 2020-02-07 Michele Merler , Cicero Nogueira dos Santos , Mauro Martino , Alfio M. Gliozzo , John R. Smith

Older adults have increasing difficulty with retrospective memory, hindering their abilities to perform daily activities and posing stress on caregivers to ensure their wellbeing. Recent developments in Artificial Intelligence (AI) and…

人机交互 · 计算机科学 2025-02-05 Natasha Maniar , Samantha W. T. Chan , Wazeer Zulfikar , Scott Ren , Christine Xu , Pattie Maes

Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. Large representation models (e.g., VLM2Vec) demonstrate strong multimodal understanding capabilities, yet they struggle with fine-grained…

计算与语言 · 计算机科学 2026-04-23 Biao Zhang , Lixin Chen , Bin Zhang , Zongwei Wang , Tong Liu , Bo Zheng

Fashionable image generation aims to synthesize images of diverse fashion prevalent around the globe, helping fashion designers in real-time visualization by giving them a basic customized structure of how a specific design preference would…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Krishna Sri Ipsit Mantri , Nevasini Sasikumar