中文
相关论文

相关论文: TalkFashion: Intelligent Virtual Try-On Assistant …

200 篇论文

Although Multimodal Large Language Models (MLLMs) have demonstrated promising versatile capabilities, their performance is still inferior to specialized models on downstream tasks, which makes adaptation necessary to enhance their utility.…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Yichi Zhang , Yinpeng Dong , Siyuan Zhang , Tianzan Min , Hang Su , Jun Zhu

Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in human-machine…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Tan-Hanh Pham , Hoang-Nam Le , Phu-Vinh Nguyen , Chris Ngo , Truong-Son Hy

Knowledge editing for large language models can offer an efficient solution to alter a model's behavior without negatively impacting the overall performance. However, the current approaches encounter issues with limited generalizability…

计算与语言 · 计算机科学 2024-04-30 Ningyu Zhang , Bozhong Tian , Siyuan Cheng , Xiaozhuan Liang , Yi Hu , Kouying Xue , Yanjie Gou , Xi Chen , Huajun Chen

Recent advancements in Virtual Try-On (VTO) have demonstrated exceptional efficacy in generating realistic images and preserving garment details, largely attributed to the robust generative capabilities of text-to-image (T2I) diffusion…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Zhenchen Wan , Yanwu Xu , Zhaoqing Wang , Feng Liu , Tongliang Liu , Mingming Gong

Image-based virtual try-on aims to fit an in-shop garment onto a clothed person image. Garment warping, which aligns the target garment with the corresponding body parts in the person image, is a crucial step in achieving this goal.…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sanhita Pathak , Vinay Kaushik , Brejesh Lall

This paper considers image-based virtual try-on, which renders an image of a person wearing a curated garment, given a pair of images depicting the person and the garment, respectively. Previous works adapt existing exemplar-based…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yisol Choi , Sangkyung Kwak , Kyungmin Lee , Hyungwon Choi , Jinwoo Shin

Tour guidance in virtual museums encourages multi-modal interactions to boost user experiences, concerning engagement, immersion, and spatial awareness. Nevertheless, achieving the goal is challenging due to the complexity of comprehending…

人机交互 · 计算机科学 2024-01-24 Zhan Wang , Lin-Ping Yuan , Liangwei Wang , Bingchuan Jiang , Wei Zeng

Fashion intelligence spans multiple tasks, i.e., retrieval, recommendation, recognition, and dialogue, yet remains hindered by fragmented supervision and incomplete fashion annotations. These limitations jointly restrict the formation of…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zhengwei Yang , Andi Long , Hao Li , Zechao Hu , Kui Jiang , Zheng Wang

In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest. These models, which combine various forms of data input, are becoming increasingly popular.…

We introduce DialogPaint, a novel framework that bridges conversational interactions with image editing, enabling users to modify images through natural dialogue. By integrating a dialogue model with the Stable Diffusion image…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Jingxuan Wei , Shiyu Wu , Xin Jiang , Yequan Wang

External tools help large language models succeed at tasks where they would otherwise typically fail. In existing frameworks, choosing tools at test time relies on naive greedy decoding, regardless of whether the model has been fine-tuned…

计算与语言 · 计算机科学 2025-09-23 Lisa Alazraki , Marek Rei

With the rapid development of e-commerce, virtual try-on technology has become an essential tool to satisfy consumers' personalized clothing preferences. Diffusion-based virtual try-on systems aim to naturally align garments with target…

多媒体 · 计算机科学 2025-04-02 Shufang Zhang , Hang Qian , Minxue Ni , Yaxuan Li , Wenxin Ding , Jun Liu

Virtual try-on system under arbitrary human poses has huge application potential, yet raises quite a lot of challenges, e.g. self-occlusions, heavy misalignment among diverse poses, and diverse clothes textures. Existing methods aim at…

计算机视觉与模式识别 · 计算机科学 2019-03-01 Haoye Dong , Xiaodan Liang , Bochao Wang , Hanjiang Lai , Jia Zhu , Jian Yin

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce…

计算与语言 · 计算机科学 2022-03-29 Yingshan Chang , Mridu Narang , Hisami Suzuki , Guihong Cao , Jianfeng Gao , Yonatan Bisk

In the industry, numerous tasks are deployed online. Traditional approaches often tackle each task separately by its own network, which leads to excessive costs for developing and scaling models, especially in the context of large language…

计算与语言 · 计算机科学 2024-11-08 Yincen Qu , Chao Ma , Xiangying Dai , Hui Zhou , Yiting Wu , Hengyue Liu

Voice assistants provide users a new way of interacting with digital products, allowing them to retrieve information and complete tasks with an increased sense of control and flexibility. Such products are comprised of several machine…

音频与语音处理 · 电气工程与系统科学 2021-08-27 Shachaf Poran , Gil Amsalem , Amit Beka , Dmitri Goldenberg

Voice dictation is an increasingly important text input modality. Existing systems that allow both dictation and editing-by-voice restrict their command language to flat templates invoked by trigger words. In this work, we study the…

计算与语言 · 计算机科学 2023-07-11 Belinda Z. Li , Jason Eisner , Adam Pauls , Sam Thomson

We showcase an application that leverages multiple agents, powered by large language models and integrated tools, to collaboratively solve complex network operation tasks across various domains. The tasks include real-time topology…

In the age of artificial intelligence (AI), providing learners with suitable and sufficient explanations of AI-based recommendation algorithm's output becomes essential to enable them to make an informed decision about it. However, the…

人机交互 · 计算机科学 2024-02-14 Hasan Abu-Rasheed , Christian Weber , Madjid Fathi

We introduce a novel framework for evaluating multimodal deep learning models with respect to their language understanding and generalization abilities. In this approach, artificial data is automatically generated according to the…

计算与语言 · 计算机科学 2017-04-18 Alexander Kuhnle , Ann Copestake