中文
相关论文

相关论文: Conversational Fashion Image Retrieval via Multitu…

200 篇论文

Convolutional Neural Networks (CNNs) have achieved superior performance on object image retrieval, while Bag-of-Words (BoW) models with handcrafted local features still dominate the retrieval of overlapping images in 3D reconstruction. In…

计算机视觉与模式识别 · 计算机科学 2018-12-11 Tianwei Shen , Zixin Luo , Lei Zhou , Runze Zhang , Siyu Zhu , Tian Fang , Long Quan

In this demo, we present Chat-to-Design, a new multimodal interaction system for personalized fashion design. Compared to classic systems that recommend apparel based on keywords, Chat-to-Design enables users to design clothes in two steps:…

人工智能 · 计算机科学 2022-07-05 Weiming Zhuang , Chongjie Ye , Ying Xu , Pengzhi Mao , Shuai Zhang

We study learning of a matching model for response selection in retrieval-based dialogue systems. The problem is equally important with designing the architecture of a model, but is less explored in existing literature. To learn a robust…

计算与语言 · 计算机科学 2019-06-12 Jiazhan Feng , Chongyang Tao , Wei Wu , Yansong Feng , Dongyan Zhao , Rui Yan

The role of social media in fashion industry has been blooming as the years have continued on. In this work, we investigate sentiment analysis for fashion related posts in social media platforms. There are two main challenges of this task.…

计算与语言 · 计算机科学 2021-11-16 Yifei Yuan , Wai Lam

To study the correlation between clothing garments and body shape, we collected a new dataset (Fashion Takes Shape), which includes images of users with clothing category annotations. We employ our multi-photo approach to estimate body…

计算机视觉与模式识别 · 计算机科学 2018-12-18 Hosnieh Sattar , Gerard Pons-Moll , Mario Fritz

We introduce a multimodal dataset where users express preferences through images. These images encompass a broad spectrum of visual expressions ranging from landscapes to artistic depictions. Users request recommendations for books or music…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Se-eun Yoon , Hyunsik Jeon , Julian McAuley

Image-based virtual try-on aims to synthesize a naturally dressed person image with a clothing image, which revolutionizes online shopping and inspires related topics within image generation, showing both research significance and…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Dan Song , Xuanpu Zhang , Juan Zhou , Weizhi Nie , Ruofeng Tong , Mohan Kankanhalli , An-An Liu

The ability to efficiently search for images is essential for improving the user experiences across various products. Incorporating user feedback, via multi-modal inputs, to navigate visual search can help tailor retrieved results to…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Surgan Jandial , Pinkesh Badjatiya , Pranit Chawla , Ayush Chopra , Mausoom Sarkar , Balaji Krishnamurthy

Image captioning models aim at connecting Vision and Language by providing natural language descriptions of input images. In the past few years, the task has been tackled by learning parametric models and proposing visual feature extraction…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

In this paper, we study the cross-modal image retrieval, where the inputs contain a source image plus some text that describes certain modifications to this image and the desired image. Prior work usually uses a three-stage strategy to…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Chunbin Gu , Jiajun Bu , Xixi Zhou , Chengwei Yao , Dongfang Ma , Zhi Yu , Xifeng Yan

Convolutional Neural Networks have been highly successful in performing a host of computer vision tasks such as object recognition, object detection, image segmentation and texture synthesis. In 2015, Gatys et. al [7] show how the style of…

计算机视觉与模式识别 · 计算机科学 2017-08-01 Prutha Date , Ashwinkumar Ganesan , Tim Oates

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Sihui Ji , Yiyang Wang , Xi Chen , Xiaogang Xu , Hao Luo , Hengshuang Zhao

Image-based virtual try-on techniques have shown great promise for enhancing the user-experience and improving customer satisfaction on fashion-oriented e-commerce platforms. However, existing techniques are currently still limited in the…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Benjamin Fele , Ajda Lampe , Peter Peer , Vitomir Štruc

This paper introduces a new challenge for image similarity search in the context of fashion, addressing the inherent ambiguity in this domain stemming from complex images. We present Referred Visual Search (RVS), a task allowing users to…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Simon Lepage , Jérémie Mary , David Picard

Recent models for cross-modal retrieval have benefited from an increasingly rich understanding of visual scenes, afforded by scene graphs and object interactions to mention a few. This has resulted in an improved matching between the visual…

计算机视觉与模式识别 · 计算机科学 2020-12-09 Andrés Mafla , Rafael Sampaio de Rezende , Lluís Gómez , Diane Larlus , Dimosthenis Karatzas

In this paper, we investigate the problem of retrieving images from a database based on a multi-modal (image-text) query. Specifically, the query text prompts some modification in the query image and the task is to retrieve images with the…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Muhammad Umer Anwaar , Egor Labintcev , Martin Kleinsteuber

The subject of conversational mining has become of great interest recently due to the explosion of social and other online media. Supplementing this explosion of text is the advancement in pre-trained language models which have helped us to…

计算与语言 · 计算机科学 2022-11-15 Nicolle Garber , Vukosi Marivate

Semantic retrieval (also known as dense retrieval) based on textual data has been extensively studied for both web search and product search application fields, where the relevance of a query and a potential target document is computed by…

信息检索 · 计算机科学 2025-02-18 Dong Liu , Esther Lopez Ramos

Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth…

计算机视觉与模式识别 · 计算机科学 2022-06-08 Jie Lei , Tamara L. Berg , Mohit Bansal

Evaluation of conversational naturalness is essential for developing human-like speech agents. However, existing speech naturalness predictors are often designed to assess utterances from a single speaker, failing to capture…

音频与语音处理 · 电气工程与系统科学 2026-03-03 Anfeng Xu , Yashesh Gaur , Naoyuki Kanda , Zhicheng Ouyang , Katerina Zmolikova , Desh Raj , Simone Merello , Anna Sun , Ozlem Kalinli
‹ 上一页 1 8 9 10 下一页 ›