中文
相关论文

相关论文: Mixing Modalities of 3D Sketching and Speech for I…

200 篇论文

Virtual assistants are becoming increasingly important speech-driven Information Retrieval platforms that assist users with various tasks. We discuss open problems and challenges with respect to modeling spoken information queries for…

信息检索 · 计算机科学 2023-04-27 Christophe Van Gysel

We introduce the novel problem of localizing all the instances of an object (seen or unseen during training) in a natural image via sketch query. We refer to this problem as sketch-guided object localization. This problem is distinctively…

计算机视觉与模式识别 · 计算机科学 2020-08-18 Aditay Tripathi , Rajath R Dani , Anand Mishra , Anirban Chakraborty

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

This paper addresses the problem of media retrieval using a multimodal query (a query which combines visual input with additional semantic information in natural language feedback). We propose a SynthTriplet GAN framework which resolves…

计算机视觉与模式识别 · 计算机科学 2021-07-01 Ivona Tautkute , Tomasz Trzcinski

The digital realm has witnessed the rise of various search modalities, among which the Image-Based Conversational Search System stands out. This research delves into the design, implementation, and evaluation of this specific system,…

信息检索 · 计算机科学 2024-04-01 Yue Zheng , Lei Yu , Junmian Chen , Tianyu Xia , Yuanyuan Yin , Shan Wang , Haiming Liu

We report experimental results associated with speech-driven text retrieval, which facilitates retrieving information in multiple domains with spoken queries. Since users speak contents related to a target collection, we produce language…

计算与语言 · 计算机科学 2016-11-15 Katunobu Itou , Atsushi Fujii , Tetsuya Ishikawa

The imitation of percussive sounds via the human voice is a natural and effective tool for communicating rhythmic ideas on the fly. Thus, the automatic retrieval of drum sounds using vocal percussion can help artists prototype drum patterns…

声音 · 计算机科学 2021-10-19 Alejandro Delgado , SkoT McDonald , Ning Xu , Charalampos Saitis , Mark Sandler

3D perception of object shapes from RGB image input is fundamental towards semantic scene understanding, grounding image-based perception in our spatially 3-dimensional real-world environments. To achieve a mapping between image views of…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Weicheng Kuo , Anelia Angelova , Tsung-Yi Lin , Angela Dai

This paper addresses cross-domain visual search, where visual queries retrieve category samples from a different domain. For example, we may want to sketch an airplane and retrieve photographs of airplanes. Despite considerable progress,…

计算机视觉与模式识别 · 计算机科学 2020-07-29 William Thong , Pascal Mettes , Cees G. M. Snoek

Generating 3D models from face sketches is an active topic of research in Computer Graphics due to its potential to tremendously facilitate the modeling of faces for both professional 3D arists and novices. Motivated by the observation that…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Nancy Iskander

Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Our goal is to synthesize humans interacting with a given 3D scene…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Kaifeng Zhao , Shaofei Wang , Yan Zhang , Thabo Beeler , Siyu Tang

We propose an interactive 3D character modeling approach from orthographic drawings (e.g., front and side views) based on 2D-space annotations. First, the system builds partial correspondences between the input drawings and generates a base…

计算机视觉与模式识别 · 计算机科学 2022-01-28 Zhengyu Huang , Haoran Xie , Tsukasa Fukusato

Conventional existing retrieval methods in remote sensing (RS) are often based on a uni-modal data retrieval framework. In this work, we propose a novel inter-modal triplet-based zero-shot retrieval scheme utilizing a sketch-based…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Ushasi Chaudhuri , Biplab Banerjee , Avik Bhattacharya , Mihai Datcu

When a toddler is presented a new toy, their instinctual behaviour is to pick it upand inspect it with their hand and eyes in tandem, clearly searching over its surface to properly understand what they are playing with. At any instance…

计算机视觉与模式识别 · 计算机科学 2020-11-04 Edward J. Smith , Roberto Calandra , Adriana Romero , Georgia Gkioxari , David Meger , Jitendra Malik , Michal Drozdzal

We present VOICE, a novel approach to science communication that connects large language models' (LLM) conversational capabilities with interactive exploratory visualization. VOICE introduces several innovative technical contributions that…

We present an architecture for integrating real-time, multimodal input into a computational agent's contextual model. Using a human-avatar interaction in a virtual world, we treat aligned gesture and speech as an ensemble where content may…

人机交互 · 计算机科学 2019-09-19 Nikhil Krishnaswamy , James Pustejovsky

This research features a deep-learning based framework to address the problem of matching a given face sketch image against a face photo database. The problem of photo-sketch matching is challenging because 1) there is large modality gap…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Seho Bae , Nizam Ud Din , Hyunkyu Park , Juneho Yi

In this paper, we extend scene understanding to include that of human sketch. The result is a complete trilogy of scene representation from three diverse and complementary modalities -- sketch, photo, and text. Instead of learning a rigid…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Pinaki Nath Chowdhury , Ayan Kumar Bhunia , Aneeshan Sain , Subhadeep Koley , Tao Xiang , Yi-Zhe Song

Spontaneous speech in the form of conversations, meetings, voice-mail, interviews, oral history, etc. is one of the most ubiquitous forms of human communication. Search engines providing access to such speech collections have the potential…

人机交互 · 计算机科学 2013-12-19 Donna Vakharia , Rachel Gibbs

Interaction plays a vital role during visual network exploration as users need to engage with both elements in the view (e.g., nodes, links) and interface controls (e.g., sliders, dropdown menus). Particularly as the size and complexity of…

人机交互 · 计算机科学 2020-05-01 Ayshwarya Saktheeswaran , Arjun Srinivasan , John Stasko