English
Related papers

Related papers: WISE: A Multimodal Search Engine for Visual Scenes…

200 papers

Multimodal retrieval is the task of aggregating information from queries across heterogeneous modalities to retrieve desired targets. State-of-the-art multimodal retrieval models can understand complex queries, yet they are typically…

Information Retrieval · Computer Science 2026-03-25 Chuong Huynh , Manh Luong , Abhinav Shrivastava

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Arjun R. Akula , Song-Chun Zhu

While search technologies have evolved to be robust and ubiquitous, the fundamental interaction paradigm has remained relatively stable for decades. With the maturity of the Brain-Machine Interface, we build an efficient and effective…

Information Retrieval · Computer Science 2021-10-18 Xuesong Chen , Ziyi Ye , Xiaohui Xie , Yiqun Liu , Weihang Su , Shuqi Zhu , Min Zhang , Shaoping Ma

Spontaneous speech in the form of conversations, meetings, voice-mail, interviews, oral history, etc. is one of the most ubiquitous forms of human communication. Search engines providing access to such speech collections have the potential…

Human-Computer Interaction · Computer Science 2013-12-19 Donna Vakharia , Rachel Gibbs

There has been a rapid growth of digitally available music data, including audio recordings, digitized images of sheet music, album covers and liner notes, and video clips. This huge amount of data calls for retrieval strategies that allow…

Information Retrieval · Computer Science 2019-02-13 Meinard Müller , Andreas Arzt , Stefan Balke , Matthias Dorfer , Gerhard Widmer

We present WISER, a new semantic search engine for expert finding in academia. Our system is unsupervised and it jointly combines classical language modeling techniques, based on text evidences, with the Wikipedia Knowledge Graph, via…

Information Retrieval · Computer Science 2019-06-11 Paolo Cifariello , Paolo Ferragina , Marco Ponza

Multimodal desire understanding, a task closely related to both emotion and sentiment that aims to infer human intentions from visual and textual cues, is an emerging yet underexplored task in affective computing with applications in social…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Wei Chen , Tongguan Wang , Feiyue Xue , Junkai Li , Hui Liu , Ying Sha

Conversational information seeking (CIS) has been recognized as a major emerging research area in information retrieval. Such research will require data and tools, to allow the implementation and study of conversational systems. This paper…

Information Retrieval · Computer Science 2019-12-20 Hamed Zamani , Nick Craswell

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Ziyi Ni , Minglun Han , Feilong Chen , Linghui Meng , Jing Shi , Pin Lv , Bo Xu

In this paper, we introduce iART: an open Web platform for art-historical research that facilitates the process of comparative vision. The system integrates various machine learning techniques for keyword- and content-based image retrieval…

Information Retrieval · Computer Science 2022-01-10 Matthias Springstein , Stefanie Schneider , Javad Rahnama , Eyke Hüllermeier , Hubertus Kohle , Ralph Ewerth

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only…

Computation and Language · Computer Science 2020-02-19 Tejas Srinivasan , Ramon Sanabria , Florian Metze

We present MetaFind, a scene-aware tri-modal compositional retrieval framework designed to enhance scene generation in the metaverse by retrieving 3D assets from large-scale repositories. MetaFind addresses two core challenges: (i)…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Zhenyu Pan , Yucheng Lu , Han Liu

We propose a novel, efficient, modular and scalable framework for content based visual media retrieval systems by leveraging the power of Deep Learning which is flexible to work both for images and videos conjointly and we also introduce an…

Machine Learning · Computer Science 2021-05-19 Ambareesh Ravi , Amith Nandakumar

Identifying user intents from natural language utterances is a crucial step in conversational systems that has been extensively studied as a supervised classification problem. However, in practice, new intents emerge after deploying an…

Computation and Language · Computer Science 2021-02-08 A. B. Siddique , Fuad Jamour , Luxun Xu , Vagelis Hristidis

Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Oriol Barbany , Michael Huang , Xinliang Zhu , Arnab Dhua

Existing vision-language methods typically support two languages at a time at most. In this paper, we present a modular approach which can easily be incorporated into existing vision-language methods in order to support many languages. We…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Donghyun Kim , Kuniaki Saito , Kate Saenko , Stan Sclaroff , Bryan A. Plummer

The ever evolving informatics technology has gradually bounded human and computer in a compact way. Understanding user behavior becomes a key enabler in many fields such as sedentary-related healthcare, human-computer interaction (HCI) and…

Human-Computer Interaction · Computer Science 2020-05-25 Yu Gu , Xiang Zhang , Zhi Liu , Fuji Ren

Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360{\deg}. However, prior approaches to visual search are limited to a static image, neglecting the physical…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Heyang Yu , Yinan Han , Xiangyu Zhang , Baiqiao Yin , Bowen Chang , Xiangyu Han , Xinhao Liu , Jing Zhang , Marco Pavone , Chen Feng , Saining Xie , Yiming Li

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Hongwei Yi , Chun-Hao P. Huang , Shashank Tripathi , Lea Hering , Justus Thies , Michael J. Black

The recent surge in artificial intelligence, particularly in multimodal processing technology, has advanced human-computer interaction, by altering how intelligent systems perceive, understand, and respond to contextual information (i.e.,…

‹ Prev 1 4 5 6 7 8 10 Next ›