English
Related papers

Related papers: OnomaCompass: A Texture Exploration Interface that…

200 papers

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Gaoyang Zhang , Bingtao Fu , Qingnan Fan , Qi Zhang , Runxing Liu , Hong Gu , Huaqi Zhang , Xinguo Liu

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Xudong Wang , Shufan Li , Konstantinos Kallidromitis , Yusuke Kato , Kazuki Kozuka , Trevor Darrell

Software engineering (SE) and requirements engineering (RE) face a significant increase in secondary studies, particularly literature reviews (LRs), due to the ever-growing number of scientific publications. Generative artificial…

Software Engineering · Computer Science 2026-02-27 Oliver Karras , Amirreza Alasti , Lena John , Sushant Aggarwal , Yücel Celik

Ever growing number of image documents available on the Internet continuously motivates research in better annotation models and more efficient retrieval methods. Formal knowledge representation of objects and events in pictures, their…

Information Retrieval · Computer Science 2017-12-06 Marko Horvat , Anton Grbin , Gordan Gledec

Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhigang Wang , Yifei Su , Chenhui Li , Dong Wang , Yan Huang , Bin Zhao , Xuelong Li

Significant progress has been made in open-vocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment,…

User interfaces provide an interactive window between physical and virtual environments. A new concept in the field of human-computer interaction is a soft user interface; a compliant surface that facilitates touch interaction through…

Human-Computer Interaction · Computer Science 2018-03-28 Chris Larson , Josef Spjut , Ross Knepper , Robert Shepherd

The digital realm has witnessed the rise of various search modalities, among which the Image-Based Conversational Search System stands out. This research delves into the design, implementation, and evaluation of this specific system,…

Information Retrieval · Computer Science 2024-04-01 Yue Zheng , Lei Yu , Junmian Chen , Tianyu Xia , Yuanyuan Yin , Shan Wang , Haiming Liu

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that…

Text-to-image models are enabling efficient design space exploration, rapidly generating images from text prompts. However, many generative AI tools are imperfect for product design applications as they are not built for the goals and…

Human-Computer Interaction · Computer Science 2025-01-22 Leah Chong , I-Ping Lo , Jude Rayan , Steven Dow , Faez Ahmed , Ioanna Lykourentzou

When people search for information about a new topic within large document collections, they implicitly construct a mental model of the unfamiliar information space to represent what they currently know and guide their exploration into the…

Information Retrieval · Computer Science 2023-02-21 Mengtian Guo , Zhilan Zhou , David Gotz , Yue Wang

This paper presents AnthropoCam, a mobile-based neural style transfer (NST) system optimized for the visual synthesis of Anthropocene environments. Unlike conventional artistic NST, which prioritizes painterly abstraction, stylizing…

Human-Computer Interaction · Computer Science 2026-01-30 Po-Hsun Chen , Ivan C. H. Liu

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highlights the challenges…

Human-Computer Interaction · Computer Science 2025-02-12 Haichuan Lin , Yilin Ye , Jiazhi Xia , Wei Zeng

Robotic systems demand accurate and comprehensive 3D environment perception, requiring simultaneous capture of photo-realistic appearance (optical), precise layout shape (geometric), and open-vocabulary scene understanding (semantic).…

Robotics · Computer Science 2025-09-10 Yinan Deng , Yufeng Yue , Jianyu Dou , Jingyu Zhao , Jiahui Wang , Yujie Tang , Yi Yang , Mengyin Fu

Object Navigation (ObjectNav) has made great progress with large language models (LLMs), but still faces challenges in memory management, especially in long-horizon tasks and dynamic scenes. To address this, we propose TopoNav, a new…

Robotics · Computer Science 2025-09-03 Peiran Liu , Qiang Zhang , Daojie Peng , Lingfeng Zhang , Yihao Qin , Hang Zhou , Jun Ma , Renjing Xu , Yiding Ji

Text-to-image diffusion models often struggle to achieve accurate semantic alignment between generated images and text prompts while maintaining efficiency for deployment on resource-constrained hardware. Existing approaches either incur…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Ziji Lu

Searching through vast libraries of sound samples can be a daunting and time-consuming task. Modern audio sample browsers use mappings between acoustic properties and visual attributes to visually differentiate displayed items. There are…

Human-Computer Interaction · Computer Science 2020-12-01 Etienne Richan , Jean Rouat

Given a pool of observations selected from a sensor stream, input data can be robustly represented, via a multiscale process, in terms of invariant concepts, and themes. Applying this to episodic natural language data, one may obtain a…

Artificial Intelligence · Computer Science 2020-10-19 Mark Burgess

Ontology engineering (OE) in large projects poses a number of challenges arising from the heterogeneous backgrounds of the various stakeholders, domain experts, and their complex interactions with ontology designers. This multi-party…

We present a novel approach to object classification and detection which requires minimal supervision and which combines visual texture cues and shape information learned from freely available unlabeled web search results. The explosion of…

Computer Vision and Pattern Recognition · Computer Science 2016-09-15 Xingchao Peng , Kate Saenko