English
Related papers

Related papers: WorldScribe: Towards Context-Aware Live Visual Des…

200 papers

Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language…

Sound · Computer Science 2024-05-07 Antonio Bevilacqua , Paolo Saviano , Alessandro Amirante , Simon Pietro Romano

Live programming provides feedback on run-time behavior by visualizing concrete values of expressions close to the source code. When using such a local perspective on run-time behavior, programmers have to mentally reconstruct the control…

Programming Languages · Computer Science 2024-03-06 Patrick Rein , Christian Flach , Stefan Ramson , Eva Krebs , Robert Hirschfeld

This work presents CLIPDraw, an algorithm that synthesizes novel drawings based on natural language input. CLIPDraw does not require any training; rather a pre-trained CLIP language-image encoder is used as a metric for maximizing…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Kevin Frans , L. B. Soros , Olaf Witkowski

Cognitive control, the ability of a system to adapt to the demands of a task, is an integral part of cognition. A widely accepted fact about cognitive control is that it is context-sensitive: Adults and children alike infer information…

Artificial Intelligence · Computer Science 2020-12-02 Rachit Dubey , Erin Grant , Michael Luo , Karthik Narasimhan , Thomas Griffiths

Indoor scene synthesis has become increasingly important with the rise of Embodied AI, which requires 3D environments that are not only visually realistic but also physically plausible and functionally diverse. While recent approaches have…

Graphics · Computer Science 2025-10-28 Yandan Yang , Baoxiong Jia , Shujie Zhang , Siyuan Huang

The powerful reasoning of modern Vision Language Models open a new frontier for advanced personalization study. However, progress in this area is critically hampered by the lack of suitable benchmarks. To address this gap, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Xia Hu , Honglei Zhuang , Brian Potetz , Alireza Fathi , Bo Hu , Babak Samari , Howard Zhou

Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals' potentially diverse needs and preferences. This research…

Human-Computer Interaction · Computer Science 2024-08-22 Rosiana Natalie , Ruei-Che Chang , Smitha Sheshadri , Anhong Guo , Kotaro Hara

Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editing remains challenging, as it requires models to precisely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Yiheng Lin , Siyu Jiao , Xiaohan Lan , Wei Zhou , Qi She , Fei Yu , Heyun Chen , Zhengwei Wang , Jinghuan Chen , Moran Li , Yingchen Yu , Zijian Feng , Yao Zhao , Yunchao Wei , Yujie Zhong

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as they operate on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Aviad Aberdam , David Bensaïd , Alona Golts , Roy Ganz , Oren Nuriel , Royee Tichauer , Shai Mazor , Ron Litman

Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recommendations, has emerged as a promising research direction…

Artificial Intelligence · Computer Science 2026-04-23 Dongding Lin , Jian Wang , Yongqi Li , Wenjie Li

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting sounds that may not…

Sound · Computer Science 2023-11-10 David Chuan-En Lin , Nikolas Martelaro

The advances in AI-enabled techniques have accelerated the creation and automation of visualizations in the past decade. However, presenting visualizations in a descriptive and generative format remains a challenge. Moreover, current…

Human-Computer Interaction · Computer Science 2024-03-28 Qing Chen , Ying Chen , Ruishi Zou , Wei Shuai , Yi Guo , Jiazhe Wang , Nan Cao

Current approaches for open-vocabulary scene graph generation (OVSGG) use vision-language models such as CLIP and follow a standard zero-shot pipeline -- computing similarity between the query image and the text embeddings for each category…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Guikun Chen , Jin Li , Wenguan Wang

We introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Seok-Young Kim , Dooyoung Kim , Woojin Cho , Hail Song , Suji Kang , Woontack Woo

In the era of digital transformation, new technological foundations and possibilities for collaboration, production as well as organization open up many opportunities to work differently in the future. The digitization of workflows results…

Human-Computer Interaction · Computer Science 2022-12-01 Enes Yigitbas , Stefan Sauer , Gregor Engels

We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify…

Sound · Computer Science 2022-07-22 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

With the maturity of visual detection techniques, we are more ambitious in describing visual content with open-vocabulary, fine-grained and free-form language, i.e., the task of image captioning. In particular, we are interested in…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zheng-Jun Zha , Daqing Liu , Hanwang Zhang , Yongdong Zhang , Feng Wu

Reading text is one of the essential needs of the visually impaired people. We developed a mobile system that can read Turkish scene and book text, using a fast gradient-based multi-scale text detection algorithm for real-time operation and…

Multimedia · Computer Science 2016-08-18 Muhammet Bastan , Hilal Kandemir , Busra Canturk

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Xu Yan , Zhengcong Fei , Shuhui Wang , Qingming Huang , Qi Tian

Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time video-to-audio…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Benjamin Klein , Kazi Ruslan Rahman , Sanchita Ghose
‹ Prev 1 8 9 10 Next ›