中文
相关论文

相关论文: WorldScribe: Towards Context-Aware Live Visual Des…

200 篇论文

People who are blind or have low vision regularly use their hands to interact with the physical world to gain access to objects' shape, size, weight, and texture. However, many rich visual features remain inaccessible through touch alone,…

Real-world environments evolve continuously, yet blind and low-vision (BLV) individuals often have limited access to understanding how they change over time. Unexpected or relocated objects, layout modifications, and content updates (e.g.,…

People frequently use speech-to-text systems to compose short texts with voice. However, current voice-based interfaces struggle to support composing more detailed, contextually complex texts, especially in scenarios where users are on the…

人机交互 · 计算机科学 2025-08-07 Hamza El Alaoui , Atieh Taheri , Yi-Hao Peng , Jeffrey P. Bigham

People watch livestreams to connect with others and learn about their hobbies. Livestreams feature multiple visual streams including the main video, webcams, on-screen overlays, and chat, all of which are inaccessible to livestream viewers…

人机交互 · 计算机科学 2023-10-12 Daniel Killough , Amy Pavel

Image editing is an iterative process that requires precise visual evaluation and manipulation for the output to match the editing intent. However, current image editing tools do not provide accessible interaction nor sufficient feedback…

人机交互 · 计算机科学 2024-08-14 Ruei-Che Chang , Yuxuan Liu , Lotus Zhang , Anhong Guo

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limits the deployment of…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

Many blind and low vision (BLV) people are excluded from professional roles that may involve visual tasks due to access barriers and persisting stigmas. Advancing generative AI systems can support BLV people through providing contextual and…

人机交互 · 计算机科学 2025-10-13 Lucy Jiang , Lotus Zhang , Leah Findlater

Audio descriptions make videos accessible to those who cannot see them by describing visual content in audio. Producing audio descriptions is challenging due to the synchronous nature of the audio description that must fit into gaps of…

人机交互 · 计算机科学 2020-10-09 Amy Pavel , Gabriel Reyes , Jeffrey P. Bigham

Constructing photorealistic virtual worlds has applications across various fields, but it often requires the extensive labor of highly trained professionals to operate conventional 3D modeling software. To democratize this process, we…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Xinhang Liu , Chi-Keung Tang , Yu-Wing Tai

We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-controlled image-to-video…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Hanlin Wang , Hao Ouyang , Qiuyu Wang , Yue Yu , Yihao Meng , Wen Wang , Ka Leong Cheng , Shuailei Ma , Qingyan Bai , Yixuan Li , Cheng Chen , Yanhong Zeng , Xing Zhu , Yujun Shen , Qifeng Chen

Several services for people with visual disabilities have emerged recently due to achievements in Assistive Technologies and Artificial Intelligence areas. Despite the growth in assistive systems availability, there is a lack of services…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Daniel Louzada Fernandes , Marcos Henrique Fonseca Ribeiro , Fabio Ribeiro Cerqueira , Michel Melo Silva

Today's video-conferencing tools support a rich range of professional and social activities, but their generic meeting environments cannot be dynamically adapted to align with distributed collaborators' needs. To enable end-user…

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

"Scene description" applications that describe visual content in a photo are useful daily tools for blind and low vision (BLV) people. Researchers have studied their use, but they have only explored those that leverage remote sighted…

人机交互 · 计算机科学 2025-03-13 Ricardo Gonzalez , Jazmin Collins , Shiri Azenkot , Cynthia Bennett

Visual scene understanding is a fundamental task in computer vision that aims to extract meaningful information from visual data. It traditionally involves disjoint and specialized algorithms for different tasks that are tailored for…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Américo Pereira , Pedro Carvalho , Luís Côrte-Real

Visual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes. Existing autoregressive methods, which rely on previous frame-sentence pairs, struggle…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Sixiao Zheng , Yanwei Fu

In the digital landscape, the ubiquity of data visualizations in media underscores the necessity for accessibility to ensure inclusivity for all users, including those with visual impairments. Current visual content often fails to cater to…

人机交互 · 计算机科学 2024-09-27 Qiang Xu , Thomas Hurtut

Current scene perception tools for Blind and Low Vision (BLV) individuals rely on spoken descriptions but lack engaging representations of visually pleasing distant environmental landscapes (Vista spaces). Our proposed Scene2Audio framework…

人机交互 · 计算机科学 2026-03-31 Chitralekha Gupta , Jing Peng , Ashwin Ram , Shreyas Sridhar , Christophe Jouffrais , Suranga Nanayakkara

This work presents PerspectroScope, a web-based system which lets users query a discussion-worthy natural language claim, and extract and visualize various perspectives in support or against the claim, along with evidence supporting each…

计算与语言 · 计算机科学 2019-06-13 Sihao Chen , Daniel Khashabi , Chris Callison-Burch , Dan Roth
‹ 上一页 1 2 3 10 下一页 ›