English
Related papers

Related papers: WorldScribe: Towards Context-Aware Live Visual Des…

200 papers

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven cross-modal retrieval, yet face significant challenges due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yingjia Xu , Jinlin Wu , Daming Gao , Zhen Chen , Yang Yang , Min Cao , Mang Ye , Zhen Lei

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the features of their…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Osamu Take , Shinnosuke Takamichi , Kentaro Seki , Yoshiaki Bando , Hiroshi Saruwatari

Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually…

Human-Computer Interaction · Computer Science 2024-05-28 Ruei-Che Chang , Chia-Sheng Hung , Bing-Yu Chen , Dhruv Jain , Anhong Guo

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing…

Human-Computer Interaction · Computer Science 2025-05-01 Ruiping Liu , Jiaming Zhang , Angela Schön , Karin Müller , Junwei Zheng , Kailun Yang , Anhong Guo , Kathrin Gerling , Rainer Stiefelhagen

Cinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph…

Computer Vision and Pattern Recognition · Computer Science 2017-08-11 Tae-Hyun Oh , Kyungdon Joo , Neel Joshi , Baoyuan Wang , In So Kweon , Sing Bing Kang

Scene understanding is essential for enhancing driver safety, generating human-centric explanations for Automated Vehicle (AV) decisions, and leveraging Artificial Intelligence (AI) for retrospective driving video analysis. This study…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Mohammed Elhenawy , Huthaifa I. Ashqar , Andry Rakotonirainy , Taqwa I. Alhadidi , Ahmed Jaber , Mohammad Abu Tami

Generating realistic 3D indoor scenes from user inputs remains a challenging problem in computer vision and graphics, requiring careful balance of geometric consistency, spatial relationships, and visual realism. While neural generation…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Mengqi Zhou , Xipeng Wang , Yuxi Wang , Zhaoxiang Zhang

We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Payam Jome Yazdian , Rachel Lagasse , Hamid Mohammadi , Eric Liu , Li Cheng , Angelica Lim

World building forms the foundation of any task that requires narrative intelligence. In this work, we focus on procedurally generating interactive fiction worlds---text-based worlds that players "see" and "talk to" using natural language.…

Artificial Intelligence · Computer Science 2020-01-29 Prithviraj Ammanabrolu , Wesley Cheung , Dan Tu , William Broniec , Mark O. Riedl

Data-driven storytelling has gained prominence in journalism and other data reporting fields. However, the process of creating these stories remains challenging, often requiring the integration of effective visualizations with compelling…

Human-Computer Interaction · Computer Science 2025-04-01 Yu Fu , Dennis Bromley , Vidya Setlur

Keeping up with research and finding related work is still a time-consuming task for academics. Researchers sift through thousands of studies to identify a few relevant ones. Automation techniques can help by increasing the efficiency and…

Information Retrieval · Computer Science 2023-09-06 Wojciech Kusa , Petr Knoth , Allan Hanbury

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Sensemaking and narrative are two inherently interconnected concepts about how people understand the world around them. Sensemaking is the process by which people structure and interconnect the information they encounter in the world with…

Artificial Intelligence · Computer Science 2022-01-26 Zev Battad , Mei Si

Accurately reporting what objects are depicted in an image is largely a solved problem in automatic caption generation. The next big challenge on the way to truly humanlike captioning is being able to incorporate the context of the image…

Computation and Language · Computer Science 2022-10-11 Sofia Nikiforova , Tejaswini Deoskar , Denis Paperno , Yoad Winter

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

Human-Computer Interaction · Computer Science 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

In today's world, time is a very important resource. In our busy lives, most of us hardly have time to read the complete news so what we have to do is just go through the headlines and satisfy ourselves with that. As a result, we might miss…

Human-Computer Interaction · Computer Science 2020-01-06 Mona teja K , Mohan Sai. S , H S S S Raviteja D , Sai Kushagra P

Immersive, stereoscopic viewing enables scientists to better analyze the spatial structures of visualized physical phenomena. However, their findings cannot be properly presented in traditional media, which lack these core attributes.…

Graphics · Computer Science 2016-11-29 Jacqueline Chu , Leonardo Ferrer , Min Shih , Kwan-Liu Ma

Recent text-to-image generation methods provide a simple yet exciting conversion capability between text and image domains. While these methods have incrementally improved the generated image fidelity and text relevancy, several pivotal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Oran Gafni , Adam Polyak , Oron Ashual , Shelly Sheynin , Devi Parikh , Yaniv Taigman

Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Shahaf Pruss , Morris Alper , Hadar Averbuch-Elor