中文
相关论文

相关论文: WorldScribe: Towards Context-Aware Live Visual Des…

200 篇论文

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven cross-modal retrieval, yet face significant challenges due to…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yingjia Xu , Jinlin Wu , Daming Gao , Zhen Chen , Yang Yang , Min Cao , Mang Ye , Zhen Lei

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the features of their…

音频与语音处理 · 电气工程与系统科学 2024-08-14 Osamu Take , Shinnosuke Takamichi , Kentaro Seki , Yoshiaki Bando , Hiroshi Saruwatari

Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually…

人机交互 · 计算机科学 2024-05-28 Ruei-Che Chang , Chia-Sheng Hung , Bing-Yu Chen , Dhruv Jain , Anhong Guo

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing…

Cinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph…

计算机视觉与模式识别 · 计算机科学 2017-08-11 Tae-Hyun Oh , Kyungdon Joo , Neel Joshi , Baoyuan Wang , In So Kweon , Sing Bing Kang

Scene understanding is essential for enhancing driver safety, generating human-centric explanations for Automated Vehicle (AV) decisions, and leveraging Artificial Intelligence (AI) for retrospective driving video analysis. This study…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Mohammed Elhenawy , Huthaifa I. Ashqar , Andry Rakotonirainy , Taqwa I. Alhadidi , Ahmed Jaber , Mohammad Abu Tami

Generating realistic 3D indoor scenes from user inputs remains a challenging problem in computer vision and graphics, requiring careful balance of geometric consistency, spatial relationships, and visual realism. While neural generation…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Mengqi Zhou , Xipeng Wang , Yuxi Wang , Zhaoxiang Zhang

We introduce MotionScript, a novel framework for generating highly detailed, natural language descriptions of 3D human motions. Unlike existing motion datasets that rely on broad action labels or generic captions, MotionScript provides…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Payam Jome Yazdian , Rachel Lagasse , Hamid Mohammadi , Eric Liu , Li Cheng , Angelica Lim

World building forms the foundation of any task that requires narrative intelligence. In this work, we focus on procedurally generating interactive fiction worlds---text-based worlds that players "see" and "talk to" using natural language.…

人工智能 · 计算机科学 2020-01-29 Prithviraj Ammanabrolu , Wesley Cheung , Dan Tu , William Broniec , Mark O. Riedl

Data-driven storytelling has gained prominence in journalism and other data reporting fields. However, the process of creating these stories remains challenging, often requiring the integration of effective visualizations with compelling…

人机交互 · 计算机科学 2025-04-01 Yu Fu , Dennis Bromley , Vidya Setlur

Keeping up with research and finding related work is still a time-consuming task for academics. Researchers sift through thousands of studies to identify a few relevant ones. Automation techniques can help by increasing the efficiency and…

信息检索 · 计算机科学 2023-09-06 Wojciech Kusa , Petr Knoth , Allan Hanbury

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Sensemaking and narrative are two inherently interconnected concepts about how people understand the world around them. Sensemaking is the process by which people structure and interconnect the information they encounter in the world with…

人工智能 · 计算机科学 2022-01-26 Zev Battad , Mei Si

Accurately reporting what objects are depicted in an image is largely a solved problem in automatic caption generation. The next big challenge on the way to truly humanlike captioning is being able to incorporate the context of the image…

计算与语言 · 计算机科学 2022-10-11 Sofia Nikiforova , Tejaswini Deoskar , Denis Paperno , Yoad Winter

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

人机交互 · 计算机科学 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

In today's world, time is a very important resource. In our busy lives, most of us hardly have time to read the complete news so what we have to do is just go through the headlines and satisfy ourselves with that. As a result, we might miss…

人机交互 · 计算机科学 2020-01-06 Mona teja K , Mohan Sai. S , H S S S Raviteja D , Sai Kushagra P

Immersive, stereoscopic viewing enables scientists to better analyze the spatial structures of visualized physical phenomena. However, their findings cannot be properly presented in traditional media, which lack these core attributes.…

图形学 · 计算机科学 2016-11-29 Jacqueline Chu , Leonardo Ferrer , Min Shih , Kwan-Liu Ma

Recent text-to-image generation methods provide a simple yet exciting conversion capability between text and image domains. While these methods have incrementally improved the generated image fidelity and text relevancy, several pivotal…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Oran Gafni , Adam Polyak , Oron Ashual , Shelly Sheynin , Devi Parikh , Yaniv Taigman

Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Shahaf Pruss , Morris Alper , Hadar Averbuch-Elor