中文
相关论文

相关论文: DescribePro: Collaborative Audio Description with …

200 篇论文

In human-AI collaboration, a central challenge is deciding whether the AI should handle a task, be deferred to a human expert, or be addressed through collaborative effort. Existing Learning to Defer approaches typically make binary choices…

人工智能 · 计算机科学 2025-05-27 Chengbo He , Bochao Zou , Junliang Xing , Jiansheng Chen , Yuanchun Shi , Huimin Ma

People with visual impairments perceive their environment non-visually and often use AI-powered assistive tools to obtain textual descriptions of visual information. Recent large vision-language model-based AI-powered tools like Be My AI…

人机交互 · 计算机科学 2024-07-15 Jingyi Xie , Rui Yu , He Zhang , Sooyeon Lee , Syed Masum Billah , John M. Carroll

We introduce Cap3D, an automatic approach for generating descriptive text for 3D objects. This approach utilizes pretrained models from image captioning, image-text alignment, and LLM to consolidate captions from multiple views of a 3D…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Tiange Luo , Chris Rockwell , Honglak Lee , Justin Johnson

Intelligent assistants that follow commands or answer simple questions, such as Siri and Google search, are among the most economically important applications of AI. Future conversational AI assistants promise even greater capabilities and…

人工智能 · 计算机科学 2020-08-28 Katya Kudashkina , Patrick M. Pilarski , Richard S. Sutton

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

Manual annotation remains the gold standard for high-quality, dense temporal video datasets, yet it is inherently time-consuming. Vision-language models can aid human annotators and expedite this process. We report on the impact of…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Juan Gutiérrez , Victor Gutiérrez , Ángel Mora , Silvia Rodriguez , José Luis Blanco

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

This paper aims to explore fundamental questions in the era when AI coding assistants like GitHub Copilot are widely adopted: what do developers truly value and criticize in AI coding assistants, and what does this reveal about their needs…

软件工程 · 计算机科学 2025-08-19 Yunbo Lyu , Zhou Yang , Jieke Shi , Jianming Chang , Yue Liu , David Lo

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Hongbin Huang , Junwei Li , Tianxin Xie , Zhuang Li , Cekai Weng , Yaodong Yang , Yue Luo , Li Liu , Jing Tang , Zhijing Shao , Zeyu Wang

Real-time captioning is a useful technique for deaf and hard-of-hearing (DHH) people to talk to hearing people. With the improvement in device performance and the accuracy of automatic speech recognition (ASR), real-time captioning is…

人机交互 · 计算机科学 2021-01-28 Kenta Yamamoto , Ippei Suzuki , Akihisa Shitara , Yoichi Ochiai

As an important component of multimedia analysis tasks, audio classification aims to discriminate between different audio signal types and has received intensive attention due to its wide applications. Generally speaking, the raw signal can…

多媒体 · 计算机科学 2020-02-25 Liang Gao , Kele Xu , Huaimin Wang , Yuxing Peng

Artificial intelligence is increasingly being integrated into professional audio production workflows, yet a gap persists between the tools developers produce and the requirements of practising sound designers. This paper investigates this…

声音 · 计算机科学 2026-05-27 Nelly Garcia , Joshua Reiss

Enforcing archival standards requires specialized expertise, and manually creating metadata descriptions for archival materials is a tedious and error-prone task. This work aims at exploring the potential of agentic AI and large language…

人工智能 · 计算机科学 2026-02-03 Jinghua Groppe , Andreas Marquet , Annabel Walz , Sven Groppe

In AI-assisted decision-making, humans often passively review AI's suggestion and decide whether to accept or reject it as a whole. In such a paradigm, humans are found to rarely trigger analytical thinking and face difficulties in…

人机交互 · 计算机科学 2025-03-13 Shuai Ma , Qiaoyi Chen , Xinru Wang , Chengbo Zheng , Zhenhui Peng , Ming Yin , Xiaojuan Ma

Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and effective method for improving alignment, typically through training-time or prompt-based…

机器学习 · 计算机科学 2025-10-01 Frédéric Berdoz , Luca A. Lanzendörfer , René Caky , Roger Wattenhofer

The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex task that requires considering diverse factors, among them,…

计算与语言 · 计算机科学 2025-08-12 Tsung-Han Wu , Joseph E. Gonzalez , Trevor Darrell , David M. Chan

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

Automatic Speech Recognition (ASR) systems have proliferated over the recent years to the point that free platforms such as YouTube now provide speech recognition services. Given the wide selection of ASR systems, we contribute to the field…

Large language models have introduced exciting new opportunities and challenges in designing and developing new AI-assisted writing support tools. Recent work has shown that leveraging this new technology can transform writing in many…

With the advancement of generative models, the synthesis of different sensory elements such as music, visuals, and speech has achieved significant realism. However, the approach to generate multi-sensory outputs has not been fully explored,…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Minheng Ni , Chenfei Wu , Huaying Yuan , Zhengyuan Yang , Ming Gong , Lijuan Wang , Zicheng Liu , Wangmeng Zuo , Nan Duan
‹ 上一页 1 8 9 10 下一页 ›