中文
相关论文

相关论文: LiveBot: Generating Live Video Comments Based on V…

200 篇论文

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices like emotions,…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Aanisha Bhattacharya , Yaman K Singla , Balaji Krishnamurthy , Rajiv Ratn Shah , Changyou Chen

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

Current dialogue systems focus more on textual and speech context knowledge and are usually based on two speakers. Some recent work has investigated static image-based dialogue. However, several real-world human interactions also involve…

计算与语言 · 计算机科学 2018-10-18 Ramakanth Pasunuru , Mohit Bansal

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

Learning tasks through videos is a dynamic way to acquire skills by witnessing entire processes. However, compared to in-person demonstrations, videos may omit tacit knowledge, including subtle details and contextual nuances. Users' unique…

人机交互 · 计算机科学 2026-03-16 Nayoung Kim , Yotam Sechayk , Zhongyi Zhou , Takeo Igarashi

Much research in recent years has focused on automatic article commenting. However, few of previous studies focus on the controllable generation of comments. Besides, they tend to generate dull and commonplace comments, which further limits…

计算与语言 · 计算机科学 2021-07-27 Linhao Zhang , Houfeng Wang

An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Xingyi Zhou , Anurag Arnab , Shyamal Buch , Shen Yan , Austin Myers , Xuehan Xiong , Arsha Nagrani , Cordelia Schmid

Tutorial videos of mobile apps have become a popular and compelling way for users to learn unfamiliar app features. To make the video accessible to the users, video creators always need to annotate the actions in the video, including what…

人机交互 · 计算机科学 2023-08-08 Sidong Feng , Chunyang Chen , Zhenchang Xing

In this paper, we introduce LiveQA, a new question answering dataset constructed from play-by-play live broadcast. It contains 117k multiple-choice questions written by human commentators for over 1,670 NBA games, which are collected from…

计算与语言 · 计算机科学 2020-10-02 Qianying Liu , Sicong Jiang , Yizhong Wang , Sujian Li

Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Sherry Yang , Jacob Walker , Jack Parker-Holder , Yilun Du , Jake Bruce , Andre Barreto , Pieter Abbeel , Dale Schuurmans

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

Video accessibility is crucial for blind and low vision users for equitable engagements in education, employment, and entertainment. Despite the availability of professional and amateur services and tools, most human-generated descriptions…

Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their correspondence to the world.…

计算与语言 · 计算机科学 2017-10-03 Stephanie Zhou , Alane Suhr , Yoav Artzi

Video live streaming is gaining prevalence among video streaming services, especially for the delivery of popular sporting events. Many objective Video Quality Assessment (VQA) models have been developed to predict the perceptual quality of…

图像与视频处理 · 电气工程与系统科学 2021-06-17 Zaixi Shang , Joshua P. Ebenezer , Alan C. Bovik , Yongjun Wu , Hai Wei , Sriram Sethuraman

The recent progress on image recognition and language modeling is making automatic description of image content a reality. However, stylized, non-factual aspects of the written description are missing from the current systems. One such…

计算机视觉与模式识别 · 计算机科学 2015-12-15 Alexander Mathews , Lexing Xie , Xuming He

Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we…

With the increased adoption of E-learning platforms, keeping online learners engaged throughout a lesson is challenging. One approach to tackle this challenge is to probe learn-ers periodically by asking questions. The paper presents an…

人机交互 · 计算机科学 2021-06-08 Ritu Gala , Revathi Vijayaraghavan , Valmik Nikam , Arvind Kiwelekar

Bots are frequently used in Github repositories to automate repetitive activities that are part of the distributed software development process. They communicate with human actors through comments. While detecting their presence is…

软件工程 · 计算机科学 2021-01-29 Mehdi Golzadeh , Alexandre Decan , Damien Legay , Tom Mens

A key challenge in the accurate prediction of viewers' emotional responses to video stimuli in real-world applications is accounting for person- and situation-specific variation. An important contextual influence shaping individuals'…

人机交互 · 计算机科学 2020-08-28 Bernd Dudzik , Joost Broekens , Mark Neerincx , Hayley Hung