中文
相关论文

相关论文: LOGO: A Long-Form Video Dataset for Group Action Q…

200 篇论文

The ability to quantify how well an action is carried out, also known as action quality assessment (AQA), has attracted recent interest in the vision community. Unfortunately, prior methods often ignore the score rubric used by human…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Abrar Majeedi , Viswanatha Reddy Gajjala , Satya Sai Srinath Namburi GNVV , Yin Li

Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Jie Ma , Min Hu , Pinghui Wang , Wangchun Sun , Lingyun Song , Hongbin Pei , Jun Liu , Youtian Du

Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models (VLMs) hold considerable promise for AQA, their actual performance in this domain remains…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Miguel Monte e Freitas , Rui Henriques , Ricardo Rei , Pedro Henrique Martins

The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Ziheng Jia , Zicheng Zhang , Jiaying Qian , Haoning Wu , Wei Sun , Chunyi Li , Xiaohong Liu , Weisi Lin , Guangtao Zhai , Xiongkuo Min

To understand human behaviors, action recognition based on videos is a common approach. Compared with image-based action recognition, videos provide much more information. Reducing the ambiguity of actions and in the last decade, many works…

计算机视觉与模式识别 · 计算机科学 2022-06-03 Fei Wu , Qingzhong Wang , Jian Bian , Haoyi Xiong , Ning Ding , Feixiang Lu , Jun Cheng , Dejing Dou

Spatio-temporal action detection is an important and challenging problem in video understanding. However, the application of the existing large-scale spatio-temporal action datasets in specific fields is limited, and there is currently no…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Fan Yang

We introduce a novel question-answering (QA) dataset using echocardiogram reports sourced from the Medical Information Mart for Intensive Care database. This dataset is specifically designed to enhance QA systems in cardiology, consisting…

人工智能 · 计算机科学 2025-03-07 Lama Moukheiber , Mira Moukheiber , Dana Moukheiiber , Jae-Woo Ju , Hyung-Chul Lee

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Yansong Tang , Jinpeng Liu , Aoyang Liu , Bin Yang , Wenxun Dai , Yongming Rao , Jiwen Lu , Jie Zhou , Xiu Li

Retrieval-augmented generation (RAG) on specialized domain datasets has shown improved performance when large language models (LLMs) are fine-tuned for generating responses to user queries. In this study, we develop a cybersecurity…

机器学习 · 计算机科学 2024-11-05 Varun Badrinath Krishna

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to investigating the…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Yansong Tang , Jiwen Lu , Jie Zhou

The prevalence of large-scale multimodal datasets presents unique challenges in assessing dataset quality. We propose a two-step method to analyze multimodal datasets, which leverages a small seed of human annotation to map each multimodal…

计算机视觉与模式识别 · 计算机科学 2023-07-11 Netta Madvil , Yonatan Bitton , Roy Schwartz

Group activity detection in soccer can be done by using either video data or player and ball trajectory data. In current soccer activity datasets, activities are labelled as atomic events without a duration. Given that the state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2020-04-23 Ryan Sanford , Siavash Gorji , Luiz G. Hafemann , Bahareh Pourbabaee , Mehrsan Javan

Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video…

图像与视频处理 · 电气工程与系统科学 2024-12-12 Qi Zheng , Yibo Fan , Leilei Huang , Tianyu Zhu , Jiaming Liu , Zhijian Hao , Shuo Xing , Chia-Ju Chen , Xiongkuo Min , Alan C. Bovik , Zhengzhong Tu

The growing demand for AI training data has transformed data annotation into a global industry, but traditional approaches relying on human annotators are often time-consuming, labor-intensive, and prone to inconsistent quality. We propose…

We live in a world filled with never-ending streams of multimodal information. As a more natural recording of the real scenario, long form audio-visual videos are expected as an important bridge for better exploring and understanding the…

多媒体 · 计算机科学 2023-06-19 Wenxuan Hou , Guangyao Li , Yapeng Tian , Di Hu

Combining sports and machine learning involves leveraging ML algorithms and techniques to extract insight from sports-related data such as player statistics, game footage, and other relevant information. However, datasets related to figure…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Wei-Yi Chen , Yi-Ling Lin , Yu-An Su , Wei-Hsin Yeh , Lun-Wei Ku

Visual Question Answering (VQA) entails answering questions about images. We introduce the first VQA dataset in which all contents originate from an authentic use case. Sourced from online question answering community forums, we call it…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Chongyan Chen , Mengchen Liu , Noel Codella , Yunsheng Li , Lu Yuan , Danna Gurari

We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a…

Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly help to accelerate research in this area by providing a…

机器学习 · 计算机科学 2023-01-31 Yanfu Yan , Ke Lu , Jian Xue , Pengcheng Gao , Jiayi Lyu

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

计算机视觉与模式识别 · 计算机科学 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta