中文
相关论文

相关论文: VREN: Volleyball Rally Dataset with Expression Not…

200 篇论文

While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of real-world tables. Additionally, existing visual table…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Iñigo Alonso , Imanol Miranda , Eneko Agirre , Mirella Lapata

Visual abstract reasoning is core to image processing. We present Valen, a unified probability-highlighting baseline that excels on both RPM (progression) and Bongard-Logo (clustering) tasks. Analysing its internals, we find solvers…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Ruizhuo Song , Beiming Yuan

With the recent progress in sports analytics, deep learning approaches have demonstrated the effectiveness of mining insights into players' tactics for improving performance quality and fan engagement. This is attributed to the availability…

机器学习 · 计算机科学 2023-06-09 Wei-Yao Wang , Yung-Chang Huang , Tsi-Ui Ik , Wen-Chih Peng

Recent advances in immersive technology have opened new possibilities in sports training, especially for activities requiring precise motor skills, such as tennis. In this paper, we present a virtual reality (VR) tennis training system…

人机交互 · 计算机科学 2025-04-23 Ryan Najami , Rami Ghannam

Benchmark datasets play an important role in evaluating Natural Language Understanding (NLU) models. However, shortcuts -- unwanted biases in the benchmark datasets -- can damage the effectiveness of benchmark datasets in revealing models'…

人机交互 · 计算机科学 2023-01-16 Zhihua Jin , Xingbo Wang , Furui Cheng , Chunhui Sun , Qun Liu , Huamin Qu

Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in…

计算与语言 · 计算机科学 2022-08-01 Suwon Shon , Ankita Pasad , Felix Wu , Pablo Brusco , Yoav Artzi , Karen Livescu , Kyu J. Han

Offline reinforcement learning has emerged as a promising technology by enhancing its practicality through the use of pre-collected large datasets. Despite its practical benefits, most algorithm development research in offline reinforcement…

机器学习 · 计算机科学 2024-10-23 Dongsu Lee , Chanin Eom , Minhae Kwon

Reinforcement Learning with Verifiable Rewards (RLVR) has been successfully applied to significantly boost the capabilities of pretrained large language models, especially in the math and logic problem domains. However, current research and…

计算与语言 · 计算机科学 2026-03-16 Konstantin Dobler , Simon Lehnerer , Federico Scozzafava , Jonathan Janke , Mohamed Ali

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated benchmarks for…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Pritam Sarkar , Ali Etemad

Sports analysis and viewing play a pivotal role in the current sports domain, offering significant value not only to coaches and athletes but also to fans and the media. In recent years, the rapid development of virtual reality (VR) and…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Wenxuan Guo , Zhiyu Pan , Ziheng Xi , Alapati Tuerxun , Jianjiang Feng , Jie Zhou

Recently, there has been an increasing number of efforts to introduce models capable of generating natural language explanations (NLEs) for their predictions on vision-language (VL) tasks. Such models are appealing, because they can provide…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Maxime Kayser , Oana-Maria Camburu , Leonard Salewski , Cornelius Emde , Virginie Do , Zeynep Akata , Thomas Lukasiewicz

Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Neha Balamurugan , Sarah Wu , Adam Chun , Gabe Gaw , Cristobal Eyzaguirre , Tobias Gerstenberg

This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable of perceiving and reasoning about specific AVR problems,…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Ke Zhu , Yu Wang , Jiangjiang Liu , Qunyi Xie , Shanshan Liu , Gang Zhang

Open-Vocabulary Multimodal Emotion Recognition (OV-MER) aims to predict emotions without being constrained by predefined label spaces, thereby enabling fine-grained emotion understanding. Unlike traditional discriminative methods, OV-MER…

人机交互 · 计算机科学 2026-05-08 Zheng Lian , Fan Zhang , Lan Chen , Yazhou Zhang , Rui Liu , Jinyang Wu , Haoyu Chen , Xiaobai Li , Xiaojiang Peng , Bin He , Jianhua Tao

Identifying significant shots in a rally is important for evaluating players' performance in badminton matches. While there are several studies that have quantified player performance in other sports, analyzing badminton data is remained…

机器学习 · 计算机科学 2021-09-15 Wei-Yao Wang , Teng-Fong Chan , Hui-Kuo Yang , Chih-Chuan Wang , Yao-Chung Fan , Wen-Chih Peng

Tracking using bio-inspired event cameras has drawn more and more attention in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The first category needs…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Xiao Wang , Shiao Wang , Chuanming Tang , Lin Zhu , Bo Jiang , Yonghong Tian , Jin Tang

The immense popularity of racket sports has fueled substantial demand in tactical analysis with broadcast videos. However, existing manual methods require laborious annotation, and recent attempts leveraging video perception models are…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Yuchen He , Zeqing Yuan , Yihong Wu , Liqi Cheng , Dazhen Deng , Yingcai Wu

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or…

计算机视觉与模式识别 · 计算机科学 2022-05-05 Wanrong Zhu , Yuankai Qi , Pradyumna Narayana , Kazoo Sone , Sugato Basu , Xin Eric Wang , Qi Wu , Miguel Eckstein , William Yang Wang

This paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relevant to the given query, from a great number of unlabeled…

计算机视觉与模式识别 · 计算机科学 2022-03-04 Jianfeng Dong , Yabing Wang , Xianke Chen , Xiaoye Qu , Xirong Li , Yuan He , Xun Wang

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we…

音频与语音处理 · 电气工程与系统科学 2024-07-19 Minsu Kim , Jeong Hun Yeo , Se Jin Park , Hyeongseop Rha , Yong Man Ro