English
Related papers

Related papers: Results of the 2024 Video Browser Showdown

200 papers

This report summarises the talks and discussions that took place over the course of the MITP Youngst@rs Colours in Darkness workshop 2023. All talks can be found at https://indico.mitp.uni-mainz.de/event/377/.

The third Pixel-level Video Understanding in the Wild (PVUW CVPR 2024) challenge aims to advance the state of art in video understanding through benchmarking Video Panoptic Segmentation (VPS) and Video Semantic Segmentation (VSS) on…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Qingfeng Liu , Mostafa El-Khamy , Kee-Bong Song

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict…

Machine Learning · Computer Science 2025-03-11 Andrew Chang , Viswadruth Akkaraju , Ray McFadden Cogliano , David Poeppel , Dustin Freeman

The query-based moment retrieval is a problem of localising a specific clip from an untrimmed video according a query sentence. This is a challenging task that requires interpretation of both the natural language query and the video…

Computer Vision and Pattern Recognition · Computer Science 2020-10-08 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä

This volume contains the proceedings of the 14th International Symposium on Games, Automata, Logics, and Formal Verification (GandALF 2023). The aim of GandALF 2023 symposium is to bring together researchers from academia and industry who…

Formal Languages and Automata Theory · Computer Science 2023-10-02 Antonis Achilleos , Dario Della Monica

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Purui Bai , Tao Wu , Jiayang Sun , Xinyue Liu , Huaibo Huang , Ran He

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching…

Volume with the Late-Breaking Abstracts submitted to the Evo* 2022 Conference, held in Madrid (Spain), from 20 to 22 of April. These papers present ongoing research and preliminary results investigating on the application of different…

Neural and Evolutionary Computing · Computer Science 2022-08-02 A. M. Mora , A. I. Esparcia-Alcázar

This paper presents the Video Super-Resolution (SR) Quality Assessment (QA) Challenge that was part of the Advances in Image Manipulation (AIM) workshop, held in conjunction with ECCV 2024. The task of this challenge was to develop an…

The contribution contains the preface to the Proceedings to the 23rd International Workshop "What Comes Beyond the Standard Models", July 04 -- July 12, 2020, Bled, Slovenia, [Virtual Workshop -- July 6.--10. 2020], Volume 1: Invited Talks…

General Physics · Physics 2022-12-16 N. S. Mankoč Borštnik , H. F. B. Nielsen , M. Y. Khlopov , D. Lukman

Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from both academia and industry. However, as new and improved LLMs are developed, existing evaluation…

Software Engineering · Computer Science 2024-06-07 Naman Jain , King Han , Alex Gu , Wen-Ding Li , Fanjia Yan , Tianjun Zhang , Sida Wang , Armando Solar-Lezama , Koushik Sen , Ion Stoica

Video object segmentation is a challenging task that serves as the cornerstone of numerous downstream applications, including video editing and autonomous driving. In this technical report, we briefly introduce the solution of our team…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Jinming Chai , Qin Ma , Junpei Zhang , Licheng Jiao , Fang Liu

This is the proceedings of the 1st International Workshop on Low Carbon Computing (LOCO 2024).

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-07 Wim Vanderbauwhede , Lauritz Thamsen , José Cano

This paper is a report of the Workshop on Simulations for Information Access (Sim4IA) workshop at SIGIR 2024. The workshop had two keynotes, a panel discussion, nine lightning talks, and two breakout sessions. Key takeaways were user…

This volume contains the papers presented at the Tenth International Workshop on Developments in Computational Models (DCM) held in Vienna, Austria on 13th July 2014, as part of the Vienna Summer of Logic. Several new models of computation…

Logic in Computer Science · Computer Science 2015-04-09 Ugo Dal Lago , Russ Harmer

This is the arXiv index for the electronic proceedings of GD 2023, which contains the peer-reviewed and revised accepted papers with an optional appendix. Proceedings (without appendices) are also to be published by Springer in the Lecture…

Computational Geometry · Computer Science 2023-09-15 Michael A. Bekos , Markus Chimani

Localizing events in videos based on semantic queries is a pivotal task in video understanding, with the growing significance of user-oriented applications like video search. Yet, current research predominantly relies on natural language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Gengyuan Zhang , Mang Ling Ada Fok , Jialu Ma , Yan Xia , Daniel Cremers , Philip Torr , Volker Tresp , Jindong Gu

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for generative models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Meiqi Wu , Zhixin Cai , Fufangchen Zhao , Xiaokun Feng , Rujing Dang , Bingze Song , Ruitian Tian , Jiashu Zhu , Jiachen Lei , Hao Dou , Jing Tang , Lei Sun , Jiahong Wu , Xiangxiang Chu , Zeming Liu , Kaiqi Huang