English
Related papers

Related papers: Script-to-Slide Grounding: Grounding Script Senten…

200 papers

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Evangelos Kazakos , Cordelia Schmid , Josef Sivic

Natural language spatial video grounding aims to detect the relevant objects in video frames with descriptive sentences as the query. In spite of the great advances, most existing methods rely on dense video frame annotations, which require…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Mengze Li , Tianbao Wang , Haoyu Zhang , Shengyu Zhang , Zhou Zhao , Jiaxu Miao , Wenqiao Zhang , Wenming Tan , Jin Wang , Peng Wang , Shiliang Pu , Fei Wu

In this paper, we study the problem of weakly-supervised temporal grounding of sentence in video. Specifically, given an untrimmed video and a query sentence, our goal is to localize a temporal segment in the video that semantically…

Computer Vision and Pattern Recognition · Computer Science 2020-01-28 Zhenfang Chen , Lin Ma , Wenhan Luo , Peng Tang , Kwan-Yee K. Wong

Temporal Sentence Grounding (TSG) aims to identify relevant moments in an untrimmed video that semantically correspond to a given textual query. Despite existing studies having made substantial progress, they often overlook the issue of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Kefan Tang , Lihuo He , Jisheng Dang , Xinbo Gao

Among numerous videos shared on the web, well-edited ones always attract more attention. However, it is difficult for inexperienced users to make well-edited videos because it requires professional expertise and immense manual labor. To…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Yu Xiong , Fabian Caba Heilbron , Dahua Lin

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Panwen Hu , Nan Xiao , Feifei Li , Yongquan Chen , Rui Huang

A vast amount of audio-visual data is available on the Internet thanks to video streaming services, to which users upload their content. However, there are difficulties in exploiting available data for supervised statistical models due to…

Multimedia · Computer Science 2019-07-30 Yasufumi Moriya , Ramon Sanabria , Florian Metze , Gareth J. F. Jones

Text-to-3D generation is a valuable technology in virtual reality and digital content creation. While recent works have pushed the boundaries of text-to-3D generation, producing high-fidelity 3D objects with inefficient prompts and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Wenqing Wang , Yun Fu

Temporal sentence grounding in videos (TSGV) aims to localize a temporal segment that semantically corresponds to a sentence query from an untrimmed video. Most current methods adopt pre-trained query-agnostic visual encoders for offline…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Allen He , Qi Liu , Kun Liu , Xinchen Liu , Wu Liu

The proliferation of online short video platforms has driven a surge in user demand for short video editing. However, manually selecting, cropping, and assembling raw footage into a coherent, high-quality video remains laborious and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Zhihui Yin , Ye Ma , Xipeng Cao , Bo Wang , Quan Chen , Peng Jiang

Temporal Sentence Grounding in Videos (TSGV) aims to detect the event timestamps described by the natural language query from untrimmed videos. This paper discusses the challenge of achieving efficient computation in TSGV models while…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Renjie Liang , Yiming Yang , Hui Lu , Li Li

Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-based agents powered by Multimodal LLMs (MLLMs) excel at tasks requiring visual perception,…

Computation and Language · Computer Science 2026-05-12 Kyudan Jung , Hojun Cho , Jooyeol Yun , Soyoung Yang , Jaehyeok Jang , Jaegul Choo

This paper aims to tackle a novel task - Temporal Sentence Grounding in Streaming Videos (TSGSV). The goal of TSGSV is to evaluate the relevance between a video stream and a given sentence query. Unlike regular videos, streaming videos are…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Tian Gan , Xiao Wang , Yan Sun , Jianlong Wu , Qingpei Guo , Liqiang Nie

Classical video quality assessment methods generate a numerical score to judge a video's perceived visual fidelity and clarity. Yet, a score fails to describe the video's complex quality dimensions, restricting its applicability. Benefiting…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Qizhi Xie , Kun Yuan , Yunpeng Qu , Jiachao Gong , Mingda Wu , Ming Sun , Chao Zhou , Jihong Zhu

Automatic assessment of code, in particular to support education, is an important feature included in several Learning Management Systems (LMS), at least to some extent. Several kinds of assessments can be designed, such as exercises asking…

Software Engineering · Computer Science 2019-11-28 Sébastien Combéfis , Guillaume de Moffarts

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence…

Computer Vision and Pattern Recognition · Computer Science 2019-11-01 Yitian Yuan , Lin Ma , Jingwen Wang , Wei Liu , Wenwu Zhu

The rapid advancements in Large Language Models (LLMs) have revolutionized educational technology, enabling innovative approaches to automated and personalized content creation. This paper introduces Slide2Text, a system that leverages LLMs…

Artificial Intelligence · Computer Science 2025-03-25 Yizhou Zhou

We introduce an approach to generating videos based on a series of given language descriptions. Frames of the video are generated sequentially and optimized by guidance from the CLIP image-text encoder; iterating through language…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Peter Schaldenbrand , Zhixuan Liu , Jean Oh

Diffusion models have achieved impressive results in generative tasks for text-to-video (T2V) synthesis. However, achieving accurate text alignment in T2V generation remains challenging due to the complex temporal dependencies across…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jaemin Kim , Bryan Sangwoo Kim , Jong Chul Ye

Generating presentation slides is a time-consuming task that urgently requires automation. Due to their limited flexibility and lack of automated refinement mechanisms, existing autonomous LLM-based agents face constraints in real-world…

Computation and Language · Computer Science 2025-02-24 Yunqing Xu , Xinbei Ma , Jiyang Qiu , Hai Zhao