English
Related papers

Related papers: GUIDE: A Guideline-Guided Dataset for Instructiona…

200 papers

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Yufan Deng , Zilin Pan , Hongyu Zhang , Xiaojie Li , Ruoqing Hu , Yufei Ding , Yiming Zou , Yan Zeng , Daquan Zhou

Self-supervised learning is an effective way for label-free model pre-training, especially in the video domain where labeling is expensive. Existing self-supervised works in the video domain use varying experimental setups to demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Akash Kumar , Ashlesha Kumar , Vibhav Vineet , Yogesh Singh Rawat

Despite widespread deployment of Large Language Models, systematic evaluation of instruction-following capabilities remains challenging. While comprehensive benchmarks exist, focused assessments that quickly diagnose specific instruction…

Computation and Language · Computer Science 2025-10-23 Richard J. Young , Brandon Gillins , Alice M. Matthews

Instructors often rely on visual actions such as pointing, marking, and sketching to convey information in educational presentation videos. These subtle visual cues often lack verbal descriptions, forcing low-vision (LV) learners to search…

Human-Computer Interaction · Computer Science 2025-08-06 Yotam Sechayk , Ariel Shamir , Amy Pavel , Takeo Igarashi

The recent rapid advancement of machine learning has been driven by increasingly powerful models with the growing availability of training data and computational resources. However, real-time decision-making tasks with limited time and…

Machine Learning · Computer Science 2024-10-22 Lingyu Zhang , Zhengran Ji , Nicholas R Waytowich , Boyuan Chen

Imitation learning field requires expert data to train agents in a task. Most often, this learning approach suffers from the absence of available data, which results in techniques being tested on its dataset. Creating datasets is a…

Machine Learning · Computer Science 2024-03-04 Nathan Gavenski , Michael Luck , Odinaldo Rodrigues

Instruction-based image editing aims to modify specific content within existing images according to user-provided instructions while preserving non-target regions. Beyond traditional object- and style-centric manipulation, text-centric…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Hui Zhang , Juntao Liu , Zongkai Liu , Liqiang Niu , Fandong Meng , Zuxuan Wu , Yu-Gang Jiang

A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed. Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of…

This work concerns video-language pre-training and representation learning. In this now ubiquitous training scheme, a model first performs pre-training on paired videos and text (e.g., video clips and accompanied subtitles) from a large…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Luowei Zhou , Jingjing Liu , Yu Cheng , Zhe Gan , Lei Zhang

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Matthew Gwilliam , Michael Cogswell , Meng Ye , Karan Sikka , Abhinav Shrivastava , Ajay Divakaran

Recent language models have achieved impressive performance in natural language tasks by incorporating instructions with task input during fine-tuning. Since all samples in the same natural language task can be explained with the same task…

Computation and Language · Computer Science 2023-11-14 Jin Myung Kwak , Minseon Kim , Sung Ju Hwang

In this paper, we suggest a novel method to help learners find relevant open educational videos to master skills demanded on the labour market. We have built a prototype, which 1) applies text classification and text mining methods on job…

Computers and Society · Computer Science 2020-05-22 Mohammadreza Tavakoli , Sherzod Hakimov , Ralph Ewerth , Gábor Kismihók

Recent advances in visual analytics have enabled us to learn from user interactions and uncover analytic goals. These innovations set the foundation for actively guiding users during data exploration. Providing such guidance will become…

Human-Computer Interaction · Computer Science 2022-07-19 Shayan Monadjemi , Sunwoo Ha , Quan Nguyen , Henry Chai , Roman Garnett , Alvitta Ottley

Facial expression captioning has found widespread application across various domains. Recently, the emergence of video Multimodal Large Language Models (MLLMs) has shown promise in general video understanding tasks. However, describing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Jiaxing Zhao , Boyuan Sun , Xiang Chen , Xihan Wei

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA).…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zhou Yu , Dejing Xu , Jun Yu , Ting Yu , Zhou Zhao , Yueting Zhuang , Dacheng Tao

Computer-Aided Design (CAD) is a time-consuming and complex process, requiring precise, long-horizon user interactions with intricate 3D interfaces. While recent advances in AI-driven user interface (UI) agents show promise, most existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Brandon Man , Ghadi Nehme , Md Ferdous Alam , Faez Ahmed

Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Michele Mazzamuto , Daniele Di Mauro , Gianpiero Francesca , Giovanni Maria Farinella , Antonino Furnari

The demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times. Despite notable advancements in the fields of video summarization and highlight detection, which can create…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yongliang Wu , Wenbo Zhu , Jiawang Cao , Yi Lu , Bozheng Li , Weiheng Chi , Zihan Qiu , Lirian Su , Haolin Zheng , Jay Wu , Xu Yang

Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ruchit Rawal , Khalid Saifullah , Miquel Farré , Ronen Basri , David Jacobs , Gowthami Somepalli , Tom Goldstein

We introduce VisIT-Bench (Visual InsTruction Benchmark), a benchmark for evaluation of instruction-following vision-language models for real-world use. Our starting point is curating 70 'instruction families' that we envision instruction…

Computation and Language · Computer Science 2023-12-27 Yonatan Bitton , Hritik Bansal , Jack Hessel , Rulin Shao , Wanrong Zhu , Anas Awadalla , Josh Gardner , Rohan Taori , Ludwig Schmidt