English
Related papers

Related papers: CutClaw: Agentic Hours-Long Video Editing via Musi…

200 papers

Translating algorithms from high-level languages like MATLAB to hardware description languages (HDLs) is a resource-intensive but necessary step for deployment on FPGAs and ASICs. While large language models (LLMs) offer a path to…

Software Engineering · Computer Science 2025-12-18 Henry Gray , Tom Yotam , Octavian Udrea

Creating data stories from raw data is challenging due to humans' limited attention spans and the need for specialized skills. Recent advancements in large language models (LLMs) offer great opportunities to develop systems with autonomous…

Human-Computer Interaction · Computer Science 2024-08-08 Leixian Shen , Haotian Li , Yun Wang , Huamin Qu

Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are…

Artificial Intelligence · Computer Science 2026-04-10 Ziyu Ma , Shidong Yang , Yuxiang Ji , Xucong Wang , Yong Wang , Yiming Hu , Tongwen Huang , Xiangxiang Chu

While AI excels at generating text, audio, images, and videos, creating interactive audio-visual content such as video games remains challenging. Current LLMs can generate JavaScript games and animations, but lack automated evaluation…

Artificial Intelligence · Computer Science 2025-08-04 Alexia Jolicoeur-Martineau

The emergence of fake news on short video platforms has become a new significant societal concern, necessitating automatic video-news-specific detection. Current detectors primarily rely on pattern-based features to separate fake news…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Yuyan Bu , Qiang Sheng , Juan Cao , Shaofei Wang , Peng Qi , Yuhui Shi , Beizhe Hu

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progress in automating…

Artificial Intelligence · Computer Science 2026-04-13 Haobo Hu , Qi Mao , Yuanhang Li , Libiao Jin

Direct prompt-based editing often fails on complex transformations because vague and subjective prompts often require nuanced understanding of what should be changed in the image. Our core intuition is that leveraging compositional image…

Machine Learning · Computer Science 2026-03-10 Subhojyoti Mukherjee , Stefano Petrangeli , Branislav Kveton , Trung Bui , Franck Dernoncourt , Arko Mukherjee

Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yiwen Song , Tomas Pfister , Yale Song

Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some…

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding, weak fine-grained…

Computation and Language · Computer Science 2026-03-31 Yifan Zhu , Xinyu Mu , Tao Feng , Zhonghong Ou , Yuning Gong , Haoran Luo

Editing long videos remains a challenging task due to the need for maintaining both global consistency and temporal coherence across thousands of frames. Existing methods often suffer from structural drift or temporal artifacts,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Zichi Liu , Yinggui Wang , Tao Wei , Chao Ma

Recent advancements in Large Language Models (LLMs) have shown significant progress in understanding complex natural language. One important application of LLM is LLM-based AI Agent, which leverages the ability of LLM as well as external…

Computation and Language · Computer Science 2024-07-19 Zelong Li , Shuyuan Xu , Kai Mei , Wenyue Hua , Balaji Rama , Om Raheja , Hao Wang , He Zhu , Yongfeng Zhang

Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Xiaoyi Zhang , Zhaoyang Jia , Zongyu Guo , Jiahao Li , Bin Li , Houqiang Li , Yan Lu

The rapid growth of online video content, especially on short video platforms, has created a growing demand for efficient video editing techniques that can condense long-form videos into concise and engaging clips. Existing automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Xiangfeng Wang , Xiao Li , Yadong Wei , Xueyu Song , Yang Song , Xiaoqiang Xia , Fangrui Zeng , Zaiyi Chen , Liu Liu , Gu Xu , Tong Xu

We address the task of zero-shot video classification for extremely fine-grained actions (e.g., Windmill Dunk in basketball), where no video examples or temporal annotations are available for unseen classes. While image-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Amir Aghdam , Vincent Tao Hu , Björn Ommer

Constructing photorealistic virtual worlds has applications across various fields, but it often requires the extensive labor of highly trained professionals to operate conventional 3D modeling software. To democratize this process, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Xinhang Liu , Chi-Keung Tang , Yu-Wing Tai

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

Multiagent Systems · Computer Science 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Music videos, as a prevalent form of multimedia entertainment, deliver engaging audio-visual experiences to audiences and have gained immense popularity among singers and fans. Creators can express their interpretations of music naturally…

Human-Computer Interaction · Computer Science 2025-04-25 Chuer Chen , Shengqi Dang , Yuqi Liu , Nanxuan Zhao , Yang Shi , Nan Cao

Recent advances demonstrate that multimodal large language models (MLLMs) exhibit strong multimodal in-context learning (ICL) capabilities, enabling them to adapt to novel vision-language tasks from a few contextual examples. However,…

Artificial Intelligence · Computer Science 2025-10-07 Honghao Fu , Yuan Ouyang , Kai-Wei Chang , Yiwei Wang , Zi Huang , Yujun Cai

Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem,…

Sound · Computer Science 2024-03-08 Xia Liang , Xingjian Du , Jiaju Lin , Pei Zou , Yuan Wan , Bilei Zhu