中文
相关论文

相关论文: OSCBench: Benchmarking Object State Change in Text…

200 篇论文

Evaluating text-to-image generative models remains a challenge, despite the remarkable progress being made in their overall performances. While existing metrics like CLIPScore work for coarse evaluations, they lack the sensitivity to…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Georgia Gabriela Sampaio , Ruixiang Zhang , Shuangfei Zhai , Jiatao Gu , Josh Susskind , Navdeep Jaitly , Yizhe Zhang

Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order. Despite the emerging advancements in interleaved generation, the…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Minqian Liu , Zhiyang Xu , Zihao Lin , Trevor Ashby , Joy Rimchala , Jiaxin Zhang , Lifu Huang

Instruction-guided video editing has emerged as a rapidly advancing research direction, offering new opportunities for intuitive content transformation while also posing significant challenges for systematic evaluation. Existing video…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Yinan Chen , Jiangning Zhang , Teng Hu , Yuxiang Zeng , Zhucun Xue , Qingdong He , Chengjie Wang , Yong Liu , Xiaobin Hu , Shuicheng Yan

Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful…

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic…

Advances in technology have led to the development of methods that can create desired visual multimedia. In particular, image generation using deep learning has been extensively studied across diverse fields. In comparison, video…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Doyeon Kim , Donggyu Joo , Junmo Kim

Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form cinematic videos into detailed, temporally grounded scripts remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Junfu Pu , Yuxin Chen , Teng Wang , Ying Shan

Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. We present a principled theoretical framework that steers…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Eric Tillmann Bill , Enis Simsar , Thomas Hofmann

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Recent text-to-video (T2V) models have demonstrated strong capabilities in producing high-quality, dynamic videos. To improve the visual controllability, recent works have considered fine-tuning pre-trained T2V models to support…

计算机视觉与模式识别 · 计算机科学 2026-02-25 June Suk Choi , Kyungmin Lee , Sihyun Yu , Yisol Choi , Jinwoo Shin , Kimin Lee

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks such as VBench have…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Dian Zheng , Ziqi Huang , Hongbo Liu , Kai Zou , Yinan He , Fan Zhang , Lulu Gu , Yuanhan Zhang , Jingwen He , Wei-Shi Zheng , Yu Qiao , Ziwei Liu

The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yupeng Chen , Penglin Chen , Xiaoyu Zhang , Yixian Huang , Qian Xie

Recent advances in text-to-video (T2V) generation with diffusion models have garnered significant attention. However, they typically perform well in scenes with a single object and motion, struggling in compositional scenarios with multiple…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Yuanhang Li , Qi Mao , Lan Chen , Zhen Fang , Lei Tian , Xinyan Xiao , Libiao Jin , Hua Wu

Adapting text-to-image (T2I) latent diffusion models (LDMs) to video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships inherent to the video data generating process.…

Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in training data. We introduce T2I-BiasBench, a unified evaluation framework of thirteen…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Nihal Jaiswal , Siddhartha Arjaria , Gyanendra Chaubey , Ankush Kumar , Aditya Singh , Anchal Chaurasiya

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Yiming Ju , Jijin Hu , Zhengxiong Luo , Haoge Deng , hanyu Zhao , Li Du , Chengwei Wu , Donglin Hao , Xinlong Wang , Tengfei Pan

Video language continual learning involves continuously adapting to information from video and text inputs, enhancing a model's ability to handle new tasks while retaining prior knowledge. This field is a relatively under-explored area, and…

人工智能 · 计算机科学 2024-12-17 Tianqi Tang , Shohreh Deldari , Hao Xue , Celso De Melo , Flora D. Salim

Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, image quality, and…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Ziwei Huang , Wanggui He , Quanyu Long , Yandi Wang , Haoyuan Li , Zhelun Yu , Fangxun Shu , Long Chan , Hao Jiang , Fei Wu , Leilei Gan

Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. However, existing video understanding benchmarks largely overlook…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Zhoufaran Yang , Yan Shu , Jing Wang , Zhifei Yang , Yan Zhang , Yu Li , Keyang Lu , Gangyan Zeng , Shaohui Liu , Yu Zhou , Nicu Sebe

Sound content creation, essential for multimedia works such as video games and films, often involves extensive trial-and-error, enabling creators to semantically reflect their artistic ideas and inspirations, which evolve throughout the…

‹ 上一页 1 8 9 10 下一页 ›