中文
相关论文

相关论文: Technical Report for CVPR 2022 LOVEU AQTC Challeng…

200 篇论文

Existing affective understanding studies have mainly focused on recognizing emotions from images, audio signals, or pre-cliped video clips, where the affective evidence is already given. This passive and clip-centered setting does not fully…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Zhen Zhang , Yuhang Yang , Yunxiang Jiang , Yuhuan Lu , Haifeng Lu , Zheng Lian , Runhao Zeng , Xiping Hu

This paper presents our system for the Multi-Task Learning (MTL) Challenge in the 4th Affective Behavior Analysis in-the-wild (ABAW) competition. We explore the research problems of this challenge from three aspects: 1) For obtaining…

计算机视觉与模式识别 · 计算机科学 2022-08-31 Tenggan Zhang , Chuanhe Liu , Xiaolong Liu , Yuchen Liu , Liyu Meng , Lei Sun , Wenqiang Jiang , Fengyuan Zhang , Jinming Zhao , Qin Jin

This paper is dedicated to team VAA's approach submitted to the Fashion-IQ challenge in CVPR 2020. Given a pair of the image and the text, we present a novel multimodal composition method, RTIC, that can effectively combine the text and the…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Minchul Shin , Yoonjae Cho , Seongwuk Hong

This technical report presents the 1st winning model for UG2+, a task in CVPR 2024 UAV Tracking and Pose-Estimation Challenge. This challenge faces difficulties in drone detection, UAV-type classification and 2D/3D trajectory estimation in…

Query performance prediction (QPP) is an important and actively studied information retrieval task, having various applications, such as query reformulation, query expansion, and retrieval system selection, among many others. The task has…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Adrian Catalin Lutu , Eduard Poesina , Radu Tudor Ionescu

Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Bozheng Li , Yongliang Wu , Yi Lu , Jiashuo Yu , Licheng Tang , Jiawang Cao , Wenqing Zhu , Yuyang Sun , Jay Wu , Wenbo Zhu

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

This technical report describes our 2nd-place solution for the ECCV 2022 YouTube-VIS Long Video Challenge. We adopt the previously proposed online video instance segmentation method IDOL for this challenge. In addition, we use pseudo labels…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Junfeng Wu , Yi Jiang , Qihao Liu , Xiang Bai , Song Bai

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yudong Yang , Jimin Zhuang , Guangzhi Sun , Changli Tang , Yixuan Li , Peihan Li , Yifan Jiang , Wei Li , Zejun Ma , Chao Zhang

This paper presents the methodologies and results of the NOWJ team's participation across all five tasks at the COLIEE 2025 competition, emphasizing advancements in the Legal Case Entailment task (Task 2). Our comprehensive approach…

Complex video object segmentation continues to face significant challenges in small object recognition, occlusion handling, and dynamic scene modeling. This report presents our solution, which ranked second in the MOSE track of CVPR 2025…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Xuqiang Cao , Linnan Zhao , Jiaxuan Zhao , Fang Liu , Puhua Chen , Wenping Ma

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image…

图像与视频处理 · 电气工程与系统科学 2025-04-18 Xin Li , Kun Yuan , Bingchen Li , Fengbin Guan , Yizhen Shao , Zihao Yu , Xijun Wang , Yiting Lu , Wei Luo , Suhang Yao , Ming Sun , Chao Zhou , Zhibo Chen , Radu Timofte , Yabin Zhang , Ao-Xiang Zhang , Tianwu Zhi , Jianzhao Liu , Yang Li , Jingwen Xu , Yiting Liao , Yushen Zuo , Mingyang Wu , Renjie Li , Shengyun Zhong , Zhengzhong Tu , Yufan Liu , Xiangguang Chen , Zuowei Cao , Minhao Tang , Shan Liu , Kexin Zhang , Jingfen Xie , Yan Wang , Kai Chen , Shijie Zhao , Yunchen Zhang , Xiangkai Xu , Hong Gao , Ji Shi , Yiming Bao , Xiugang Dong , Xiangsheng Zhou , Yaofeng Tu , Ying Liang , Yiwen Wang , Xinning Chai , Yuxuan Zhang , Zhengxue Cheng , Yingsheng Qin , Yucai Yang , Rong Xie , Li Song , Wei Sun , Kang Fu , Linhan Cao , Dandan Zhu , Kaiwei Zhang , Yucheng Zhu , Zicheng Zhang , Menghan Hu , Xiongkuo Min , Guangtao Zhai , Zhi Jin , Jiawei Wu , Wei Wang , Wenjian Zhang , Yuhai Lan , Gaoxiong Yi , Hengyuan Na , Wang Luo , Di Wu , MingYin Bai , Jiawang Du , Zilong Lu , Zhenyu Jiang , Hui Zeng , Ziguan Cui , Zongliang Gan , Guijin Tang , Xinglin Xie , Kehuan Song , Xiaoqiang Lu , Licheng Jiao , Fang Liu , Xu Liu , Puhua Chen , Ha Thu Nguyen , Katrien De Moor , Seyed Ali Amirshahi , Mohamed-Chaker Larabi , Qi Tang , Linfeng He , Zhiyong Gao , Zixuan Gao , Guohua Zhang , Zhiye Huang , Yi Deng , Qingmiao Jiang , Lu Chen , Yi Yang , Xi Liao , Nourine Mohammed Nadir , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Meiqin Liu , Chao Yao , Yao Zhao

Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing is on the rise. Building on text-to-image models like Stable…

In the last few years, we have witnessed a renewed and fast-growing interest in continual learning with deep neural networks with the shared objective of making current AI systems more adaptive, efficient and autonomous. However, despite…

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Junjie Zhou , Yan Shu , Bo Zhao , Boya Wu , Zhengyang Liang , Shitao Xiao , Minghao Qin , Xi Yang , Yongping Xiong , Bo Zhang , Tiejun Huang , Zheng Liu

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu

Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Yuan Xie , Tianshui Chen , Zheng Ge , Lionel Ni

In this paper, we describe the solution to the QQ Browser 2021 Ai Algorithm Competition (AIAC) Track 1. We use the multi-modal transformer model for the video embedding extraction. In the pretrain phase, we train the model with three tasks,…

计算机视觉与模式识别 · 计算机科学 2021-11-03 Zhuoran Ma , Majing Lou , Xuan Ouyang

This paper describes the architecture of our system developed for Task 3 of SemEval-2024: Multimodal Emotion-Cause Analysis in Conversations. Our project targets the challenges of subtask 2, dedicated to Multimodal Emotion-Cause Pair…

计算与语言 · 计算机科学 2025-01-30 Meng Luo , Han Zhang , Shengqiong Wu , Bobo Li , Hong Han , Hao Fei

Facial behavior analysis is a broad topic with various categories such as facial emotion recognition, age, and gender recognition. Many studies focus on individual tasks while the multi-task learning approach is still an open research issue…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Dang-Khanh Nguyen , Sudarshan Pant , Ngoc-Huynh Ho , Guee-Sang Lee , Soo-Huyng Kim , Hyung-Jeong Yang