English
Related papers

Related papers: Scalable Video-to-Dataset Generation for Cross-Pla…

200 papers

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Ruiyuan Lyu , Jingli Lin , Tai Wang , Shuai Yang , Xiaohan Mao , Yilun Chen , Runsen Xu , Haifeng Huang , Chenming Zhu , Dahua Lin , Jiangmiao Pang

The visual world around us constantly evolves, from real-time news and social media trends to global infrastructure changes visible through satellite imagery and augmented reality enhancements. However, Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Mingyang Fu , Yuyang Peng , Dongping Chen , Zetong Zhou , Benlin Liu , Yao Wan , Zhou Zhao , Philip S. Yu , Ranjay Krishna

Deep learning techniques have enabled the emergence of state-of-the-art models to address object detection tasks. However, these techniques are data-driven, delegating the accuracy to the training dataset which must resemble the images in…

Computer Vision and Pattern Recognition · Computer Science 2020-07-13 Vinicius F. Arruda , Thiago M. Paixão , Rodrigo F. Berriel , Alberto F. De Souza , Claudine Badue , Nicu Sebe , Thiago Oliveira-Santos

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and process multimodal data have grown. This survey provides a…

Artificial Intelligence · Computer Science 2025-09-16 Biao Wu , Yanda Li , Zhiwei Zhang , Yunchao Wei , Meng Fang , Ling Chen

Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% J&F) on existing datasets. However, since the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Henghui Ding , Chang Liu , Shuting He , Xudong Jiang , Philip H. S. Torr , Song Bai

With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models. However, most benchmarks predominantly assess…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Kunchang Li , Yali Wang , Yinan He , Yizhuo Li , Yi Wang , Yi Liu , Zun Wang , Jilan Xu , Guo Chen , Ping Luo , Limin Wang , Yu Qiao

Vision-language Navigation (VLN) tasks require an agent to navigate step-by-step while perceiving the visual observations and comprehending a natural language instruction. Large data bias, which is caused by the disparity ratio between the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Chong Liu , Fengda Zhu , Xiaojun Chang , Xiaodan Liang , Zongyuan Ge , Yi-Dong Shen

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Zirui Pan , Xin Wang , Yipeng Zhang , Hong Chen , Kwan Man Cheng , Yaofei Wu , Wenwu Zhu

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

Artificial Intelligence · Computer Science 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

Vision-and-language navigation (VLN) requires an embodied agent to navigate in realistic 3D environments using natural language instructions. Existing VLN methods suffer from training on small-scale environments or unreasonable…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Kunyang Lin , Peihao Chen , Diwei Huang , Thomas H. Li , Mingkui Tan , Chuang Gan

The advancement of mobile GUI agents has opened new opportunities for automating tasks on mobile devices. Training these agents requires large-scale high-quality data, which is prohibitively expensive when relying on human labor. Given the…

Artificial Intelligence · Computer Science 2025-05-21 Wenhao Wang , Mengying Yuan , Zijie Yu , Guangyi Liu , Rui Ye , Tian Jin , Siheng Chen , Yanfeng Wang

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Aimon Rahman , Jiang Liu , Ze Wang , Ximeng Sun , Jialian Wu , Xiaodong Yu , Yusheng Su , Vishal M. Patel , Zicheng Liu , Emad Barsoum

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Haodong Hong , Yanyuan Qiao , Sen Wang , Jiajun Liu , Qi Wu

Most existing sign language translation (SLT) datasets are limited in scale, lack multilingual coverage, and are costly to curate due to their reliance on expert annotation and controlled recording setup. Recently, Vision Language Models…

Computation and Language · Computer Science 2025-10-30 Shakib Yazdani , Yasser Hamidullah , Cristina España-Bonet , Josef van Genabith

We present RAVEN an adaptive AI agent framework designed for multimodal entity discovery and retrieval in large-scale video collections. Synthesizing information across visual, audio, and textual modalities, RAVEN autonomously processes…

Information Retrieval · Computer Science 2025-04-10 Kevin Dela Rosa

Large Language Model (LLM) agents are rapidly improving to handle increasingly complex web-based tasks. Most of these agents rely on general-purpose, proprietary models like GPT-4 and focus on designing better prompts to improve their…

Computation and Language · Computer Science 2024-12-06 Junhong Shen , Atishay Jain , Zedian Xiao , Ishan Amlekar , Mouad Hadji , Aaron Podolny , Ameet Talwalkar

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

Utilizing Graphic User Interface (GUI) for human-computer interaction is essential for accessing a wide range of digital tools. Recent advancements in Vision Language Models (VLMs) highlight the compelling potential to develop versatile…

Artificial Intelligence · Computer Science 2025-06-02 Wentong Chen , Junbo Cui , Jinyi Hu , Yujia Qin , Junjie Fang , Yue Zhao , Chongyi Wang , Jun Liu , Guirong Chen , Yupeng Huo , Yuan Yao , Yankai Lin , Zhiyuan Liu , Maosong Sun

The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which will help improve traffic safety. However, analyzing footage…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ruixuan Zhang , Beichen Wang , Juexiao Zhang , Zilin Bian , Chen Feng , Kaan Ozbay

Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the way to expedite this…

Artificial Intelligence · Computer Science 2025-09-26 Yidan Zhang , Mutian Xu , Yiming Hao , Kun Zhou , Jiahao Chang , Xiaoqiang Liu , Pengfei Wan , Hongbo Fu , Xiaoguang Han
‹ Prev 1 4 5 6 7 8 10 Next ›