English
Related papers

Related papers: UniT3D: A Unified Transformer for 3D Dense Caption…

200 papers

Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly trained on 2D image data and therefore often fail to capture 3D spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Byungwoo Jeon , Dongyoung Kim , Huiwon Jang , Insoo Kim , Jinwoo Shin

Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ayush Jain , Nikolaos Gkanatsios , Ishita Mediratta , Katerina Fragkiadaki

Multi-task learning has recently emerged as a promising solution for a comprehensive understanding of complex scenes. In addition to being memory-efficient, multi-task models, when appropriately designed, can facilitate the exchange of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Ivan Lopes , Tuan-Hung Vu , Raoul de Charette

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Visual grounding is a promising path toward more robust and accurate Natural Language Processing (NLP) models. Many multimodal extensions of BERT (e.g., VideoBERT, LXMERT, VL-BERT) allow a joint modeling of texts and images that lead to…

Computation and Language · Computer Science 2021-03-26 Damien Sileo

In this study, we explore the efficacy of advanced pre-trained architectures, such as Vision Transformers (ViT), ConvNeXt, and Swin Transformers in enhancing Federated Domain Generalization. These architectures capture global contextual…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Avi Deb Raha , Apurba Adhikary , Mrityunjoy Gain , Yu Qiao , Choong Seon Hong

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Chenjie Cao , Jingkai Zhou , Shikai Li , Jingyun Liang , Chaohui Yu , Fan Wang , Xiangyang Xue , Yanwei Fu

Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Cheng Chen , Yudong Zhu , Zhenshan Tan , Qingrong Cheng , Xin Jiang , Qun Liu , Xiaodong Gu

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys

Transformer-based end-to-end speech recognition has achieved great success. However, the large footprint and computational overhead make it difficult to deploy these models in some real-world applications. Model compression techniques can…

Computation and Language · Computer Science 2023-03-15 Yifan Peng , Jaesong Lee , Shinji Watanabe

Thanks to its precise spatial referencing, 3D point cloud visual grounding is essential for deep understanding and dynamic interaction in 3D environments, encompassing 3D Referring Expression Comprehension (3DREC) and Segmentation (3DRES).…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Haojia Lin , Yongdong Luo , Xiawu Zheng , Lijiang Li , Fei Chao , Taisong Jin , Donghao Luo , Yan Wang , Liujuan Cao , Rongrong Ji

We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded…

Computation and Language · Computer Science 2017-07-10 Desmond Elliott , Ákos Kádár

In the current state of 3D object detection research, the severe scarcity of annotated 3D data, substantial disparities across different data modalities, and the absence of a unified architecture, have impeded the progress towards the goal…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Zhenyu Wang , Yali Li , Taichi Liu , Hengshuang Zhao , Shengjin Wang

Multimodal large language models (MLLMs) have made significant progress in vision-language understanding, yet effectively aligning different modalities remains a fundamental challenge. We present a framework that unifies multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Wanpeng Zhang , Yicheng Feng , Hao Luo , Yijiang Li , Zihao Yue , Sipeng Zheng , Zongqing Lu

Depth prediction is a critical problem in robotics applications especially autonomous driving. Generally, depth prediction based on binocular stereo matching and fusion of monocular image and laser point cloud are two mainstream methods.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Guancheng Chen , Junli Lin , Huabiao Qin

Multiple modalities can provide more valuable information than single one by describing the same contents in various ways. Hence, it is highly expected to learn effective joint representation by fusing the features of different modalities.…

Computer Vision and Pattern Recognition · Computer Science 2018-10-09 Di Hu , Feiping Nie , Xuelong Li

Training models to apply common-sense linguistic knowledge and visual concepts from 2D images to 3D scene understanding is a promising direction that researchers have only recently started to explore. However, it still remains understudied…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Alexandros Delitzas , Maria Parelli , Nikolas Hars , Georgios Vlassis , Sotirios Anagnostidis , Gregor Bachmann , Thomas Hofmann

Pre-trained language models have been shown to improve performance in many natural language tasks substantially. Although the early focus of such models was single language pre-training, recent advances have resulted in cross-lingual and…

Computation and Language · Computer Science 2021-04-22 Ozan Caglayan , Menekse Kuyu , Mustafa Sercan Amac , Pranava Madhyastha , Erkut Erdem , Aykut Erdem , Lucia Specia

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introduces an implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 An-Chieh Cheng , Yang Fu , Yatai Ji , Ligeng Zhu , Guanqi Zhan , Zhuoyang Zhang , Zhaojing Yang , Song Han , Yao Lu , Pavlo Molchanov , Vidya Nariyambut Murali , Jan Kautz , Xiaolong Wang , Hongxu Yin , Sifei Liu

Visual Question Answering (VQA) and Image Captioning (CAP), which are among the most popular vision-language tasks, have analogous scene-text versions that require reasoning from the text in the image. Despite their obvious resemblance, the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Roy Ganz , Oren Nuriel , Aviad Aberdam , Yair Kittenplon , Shai Mazor , Ron Litman
‹ Prev 1 8 9 10 Next ›