English
Related papers

Related papers: SKALD: Learning-Based Shot Assembly for Coherent M…

200 papers

We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent past but has primarily been limited to visual modality and to…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Kranti Kumar Parida , Neeraj Matiyali , Tanaya Guha , Gaurav Sharma

Traditional indoor scene synthesis methods often take a two-step approach: object selection and object arrangement. Current state-of-the-art object selection approaches are based on convolutional neural networks (CNNs) and can produce…

Graphics · Computer Science 2020-03-17 Yu He , Yun Cai , Yuan-Chen Guo , Zheng-Ning Liu , Shao-Kui Zhang , Song-Hai Zhang , Hong-Bo Fu , Sheng-Yong Chen

Prior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practical applications such as video alignment have strong demand…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Minghao Chen , Fangyun Wei , Chong Li , Deng Cai

Sparse sampling schemes have the potential to dramatically reduce image acquisition time while simultaneously reducing radiation damage to samples. However, for a sparse sampling scheme to be useful it is important that we are able to…

Computer Vision and Pattern Recognition · Computer Science 2017-03-16 G. M. Dilshan P. Godaliyadda , Dong Hye Ye , Michael D. Uchic , Michael A. Groeber , Gregery T. Buzzard , Charles A. Bouman

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Amit Eliav , Sharon Gannot

Advertisement video editing aims to automatically edit advertising videos into shorter videos while retaining coherent content and crucial information conveyed by advertisers. It mainly contains two stages: video segmentation and segment…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Yolo Yunlong Tang , Siting Xu , Teng Wang , Qin Lin , Qinglin Lu , Feng Zheng

Multimodal Large Language Models (MLLMs) adapt to visual tasks via in-context learning (ICL), which relies heavily on demonstration quality. The dominant demonstration selection strategy is unsupervised k-Nearest Neighbor (kNN) search.…

Machine Learning · Computer Science 2026-03-31 Eugene Lee , Yu-Chi Lin , Jiajie Diao

Recent advancements in video large multimodal models (LMMs) have significantly improved their video understanding and reasoning capabilities. However, their performance drops on out-of-distribution (OOD) tasks that are underrepresented in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Kangsan Kim , Geon Park , Youngwan Lee , Woongyeong Yeo , Sung Ju Hwang

Multimodal learning enhances the performance of various machine learning tasks by leveraging complementary information across different modalities. However, existing methods often learn multimodal representations that retain substantial…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Tong Zhang , Shu Shen , C. L. Philip Chen

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Davide Berghi , Philip J. B. Jackson

Humans can watch a continuous video stream and effortlessly perform continual acquisition and transfer of new knowledge with minimal supervision yet retaining previously learnt experiences. In contrast, existing continual learning (CL)…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Jay Zhangjie Wu , David Junhao Zhang , Wynne Hsu , Mengmi Zhang , Mike Zheng Shou

Continual Learning (CL) aims to enable models to sequentially learn multiple tasks without forgetting previous knowledge. Recent studies have shown that optimizing towards flatter loss minima can improve model generalization. However,…

Machine Learning · Computer Science 2026-01-13 Yanan Chen , Tieliang Gong , Yunjiao Zhang , Wen Wen

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset's action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jiahui Zhang , Yurui Chen , Yueming Xu , Ze Huang , Yanpeng Zhou , Yu-Jie Yuan , Xinyue Cai , Guowei Huang , Xingyue Quan , Hang Xu , Li Zhang

Text-driven motion generation has advanced significantly with the rise of denoising diffusion models. However, previous methods often oversimplify representations for the skeletal joints, temporal frames, and textual words, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Seokhyeon Hong , Chaelin Kim , Serin Yoon , Junghyun Nam , Sihun Cha , Junyong Noh

Text-based Visual Question Answering (TextVQA) aims at answering questions about the text in images. Most works in this field focus on designing network structures or pre-training tasks. All these methods list the OCR texts in reading order…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Chengyang Fang , Jiangnan Li , Liang Li , Can Ma , Dayong Hu

Interpreting the internal reasoning of vision-language models is essential for deploying AI in safety-critical domains. Concept-based explainability provides a human-aligned lens by representing a model's behavior through semantically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ehud Gordon , Meir Yossef Levi , Guy Gilboa

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Yash Kant , Ziyi Wu , Michael Vasilkovsky , Guocheng Qian , Jian Ren , Riza Alp Guler , Bernard Ghanem , Sergey Tulyakov , Igor Gilitschenski , Aliaksandr Siarohin

Precise, object-aware control over visual content is essential for advanced image editing and compositional generation. Yet, most existing approaches operate on entire images holistically, limiting the ability to isolate and manipulate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Fangyi Chen , Yaojie Shen , Lu Xu , Ye Yuan , Shu Zhang , Yulei Niu , Longyin Wen

Continual semantic segmentation (CSS) based on incremental learning (IL) is a great endeavour in developing human-like segmentation models. However, current CSS approaches encounter challenges in the trade-off between preserving old…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Bo Yuan , Danpei Zhao , Zhenwei Shi

We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multiple image patches often correspond to single words, we…