English
Related papers

Related papers: STSA: Spatial-Temporal Semantic Alignment for Visu…

200 papers

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track. Existing state-of-the-art V2C models…

Computation and Language · Computer Science 2024-07-03 Gaoxiang Cong , Yuankai Qi , Liang Li , Amin Beheshti , Zhedong Zhang , Anton van den Hengel , Ming-Hsuan Yang , Chenggang Yan , Qingming Huang

It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results.…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Xiao Lu , Yihong Cao , Sheng Liu , Chengjiang Long , Zipei Chen , Xuanyu Zhou , Yimin Yang , Chunxia Xiao

Semantic segmentation is an important task for intelligent vehicles to understand the environment. Current deep learning methods require large amounts of labeled data for training. Manual annotation is expensive, while simulators can…

Computer Vision and Pattern Recognition · Computer Science 2022-08-24 Weihao Yan , Yeqiang Qian , Chunxiang Wang , Ming Yang

Multimodal aspect-based sentiment analysis (MABSA) aims to understand opinions in a granular manner, advancing human-computer interaction and other fields. Traditionally, MABSA methods use a joint prediction approach to identify aspects and…

Computation and Language · Computer Science 2024-06-14 Shezheng Song , Shasha Li , Shan Zhao , Chengyu Wang , Xiaopeng Li , Jie Yu , Qian Wan , Jun Ma , Tianwei Yan , Wentao Ma , Xiaoguang Mao

Simultaneous Localization and Mapping (SLAM) is one of the most essential techniques in many real-world robotic applications. The assumption of static environments is common in most SLAM algorithms, which however, is not the case for most…

Robotics · Computer Science 2022-05-17 Han Wang , Jing Ying Ko , Lihua Xie

In semi-supervised semantic segmentation (SSSS), data augmentation plays a crucial role in the weak-to-strong consistency regularization framework, as it enhances diversity and improves model generalization. Recent strong augmentation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Lingyan Ran , Yali Li , Tao Zhuo , Shizhou Zhang , Yanning Zhang

With information from multiple input modalities, sensor fusion-based algorithms usually out-perform their single-modality counterparts in robotics. Camera and LIDAR, with complementary semantic and depth information, are the typical choices…

Computer Vision and Pattern Recognition · Computer Science 2022-07-11 Akio Kodaira , Yiyang Zhou , Pengwei Zang , Wei Zhan , Masayoshi Tomizuka

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image…

Machine Learning · Computer Science 2025-03-18 Xulin Fan , Heting Gao , Ziyi Chen , Peng Chang , Mei Han , Mark Hasegawa-Johnson

Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding. Despite the progress achieved by prior architectures like CNN and Transformer models,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Yicheng Qiu , Keiji Yanai

Vision transformer has demonstrated great potential in abundant vision tasks. However, it also inevitably suffers from poor generalization capability when the distribution shift occurs in testing (i.e., out-of-distribution data). To…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Xin Li , Cuiling Lan , Guoqiang Wei , Zhibo Chen

Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, most existing approaches rely on a single static reference image to condition the entire…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhicheng Zhang , Lei Wang , Yu Zhang , Yongsheng Gao

Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video fine-tuning. However, fine-tuning heavy modules may…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yang Liu , Qianqian Xu , Peisong Wen , Siran Dai , Xilin Zhao , Qingming Huang

Multi-label image recognition is a fundamental task in computer vision. Recently, Vision-Language Models (VLMs) have made notable advancements in this area. However, previous methods fail to effectively leverage the rich knowledge in…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Hao Tan , Zichang Tan , Jun Li , Jun Wan , Zhen Lei , Stan Z. Li

The goal of sign language recognition (SLR) is to help those who are hard of hearing or deaf overcome the communication barrier. Most existing approaches can be typically divided into two lines, i.e., Skeleton-based and RGB-based methods,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Xiaolong Shen , Zhedong Zheng , Yi Yang

In many video restoration/translation tasks, image processing operations are na\"ively extended to the video domain by processing each frame independently, disregarding the temporal connection of the video frames. This disregard for the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Muhammad Kashif Ali , Dongjin Kim , Tae Hyun Kim

Satellite communications face severe bottlenecks in supporting high-fidelity synchronized audiovisual services, as conventional schemes struggle with cross-modal coherence under fluctuating channel conditions, limited bandwidth, and long…

Image and Video Processing · Electrical Eng. & Systems 2026-03-12 Fangyu Liu , Peiwen Jiang , Wenjin Wang , Chao-Kai Wen , Xiao Li , Shi Jin

Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal manipulation. Typically, existing methods mainly embed spatiotemporal knowledge into visual…

Robotics · Computer Science 2026-04-21 Chuanhao Ma , Hanyu Zhou , Shihan Peng , Yan Li , Tao Gu , Luxin Yan

The objective of this work is the effective extraction of spatial and dynamic features for Continuous Sign Language Recognition (CSLR). To accomplish this, we utilise a two-pathway SlowFast network, where each pathway operates at distinct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Junseok Ahn , Youngjoon Jang , Joon Son Chung

Video prediction yields future frames by employing the historical frames and has exhibited its great potential in many applications, e.g., meteorological prediction, and autonomous driving. Previous works often decode the ultimate…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Ping Li , Chenhan Zhang , Zheng Yang , Xianghua Xu , Mingli Song

Semantic segmentation from RGB cameras is essential to the perception of autonomous flying vehicles. The stability of predictions through the captured videos is paramount to their reliability and, by extension, to the trustworthiness of the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Cédric Vincent , Taehyoung Kim , Henri Meeß