English
Related papers

Related papers: Bridging the Pose-Semantic Gap: A Cascade Framewor…

200 papers

Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and semantic features in decoder, which is crucial for generating…

Computation and Language · Computer Science 2023-05-24 Tian-Hao Zhang , Hai-Bo Qin , Zhi-Hao Lai , Song-Lu Chen , Qi Liu , Feng Chen , Xinyuan Qian , Xu-Cheng Yin

Human keypoint detection from a single image is very challenging due to occlusion, blur, illumination and scale variance. In this paper, we address this problem from three aspects by devising an efficient network structure, proposing three…

Computer Vision and Pattern Recognition · Computer Science 2021-05-25 Jing Zhang , Zhe Chen , Dacheng Tao

Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on pure appearance matching. Due to the complexity of motion…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Siyuan Li , Lei Ke , Yung-Hsu Yang , Luigi Piccinelli , Mattia Segù , Martin Danelljan , Luc Van Gool

A unique aspect of human visual understanding is the ability to flexibly interpret abstract concepts: acquiring lifted rules explaining what they symbolize, grounding them across familiar and unfamiliar contexts, and making predictions or…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Joy Hsu , Jiayuan Mao , Joshua B. Tenenbaum , Noah D. Goodman , Jiajun Wu

Stance detection seeks to identify the viewpoints of individuals either in favor or against a given target or a controversial topic. Current advanced neural models for stance detection typically employ fully parametric softmax classifiers.…

Machine Learning · Computer Science 2024-06-21 Yinghan Cheng , Qi Zhang , Chongyang Shi , Liang Xiao , Shufeng Hao , Liang Hu

Semantic segmentation and 3D reconstruction are two fundamental tasks in remote sensing, typically treated as separate or loosely coupled tasks. Despite attempts to integrate them into a unified network, the constraints between the two…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Chen Chen , Liangjin Zhao , Yuanchun He , Yingxuan Long , Kaiqiang Chen , Zhirui Wang , Yanfeng Hu , Xian Sun

Temporal grounding is the task of locating a specific segment from an untrimmed video according to a query sentence. This task has achieved significant momentum in the computer vision community as it enables activity grounding beyond…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Juncheng Li , Siliang Tang , Linchao Zhu , Wenqiao Zhang , Yi Yang , Tat-Seng Chua , Fei Wu , Yueting Zhuang

Trajectory prediction aims to predict the movement trend of the agents like pedestrians, bikers, vehicles. It is helpful to analyze and understand human activities in crowded spaces and widely applied in many areas such as surveillance…

Computer Vision and Pattern Recognition · Computer Science 2022-02-18 Beihao Xia , Conghao Wong , Qinmu Peng , Wei Yuan , Xinge You

Automated analysis of mouse behaviours is crucial for many applications in neuroscience. However, quantifying mouse behaviours from videos or images remains a challenging problem, where pose estimation plays an important role in describing…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Feixiang Zhou , Zheheng Jiang , Zhihua Liu , Fang Chen , Long Chen , Lei Tong , Zhile Yang , Haikuan Wang , Minrui Fei , Ling Li , Huiyu Zhou

With the rapid advancement of image generation techniques, robust forgery detection has become increasingly imperative to ensure the trustworthiness of digital media. Recent research indicates that the learned semantic concepts of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ziye Wang , Minghang Yu , Chunyan Xu , Zhen Cui

Image-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the fine-grained cross-modal correspondence between image and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-14 Zhong Ji , Kexin Chen , Haoran Wang

Public spaces such as transport hubs, city centres, and event venues require timely and reliable detection of potentially violent behaviour to support public safety. While automated video analysis has made significant progress, practical…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ganen Sethupathy , Lalit Dumka , Jan Schagen

Knowledge-based conversational question answering (KBCQA) confronts persistent challenges in resolving coreference, modeling contextual dependencies, and executing complex logical reasoning. Existing approaches often suffer from…

Computation and Language · Computer Science 2026-05-27 Hao Wang , Jialun Zhong , Changcheng Wang , Zhujun Nie , Zheng Li , Shunyu Yao , Yanzeng Li , Xinchi Li

This paper proposes a strategy for the detection and triangulation of structural anomalies in solid media. The method revolves around the construction of sparse representations of the medium's dynamic response, obtained by learning…

Computer Vision and Pattern Recognition · Computer Science 2014-05-13 Jeffrey M. Druce , Jarvis D. Haupt , Stefano Gonella

The exponential growth of video content has created an urgent need for efficient multimodal moment retrieval systems. However, existing approaches face three critical challenges: (1) fixed-weight fusion strategies fail across cross modal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Toan Le Ngo Thanh , Phat Ha Huu , Tan Nguyen Dang Duy , Thong Nguyen Le Minh , Anh Nguyen Nhu Tinh

Community detection plays a pivotal role in uncovering closely connected subgraphs, aiding various real-world applications such as recommendation systems and anomaly detection. With the surge of rich information available for entities in…

Social and Information Networks · Computer Science 2024-11-05 Anran Zhang , Xingfen Wang , Yuhan Zhao

Retrieving unlabeled videos by textual queries, known as Ad-hoc Video Search (AVS), is a core theme in multimedia data management and retrieval. The success of AVS counts on cross-modal representation learning that encodes both query…

Computer Vision and Pattern Recognition · Computer Science 2020-11-25 Xirong Li , Fangming Zhou , Chaoxi Xu , Jiaqi Ji , Gang Yang

Semantic correspondence remains a challenging task for establishing correspondences between a pair of images with the same category or similar scenes due to the large intra-class appearance. In this paper, we introduce a novel problem…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Hailong Jin , Huiying Li

Panoptic Narrative Detection (PND) and Segmentation (PNS) are two challenging tasks that involve identifying and locating multiple targets in an image according to a long narrative description. In this paper, we propose a unified and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Haowei Wang , Jiayi Ji , Tianyu Guo , Yilong Yang , Yiyi Zhou , Xiaoshuai Sun , Rongrong Ji

Remote sensing (RS) image-text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image-text pairs, which hinder the learning of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Pengxiang Ouyang , Qing Ma , Zheng Wang , Cong Bai
‹ Prev 1 8 9 10 Next ›