English
Related papers

Related papers: SynMotion: Semantic-Visual Adaptation for Motion C…

200 papers

Our objective is video retrieval based on natural language queries. In addition, we consider the analogous problem of retrieving sentences or generating descriptions given an input video. Recent work has addressed the problem by embedding…

Computer Vision and Pattern Recognition · Computer Science 2016-08-09 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä , Naokazu Yokoya

Motion is an important signal for agents in dynamic environments, but learning to represent motion from unlabeled video is a difficult and underconstrained problem. We propose a model of motion based on elementary group properties of…

Computer Vision and Pattern Recognition · Computer Science 2018-02-27 Andrew Jaegle , Stephen Phillips , Daphne Ippolito , Kostas Daniilidis

Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yao-Hung Hubert Tsai , Liang-Kang Huang , Ruslan Salakhutdinov

In cross-domain retrieval, a model is required to identify images from the same semantic category across two visual domains. For instance, given a sketch of an object, a model needs to retrieve a real image of it from an online store's…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Samarth Mishra , Carlos D. Castillo , Hongcheng Wang , Kate Saenko , Venkatesh Saligrama

In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking,…

Computer Vision and Pattern Recognition · Computer Science 2021-03-24 Sijie Song , Xudong Lin , Jiaying Liu , Zongming Guo , Shih-Fu Chang

Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has received growing research interest. Existing methods often…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Shang Chai , Zihang Lin , Min Zhou , Xubin Li , Liansheng Zhuang , Houqiang Li

Styled motion in-betweening is crucial for computer animation and gaming. However, existing methods typically encode motion styles by modeling whole-body motions, often overlooking the representation of individual body parts. This…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Minyue Dai , Ke Fan , Bin Ji , Haoran Xu , Haoyu Zhao , Junting Dong , Jingbo Wang , Bo Dai

We present a framework that allows users to incorporate the semantics of their domain knowledge for topic model refinement while remaining model-agnostic. Our approach enables users to (1) understand the semantic space of the model, (2)…

Human-Computer Interaction · Computer Science 2019-08-02 Mennatallah El-Assady , Rebecca Kehlbeck , Christopher Collins , Daniel Keim , Oliver Deussen

We introduce Programmatic Motion Concepts, a hierarchical motion representation for human actions that captures both low-level motion and high-level description as motion concepts. This representation enables human motion description,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Sumith Kulal , Jiayuan Mao , Alex Aiken , Jiajun Wu

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuanhao Cai , He Zhang , Xi Chen , Jinbo Xing , Yiwei Hu , Yuqian Zhou , Kai Zhang , Zhifei Zhang , Soo Ye Kim , Tianyu Wang , Yulun Zhang , Xiaokang Yang , Zhe Lin , Alan Yuille

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Miaowei Wang , Qingxuan Yan , Zhi Cao , Yayuan Li , Oisin Mac Aodha , Jason J. Corso , Amir Vaxman

Diverse and extensive work has recently been conducted on text-conditioned human motion generation. However, progress in the reverse direction, motion captioning, has seen less comparable advancement. In this paper, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Karim Radouane , Julien Lagarde , Sylvie Ranwez , Andon Tchechmedjiev

Unsupervised Domain Adaptation has been an efficient approach to transferring the semantic segmentation model across data distributions. Meanwhile, the recent Open-vocabulary Semantic Scene understanding based on large-scale vision language…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Thanh-Dat Truong , Utsav Prabhu , Dongyi Wang , Bhiksha Raj , Susan Gauch , Jeyamkondan Subbiah , Khoa Luu

Existing multi-object tracking algorithms typically fail to adequately address the issues in low-quality videos, resulting in a significant decline in tracking performance when image quality deteriorates in real-world scenarios. This…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jun Du

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence…

Computer Vision and Pattern Recognition · Computer Science 2019-11-01 Yitian Yuan , Lin Ma , Jingwen Wang , Wei Liu , Wenwu Zhu

Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Peng Jin , Yang Wu , Yanbo Fan , Zhongqian Sun , Yang Wei , Li Yuan

Semantic communication aims to transmit information most relevant to a task rather than raw data, offering significant gains in communication efficiency for applications such as telepresence, augmented reality, and remote sensing. Recent…

Machine Learning · Computer Science 2025-12-18 Matin Mortaheb , Erciyes Karakaya , Sennur Ulukus

Personalizing generative text-to-image models has seen remarkable progress, but extending this personalization to text-to-video models presents unique challenges. Unlike static concepts, personalizing text-to-video models has the potential…

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jianrong Zhang , Hehe Fan , Yi Yang

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Xiaoqi Wang , Yi Wang , Lap-Pui Chau
‹ Prev 1 8 9 10 Next ›