中文
相关论文

相关论文: VGSG: Vision-Guided Semantic-Group Network for Tex…

200 篇论文

Zero-shot learning uses semantic attributes to connect the search space of unseen objects. In recent years, although the deep convolutional network brings powerful visual modeling capabilities to the ZSL task, its visual features have…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Yang Hu , Guihua Wen , Adriane Chapman , Pei Yang , Mingnan Luo , Yingxue Xu , Dan Dai , Wendy Hall

Vision Transformers (ViTs) have achieved impressive results in large-scale image classification. However, when training from scratch on small datasets, there is still a significant performance gap between ViTs and Convolutional Neural…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Dongjing Shan , guiqiang chen

Parse graphs boost human pose estimation (HPE) by integrating context and hierarchies, yet prior work mostly focuses on single modality modeling, ignoring the potential of multimodal fusion. Notably, language offers rich HPE priors like…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Shibang Liu , Xuemei Xie , Guangming Shi

Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features…

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Rui Liu , Shuwei He , Yifan Hu , Haizhou Li

Visual-Semantic Embedding (VSE) is a prevalent approach in image-text retrieval by learning a joint embedding space between the image and language modalities where semantic similarities would be preserved. The triplet loss with…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Hong Xuan , Xi Chen

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Ning Cheng , You Li , Jing Gao , Bin Fang , Jinan Xu , Wenjuan Han

Conventional approaches to image-text retrieval mainly focus on indexing visual objects appearing in pictures but ignore the interactions between these objects. Such objects occurrences and interactions are equivalently useful and important…

计算机视觉与模式识别 · 计算机科学 2021-06-07 Manh-Duy Nguyen , Binh T. Nguyen , Cathal Gurrin

Task-oriented grasping (TOG) refers to the problem of predicting grasps on an object that enable subsequent manipulation tasks. To model the complex relationships between objects, tasks, and grasps, existing methods incorporate semantic…

机器人学 · 计算机科学 2023-09-21 Chao Tang , Dehao Huang , Wenqi Ge , Weiyu Liu , Hong Zhang

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing methods use a…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Yao Zhang , Zhuchenyang Liu , Yanlan He , Thomas Ploetz , Yu Xiao

We introduce Synthetic Visual Genome 2 (SVG2), a large-scale panoptic video scene graph dataset. SVG2 contains over 636K videos with 6.6M objects, 52.0M attributes, and 6.7M relations, providing an order-of-magnitude increase in scale and…

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Ming Dai , Lingfeng Yang , Yihao Xu , Zhenhua Feng , Wankou Yang

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

计算与语言 · 计算机科学 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Incorporating knowledge graph as side information has become a new trend in recommendation systems. Recent studies regard items as entities of a knowledge graph and leverage graph neural networks to assist item encoding, yet by considering…

信息检索 · 计算机科学 2022-11-22 Lingyun Lu , Bang Wang , Zizhuo Zhang , Shenghao Liu , Han Xu

In recent years, developing AI for robotics has raised much attention. The interaction of vision and language of robots is particularly difficult. We consider that giving robots an understanding of visual semantics and language semantics…

机器人学 · 计算机科学 2021-05-26 Cheng Yu Tsai , Mu-Chun Su

Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Yating Liu , Yaowei Li , Zimo Liu , Wenming Yang , Yaowei Wang , Qingmin Liao

Temporal sentence grounding (TSG) is an important yet challenging task in multimedia information retrieval. Although previous TSG methods have achieved decent performance, they tend to capture the selection biases of frequently appeared…

计算机视觉与模式识别 · 计算机科学 2022-07-28 Daizong Liu , Xiaoye Qu , Wei Hu

The goal of Text-to-Image Person Retrieval (TIPR) is to retrieve specific person images according to the given textual descriptions. A primary challenge in this task is bridging the substantial representational gap between visual and…

计算与语言 · 计算机科学 2025-01-20 Delong Liu , Haiwen Li , Zhicheng Zhao , Yuan Dong

We propose a novel Connectionist Text Proposal Network (CTPN) that accurately localizes text lines in natural image. The CTPN detects a text line in a sequence of fine-scale text proposals directly in convolutional feature maps. We develop…

计算机视觉与模式识别 · 计算机科学 2016-10-02 Zhi Tian , Weilin Huang , Tong He , Pan He , Yu Qiao

Instructed Visual Segmentation (IVS) tasks require segmenting objects in images or videos based on natural language instructions. While recent multimodal large language models (MLLMs) have achieved strong performance on IVS, their inference…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Wenhui Zhu , Xiwen Chen , Zhipeng Wang , Shao Tang , Sayan Ghosh , Xuanzhao Dong , Rajat Koner , Yalin Wang
‹ 上一页 1 8 9 10 下一页 ›