中文
相关论文

相关论文: Tree-based Text-Vision BERT for Video Search in Ba…

200 篇论文

Safe autonomous agents and mobile robots need fast real time 3D perception, especially for vulnerable road users (VRUs) such as pedestrians. We introduce a new bird's eye view (BEV) encoding, which maps the full 3D LiDAR point cloud into a…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Mohammad Khoshkdahan , Alexey Vinel

Automatically describing video content with natural language has been attracting much attention in CV and NLP communities. Most existing methods predict one word at a time, and by feeding the last generated word back as input at the next…

计算机视觉与模式识别 · 计算机科学 2019-11-06 Huanhou Xiao , Jinglun Shi

In this paper, we first tackle the problem of pedestrian attribute recognition by video-based approach. The challenge mainly lies in spatial and temporal modeling and how to integrating them for effective and dynamic pedestrian…

计算机视觉与模式识别 · 计算机科学 2019-10-29 Zhiyuan Chen , Annan Li , Yunhong Wang

Recently, the Metaverse is becoming increasingly attractive, with millions of users accessing the many available virtual worlds. However, how do users find the one Metaverse which best fits their current interests? So far, the search…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Ali Abdari , Alex Falcon , Giuseppe Serra

Video semantic segmentation requires to utilize the complex temporal relations between frames of the video sequence. Previous works usually exploit accurate optical flow to leverage the temporal relations, which suffer much from heavy…

计算机视觉与模式识别 · 计算机科学 2021-09-14 Hao Wang , Weining Wang , Jing Liu

In the rapidly evolving field of e-commerce, the effectiveness of search re-ranking models is crucial for enhancing user experience and driving conversion rates. Despite significant advancements in feature representation and model…

信息检索 · 计算机科学 2024-08-13 Enqiang Xu , Xinhui Li , Zhigong Zhou , Jiahao Ji , Jinyuan Zhao , Dadong Miao , Songlin Wang , Lin Liu , Sulong Xu

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Yuming Qiao , Fanyi Wang , Jingwen Su , Yanhao Zhang , Yunjie Yu , Siyu Wu , Guo-Jun Qi

In this report, we present the Baidu-UTS submission to the EPIC-Kitchens Action Recognition Challenge in CVPR 2019. This is the winning solution to this challenge. In this task, the goal is to predict verbs, nouns, and actions from the…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Xiaohan Wang , Yu Wu , Linchao Zhu , Yi Yang

Audio-visual learning seeks to enhance the computer's multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and…

人机交互 · 计算机科学 2023-07-31 Zheng Zhang , Zheng Ning , Chenliang Xu , Yapeng Tian , Toby Jia-Jun Li

With rapidly evolving internet technologies and emerging tools, sports related videos generated online are increasing at an unprecedentedly fast pace. To automate sports video editing/highlight generation process, a key task is to precisely…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Xin Zhou , Le Kang , Zhiyu Cheng , Bo He , Jingyu Xin

Visual-Semantic Embedding (VSE) networks can help search engines better understand the meaning behind visual content and associate it with relevant textual information, leading to more accurate search results. VSE networks can be used in…

多媒体 · 计算机科学 2023-11-02 Yan Gong , Georgina Cosma

Image-text retrieval is a central problem for understanding the semantic relationship between vision and language, and serves as the basis for various visual and language tasks. Most previous works either simply learn coarse-grained…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Chong Liu , Yuqi Zhang , Hongsong Wang , Weihua Chen , Fan Wang , Yan Huang , Yi-Dong Shen , Liang Wang

This paper is interested in investigating whether human gaze signals can be leveraged to improve state-of-the-art search engine performance and how to incorporate this new input signal marked by human attention into existing neural…

信息检索 · 计算机科学 2022-07-06 Sibo Dong , Justin Goldstein , Grace Hui Yang

Neural networks have proved to be very robust at processing unstructured data like images, text, videos, and audio. However, it has been observed that their performance is not up to the mark in tabular data; hence tree-based models are…

机器学习 · 计算机科学 2022-04-25 Tushar Sarkar

The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural movement or deformation, reverse playback, and motionless…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Yuta Oshima , Masahiro Suzuki , Yutaka Matsuo , Hiroki Furuta

Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Jian Chen , JinZe Lv , Zi Long , XiangHua Fu

The rapid expansion of multimedia content has made accurately retrieving relevant videos from large collections increasingly challenging. Recent advancements in text-video retrieval have focused on cross-modal interactions, large-scale…

计算与语言 · 计算机科学 2024-10-17 Donghoon Han , Eunhwan Park , Gisang Lee , Adam Lee , Nojun Kwak

As bird's-eye-view (BEV) semantic segmentation is simple-to-visualize and easy-to-handle, it has been applied in autonomous driving to provide the surrounding information to downstream tasks. Inferring BEV semantic segmentation conditioned…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Naiyu Fang , Lemiao Qiu , Shuyou Zhang , Zili Wang , Kerui Hu , Kang Wang

The growing popularity of generative language models has amplified interest in interactive methods to guide model outputs. Prompt refinement is considered one of the most effective means to influence output among these methods. We identify…

As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zechen Bai , Tianjun Xiao , Tong He , Pichao Wang , Zheng Zhang , Thomas Brox , Mike Zheng Shou