中文
相关论文

相关论文: Cross-view Semantic Alignment for Livestreaming Pr…

200 篇论文

Referring Image Segmentation (RIS) aims at segmenting the target object from an image referred by one given natural language expression. The diverse and flexible expressions as well as complex visual contents in the images raise the RIS…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Yang Jiao , Zequn Jie , Weixin Luo , Jingjing Chen , Yu-Gang Jiang , Xiaolin Wei , Lin Ma

Existing human recognition systems often rely on separate, specialized models for face and body analysis, limiting their effectiveness in real-world scenarios where pose, visibility, and context vary widely. This paper introduces SapiensID,…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Minchul Kim , Dingqiang Ye , Yiyang Su , Feng Liu , Xiaoming Liu

Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressive progress in processing diverse inputs, it remains unclear…

Recently, there has been increasing interest in multimodal applications that integrate text with other modalities, such as images, audio and video, to facilitate natural language interactions with multimodal AI systems. While applications…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Roger Ferrod , Luigi Di Caro , Dino Ienco

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Boyuan Sun , Bowen Yin , Yuanming Li , Xihan Wei , Qibin Hou

Live streaming platforms require real-time monitoring and reaction to social signals, utilizing partial and asynchronous evidence from video, text, and audio. We propose StreamSense, a streaming detector that couples a lightweight streaming…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Han Wang , Deyi Ji , Lanyun Zhu , Jiebo Luo , Roy Ka-Wei Lee

Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream…

信息检索 · 计算机科学 2026-01-13 Sayak Chakrabarty , Souradip Pal

Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, where decision…

机器学习 · 计算机科学 2026-04-20 Yide Ran , Jianwen Xie , Minghui Wang , Wenjin Zheng , Denghui Zhang , Chuan Li , Zhaozhuo Xu

Live streaming is becoming an increasingly popular trend of sales in E-commerce. The core of live-streaming sales is to encourage customers to purchase in an online broadcasting room. To enable customers to better understand a product…

信息检索 · 计算机科学 2021-09-16 Guohai Xu , Hehong Chen , Feng-Lin Li , Fu Sun , Yunzhou Shi , Zhixiong Zeng , Wei Zhou , Zhongzhou Zhao , Ji Zhang

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Shreya Rajpal , Michal Golovanevsky , Carsten Eickhoff

Recommender systems (RSs) are essential for e-commerce platforms to help meet the enormous needs of users. How to capture user interests and make accurate recommendations for users in heterogeneous e-commerce scenarios is still a continuous…

信息检索 · 计算机科学 2020-12-17 Yuting Chen , Yanshi Wang , Yabo Ni , An-Xiang Zeng , Lanfen Lin

Large vision-language contrastive models (VLCMs), such as CLIP, have become foundational, demonstrating remarkable success across a variety of downstream tasks. Despite their advantages, these models, akin to other foundational systems,…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Haocheng Dai , Sarang Joshi

E-commerce provides rich multimodal data that is barely leveraged in practice. One aspect of this data is a category tree that is being used in search and recommendation. However, in practice, during a user's session there is often a…

This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text Semantic Generation (TSG), Infrared Feature Embedding (IFE), and High-level Semantic…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Xiaomei Yang , Xizhan Gao , Sijie Niu , Fa Zhu , Guang Feng , Xiaofeng Qu , David Camacho

In streaming Reinforcement Learning (RL), transitions are observed and discarded immediately after a single update. While this minimizes resource usage for on-device applications, it makes agents notoriously sample-inefficient, since…

机器学习 · 计算机科学 2026-02-11 Nilaksh , Antoine Clavaud , Mathieu Reymond , François Rivest , Sarath Chandar

Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones. To address…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Chaeyun Kim , Seunghoon Yi , Yejin Kim , Yohan Jo , Joonseok Lee

Retrieving clothes which are worn in social media videos (Instagram, TikTok) is the latest frontier of e-fashion, referred to as "video-to-shop" in the computer vision literature. In this paper we present MovingFashion, the first publicly…

计算机视觉与模式识别 · 计算机科学 2021-10-15 Marco Godi , Christian Joppi , Geri Skenderi , Marco Cristani

Real-world reasoning often requires combining information across modalities, connecting textual context with visual cues in a multi-hop process. Yet, most multimodal benchmarks fail to capture this ability: they typically rely on single…

机器学习 · 计算机科学 2026-04-03 Junyoung Sung , Seungwoo Lyu , Minjun Kim , Sumin An , Arsha Nagrani , Paul Hongsuck Seo

Existing two-stream models, such as CLIP, encode images and text through independent representations, showing good performance while ensuring retrieval speed, have attracted attention from industry and academia. However, the single…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Wanqing Cui , Rui Cheng , Jiafeng Guo , Xueqi Cheng

Cross-lingual cross-modal retrieval has garnered increasing attention recently, which aims to achieve the alignment between vision and target language (V-T) without using any annotated V-T data pairs. Current methods employ machine…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Yabing Wang , Fan Wang , Jianfeng Dong , Hao Luo