中文
相关论文

相关论文: Cross-Modal Content Optimization for Steering Web …

200 篇论文

Humans can collaborate and complete tasks based on visual signals and instruction from the environment. Training such a robot is difficult especially due to the understanding of the instruction and the complicated environment. Previous…

人工智能 · 计算机科学 2023-05-12 Kairui Zhou

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLMs, the associated safety issues have become increasingly…

计算与语言 · 计算机科学 2025-03-19 Yongqi Li , Lu Yang , Jian Wang , Runyang You , Wenjie Li , Liqiang Nie

Multimodal Large Language Models (MLLMs) often exhibit significant modality preference, which is a tendency to favor one modality over another. Depending on the input, they may over-rely on linguistic priors relative to visual evidence, or…

机器学习 · 计算机科学 2026-02-16 Zihan Huang , Xintong Li , Rohan Surana , Tong Yu , Rui Wang , Julian McAuley , Jingbo Shang , Junda Wu

Unlike traditional vision-only models, vision language models (VLMs) offer an intuitive way to access visual content through language prompting by combining a large language model (LLM) with a vision encoder. However, both the LLM and the…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Paul Gavrikov , Jovita Lukasik , Steffen Jung , Robert Geirhos , M. Jehanzeb Mirza , Margret Keuper , Janis Keuper

We consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Georgios Georgakis , Karl Schmeckpeper , Karan Wanchoo , Soham Dan , Eleni Miltsakaki , Dan Roth , Kostas Daniilidis

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content…

密码学与安全 · 计算机科学 2026-05-07 Jie Zhang , Pura Peetathawatchai , Florian Tramèr , Avital Shafran

Learning from human involvement aims to incorporate the human subject to monitor and correct agent behavior errors. Although most interactive imitation learning methods focus on correcting the agent's action at the current state, they do…

机器学习 · 计算机科学 2025-10-17 Haoyuan Cai , Zhenghao Peng , Bolei Zhou

Designing reward functions for continuous-control robotics often leads to subtle misalignments or reward hacking, especially in complex tasks. Preference-based RL mitigates some of these pitfalls by learning rewards from comparative…

Large Vision-Language Models (LVLMs) unlock powerful multimodal reasoning but also expand the attack surface, particularly through adversarial inputs that conceal harmful goals in benign prompts. We propose SHIELD, a lightweight,…

计算与语言 · 计算机科学 2025-10-16 Juan Ren , Mark Dras , Usman Naseem

Large Vision-Language Models (LVLMs), such as GPT-4o and LLaVA, have recently witnessed remarkable advancements and are increasingly being deployed in real-world applications. However, inheriting the sensitivity of visual neural networks,…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Chaohu Liu , Tianyi Gui , Yu Liu , Linli Xu

Large language models (LLMs) have shown remarkable capabilities in dialogue generation and reasoning, yet their effectiveness in eliciting user-known but concealed information in open-ended conversations remains limited. In many interactive…

机器学习 · 计算机科学 2026-04-16 Tao Wang , Jingyao Lu , Xibo Wang , Haonan Huang , Su Yao , Zhiqiang Hu , Xingyan Chen , Enmao Diao

Vision-language Navigation (VLN) tasks require an agent to navigate step-by-step while perceiving the visual observations and comprehending a natural language instruction. Large data bias, which is caused by the disparity ratio between the…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Chong Liu , Fengda Zhu , Xiaojun Chang , Xiaodan Liang , Zongyuan Ge , Yi-Dong Shen

Model Predictive Control (MPC) is a widely adopted control paradigm that leverages predictive models to estimate future system states and optimize control inputs accordingly. However, while MPC excels in planning and control, it lacks the…

机器人学 · 计算机科学 2025-04-08 Jiaming Chen , Wentao Zhao , Ziyu Meng , Donghui Mao , Ran Song , Wei Pan , Wei Zhang

We address the problem of policy selection in contextual stochastic optimization (CSO), where covariates are available as contextual information and decisions must satisfy hard feasibility constraints. In many CSO settings, multiple…

机器学习 · 计算机科学 2026-05-29 Caio de Prospero Iglesias , Kimberly Villalobos Carballo , Dimitris Bertsimas

While personalized recommender systems excel at content discovery, they frequently expose users to undesirable or discomforting information, highlighting the critical need for user-centric filtering tools. Current methods leveraging Large…

信息检索 · 计算机科学 2026-04-21 Chi Zhang , Zhipeng Xu , Jiahao Liu , Dongsheng Li , Hansu Gu , Peng Zhang , Ning Gu , Tun Lu

As large language models (LLMs) continue to evolve, their potential use in automating cyberattacks becomes increasingly likely. With capabilities such as reconnaissance, exploitation, and command execution, LLMs could soon become integral…

密码学与安全 · 计算机科学 2024-10-22 Daniel Ayzenshteyn , Roy Weiss , Yisroel Mirsky

Recently, driven by advancements in Multimodal Large Language Models (MLLMs), Vision Language Action Models (VLAMs) are being proposed to achieve better performance in open-vocabulary scenarios for robotic manipulation tasks. Since…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Hao Cheng , Erjia Xiao , Yichi Wang , Chengyuan Yu , Mengshu Sun , Qiang Zhang , Jiahang Cao , Yijie Guo , Ning Liu , Kaidi Xu , Jize Zhang , Chao Shen , Philip Torr , Jindong Gu , Renjing Xu

Vision-and-Language Navigation (VLN) is a natural language grounding task where agents have to interpret natural language instructions in the context of visual scenes in a dynamic environment to achieve prescribed navigation goals.…

计算与语言 · 计算机科学 2019-06-03 Haoshuo Huang , Vihan Jain , Harsh Mehta , Jason Baldridge , Eugene Ie

The existing image manipulation localization (IML) models mainly relies on visual cues, but ignores the semantic logical relationships between content features. In fact, the content semantics conveyed by real images often conform to human…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Songlin Li , Zhiqing Guo , Yuanman Li , Zeyu Li , Yunfeng Diao , Gaobo Yang , Liejun Wang

Benefiting from the powerful capabilities of Large Language Models (LLMs), pre-trained visual encoder models connected to LLMs form Vision Language Models (VLMs). However, recent research shows that the visual modality in VLMs is highly…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Zhendong Liu , Yuanbi Nie , Yingshui Tan , Jiaheng Liu , Xiangyu Yue , Qiushi Cui , Chongjun Wang , Xiaoyong Zhu , Bo Zheng