中文
相关论文

相关论文: Structural Rigidity and the 57-Token Predictive Wi…

200 篇论文

An interactive instruction following task has been proposed as a benchmark for learning to map natural language instructions and first-person vision into sequences of actions to interact with objects in 3D environments. We found that an…

人工智能 · 计算机科学 2022-11-16 Kazutoshi Shinoda , Yuki Takezawa , Masahiro Suzuki , Yusuke Iwasawa , Yutaka Matsuo

Hub importance scores in multilayer networks persist more strongly between functionally similar layers than dissimilar ones. We call this the Functional Proximity Law and test it across 23 pre-registered experiments: 13 canonical domains…

社会与信息网络 · 计算机科学 2026-05-26 Vladi Ivanov

Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Yuan Tian , Bing Hu , Fang Wu , Xiaomin Li , Binghang Lu , Neil Zhenqiang Gong

Existing approaches for predictive process monitoring are sub-symbolic, meaning that they learn correlations between descriptive features and a target feature fully based on data, e.g., predicting the surgical needs of a patient based on…

人工智能 · 计算机科学 2026-04-01 Fabrizio De Santis , Gyunam Park , Wil M. P. van der Aalst , Francesco Zanichelli

Artificial intelligence systems are increasingly deployed in domains that shape human behaviour, institutional decision-making, and societal outcomes. Existing responsible AI and governance efforts provide important normative principles but…

人工智能 · 计算机科学 2025-12-19 Otman A. Basir

Smart grid infrastructure needs improved resilience and preventive maintenance through more accurate predictions. Current methodologies lack accurate representation of spatio-temporal-causal interdependencies and class imbalance in failure…

系统与控制 · 电气工程与系统科学 2026-01-07 Anh Le , Phat K. Huynh , Om P. Yadav , Harun Pirim , Chau Le , Trung Q. Le

As LLM-based systems increasingly operate as agents embedded within human social and technical systems, alignment can no longer be treated as a property of an isolated model, but must be understood in relation to the environments in which…

While large transformer models excel in predictive performance, their lack of interpretability restricts their usefulness in high-stakes domains. To remedy this, we propose the Generalized Induction-Head Model (GIM), an interpretable model…

计算与语言 · 计算机科学 2025-10-31 Eunji Kim , Sriya Mantena , Weiwei Yang , Chandan Singh , Sungroh Yoon , Jianfeng Gao

Regression and Bayesian accounts of in-context learning (ICL) explain how demonstrations can induce predictors, while mechanistic analyses often identify compact activation directions that steer prompted behavior. However, it remains…

机器学习 · 计算机科学 2026-05-20 Wei Tang , Xinyan Jiang , Fakhri Karray , Lijie Hu

Hallucinations in deployed language models can have real consequences for downstream decisions in domains such as healthcare, legal, and financial services. In production, detection has to run on what the deployed system can see: the query,…

人工智能 · 计算机科学 2026-05-11 Javier Marín

AI-native wireless receivers based on deep learning exhibit remarkable performance under stationary channel conditions, yet their resilience to distributional shifts remains poorly characterized by conventional metrics such as bit error…

信息论 · 计算机科学 2026-05-25 Christo Kurisummoottil Thomas , Emilio Calvanese Strinati

Evaluating LLM forecasting capabilities is constrained by a fundamental tension: prospective evaluation offers methodological rigor but prohibitive latency, while retrospective forecasting (RF) -- evaluating on already-resolved events --…

计算与语言 · 计算机科学 2026-01-21 Zehan Li , Yuxuan Wang , Ali El Lahib , Ying-Jieh Xia , Xinyu Pi

Binary vulnerability analysis is increasingly performed by LLM-based agents in an iterative, multi-pass manner, with the model as the core decision-maker. However, how such systems organize exploration over hundreds of reasoning steps…

人工智能 · 计算机科学 2026-03-20 Qiang Li , XiangRui Zhang , Haining Wang

Large Vision-Language Models (LVLMs) enable sophisticated reasoning over images and videos, yet their inference is hindered by a systemic efficiency barrier known as visual token dominance. This overhead is driven by a multi-regime…

计算与语言 · 计算机科学 2026-04-15 Jun Zhang , Yicheng Ji , Feiyang Ren , Yihang Li , Bowen Zeng , Zonghao Chen , Ke Chen , Lidan Shou , Gang Chen , Huan Li

When researchers iteratively refine ideas with large language models, do the models preserve fidelity to the original objective? We introduce DriftBench, a benchmark for evaluating constraint adherence in multi-turn LLM-assisted scientific…

计算与语言 · 计算机科学 2026-05-05 Garvin Kruthof

Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliability, subgroup equity, threshold sensitivity, or operational feasibility. We propose the…

机器学习 · 计算机科学 2026-05-14 Rohith Reddy Bellibatlu

Structural condition identification based on monitoring data is important for automatic civil infrastructure asset management. Nevertheless, the monitoring data is almost always insufficient, because the real-time monitoring data of a…

计算工程、金融与科学 · 计算机科学 2023-07-31 Nengxin Bao , Tong Zhang , Ruizhi Huang , Suryakanta Biswal , Jingyong Su , Ying Wang

Despite extensive investment in artificial intelligence, 95% of enterprises report no measurable profit impact from AI deployments (MIT, 2025). In this theoretical paper, we argue that this gap reflects paradigmatic lock-in that channels AI…

计算机与社会 · 计算机科学 2025-09-15 Diana A. Wolfe , Alice Choe , Fergus Kidd

We introduce Refusal Steering, an inference-time method to exercise fine-grained control over Large Language Models refusal behaviour on politically sensitive topics without retraining. We replace fragile pattern-based refusal detection…

计算与语言 · 计算机科学 2026-02-25 Iker García-Ferrero , David Montero , Roman Orus

The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct dimensions: confidence calibration (T1-CC), epistemic vigilance (T2-EV), knowledge…

人工智能 · 计算机科学 2026-05-12 Rafael C. T. Oliveira