中文
相关论文

相关论文: THEIA: Learning Complete Kleene Three-Valued Logic…

200 篇论文

Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where the student already has strong reasoning ability, naive continual distillation often yields…

计算与语言 · 计算机科学 2026-05-22 Zhanming Shen , Jiaqi Hu , Zeyu Qin , Hao Chen , Wentao Ye , Zenan Huang , Yihong Zhuang , Guoshan Lu , Junlin Zhou , Junbo Zhao

Training Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive. Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger…

机器学习 · 计算机科学 2025-02-20 Yifei Yang , Zouying Cao , Xinbei Ma , Yao Yao , Libo Qin , Zhi Chen , Hai Zhao

Truly reliable AI requires more than simply scaling up knowledge; it demands the ability to know what it knows and when it does not. Yet recent research shows that even the best LLMs misjudge their own competence in more than one in five…

计算与语言 · 计算机科学 2025-10-14 Sahil Kale , Devendra Singh Dhami

The intrinsic alignments (IA) of galaxies, a key contaminant in weak lensing analyses, arise from correlations in galaxy shapes driven by tidal interactions and galaxy formation processes. Accurate IA modeling is essential for robust…

宇宙学与河外天体物理 · 物理学 2025-12-03 Sneh Pandya , Yuanyuan Yang , Nicholas Van Alfen , Jonathan Blazek , Robin Walters

Three classes of algorithms to learn the structure of Bayesian networks from data are common in the literature: constraint-based algorithms, which use conditional independence tests to learn the dependence structure of the data; score-based…

统计方法学 · 统计学 2021-02-10 Marco Scutari , Catharina Elisabeth Graafland , José Manuel Gutiérrez

The discovery of advanced metallic alloys is hindered by vast composition spaces, competing property objectives, and real-world constraints on manufacturability. Here we introduce MATAI, a generalist machine learning framework for property…

We introduce a principled approach for unsupervised structure learning of deep neural networks. We propose a new interpretation for depth and inter-layer connectivity where conditional independencies in the input distribution are encoded…

机器学习 · 统计学 2018-10-18 Raanan Y. Rohekar , Shami Nisimov , Yaniv Gurwicz , Guy Koren , Gal Novik

We propose Self-Supervised Implicit Attention (SSIA), a new approach that adaptively guides deep neural network models to gain attention by exploiting the properties of the models themselves. SSIA is a novel attention mechanism that does…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Jinyi Wu , Xun Gong , Zhemin Zhang

Despite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which causes them to overlook information at certain positions. Whether these failures stem from an…

人工智能 · 计算机科学 2026-04-22 Meiru Zhang , Zaiqiao Meng , Nigel Collier

Radiographic grading of knee osteoarthritis (KOA) with the Kellgren-Lawrence (KL) system is limited by inter-reader variability and the opacity of current deep learning approaches, which predict KL grades directly from images without…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Azmul A. Irfan , Nur Ahmad Khatim , Alfan Alfian Irfan , Achmad Zaki , Erike A. Suwarsono , Mansur M. Arief

This paper introduces TEDI (Truthful, Expressive, and Dimension-Insensitive approach), a discretization-free algorithm to learn truthful and utility-maximizing mechanisms. Existing learning-based approaches often rely on discretization of…

计算机科学与博弈论 · 计算机科学 2025-07-01 Yunxuan Ma , Siqiang Wang , Zhijian Duan , Yukun Cheng , Xiaotie Deng

Neural networks trained with gradient-based methods exhibit a strong simplicity bias: they learn simpler statistical features of their data before moving to more complex features. Previous analyses of this phenomenon have largely focused on…

机器学习 · 统计学 2026-05-19 Fabiola Ricci , Claudia Merger , Sebastian Goldt

Artificial intelligence (AI) has evolved into an ecosystem of specialized "species," each with unique strengths. We analyze two: DeepSeek-V3, a 671-billion-parameter Mixture of Experts large language model (LLM) exemplifying scale-driven…

机器学习 · 计算机科学 2025-06-23 Joseph Geraci , Bessi Qorri , Christian Cumbaa , Mike Tsay , Paul Leonczyk , Luca Pani

Understanding whether deep neural networks are effectively optimized remains challenging, as training occurs in highly nonconvex landscapes and standard metrics provide limited visibility into layer-wise learning quality. This challenge is…

机器学习 · 计算机科学 2026-05-05 Arian Eamaz , Farhang Yeganegi , Mojtaba Soltanalian

We assess the robustness of the two highest rungs of the "cosmic distance ladder" for Type Ia supernovae and the determination of the Hubble-Lema\^itre constant. In this analysis, we hold fixed Rung 1 as the distance to the LMC determined…

宇宙学与河外天体物理 · 物理学 2020-10-22 Mario Hamuy , Régis Cartier , Carlos Contreras , Nicholas B. Suntzeff

A central question in the LLM debate is whether transformers can infer rules absent from training, or whether apparent generalisation reduces to similarity-based interpolation over observed examples. We test a strong interpolation-only…

机器学习 · 计算机科学 2026-03-19 Andy Gray

Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embedding space. However, comparing such spaces across models remains difficult: changes in…

Growing concerns over data privacy and security highlight the importance of machine unlearning--removing specific data influences from trained models without full retraining. Techniques like Membership Inference Attacks (MIAs) are widely…

机器学习 · 计算机科学 2025-06-09 Cheng-Long Wang , Qi Li , Zihang Xiang , Yinzhi Cao , Di Wang

The rapid integration of AI into education has prioritized capability over trustworthiness, creating significant risks. Real-world deployments reveal that even advanced models are insufficient without extensive architectural scaffolding to…

计算机与社会 · 计算机科学 2026-01-13 Abu Syed

For a three-cell constant cellular interfering network, a new property of alignment is identified, i.e., interference alignment (IA) solution obtained in an user-cooperation scenario can also be applied in a non-cooperation environment. By…

信息论 · 计算机科学 2012-03-02 Yanjun Ma , Jiandong Li , Rui Chen , Qin Liu