English
Related papers

Related papers: Depth Registers Unlock W4A4 on SwiGLU: A Reader/Ge…

200 papers

The tremendous success of deep neural networks has motivated the need to better understand the fundamental properties of these networks, but many of the theoretical results proposed have only been for shallow networks. In this paper, we…

Data Structures and Algorithms · Computer Science 2020-02-20 Rajesh Jayaram , David P. Woodruff , Qiuyi Zhang

We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconductor manufacturing, we propose a three-stage synthesis pipeline incorporating structured…

Artificial Intelligence · Computer Science 2026-05-12 Ke Xu , Zhongyuan Lian

We investigate the influence of higher twist corrections to deep inelastic structure functions in the low-$Q^2$ and small-$x$ HERA region. We review the general features of the lowest-order QCD diagrams which contribute to twist-4 at…

High Energy Physics - Phenomenology · Physics 2008-11-26 J. Bartels , K. Golec-Biernat , K. Peters

Semantic segmentation is important for many real-world systems, e.g., autonomous vehicles, which predict the class of each pixel. Recently, deep networks achieved significant progress w.r.t. the mean Intersection-over Union (mIoU) with the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-12 Xiaofeng Liu , Yimeng Zhang , Xiongchang Liu , Song Bai , Site Li , Jane You

Linear layers hold most of a transformer's parameters. We replace each linear layer with one that stores $K$ out of $mn$ two-dimensional DCT coefficients per weight matrix and reconstructs the full matrix through an inverse DCT at every…

Performance · Computer Science 2026-04-10 Mohamed Amine Bergach

For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging than one-shot generation. We study PPO-trained agents using MiniGrid as core evaluation…

Machine Learning · Computer Science 2026-05-29 Youting Wang , Yuan Tang , Bowen Liu , Xuan Liu , Dingyan Shang

Zero-ablation -- replacing token activations with zero vectors -- is widely used to probe token function in vision transformers. Register zeroing in DINOv2+registers and DINOv3 produces large drops (up to $-36.6$\,pp classification,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Felipe Parodi , Jordan Matelsky , Melanie Segado

Inverse scattering problems, such as those in electromagnetic imaging using phaseless data (PD-ISPs), involve imaging objects using phaseless measurements of wave scattering. Such inverse problems can be highly non-linear and ill-posed…

Signal Processing · Electrical Eng. & Systems 2022-12-07 Samruddhi Deshmukh , Amartansh Dubey , Ross Murch

We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks. While the approximation power of neural networks is now relatively well understood, training very deep architectures…

Machine Learning · Computer Science 2026-04-03 Shijun Zhang , Zuowei Shen , Yuesheng Xu

Naively trained Deep Reinforcement Learning agents may fail to satisfy vital safety constraints. To avoid costly retraining, we may desire to repair a previously trained reinforcement learning agent to obviate unsafe behaviour. We devise a…

Machine Learning · Computer Science 2024-05-27 David Boetius , Stefan Leue

Over the past decade, Deep Learning (DL) has become an integral part of our daily lives. This surge in DL usage has heightened the need for developing reliable DL software systems. Given that fault localization is a critical task in…

Software Engineering · Computer Science 2024-11-14 Mohammad Mehdi Morovati , Amin Nikanjam , Foutse Khomh

The MXFP4 microscaling format, which partitions tensors into blocks of 32 elements sharing an E8M0 scaling factor, has emerged as a promising substrate for efficient LLM inference, backed by native hardware support on NVIDIA Blackwell…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Haokun Lin , Xinle Jia , Haobo Xu , Bingchen Yao , Xianglong Guo , Yichen Wu , Zhichao Lu , Ying Wei , Qingfu Zhang , Zhenan Sun

Low-rank adaptation~(LoRA) has recently gained much interest in fine-tuning foundation models. It effectively reduces the number of trainable parameters by incorporating low-rank matrices $A$ and $B$ to represent the weight change, i.e.,…

Machine Learning · Computer Science 2024-05-07 Ziqi Gao , Qichao Wang , Aochuan Chen , Zijing Liu , Bingzhe Wu , Liang Chen , Jia Li

Deep neural networks can struggle to learn continually in the face of non-stationarity. This phenomenon is known as loss of plasticity. In this paper, we identify underlying principles that lead to plastic algorithms. In particular, we…

Machine Learning · Computer Science 2024-10-29 Alex Lewandowski , Dale Schuurmans , Marlos C. Machado

Embryo fragmentation is a morphological indicator critical for evaluating developmental potential in In Vitro Fertilization (IVF). However, manual grading is subjective and inefficient, while existing deep learning solutions often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Ming-Jhe Lee

Spin qubits in semiconductor structures bring the promise of large-scale 2D integration, with the possibility to incorporate the control electronics on the same chip. In order to perform error correction on this platform, the characteristic…

Quantum Physics · Physics 2024-07-10 Bence Hetényi , James R. Wootton

Background: Accurate deformable image registration (DIR) is required for contour propagation and dose accumulation in MR-guided adaptive radiotherapy (MRgART). This study trained and evaluated a deep learning DIR method for domain invariant…

Structured width pruning of GLU-MLP layers, guided by the Maximum Absolute Weight (MAW) criterion, reveals a systematic dichotomy in how reducing the expansion ratio affects different model capabilities. While performance on tasks relying…

Computation and Language · Computer Science 2026-05-07 Pere Martra

Transformers trained on modular arithmetic exhibit sharp transitions between memorization, generalization, and collapse. We show that weight decay acts as a scalar empirical control parameter for these regimes, and introduce two cheap…

Machine Learning · Computer Science 2026-05-21 Lucky Verma

Analysis of longitudinal changes in imaging studies often involves both segmentation of structures of interest and registration of multiple timeframes. The accuracy of such analysis could benefit from a tailored framework that jointly…

Computer Vision and Pattern Recognition · Computer Science 2021-02-25 Bo Li , Wiro J. Niessen , Stefan Klein , M. Arfan Ikram , Meike W. Vernooij , Esther E. Bron