English
Related papers

Related papers: No Free Swap: Protocol-Dependent Layer Redundancy …

200 papers

For years the model performance in machine learning obeyed a power-law relationship with the model size. For the consideration of parameter efficiency, recent studies focus on increasing model depth rather than width to achieve better…

Computation and Language · Computer Science 2023-05-11 Ye Lin , Shuhan Zhou , Yanyang Li , Anxiang Ma , Tong Xiao , Jingbo Zhu

Deploying transformer models in practice is challenging due to their inference cost, which scales quadratically with input sequence length. To address this, we present a novel Learned Token Pruning (LTP) method which adaptively removes…

Computation and Language · Computer Science 2022-06-06 Sehoon Kim , Sheng Shen , David Thorsley , Amir Gholami , Woosuk Kwon , Joseph Hassoun , Kurt Keutzer

The next generation wireless systems will face stringent new requirements, including ultra-low latency, high data rates and enhanced reliability. Large Intelligent Surfaces, is one proposed solution that has the potential to solve these…

Signal Processing · Electrical Eng. & Systems 2025-10-16 Lina Tinnerberg , Dumitra Iancu , Ove Edfors , Liang Liu , Juan Vidal Alegría

Consider a MIMO interference channel whereby each transmitter and receiver are equipped with multiple antennas. The basic problem is to design optimal linear transceivers (or beamformers) that can maximize system throughput. The recent work…

Information Theory · Computer Science 2010-09-20 Meisam Razaviyayn , Maziar Sanjabi , Zhi-Quan Luo

Overparameterized transformer networks have obtained state of the art results in various natural language processing tasks, such as machine translation, language modeling, and question answering. These models contain hundreds of millions of…

Machine Learning · Computer Science 2019-09-26 Angela Fan , Edouard Grave , Armand Joulin

This work considers load-balance control among the relays under the secure transmission protocol via relay cooperation in two-hop wireless networks without the information of both eavesdropper channels and locations. The available two-hop…

Information Theory · Computer Science 2013-01-01 Yulong Shen , Xiaohong Jiang , Jianfeng Ma

In response to the development of recent efficient dense layers, this paper shows that something as simple as replacing linear components in pointwise convolutions with structured linear decompositions also produces substantial gains in the…

Machine Learning · Statistics 2019-06-04 Gavin Gray , Elliot J. Crowley , Amos Storkey

An infinite hierarchy of layering transitions exists for model polymers in solution under poor solvent or low temperatures and near an attractive surface. A flat histogram stochastic growth algorithm known as FlatPERM has been used on a…

Statistical Mechanics · Physics 2009-11-10 J. Krawczyk , A. L. Owczarek , T. Prellberg , A. Rechnitzer

Learning, prediction, and compression are intimately connected: a model that accurately predicts the next symbol in a sequence can be coupled with a source coder to compress that sequence near its information-theoretic limit. When tokenized…

Information Theory · Computer Science 2026-05-05 Vishnu Teja Kunde , Jean-Francois Chamberland , Krishna R. Narayanan , Jamison Ebert

Although large language models (LLMs) have achieved remarkable success across various domains, their considerable scale necessitates substantial computational resources, posing significant challenges for deployment in resource-constrained…

Machine Learning · Computer Science 2024-11-26 Yao Lu , Hao Cheng , Yujie Fang , Zeyu Wang , Jiaheng Wei , Dongwei Xu , Qi Xuan , Xiaoniu Yang , Zhaowei Zhu

Transformers do not scale very well to long sequence lengths largely because of quadratic self-attention complexity. In the recent months, a wide spectrum of efficient, fast Transformers have been proposed to tackle this problem, more often…

Machine Learning · Computer Science 2020-11-10 Yi Tay , Mostafa Dehghani , Samira Abnar , Yikang Shen , Dara Bahri , Philip Pham , Jinfeng Rao , Liu Yang , Sebastian Ruder , Donald Metzler

This work investigates distillation methods for large language models (LLMs) with the goal of developing compact models that preserve high performance. Several existing approaches are reviewed, with a discussion of their respective…

Computation and Language · Computer Science 2025-11-10 Grigory Kovalev , Mikhail Tikhomirov

The possible paralelism existing between phase transitions and fracture in disordered materials, is discussed using the well-known Fiber Bundle Models and a probabilistic approach suited to smooth fluctuations near the critical point. Two…

Statistical Mechanics · Physics 2009-11-07 Y. Moreno , J. B. Gomez , A. F. Pacheco

Recently much attention has been paid to the study of the robustness of interdependent and multiplex networks and, in particular, networks of networks. The robustness of interdependent networks can be evaluated by the size of a mutually…

Statistical Mechanics · Physics 2015-06-18 Ginestra Bianconi , Sergey N. Dorogovtsev

The rapid development in the performance of large language models (LLMs) is accompanied by the escalation of model size, leading to the increasing cost of model training and inference. Previous research has discovered that certain layers in…

Computation and Language · Computer Science 2024-10-14 Fangwei Zhu , Dian Li , Jiajun Huang , Gang Liu , Hui Wang , Zhifang Sui

There is evidence that transformers offer state-of-the-art recognition performance on tasks involving overhead imagery (e.g., satellite imagery). However, it is difficult to make unbiased empirical comparisons between competing deep…

Computer Vision and Pattern Recognition · Computer Science 2022-11-02 Francesco Luzi , Aneesh Gupta , Leslie Collins , Kyle Bradbury , Jordan Malof

Since its inception in "Attention Is All You Need", transformer architecture has led to revolutionary advancements in NLP. The attention layer within the transformer admits a sequence of input tokens $X$ and makes them interact through…

Machine Learning · Computer Science 2024-02-23 Davoud Ataee Tarzanagh , Yingcong Li , Christos Thrampoulidis , Samet Oymak

Recent work on permutation-based model merging has shown impressive low- or zero-barrier mode connectivity between models from completely different initializations. However, this line of work has not yet extended to the Transformer…

Computation and Language · Computer Science 2024-12-17 Neha Verma , Maha Elbayad

Transactional data may be represented as a bipartite graph $G:=(L \cup R, E)$, where $L$ denotes agents, $R$ denotes objects visible to many agents, and an edge in $E$ denotes an interaction between an agent and an object. Unsupervised…

Combinatorics · Mathematics 2019-07-22 R W R Darling , Cheyne Homberger

From generation to generation there are increasing requirements for wireless standards both in terms of spectral and energy efficiency. While up to now the layered wireless transceiver architecture worked allowing for, e.g., separation of…

Networking and Internet Architecture · Computer Science 2023-06-29 Pawel Kryszkiewicz , Pawel Sroka , Marcin Hoffmann , Marcin Wachowiak