English
Related papers

Related papers: Three-Phase Transformer

200 papers

Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yutong Chen , Yiming Wang , Xucong Zhang , Sergey Prokudin , Siyu Tang

We propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere. In nGPT, all vectors forming the embeddings, MLP, attention matrices and hidden states are unit norm…

Machine Learning · Computer Science 2025-04-25 Ilya Loshchilov , Cheng-Ping Hsieh , Simeng Sun , Boris Ginsburg

To obtain high-quality positron emission tomography (PET) images while minimizing radiation exposure, various methods have been proposed for reconstructing standard-dose PET (SPET) images from low-dose PET (LPET) sinograms directly.…

Image and Video Processing · Electrical Eng. & Systems 2023-08-11 Jiaqi Cui , Pinxian Zeng , Xinyi Zeng , Peng Wang , Xi Wu , Jiliu Zhou , Yan Wang , Dinggang Shen

Transformer-based 3D human pose estimation methods suffer from high computational costs due to the quadratic complexity of self-attention with respect to sequence length. Additionally, pose sequences often contain significant redundancy…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Zenghao Zheng , Lianping Yang , Hegui Zhu , Mingrui Ye

Standard Transformers have a fixed computational depth, fundamentally limiting their ability to generalize to tasks requiring variable-depth reasoning, such as multi-hop graph traversal or nested logic. We propose a depth-recurrent…

Machine Learning · Computer Science 2026-03-24 Hung-Hsuan Chen

Transformers process tokens in parallel but are temporally shallow: at position $t$, each layer attends to key-value pairs computed based on the previous layer, yielding a depth capped by the number of layers. Recurrent models offer…

Machine Learning · Computer Science 2026-04-24 Costin-Andrei Oncescu , Depen Morwani , Samy Jelassi , Alexandru Meterez , Mujin Kwun , Sham Kakade

Every existing method for compressing 3D Gaussian Splatting, NeRF, or transformer-based 3D reconstructors requires learning a data-dependent codebook through per-scene fine-tuning. We show this is unnecessary. The parameter vectors that…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Jae Joong Lee

The Transformer architecture has gained growing attention in graph representation learning recently, as it naturally overcomes several limitations of graph neural networks (GNNs) by avoiding their strict structural inductive biases and…

Machine Learning · Statistics 2022-06-14 Dexiong Chen , Leslie O'Bray , Karsten Borgwardt

A widely cited result by Dong et al. (2021) showed that Transformers built from self-attention alone, without skip connections or feed-forward layers, suffer from rapid rank collapse: all token representations converge to a single…

Machine Learning · Computer Science 2026-04-28 Giansalvo Cirrincione

Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This keeps the field isolated from the broader Transformer ecosystem, limiting…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Kadir Yilmaz , Adrian Kruse , Tristan Höfer , Daan de Geus , Bastian Leibe

Many complex physical systems admit natural decomposition into an exactly solvable component and a perturbative correction. Rather than training neural networks to learn complete trajectories from scratch, we introduce Neural Network…

Computational Physics · Physics 2025-12-02 Zhenhao Chen , Mutian Shen , Boris Fain , Zohar Nussinov

This paper presents a memory efficient, high throughput parallel lifting based running three dimensional discrete wavelet transform (3-D DWT) architecture. 3-D DWT is constructed by combining the two spatial and four temporal processors.…

Hardware Architecture · Computer Science 2015-09-16 Batta Kota Naga Srinivasarao , Indrajit Chakrabarti

We propose DepthTCM, a physics-aware end-to-end framework for depth map compression. In our framework of DepthTCM, the high-bit depth map is first converted to a conventional 3-channel image representation losslessly using a method inspired…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Young-Seo Chang , Yatong An , Jae-Sang Hyun

Resonant rectifier topologies would be a promising candidate for achieving simple, compact, and reliable single-stage wireless power transfer (WPT) receiver if not for the lack of good DC regulation capability. This paper investigates the…

Systems and Control · Electrical Eng. & Systems 2021-04-10 Kerui Li , Siew Chong Tan , Ron Shu Yuen Hui

The highly popular Transformer architecture, based on self-attention, is the foundation of large pretrained models such as BERT, that have become an enduring paradigm in NLP. While powerful, the computational resources and time required to…

Computation and Language · Computer Science 2021-08-31 Ran Tian , Joshua Maynez , Ankur P. Parikh

Transformer has been widely adopted in Neural Machine Translation (NMT) because of its large capacity and parallel training of sequence generation. However, the deployment of Transformer is challenging because different scenarios require…

Computation and Language · Computer Science 2021-06-21 Peng Gao , Shijie Geng , Yu Qiao , Xiaogang Wang , Jifeng Dai , Hongsheng Li

We present in this paper a new architecture, the Pattern Attention Transformer (PAT), that is composed of the new doughnut kernel. Compared with tokens in the NLP field, Transformer in computer vision has the problem of handling the high…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 WenYuan Sheng

Transformer terminal equivalents obtained via admittance measurements are suitable for simulating high-frequency transient interaction between the transformer and the network. This paper augments the terminal equivalent approach with a…

Computational Engineering, Finance, and Science · Computer Science 2016-11-22 Bjorn Gustavsen , Alvaro Portillo , Rodrigo Ronchi , Asgeir Mjelve

According to the classification using projective representations of the SO(3) group, there exist two topologically distinct gapped phases in spin-1 chains. The symmetry-protected topological (SPT) phase possesses half-integer projective…

Strongly Correlated Electrons · Physics 2014-01-09 Wei Li , Andreas Weichselbaum , Jan von Delft

We investigate whether the Feed-Forward Network (FFN) sublayer in a decoder-only transformer can be replaced by an explicit learned memory graph while preserving the surrounding autoregressive architecture. The proposed Graph Memory…

Machine Learning · Computer Science 2026-05-29 Nicola Zanarini , Niccolò Ferrari , Evelina Lamma
‹ Prev 1 2 3 10 Next ›