English
Related papers

Related papers: Training new operators - the first six months

200 papers

While pretrained models such as BERT have shown large gains across natural language understanding tasks, their performance can be improved by further training the model on a data-rich intermediate task, before fine-tuning it on a target…

Recent reinforcement learning (RL) algorithms have demonstrated impressive results in simulated driving environments. However, autonomous vehicles trained in simulation often struggle to work well in the real world due to the fidelity gap…

Robotics · Computer Science 2025-01-17 Sang-Hyun Lee , Daehyeok Kwon , Seung-Woo Seo

The Fermilab Linac delivers 400 MeV H- beam to the rest of the accelerator chain. Providing stable intensity, energy, and emittance is key since it directly affects downstream machines. To operate high current beam, accelerators must…

Accelerator Physics · Physics 2023-07-11 R. Sharankova , M. Mwaniki , K. Seiya , M. Wesley

Pre-training techniques play a crucial role in deep learning, enhancing models' performance across a variety of tasks. By initially training on large datasets and subsequently fine-tuning on task-specific data, pre-training provides a solid…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Fulong Ma , Guoyang Zhao , Weiqing Qi , Ming Liu , Jun Ma

Tear film (TF) breakup is a key driver of understanding dry eye disease, yet estimating TF thickness and osmolarity from fluorescence (FL) imaging typically requires solving computationally expensive inverse problems. We propose an operator…

Numerical Analysis · Mathematics 2026-01-14 Qinying Chen , Arnab Roy , Tobin A. Driscoll

While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this training data bottleneck,…

Computation and Language · Computer Science 2016-11-02 Yingce Xia , Di He , Tao Qin , Liwei Wang , Nenghai Yu , Tie-Yan Liu , Wei-Ying Ma

Zeroth-order (ZO) fine-tuning is attractive for large language models because it replaces backpropagation with forward objective evaluations. Existing implementations nevertheless execute ZO algorithms inside conventional training loops,…

Machine Learning · Computer Science 2026-05-28 Zelin Li , Caiwen Ding

It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Asher Trockman , J. Zico Kolter

Offline reinforcement learning (ORL) holds great promise for robot learning due to its ability to learn from arbitrary pre-generated experience. However, current ORL benchmarks are almost entirely in simulation and utilize contrived…

Robotics · Computer Science 2022-10-14 Gaoyue Zhou , Liyiming Ke , Siddhartha Srinivasa , Abhinav Gupta , Aravind Rajeswaran , Vikash Kumar

The development of state-of-the-art large language models is commonly understood as a two-stage process involving pre-training and post-training. We point out the need for an additional intermediate stage called reinforcement mid-training…

Computation and Language · Computer Science 2025-09-30 Yijun Tian , Shaoyu Chen , Zhichao Xu , Yawei Wang , Jinhe Bi , Peng Han , Wei Wang

NLP is currently dominated by general-purpose pretrained language models like RoBERTa, which achieve strong performance on NLU tasks through pretraining on billions of words. But what exact knowledge or skills do Transformer LMs learn from…

Computation and Language · Computer Science 2020-11-11 Yian Zhang , Alex Warstadt , Haau-Sing Li , Samuel R. Bowman

Generalizing deep learning models to unknown target domain distribution with low latency has motivated research into test-time training/adaptation (TTT/TTA). Existing approaches often focus on improving test-time training performance under…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yushu Li , Xun Xu , Yongyi Su , Kui Jia

We report a course on teaching in physics lab for teachers enrolled in Formative Active Training, which actually allows to obtain the teacher qualification in Italy. The course was designed with the purpose of showing in practice what means…

Physics Education · Physics 2016-01-08 Vera Montalbano , Roberto Benedetti

Effective exoskeleton assistance requires co-adaptation: as the device alters joint dynamics, the user reorganizes neuromuscular coordination, creating a non-stationary learning problem. Most learning-based approaches do not explicitly…

Robotics · Computer Science 2026-03-10 Yifei Yuan , Ghaith Androwis , Xianlian Zhou

While machine-learned interatomic potentials have become a mainstay for modeling materials, designing training sets that lead to robust potentials is challenging. Automated methods, such as active learning and on-the-fly learning, construct…

Materials Science · Physics 2023-10-13 Jason A. Meziere , Yu Luo , Yi Zia , LK Beland , MR Daymond , Gus L. W. Hart

This thesis examines self-attention training through the lens of Optimal Transport (OT) and develops an OT-based alternative for tabular classification. The study tracks intermediate projections of the self-attention layer during training…

Machine Learning · Statistics 2026-02-19 Alessandro Quadrio , Antonio Candelieri

We study the mirror-field interaction in several frameworks: when it is driven, when it is affected by an environment and when a two-level atom is introduced in the cavity. By using operator techniques we show how these problems may be…

Quantum Physics · Physics 2015-06-02 C. Ventura-Velázquez , B. M. Rodríguez-Lara , H. M. Moya-Cessa

We present a Real-Time Operator Takeover (RTOT) paradigm that enables operators to seamlessly take control of a live visuomotor diffusion policy, guiding the system back to desirable states or providing targeted corrective demonstrations.…

Robotics · Computer Science 2026-04-01 Marco Moletta , Michael C. Welle , Nils Ingelhag , Jesper Munkeby , Danica Kragic

The NML cryogenic plant cools two individually cryostated superconducting radio frequency (SRF) capture cavities and one prototype ILC cryomodule with eight SRF cavities. This complex accelerates electrons at 150 MeV for the Integrable…

Recent studies have proven that the training of neural machine translation (NMT) can be facilitated by mimicking the learning process of humans. Nevertheless, achievements of such kind of curriculum learning rely on the quality of…

Computation and Language · Computer Science 2022-10-20 Yu Wan , Baosong Yang , Derek F. Wong , Yikai Zhou , Lidia S. Chao , Haibo Zhang , Boxing Chen
‹ Prev 1 4 5 6 7 8 10 Next ›