English
Related papers

Related papers: Separating Principles Below WKL0

200 papers

With the rapid development of deep learning, the increasing complexity and scale of parameters make training a new model increasingly resource-intensive. In this paper, we start from the classic convolutional neural network (CNN) and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jiacong Hu , Jing Gao , Jingwen Ye , Yang Gao , Xingen Wang , Zunlei Feng , Mingli Song

Current research highlights the great potential of Large Language Models (LLMs) for constructing Scholarly Knowledge Graphs (SKGs). One particularly complex step in this process is relation extraction, aimed at identifying suitable…

Information Retrieval · Computer Science 2025-02-18 Sandra Schaftner

Large language models (LLMs) have proven to be highly effective for solving complex reasoning tasks. Surprisingly, their capabilities can often be improved by iterating on previously generated solutions. In this context, a reasoning plan…

Artificial Intelligence · Computer Science 2025-12-05 MohammadHossein Bateni , Vincent Cohen-Addad , Yuzhou Gu , Silvio Lattanzi , Simon Meierhans , Christopher Mohri

Large language models (LLMs) can generate code rapidly but remain unreliable for scientific algorithms whose correctness depends on structural assumptions rarely explicit in the source literature. We introduce a multi-stage LLM-assisted…

Computational Physics · Physics 2026-04-13 Yi Zhou

Large language models (LLMs) such as GPT-4 sometimes appear to be creative, solving novel tasks often with a few demonstrations in the prompt. These tasks require the models to generalize on distributions different from those from training…

Computation and Language · Computer Science 2024-12-31 Jiajun Song , Zhuoyan Xu , Yiqiao Zhong

In this paper we give a new proof of the Ne\v{s}et\v{r}il-R\"odl Theorem, a deep result of discrete mathematics which is one of the cornerstones of the structural Ramsey theory. In contrast to the well-known proofs which employ intricate…

Category Theory · Mathematics 2017-08-08 Dragan Masulovic

We consider high temperature KMS states for quantum spin systems on a lattice. We prove a large deviation principle for the distribution of empirical averages $\frac{1}{|\Lambda|} \sum_{i\in\Lambda} X_i$, where the $X_i$'s are copies of a…

Mathematical Physics · Physics 2009-11-10 K. Netocny , F. Redig

Discriminative pre-trained language models (PrLMs) can be generalized as denoising auto-encoders that work with two procedures, ennoising and denoising. First, an ennoising process corrupts texts with arbitrary noising functions to…

Computation and Language · Computer Science 2022-10-12 Zhuosheng Zhang , Hai Zhao , Ming Zhou

Finding a logical formula that separates positive and negative examples given in the form of labeled data items is fundamental in applications such as concept learning, reverse engineering of database queries, generating referring…

Logic in Computer Science · Computer Science 2022-08-18 Jean Christoph Jung , Carsten Lutz , Hadrien Pulcini , Frank Wolter

We study the pigeonhole principle for $\Sigma_2$-definable injections with domain twice as large as the codomain, and the weak K\"onig lemma for $\Delta^0_2$-definable trees in which every level has at least half of the possible nodes. We…

Logic · Mathematics 2019-12-10 David Belanger , Chitat Chong , Wei Wang , Tin Lok Wong , Yue Yang

Machine learning algorithms typically assume independent and identically distributed samples in training and at test time. Much work has shown that high-performing ML classifiers can degrade significantly and provide overly-confident, wrong…

Computation and Language · Computer Science 2023-03-09 Jie Ren , Jiaming Luo , Yao Zhao , Kundan Krishna , Mohammad Saleh , Balaji Lakshminarayanan , Peter J. Liu

We consider the problem of learning a nonlinear function over a network of learners in a fully decentralized fashion. Online learning is additionally assumed, where every learner receives continuous streaming data locally. This learning…

Machine Learning · Computer Science 2021-03-01 Jeongmin Chae , Songnam Hong

We propose a method to teach multiple large language models (LLM) to collaborate by interleaving their generations at the token level. We model the decision of which LLM generates the next token as a latent variable. By optimizing the…

Computation and Language · Computer Science 2024-08-28 Shannon Zejiang Shen , Hunter Lang , Bailin Wang , Yoon Kim , David Sontag

Training deep neural networks (DNNs) is an important and challenging optimization problem in machine learning due to its non-convexity and non-separable structure. The alternating minimization (AM) approaches split the composition structure…

Machine Learning · Computer Science 2023-04-05 Jintao Xu , Chenglong Bao , Wenxun Xing

Large Language Models (LLMs) can solve previously intractable tasks given only natural-language instructions and a few examples, but they remain difficult to steer precisely and lack a key capability for building reliable software at scale:…

Programming Languages · Computer Science 2026-03-19 Jonathan Laurent , André Platzer

The paper addresses the multiple kernel learning (MKL) problem for one-class classification (OCC). For this purpose, based on the Fisher null-space one-class classification principle, we present a multiple kernel learning algorithm where a…

Machine Learning · Computer Science 2021-09-28 Shervin Rahimzadeh Arashloo

A major theme in arithmetic combinatorics is proving multiple recurrence results on semigroups (such as Szemer\'edi's theorem) and this can often be done using methods of ergodic Ramsey theory. What usually lies at the heart of such proofs…

Logic · Mathematics 2017-04-18 Anush Tserunyan

This paper describes an algorithm for the compilation of a two (or more) level orthographic or phonological rule notation into finite state transducers. The notation is an alternative to the standard one deriving from Koskenniemi's work: it…

cmp-lg · Computer Science 2008-02-03 Edmund Grimley-Evans , George Anton Kiraz , Stephen G. Pulman

We present a novel nonnegative tensor decomposition method, called Legendre decomposition, which factorizes an input tensor into a multiplicative combination of parameters. Thanks to the well-developed theory of information geometry, the…

Machine Learning · Statistics 2020-01-29 Mahito Sugiyama , Hiroyuki Nakahara , Koji Tsuda

Large Language Models (LLMs) have demonstrated strong generalization capabilities across a wide range of natural language processing (NLP) tasks. However, they exhibit notable weaknesses in character-level string manipulation, struggling…

Computation and Language · Computer Science 2025-03-28 Zhen Xiong , Yujun Cai , Bryan Hooi , Nanyun Peng , Zhecheng Li , Yiwei Wang
‹ Prev 1 4 5 6 7 8 10 Next ›