Related papers: BERT in Plutarch's Shadows
This paper presents yet another personal reflection on one the most important concepts in both science and the humanities: time. This elusive notion has been not only bothering philosophers since Plato and Aristotle. It goes throughout…
BERT is a widely used pre-trained model in natural language processing. However, since BERT is quadratic to the text length, the BERT model is difficult to be used directly on the long-text corpus. In some fields, the collected text data…
As for Shakespeare, a hard-fought debate has emerged about Moli\`ere, a supposedly uneducated actor who, according to some, could not have written the masterpieces attributed to him. In the past decades, the century-old thesis according to…
We present and explain two unpublished remarks of Stefano Berardi connected to game semantics.
This article is the text of a commentary on a talk delivered by Mark Textor entitled 'Brentano's Positing Theory of Existence' in December 2015. It contains ideas on implementing Textor's Neo-Brentanian theory of existence in a natural…
We investigate how well BERT performs on predicting factuality in several existing English datasets, encompassing various linguistic constructions. Although BERT obtains a strong performance on most datasets, it does so by exploiting common…
In this paper I study the connection between logic and metaphysics in Plato's participation theory, from the structural properties of the latter. Although Plato was the first ever to formulate the contradiction principle explicitly (in the…
Flint is a frame-based and action-centered language developed by Van Doesburg et al. to capture and compare different interpretations of sources of norms (e.g. laws or regulations). The aim of this research is to investigate whether Flint…
We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE - 100 CE). Due to the tablets' deterioration, scholars often rely on contextual…
The rise of language models such as BERT allows for high-quality text paraphrasing. This is a problem to academic integrity, as it is difficult to differentiate between original and machine-generated content. We propose a benchmark…
Text generation has made significant advances in the last few years. Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate poorly with human judgments. We propose BLEURT, a learned…
A C.R. note by Alano Ancona from 1980 is reexamined. Within his line of ideas in the proof two new theorems are constructed. No claim of originality. These are put into the context of later research. One far reaching conjecture is given.…
This is a 1-page comment on a wrong paper that recently appeared in PRL (Phys. Rev. Lett. 86 (23), 5393 (2001), also quant-ph/0101004). The authors claim to have shown that using a quantum computer gives an "exponential advantage" for…
This study presents EgyBERT, an Arabic language model pretrained on 10.4 GB of Egyptian dialectal texts. We evaluated EgyBERT's performance by comparing it with five other multidialect Arabic language models across 10 evaluation datasets.…
In this paper dedicated to the memory of Walter Philipp, we formalize the rules of classical$\to$ quantum correspondence and perform a rigorous mathematical analysis of the assumptions in Bell's NO-GO arguments.
We give a brief historical overview of the famous Pythagoras' theorem and Pythagoras. We present a simple proof of the result and dicsuss some extensions. We follow \cite{thales}, \cite{wiki} and \cite{wiki2} for the historical comments and…
Recent work on evaluating grammatical knowledge in pretrained sentence encoders gives a fine-grained view of a small number of phenomena. We introduce a new analysis dataset that also has broad coverage of linguistic phenomena. We annotate…
Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets…
While multilingual language models can improve NLP performance on low-resource languages by leveraging higher-resource languages, they also reduce average performance on all languages (the 'curse of multilinguality'). Here we show another…
The ubiquity of the contemporary language understanding tasks gives relevance to the development of generalized, yet highly efficient models that utilize all knowledge, provided by the data source. In this work, we present SocialBERT - the…