Related papers: BERT in Plutarch's Shadows
The oldest (c. 4000 BC) undeciphered language is the Old European Script known from approximately 940 inscribed objects (82% of inscriptions on pottery) found in excavations in the Vinca-Tordos region Transylvania. Also, it is not known for…
This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…
We expand the substantive terrain of QI's reach by illuminating a body of political theory that to date has been elaborated in strictly classical language and formalisms but has complex features that seem to merit generalizations of the…
Online discourse is often perceived as polarized and unproductive. While some conversational discourse parsing frameworks are available, they do not naturally lend themselves to the analysis of contentious and polarizing discussions.…
We introduce an original model of quantum phenomena, a model that provides a picture of a "deep structure", an "underlying pattern" of quantum dynamics. We propose that the source of a particle and all of that particle's possible detectors…
I present Lepton (Letter Prediction), a fine-tuned BERT classifier that predicts whether a title in a Classical Chinese wenji table of contents is a personal letter or a closely confusable preface (particularly the farewell-preface). Lepton…
Recent advancements in NLP have spurred significant interest in analyzing social media text data for identifying linguistic features indicative of mental health issues. However, the domain of Expressive Narrative Stories (ENS)-deeply…
We explore the influence and interconnectivity of philosophical thinkers within the Wikipedia knowledge network. Using a dataset of 237 articles dedicated to philosophers across nine different language editions (Arabic, Chinese, English,…
Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we…
This paper revisits Buridan's Bridge paradox (Sophismata, chapter 8, Sophism 17), itself close kin to the Liar paradox, a version of which also appears in Bradwardine's Insolubilia. Prompted by the occurrence of the paradox in Cervantes's…
We study the use of BERT for non-factoid question-answering, focusing on the passage re-ranking task under varying passage lengths. To this end, we explore the fine-tuning of BERT in different learning-to-rank setups, comprising both…
Word meaning changes over time, depending on linguistic and extra-linguistic factors. Associating a word's correct meaning in its historical context is a central challenge in diachronic research, and is relevant to a range of NLP tasks,…
Many of the most familiar features of our everyday environment, and some of our basic notions about it, stem from Relativistic Quantum Field Theory (RQFT). We argue in particular that the origin of common names, verbs, adjectives such as…
It has been shown at other occasions that recent results of modern physics can be used to shed some more light onto the foundations of the world, provided the actual task of philosophy is being re-interpreted in terms of a theory which is…
Recent works show that pre-trained language models (PTLMs), such as BERT, possess certain commonsense and factual knowledge. They suggest that it is promising to use PTLMs as "neural knowledge bases" via predicting masked words.…
Contextual pretrained language models, such as BERT (Devlin et al., 2019), have made significant breakthrough in various NLP tasks by training on large scale of unlabeled text re-sources.Financial sector also accumulates large amount of…
The Arabic language is a morphologically rich language with relatively few resources and a less explored syntax compared to English. Given these limitations, Arabic Natural Language Processing (NLP) tasks like Sentiment Analysis (SA), Named…
Contrastive learning has shown great potential in unsupervised sentence embedding tasks, e.g., SimCSE. However, We find that these existing solutions are heavily affected by superficial features like the length of sentences or syntactic…
Transformer-based masked language models such as BERT, trained on general corpora, have shown impressive performance on downstream tasks. It has also been demonstrated that the downstream task performance of such models can be improved by…
During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental…