English
Related papers

Related papers: Richer Countries and Richer Representations

200 papers

Abstractive text summarization is surging with the number of training samples to cater to the needs of the deep learning models. These models tend to exploit the training data representations to attain superior performance by improving the…

Computation and Language · Computer Science 2023-12-21 Yash Kumar Atri , Vikram Goyal , Tanmoy Chakraborty

Adoption of recently developed methods from machine learning has given rise to creation of drug-discovery knowledge graphs (KG) that utilize the interconnected nature of the domain. Graph-based modelling of the data, combined with KG…

Machine Learning · Computer Science 2022-07-27 Stephen Bonner , Ufuk Kirik , Ola Engkvist , Jian Tang , Ian P Barrett

We present a model of speech perception which takes into account effects of correlations between sounds. Words in this model correspond to the attractors of a suitably chosen descent dynamics. The resulting lexicon is rich in short words,…

Statistical Mechanics · Physics 2025-02-28 Jean-Marc Luck , Anita Mehta

Gender bias in Language and Vision datasets and models has the potential to perpetuate harmful stereotypes and discrimination. We analyze gender bias in two Language and Vision datasets. Consistent with prior work, we find that both…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Sophia Harrison , Eleonora Gualdoni , Gemma Boleda

Recent language models generate false but plausible-sounding text with surprising frequency. Such "hallucinations" are an obstacle to the usability of language-based AI systems and can harm people who rely upon their outputs. This work…

Computation and Language · Computer Science 2024-03-21 Adam Tauman Kalai , Santosh S. Vempala

We analyse a preferential urn model with randomness using the replica method. The preferential urn model is a stochastic model based on the concept "the rich get richer." The replica analysis clarifies that the preferential urn model with…

Disordered Systems and Neural Networks · Physics 2009-11-11 Jun Ohkubo , Muneki Yasuda , Kazuyuki Tanaka

Vector quantization is a technique in machine learning that discretizes continuous representations into a set of discrete vectors. It is widely employed in tokenizing data representations for large language models, diffusion models, and…

Machine Learning · Computer Science 2026-03-19 Wenhao Zhao , Qiran Zou , Rushi Shah , Yudi Wu , Zhouhan Lin , Dianbo Liu

Most reinforcement learning methods operate on propositional representations of the world state. Such representations are often intractably large and generalize poorly. Using a deictic representation is believed to be a viable alternative:…

Machine Learning · Computer Science 2013-01-07 Sarah Finney , Natalia Gardiol , Leslie Pack Kaelbling , Tim Oates

Sentence encoders map sentences to real valued vectors for use in downstream applications. To peek into these representations - e.g., to increase interpretability of their results - probing tasks have been designed which query them for…

Computation and Language · Computer Science 2020-10-29 Steffen Eger , Johannes Daxenberger , Iryna Gurevych

Politicians world-wide frequently promise a better life for their citizens. We find that the probability that a country will increase its {\it per capita} GDP ({\it gdp}) rank within a decade follows an exponential distribution with decay…

General Finance · Quantitative Finance 2012-09-14 Boris Podobnik , Davor Horvatic , Dror Y. Kenett , H. Eugene Stanley

We investigate how embedding dimension affects the emergence of an internal "world model" in a transformer trained with reinforcement learning to perform bubble-sort-style adjacent swaps. Models achieve high accuracy even with very small…

Machine Learning · Computer Science 2025-10-22 Brady Bhalla , Honglu Fan , Nancy Chen , Tony Yue YU

Representing words by vectors, or embeddings, enables computational reasoning and is foundational to automating natural language tasks. For example, if word embeddings of similar words contain similar values, word similarity can be readily…

Computation and Language · Computer Science 2022-02-02 Carl Allen

Pre-trained language models (LMs) may perpetuate biases originating in their training corpus to downstream models. We focus on artifacts associated with the representation of given names (e.g., Donald), which, depending on the corpus, may…

Computation and Language · Computer Science 2020-09-17 Vered Shwartz , Rachel Rudinger , Oyvind Tafjord

The representation space of pretrained Language Models (LMs) encodes rich information about words and their relationships (e.g., similarity, hypernymy, polysemy) as well as abstract semantic notions (e.g., intensity). In this paper, we…

Computation and Language · Computer Science 2023-06-02 Qing Lyu , Marianna Apidianaki , Chris Callison-Burch

Language encodes societal beliefs about social groups through word patterns. While computational methods like word embeddings enable quantitative analysis of these patterns, studies have primarily examined gradual shifts in Western…

Computation and Language · Computer Science 2026-01-30 Yuxi Ma , Yongqian Peng , Yixin Zhu

In the last 150 years, CO2 concentration in the atmosphere has increased from 280 parts per million to 400 parts per million. This has caused an increase in the average global temperatures by nearly 0.7 degree centigrade due to the…

Signal Processing · Electrical Eng. & Systems 2020-06-28 Sandeep Kumar , Pranab K. Muhuri

Ensembling is commonly regarded as an effective way to improve the general performance of models in machine learning, while also increasing the robustness of predictions. When it comes to algorithmic fairness, heterogeneous ensembles,…

Machine Learning · Computer Science 2025-01-27 Estanislao Claucich , Sara Hooker , Diego H. Milone , Enzo Ferrante , Rodrigo Echeveste

Understanding what constitutes high-quality pre-training data remains a central question in language model training. In this work, we investigate whether benchmark performance is primarily driven by the degree of statistical pattern overlap…

Computation and Language · Computer Science 2026-02-12 Woojin Chung , Jeonghoon Kim

A society or country with income equally distributed among its people is truly a fiction! The phenomena of socioeconomic inequalities have been plaguing mankind from times immemorial. We are interested in gaining an insight about the…

General Finance · Quantitative Finance 2018-08-07 Kiran Sharma , Subhradeep Das , Anirban Chakraborti

Pre-trained large foundation models play a central role in the recent surge of artificial intelligence, resulting in fine-tuned models with remarkable abilities when measured on benchmark datasets, standard exams, and applications. Due to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Shaeke Salman , Md Montasir Bin Shams , Xiuwen Liu