中文
相关论文

相关论文: Testing different Log Bases For Vector Model Weigh…

200 篇论文

To compare autoregressive language models at scale, we propose using log-likelihood vectors computed on a predefined text set as model features. This approach has a solid theoretical basis: when treated as model coordinates, their squared…

计算与语言 · 计算机科学 2025-06-03 Momose Oyama , Hiroaki Yamagiwa , Yusuke Takase , Hidetoshi Shimodaira

In most recent studies, gender bias in document ranking is evaluated with the NFaiRR metric, which measures bias in a ranked list based on an aggregation over the unbiasedness scores of each ranked document. This perspective in measuring…

计算与语言 · 计算机科学 2024-03-12 Amin Abolghasemi , Leif Azzopardi , Arian Askari , Maarten de Rijke , Suzan Verberne

Item factor analysis (IFA) refers to the factor models and statistical inference procedures for analyzing multivariate categorical data. IFA techniques are commonly used in social and behavioral sciences for analyzing item-level response…

统计方法学 · 统计学 2020-04-17 Yunxiao Chen , Siliang Zhang

The main aim of an information retrieval system is to extract appropriate information from an enormous collection of data based on users need. The basic concept of the information retrieval system is that when a user sends out a query, the…

信息检索 · 计算机科学 2020-12-17 Abdulmalik Johar

The aim of idiomify is to build a collocation-supplemented reverse dictionary of idioms for the non-native learners of English. We aim to do so because the reverse dictionary could help the non-natives explore idioms on demand, and the…

计算与语言 · 计算机科学 2022-04-13 Eu-Bin Kim

Recognition and retrieval of textual content from the large document collections have been a powerful use case for the document image analysis community. Often the word is the basic unit for recognition as well as retrieval. Systems that…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Siddhant Bansal , Praveen Krishnan , C. V. Jawahar

This paper proposes an easy-to-use method for one-class classification: Repeated Element-wise Folding (REF). The algorithm consists of repeatedly standardizing and applying an element-wise folding operation on the one-class training data.…

机器学习 · 计算机科学 2025-06-19 Jenni Raitoharju

We propose an embarrassingly simple method -- instance-aware repeat factor sampling (IRFS) to address the problem of imbalanced data in long-tailed object detection. Imbalanced datasets in real-world object detection often suffer from a…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Burhaneddin Yaman , Tanvir Mahmud , Chun-Hao Liu

A common approach for knowledge-base entity search is to consider an entity as a document with multiple fields. Models that focus on matching query terms in different fields are popular choices for searching such entity representations. An…

信息检索 · 计算机科学 2019-07-03 Shahrzad Naseri , Sheikh Muhammad Sarwar , James Allan

Toxic online content has become a major issue in today's world due to an exponential increase in the use of internet by people of different cultures and educational background. Differentiating hate speech and offensive language is a key…

计算与语言 · 计算机科学 2018-09-25 Aditya Gaydhani , Vikrant Doma , Shrikant Kendre , Laxmi Bhagwat

In context of document classification, where in a corpus of documents their label tags are readily known, an opportunity lies in utilizing label information to learn document representation spaces with better discriminative properties. To…

计算与语言 · 计算机科学 2014-07-28 Ivan Ivek

Stochastic Discount Factor (SDF) models provide a unified framework for asset pricing and risk assessment, yet traditional formulations struggle to incorporate unstructured textual information. We introduce NewsNet-SDF, a novel deep…

投资组合管理 · 定量金融 2025-05-13 Shunyao Wang , Ming Cheng , Christina Dan Wang

Fusing and ranking multimodal information remains always a challenging task. A robust decision-level fusion method should not only be dynamically adaptive for assigning weights to each representation but also incorporate inter-relationships…

信息检索 · 计算机科学 2018-11-29 Dimitris Gkoumas , Dawei Sogn

In this paper we study lifted inference for the Weighted First-Order Model Counting problem (WFOMC), which counts the assignments that satisfy a given sentence in first-order logic (FOL); it has applications in Statistical Relational…

人工智能 · 计算机科学 2019-11-12 Eric Gribkoff , Guy Van den Broeck , Dan Suciu

Social media has become a very popular source of information. With this popularity comes an interest in systems that can classify the information produced. This study tries to create such a system detecting irony in Twitter users. Recent…

计算与语言 · 计算机科学 2023-11-09 Tibor L. R. Krols , Marie Mortensen , Ninell Oldenburg

Keyword extraction is the process of identifying the words or phrases that express the main concepts of text to the best of one's ability. Electronic infrastructure creates a considerable amount of text every day and at all times. This…

计算与语言 · 计算机科学 2021-10-04 Aidin Zehtab-Salmasi , Mohammad-Reza Feizi-Derakhshi , Mohamad-Ali Balafar

The recent advancements of the Semantic Web and Linked Data have changed the working of the traditional web. There is significant adoption of the Resource Description Framework (RDF) format for saving of web-based data. This massive…

数据库 · 计算机科学 2020-09-24 Waqas Ali , Muhammad Saleem , Bin Yao , Aidan Hogan , Axel-Cyrille Ngonga Ngomo

Large language models (LLMs) for formal theorem proving have become a prominent research focus. At present, the proving ability of these LLMs is mainly evaluated through proof pass rates on datasets such as miniF2F. However, this evaluation…

人工智能 · 计算机科学 2025-02-04 Jianyu Zhang , Yongwang Zhao , Long Zhang , Jilin Hu , Xiaokun Luan , Zhiwei Xu , Feng Yang

Case Base Reasoning (CBR) is a case solving technique based on experience in cases that have occurred before with the highest similarity. CBR is used to search for practical work titles. TF-IDF is applied to process the vectorization of…

计算与语言 · 计算机科学 2025-08-29 Agung Sukrisna Jaya , Osvari Arsalan , Danny Matthew Saputra

We study the performance of Arabic text classification combining various techniques: (a) tfidf vs. dependency syntax, for feature selection and weighting; (b) class association rules vs. support vector machines, for classification. The…

计算与语言 · 计算机科学 2014-10-21 Yannis Haralambous , Yassir Elidrissi , Philippe Lenca
‹ 上一页 1 8 9 10 下一页 ›