中文
相关论文

相关论文: Peacock: Learning Long-Tail Topic Features for Ind…

200 篇论文

Hierarchical latent tree analysis (HLTA) is recently proposed as a new method for topic detection. It differs fundamentally from the LDA-based methods in terms of topic definition, topic-document relationship, and learning method. It has…

机器学习 · 计算机科学 2015-08-06 Peixian Chen , Nevin L. Zhang , Leonard K. M. Poon , Zhourong Chen

Anomaly detection (AD) identifies the defect regions of a given image. Recent works have studied AD, focusing on learning AD without abnormal images, with long-tailed distributed training data, and using a unified model for all classes. In…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Chiao-An Yang , Kuan-Chuan Peng , Raymond A. Yeh

Large language models (LLMs) face persistent challenges when handling long-context tasks, most notably the lost in the middle issue, where information located in the middle of a long input tends to be underutilized. Some existing methods…

人工智能 · 计算机科学 2025-10-22 Song Yu , Xiaofei Xu , Ke Deng , Li Li , Lin Tian

Familia is an open-source toolkit for pragmatic topic modeling in industry. Familia abstracts the utilities of topic modeling in industry as two paradigms: semantic representation and semantic matching. Efficient implementations of the two…

信息检索 · 计算机科学 2017-08-01 Di Jiang , Zeyu Chen , Rongzhong Lian , Siqi Bao , Chen Li

While generative models such as Latent Dirichlet Allocation (LDA) have proven fruitful in topic modeling, they often require detailed assumptions and careful specification of hyperparameters. Such model complexity issues only compound when…

计算与语言 · 计算机科学 2018-09-05 Ryan J. Gallagher , Kyle Reing , David Kale , Greg Ver Steeg

A U.S. Senator from South Dakota donated documents that were accumulated during his service as a house representative and senator to be housed at the Bridges library at South Dakota State University. This project investigated the utility of…

信息检索 · 计算机科学 2019-04-30 Damon Bayer , Semhar Michael

Topic Modelling (TM) is from the research branches of natural language understanding (NLU) and natural language processing (NLP) that is to facilitate insightful analysis from large documents and datasets, such as a summarisation of main…

计算与语言 · 计算机科学 2023-04-19 Bernadeta Griciūtė , Lifeng Han , Goran Nenadic

Topic modeling analyzes documents to learn meaningful patterns of words. However, existing topic models fail to learn interpretable topics when working with large and heavy-tailed vocabularies. To this end, we develop the Embedded Topic…

信息检索 · 计算机科学 2019-07-12 Adji B. Dieng , Francisco J. R. Ruiz , David M. Blei

Reinforcement learning(RL) post-training has become essential for aligning large language models (LLMs), yet its efficiency is increasingly constrained by the rollout phase, where long trajectories are generated token by token. We identify…

This paper proposes a new methodology to study sequential corpora by implementing a two-stage algorithm that learns time-based topics with respect to a scale of document positions and introduces the concept of Topic Scaling which ranks…

信息检索 · 计算机科学 2021-04-05 Sami Diaf , Ulrich Fritsche

While there is an increased discourse on large language models (LLMs) like ChatGPT and DeepSeek, there is no comprehensive understanding of how users of online platforms, like Reddit, perceive these models. This is an important omission…

社会与信息网络 · 计算机科学 2025-02-27 Krishnaveni Katta

In this paper, we compare different methods to extract skill requirements from job advertisements. We consider three top-down methods that are based on expert-created dictionaries of keywords, and a bottom-up method of unsupervised topic…

综合经济学 · 经济学 2022-07-27 Ziqiao Ao , Gergely Horvath , Chunyuan Sheng , Yifan Song , Yutong Sun

Long-tailed classification poses a challenge due to its heavy imbalance in class probabilities and tail-sensitivity risks with asymmetric misprediction costs. Recent attempts have used re-balancing loss and ensemble methods, but they are…

机器学习 · 计算机科学 2023-03-22 Bolian Li , Ruqi Zhang

Long-tailed data distributions pose challenges for a variety of domains like e-commerce, finance, biomedical science, and cyber security, where the performance of machine learning models is often dominated by head categories while tail…

机器学习 · 计算机科学 2024-10-31 Haohui Wang , Weijie Guan , Jianpeng Chen , Zi Wang , Dawei Zhou

Artificial intelligence (AI) is widely deployed to solve problems related to marketing attribution and budget optimization. However, AI models can be quite complex, and it can be difficult to understand model workings and insights without…

计算与语言 · 计算机科学 2024-04-23 Yilin Gao , Sai Kumar Arava , Yancheng Li , James W. Snyder

Mapping a search query to a set of relevant categories in the product taxonomy is a significant challenge in e-commerce search for two reasons: 1) Training data exhibits severe class imbalance problem due to biased click behavior, and 2)…

信息检索 · 计算机科学 2021-05-11 Ali Ahmadvand , Surya Kallumadi , Faizan Javed , Eugene Agichtein

Topic modeling is a widely used technique for revealing underlying thematic structures within textual data. However, existing models have certain limitations, particularly when dealing with short text datasets that lack co-occurring words.…

人工智能 · 计算机科学 2023-12-18 Han Wang , Nirmalendu Prakash , Nguyen Khoi Hoang , Ming Shan Hee , Usman Naseem , Roy Ka-Wei Lee

We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over…

信息检索 · 计算机科学 2012-07-19 Michal Rosen-Zvi , Thomas Griffiths , Mark Steyvers , Padhraic Smyth

DeepSeek, a Chinese Artificial Intelligence (AI) startup, has released their V3 and R1 series models, which attracted global attention due to their low cost, high performance, and open-source advantages. This paper begins by reviewing the…

人工智能 · 计算机科学 2025-07-15 Luolin Xiong , Haofen Wang , Xi Chen , Lu Sheng , Yun Xiong , Jingping Liu , Yanghua Xiao , Huajun Chen , Qing-Long Han , Yang Tang

Topic modeling plays a vital role in uncovering hidden semantic structures within text corpora, but existing models struggle in low-resource settings where limited target-domain data leads to unstable and incoherent topic inference. We…

计算与语言 · 计算机科学 2025-06-10 Pritom Saha Akash , Kevin Chen-Chuan Chang
‹ 上一页 1 8 9 10 下一页 ›