中文
相关论文

相关论文: JobHop: A Large-Scale Dataset of Career Trajectori…

200 篇论文

In this study, we focused on proposing an optimal clustering mechanism for the occupations defined in the well-known US-based occupational database, O*NET. Even though all occupations are defined according to well-conducted surveys in the…

机器学习 · 计算机科学 2025-07-11 Iago Xabier Vázquez García , Damla Partanaz , Emrullah Fatih Yetkin

In this paper, we use machine learning techniques to explore the H-1B application dataset disclosed by the Department of Labor (DOL), from 2008 to 2018, in order to provide more stylized facts of the international workers in US labor…

应用统计 · 统计学 2019-04-25 Barry Ke , Angela Qiao

Today, big data is generated from many sources and there is a huge demand for storing, managing, processing, and querying on big data. The MapReduce model and its counterpart open source implementation Hadoop, has proven itself as the de…

分布式、并行与集群计算 · 计算机科学 2014-08-04 Saeed Shahrivari , Saeed Jalili

We present a new challenge to examine whether large language models understand social norms. In contrast to existing datasets, our dataset requires a fundamental understanding of social norms to solve. Our dataset features the largest set…

计算与语言 · 计算机科学 2024-05-24 Ye Yuan , Kexin Tang , Jianhao Shen , Ming Zhang , Chenguang Wang

This Work in Progress Research paper departs from the recent, turbulent changes in global societies, forcing many citizens to re-skill themselves to (re)gain employment. Learners therefore need to be equipped with skills to be autonomous…

计算机与社会 · 计算机科学 2020-06-02 Mohammadreza Tavakoli , Ali Faraji , Stefan T. Mol , Gábor Kismihók

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root system that…

计算与语言 · 计算机科学 2024-02-29 Yang Liu , Jiahuan Cao , Chongyu Liu , Kai Ding , Lianwen Jin

In recent years, Large Language Models (LLMs) have emerged as transformative tools across numerous domains, impacting how professionals approach complex analytical tasks. This systematic mapping study comprehensively examines the…

计算机与社会 · 计算机科学 2025-08-19 Sai Sanjna Chintakunta , Nathalia Nascimento , Everton Guimaraes

Large datasets have become commonplace in NLP research. However, the increased emphasis on data quantity has made it challenging to assess the quality of data. We introduce Data Maps---a model-based tool to characterize and diagnose…

计算与语言 · 计算机科学 2020-10-16 Swabha Swayamdipta , Roy Schwartz , Nicholas Lourie , Yizhong Wang , Hannaneh Hajishirzi , Noah A. Smith , Yejin Choi

Standard Occupational Classifiers (SOC) are systems used to categorize and classify different types of jobs and occupations based on their similarities in terms of job duties, skills, and qualifications. Integrating these facets with Big…

计算与语言 · 计算机科学 2025-12-01 Sidharth Rony , Jack Patman

Job search through online matching engines nowadays are very prominent and beneficial to both job seekers and employers. But the solutions of traditional engines without understanding the semantic meanings of different resumes have not kept…

计算与语言 · 计算机科学 2016-07-27 Yiou Lin , Hang Lei , Prince Clement Addo , Xiaoyu Li

Fake job postings have become prevalent in the online job market, posing significant challenges to job seekers and employers. Despite the growing need to address this problem, there is limited research that leverages deep learning…

机器学习 · 计算机科学 2023-04-06 Aravind Sasidharan Pillai

Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent on building such large-scale datasets around the globe.…

As retrieval models converge on generic benchmarks, the pressing question is no longer "who scores higher" but rather "where do systems fail, and why?" Person-job matching is a domain that urgently demands such diagnostic capability -- it…

信息检索 · 计算机科学 2026-03-19 Guangzhi Wang , Xiaohui Yang , Kai Li , Jiawen He , Kai Yang , Ruixuan Zhang , Zhi Liu

As a representative information retrieval task, site recommendation, which aims at predicting the optimal sites for a brand or an institution to open new branches in an automatic data-driven way, is beneficial and crucial for brand…

信息检索 · 计算机科学 2023-07-04 Xinhang Li , Xiangyu Zhao , Yejing Wang , Yu Liu , Yong Li , Cheng Long , Yong Zhang , Chunxiao Xing

The emergence of new and disruptive technologies makes the economy and labor market more unstable. To overcome this kind of uncertainty and to make the labor market more comprehensible, we must employ labor market intelligence techniques,…

综合经济学 · 经济学 2024-09-04 Seyed Mohammad Ali Jafari , Ehsan Chitsaz

Efficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively. Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document…

计算与语言 · 计算机科学 2025-05-27 Ziliang Wang , Xuhui Zheng , Kang An , Cijun Ouyang , Jialu Cai , Yuhang Wang , Yichao Wu

In this paper we describe two bootstrap methods for massive data sets. Naive applications of common resampling methodology are often impractical for massive data sets due to computational burden and due to complex patterns of inhomogeneity.…

应用统计 · 统计学 2013-01-14 S. N. Lahiri , C. Spiegelman , J. Appiah , L. Rilett

Recently, due to the ubiquity and supremacy of E-recruitment platforms, job recommender systems have been largely studied. In this paper, we tackle the next job application problem, which has many practical applications. In particular, we…

信息检索 · 计算机科学 2021-11-23 Jun Zhu , Gautier Viaud , Céline Hudelot

Developing new ideas and algorithms in the fields of graph processing and relational learning requires public datasets. While Wikidata is the largest open source knowledge graph, involving more than fifty million entities, it is larger than…

机器学习 · 计算机科学 2019-10-07 Armand Boschin , Thomas Bonald