中文
相关论文

相关论文: Evaluating Progress in Web3 Grants: Introducing th…

200 篇论文

Citation metrics are analytic measures used to evaluate the usage, impact and dissemination of scientific research. Traditionally, citation metrics have been independently measured at each level of the publication pyramid, namely at the…

数字图书馆 · 计算机科学 2017-07-04 Jacques Balayla

Rigorous evaluation of domain-specific language models requires benchmarks that are comprehensive, contamination-resistant, and maintainable. Static, manually curated datasets do not satisfy these properties. We present a graph-based…

人工智能 · 计算机科学 2026-05-18 Jessica M. Lundin , Usman Nasir Nakakana , Guillaume Chabot-Couture

We explore the evolving efficacy of three generative pre-trained transformer (GPT) models in generating answers for multiple-choice questions (MCQ) from introductory and intermediate Python programming courses in higher education. We focus…

计算机与社会 · 计算机科学 2023-11-17 Jaromir Savelka , Arav Agarwal , Christopher Bogart , Majd Sakr

The Web3 ecosystem is highly fragmented, making seamless integration difficult for over a billion Web2 businesses, enterprises, and AI protocols. As blockchains, rollups, and app-specific chains expand, cross-chain interactions remain…

密码学与安全 · 计算机科学 2025-03-21 Hardik Gajera , Akhil Reddy , Bhagath Reddy

As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to…

Multimodal Large Language models (MLLMs) have shown promise in web-related tasks, but evaluating their performance in the web domain remains a challenge due to the lack of comprehensive benchmarks. Existing benchmarks are either designed…

计算与语言 · 计算机科学 2024-04-10 Junpeng Liu , Yifan Song , Bill Yuchen Lin , Wai Lam , Graham Neubig , Yuanzhi Li , Xiang Yue

Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high-quality training data while maintaining annotation…

This paper presents a novel approach named Persona-Grouping-Intelligence (PGI), which has been crafted to tackle the challenges posed by GPT models when applied to real-world business issues. PGI leverages the inherent capabilities of the…

人工智能 · 计算机科学 2023-08-28 Aline Ioste

Research grants have played an important role in seeding and promoting fundamental research projects worldwide. There is a growing demand for developing and delivering scientific influence analysis as a service on research grant…

社会与信息网络 · 计算机科学 2019-08-26 Yuming Wang , Yanbo Long , Lai Tu , Ling Liu

Text-to-Image (TTI) systems often support people during ideation, the early stages of a creative process when exposure to a broad set of relevant images can help explore the design space. Since ideation is an important subclass of TTI…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Negar Arabzadeh , Fernando Diaz , Junfeng He

Social platforms, and the online communities that use them, are evolving at a rapid pace. As a result, research and development regarding how to moderate online communities is being out-paced. In this paper, we present a novel framework…

人机交互 · 计算机科学 2022-06-27 Tanvi Bajpai , Drshika Asher , Anwesa Goswami , Eshwar Chandrasekharan

Modern applications increasingly interact with web APIs -- reusable components, deployed and operated outside the application, and accessed over the network. Their existence, arguably, spurs application innovations, making it easy to…

软件工程 · 计算机科学 2020-07-07 David Bermbach , Erik Wittern

The rapid emergence of Large Language Models (LLMs) presents both opportunities and challenges for programming education. While students increasingly use generative AI tools, direct access often hinders the learning process by providing…

人工智能 · 计算机科学 2026-03-31 Thomas Van Mullem , Bart Mesuere , Peter Dawyndt

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains…

计算与语言 · 计算机科学 2024-01-09 Vahid Ashrafimoghari , Necdet Gürkan , Jordan W. Suchow

Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting have eroded the authority of standard benchmarks, leaving the…

人工智能 · 计算机科学 2026-02-13 Longyuan Zhu , Hairan Hua , Linlin Miao , Bing Zhao

In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed…

As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and…

计算与语言 · 计算机科学 2025-12-09 Xiaojun Wu , Junxi Liu , Huanyi Su , Zhouchi Lin , Yiyan Qi , Chengjin Xu , Jiajun Su , Jiajie Zhong , Fuwei Wang , Saizhuo Wang , Fengrui Hua , Jia Li , Jian Guo

Web3 is leading a wave of the next generation of web services that even many Web2 applications are keen to ride. However, the lack of Web3 background for Web2 developers hinders easy and effective access and transition. On the other hand,…

软件工程 · 计算机科学 2023-10-17 Guangsheng Yu , Xu Wang , Qin Wang , Tingting Bi , Yifei Dong , Ren Ping Liu , Nektarios Georgalas , Andrew Reeves

Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape, and it is becoming clear that the quality of automatic evaluation metrics is not keeping up with the pace of development of generative models. We aim to improve…

计算与语言 · 计算机科学 2023-10-24 Andrea Sottana , Bin Liang , Kai Zou , Zheng Yuan

Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on benchmarks and fixed datasets, frequently failing to reflect…