中文
相关论文

相关论文: Open4Business(O4B): An Open Access Dataset for Sum…

200 篇论文

Summarizing book-length documents (>100K tokens) that exceed the context window size of large language models (LLMs) requires first breaking the input document into smaller chunks and then prompting an LLM to merge, update, and compress…

计算与语言 · 计算机科学 2024-04-16 Yapei Chang , Kyle Lo , Tanya Goyal , Mohit Iyyer

Text summarization aims to condense long documents and retain key information. Critical to the success of a summarization model is the faithful inference of latent representations of words or tokens in the source documents. Most recent…

计算与语言 · 计算机科学 2022-03-16 Bo Pang , Erik Nijkamp , Wojciech Kryściński , Silvio Savarese , Yingbo Zhou , Caiming Xiong

Due to its promise to alleviate information overload, text summarization has attracted the attention of many researchers. However, it has remained a serious challenge. Here, we first prove empirical limits on the recall (and F1-scores) of…

计算与语言 · 计算机科学 2018-03-23 Rakesh Verma , Daniel Lee

Pretraining has proven successful in Document Intelligence tasks where deluge of documents are used to pretrain the models only later to be finetuned on downstream tasks. One of the problems of the pretraining approaches is the inconsistent…

计算机视觉与模式识别 · 计算机科学 2022-03-01 Ali Furkan Biten , Rubèn Tito , Lluis Gomez , Ernest Valveny , Dimosthenis Karatzas

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference…

计算与语言 · 计算机科学 2025-10-29 Frederik Broy , Maike Züfle , Jan Niehues

Multi-document summarization is the process of automatically generating a concise summary of multiple documents related to the same topic. This summary can help users quickly understand the key information from a large collection of…

计算与语言 · 计算机科学 2023-12-20 Charles Rajan , Nishit Asnani , Shreya Singh

In the rapidly evolving field of artificial intelligence (AI), mapping innovation patterns and understanding effective technology transfer from research to applications are essential for economic growth. However, existing data…

数据库 · 计算机科学 2025-06-02 Haixing Gong , Hui Zou , Xingzhou Liang , Shiyuan Meng , Pinlong Cai , Xingcheng Xu , Jingjing Qu

Accurate extraction of key information from 2D engineering drawings is crucial for high-precision manufacturing. Manual extraction is slow and labor-intensive, while traditional Optical Character Recognition (OCR) techniques often struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Muhammad Tayyab Khan , Zane Yong , Lequn Chen , Jun Ming Tan , Wenhe Feng , Seung Ki Moon

The business model represents an increasingly important management concept. However, progress in research related to the concept is currently inhibited from inconsistencies in terms of formalizing and therewith also empirically measuring…

计算机与社会 · 计算机科学 2015-03-05 Fredrik Hacklin , Nobuaki Minato , Toma Kobayashi

In this study, we introduce Orion-14B, a collection of multilingual large language models with 14 billion parameters. We utilize a data scheduling approach to train a foundational model on a diverse corpus of 2.5 trillion tokens, sourced…

计算与语言 · 计算机科学 2024-01-24 Du Chen , Yi Huang , Xiaopu Li , Yongqiang Li , Yongqiang Liu , Haihui Pan , Leichao Xu , Dacheng Zhang , Zhipeng Zhang , Kun Han

The recognition of dataset names is a critical task for automatic information extraction in scientific literature, enabling researchers to understand and identify research opportunities. However, existing corpora for dataset mention…

计算与语言 · 计算机科学 2023-10-06 Huitong Pan , Qi Zhang , Eduard Dragut , Cornelia Caragea , Longin Jan Latecki

Legal documents are often long, dense, and difficult to comprehend, not only for laypeople but also for legal experts. While automated document summarization has great potential to improve access to legal knowledge, prevailing task-based…

计算与语言 · 计算机科学 2026-03-24 Tsz Fung Pang , Maryam Berijanian , Thomas Orth , Breanna Shi , Charlotte S. Alexander

Abstractive document summarization is usually modeled as a sequence-to-sequence (Seq2Seq) learning problem. Unfortunately, training large Seq2Seq based summarization models on limited supervised summarization data is challenging. This paper…

计算与语言 · 计算机科学 2020-10-13 Yanyan Zou , Xingxing Zhang , Wei Lu , Furu Wei , Ming Zhou

In Open Source Software, the source code and any other resources available in a project can be viewed or reused by anyone subject to often permissive licensing restrictions. In contrast to some studies of dependency-based reuse supported…

软件工程 · 计算机科学 2024-02-12 Mahmoud Jahanshahi , Audris Mockus

To ensure the fairness and trustworthiness of machine learning (ML) systems, recent legislative initiatives and relevant research in the ML community have pointed out the need to document the data used to train ML models. Besides,…

机器学习 · 计算机科学 2024-12-18 Joan Giner-Miguelez , Abel Gómez , Jordi Cabot

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a…

计算与语言 · 计算机科学 2025-06-03 Yifan Peng , Shakeel Muhammad , Yui Sudo , William Chen , Jinchuan Tian , Chyi-Jiunn Lin , Shinji Watanabe

The large volumes of structured data currently available, from Web tables to open-data portals and enterprise data, open up new opportunities for progress in answering many important scientific, societal, and business questions. However,…

信息检索 · 计算机科学 2021-09-01 Sonia Castelo , Rémi Rampin , Aécio Santos , Aline Bessa , Fernando Chirigati , Juliana Freire

We present PeerSum, a new MDS dataset using peer reviews of scientific publications. Our dataset differs from the existing MDS datasets in that our summaries (i.e., the meta-reviews) are highly abstractive and they are real summaries of the…

信息检索 · 计算机科学 2022-09-30 Miao Li , Jianzhong Qi , Jey Han Lau

This paper presents Summary Workbench, a new tool for developing and evaluating text summarization models. New models and evaluation measures can be easily integrated as Docker-based plugins, allowing to examine the quality of their…

计算与语言 · 计算机科学 2022-10-19 Shahbaz Syed , Dominik Schwabe , Martin Potthast

Search agents, which integrate language models (LMs) with web search, are becoming crucial for answering complex user queries. Constructing training datasets for deep research tasks, involving multi-step retrieval and reasoning, remains…

计算与语言 · 计算机科学 2026-04-03 Nandan Thakur , Zijian Chen , Xueguang Ma , Jimmy Lin
‹ 上一页 1 8 9 10 下一页 ›