中文
相关论文

相关论文: IPO-Mine: A Toolkit and Dataset for Section-Struct…

200 篇论文

Offline evaluation of recommender systems is often affected by hidden, under-documented choices in data preparation. Seemingly minor decisions in filtering, handling repeats, cold-start treatment, and splitting strategy design can…

信息检索 · 计算机科学 2026-02-24 Anna Volodkevich , Dmitry Anikin , Danil Gusak , Anton Klenitskiy , Evgeny Frolov , Alexey Vasilev

A lot of research has been devoted to identity documents analysis and recognition on mobile devices. However, no publicly available datasets designed for this particular problem currently exist. There are a few datasets which are useful for…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Vladimir V. Arlazarov , Konstantin Bulatov , Timofey Chernov , Vladimir L. Arlazarov

Most existing text summarization datasets are compiled from the news domain, where summaries have a flattened discourse structure. In such datasets, summary-worthy content often appears in the beginning of input articles. Moreover, large…

计算与语言 · 计算机科学 2019-06-11 Eva Sharma , Chen Li , Lu Wang

Large language models (LLMs) have shown impressive performance on general-purpose tasks, yet adapting them to specific domains remains challenging due to the scarcity of high-quality domain data. Existing data synthesis tools often struggle…

计算与语言 · 计算机科学 2025-07-08 Ziyang Miao , Qiyu Sun , Jingyuan Wang , Yuchen Gong , Yaowei Zheng , Shiqi Li , Richong Zhang

The broad goal of information extraction is to derive structured information from unstructured data. However, most existing methods focus solely on text, ignoring other types of unstructured data such as images, video and audio which…

计算与语言 · 计算机科学 2017-12-01 Robert L. Logan , Samuel Humeau , Sameer Singh

Within the Private Equity (PE) market, the event of a private company undertaking an Initial Public Offering (IPO) is usually a very high-return one for the investors in the company. For this reason, an effective predictive model for the…

机器学习 · 计算机科学 2019-11-26 Giuseppe C. Calafiore , Marisa H. Morales , Vittorio Tiozzo , Giulia Fracastoro , Serge Marquie

Data-driven materials discovery requires large-scale experimental datasets, yet most of the information remains trapped in unstructured literature. Existing extraction efforts often focus on a limited set of features and have not addressed…

计算与语言 · 计算机科学 2025-10-08 Xin Wang , Anshu Raj , Matthew Luebbe , Haiming Wen , Shuozhi Xu , Kun Lu

India's Right to Information Act, 2005 gives every citizen the right to demand information from public authorities, yet in practice most people cannot make sense of the dense administrative language used in Central Information Commission…

计算与语言 · 计算机科学 2026-05-19 Joy Bose

This technical report introduces Docling, an easy to use, self-contained, MIT-licensed open-source package for PDF document conversion. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table…

Data profiling plays a critical role in understanding the structure of complex datasets and supporting numerous downstream tasks, such as social media analytics and financial fraud detection. While existing research predominantly focuses on…

人机交互 · 计算机科学 2025-03-11 Yanwei Huang , Yan Miao , Di Weng , Adam Perer , Yingcai Wu

Improving the accessibility and automation capabilities of mobile devices can have a significant positive impact on the daily lives of countless users. To stimulate research in this direction, we release a human-annotated dataset with…

人机交互 · 计算机科学 2022-10-07 Srinivas Sunkara , Maria Wang , Lijuan Liu , Gilles Baechler , Yu-Chung Hsiao , Jindong , Chen , Abhanshu Sharma , James Stout

Visually rich documents (e.g. leaflets, banners, magazine articles) are physical or digital documents that utilize visual cues to augment their semantics. Information contained in these documents are ad-hoc and often incomplete. Existing…

机器学习 · 计算机科学 2024-04-02 Ritesh Sarkhel , Arnab Nandi

High-quality and open datasets remain a major bottleneck for text-to-image (T2I) fine-tuning. Despite rapid progress in model architectures and training pipelines, most publicly available fine-tuning datasets suffer from low resolution,…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Xu Ma , Yitian Zhang , Qihua Dong , Yun Fu

Intellectual property protection(IPP) have received more and more attention recently due to the development of the global e-commerce platforms. brand recognition plays a significant role in IPP. Recent studies for brand recognition and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Xuan Jin , Wei Su , Rong Zhang , Yuan He , Hui Xue

Context: When software is released publicly, it is common to include with it either the full text of the license or licenses under which it is published, or a detailed reference to them. Therefore public licenses, including FOSS (free, open…

软件工程 · 计算机科学 2023-08-23 Jesús M. González-Barahona , Sergio Montes-Leon , Gregorio Robles , Stefano Zacchiroli

Generating professional financial reports is a labor-intensive and intellectually demanding process that current AI systems struggle to fully automate. To address this challenge, we introduce FinSight (Financial InSight), a novel multi…

计算与语言 · 计算机科学 2025-10-21 Jiajie Jin , Yuyao Zhang , Yimeng Xu , Hongjin Qian , Yutao Zhu , Zhicheng Dou

Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, they remain…

人工智能 · 计算机科学 2026-05-29 Bangbang Zhou , Hangdi Xing , Yifan Chen , Jianjun Xu , Qi Zheng , Feiyu Gao , Zhibo Yang , Shuai Bai , Ming Yan , Jieping Ye , Hongtao Xie

We focus on electronic theses and dissertations (ETDs), aiming to improve access and expand their utility, since more than 6 million are publicly available, and they constitute an important corpus to aid research and education across…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Sampanna Yashwant Kahu , William A. Ingram , Edward A. Fox , Jian Wu

From small screenshots to large videos, documents take up a bulk of space in a modern smartphone. Documents in a phone can accumulate from various sources, and with the high storage capacity of mobiles, hundreds of documents are accumulated…

计算机视觉与模式识别 · 计算机科学 2021-01-07 Sugam Garg , Harichandana , Sumit Kumar

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

计算与语言 · 计算机科学 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec