English
Related papers

Related papers: AICC: Parse HTML Finer, Make Models Better -- A 7.…

200 papers

The recently introduced Controlled Text Reduction (CTR) task isolates the text generation step within typical summarization-style tasks. It does so by challenging models to generate coherent text conforming to pre-selected content within…

Computation and Language · Computer Science 2024-02-27 Aviv Slobodkin , Avi Caciularu , Eran Hirsch , Ido Dagan

AIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Fan Yang , Ru Zhen , Jianing Wang , Yanhao Zhang , Haoxiang Chen , Haonan Lu , Sicheng Zhao , Guiguang Ding

Efficient distributed numerical word representation models (word embeddings) combined with modern machine learning algorithms have recently yielded considerable improvement on automatic document classification tasks. However, the…

Computation and Language · Computer Science 2018-09-07 Roger A. Stein , Patricia A. Jaques , Joao F. Valiati

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely on a few…

Information Retrieval · Computer Science 2022-08-30 Ritesh Sarkhel , Binxuan Huang , Colin Lockard , Prashant Shiralkar

Large Language Models improve with increasing amounts of high-quality training data. However, leveraging larger datasets requires balancing quality, quantity, and diversity across sources. After evaluating nine baseline methods under both…

Computation and Language · Computer Science 2025-01-27 William Held , Bhargavi Paranjape , Punit Singh Koura , Mike Lewis , Frank Zhang , Todor Mihaylov

Large language models (LLMs) have become integral to various real-world applications, leveraging massive, web-sourced datasets like Common Crawl, C4, and FineWeb for pretraining. While these datasets provide linguistic data essential for…

Computation and Language · Computer Science 2025-08-14 Sai Krishna Mendu , Harish Yenala , Aditi Gulati , Shanu Kumar , Parag Agrawal

Click-Through Rate (CTR) prediction is a crucial task in recommendation systems, online searches, and advertising platforms, where accurately capturing users' real interests in content is essential for performance. However, existing methods…

We explore using multilingual document embeddings for nearest neighbor mining of parallel data. Three document-level representations are investigated: (i) document embeddings generated by simply averaging multilingual sentence embeddings;…

Computation and Language · Computer Science 2019-07-02 Mandy Guo , Yinfei Yang , Keith Stevens , Daniel Cer , Heming Ge , Yun-Hsuan Sung , Brian Strope , Ray Kurzweil

Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understanding…

Artificial Intelligence · Computer Science 2025-07-16 Chao Deng , Jiale Yuan , Pi Bu , Peijie Wang , Zhong-Zhi Li , Jian Xu , Xiao-Hui Li , Yuan Gao , Jun Song , Bo Zheng , Cheng-Lin Liu

Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, we introduce an LLM-based line-level filtering method to…

Computation and Language · Computer Science 2025-01-14 Erik Henriksson , Otto Tarkka , Filip Ginter

For digital infrastructure to be safe, compatible, and standards-aligned, automated communication protocol compliance verification is crucial. Nevertheless, current rule-based systems are becoming less and less effective since they are…

Cryptography and Security · Computer Science 2026-04-07 Mohammad Wali Ur Rahman , Martin Manuel Lopez , Lamia Tasnim Mim , Carter Farthing , Julius Battle , Kathryn Buckley , Salim Hariri

Code optimization is a challenging task requiring a substantial level of expertise from developers. Nonetheless, this level of human capacity is not sufficient considering the rapid evolution of new hardware architectures and software…

Multimodal large language models (MLLMs) have streamlined front-end interface development by automating code generation. However, these models also introduce challenges in ensuring code quality. Existing approaches struggle to maintain both…

Software Engineering · Computer Science 2025-06-17 Yunnong Chen , Shixian Ding , YingYing Zhang , Wenkai Chen , Jinzhou Du , Lingyun Sun , Liuqing Chen

Multimedia or spoken content presents more attractive information than plain text content, but the former is more difficult to display on a screen and be selected by a user. As a result, accessing large collections of the former is much…

Computation and Language · Computer Science 2017-01-03 Wei Fang , Jui-Yang Hsu , Hung-yi Lee , Lin-Shan Lee

With the rapid development of Internet technology, people have more and more access to a variety of web page resources. At the same time, the current rapid development of deep learning technology is often inseparable from the huge amount of…

Information Retrieval · Computer Science 2022-10-27 Bowen Yu , Junping Du , Yingxia Shao

Designing pre-training objectives that more closely resemble the downstream tasks for pre-trained language models can lead to better performance at the fine-tuning stage, especially in the ad-hoc retrieval area. Existing pre-training…

Information Retrieval · Computer Science 2021-08-24 Zhengyi Ma , Zhicheng Dou , Wei Xu , Xinyu Zhang , Hao Jiang , Zhao Cao , Ji-Rong Wen

Temporal information extraction (TIE) has attracted a great deal of interest over the last two decades, leading to the development of a significant number of datasets. Despite its benefits, having access to a large volume of corpora makes…

Computation and Language · Computer Science 2023-11-27 Hugo Sousa , Alípio Jorge , Ricardo Campos

LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click, resize, or gameplay. Evaluation from screenshots can miss these failures, and filtering…

Software Engineering · Computer Science 2026-05-27 Jiajun Wu , Jian Yang , Tuney Zheng , Wei Zhang , Haowen Wang , Yihang Lou , Xianglong Liu

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliably distinguishing between human-authored and LLM-generated…

Computation and Language · Computer Science 2024-12-18 Zhen Tao , Yanfang Chen , Dinghao Xi , Zhiyu Li , Wei Xu

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential for real-world practice. While recent interactive benchmarks…

Artificial Intelligence · Computer Science 2026-01-21 Chuhan Qiao , Jianghua Huang , Daxing Zhao , Ziding Liu , Yanjun Shen , Bing Cheng , Wei Lin , Kai Wu
‹ Prev 1 8 9 10 Next ›