English
Related papers

Related papers: NVIDIA Nemotron Parse 1.1

200 papers

Table of contents (ToC) extraction aims to extract headings of different levels in documents to better understand the outline of the contents, which can be widely used for document understanding and information retrieval. Existing works…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Pengfei Hu , Zhenrong Zhang , Jianshu Zhang , Jun Du , Jiajia Wu

Implicit neural representations for videos (NeRV) have shown strong potential for video compression. However, applying NeRV to high-resolution 360-degree videos causes high memory usage and slow decoding, making real-time applications…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Daichi Arai , Kyohei Unno , Yasuko Sugito , Yuichi Kusakabe

Reading dense text and locating objects within images are fundamental abilities for Large Vision-Language Models (LVLMs) tasked with advanced jobs. Previous LVLMs, including superior proprietary models like GPT-4o, have struggled to excel…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Ya-Qi Yu , Minghui Liao , Jiwen Zhang , Jihao Wu

Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Expert (MoE)…

Multimedia · Computer Science 2026-03-09 Kin Wai Lau , Yasar Abbas Ur Rehman , Lai-Man Po , Pedro Porto Buarque de Gusmão

Many real-world applications involve the use of Optical Character Recognition (OCR) engines to transform handwritten images into transcripts on which downstream Natural Language Processing (NLP) models are applied. In this process, OCR…

Computation and Language · Computer Science 2021-07-16 Guowei Xu , Wenbiao Ding , Weiping Fu , Zhongqin Wu , Zitao Liu

Redundancy of visual tokens in multi-modal large language models (MLLMs) significantly reduces their computational efficiency. Recent approaches, such as resamplers and summarizers, have sought to reduce the number of visual tokens, but at…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yimu Wang , Mozhgan Nasr Azadani , Sean Sedwards , Krzysztof Czarnecki

Conventional Optical Character Recognition (OCR) systems are challenged by variant invoice layouts, handwritten text, and low-quality scans, which are often caused by strong template dependencies that restrict their flexibility across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Khushi Khanchandani , Advait Thakur , Akshita Shetty , Chaitravi Reddy , Ritisa Behera

Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jingyu Lei , Gaoang Wang , Der-Horng Lee

The resurgence of convolutional neural networks (CNNs) in visual recognition tasks, exemplified by ConvNeXt, has demonstrated their capability to rival transformer-based architectures through advanced training methodologies and ViT-inspired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Quan Bi Pay , Vishnu Monn Baskaran , Junn Yong Loo , KokSheik Wong , Simon See

Detecting and recognizing text in natural scene images is a challenging, yet not completely solved task. In re- cent years several new systems that try to solve at least one of the two sub-tasks (text detection and text recognition) have…

Computer Vision and Pattern Recognition · Computer Science 2017-07-28 Christian Bartz , Haojin Yang , Christoph Meinel

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer…

Computation and Language · Computer Science 2025-09-09 NVIDIA , : , Aaron Blakeman , Aarti Basant , Abhinav Khattar , Adithya Renduchintala , Akhiad Bercovich , Aleksander Ficek , Alexis Bjorlin , Ali Taghibakhshi , Amala Sanjay Deshmukh , Ameya Sunil Mahabaleshwarkar , Andrew Tao , Anna Shors , Ashwath Aithal , Ashwin Poojary , Ayush Dattagupta , Balaram Buddharaju , Bobby Chen , Boris Ginsburg , Boxin Wang , Brandon Norick , Brian Butterfield , Bryan Catanzaro , Carlo del Mundo , Chengyu Dong , Christine Harvey , Christopher Parisien , Dan Su , Daniel Korzekwa , Danny Yin , Daria Gitman , David Mosallanezhad , Deepak Narayanan , Denys Fridman , Dima Rekesh , Ding Ma , Dmytro Pykhtar , Dong Ahn , Duncan Riach , Dusan Stosic , Eileen Long , Elad Segal , Ellie Evans , Eric Chung , Erick Galinkin , Evelina Bakhturina , Ewa Dobrowolska , Fei Jia , Fuxiao Liu , Gargi Prasad , Gerald Shen , Guilin Liu , Guo Chen , Haifeng Qian , Helen Ngo , Hongbin Liu , Hui Li , Igor Gitman , Ilia Karmanov , Ivan Moshkov , Izik Golan , Jan Kautz , Jane Polak Scowcroft , Jared Casper , Jarno Seppanen , Jason Lu , Jason Sewall , Jiaqi Zeng , Jiaxuan You , Jimmy Zhang , Jing Zhang , Jining Huang , Jinze Xue , Jocelyn Huang , Joey Conway , John Kamalu , Jon Barker , Jonathan Cohen , Joseph Jennings , Jupinder Parmar , Karan Sapra , Kari Briski , Kateryna Chumachenko , Katherine Luna , Keshav Santhanam , Kezhi Kong , Kirthi Sivamani , Krzysztof Pawelec , Kumar Anik , Kunlun Li , Lawrence McAfee , Leon Derczynski , Lindsey Pavao , Luis Vega , Lukas Voegtle , Maciej Bala , Maer Rodrigues de Melo , Makesh Narsimhan Sreedhar , Marcin Chochowski , Markus Kliegl , Marta Stepniewska-Dziubinska , Matthieu Le , Matvei Novikov , Mehrzad Samadi , Michael Andersch , Michael Evans , Miguel Martinez , Mike Chrzanowski , Mike Ranzinger , Mikolaj Blaz , Misha Smelyanskiy , Mohamed Fawzy , Mohammad Shoeybi , Mostofa Patwary , Nayeon Lee , Nima Tajbakhsh , Ning Xu , Oleg Rybakov , Oleksii Kuchaiev , Olivier Delalleau , Osvald Nitski , Parth Chadha , Pasha Shamis , Paulius Micikevicius , Pavlo Molchanov , Peter Dykas , Philipp Fischer , Pierre-Yves Aquilanti , Piotr Bialecki , Prasoon Varshney , Pritam Gundecha , Przemek Tredak , Rabeeh Karimi , Rahul Kandu , Ran El-Yaniv , Raviraj Joshi , Roger Waleffe , Ruoxi Zhang , Sabrina Kavanaugh , Sahil Jain , Samuel Kriman , Sangkug Lym , Sanjeev Satheesh , Saurav Muralidharan , Sean Narenthiran , Selvaraj Anandaraj , Seonmyeong Bak , Sergey Kashirsky , Seungju Han , Shantanu Acharya , Shaona Ghosh , Sharath Turuvekere Sreenivas , Sharon Clay , Shelby Thomas , Shrimai Prabhumoye , Shubham Pachori , Shubham Toshniwal , Shyamala Prayaga , Siddhartha Jain , Sirshak Das , Slawek Kierat , Somshubra Majumdar , Song Han , Soumye Singhal , Sriharsha Niverty , Stefania Alborghetti , Suseella Panguluri , Swetha Bhendigeri , Syeda Nahida Akter , Szymon Migacz , Tal Shiri , Terry Kong , Timo Roman , Tomer Ronen , Trisha Saar , Tugrul Konuk , Tuomas Rintamaki , Tyler Poon , Ushnish De , Vahid Noroozi , Varun Singh , Vijay Korthikanti , Vitaly Kurin , Wasi Uddin Ahmad , Wei Du , Wei Ping , Wenliang Dai , Wonmin Byeon , Xiaowei Ren , Yao Xu , Yejin Choi , Yian Zhang , Ying Lin , Yoshi Suhara , Zhiding Yu , Zhiqi Li , Zhiyu Li , Zhongbo Zhu , Zhuolin Yang , Zijia Chen

The exponential growth of Large Multimodal Models (LMMs) has driven advancements in cross-modal reasoning but at significant computational costs. In this work, we focus on visual language models. We highlight the redundancy and inefficiency…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Yasmine Omri , Parth Shroff , Thierry Tambe

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful semantics. To this…

Multimedia · Computer Science 2023-05-05 Yuanyuan Liu , Haoyu Zhang , Yibing Zhan , Zijing Chen , Guanghao Yin , Lin Wei , Zhe Chen

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectural simplicity and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ziyao Wang , Chen Chen , Jingtao Li , Weiming Zhuang , Jiabo Huang , Ang Li , Lingjuan Lyu

Parsing sentences to linguistically-expressive semantic representations is a key goal of Natural Language Processing. Yet statistical parsing has focused almost exclusively on bilexical dependencies or domain-specific logical forms. We…

Computation and Language · Computer Science 2017-07-28 Jan Buys , Phil Blunsom

Computing at the edge offers intriguing possibilities for the development of autonomy and artificial intelligence. The advancements in autonomous technologies and the resurgence of computer vision have led to a rise in demand for fast and…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Martina Lofqvist , José Cano

Scientific knowledge is predominantly stored in books and scientific journals, often in the form of PDFs. However, the PDF format leads to a loss of semantic information, particularly for mathematical expressions. We propose Nougat (Neural…

Machine Learning · Computer Science 2023-08-28 Lukas Blecher , Guillem Cucurull , Thomas Scialom , Robert Stojnic

There has been a growing trend in compressing and transmitting videos from terminals for machine vision tasks. Nevertheless, most video coding optimization method focus on minimizing distortion according to human perceptual metrics,…

Multimedia · Computer Science 2025-12-18 Fei Zhao , Mengxi Guo , Shijie Zhao , Junlin Li , Li Zhang , Xiaodong Xie

This paper revisits the Neonatal Convolutional Neural Network (N-CNN) by optimizing its hyperparameters and evaluating how they affect its classification metrics, explainability and reliability, discussing their potential impact in clinical…

Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the…

Computation and Language · Computer Science 2026-01-22 Surapon Nonesung , Natapong Nitarach , Teetouch Jaknamon , Pittawat Taveekitworachai , Kunat Pipatanakul
‹ Prev 1 4 5 6 7 8 10 Next ›