Distributed, Parallel, and Cluster Computing · Computer Science
Active Access: A Mechanism for High-Performance Distributed Data-Centric Computations
Maciej Besta, Torsten Hoefler
2020-10-20
Distributed, Parallel, and Cluster Computing · Computer Science
Fault Tolerance for Remote Memory Access Programming Models
Maciej Besta, Torsten Hoefler
2020-10-20
Distributed, Parallel, and Cluster Computing · Computer Science
Leveraging MPI RMA to optimise halo-swapping communications in MONC on Cray machines
Nick Brown, Michael Bareford, Michèle Weiland
2020-10-27
Emerging Technologies · Computer Science
Machine Unlearning and Continual Learning in Hybrid Resistive Memory Neuromorphic Systems
Ning Lin, Jichang Yang, Yangu He, Zijian Ye +11
2026-02-10
Distributed, Parallel, and Cluster Computing · Computer Science
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
Anderson de Lima Luiz, Shubham Vijay Kurlekar, Munir Georges
2025-08-26
Artificial Intelligence · Computer Science
MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents
Zeyu Zhang, Quanyu Dai, Xu Chen, Rui Li +2
2025-05-06
Distributed, Parallel, and Cluster Computing · Computer Science
Dynamic reconfiguration for malleable applications using RMA
Iker Martín-Álvarez, José I. Aliaga, Maribel Castillo
2026-01-28
Performance · Computer Science
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
George Karfakis, Faraz Tahmasebi, Binglu Chen, Lime Yao +4
2025-12-23
Machine Learning · Computer Science
Deep Unrolling Networks with Recurrent Momentum Acceleration for Nonlinear Inverse Problems
Qingping Zhou, Jiayu Qian, Junqi Tang, Jinglai Li
2024-04-02
Machine Learning · Computer Science
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
Kehao Zhang, Shangtong Gui, Sheng Yang, Wei Chen +1
2026-02-24
Distributed, Parallel, and Cluster Computing · Computer Science
Enabling Highly-Scalable Remote Memory Access Programming with MPI-3 One Sided
Robert Gerstenberger, Maciej Besta, Torsten Hoefler
2020-07-01
Hardware Architecture · Computer Science
High Area/Energy Efficiency RRAM CNN Accelerator with Kernel-Reordering Weight Mapping Scheme Based on Pattern Pruning
Songming Yu, Yongpan Liu, Lu Zhang, Jingyu Wang +4
2020-10-14
Distributed, Parallel, and Cluster Computing · Computer Science
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
Tomer Bitan, Tal Kadosh, Erel Kaplan, Shira Meiri +4
2025-09-16
Programming Languages · Computer Science
A Verified High-Performance Composable Object Library for Remote Direct Memory Access (Extended Version)
Guillaume Ambal, George Hodgkins, Mark Madler, Gregory Chockler +4
2025-10-14
Distributed, Parallel, and Cluster Computing · Computer Science
Storm: a fast transactional dataplane for remote data structures
Stanko Novakovic, Yizhou Shan, Aasheesh Kolli, Michael Cui +6
2019-02-08
Distributed, Parallel, and Cluster Computing · Computer Science
Quo Vadis MPI RMA? Towards a More Efficient Use of MPI One-Sided Communication
Joseph Schuchart, Christoph Niethammer, José Gracia, George Bosilca
2021-11-17
Distributed, Parallel, and Cluster Computing · Computer Science
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
Dong Xu, Han Meng, Xinyu Chen, Dengcheng Zhu +10
2026-05-08
Hardware Architecture · Computer Science
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
Shangyi Shi, Husheng Han, Jianan Mu, Xinyao Zheng +6
2025-04-01
Distributed, Parallel, and Cluster Computing · Computer Science
Enabling Scientific Workflow Scheduling Research in Non-Uniform Memory Access Architectures
Aurelio Vivas, Harold Castro
2025-11-26
Hardware Architecture · Computer Science
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
Dowon Kim, MinJae Lee, Janghyeon Kim, HyuckSung Kwon +7
2025-11-04
Machine Learning · Computer Science
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Bo Wu, Sid Wang, Yunhao Tang, Jia Ding +10
2025-06-03
Robotics · Computer Science
PRAM-R: A Perception-Reasoning-Action-Memory Framework with LLM-Guided Modality Routing for Adaptive Autonomous Driving
Yi Zhang, Xian Zhang, Saisi Zhao, Yinglei Song +3
2026-03-05