TBD
TBD
Workshop · Co-located with ICCAD 2026
The rapid growth of large language models (LLMs) is creating unprecedented demands on memory capacity, bandwidth, energy efficiency, and scalability. As LLM inference scales toward long-context reasoning, high-throughput serving, and edge deployment, memory access and data movement have become major system bottlenecks, exposing the limitations of conventional GPU-centric architectures. Emerging paradigms such as processing-in-memory (PIM), near-memory and near-storage computing, 3D-stacked PIM, emerging non-volatile memory (NVM)-based computing architectures, heterogeneous computing, and hardware–software co-design offer promising solutions for scalable and energy-efficient LLM inference.
This workshop brings together researchers and practitioners across architecture, hardware design, EDA, systems, and AI to discuss the future of memory-centric computing for LLM inference. Topics include 2D/3D PIM, near-memory and near-storage computing, emerging NVM-based computing architectures, heterogeneous CPU/GPU/NPU–PIM systems, memory hierarchy, data movement and KV-cache optimization, compiler and runtime support, hardware–software co-design, and EDA and design-space exploration for memory-centric AI. By fostering interactions across architecture, hardware design, memory, systems, EDA, and AI communities, the workshop aims to identify key challenges, emerging opportunities, and future research directions for next-generation memory-centric AI computing infrastructures.
TBD
TBD
TBD
TBD
TBD
TBD
TBD
TBD
TBD
Click any session with a speaker to view details.
Introduction and opening remarks
Break (15 min)
TBD