Paper title: SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures. We propose SPIMOE, a reasoning-aware algorithm-architecture co-design framework that exploits hybrid sparsity for efficient long-reasoning Mixture-of-Experts (MoE) inference on heterogeneous PIM architectures.
SPIMOE combines adaptive expert routing and block-sparse attention with physical KV-cache eviction, and maps Attention and MoE FFN onto heterogeneous SRAM-PIM and HBM-PIM substrates. It further uses static expert mapping and dynamic sub-batch scheduling to balance workloads and overlap execution. Evaluations report up to 8.35x speedup over an NVIDIA A100 GPU and 3.33x over PIMoE while maintaining reasoning accuracy comparable to full-precision baselines.
Paper title: Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing. We propose Focus-dLLM for long-context diffusion LLM inference. It uses confidence-guided context focusing and sink-aware pruning to reduce redundant bidirectional attention without retraining.
GitHub: Longxmas/Focus-dLLM
Paper title: SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning. We propose SlimInfer for long-context LLM inference. It uses dynamic block-wise token pruning and a predictor-free asynchronous KV cache manager to reduce prefill latency, memory pressure, and I/O overhead.
GitHub: Longxmas/SlimInfer
Paper title: Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms. We introduce A3GNN for adaptive GNN training on CPU-GPU platforms. It combines locality-aware sampling and fine-grained scheduling to balance throughput, memory footprint, and training quality on heterogeneous systems.
GitHub: BUAA-CI-LAB/A3GNN
Paper title: Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design. We introduce Finesse, a software-hardware co-design framework for pairing-based cryptography. It integrates compiler support, simulation, and parameterized pipelined hardware to speed up accelerator design and execution.
GitHub: BUAA-CI-LAB/Finesse