FlashMLA
Efficient Multi-head Latent Attention Kernels
FlashMLA is DeepSeek's open-source library of optimized attention kernels for accelerating DeepSeek model inference/training workloads, including dense and sparse MLA/MHA kernels for prefill and decoding on NVIDIA GPU architectures.

Recent stories
0 linked stories
No linked stories yet.