FlashQLA
High-performance linear attention kernels built on TileLang.
A TileLang-based stack of high-performance linear-attention kernels, intended to accelerate forward and backward passes for AI deployments, including edge and long-context use cases.

Recent stories
1 linked story