DFlash
Diffusion-based speculative decoding for faster LLM inference
Speculative-decoding inference acceleration software for large language models. It uses trained draft/speculator models and is integrated with serving and local-inference stacks including SGLang and llama.cpp to accelerate token generation without changing output.

Recent stories
0 linked stories
No linked stories yet.