- [2026-06-10] DefTruth, Butterfingrz (2026). FFPA: Efficient Flash Prefill Attention for Large Head Dimensions via Split-D. Zenodo, 2026. 🎉🎉🎉
xlite-dev
Pinned Loading
Repositories
- ffpa-attn Public
Fast and Memory-Efficient Exact Attention for Large Headdim, 1.5x~6x speedup over PyTorch SDPA.
- diffusers Public Forked from huggingface/diffusers
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.
- cutlass Public Forked from NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
- LeetCUDA Public
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
- sglang Public Forked from sgl-project/sglang
SGLang is a fast serving framework for large language models and vision language models.
- CuTe Public Forked from NVlabs/CuTe
Reference implementation and examples of the CuTe Layout representation and algebra.
- Awesome-LLM-Inference Public
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
- .github Public
Top languages
Loading…
Most used topics
Loading…
