Out-of-tree vLLM plugin and open format spec for NVFP4-CB / FP8-CB product-codebook weights — 2-6 bit-per-weight LLM quantization served on native Blackwell tensor cores.
triton moe cuda-kernels quantization vector-quantization blackwell fp8 llm-serving vllm llm-inference fp4 nvfp4 dgx-spark codebook-quantization
-
Updated
Aug 5, 2026 - Python