Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors export, validated on Qwen3.6-35B-A3B MoE with MTP speculative decoding.
// readme
PrismaQuant
Mixed-precision LLM quantization that chooses the right format for every weight matrix, selected on real end-to-end KL — shipped as artifacts that stock inference engines serve on an unforked runtime.
PrismaQuant’s allocator is AURA (Production-Faithful KL–Fisher Allocation): a per-Linear cost model built from KL-Fisher probes of the full model multiplied against the production-rendered weight error — the exact bytes that ship — solved as a multi-choice knapsack, and gated by real KL measured on a held-out split before anything is published. Three output containers:
compressed-tensors— vanilla vLLM serves it natively (vllm serve $WORK_DIR/exported), no custom kernels. NVFP4 / FP8 / BF16 per Linear, CUTLASS kernels on Blackwell.- GGUF — llama.cpp and vLLM (via the GGUF plugin) serve the same file, again with no PrismaQuant kernels. Full k-quant + IQ menu, per-tensor mixed, imatrix-weighted. This is how a 295B MoE fits on one 128 GB box.
- Tessera — the Tessera trellis wire, served by Tessera’s own out-of-tree vLLM plugin (
tessera.serving,quant_method = "tessera"), using native kernels only. Still an unforked vLLM: install…