KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

Department of Electrical Engineering, POSTECH, South Korea.
NeurIPS 2026
*Corresponding author
KV-COBRA overview

(a) The spectrum sets (r*, b*): one head against the layer's 7 others. (b) Heads differ: per-head distortion at a uniform 2 bpd spans 100× across 256 KV heads. (c) The Hadamard rotation flattens the per-channel variance of the retained subspace.

TL;DR. The bottleneck of low-bit KV compression is allocation across heads, not the kernel itself. KV-COBRA picks the retained rank and the bit-width of every head jointly, redistributes the budget across heads, and reuses the standard rotate-and-quantize kernel. It stays accurate down to 0.5 bits per dimension with no per-token overhead.
100×
spread of per-head distortion at a uniform 2 bpd across the 256 KV heads of LLaMA-3.1-8B
0.5 bpd
KV-COBRA stays accurate where uniform-allocation baselines collapse
1.8×
faster attention primitive than FP16 SDPA at r = 16, from 32k to 128k tokens
0
per-token overhead: the C1+C2 allocation runs once at calibration, in under 200 ms

Abstract

What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head—using only standard low-rank projection and scalar quantization—dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint K+V compression. On perplexity, zero-shot, and long-context benchmarks from 0.5 to 4 bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.

How KV-COBRA Works

Rank r and bit-width b are usually global hyperparameters, yet heads differ widely in how their key spectrum concentrates. A single (r, b) wastes bits on concentrated heads and starves diffuse ones, and the two choices are coupled. KV-COBRA treats this as a rate-distortion allocation problem with the per-head distortion

D(r, b) = ∑i>r wi + q(b) ∑i≤r wi,   q(b) = 2−2b/12

KV-COBRA pipeline

Calibration. Per-head SVD of prefill keys; a two-level C1+C2 allocator freezes the basis, rank r* and bit-width b* of every head. Inference. Each key is rotated and truncated with H·VT and quantized with its head's b*-bit quantizer. The kernel is shared by all heads; only the allocation changes.

C1, per head. For a head budget B, enumerate the even ranks r, set b ≈ B/r clipped to [2, 8], and keep the pair with the smallest D, an O(d/2) search. At the optimum, the projection loss saved by two more directions matches the quantization loss added on the retained ones.

C2, across heads. With C1 as an oracle, move budget between heads until their marginal distortions equalize, the water-filling condition of rate-distortion theory. A damped update converges in five rounds.

Hadamard rotation. A Walsh–Hadamard transform with a fixed random sign flip, folded into the projection, flattens the steeply decaying variances of the retained coordinates, so one uniform quantizer per head suffices.

Attention-KL reordering. A second-order bound on the softmax KL gives each direction the weight wi = σ2Q,iσ2K,i instead of the key eigenvalue alone. Sorting the basis by this weight before C1/C2 gives KV-COBRAKL, the main method. With wi = λi the same solver gives KV-COBRAMSE.

Results

LLaMA-3.1-8B Mistral-7B-v0.3 Qwen2.5-7B-Instruct
1 bit / dim FP16MSEKL FP16MSEKL FP16MSEKL
Perplexity (↓) 7.6724.7620.26 13.1527.1815.90 9.5012.9713.20
Zero-shot acc. (↑) 55.451.551.6 56.354.655.1 54.151.752.5
LongBench F1 (↑) 15.809.2511.36 11.549.5510.72 10.649.109.53

Keys compressed pre-RoPE to 1 bit per dimension, mean over 5 seeds. MSE = KV-COBRAMSE, KL = KV-COBRAKL. Perplexity averages WikiText-2, PTB and C4; zero-shot averages ARC-C, HellaSwag, PIQA, WinoGrande and MMLU; LongBench averages five QA tasks. Bold marks the better variant.

Results across bit rates

Main k-only results from 0.5 to 4 bits per dimension against KIVI, KVQuant, GEAR, TurboQuant, SVDq and KQ-SVD. Dashed lines mark FP16; error bars are ±1 SD over 5 seeds.

Above 2 bpd the allocator barely matters. Below it, the methods separate: integer-only quantizers cannot reach the regime, uniform-budget SVD baselines degrade sharply, and KV-COBRAKL degrades most gracefully, nearly doubling the LongBench F1 of the integer-only baselines at a nominal 1-bpd budget. KV-COBRAMSE wins on perplexity and KV-COBRAKL on LongBench, the split the two objectives predict. The pattern carries over to LLaMA-3.1-70B, Qwen2.5-72B and Mixtral-8x7B.

No Single (r, b) Fits All Heads

Per-head rank and bit-width chosen by C1

Per-head (r*, b*) selected by C1 on three models, colored by target bpd.

The optima spread over many operating points and shift with the budget and the model. This is the freedom a global hyperparameter throws away, and the reason per-head allocation pays off most at low bit rates.

Efficiency

Context r = 16, b = 2 r = 16, b = 4 r = 32, b = 2 r = 64, b = 2
32k1.83×1.84×1.77×1.40×
64k1.79×1.85×1.82×1.48×
128k1.89×1.88×1.85×1.49×

Per-layer attention primitive vs. FP16 SDPA (Hq = 32, d = 128, CUDA graphs, RTX 6000 Ada).

The cache is stored as packed b-bit codes with the rotation folded into the basis, so dequantization and the inverse rotation run in one fused pass. All allocation work happens once: the C1+C2 solve takes under 200 ms on 7–8B models, and the full calibration of LLaMA-3.1-70B takes 18.3 s.

BibTeX

@article{ha2026kv,
  title={KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation},
  author={Ha, Sihyeon and Lee, Jaeho and Jeon, Yo-Seb},
  journal={arXiv preprint arXiv:2609.24298},
  year={2026}
}