kv_cache_decode
mediumruntimePython
results
Measured by wall-clock runtime in seconds — lower is better.
38.0s2.8s
Usage
Run the reference answer to verify your environment is set up correctly
$ harbor run -p tasks/kv_cache_decode
Test model
$ harbor run -p tasks/kv_cache_decode \ -a claude-code -m claude-opus-4-6
Description
Simulate autoregressive decoding of 256 tokens for a batch of 8 sequences. At each step the baseline concatenates Q, K, V for all past tokens and recomputes full multi-head attention from scratch. The reference appends only the new K, V to a persistent cache and computes single-row attention against the full cache. Output logits must match within 1e-4.
Files
path
permission
/app/solve.py✎ Edit/app/model_weights.npzRead-only/app/main.pyRead-onlyRules
- 01Edit /app/solve.py only.
- 02Allowed imports: numpy only. No PyTorch, JAX, or compiled extensions.
- 03Output logits must match reference within 1e-4 relative tolerance.
Tags
LLMKV-cacheattentioninferencetransformerdecoding