kv_cache_decode

mediumruntimePython
results

Measured by wall-clock runtime in seconds — lower is better.

38.0s2.8s

Usage

Run the reference answer to verify your environment is set up correctly

$ harbor run -p tasks/kv_cache_decode

Test model

$ harbor run -p tasks/kv_cache_decode \
  -a claude-code -m claude-opus-4-6

Description

Simulate autoregressive decoding of 256 tokens for a batch of 8 sequences. At each step the baseline concatenates Q, K, V for all past tokens and recomputes full multi-head attention from scratch. The reference appends only the new K, V to a persistent cache and computes single-row attention against the full cache. Output logits must match within 1e-4.

Files

path
permission
/app/solve.py✎ Edit
/app/model_weights.npzRead-only
/app/main.pyRead-only

Rules

  • 01Edit /app/solve.py only.
  • 02Allowed imports: numpy only. No PyTorch, JAX, or compiled extensions.
  • 03Output logits must match reference within 1e-4 relative tolerance.

Tags

LLMKV-cacheattentioninferencetransformerdecoding