flash_attention
hardruntimeC
results
Measured by wall-clock runtime in seconds — lower is better.
0.75s0.10s
sota
Usage
Run the reference answer to verify your environment is set up correctly
$ harbor run -p tasks/flash_attention
Test model
$ harbor run -p tasks/flash_attention \ -a claude-code -m claude-opus-4-6
Description
Implement an efficient scaled dot-product attention computation (output = softmax(QK^T / sqrt(d)) V) for sequence length 4096 and head dimension 64 in float32. The naive approach materializes a full 128 MB score matrix that blows out the CPU cache, so the key challenge is computing attention in a tiled, memory-efficient manner.
Files
path
permission
/app/solve.c✎ Edit/app/solve.hRead-only/app/main.cRead-only/app/MakefileRead-onlyRules
- 01Edit /app/solve.c only.
- 02SIMD intrinsics (<immintrin.h>) are allowed.
- 03No external libraries. Single-threaded only.
- 04Wrong results score 0.
Tags
attentiontilingAVX2online-softmax
Model Results
Click a row to view its trajectory in Live Lab
model
reward
score
0.850
Kimi-K2.6
0.650
0.570
Hunyuan-3-Preview
0.480
Grok-4-20
0.440
MiMo-V2.5-Pro
0.430
0.430
0.390
Qwen-3.6-Plus
0.350
DeepSeek-V4-Pro
0.330
0.320