flash_attention

hardruntimeC
results

Measured by wall-clock runtime in seconds — lower is better.

0.75s0.10s
sota
Claude-Opus-4.6
reward 0.850

Usage

Run the reference answer to verify your environment is set up correctly

$ harbor run -p tasks/flash_attention

Test model

$ harbor run -p tasks/flash_attention \
  -a claude-code -m claude-opus-4-6

Description

Implement an efficient scaled dot-product attention computation (output = softmax(QK^T / sqrt(d)) V) for sequence length 4096 and head dimension 64 in float32. The naive approach materializes a full 128 MB score matrix that blows out the CPU cache, so the key challenge is computing attention in a tiled, memory-efficient manner.

Files

path
permission
/app/solve.c✎ Edit
/app/solve.hRead-only
/app/main.cRead-only
/app/MakefileRead-only

Rules

  • 01Edit /app/solve.c only.
  • 02SIMD intrinsics (<immintrin.h>) are allowed.
  • 03No external libraries. Single-threaded only.
  • 04Wrong results score 0.

Tags

attentiontilingAVX2online-softmax

Model Results

Click a row to view its trajectory in Live Lab

model
reward
score
Claude-Opus-4.6
0.850
Kimi-K2.6
0.650
GLM-5
0.570
Hunyuan-3-Preview
0.480
Grok-4-20
0.440
MiMo-V2.5-Pro
0.430
MiniMax-M2.7
0.430
Gemini-3.1-Pro
0.390
Qwen-3.6-Plus
0.350
DeepSeek-V4-Pro
0.330
GPT-5.4
0.320