speculative_decode

mediumruntimePython
results

Measured by wall-clock runtime in seconds — lower is better.

52.0s4.5s

Usage

Run the reference answer to verify your environment is set up correctly

$ harbor run -p tasks/speculative_decode

Test model

$ harbor run -p tasks/speculative_decode \
  -a claude-code -m claude-opus-4-6

Description

Implement speculative decoding for 512 generation steps. A fast draft model (2-layer, d=512) proposes 8 candidate tokens. The target model (8-layer, d=2048) must verify which prefix to accept. The baseline runs the target model autoregressively on each draft token. The reference processes all 8 draft tokens in a single batched forward pass and uses rejection sampling to determine the acceptance length.

Files

path
permission
/app/solve.py✎ Edit
/app/main.pyRead-only
/app/draft_model.npzRead-only
/app/target_model.npzRead-only

Rules

  • 01Edit /app/solve.py only.
  • 02Allowed imports: numpy only.
  • 03Generated text must match reference (identical acceptance decisions). Wrong results score 0.

Tags

LLMspeculative-decodingdraft-modelverificationinference