scaling_law

mediumperplexityPython
results

Measured by test-set perplexity — lower is better.

95.023.0
sota
Claude-Opus-4.6
reward 0.850

Usage

Run the reference answer to verify your environment is set up correctly

$ harbor run -p tasks/scaling_law

Test model

$ harbor run -p tasks/scaling_law \
  -a claude-code -m claude-opus-4-6

Description

Train a language model from scratch on WikiText-103 to achieve the lowest possible test perplexity within a fixed compute budget on a single H100 GPU. You can change the model architecture, size, training hyperparameters, precision, and any other aspect of the training script. The checkpoint must be compatible with the fixed evaluation pipeline (block_size >= 1024).

Files

path
permission
/app/train.py✎ Edit
/app/train.shRead-only
/app/evaluate_local.pyRead-only
/data/wikitext103/train.ptRead-only

Rules

  • 01Edit only /app/train.py.
  • 02No external network access.
  • 03Single GPU (H100), 8 CPUs, 32GB RAM. Time budget: 12 hours.
  • 04Checkpoint must contain model_state_dict and model_config; block_size >= 1024.

Tags

litgpttransformerscaling-lawperplexitycompute-optimal

Model Results

Click a row to view its trajectory in Live Lab

model
reward
score
Claude-Opus-4.6
0.850
MiMo-V2.5-Pro
0.780
Kimi-K2.6
0.650
Qwen-3.6-Plus
0.630
GLM-5
0.610
MiniMax-M2.7
0.430
GPT-5.4
0.330
DeepSeek-V4-Pro
0.300
Gemini-3.1-Pro
0.140
Grok-4-20
0.000
Hunyuan-3-Preview
0.000