moe_routing
mediumruntimePython
results
Measured by wall-clock runtime in seconds — lower is better.
35.0s3.0s
Usage
Run the reference answer to verify your environment is set up correctly
$ harbor run -p tasks/moe_routing
Test model
$ harbor run -p tasks/moe_routing \ -a claude-code -m claude-opus-4-6
Description
Process a batch of 32 sequences × 512 tokens through an MoE FFN layer with 64 experts (d_model=2048, d_ff=4096). Each token is routed to its top-2 experts by a learned gating network. The baseline iterates over each token, computes gating scores, selects top-2, runs two expert FFNs, and blends outputs. Output must match reference within 1e-5.
Files
path
permission
/app/solve.py✎ Edit/app/main.pyRead-onlyRules
- 01Edit /app/solve.py only.
- 02Allowed imports: numpy only. No PyTorch, JAX, or compiled extensions.
- 03Output must match reference within 1e-5. Wrong results score 0.
Tags
MoEmixture-of-expertstop-k-routingsparseLLM