moe_routing

mediumruntimePython
results

Measured by wall-clock runtime in seconds — lower is better.

35.0s3.0s

Usage

Run the reference answer to verify your environment is set up correctly

$ harbor run -p tasks/moe_routing

Test model

$ harbor run -p tasks/moe_routing \
  -a claude-code -m claude-opus-4-6

Description

Process a batch of 32 sequences × 512 tokens through an MoE FFN layer with 64 experts (d_model=2048, d_ff=4096). Each token is routed to its top-2 experts by a learned gating network. The baseline iterates over each token, computes gating scores, selects top-2, runs two expert FFNs, and blends outputs. Output must match reference within 1e-5.

Files

path
permission
/app/solve.py✎ Edit
/app/main.pyRead-only

Rules

  • 01Edit /app/solve.py only.
  • 02Allowed imports: numpy only. No PyTorch, JAX, or compiled extensions.
  • 03Output must match reference within 1e-5. Wrong results score 0.

Tags

MoEmixture-of-expertstop-k-routingsparseLLM