llm_online_serving
hardserving scorePython
results
Measured by combined throughput and latency score — higher is better.
1.0×1.40×
sota
Grok-4-20
reward 0.340Usage
Run the reference answer to verify your environment is set up correctly
$ harbor run -p tasks/llm_online_serving
Test model
$ harbor run -p tasks/llm_online_serving \ -a claude-code -m claude-opus-4-6
Description
Optimize the SimpleLLM serving engine to maximize throughput and minimize latency when serving a 21B Mixture-of-Experts language model on an H100 GPU. The engine uses async request queuing, continuous batching, and slot-based KV caching. Score is based on a combined throughput and completion time ratio across 96 Poisson-arrival requests.
Files
path
permission
/app/llm.py✎ Edit/app/benchmark.pyRead-onlyRules
- 01Edit /app/llm.py only.
- 02Do NOT modify /app/model/, /app/kernels/, or /orig/.
- 03No external network access.
- 04Single GPU (H100 80GB). Time budget: 2 hours.
- 05Correctness check (pytest /tests/test_state.py) must pass.
Tags
llm-servingcuda-graphscontinuous-batchingmoeslot-management
Model Results
Click a row to view its trajectory in Live Lab
model
reward
score
Grok-4-20
0.340
0.310
Qwen-3.6-Plus
0.280
DeepSeek-V4-Pro
0.040
0.030
MiMo-V2.5-Pro
0.010
0.000
0.000
Kimi-K2.6
0.000
Hunyuan-3-Preview
0.000
0.000