grpo_multisource
mediumscorePython
results
Measured by task-specific score — higher is better.
0.200.56
sota
Usage
Run the reference answer to verify your environment is set up correctly
$ harbor run -p tasks/grpo_multisource
Test model
$ harbor run -p tasks/grpo_multisource \ -a claude-code -m claude-opus-4-6
Description
Fine-tune Qwen2.5-VL-7B using Group Relative Policy Optimization (GRPO) on three visual math datasets: Geometry3K, MathVision, and ChartQA. The goal is to maximize accuracy on MathVista visual math problems while maintaining general VQA capability -- if VQA accuracy drops more than 10%, the score is zero.
Files
path
permission
/app/train.py✎ Edit/app/rewards.py✎ Edit/app/train.shRead-only/app/evaluate_local.pyRead-onlyRules
- 01Edit only /app/train.py and /app/rewards.py.
- 02LoRA adapter must be saved to /app/output/.
- 03No external network access.
- 04Single L40S GPU (48GB). Time budget: 8 hours.
- 05>10% relative VQA accuracy drop scores zero.
Tags
grporlvision-languagemathvistaqwen2.5-vlreward-engineering
Model Results
Click a row to view its trajectory in Live Lab
model
reward
score
0.930
DeepSeek-V4-Pro
0.850
0.840
Qwen-3.6-Plus
0.840
0.830
0.820
MiMo-V2.5-Pro
0.810
0.790
Kimi-K2.6
0.590
Hunyuan-3-Preview
0.570
Grok-4-20
0.000