All Tasks
23 tasks across model development, system optimization, and puzzle challenges. Each task presents a working but suboptimal program — the agent must optimize it as far as possible.
View on GitHubModel Development(5)
End-to-end LLM pipeline — from pretraining to posttraining, data curation to serving.
Pre-training(1)
Post-training(3)
grpo_multisource
PythonFine-tune Qwen2.5-VL-7B with GRPO on multi-source visual math data (Geometry3K, MathVision, ChartQA) to maximize MathVista accuracy.
data_select_ifeval
PythonSelect up to 5,000 training samples from a 50k mixed data pool to maximize IFEval prompt-level strict accuracy after LoRA fine-tuning Qwen2.5-3B-Instruct.
multilingual_ocr
PythonFine-tune DeepSeek-OCR (3B) with Unsloth LoRA to minimize Character Error Rate on Persian and Bengali synthetic text images.
Inference Serving(1)
System Optimization(12)
Low-level performance optimization across algorithms, data structures, and systems.
flash_attention
CCompute scaled dot-product attention for n=4096, d=64, float32. The naive baseline allocates a full 128MB score matrix. Key optimization: Flash Attention tiling with online softmax.
bm25_search_go
GoOptimize a Go BM25 search engine to compute exact top-10 results across a synthetic corpus as fast as possible. Goroutines are allowed; standard library only.
aes128_ctr
COptimize AES-128 in CTR mode to encrypt 256 MiB of data as fast as possible. Output must match NIST SP 800-38A test vectors.
bvh_raytracer
C++Build a BVH acceleration structure in C++ to reduce ray-triangle intersection tests from O(N) brute-force to O(log N) per ray for a 638×638 scene with 4096 triangles.
concurrent_kv_wal
GoOptimize a WAL-backed in-memory key-value store in Go running a deterministic multi-phase workload with 4 concurrent goroutines.
fft_rust
RustImplement a fast FFT in Rust for a length-32768 real signal. Replace the naive O(n²) DFT with an iterative Cooley-Tukey FFT with precomputed twiddle factors.
gaussian_blur
CApply a 17×17 Gaussian blur (sigma=3.0) to a 4096×4096 grayscale image 5 times. Key avenues: separable filter decomposition, SIMD/AVX2, cache-friendly access.
hash_join
COptimize a C hash join between a 20K-row build table and 5M-row probe table. Replace the O(R×S) nested-loop baseline with an open-addressing hash table.
regex_engine
RustImplement a fast regex engine in Rust that compiles patterns and searches 100,000 haystacks. Replace the recursive AST-walking NFA with a bytecode NFA using bitset-based active-state tracking.
sstable_compaction_rs
RustOptimize an LSM-style SSTable compaction pipeline in Rust that merges prefix-compressed sorted tables, applies LSM visibility rules, and emits a new sorted output SSTable.
radix_sort
CSort 50 million random 32-bit unsigned integers in C as fast as possible. Replace the stdlib qsort baseline with a 2-pass LSD radix sort.
sha256_throughput
COptimize SHA-256 in C to hash a 512 MiB buffer as fast as possible. Key optimization: runtime CPUID dispatch to Intel SHA-NI intrinsics.
Puzzle & Challenge(6)
Algorithmic puzzles, code golf, and discrete optimization challenges with unique scoring.
discover_sorting
PythonGenerate a correct 16-input sorting network with as few comparators as possible. Correctness is verified exhaustively on all 2^16 binary inputs.
fredkin_sort_network
TextRewrite a reversible circuit to sort four 2-bit values using Fredkin/Toffoli gates with as few gates as possible. All scratch wires must be restored to 0.
stack_machine_golf
TextRewrite a stack machine program to compute a 256-element integer dot product with as few executed instructions as possible. Key technique: loop unrolling.
vliw_scheduler
CImplement a VLIW instruction scheduler in C that packs 3,000 operations into 3-slot bundles (ALU/MUL/MEM) minimizing total cycles.
smallest_game_player
PythonBuild a model with the fewest learnable parameters that achieves ≥95% accuracy on perfect-play positions from 4×4 gravity Connect-3.
toy_isa_opt
TextRewrite a PINC ISA assembly program to minimize simulated cycle count for a 512-element dot product. Key technique: 4× loop unrolling to hide 5-cycle latency.