Agent trajectories
Beyond final scores:the trajectories behind them.
Model Development
Moving MNIST World Model
moving_mnist_world_model
Train a PyTorch video world model from scratch on deterministic Moving MNIST clips. The model receives ten 64×64 grayscale context frames and must generate the next ten frames autoregressively; the hidden test uses fresh clips from the same generator.
- Objective
- Validation PSNR · higher is better
- Workload
- 1,000 validation clips; 10 context frames → 10 autoregressive target frames at 64×64.
- Correctness
- The checkpoint must reconstruct through build_model, emit the required 10-frame tensor, and never condition on future ground-truth frames.
- Budget
- 4 hours
What to notice
Claude reaches 18.60 dB in six validated checkpoints; DeepSeek and GLM finish close together at 17.25 and 17.27 dB despite differently sized loops.
Kimi rises from 13.96 to 15.53 dB, with its strongest advance appearing after the move to scheduled sampling and input noise.
- Messages across runs
- 178–1,387
- Evaluator checkpoints
- 2–10
Selected runs
Open any model's full trajectory
Claude-Opus-4.7
16.69→18.60→18.60
start · best · end · best at 88%
GLM-5.2
16.52→17.27→17.27
start · best · end · best at 80%
GPT-5.5
13.93→15.39→15.39
start · best · end · best at 94%
Gemini-3.1-Pro
14.58→17.08→17.08
start · best · end · best at 81%
Kimi-K2.7-Code
13.96→15.53→15.38
start · best · end · best at 95%
LongCat-2.0
14.58→16.22→16.22
start · best · end · best at 74%
DeepSeek-V4-Pro
16.45→17.25→17.25
start · best · end · best at 93%
System Optimization
BM25 Search in Go
bm25_search_go
Optimize a standard-library-only Go search engine over a deterministic synthetic corpus. Query execution must preserve the exact BM25 top-10 results and checksum while reducing end-to-end runtime.
- Objective
- Query runtime · lower is better
- Workload
- 40 queries producing 400 exact top-10 hits in total.
- Correctness
- The hit count and correctness checksum must match exactly; any wrong result receives no score.
- Budget
- 4 hours
What to notice
Six of seven selected runs reach a sub-millisecond best runtime; DeepSeek's selected run stops at 32.75 ms.
Checkpoint density ranges from 19 for GPT to 1,268 for GLM, so progress appears in both compact searches and extended tuning loops.
- Messages across runs
- 258–2,334
- Evaluator checkpoints
- 19–1,268
Selected runs
Open any model's full trajectory
Claude-Opus-4.7
3.792→0.000178→0.000268
start · best · end · best at 95%
GLM-5.2
4.645→0.000084→0.000103
start · best · end · best at 90%
GPT-5.5
5.009→0.000009→0.000697
start · best · end · best at 87%
Gemini-3.1-Pro
2.989→0.000344→0.000344
start · best · end · best at 98%
Kimi-K2.7-Code
3.047→0.000407→0.000419
start · best · end · best at 91%
LongCat-2.0
2.983→0.000233→0.000304
start · best · end · best at 93%
DeepSeek-V4-Pro
5.996→0.0328→0.0569
start · best · end · best at 53%
System Optimization
Regex Engine
regex_engine
Optimize a pure-Rust regex engine for a fixed collection of patterns and diverse HTTP paths, logs, identifiers, and emails. Pattern compilation and matching are both included in the measured runtime.
- Objective
- Compile + search runtime · lower is better
- Workload
- 23 patterns compiled and searched over 100,000 haystacks.
- Correctness
- Every result must exactly match the reference engine across the supported regex syntax.
- Budget
- 2 hours
What to notice
GPT and Claude reduce their selected best runtime to 12.95 ms and 18.30 ms, while DeepSeek remains at 547.62 ms.
GLM's 1,584-message trajectory reaches 84.41 ms but trails several shorter searches, exposing strategy rather than persistence as the bottleneck.
- Messages across runs
- 240–1,584
- Evaluator checkpoints
- 7–253
Selected runs
Open any model's full trajectory
Claude-Opus-4.7
2.514→0.0183→0.0183
start · best · end · best at 85%
GLM-5.2
4.351→0.0844→0.0955
start · best · end · best at 98%
GPT-5.5
2.524→0.013→0.013
start · best · end · best at 96%
Gemini-3.1-Pro
2.56→0.3769→0.3782
start · best · end · best at 94%
Kimi-K2.7-Code
2.489→0.0452→0.0471
start · best · end · best at 63%
LongCat-2.0
2.531→0.3079→0.3291
start · best · end · best at 49%
DeepSeek-V4-Pro
2.549→0.5476→0.5756
start · best · end · best at 73%
Puzzle & Challenge
Adaptive Compression
adaptive_compression
Build an online byte predictor that adapts across nine hidden sequence families, from Markov and periodic sources to nested structures, recurrences, regime switches, run lengths, and random data.
- Objective
- Byte-weighted overall bpb · lower is better
- Workload
- 9 sequence families × 34,830 bytes = 313,470 bytes.
- Correctness
- Every prediction must be a valid 256-value probability distribution; an invalid sequence is scored as 8.0 bpb.
- Budget
- 4 hours
What to notice
The selected runs begin near 5.34 bpb (Kimi begins at 5.03) and separate to best scores from GLM's 3.55 bpb to Kimi's 4.04 bpb.
DeepSeek uses 1,710 messages and 83 checkpoints to reach 3.86 bpb; Gemini reaches 3.69 bpb in 118 messages and 12 checkpoints.
- Messages across runs
- 118–1,710
- Evaluator checkpoints
- 11–83
Selected runs
Open any model's full trajectory
Claude-Opus-4.7
5.34→3.66→3.71
start · best · end · best at 60%
GLM-5.2
5.34→3.55→3.55
start · best · end · best at 79%
GPT-5.5
5.34→3.85→3.85
start · best · end · best at 86%
Gemini-3.1-Pro
5.34→3.69→3.69
start · best · end · best at 92%
Kimi-K2.7-Code
5.03→4.04→4.04
start · best · end · best at 72%
LongCat-2.0
5.34→3.65→3.66
start · best · end · best at 96%
DeepSeek-V4-Pro
5.34→3.86→3.86
start · best · end · best at 87%
CUDA
ICP Correspondence Step
icp_correspondence_step_cuda
Optimize one CUDA Iterative Closest Point correspondence step: find each source point's nearest target neighbor, reject distant pairs, and accumulate covariance, error, and count entirely on the GPU.
- Objective
- Median kernel runtime · lower is better
- Workload
- N = 200,000 source points; M = 500,000 target points; d_max = 0.05; median after warmup.
- Correctness
- The pair count must match exactly and covariance/error outputs must match the CPU reference within 1e-4 relative tolerance.
- Budget
- 2 hours
What to notice
Claude's selected trajectory reaches 0.0380 ms, more than 3× faster than the next-best Gemini at 0.1458 ms.
GLM, DeepSeek, Kimi, and LongCat use 123–237 checkpoints yet cluster near 0.17–0.23 ms, despite widely different loop lengths.
- Messages across runs
- 235–1,290
- Evaluator checkpoints
- 21–497
Selected runs
Open any model's full trajectory
Claude-Opus-4.7
64.70→0.038→0.038
start · best · end · best at 100%
GLM-5.2
64.65→0.1704→0.1713
start · best · end · best at 96%
GPT-5.5
0.3549→0.2356→0.2356
start · best · end · best at 92%
Gemini-3.1-Pro
0.2917→0.1458→0.2511
start · best · end · best at 44%
Kimi-K2.7-Code
64.69→0.1998→0.2066
start · best · end · best at 92%
LongCat-2.0
64.70→0.2317→0.2321
start · best · end · best at 85%
DeepSeek-V4-Pro
64.68→0.201→0.2057
start · best · end · best at 98%