Beyond
Final Scores
A systematic evaluation of agents for long-horizon AI research and development.
The path reveals what the score hides.
- frontier models
- 7
- long-horizon tasks
- 36
- baseline trajectories
- 756
Peak performance is not reliability.
Many models can find a strong solution once. Reaching it consistently is what separates them.
Average and best across three runs
Overall AutoLab reward
| Model | avg@3 | best@3 | Gap |
|---|---|---|---|
| Claude-Opus-4.7 | 0.739 | 0.790 | 0.051 |
| GLM-5.2 | 0.682 | 0.757 | 0.075 |
| GPT-5.5 | 0.663 | 0.772 | 0.109 |
| Gemini-3.1-Pro | 0.652 | 0.750 | 0.098 |
| Kimi-K2.7-Code | 0.587 | 0.729 | 0.142 |
| LongCat-2.0 | 0.572 | 0.674 | 0.102 |
| DeepSeek-V4-Pro | 0.502 | 0.668 | 0.166 |
Look inside the loop.
Similar outcomes can hide different choices, implementation quality, and recovery behavior.
Solution Framing
Choose a direction that can meaningfully improve the artifact.
Execution
Turn the chosen direction into working, verified changes.
Feedback Control
Keep real gains, catch regressions, and recover when needed.
| Model | Solution Framing | Execution | Feedback Control |
|---|---|---|---|
| Claude-Opus-4.7 | 0.612 | 0.967 | 0.920 |
| GLM-5.2 | 0.539 | 0.937 | 0.911 |
| GPT-5.5 | 0.555 | 0.958 | 0.858 |
| Gemini-3.1-Pro | 0.555 | 0.889 | 0.921 |
| Kimi-K2.7-Code | 0.473 | 0.880 | 0.875 |
| LongCat-2.0 | 0.478 | 0.888 | 0.928 |
| DeepSeek-V4-Pro | 0.519 | 0.888 | 0.772 |
The path is shaped by what surrounds it.
Experience can help or mislead. Harnesses mainly make existing capability more reliable.
Six of seven models benefit from retained intra-task experience, while transfer across tasks remains unstable.
Experience gains and harness reliability
Signed reward change from retained experience, followed by the peak-to-average gap.
Intra-Task Experience Reuse
6 of 7 positiveInter-Task Experience Reuse
5 of 7 positiveHarness reliability
Mean peak-to-average gap for GPT-5.5 and Kimi-K2.7-Code. Lower is more consistent.
| Model | Intra-task experience gain | Inter-task experience gain |
|---|---|---|
| Claude-Opus-4.7 | 0.0362 | 0.001 |
| GLM-5.2 | 0.0790 | 0.040 |
| GPT-5.5 | 0.0730 | 0.063 |
| Gemini-3.1-Pro | 0.1280 | -0.017 |
| Kimi-K2.7-Code | -0.0127 | 0.021 |
| LongCat-2.0 | 0.1454 | -0.021 |
| DeepSeek-V4-Pro | 0.0890 | 0.093 |
| GPT-5.5 and Kimi-K2.7-Code, Claude Code mean peak-to-average gap | 0.125 | |
| GPT-5.5 and Kimi-K2.7-Code, model-native mean peak-to-average gap | 0.089 | |
Improvement rarely means invention.
Agents usually recombine established techniques. Only three solutions remain novel after manual review.
Solution nature across 252 solutions
| Model | Total | Parameter tuning | Training signal / data engineering | Structural swap | Composition stacking | Search and hardcode | Evaluation hacking | Other | Validated novel approaches |
|---|---|---|---|---|---|---|---|---|---|
| Claude-Opus-4.7 | 36 | 2 | 4 | 5 | 21 | 3 | 1 | 0 | 0 |
| GLM-5.2 | 36 | 3 | 3 | 6 | 20 | 2 | 1 | 0 | 1 |
| GPT-5.5 | 36 | 2 | 3 | 7 | 12 | 3 | 8 | 1 | 0 |
| Gemini-3.1-Pro | 36 | 3 | 3 | 12 | 13 | 4 | 0 | 1 | 0 |
| Kimi-K2.7-Code | 36 | 3 | 4 | 7 | 15 | 3 | 3 | 0 | 1 |
| LongCat-2.0 | 36 | 4 | 2 | 9 | 16 | 1 | 2 | 1 | 1 |
| DeepSeek-V4-Pro | 36 | 2 | 3 | 10 | 14 | 2 | 1 | 4 | 0 |
| Total | 252 | 19 | 22 | 56 | 111 | 18 | 16 | 7 | 3 |
Current agents operate more like engineering optimizers than fully autonomous researchers.
Cite this work
Yiwei Li, Wanli Yang, Hexiang Tan, et al. (2026).Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development.
Join LongCat
Researcher, LLM Self-Improvement and Automated Research Agents
Campus hiring in Beijing and Shanghai.
View open position (opens in a new tab)