SWE-bench Verified
Repository-level issue resolution with executable validation.
Can better data teach a model to do more?
RSIBench-Data tests whether a frontier researcher agent can turn model failures into a stronger training-data strategy under a shared, bounded post-training interface.
Researchers submit data strategies and permitted configurations. Every candidate then passes through the same target model, training backend, serving path, sandbox, verifier, and resource budget.
Experiment record with five variables. Four hold a single fixed mark in one vertical column. Data strategy shows multiple candidate marks, with short guides into a comparable evidence axis below the locked rows.
The agent may diagnose, propose, train, evaluate, and revise within a fixed budget.
A spiral of about one and a third turns through Diagnose, Propose, Train, Evaluate and Revise. Propose appears at two radii; the path ends at the revised Propose marker.
The same researcher protocol spans repository-grounded coding, long-horizon terminal work, graduate science, and competition mathematics.
Repository-level issue resolution with executable validation.
Cross-language repository tasks and diverse code ecosystems.
Harder professional software tasks with sparse success.
Long-horizon terminal interaction with sandbox verification.
Graduate-level, expert-written scientific reasoning questions.
Competition mathematics with exact, verifiable answers.
These original research figures are preserved unchanged in this redesign, without cropping, recoloring, or animated reinterpretation.
Current researcher agents often find stronger data strategies after feedback, then continue searching past the best candidate or fail to preserve it.
settings improve beyond the first valid attempt during the research trajectory.
post-peak searches finish below their historical best candidate.
The strongest agent depends on the benchmark. The result is a research landscape, not one universal winner.
| Benchmark | Agent | Official | Time (h) | Tinker cost ($) |
|---|---|---|---|---|
| SWE-bench Verified | Base model | 12.00% | — | — |
| Claude Code · Opus-4.8 | 46.00% | 10.21 | 195.90 | |
| Claude Code · Sonnet-5 | 35.00% | 14.91 | 181.70 | |
| Codex · gpt-5.6-sol | 33.00% | 4.41 | 55.61 | |
| Codex · gpt-5.6-terra | 42.00% | 2.80 | 59.79 | |
| SWE-bench Multilingual | Base model | 7.00% | — | — |
| Claude Code · Opus-4.8 | 5.00% | 12.68 | 195.70 | |
| Claude Code · Sonnet-5 | 22.00% | 14.20 | 363.77 | |
| Codex · gpt-5.6-sol | 15.00% | 5.99 | 78.64 | |
| Codex · gpt-5.6-terra | 6.00% | 5.18 | 56.87 | |
| SWE-bench Pro | Base model | 0.00% | — | — |
| Claude Code · Opus-4.8 | 2.00% | 2.19 | 17.54 | |
| Claude Code · Sonnet-5 | 4.00% | 9.54 | 45.63 | |
| Codex · gpt-5.6-sol | 9.00% | 11.61 | 300.49 | |
| Codex · gpt-5.6-terra | 1.00% | 3.85 | 34.39 | |
| GPQA Diamond | Base model | 61.00% | — | — |
| Claude Code · Opus-4.8 | 56.00% | 5.98 | 27.92 | |
| Claude Code · Sonnet-5 | 52.00% | 6.43 | 16.03 | |
| Codex · gpt-5.6-sol | 65.00% | 2.42 | 10.37 | |
| Codex · gpt-5.6-terra | 64.00% | 1.14 | 4.80 | |
| AIME 2026 | Base model | 30.00% | — | — |
| Claude Code · Opus-4.8 | 40.83% | 4.89 | 61.36 | |
| Claude Code · Sonnet-5 | 49.17% | 9.29 | 121.21 | |
| Codex · gpt-5.6-sol | 53.33% | 8.90 | 63.87 | |
| Codex · gpt-5.6-terra | 33.33% | 1.85 | 8.50 | |
| Terminal-Bench 2.0 | Base model | 1.12% | — | — |
| Claude Code · Opus-4.8 | 10.11% | 8.99 | 175.63 | |
| Claude Code · Sonnet-5 | 5.62% | 8.87 | 157.85 | |
| Codex · gpt-5.6-sol | 20.22% | 9.56 | 69.07 | |
| Codex · gpt-5.6-terra | 12.36% | 6.67 | 186.96 |
Current release / RSIBench-Data
RSIBench-Data exercises the currently unlocked surface in a larger program that will gradually open Algorithm, Agent / Harness, Arch, and eventually their interactions.