RSIBench-Data

Can better data teach a model to do more?

RSIBench-Data tests whether a frontier researcher agent can turn model failures into a stronger training-data strategy under a shared, bounded post-training interface.

4 researcher agents
6 target benchmarks
Qwen3.5 target model
16 h nominal budget

Change the experience. Keep the comparison.

Researchers submit data strategies and permitted configurations. Every candidate then passes through the same target model, training backend, serving path, sandbox, verifier, and resource budget.

Experiment record with five variables. Four hold a single fixed mark in one vertical column. Data strategy shows multiple candidate marks, with short guides into a comparable evidence axis below the locked rows.

Evidence changes the next proposal.

The agent may diagnose, propose, train, evaluate, and revise within a fixed budget.

A spiral of about one and a third turns through Diagnose, Propose, Train, Evaluate and Revise. Propose appears at two radii; the path ends at the revised Propose marker.

One interface, six different kinds of failure.

The same researcher protocol spans repository-grounded coding, long-horizon terminal work, graduate science, and competition mathematics.

SWE-bench Verified

Repository-level issue resolution with executable validation.

SWE-bench Multilingual

Cross-language repository tasks and diverse code ecosystems.

SWE-bench Pro

Harder professional software tasks with sparse success.

Terminal-Bench 2.0

Long-horizon terminal interaction with sandbox verification.

GPQA Diamond

Graduate-level, expert-written scientific reasoning questions.

AIME 2026

Competition mathematics with exact, verifiable answers.

The original research figures, in full.

These original research figures are preserved unchanged in this redesign, without cropping, recoloring, or animated reinterpretation.

Comparison of existing agentic post-training benchmarks, where many variables change, and RSIBench-Data, which fixes training infrastructure and opens the data research loop.
Motivation figure from the RSIBench-Data study, preserved without alteration.
Framework diagram showing fixed infrastructure, research inputs, the data-synthesis loop, and independent official evaluation.
Framework figure from the RSIBench-Data study, preserved without alteration.

Discovery is real. Reliability is not.

Current researcher agents often find stronger data strategies after feedback, then continue searching past the best candidate or fail to preserve it.

Discovery 14 / 24

settings improve beyond the first valid attempt during the research trajectory.

Reliability gap 18 / 23

post-peak searches finish below their historical best candidate.

Selection score trajectories across valid attempts for six benchmarks.
Figure 03 / Valid attempts in chronological order. Stars mark the first candidate that reaches each run's historical best.

No single researcher wins every task.

The strongest agent depends on the benchmark. The result is a research landscape, not one universal winner.

Official performance and resource use for the complete researcher-agent by benchmark matrix.
Benchmark Agent Official Time (h) Tinker cost ($)
SWE-bench Verified Base model 12.00%
Claude Code · Opus-4.8 46.00% 10.21 195.90
Claude Code · Sonnet-5 35.00% 14.91 181.70
Codex · gpt-5.6-sol 33.00% 4.41 55.61
Codex · gpt-5.6-terra 42.00% 2.80 59.79
SWE-bench Multilingual Base model 7.00%
Claude Code · Opus-4.8 5.00% 12.68 195.70
Claude Code · Sonnet-5 22.00% 14.20 363.77
Codex · gpt-5.6-sol 15.00% 5.99 78.64
Codex · gpt-5.6-terra 6.00% 5.18 56.87
SWE-bench Pro Base model 0.00%
Claude Code · Opus-4.8 2.00% 2.19 17.54
Claude Code · Sonnet-5 4.00% 9.54 45.63
Codex · gpt-5.6-sol 9.00% 11.61 300.49
Codex · gpt-5.6-terra 1.00% 3.85 34.39
GPQA Diamond Base model 61.00%
Claude Code · Opus-4.8 56.00% 5.98 27.92
Claude Code · Sonnet-5 52.00% 6.43 16.03
Codex · gpt-5.6-sol 65.00% 2.42 10.37
Codex · gpt-5.6-terra 64.00% 1.14 4.80
AIME 2026 Base model 30.00%
Claude Code · Opus-4.8 40.83% 4.89 61.36
Claude Code · Sonnet-5 49.17% 9.29 121.21
Codex · gpt-5.6-sol 53.33% 8.90 63.87
Codex · gpt-5.6-terra 33.33% 1.85 8.50
Terminal-Bench 2.0 Base model 1.12%
Claude Code · Opus-4.8 10.11% 8.99 175.63
Claude Code · Sonnet-5 5.62% 8.87 157.85
Codex · gpt-5.6-sol 20.22% 9.56 69.07
Codex · gpt-5.6-terra 12.36% 6.67 186.96

Current release / RSIBench-Data

One controlled surface, not the final system.

RSIBench-Data exercises the currently unlocked surface in a larger program that will gradually open Algorithm, Agent / Harness, Arch, and eventually their interactions.