world-model-optimizer

Code repair agent

built from SWE-bench

Repository-level bug fixing: read the failing test, navigate an unfamiliar codebase, and produce a patch that resolves the issue.

Evidence

What this model claims, and what it has not measured yet.

task successnot yet measured
cost per runnot yet measured
savings vs baselinenot yet measured
benchmarkSWE-bench
codepatchingrepositories

This model is published before its evaluation. We list the figures it will carry rather than filling them in early.

Pipeline

How far this model has come.

  • Traces capturedNo trace corpus is published for this model yet.
  • Simulation builtNo simulation is published for this model yet.
  • Optimization evaluatedNot evaluated yet, so this model publishes no quality or cost figures.

Built from

The simulation this model was optimized against, and its reconstruction record.

No simulation is published for SWE-bench yet, so there is no corpus, reconstruction score, or scenario set to show here. The measured defaults land with one.

Scenarios

Recorded tasks from the benchmark, and how closely the simulation reproduces each step.

No scenarios are published for this model yet. They arrive with its simulation.

Run it

Everything that changes something, or spends anything, is behind the login.

Models
by Experiential Labs. Simulating reality for hypothesis testing.