AWS
Preprint · 2026 | arXiv:2609.38349

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

MILO (Meta-evolutionary Island Orchestration) is a self-adaptive evolutionary search framework for automated harness discovery. Given an agent, it discovers a stronger harness for its model without retraining. Multiple agentic mutators, each evolving its own island of candidates, rewrite whole harnesses, including their control flow, to fit the target task distribution, orchestrated by a meta-agent that grafts what works between islands, swaps in fresh mutators when one runs dry, and reshapes the curriculum, so the search itself keeps evolving toward the strongest harness.

MILO’s harnesses outperform eight expert-built and six searched harnesses on command-line, paper-replication and software-engineering tasks. On Terminal-Bench 2.1 they reach 86.1%, above the leaderboard’s top entry, with 26% fewer tokens than their seed. On EinsteinArena’s open math problems, MILO sets three new records, surpassing AlphaEvolve, TTT-Discover and EvoX.

Prithwish Jana1,*,†, Mononito Goswami2, Hao Liu2, Xinyu Li3,†, Langlin Huang4,†, Zhehui Huang2, Zhishen Huang2, Patrick Blöbaum2, Anoop Deoras2, Purak Jain2,‡, Nikos Kanakaris2,*,‡, Sahika Genc2,*,‡

1Georgia Institute of Technology 2AWS AI Labs 3Carnegie Mellon University 4Washington University in St. Louis

Georgia Institute of Technology Amazon Web Services Carnegie Mellon University Washington University in St. Louis

*Correspondence to pjana7@gatech.edu, nikosk@amazon.com or sahika@amazon.com.
†Work done while at AWS AI Labs.
‡Senior co-authorship.

Overview of MILO: an inner loop evolves harnesses across islands (parent selection, mutation, fitness evaluation, child acceptance, progress monitoring) and an outer orchestrator loop restructures memory, reassigns mutators and reshapes the curriculum when an island stalls.
Fig. 1Overview of MILO, a meta-evolutionary harness-discovery framework. Given an agent, MILO discovers a stronger harness for its model via two coupled loops: an inner loop that evolves harnesses and an outer loop that evolves the search strategy. In the inner loop, (1) a parent is drawn from the island’s admitted candidates, (2) its mutator agent rewrites the whole harness from the parent’s failure traces and the island’s lineage memory, including rejected candidates, and (3–4) the child is scored and admitted on multi-objective Pareto gain. (5) When progress stalls, the outer loop’s orchestrator agent diagnoses the island and revises its mutator, memory, and curriculum. Interactive: click a step on the left or any box in the diagram to play that part; Space pauses, arrow keys step through.tap a step on the left or any box in the diagram to play that part; swipe sideways to see the whole figure. Press and hold static figure to see the original from the paper.
Terminal-Bench 2.1 86.1% RR@5 Above the official leaderboard’s top entry (83.8%), with Opus 4.8 as the backbone.
Gain over initial harness +12.0 · +28.3 · +10.3 Resolution-rate gains on Terminal-Bench 2.1, PaperBench and DeepSWE; the best prior search manages +4.5, +18.3 and 0.
Cheaper, not just better 26% fewer tokens MILO’s Terminal-Bench harness spends 0.74× the tokens of the harness it started from.
Open math problems 3 new records Tighter best-known bounds on EinsteinArena, surpassing AlphaEvolve, TTT-Discover and EvoX.
01

Overview

Modern agentic systems pair a model with a harness: the software layer that executes the model’s actions in a stateful environment and governs its scaffolding and control flow. On long-horizon tasks the harness shapes agent performance as much as the model does. Yet harness engineering is still artisanal, and the effort must be repeated every time the model changes.

Model ℳA black box with inference access only. It stays fixed: MILO never retrains it.
what MILO evolvesHarness hPrompts, tools, context and memory, plus the control flow: spawning sub-agents, intercepting tool calls, verifying outputs, deciding when to stop.
EnvironmentA container, repository or terminal, scored by a hidden verifier: a test suite or a rubric.
35.2%→49.6%
Same model, different harness. On Terminal-Bench, GPT-5 solves 35.2% of tasks with the Terminus 2 harness and 49.6% with Codex, while consuming 35% fewer tokens. A harness can matter as much as the model, and it is far cheaper to change than model weights. Yet production harnesses take teams months to build, and because harnesses are model-specific, the work is redone with every new model.

This motivates automated harness discovery (AHD): given an agent, discover a stronger harness for its model. AHD is usually cast as LLM-driven evolutionary search, a loop in which an LLM-based mutator proposes harnesses, an oracle scores them, and the results guide the next round. Because the loop is model-agnostic, rerunning it per model automates the perpetual re-tuning. Existing search methods explore this space poorly, however: most optimize only parts of the harness, such as prompts or skills, while the rest fix their search strategy in advance and inherit the exploitative bias of their LLM mutators.

MILO (Meta-evolutionary Island Orchestration) co-evolves the harness and its own search strategy. A hierarchical lineage memory of island-based trees keeps rejected mutations as negative evidence; per-island mutator agents combine global search history with feedback on parent weaknesses to rewrite entire harnesses; and an orchestrator agent adapts the search by grafting and speciating lineages, reassigning mutators, and revising the curriculum. With gpt-oss-120b it lifts DeepSWE pass-rate to 15.6%, 20× its initial harness, while prior search stays below 1%.

Problem statement · Automated harness discovery

Given an agent ⟨ℳ, h⟩ and a target task distribution 𝒯, discover a harness h∗ that advances the agent’s accuracy–cost frontier on 𝒯, with the model ℳ fixed.

InputsThe agent, optionally with further seed harnesses; a benchmark sampled from 𝒯 and split into search and validation tasks; constraints requiring a runnable agent and no task-specific logic; a budget of rounds or wall-clock time.
Search spaceCountably infinite: any harness with a runnable entry point and editable source, whether a Python program, a YAML file, or a tool and skill repository.
OutputThe harness that most advances the accuracy–cost frontier on the validation split, ideally surpassing the strongest hand-built agents.

Four shortcomings of existing search, four design choices

Applied to harness discovery, prior evolutionary search falls short in four ways. MILO’s key insight is to evolve the complete search strategy alongside the harness, through four matching design choices.

1 · Search-space exploration

Islands that share ideas but never compete

×An LLM mutator exploits far more than it explores, so populations collapse onto variants of a few designs. Prior methods confine edits to prompts or skills, or rewrite whole harnesses with a single fixed, exploitation-biased mutator.

✓Each island grows from its own seed with its own mutator agent. Islands exchange ideas through grafts but never compete for survival, so a mutator’s bias stays confined. In one run the weakest seed produced the best harness.

2 · Sample-efficient search

A memory that keeps its failures

×Scoring one harness takes hours, so every evaluation must inform the search. Yet most methods are failure-blind: they keep only survivors, whether a single candidate, a Pareto front or the fittest per MAP-Elites cell, so failed directions get retried.

✓The lineage memory is append-only: every seed-to-candidate path is kept, rejected children included, annotated with the edit that produced it and its outcome. Successes show which edits paid off; failures prune later mutations and push mutators toward bolder structural changes.

3 · Self-adaptive search

An orchestrator that rewrites the strategy

×The strategy decides which parents are selected, how they are mutated and what is kept. Most methods fix it in advance and cannot escape local optima; the most adaptive evolve only parent selection and mutation.

✓When an island stalls, the orchestrator diagnoses it and intervenes: reassigning its mutator, grafting from a stronger island, speciating a crowded-out niche, or reshaping the curriculum toward high-regret tasks.

4 · Multi-objective fitness

Admission by Pareto gain

×A harness sets not only whether an agent succeeds but also its tokens and latency. Yet almost all search methods optimize accuracy alone and drift toward costly harnesses; none optimizes latency.

✓A child is admitted only if it grows its island’s Pareto front in accuracy, tokens and latency. The discovered harnesses are both more accurate and cheaper to run: on Terminal-Bench, 0.74× the seed’s tokens with Opus 4.8 and 0.70× with gpt-oss-120b.

02

Method

MILO is a strategy-level evolutionary framework: it searches for a solver (a harness) that must generalize across a task distribution, rather than a solution to a single task. Three interacting components drive the search, coupled by a five-stage evolution loop.

Component i

Hierarchical lineage memory

A forest of per-island trees. Each node is a harness annotated with its accuracy, cost, failure evidence and admission verdict; each edge is the source-code patch that produced the child and the gains it brought. The memory is append-only: rejected children stay as negative evidence.

Self-adaptive. Graft breeds a stalled island’s best harness with a donor island’s, adding a cross-island edge so the destination inherits the donor’s capability; Speciate promotes a crowded-out subtree to a new island of its own.

Component ii

Evidence-driven mutator agents

Mutators are coding agents that reason over many tool-calling turns. Each receives the parent’s source, its failure evidence (traces and verifier reports) and a snapshot of the island’s lineage, and follows a diagnose-first workflow: find why the parent fails, then apply anything from a local fix to a control-flow rewrite. A pool of heterogeneous agents, each pairing a different model with a different scaffold, gives islands different mutation tendencies.

Self-adaptive. The island–mutator assignment carries over between rounds until Reassign hands a stalled island a different mutator whose tendencies fit its diagnosed need.

Component iii

Meta-evolutionary orchestrator

A single agent that observes all islands. Triggered when an island stalls for P rounds, it diagnoses the bottleneck against the whole search configuration: the memory, the island–mutator assignment and the curriculum of search tasks. It intervenes with at most two moves.

Self-adaptive. Besides Graft, Speciate and Reassign, Curriculum replaces the search set with high-regret tasks, those a few harnesses solve and most fail, where search helps most. By rewriting the configuration, the search adapts its memory, mutators and curriculum together.

The evolution loop every island advances one generation per round

Parent selection

Softmax over the island’s admitted set, trading dominance (exploitation) against behavioural similarity on per-task pass rates (exploration).

Mutation

The island’s mutator agent rewrites the parent into a child harness, guided by the parent’s failure evidence and the lineage snapshot.

Fitness evaluation

k attempts per task on the search split. Accuracy, tokens and latency form a multi-objective fitness vector.

Child acceptance

Admitted only if it grows the island’s Pareto hypervolume. Admitted or not, the child is appended to the lineage.

Progress monitoring

Held-out accuracy is tracked. A stall within patience P escalates the island to the orchestrator.

Diagnose, then intervene four diagnoses, four actions

Diagnosis
Dry mutator
↓
Reassign(j, μ)

Children are repeatedly rejected, inert, or minor variations of failed edits. The island gets a new mutator from the pool.

Diagnosis
Exhausted lineage
↓
Graft(donor → dst)

No admitted child although another island solves tasks this one fails. Both islands’ best harnesses are bred as co-parents.

Diagnosis
Crowded-out niche
↓
Speciate(h, μ)

A dominant lineage crowds out a distinct subtree that solves different tasks. The subtree becomes a new island with its own mutator.

Diagnosis
Whole-population plateau
↓
Curriculum(ℬ′)

All islands approach the same ceiling. The search set is replaced with high-regret frontier tasks: solved by a few harnesses, failed by most.

MILO is model-agnostic (it improves any agent regardless of its backbone) and domain-agnostic (it only scores and edits source code). It also extends to instance-level discovery: rather than evolving a solution program directly, MILO evolves the harness of the agent that writes it, pairing starting constructions for a scientific problem with optimizer programs for the agent to edit. That extension is what set the EinsteinArena records below.

03

See it run

Two interactive demos, each a real run played back from its own records. First, the harness search from the paper: the Terminal-Bench 2.1 run with gpt-oss-120b, replayed round by round. Three islands start from three expert-designed seeds, stall together for nine rounds, and the orchestrator’s grafts and mutator reassignments pull the population from 45.0% to 59.9%.

Interactive demo · a real run, played back from its records

Three islands, nine orchestrator calls, 93 candidate harnesses

Each lane is an island’s lineage tree (x = round). Filled dots are admitted harnesses, hollow crossed dots are rejected children kept as negative evidence, and dashed arrows are grafts bred from another island’s best harness (green once admitted). The right panel tracks each island’s best pass-rate. Every event is read from the run’s own records.

Download clip (MP4)
Island 1Island 2Island 3 Rejected childGraft admittedGraft rejectedOrchestrator invoked
Mutator codes: cc Claude Code (Opus 4.8), cx Codex (GPT-5.5), d:g DeepAgents (GPT-5.5), d:q DeepAgents (Qwen3-Coder-480B). The orchestrator notes are condensed from the diagnosis it wrote at each invocation. Hover any dot for the candidate’s pass-rate, verdict and token cost. Scrub the timeline to step through rounds.

Second, scientific discovery: MILO’s instance-level variant evolving optimizer-writing agents on two open problems from EinsteinArena, the third autocorrelation inequality and the Erdős minimum overlap problem (records in section 06). Watch the objective ratchet down over wall-clock days until it passes the arena’s best-known bound; switch problems under the chart.

Interactive demo · a real run, played back from its records

Watch a record fall: one search over wall-clock days

Replayed from the run directories. Each hollow circle is one evolved harness evaluated on 30 optimizer-writing tasks (its best objective); each blue dot is a construction that beat the arena’s frozen #1, timestamped by its file on disk. The red line is the best objective so far; the right panel draws the record construction over the arena entry the agent started from.

The third-autocorrelation run began on 2026-09-02 and dipped below the arena’s #1 within two days, but the record itself took 12.7 days and 41 harness evaluations; 403 sub-record constructions were saved along the way. The Erdős run also crossed the #1 within two days and set its record on day 17, its last day. Timestamps are UTC.
04

Results

Three baseline tiers on three long-horizon benchmarks, with a frontier and an open-weight backbone. Every automated search method starts from the same three expert-designed seeds and gets the same 72-hour budget.

3+1 benchmarksTerminal-Bench 2.1 (89 command-line tasks), PaperBench-CodeDev (20 ICML papers), DeepSWE (113 tasks across 91 repos); Frontier-Bench (70 tasks) for transfer only.
8+1 harnessesCline, DeepAgents, Goose, Mini-SWE-Agent, OpenCode, OpenHands, Qwen-Code, Terminus-2, plus a minimal read/write/bash harness.
6search methodsInstance-level: OpenEvolve, ShinkaEvolve, EvoX. Strategy-level: A-Evolve, GEPA, Meta-Harness. All seeded from the same three harnesses (B10–B12).
2backbonesOpus 4.8 (frontier) and gpt-oss-120b (open-weight). Metrics: pass@k, pass rate PR@k and resolution rate RR@k, with k = 5 on Terminal-Bench and 3 elsewhere.
Backbone
Backbone: Opus 4.8

Leaderboard. 17 harnesses on Terminal-Bench 2.1, PaperBench and DeepSWE with Opus 4.8. Click a column to sort.

IDDesignTerminal-Bench 2.1PaperBenchDeepSWE
Minimal harness
0Minimal harnessexpert 85.2  74.9±2.7 68.0±3.1 65.7±1.5 15.0±0.0 23.9  20.3±3.1 13.6±2.7
State-of-the-art harnesses (expert-designed)
1Clineexpert 87.5  80.0±2.1 71.4±3.0 7.0±0.3 1.7±1.7 16.8  14.0±3.8 5.6±2.5
2DeepAgentsexpert 83.0  76.8±2.2 70.9±2.6 75.1±0.5 31.7±3.3 77.9  85.7±2.9 54.9±4.1
3Gooseexpert 89.8  83.0±1.9 74.5±2.7 66.4±1.3 8.3±1.7 27.4  24.2±5.1 9.1±3.2
4Mini-SWE-Agentexpert 89.9  83.2±2.0 73.9±2.9 57.2±0.7 3.3±1.7 43.4  65.8±2.9 25.7±3.4
5OpenCodeexpert 87.5  80.9±1.8 72.7±2.6 64.8±2.3 8.3±4.4 61.9  84.8±2.1 40.4±4.1
6OpenHandsexpert 88.6  79.4±2.2 72.5±2.7 66.0±0.5 13.3±4.4 76.1  91.6±1.3 53.7±4.1
7Qwen-Codeexpert 88.6  81.9±2.2 73.9±2.9 62.4±0.4 8.3±1.7 65.5  86.3±2.3 38.6±4.2
8Terminus-2expert 86.5  80.0±2.3 68.3±3.3 61.4±0.1 3.3±1.7 49.6  83.5±2.3 30.4±3.6
Automatically discovered harnesses (SoTA evolutionary search vs. MILO)
9Best-of-3 (seed)expert 87.5  79.1±2.0 74.1±2.4 74.5±0.8 26.7±6.0 77.0  90.6±2.2 59.0±3.8
10GEPA (prompt)auto 92.0  82.8±2.2 78.6±2.5 79.3±0.5 33.3±4.4 73.5  83.6±3.3 53.7±4.1
11GEPA (optimize anything)auto 87.5  79.1±2.0 74.1±2.4 82.3±0.3 35.0±2.9 77.9  89.4±2.0 58.7±3.8
12A-Evolve (skill, memory)auto 92.0  84.4±2.2 78.2±2.7 74.5±0.8 26.7±6.0 66.4  90.9±3.6 57.1±4.0
13OpenEvolveauto 88.6  81.9±2.0 76.1±2.6 71.1±0.6 23.3±1.7 77.9  91.8±1.9 53.7±4.2
14ShinkaEvolveauto 89.8  84.1±2.0 77.7±2.6 71.4±1.7 15.0±2.9 78.8  89.8±2.1 56.0±4.1
15EvoXauto 89.8  83.0±2.0 76.8±2.5 73.7±0.7 25.0±2.9 78.8  90.0±2.2 58.7±4.1
16Meta-Harnessauto 92.0  82.9±2.2 78.2±2.7 85.6±1.9 45.0±2.9 77.0  90.6±2.2 59.0±3.8
★MILOauto 93.2  90.4±1.3 86.1±2.0 88.6±0.5 55.0±5.0 86.7  96.4±0.9 69.3±3.4
Backbone: gpt-oss-120b

Leaderboard. 17 harnesses on Terminal-Bench 2.1, PaperBench and DeepSWE with gpt-oss-120b. Click a column to sort.

IDDesignTerminal-Bench 2.1PaperBenchDeepSWE
Minimal harness
0Minimal harnessexpert 35.2  38.6±2.1 18.6±2.5 14.4±2.4 0.0±0.0 0.0  0.0±0.0 0.0±0.0
State-of-the-art harnesses (expert-designed)
1Clineexpert 36.4  47.9±1.8 23.4±2.4 7.0±0.6 0.0±0.0 0.0  0.0±0.0 0.0±0.0
2DeepAgentsexpert 29.5  36.8±2.3 15.5±2.4 13.8±0.3 0.0±0.0 0.0  1.2±0.8 0.0±0.0
3Gooseexpert 40.9  47.6±2.1 24.1±2.6 15.0±0.6 0.0±0.0 0.0  0.5±0.5 0.0±0.0
4Mini-SWE-Agentexpert 44.3  53.6±1.8 29.8±2.5 8.4±0.5 0.0±0.0 0.0  2.9±1.1 0.0±0.0
5OpenCodeexpert 34.1  40.0±2.1 16.8±2.6 11.1±0.8 0.0±0.0 0.0  2.0±1.0 0.0±0.0
6OpenHandsexpert 52.3  48.6±2.4 31.2±3.0 17.1±0.9 0.0±0.0 0.9  11.0±2.3 0.3±0.6
7Qwen-Codeexpert 26.1  37.2±1.9 16.6±2.2 9.7±0.6 0.0±0.0 0.9  3.3±1.4 0.3±0.6
8Terminus-2expert 29.5  42.4±2.0 17.5±2.4 6.2±0.4 0.0±0.0 0.9  1.6±0.9 0.3±0.6
Automatically discovered harnesses (SoTA evolutionary search vs. MILO)
9Best-of-3 (seed)expert 39.8  45.3±2.5 23.4±2.5 13.0±0.7 0.0±0.0 0.0  0.8±0.6 0.0±0.0
10GEPA (prompt)auto 39.8  44.2±2.3 22.7±2.6 14.7±1.0 0.0±0.0 0.0  0.8±0.6 0.0±0.0
11GEPA (optimize anything)auto 39.8  45.3±2.5 23.4±2.5 15.5±0.7 0.0±0.0 0.0  0.8±0.6 0.0±0.0
12A-Evolve (skill, memory)auto 43.2  44.0±2.4 25.9±2.6 13.0±0.7 0.0±0.0 0.0  0.8±0.6 0.0±0.0
13OpenEvolveauto 39.8  51.1±2.2 25.2±2.5 13.1±1.1 0.0±0.0 0.0  0.9±0.7 0.0±0.0
14ShinkaEvolveauto 44.3  53.5±2.1 28.9±2.7 12.7±0.2 0.0±0.0 0.0  0.4±0.5 0.0±0.0
15EvoXauto 44.3  51.5±2.2 29.5±2.5 13.0±0.7 0.0±0.0 0.0  0.6±0.5 0.0±0.0
16Meta-Harnessauto 40.9  51.0±1.9 26.8±2.3 17.8±1.9 0.0±0.0 0.0  0.8±0.6 0.0±0.0
★MILOauto 60.2  57.1±2.0 34.1±3.1 21.6±0.6 0.0±0.0 1.8  15.6±1.7 0.6±0.8

Bold = column best, underline = second best. Grey italics: the search found no fitness gain over Best-of-3, so the seed is reported. The thin bar under each value shows its magnitude on a 0–100 scale. ± are task-clustered 95% confidence intervals (standard error over runs on PaperBench). Best-of-3 is the best of the three seeds without any search.

F1Off-the-shelf harnesses do not consistently beat the minimal harness.

With Opus 4.8, seven of eight state-of-the-art harnesses underperform the minimal harness on PaperBench RR@3, while six exceed it on DeepSWE by up to 41.3 points. Gains also fail to transfer across backbones: all eight beat the minimal harness on Terminal-Bench with Opus 4.8, but four fall below it with gpt-oss-120b.

F2MILO discovers higher-performing harnesses than SoTA evolutionary search.

With Opus 4.8 it improves RR over Best-of-3 by +12.0, +28.3 and +10.3 on the three benchmarks, against +4.5 (GEPA), +18.3 (Meta-Harness) and 0 for the best prior method. With gpt-oss-120b it improves PR by +11.8, +8.6 and +14.8, against +8.2, +4.8 and +0.1.

F3MILO explores more of the harness space.

Methods that mutate only prompts (GEPA) or skills and memory (A-Evolve) gain at most +6.6 RR@3 on PaperBench. Whole-harness methods often settle for targeted edits: ShinkaEvolve’s and EvoX’s best harnesses differ from the seed only in prompts and rubrics. MILO’s best harness rewrites the control flow: produce a minimal valid solution first, triage as the deadline nears, and stop only after independent verification.

F4MILO raises accuracy while lowering inference cost.

Its Terminal-Bench harness with Opus 4.8 uses 0.74× the tokens of Best-of-3 while gaining +12.0 RR@5. With gpt-oss-120b every prior method’s best harness costs more than Best-of-3 (EvoX spends 1.46× the tokens for +6.1), whereas MILO gains +10.7 at 0.70× the tokens.

Where MILO sits among 17 harnesses

Every harness on each benchmark, by pass rate or resolution rate. The tinted band spans the full range from the weakest to the strongest harness; MILO is the star at its right end.

Minimal harnessExpert-designed SoTA harnessBest-of-3 seedAutomatically discovered (SoTA search)MILO (ours)
Opus 4.8
Pass rate
020406080100Pass rate (%)Terminal-Bench 2.1PR@590.474.9PaperBenchPR@388.67.0DeepSWEPR@396.414.0
Resolution rate
020406080100Resolution rate (%)Terminal-Bench 2.1RR@586.168.0PaperBenchRR@355.0DeepSWERR@369.35.6
gpt-oss-120b
Pass rate
020406080100Pass rate (%)Terminal-Bench 2.1PR@557.136.8PaperBenchPR@321.66.2DeepSWEPR@3715.6
Resolution rate
020406080100Resolution rate (%)Terminal-Bench 2.1RR@534.115.5PaperBenchRR@3150.0DeepSWERR@3130.6
Hover a mark for the harness name. Coincident harnesses pack around the row line so none is hidden. Pass rate gives per-attempt partial credit; resolution rate counts full solves, which with gpt-oss-120b are rare on PaperBench and DeepSWE. Cells the paper greys out (no gain over the seed) are omitted since they duplicate Best-of-3.

Accuracy against cost per attempt

Mean tokens and wall-clock time per attempt against resolution rate, with Opus 4.8. Numbers are the harness IDs from the leaderboard. The faint red rules mark MILO: nothing sits to its right.

Terminal-Bench 2.10.40.60.811.260708090Tokens (M)RR@5 (%)01234567↑ 1.9M89101213141516MILO510152060708090Latency (min)RR@5 (%)0123456789101213141516MILODeepSWE05101520250204060Tokens (M)RR@3 (%)0123456789101213141516MILO01530450204060Latency (min)RR@3 (%)0123456789101213141516MILO
On Terminal-Bench, MILO reaches the highest resolution rate at a token budget below every automatically discovered baseline and below its own seed, at the price of longer wall-clock time. Qwen-Code (7) sits off the token scale at 1.9M tokens per attempt. On DeepSWE, MILO’s +10.3 over Best-of-3 comes at essentially the seed’s token cost.

Every mechanism helps; the orchestrator helps most

Search-mechanism ablation on Terminal-Bench 2.1 with Opus 4.8. Each configuration adds one mechanism to the one on its left, so neighbouring columns isolate that mechanism.

020406080100RR@5 on Terminal-Bench 2.1 (%)76.4A79.1+2.7B79.3+0.2C80.5+1.2D86.1+5.6EMutator agent–✓✓✓✓Lineage memory––✓✓✓Multiple islands–––✓✓Orchestrator––––✓
(A) a single LLM call proposes each child from the current best over a flat archive; (B) an agentic mutator that inspects failure evidence; (C) the hierarchical lineage memory; (D) islands evolving in parallel; (E) the orchestrator, i.e. full MILO. RR@5 rises at every step, 76.4 → 79.1 → 79.3 → 80.5 → 86.1.
Robust to a tight output cap · Terminal-Bench 2.1, Opus 4.8, 4K tokens per call
34.4%81.1%+46.7
RR@5 of the best seed (Best-of-3) versus the harness MILO evolves under the same cap. Without a cap the same run reaches 86.1%.
An open-weight model, unlocked · DeepSWE, gpt-oss-120b
0.8%15.6%20×
Pass-rate of the initial harness versus MILO’s. Every prior search method stays below 1%.
Reusability · Frontier-Bench, without further search
5.2%13.8%2.7×
RR@3 of Mini-SWE-Agent, the strongest SoTA harness on Terminal-Bench by pass rate, versus MILO’s Terminal-Bench-evolved harness run unchanged.
05

Search dynamics

Why does MILO keep improving where other methods plateau? Its run on Terminal-Bench 2.1 with gpt-oss-120b shows the mechanism: every island stalls for nine rounds, and the orchestrator’s interventions pull the population out. The interactive demo in section 03 steps through this same run round by round; here is the summary view from the paper.

Stall, diagnose, intervene

Population pass-rate per island (top) and the orchestrator’s interventions (bottom) over 28 rounds. Vertical guides mark orchestrator invocations; the amber band is the stall.

Island 1 (seed 45.0)Island 2 (seed 44.0)Island 3 (seed 39.2)Population best so farGraft admittedGraft rejected
all islands stalled (R6–R14)4044485256600481216202428Population pass-rate (%)Evolution round45.044.039.259.9Isl. 1 56.7Isl. 2 55.6Isl. 3 59.9Orchestrator events123cc→cxcx→d:gcc→d:qcc→cxd:q→cc
All three islands are flat in rounds 6–14. Early grafts are rejected but persist in memory as negative examples. The admitted graft from island 3 into island 1 at round 15, mutator reassignments, and the grafts from island 1 at round 21 resume progress. The weakest seed, island 3 (39.2%), sets the population best of 59.9% at round 27: preserved diversity pays off. Mutator codes: cc Claude Code (Opus 4.8), cx Codex (GPT-5.5), d:g DeepAgents (GPT-5.5), d:q DeepAgents (Qwen3-Coder-480B).

MILO against SoTA evolutionary search, round by round

Best-so-far on the search split for each method, same seeds, same 72-hour budget. MILO’s curve is repeated faintly in every panel for reference.

Baseline search methodMILO (ours)MILO, repeated for reference
Pass-rate
444852566007142128GEPA (prompt)45.6%07142128OpenEvolve51.1%07142128ShinkaEvolve52.9%444852566007142128EvoX52.2%Evolution round07142128Meta-Harness51.0%Evolution round07142128MILO (ours)59.9%Evolution roundPass-rate on the search split (%)
Tokens
8111407142128GEPA (prompt)11.7K07142128OpenEvolve13.2K07142128ShinkaEvolve15.0K8111407142128EvoX15.6KEvolution round07142128Meta-Harness13.1KEvolution round07142128MILO (ours)7.5KEvolution roundMean tokens per attempt (K)
Latency
68101207142128GEPA (prompt)8.3 min07142128OpenEvolve7.2 min07142128ShinkaEvolve7.4 min68101207142128EvoX9.2 minEvolution round07142128Meta-Harness9.1 minEvolution round07142128MILO (ours)8.0 minEvolution roundMean latency per attempt (min)
Every baseline makes a few early gains and then plateaus, never recovering. MILO stalls too, then recovers, and ends with the highest pass-rate at the lowest token cost: 7.5K tokens per attempt against 10.7K for the seed and 11.8–15.6K for the baselines.
06

Reusability and scientific discovery

A harness evolved on one benchmark should not overfit to it, and a strategy-level framework should be able to drive instance-level discovery. MILO passes both tests.

Transfer to Frontier-Bench no further search

The generality constraint rejects task-specific hard-coding during search, and a manual inspection of the evolved harnesses found none. As a stronger test, the Terminal-Bench-evolved harness was run unchanged on Frontier-Bench, the harder successor with 70 disjoint tasks and time limits up to eight hours. In RR@3 it beats Best-of-3 by +3.8 and Mini-SWE-Agent, the strongest SoTA harness on Terminal-Bench by pass rate, by +8.6.

Harnesspass@3PR@3RR@3
Mini-SWE-Agent 12.9  44.0±3.3 5.2±2.8
Best-of-3 21.4  48.4±2.7 10.0±3.4
MILO (TB2.1-evolved) 24.3  51.4±2.0 13.8±3.2

Reusability on Frontier-Bench (RR@3 etc.). Bars are on a 0–100 scale; bold marks the column best.

New records on EinsteinArena instance-level MILO, Opus 5 backbone

EinsteinArena is a leaderboard of open mathematical problems. Here MILO evolves the harness of an agent that edits optimizer programs: each task pairs a frozen leaderboard construction with a parent optimizer, the agent writes a revised optimizer, and the verifier scores its gain. Three problems yielded constructions below the prior best, each confirmed by the arena’s own verifier and each clearing the arena’s minimum improvement for a new #1. Bold digits mark where MILO’s value departs from the prior best.

Open problemObjective (lower is better)AlphaEvolveTTT-DiscoverEvoXPrior best (arena leader, Sep 2026)MILO (ours)Margin
Erdős minimum overlapminimize maxk ∫ h(x)(1 − h(x+k)) dx0.3809240.3808753—0.3808586CodexProLong, 2026-08-150.38085681.8 × 10⁻⁶min. for a new #1: 10⁻⁷
First autocorrelation inequalityminimize max(f⋆f) / (∫f)², f ≥ 01.50321.5028629—1.50274365CodexProLong, 2026-08-141.502743605.1 × 10⁻⁸min. for a new #1: 10⁻⁸
Third autocorrelation inequalityminimize |max(f⋆f)| / (∫f)², f signed1.4557—1.45581.4508066Poolish, 2026-08-231.44888601.9 × 10⁻³min. for a new #1: 10⁻⁵

This search is replayed day by day in the Demos section, alongside the harness-search demo.

All three are minimization problems, so lower is better. AlphaEvolve and TTT-Discover values are the arena’s baseline entries; EvoX’s third-autocorrelation value is its published result on the arena’s discretized objective. Prior best is the live arena leader, unchanged between 2026-09-09 and 2026-09-22. Two independent verifier implementations agree to within 2×10−16.

07

Where MILO sits

LLM-driven evolutionary search, sorted by how much of its own strategy it adapts online. Most methods fix their search strategy in advance; the few that adapt it are mostly instance-level.

Stage 0
Open-loop test-time scaling
No executed outcome feeds the next attempt, so gains cannot compound.
Best-of-NChain-of-ThoughtTree-of-Thoughts
Stage I
Iterative refinement
Executed feedback steers one greedy lineage; no diversity-preserving population.
ReflexionEurekaAIDESelf-HarnessA-EvolveAHE
Stage II
Population search
An archive preserves diversity, but every search knob follows a preset schedule.
ELMFunSearchAlphaEvolveOpenEvolveAutoHarnessGEPAMeta-Harness
Stage III
Adaptive strategy
One or more components are retuned online, but by a fixed, hand-coded rule.
ShinkaEvolveEvoControlAdaEvolveCORALDarwinX
Stage IV
Self-adaptive strategy
The adaptation rule is no longer fixed: the search rewrites its own strategy when progress stalls.
EvoXMILO (ours)

blue strategy-level discovery (evolves a solver such as a harness); plain chips are instance-level (evolve a single solution).

08

BibTeX

@article{jana2026milo,
  title   = {{MILO}: Automated Harness Discovery via Orchestrated Multi-Agent Evolution},
  author  = {Jana, Prithwish and Goswami, Mononito and Liu, Hao and Li, Xinyu and
             Huang, Langlin and Huang, Zhehui and Huang, Zhishen and Bl{\"o}baum, Patrick and
             Deoras, Anoop and Jain, Purak and Kanakaris, Nikos and Genc, Sahika},
  journal = {arXiv preprint arXiv:2609.38349},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.38349}
}

Acknowledgements. The authors thank Amir Tahmasbi, Karen Hovsepian, Xing Niu, Lecheng (Jerry) Kong, Like Hui and Narayanan Sadagopan for helpful feedback and discussions throughout the project.