10/10
Main development gate, 5 runs per repo
12/12
Four smaller gates, 3 runs each
0/9
Frozen held-out set of then-unseen repos
3 of 3
Held-out repos that later passed once in development
Development closed · October 2026
Give it a Flask repo and a recipe; it plans the target architecture, rewrites existing files and creates the new modules the migration requires, verifies against the repo's own tests in a network-off Docker sandbox, recovers from failures under bounded budgets, and reports honestly, including when it fails.
AI migration demos usually stop once the model has written the code. Portage only calls a migration done when the repo's own tests pass. It moves Flask apps to FastAPI: it plans the new modules the migration needs, writes the code, runs the tests in a sandbox with no network, and repairs what breaks. Progress is saved to Postgres after every step, and other coding agents can use the same sandbox through MCP.
It passed its development gates, then went 0 for 9 on three repos it had never seen. I publish both side by side. That gap is the main limitation, and I closed development in October 2026 rather than keep tuning against the same repos.
Below is the core durability claim, live: the worker is killed mid-migration, and a restarted worker resumes from the Postgres checkpoint instead of starting over.
reproduce: bash scripts/demo_kill_resume.sh · stricter: scripts/dod_check.sh
A submitted job runs this LangGraph graph, checkpointing to Postgres after every node: kill the worker mid-run and a restarted worker resumes from the last completed node. When Verify fails, Recover classifies the failure and routes back: targeted rollback + regenerate to Execute, replan to Plan for planner misses, or give up to Integrate once budgets are exhausted, reporting an honest red rather than a gamed green.
v1 ships one recipe: Flask → FastAPI. That target is deliberate. Routing decorators, request/response handling, blueprints → routers, error handlers, app factories, and ambient request context (g, session) need understanding, not mechanical rewriting; deterministic codemods cannot do this reliably. The architecture is recipe-pluggable; the evidence is recipe-specific by design. The capability that unlocked the hard repos: some migrations are unreachable by rewriting existing files. Flask's g/session have no FastAPI equivalent. A correct port needs a new request-context module, a test-compatibility surface, a rendering layer, and every consumer wired to them coherently. Portage plans those artifacts with a bounded architect call, freezes their contracts before generation, compiles the deterministic parts itself, and enforces that a framework-shaped capability is only valid when the plan owns and implements it, so a model can't reference a helper it wishes existed. A second wave, coherent-cut preservation, closed the gap the first one left open. One bad file inside an otherwise-correct migration used to trigger a full rollback of every file in its verification cut, so a single local mistake could sink a ten-file run. Recover now checkpoints the last coherent state before a targeted repair and restores that on failure instead of the whole migration, and one shared gate (caller, capability, import-direction, cycle, and contract checks) runs identically across every generation path: first draft, contract repair, and targeted repair alike. That is what took watchlist, a Flask-SQLAlchemy app that had never gone green, to autonomous 15/15, and pushed flaskr to a 5-for-5 reliability gate. Then the first frozen held-out evaluation supplied the correction. On three repositories that had never been migrated during development, Portage scored 0/9 strict green. The engine failed honestly: five trees restored coherently, four stayed migrated-but-red, zero were hybrid. But the recipe did not generalize. The project now has both halves of a credible result, strong development convergence and a measured unseen-repository gap, and publishes them together. A later forensic audit then corrected part of that failure explanation. The protected ws-example test file was byte-identical; a truncated read of a long protected test file caused a false oracle-integrity alarm. That correction changes the explanation, not the 0/9, because each sample was independently red for migration reasons. One core engine, two interfaces. Autonomous mode: `portage migrate <repo> --watch` drives the full graph. Co-pilot mode: Claude Code / Cursor call verify_patch_in_sandbox, repo_graph, and blast_radius over MCP, the same verified primitives the eval numbers were measured on. The dashboard is the observability and proof surface, not the front door.
The portage console script is a thin httpx client over the REST API; it never touches the DB or queue directly, the same boundary the dashboard respects. migrate --watch streams live task transitions; status, jobs, report and diff <job-id> (add --stat for a summary) cover inspection.
The exit-code contract belongs to the migration, not to every command that can print something about it. A completed status reflects the strict result: 0 only for an honest green, 1 for finished but not complete-and-green, 2 for usage or infrastructure. On a job that is still running, status can return 0 because the query succeeded, and jobs listing rows means the listing worked, not that a migration passed. Read the inspection commands as inspection: the verdict is the strict outcome they report, not their exit status.
The MCP server exposes the verified core so another AI agent can test its own work before writing to the caller's tree: verify_patch_in_sandbox copies the repo, applies a unified diff, runs the tests network-off, and returns structured pass/fail with failing test names, never mutating the caller's files. repo_graph and blast_radius give it structural awareness. Here is what a blast-radius query actually computes:
The blast_radius primitive in action: when db.py changes, Portage walks the structural code graph outward, direct callers first (hop 1) and then their dependents (hop 2), and selects only the tests that cover the impacted set. Verify uses this to iterate fast; the final honesty bar still runs the full suite. The same query is exposed to co-pilot agents as the blast_radius MCP tool.
Next.js App Router, REST only; the frontend never owns schema. Jobs list with launch form, job detail with live pipeline route, per-file diffs, and attempt tier/model timelines, and a public /eval leaderboard rendering per repo×scenario green rates, mean±variance, cost, and recovery straight from the runs/metrics tables.
A bounded architect call proposes new target-architecture modules; a deterministic contract compiler fills in what the engine already derives; contracts freeze before generation and bind every retry, escalation, replan, and resume. Created files get the same ordering, diffs, rollback, and cost accounting as rewrites.
A Flask-shaped capability (test_client, app_context, g, session) is accepted only when a frozen plan artifact owns and implements it, checked receiver-aware. "The model referenced a module it wished existed" becomes a pre-sandbox rejection, not a silent runtime failure.
LangGraph Postgres checkpointer after every node; worker lease with heartbeat. Kill the worker mid-run and a restarted one resumes: Ingest runs once, Execute skips already-applied files via content hashes.
When a failure traces to one file, Portage repairs only that file (a stray .decode() fixed for $0.011 without regenerating its ten-file batch). Since September, a failure it can't pin to one file stops the run with the draft, diff and checkpoints kept, instead of paying to regenerate a batch on a guess. That is cost control, not a measured gain in success rate.
A failed targeted repair restores the last known-coherent checkpoint, not the original sources, so one bad file can no longer roll back the nine correct ones beside it. The single highest-leverage fix in the project: it converted watchlist and flaskr from occasional greens into repeatable ones.
Test files are protected artifacts: names, assertions, raises/parametrize/skip structure and fixture lifecycles are frozen at Plan; only sanctioned plumbing may differ. The guard also had to survive an audit of itself. A held-out 0.75 integrity score looked like deleted tests, but the files were byte-identical and both readers had truncated a long protected test file. Fixed, regression-covered, and the run stayed red on its own merits.
Green cannot be gamed by skip-and-continue, empty diffs, or all-skipped suites. Report reloads task truth from Postgres, Integrate recomputes the diff, Verify requires passed > 0, and engine errors count against the score.
First N attempts use the driver model tier; later attempts escalate. Every attempt lands in attempts_log with tier, model, tokens, and USD, so "how often does escalation rescue?" is a SQL query.
Every LLM call's tokens and USD recorded per attempt, summed per job, averaged per eval cell, with retries, escalations, and architect calls included. Cost scales with recovery, and that relationship is part of the result.
Three repositories were frozen, baseline-vetted, and unseen at the time R5 v1 ran. It ran once from a pinned commit and scored 0/9. No failed sample was renamed, replaced, or rerun, and the result is published beside the development gates rather than behind them. All three became development inputs once those findings shaped the work, so their later greens cannot supply held-out evidence and the next generalization test needs a genuinely fresh frozen corpus.
Development performance and unseen-repository performance answer different questions. These are disclosed development milestones and one frozen held-out evaluation, not a claim that every regression gate is green on the current tree.
Green requires the full suite passing, every planned task done, zero skips, oracle integrity 1.0, and a tree_state of migrated: a run that recovery rolls back to original sources passes the original suite and still scores red.
| Evidence set | Result | What it means |
|---|---|---|
| Flaskr + Watchlist · disclosed K=5 v4 development gate | 10/10 green | Flaskr 24/24 tests and Watchlist 15/15 per successful run. Earlier gate generations stay disclosed. |
| Items, RESTX, Structural, Minimal · their K=3 development gates | 12/12 green | Four independently scoped gates, not the entire general-preservation suite. |
| Seven-repository autonomous development confirmation | 6/7 green | One sample per repository. Microblog was red. |
| Microblog accepted-plan replay | 4/4 tests · 26/26 tasks | Historical replay diagnostic. Excluded from autonomous rates. |
| Frozen R5 v1 · three then-unseen repositories, K=3 each | 0/9 green | Ran exactly once as the frozen held-out evaluation. The recipe had not generalized. |
| Post-R5 ws-example · development K1 | 42/42 tests · 5/5 tasks | One strict autonomous green after becoming a development input. |
| Post-R5 Silicon · development K1 | 34/34 tests · 14/14 tasks | One strict autonomous green after becoming a development input. |
| Post-R5 flask-email-login · development K1 | 18/18 tests · 15/15 tasks | One strict autonomous green after becoming a development input. |
| Latest Microblog preservation replay · after the post-R5 fixes | red · 1/26 tasks | Checked whether the fixes still held on a development repo. They did not; the original code was restored cleanly. 39 calls, $2.056. |
Each of the three post-R5 development K1 greens has tree_state=migrated and oracle integrity 1.0. They do not replace R5 v1 and do not count as held-out evidence.
The held-out set, in full
R5 v1 ran exactly once as the frozen held-out evaluation, from commit 3b25ee9 against corpus/heldout.toml, with the frozen offline sandbox and GPT-4o on both model tiers. It scored 0/9 strict autonomous green. Five runs restored coherently, four stayed migrated but red, and zero produced hybrid trees. Those failures showed that the development performance had not generalized.
| Unseen repo | Baseline | K=3 | Dominant failure |
|---|---|---|---|
| ws-example | 42/42 | 0/3 | generated test-client facade shadowed FastAPI route decorators; two samples stalled at 13/42 |
| silicon | 34/34 | 0/3 | invalid generated signatures; a raw FastAPI object constructed instead of the frozen facade |
| flask-email-login | 18/18 | 0/3 | architect missed the required context owner; the fallback left CSRF and mail providers as None |
The scoring machinery held even though the recipe failed
This is the part worth reading. Rejected cuts restored the original suite, and those restored passes contributed exactly zero migration score. Trees came back 4 migrated / 5 restored-coherent / 0 hybrid. All nine jobs produced durable reports with no missing run rows: 119 LLM calls, 19 recovery visits, $3.8643, architect acceptance 6/9.
A later forensic audit corrected part of the failure explanation. The protected ws-example test file was byte-identical; a truncated read of a long protected test file caused a false oracle-integrity alarm. That correction changes the explanation, not the 0/9. Each sample was independently red for migration reasons: two stalled at 13/42 behind the shadowed route decorators, and one restored the original tree with tasks still incomplete.
Once those findings influenced development, all three repositories became development inputs permanently. Their later successful runs cannot supply held-out evidence. The original result stays visible. Fresh held-out evaluation is parked, and any future test of generalization has to keep this 0/9 beside it and use repositories nobody has touched.
What convergence looks like when it works
flaskr, the canonical Flask tutorial app (templates + factory + auth + SQLite + Click CLI), went from never green in any grid to 24/24 tests, 12/12 tasks and zero recovery, and held 5/5 at K=5. One stored flaskr job cost $0.154 including retries and planning; that is a single run, not a rate across the full history. watchlist, a Flask-SQLAlchemy app that had never gone green, held 5/5 at 15/15 tests. Both needed new modules to exist; the engine designed and wired them. What made them repeatable rather than occasional was coherent-cut preservation.
Where it stopped
I closed active development on October 8, 2026. That is a stopping point I chose, not a claim the problem is solved: Portage is not production-ready or proven on repositories it hasn't seen. The code and evidence stay up, and I'll pick it back up only for a genuinely new idea, not more of the same tuning.
The fixes after the held-out run were meant to keep every earlier result passing, and Microblog's replay is still red, so that work stopped incomplete. Gates and the held-out run are from July 2026, the three later greens from August, the Microblog replay from September. No migration has been rerun since.
Every implementation choice, numbered, from the queue claim to the sandbox runtime.
Friction
Takeaways
portage migrate --watch · exit 0 = honest green
verify_patch_in_sandbox · repo_graph · blast_radius
jobs · recovery timelines · public /eval
Every moving part explained: the animated architecture, graph-node lifecycle, checkpoint and lease mechanics, sandbox anti-gaming predicates, artifact-producing plans, the recovery strategies, K-run eval methodology with non-claims, and the ten-category failure taxonomy with evidence.