On Sept. 26, three researchers shared a new system called Planner-as-Router (PaR) that cuts AI agent costs by 44 percent [1]. Instead of building a separate tool to pick models, PaR assigns a cheap or smart model to each step while making the initial plan [1]. Top frontier models cost 25 times more per word than small ones, and that price gap piles up fast when tools link many steps together [1].
The Router Moves Upfront
I have watched agent builders fight high cloud bills all month, and the real shift here is exactly where the routing choice happens. Most teams build a separate router that checks one step at a time to pick a model [1]. PaR flips this by assigning each small job a model tier — small, mid, or frontier — before any agent starts working [1].
Why it matters: Picking the right model is now a core design choice, not just a quick fix to save money later. Teams do not need to train a separate router model or gather data to teach it [1]. This lets the system see how tasks link together right from the start [1].
By the numbers:
-
25×: The cost jump from a small model to a top-tier frontier model for the exact same amount of text [1].
-
44%: The total cost drop PaR hits compared to using the biggest model for every step [1].
-
1,157: The total number of tests run across eight routers and 54 business tasks [1].
-
3: The number of test seeds used to grade tasks by running code against live Structured Query Language (SQL) and MongoDB databases [1].

A Fixed Backup Plan
What strikes me here is how this open-source system looks like classic software design, rather than a complex machine learning trick [1]. The main planner gives a tier to every small job right away, and the dispatcher runs them in the exact order they need to happen [1]. If a job fails, the dispatcher simply bumps it up to a smarter model and tries again [1].
What I’d watch:
-
Plan-time sorting: Whether picking models before the work starts holds up well on messy, real-world jobs [1].
-
The escalating dispatcher: How this fixed backup plan — bumping a failed job to a smarter model — compares to messy, custom retry loops [1].
-
The compounding penalty: The paper warns that using cheap models early on might quietly multiply errors across long chains of tasks [1]. I want to see builders test this at scale before they trust the 44 percent savings [1].
The catch: The 44 percent cost cut acts as a ceiling, not a floor, because it comes with a 2.9-point drop in accuracy [1]. The paper honestly notes that several of these gaps sit inside a wide ±6-point range of doubt [1]. My read: The real win is moving the model choice upfront to see how tasks link together, while finally naming the risk of piled-up errors that most cheap-router pitches ignore [1].
Go deeper:
-
The Big Shift: A new paper called Planner-as-Router (PaR) cuts AI agent costs by 44 percent by assigning a small, mid, or frontier model to each step during the planning phase.
-
Why It Matters: Moving this choice upfront removes the need to train and keep a separate router model, making cost control a core design choice instead of a late fix.
-
What I’d Watch: Whether the cost savings hold up in the real world without causing a chain of hidden errors.
-
Plan-time tiering: Giving a model tier to each small job before the work starts so the system sees how tasks link together.
-
Escalating dispatcher: Running tasks in order and automatically bumping failed jobs to a smarter model tier.
-
Compounding penalty: The risk that using cheap models for early jobs quietly multiplies errors across long chains of tasks.
-
The Catch: The 44 percent savings cost 2.9 accuracy points, and several test results sit within a wide ±6-point range of doubt.
Related reading
-
OpenAI just made the agent loop a commodity — more on AI Stack & Tool TCO
-
MCP Skills Extension — more on Autonomous & Agentic Workflows:
Sources
[1] Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows — arXiv:2609.32917 (submitted Sept 26, 2026) — https://arxiv.org/abs/2609.32917