Skip to content

routing.yaml

routing.yaml sits next to config.yaml and decides which agent, model and effort runs a session, and what the intake agent is told about your preferences. Gluon writes this file when there is none and replaces it with a newer default only while you haven’t edited it. gluon routing check validates yours, gluon routing path says where it is and gluon routing default prints the file below.

# Gluon routing: which agent runs a session. Edit freely.
#
# Model facts (efforts each model accepts, its default effort, its Bedrock id)
# are in Gluon's catalog, not here. `gluon --help` lists them. Which models
# serve which level, and in what priority, is here: `rank`.
#
# How a session is routed:
# 1. The intake agent writes the spec, then picks every type the session
# covers (e.g. debug + test).
# 2. For each type it judges, using the type's example lists below, how far
# to move the model (-2 to +2 levels), the effort (0 to +2 steps) and
# whether to switch the mode.
# 3. The session gets the strongest of each across its types:
# plan > build > explore, the highest level, the highest effort.
# 4. The first connected, allowed model in that level's `rank` list that can
# give the needed effort wins. If the level has none, the next level up,
# with one step less effort per level it rounded up.
# 5. If one of your `prefer` notes applies, the intake agent either puts a
# harness first or folds it into the steps, and quotes the note.
version: 3
# Which models serve each level, in priority order (first = preferred), as
# harness/model. A model belongs to the level it's listed under; a model in no
# list is never routed to, but you can still name it. `gluon routing check`
# flags typos.
# light small, mechanical, well-specified edits; quick answers
# standard everyday, well-bounded work: features, fixes, reviews
# strong harder work: some ambiguity, several modules, tricky bugs
# extra ambiguous, cross-cutting or hard-to-debug work; design
# frontier the hardest and longest tasks
rank:
light: [claude-code/haiku, codex/gpt-6-luna, kimi-code/kimi-k2.7-code]
# standard: cheapest first, by cache-read price per 1M tokens (Oct 2026):
# Contributor $0.002 (opt-in), DeepSeek Flash $0.006, Gemini 3.8 Flash $0.075,
# Muse Spark 1.3 $0.15, Kimi K3 $0.30, Grok 4.7 $0.50
standard: [opencode/muse-spark-1.3-contributor, opencode/deepseek-flash, antigravity/gemini-3.8-flash,
opencode/muse-spark-1.3, kimi-code/kimi-k3, grok-build/grok-4.7]
strong: [claude-code/sonnet, codex/gpt-6.1-sol]
extra: [claude-code/opus]
frontier: [claude-code/fable, codex/gpt-6-astra]
# Hard limits. These are never crossed, not even in alternatives.
limits:
never_harnesses: []
never_models: [] # model or harness/model
max_model: frontier # light | standard | strong | extra | frontier
max_effort: max # low | medium | high | xhigh | max. Max is used only when you pin it (e.g. claude-code/opus@max) or when it is a model's own default (Muse, DeepSeek Flash)
# Muse Spark 1.3 Contributor is much cheaper than Muse Spark 1.3, but it shares
# your session's code and prompts with Meta. Off unless you set this to true.
allow_muse_contributor: false
# Soft preferences, in plain words. Each one can mean "within each level, try
# this harness first" or "a step up or down in model or effort" for the
# sessions it describes. They never override `limits`. One note per line (put
# a note in quotes if it contains a colon):
prefer:
# - Use codex for test sessions.
# - Anything under src/pty is delicate, so go one level stronger.
# Instructions for the intake agent: how it talks, what it always or never
# asks, and what every spec must say. Write them as a block, for example:
# instructions: |
# Talk to me in Hebrew.
# Don't ask about tests; I always want them.
instructions:
# Session types. To add a type, copy an entry and change it.
# means what the session is about; the intake agent classifies by this
# mode explore (read-only) | plan (agree a plan first) | build
# model light | standard | strong | extra | frontier (see rank): where routing starts
# effort low | medium | high, relative to the chosen model's own default:
# medium = the model's default, high = one step above it (every
# built-in type uses medium)
# The lists below are EXAMPLES for the intake agent, not a checklist. It
# judges the steps itself; the examples show what kind of thing moves each one.
# lighter_model_when the model can go down (to -2; never below light)
# stronger_model_when the model should go up (to +2)
# more_effort_when the effort should go up (to +2)
# plan_when / build_when / explore_when
# the type should switch to that mode
# ask what the spec needs from the developer when the repo doesn't say it
# use optional; always use this agent for the type, e.g. codex/gpt-6-luna@low
# (empty = route normally)
# Every type lists every field, even when empty, so they all read the same.
#
# Model or effort? They solve different problems:
# A stronger model is for work that is HARD TO FIGURE OUT: the answer needs
# better priors and judgment because the space can't be searched, either
# because it is too big or because nothing gives quick feedback. Examples are
# ambiguous requirements, design and architecture choices, novel algorithms,
# concurrency, security, and bugs you can't reproduce. In benchmarks the
# stronger model mainly helps on knowledge-heavy work (HLE, SciCode), and
# across tiers it sets the ceiling: Luna at max never reaches a top model's
# medium on agentic tasks.
# More effort is for work that MORE WORK CAN SOLVE: more reading and
# searching, more checks, more tests, more runs. Examples are long call chains
# to trace, many call sites to update, weak tests to compensate for, flaky
# failures to rerun, and claims to verify. Among the top models, effort moves
# agentic results (Terminal-Bench 4.0) more than switching model does.
# Up to xhigh, each effort step costs 1.3-3.6x the thinking. Max costs 3-38x
# more for 0-4 points, so it is used only when you ask for it, or when it is
# a model's own default (Muse, where max costs no more latency than xhigh).
# Every model starts at its own sweet spot (defaultEffort in Gluon's
# catalog); +2 effort is for when there is clearly a lot to check.
# A lighter model is for work that is SMALL AND SPECIFIED: the exact change is
# known, it is mechanical, or it stays in one file. Small size never makes a
# hard problem easy: if something in it is hard to figure out, don't go down.
# Something that is both hard to figure out and a lot to check raises both.
types:
understand:
means: Explain something that already has an answer, such as how this code, a library or a protocol works. Deciding an open question is research; judging whether something is right is review.
mode: explore
model: light
effort: medium
lighter_model_when: []
stronger_model_when: [subtle or theory-heavy subject]
more_effort_when: [many modules or sources to read, long call chains to trace, code to run to confirm the answer, "a decision hinges on the answer, so it must be verified"]
plan_when: []
build_when: []
explore_when: []
ask: [what they will do with it, their level]
use:
research:
means: Act as a research scientist to settle an open question. Form hypotheses, compare approaches, weigh architectures, test feasibility, and look for evidence against your own conclusion, or think it through together with the developer. The output is findings with their evidence and confidence, and a recommendation. Building the thing is feature; improving this code against a metric is optimize.
mode: explore
model: strong
effort: medium
lighter_model_when: []
stronger_model_when: [candidates to discover, architecture-scale trade-offs, novel domain]
more_effort_when: [experiments on their code or data, many candidates to compare, rigorous analysis wanted, "high-stakes decision, so claims must be verified"]
plan_when: []
build_when: [experiments on their code or data, a report file in the repo]
explore_when: []
ask: [the decision it informs, 'what "best" means', reading only or experiments]
use:
review:
means: Judge existing changes (a diff, branch or PR). The output is feedback; nothing is edited.
mode: explore
model: standard
effort: medium
lighter_model_when: [small diff, style-only focus]
stronger_model_when: [security or concurrency, design judgment needed]
more_effort_when: [large diff, "weak tests, so behaviour must be checked by hand", "data-loss risk, so data paths must be traced"]
plan_when: []
build_when: []
explore_when: []
ask: ["focus: correctness, security, style or tests"]
use:
feature:
means: Add or change behaviour, including a new project or module from nothing, and one-off scripts or data tasks (e.g. a script that cleans a CSV, a query against a database).
mode: build
model: standard
effort: medium
lighter_model_when: [exact change specified, single file, mechanical change, template or generator given]
stronger_model_when: [ambiguous behaviour, algorithmic novelty, open stack choices, cross-cutting design decisions, security-sensitive code]
more_effort_when: [spans layers or modules, weak tests, many constraints to preserve, "auth, payments or data", security-sensitive foundation]
plan_when: [ambiguous behaviour, spans layers or modules, hard to reverse, starting from nothing]
build_when: []
explore_when: []
ask: [scope limits, behaviour the repo doesn't settle, "for something new: stack, structure and conventions"]
use:
debug:
means: Restore intended behaviour, including "fix the failing CI" and "run the tests and fix what fails". For PR review comments, use the types the comments call for.
mode: build
model: standard
effort: medium
lighter_model_when: [cause already known, single file]
stronger_model_when: [concurrent or distributed, can't reproduce locally, security-sensitive code]
more_effort_when: [unknown cause, fails only sometimes, large surface to search, "auth, payments or data"]
plan_when: [the fix needs a design change]
build_when: []
explore_when: [diagnose only]
ask: [expected vs actual, fix or diagnose only]
use:
refactor:
means: Same behaviour, better structure. When a dependency or platform change drives it, it is chore.
mode: build
model: standard
effort: medium
lighter_model_when: [mechanical rename or move, single file]
stronger_model_when: [target structure to design, stateful or concurrent code]
more_effort_when: [spans many modules, public API with many callers, weak or no tests]
plan_when: [spans many modules, public API changes]
build_when: []
explore_when: []
ask: [behaviour and API to preserve]
use:
optimize:
means: Same behaviour, better on a metric (speed, memory, cost, size).
mode: build
model: standard
effort: medium
lighter_model_when: [single known hotspot with a benchmark]
stronger_model_when: [architecture change rather than tuning, no clear hotspot]
more_effort_when: [no benchmark, noisy measurements, correctness-sensitive code]
plan_when: [architecture change rather than tuning]
build_when: []
explore_when: []
ask: [metric, target, acceptable trade-offs]
use:
test:
means: Write, improve or run tests, including running them only to report the results. Fixing the code that fails is debug.
mode: build
model: standard
effort: medium
lighter_model_when: [similar tests to mirror, run and report only]
stronger_model_when: ["hard-to-test code (time, concurrency, heavy mocking)", no conventions to mirror]
more_effort_when: [flaky vs real failures to tell apart, tests must fail before a fix, tests that need many runs to trust]
plan_when: []
build_when: []
explore_when: []
ask: ["level: unit, integration or e2e", may production code change]
use:
docs:
means: Write or update documentation.
mode: build
model: light
effort: medium
lighter_model_when: []
stronger_model_when: [concepts that need deep understanding to explain]
more_effort_when: [broad surface to read and check against the code, "public-facing, so every claim checked against the code"]
plan_when: []
build_when: []
explore_when: []
ask: [audience, scope]
use:
chore:
means: Keep the project running and moving. This covers dependency and framework upgrades, framework migrations (wide but mechanical) and database or data migrations (risky, touch data). It also covers CI, build and config, infrastructure (Terraform, Kubernetes, deploy scripts), merge conflicts and rebases, and housekeeping. When code quality drives the change rather than a dependency or platform, it is refactor.
mode: build
model: standard
effort: medium
lighter_model_when: [patch or minor version bump, single config change, text-only conflicts]
stronger_model_when: [conflicting logic not just text, poorly documented upgrade path]
more_effort_when: [major version jump, many call sites, weak tests, multi-platform builds, data migration, touches production infrastructure]
plan_when: [major version jump, data migration, touches production infrastructure]
build_when: []
explore_when: []
ask: [what must keep working, how to roll back]
use:
other:
means: Only when no type above fits. The spec's goal says what kind of session it is.
mode: build
model: standard
effort: medium
lighter_model_when: [exact change specified, single file]
stronger_model_when: [ambiguous goal, unfamiliar domain]
more_effort_when: [spans many modules, weak tests, "auth, payments or data"]
plan_when: [ambiguous goal, hard to reverse]
build_when: []
explore_when: [no changes wanted]
ask: [what a good outcome looks like]
use: