Methodology
How role quality and benchmark-vs-cost Ranking Value are produced.
Two views of the same role score
Quality ignores price and ranks only the role-quality score. Balanced divides that same quality by CommandCode Cost per Task. Higher is better in both modes.
Role Quality = Universal Role ScoreRole Quality = ⅔ Universal Role Score + ⅓ CAI*Balanced Score = Role Quality / CommandCode Cost per TaskMinimize CommandCode Cost per Task subject to Role Quality ≥ 80% of the best available role qualityThe Budget lineup is only a recommended shortlist, not a third full-table ranking. It also applies the current-generation and model-family-diversity constraints used by the recommended lineups.
Free-model external-evidence overrides
The recommended lineups may include a CommandCode model that is still UNSCORED when it is explicitly free and there is enough external evidence to justify a narrow role recommendation. This is an editorial override only: unresolved models remain excluded from the full Quality and Balanced rankings until a verified Artificial Analysis family mapping exists.
Laguna S 2.1 is used as a free coding override. CommandCode exposes it as free with a 256K served context; Poolside documents a native 1M context and reports 70.2% Terminal-Bench 2.1, 78.5% SWE-bench Multilingual, 59.4% SWE-bench Pro and 40.4% DeepSWE. Poolside model page · model card and benchmark table.
Ox Alpha is used only in the Budget shortlist and marked EXPERIMENTAL. The CommandCode row is free with a 1M context window, but the model remains stealth and unresolved, so no synthetic benchmark score is assigned to it and it is not used in Balanced.
Portfolio assignment
The eight recommendation cards are selected as a portfolio rather than as eight independent winners. The same exact model may occupy two seats, and commercial rows sharing the same Artificial Analysis model identity may coexist. Diversity is enforced at the family level: each lineup must span 5–8 distinct model families, and no family may occupy more than two seats. Six, seven or eight families are all acceptable; five remains valid when deliberately justified. No manual lineup is forced toward a particular target.
Balanced currently requires Claude Opus 5, GPT-5.6 Luna and GPT-5.6 Sol somewhere in the portfolio. Their roles are not permanently fixed. The current assignment uses Muse Spark 1.2 Contributor for normal Implementer, Luna for Code Reviewer and Sol for Implementer Heavy. Sol is explicitly forbidden for normal Implementer because routine implementation is the highest-volume seat and cheaper strong models provide better economics. Claude Sonnet 5 remains eligible but is not mandatory.
Implementation routing should be interpreted as a cost-per-completed-task policy: use Muse for routine high-volume implementation, Luna for review, correct cheaply on the normal tier, and escalate genuinely hard or repeatedly failing work to Sol. The public benchmark does not yet measure retry probability directly, so this escalation rule is a documented operating policy rather than a measured benchmark component.
The coding adjustment applies only to Implementer, Implementer Heavy and Code Reviewer.
Security Reviewer tier policy
The numerical Security Reviewer score remains a universal proxy based on SciCode, GPQA, HLE, Non-Hallucination and LCR; it is not a dedicated cybersecurity benchmark. Public cyber benchmarks reviewed so far do not provide sufficiently broad, comparable model coverage to justify a universal Security* adjustment.
The recommendation cards therefore use explicit security tiers: Quality → Claude Fable 5, Balanced → Claude Opus 5, Budget → Muse Spark 1.2 Contributor. These are policy anchors and Security Reviewer changes always require human review. Octopus has no separate security-heavy role; implementer-heavy is a coding/implementation role.
CommandCode Cost per Task
Price per million tokens is not used as the primary denominator. Artificial Analysis measures the token mix consumed per Intelligence Index task. Those measured non-cached input, cache-read, cache-write and output tokens are repriced with the current effective CommandCode Max rates.
task cost = input tokens × CC input + cache-read × CC cache-read + cache-write × CC cache-write + output × CC output
If CommandCode exposes no separate cache rate, normal input price is used. Active CommandCode promotions are retained.
CAI*: observed when possible, estimated otherwise
For missing Coding Agent Index values, the estimate is deliberately simple:
CAI* = 50% Ridge prediction + 50% inverse-distance 5NN prediction
Ridge (λ=1) uses SciCode, GPQA, HLE, LCR, GDPval, Omniscience Accuracy, Omniscience Reliability and log(Output Tokens/Task). 5NN uses SciCode, GPQA, HLE and LCR, standardized across the observed CAI training set, with Euclidean distance and inverse-distance weighting over the five nearest models.
The estimator is trained only on unique model families with observed AA Coding Agent results. Commercial CommandCode variants never count as extra training examples.
Coverage and failure rules
All components of the Universal Role Score, CommandCode Cost per Task and final CAI* must cover 100% of the scored model-family universe. Stealth/unresolved rows remain visible as UNSCORED. The daily job also reruns reverse validation and fails closed if estimator/ranking guardrails are breached.
Raw model mappings, observed/estimated CAI*, Ridge/5NN components, role scores, task costs and final Ranking Values are all preserved in the public JSON snapshot.
Swap evaluation
Every data refresh evaluates possible single-seat replacements against the current lineup policy. Quality candidates must clear a minimum absolute-quality improvement; Budget candidates must clear a minimum cost-reduction threshold while preserving the 80% quality floor. An automatic-safe candidate must also preserve or increase the current distinct-family count. Security Reviewer is always review-only because it is a policy anchor. Balanced candidates are always review-only because a better quality-per-cost number can still reduce absolute quality or role fit.
External-evidence overrides and multi-seat rotations are also always review-only. The evaluator never changes the portfolio by itself: applyAutomatically remains false. New models must first receive an explicit family classification before they are eligible for automated swap evaluation.