Methodology

How role quality and benchmark-vs-cost Ranking Value are produced.

Loading methodology…

Two views of the same role score

Quality ignores price and ranks only the role-quality score. Balanced divides that same quality by CommandCode Cost per Task. Higher is better in both modes.

Non-coding role qualityRole Quality = Universal Role Score
Coding role qualityRole Quality = ⅔ Universal Role Score + ⅓ CAI*
Balanced modeBalanced Score = Role Quality / CommandCode Cost per Task
Budget lineupMinimize CommandCode Cost per Task subject to Role Quality ≥ 80% of the best available role quality

The Budget lineup is only a recommended shortlist, not a third full-table ranking. It also applies the current-generation and model-family-diversity constraints used by the recommended lineups.

Free-model external-evidence overrides

The recommended lineups may include a CommandCode model that is still UNSCORED when it is explicitly free and there is enough external evidence to justify a narrow role recommendation. This is an editorial override only: unresolved models remain excluded from the full Quality and Balanced rankings until a verified Artificial Analysis family mapping exists.

Laguna S 2.1 is used as a free coding override. CommandCode exposes it as free with a 256K served context; Poolside documents a native 1M context and reports 70.2% Terminal-Bench 2.1, 78.5% SWE-bench Multilingual, 59.4% SWE-bench Pro and 40.4% DeepSWE. Poolside model page · model card and benchmark table.

Ox Alpha is used only in the Budget shortlist and marked EXPERIMENTAL. The CommandCode row is free with a 1M context window, but the model remains stealth and unresolved, so no synthetic benchmark score is assigned to it and it is not used in Balanced.

Portfolio assignment

The eight recommendation cards are selected as a portfolio rather than as eight independent winners. The same exact model may occupy two seats, and commercial rows sharing the same Artificial Analysis model identity may coexist. Diversity is enforced at the family level: each lineup must span 5–8 distinct model families, and no family may occupy more than two seats. Six, seven or eight families are all acceptable; five remains valid when deliberately justified. No manual lineup is forced toward a particular target.

Balanced currently requires Claude Opus 5, GPT-5.6 Luna and GPT-5.6 Sol somewhere in the portfolio. Their roles are not permanently fixed. The current assignment uses Muse Spark 1.2 Contributor for normal Implementer, Luna for Code Reviewer and Sol for Implementer Heavy. Sol is explicitly forbidden for normal Implementer because routine implementation is the highest-volume seat and cheaper strong models provide better economics. Claude Sonnet 5 remains eligible but is not mandatory.

Implementation routing should be interpreted as a cost-per-completed-task policy: use Muse for routine high-volume implementation, Luna for review, correct cheaply on the normal tier, and escalate genuinely hard or repeatedly failing work to Sol. The public benchmark does not yet measure retry probability directly, so this escalation rule is a documented operating policy rather than a measured benchmark component.

The coding adjustment applies only to Implementer, Implementer Heavy and Code Reviewer.

Security Reviewer tier policy

The numerical Security Reviewer score remains a universal proxy based on SciCode, GPQA, HLE, Non-Hallucination and LCR; it is not a dedicated cybersecurity benchmark. Public cyber benchmarks reviewed so far do not provide sufficiently broad, comparable model coverage to justify a universal Security* adjustment.

The recommendation cards therefore use explicit security tiers: Quality → Claude Fable 5, Balanced → Claude Opus 5, Budget → Muse Spark 1.2 Contributor. These are policy anchors and Security Reviewer changes always require human review. Octopus has no separate security-heavy role; implementer-heavy is a coding/implementation role.

CommandCode Cost per Task

Price per million tokens is not used as the primary denominator. Artificial Analysis measures the token mix consumed per Intelligence Index task. Those measured non-cached input, cache-read, cache-write and output tokens are repriced with the current effective CommandCode Max rates.

task cost = input tokens × CC input + cache-read × CC cache-read + cache-write × CC cache-write + output × CC output

If CommandCode exposes no separate cache rate, normal input price is used. Active CommandCode promotions are retained.

CAI*: observed when possible, estimated otherwise

For missing Coding Agent Index values, the estimate is deliberately simple:

CAI* = 50% Ridge prediction + 50% inverse-distance 5NN prediction

Ridge (λ=1) uses SciCode, GPQA, HLE, LCR, GDPval, Omniscience Accuracy, Omniscience Reliability and log(Output Tokens/Task). 5NN uses SciCode, GPQA, HLE and LCR, standardized across the observed CAI training set, with Euclidean distance and inverse-distance weighting over the five nearest models.

The estimator is trained only on unique model families with observed AA Coding Agent results. Commercial CommandCode variants never count as extra training examples.

Coverage and failure rules

All components of the Universal Role Score, CommandCode Cost per Task and final CAI* must cover 100% of the scored model-family universe. Stealth/unresolved rows remain visible as UNSCORED. The daily job also reruns reverse validation and fails closed if estimator/ranking guardrails are breached.

Raw model mappings, observed/estimated CAI*, Ridge/5NN components, role scores, task costs and final Ranking Values are all preserved in the public JSON snapshot.

Swap evaluation

Every data refresh evaluates possible single-seat replacements against the current lineup policy. Quality candidates must clear a minimum absolute-quality improvement; Budget candidates must clear a minimum cost-reduction threshold while preserving the 80% quality floor. An automatic-safe candidate must also preserve or increase the current distinct-family count. Security Reviewer is always review-only because it is a policy anchor. Balanced candidates are always review-only because a better quality-per-cost number can still reduce absolute quality or role fit.

External-evidence overrides and multi-seat rotations are also always review-only. The evaluator never changes the portfolio by itself: applyAutomatically remains false. New models must first receive an explicit family classification before they are eligible for automated swap evaluation.