Evaluations
2 completed runsQuality evidence and learning (memo §7.6): golden datasets replayed through the real pipeline, versioned scorecards, and staged rollouts that gate every change to catalog quality scores.
Rollout pipeline
Router versions move shadow → canary → full. Quality scores are written to the model catalog only when a rollout reaches full — nothing else may mutate them.
Canary workspaces only; scorecards observed against live traffic.
Evidence: run run_1588ba8odt over Extraction JSON set (18 graded items)
Scorecard matrix
Model × dataset mean score from the latest run of each pair. Bars encode the deterministic grade (JSON validity, keyword coverage, length sanity).
| Model | Coding golden setcoding | Summarization golden setsummarization | Extraction JSON setextraction | Mean |
|---|---|---|---|---|
| anthropic/claude-haiku | 0.95pass 100% · 8 items · 1.79 s | not evaluated | not evaluated | 0.95 |
| anthropic/claude-sonnet | 1.00pass 100% · 8 items · 2.24 s | not evaluated | 0.98pass 100% · 6 items · 1.61 s | 0.99 |
| google/gemini-3-pro | not evaluated | not evaluated | 0.94pass 100% · 6 items · 2.24 s | 0.94 |
| openai/gpt-5.2 | 0.92pass 100% · 8 items · 1.73 s | not evaluated | 1.00pass 100% · 6 items · 1.96 s | 0.96 |
Golden datasets
Organization-layer evidence (memo §7.6): curated prompts with expected traits. Running one replays every item through the real gateway pipeline under @company/general-balanced.
Golden coding tasks: implementation, refactoring, debugging and review prompts with expected topical coverage.
Golden summarization tasks: incident notes, meeting minutes and policy digests with concise-answer length windows.
Structured extraction tasks graded primarily on JSON validity plus field-name coverage.
Evaluation runs
Every run is versioned evidence: dataset, pinned models, graded items, settled cost and the rollout it proposed.
| Run | Dataset | Models | Status | Graded | Mean score | Cost | Rollout | Triggered by | Completed |
|---|---|---|---|---|---|---|---|---|---|
| run_1588ba8odt | Extraction JSON set | openai/gpt-5.2google/gemini-3-proanthropic/claude-sonnet | completed | 18 | 0.97 | $0.06 | canary | seed | 2d ago |
| run_gnmqri1dvt | Coding golden set | anthropic/claude-sonnetopenai/gpt-5.2anthropic/claude-haiku | completed | 24 | 0.95 | $0.09 | — | seed | 9d ago |
Recent feedback
Runtime evidence layer (memo §7.6): end-user ratings captured via POST /v1/feedback and joined to the generation ledger.
Fast and grounded; keep routing these to this model.
Answer ignored the JSON schema we asked for.
Accurate and well structured — exactly what the runbook needed.