Straits Industriesproduction
Search models, nodes, traces…⌘K
SI

Evaluations

2 completed runs

Quality evidence and learning (memo §7.6): golden datasets replayed through the real pipeline, versioned scorecards, and staged rollouts that gate every change to catalog quality scores.

Golden datasets
321 graded items
Scorecards
64 models covered
Mean score
0.96across all scorecards
Feedback
475% positive

Rollout pipeline

Router versions move shadow → canary → full. Quality scores are written to the model catalog only when a rollout reaches full — nothing else may mutate them.

router_1.0+eval.run_1588ba8odtproposalro_4ygphfac0l
shadow2d agocanary1d agofull

Canary workspaces only; scorecards observed against live traffic.

openai/gpt-5.2extraction0.94google/gemini-3-proextraction0.91anthropic/claude-sonnetextraction0.94

Evidence: run run_1588ba8odt over Extraction JSON set (18 graded items)

Scorecard matrix

Model × dataset mean score from the latest run of each pair. Bars encode the deterministic grade (JSON validity, keyword coverage, length sanity).

ModelCoding golden setcodingSummarization golden setsummarizationExtraction JSON setextractionMean
anthropic/claude-haiku0.95pass 100% · 8 items · 1.79 snot evaluatednot evaluated0.95
anthropic/claude-sonnet1.00pass 100% · 8 items · 2.24 snot evaluated0.98pass 100% · 6 items · 1.61 s0.99
google/gemini-3-pronot evaluatednot evaluated0.94pass 100% · 6 items · 2.24 s0.94
openai/gpt-5.20.92pass 100% · 8 items · 1.73 snot evaluated1.00pass 100% · 6 items · 1.96 s0.96

Golden datasets

Organization-layer evidence (memo §7.6): curated prompts with expected traits. Running one replays every item through the real gateway pipeline under @company/general-balanced.

Coding golden set
ds_coding_golden
coding

Golden coding tasks: implementation, refactoring, debugging and review prompts with expected topical coverage.

8 graded items
Summarization golden set
ds_summarization_golden
summarization

Golden summarization tasks: incident notes, meeting minutes and policy digests with concise-answer length windows.

7 graded items
Extraction JSON set
ds_extraction_json
extraction

Structured extraction tasks graded primarily on JSON validity plus field-name coverage.

6 graded items

Evaluation runs

Every run is versioned evidence: dataset, pinned models, graded items, settled cost and the rollout it proposed.

RunDatasetModelsStatusGradedMean scoreCostRolloutTriggered byCompleted
run_1588ba8odtExtraction JSON setopenai/gpt-5.2google/gemini-3-proanthropic/claude-sonnetcompleted180.97$0.06canaryseed2d ago
run_gnmqri1dvtCoding golden setanthropic/claude-sonnetopenai/gpt-5.2anthropic/claude-haikucompleted240.95$0.09seed9d ago

Recent feedback

Runtime evidence layer (memo §7.6): end-user ratings captured via POST /v1/feedback and joined to the generation ledger.

gen_bxa9cnizlePlatform Admin+1

Fast and grounded; keep routing these to this model.

2h ago
gen_u1l3vrtgkkOps Copilot Agent+1
1d ago
gen_xa3lbqedvmCustomer Portal API−1

Answer ignored the JSON schema we asked for.

2d ago
gen_9z984r0d2gPlatform Admin+1

Accurate and well structured — exactly what the runbook needed.

3d ago