Canary Promotion
Eval-gated model swap and rollback at the LLM gateway: change records, traffic split, and when canary becomes stable.
Eval-gated model swap and rollback at the LLM gateway: change records, traffic split, and when canary becomes stable.
Versioned registry of approved model endpoints per profile, task, data class, and region: the Plane ③ source of truth.
How to evaluate the Reasoning plane — faithfulness to context, conclusion quality, tool selection, and multi-step logic.
Curated third-party articles, guides, and tool docs on LLM and agent evaluation — mapped to the Eval Framework Blueprint series.
Per-call filter, score, and pick (or abstain) at the LLM gateway: task type, failover, cost caps, and residency.
Every plan or synthesize call goes through Plane ③, then validate proposals against the manifest and gate side effects at the PEP.
How to deploy LLM-as-judge for plane-aware evaluation — rubric design, judge selection, bias controls, and calibration against human ground truth.
Hub for Plane ③ model routing playbooks: capability matrix, gateway task routing, and eval-gated canary promotion at the LLM gateway.
Tool schemas only, proposal-not-permission, and keeping authority out of the model boundary.