Evidence Readiness Dashboard
A snapshot-based evidence surface that shows readiness, missingness, blocked decisions, and the next unblock action without implying benchmark results or live monitoring.
Evidence status
Evidence Readiness layer operational / comparison blocked
Source of truth: GPT KB + Git. Refresh cadence: manual snapshot only.
Cohort boundary
Narrow new-record cohort
Legacy rows and non-joinable historical exceptions are excluded by construction.
- The page shows what exists, what is missing, and what decision is blocked.
- The page is static and public-safe; it does not fetch live data.
- The page uses a narrow new-record cohort and excludes legacy / non-joinable rows by construction.
- The page keeps comparison readiness separate from publish readiness.
- No benchmark proof.
- No provider superiority claim.
- No cost-saving claim.
- No routing correctness claim.
- No decision-quality claim.
- No live observability claim.
As-of timestamp: 2026-06-03
Refresh cadence: manual release snapshot only
Exclusions: legacy rows, non-joinable records, raw receipts, and any data outside the narrow evidence cohort.
Future comparison candidate: OpenRouter profile vs later Codex receipts, but only after the evidence cohort becomes joinable.
The readout now separates historical telemetry from comparison-ready evidence.
Completed:
- Schema v0.7 implemented and frozen
- JSONL readout tool created and working
- Historical Codex receipts reclassified as semi_structured backfill records
- 3 historical receipts excluded from comparison by design
- No usage or cost data was invented
- Comparison remains blocked until future records meet evidence contract
Current value:
- Telemetry is no longer treated as automatically usable
- Historical records clearly separated from future evidence cohorts
- Missingness is visible instead of hidden
- Fake completeness is prevented
Total receipts
3
Dashboard-grade
0
Blocked state
0
Excluded
3
Comparison-ready
0
Top blocker
Historical/semi_structured backfill records are not comparison candidates
| Metric | Numerator | Denominator | % | Source | Confidence | Decision value |
|---|---|---|---|---|---|---|
| Join completeness | 0 new records with required join keys | 3 new records in the P1 scope | 0% | P1.1 metrics + blocker matrix | High | Comparison stays blocked until joinability exists. |
| Routing metadata coverage | 0 new records with routing_decision + routing_reason | 3 new records in the P1 scope | 0% | P1.1 blocker refresh | High | Routing evidence remains incomplete and comparison stays blocked. |
| Usage / cost provenance coverage | 0 records with explicit usage and cost provenance | 3 records where model usage is applicable | 0% | Receipt schema + runner telemetry | High | Cost/value comparison stays blocked until provenance is explicit. |
| Validation artifact coverage | 1 clearly applicable record with validation artifact ref | 2 records requiring validation | 50% | Validation summaries + receipt refs | Medium | Blocked states can be shown, but comparison is not ready. |
| Decision metadata coverage | 0 new records with business_question + owner_role + exactly one decision field | 3 new records in the P1 scope | 0% | Task spec + receipt context | High | Decision mapping remains blocked until each record carries one clear outcome. |
- Join keys are still missing on the new-record cohort.
- Routing metadata is still incomplete on the new-record cohort.
- Usage / cost provenance is still incomplete on applicable records.
- Decision metadata still needs one clear outcome per record.
- Blocked states are visible instead of being hidden.
- Validation artifacts are explicitly referenced where applicable.
- Future records can be classified as dashboard-grade, blocked, or excluded.
- The page keeps publish readiness separate from comparison readiness.
Operational comparison
Blocked. The new-record cohort is still missing join keys, routing metadata, and decision metadata.
Next unblock action
Keep future records on the minimal task spec and enforce joinable receipts.
Benchmark claim
Blocked. This surface is evidence readiness only and does not prove model superiority or ROI.
Next unblock action
Keep the claim boundary fixed to readiness and missingness.
Provider comparison
Blocked. OpenRouter profile vs later Codex receipts remains a future candidate only.
Next unblock action
Collect future records with explicit routing and provenance fields first.
Live observability
Blocked. This page is a static snapshot, not a runtime telemetry surface.
Next unblock action
Keep the refresh cadence manual and snapshot-based.
Owner: Data Evidence / Routing. The next unblock is to keep future tasks on the minimal task spec and carry the join keys, routing metadata, provenance, validation, and decision metadata through to the receipt.
Review gate: GPT / Opus boundary review for any public claim, then Sararin for final strategic approval.
This surface is intentionally narrow. It is a public-safe evidence discipline snapshot, not a live monitoring system.
Keep the dashboard static, keep the cohort narrow, and keep the comparison path blocked until evidence quality is strong enough to justify it.