Matched labels
Both G10 finalists reached 99.8% agreement with the experimental teacher-silver reference and 100% prediction coverage on the shared 500-patient panel.
Project overview
The project tests how rules, machine learning, routing, and language-model components can be compared without losing coverage, auditability, cost records, or the limits of the evidence.
How to use this site
The catalog defines the available families. The graph explorer shows how one declared configuration is assembled. The tournament and figure pages then show the recorded aggregate results. The workflow builder is for design proposals; it does not invent performance estimates for combinations that have not been run.
Framework
Translate each source into a common evidence contract.
Use transparent rules and local models for decisive cases.
Escalate records with structural uncertainty.
Apply the selected agent or bounded program.
Record coverage, cost, probability quality, and failures.
Figure 1. A declared configuration selects exact implementations for each stage and records their dependencies.
Current finding
All results here are comparisons with an experimental teacher-silver reference on synthetic records. They are not estimates of clinical diagnostic accuracy.
Both G10 finalists reached 99.8% agreement with the experimental teacher-silver reference and 100% prediction coverage on the shared 500-patient panel.
The mixed-model G10 cascade had the lower Brier score: 0.0025 compared with 0.0103 for the all-Haiku configuration.
The all-Haiku G10 cascade was less expensive under the reconstructed cold-equivalent cost. The practical result is a frontier, not one universal winner.