Project overview

A modular, cost-aware framework for agentic computable phenotyping

The project tests how rules, machine learning, routing, and language-model components can be compared without losing coverage, auditability, cost records, or the limits of the evidence.

How to use this site

Start with the question, then inspect the exact configuration

The catalog defines the available families. The graph explorer shows how one declared configuration is assembled. The tournament and figure pages then show the recorded aggregate results. The workflow builder is for design proposals; it does not invent performance estimates for combinations that have not been run.

Framework

One common spine; replaceable, versioned components

1

Harmonize

Translate each source into a common evidence contract.

2

Screen

Use transparent rules and local models for decisive cases.

3

Route

Escalate records with structural uncertainty.

4

Adjudicate

Apply the selected agent or bounded program.

5

Audit

Record coverage, cost, probability quality, and failures.

Figure 1. A declared configuration selects exact implementations for each stage and records their dependencies.

Current finding

Similar labels did not imply an identical system

All results here are comparisons with an experimental teacher-silver reference on synthetic records. They are not estimates of clinical diagnostic accuracy.

1

Matched labels

Both G10 finalists reached 99.8% agreement with the experimental teacher-silver reference and 100% prediction coverage on the shared 500-patient panel.

2

Probability quality

The mixed-model G10 cascade had the lower Brier score: 0.0025 compared with 0.0103 for the all-Haiku configuration.

3

Operating burden

The all-Haiku G10 cascade was less expensive under the reconstructed cold-equivalent cost. The practical result is a frontier, not one universal winner.