Validated benchmarks
Every number here is a live re-run against the current engine, checked against a label-shuffle null — not a leaderboard cherry-pick. Structure in, deterministic out.
| Task | Metric | Value | n |
|---|---|---|---|
| Taste, 6-class (full ChemTastesDB) — vs a shuffled-label null | accuracy / null | 70.7% / 33.6% | 3,517 |
| E-tongue bitterness (RCFTR contracted correlation) | Pearson r / R² | 0.98 / 0.95 | 20 |
| Umami peptides vs trained iUmami-SCM (0.61) | MCC | 0.68 | 88 |
| Off-note flag (beany / bitter / metallic) | balanced accuracy | 79% | 3,517 |
| Aroma family, on brand-new compounds | correct family | 72.5% | 40 |
| Allergen family (structure → family) | exact / wrong-family | 94% / 0% | 31 |
| Protein melting temp (DSC Td), cross-family | MAE | 6.55°C | 22 |
| Protein quality (DIAAS screen, composite isolate) | MAE vs measured | 0.15 | 10 |
| CD secondary structure (β-sheet / helix) | Spearman ρ | 0.94 / 0.71 | 6 |
| Isoelectric point (pI) | Spearman ρ / MAE | 0.80 / 1.04 | 4 |
| Intact mass & ε280 | vs reference method | matches exactly | — |
Every accuracy is re-run against a null: we shuffle the labels 300–2,000× and recompute. If the real score doesn’t clearly beat the shuffled one, it doesn’t go here. That’s why the taste figure is 70.7% vs a 33.6% shuffled null (p=0.0033) on the full set — a claim a data scientist can check, not a headline.
A screen, not an assay. When a call is structurally ambiguous the platform says provisional — confirm by panel rather than guess; when a protein fold is outside its tested scope it declines and still returns the biophysical screen. Certified numbers (compliance DIAAS, allergen assays) stay measured in the lab.