Reproduced live · checked vs a shuffle-null

Validated benchmarks

Every number here is a live re-run against the current engine, checked against a label-shuffle null — not a leaderboard cherry-pick. Structure in, deterministic out.

TaskMetricValuen
Taste, 6-class (full ChemTastesDB) — vs a shuffled-label nullaccuracy / null70.7% / 33.6%3,517
E-tongue bitterness (RCFTR contracted correlation)Pearson r / R²0.98 / 0.9520
Umami peptides vs trained iUmami-SCM (0.61)MCC0.6888
Off-note flag (beany / bitter / metallic)balanced accuracy79%3,517
Aroma family, on brand-new compoundscorrect family72.5%40
Allergen family (structure → family)exact / wrong-family94% / 0%31
Protein melting temp (DSC Td), cross-familyMAE6.55°C22
Protein quality (DIAAS screen, composite isolate)MAE vs measured0.1510
CD secondary structure (β-sheet / helix)Spearman ρ0.94 / 0.716
Isoelectric point (pI)Spearman ρ / MAE0.80 / 1.044
Intact mass & ε280vs reference methodmatches exactly
How we test

Every accuracy is re-run against a null: we shuffle the labels 300–2,000× and recompute. If the real score doesn’t clearly beat the shuffled one, it doesn’t go here. That’s why the taste figure is 70.7% vs a 33.6% shuffled null (p=0.0033) on the full set — a claim a data scientist can check, not a headline.

Where it defers

A screen, not an assay. When a call is structurally ambiguous the platform says provisional — confirm by panel rather than guess; when a protein fold is outside its tested scope it declines and still returns the biophysical screen. Certified numbers (compliance DIAAS, allergen assays) stay measured in the lab.

The set keeps growing. Each month we add new blind studies straight from the published literature — run structure-only, then checked against their own shuffle-null before they join the set. More validated ground = more of your lab campaign we can triage from one sequence or structure.