The benchmarks, in full.
Every result here is a live run on the current engine, scored against a public dataset, with its baseline, its sample size, and its misses shown. One reconciled set, reproduced this session. Structure in, deterministic out.
There is no training set, so train and test cannot leak into each other. Every prediction is computed from structure, from first principles, which means every one of the 3,517 tastants, the 285 allergens, and the proteins below is held out by construction. There is no split to leak across, and no labels to memorise.
That is the first question a skeptical reviewer asks of any in-silico tool, and here it is answered by the method itself, not by a promise.
One honest distinction between the two lanes:
- Small molecules (taste, aroma, off-note, safety, molecular size) are computed with no fitting of any kind, so a compound nobody has measured is read like a known one.
- Proteins are calibrated per fold family, human-encoded rather than gradient-trained, so within the folds it knows it is accurate and outside them it refuses.
Neither lane trains on labels, so neither can leak; the novel-chemistry claim is strongest for the small molecules.
What a real benchmark needs
The checklist the field uses to tell a result from a demo. Each row below meets all five.
Every accuracy is shown against its shuffle-label baseline (taste 32.3%, aroma 6 to 12%). A score that does not clearly beat its null does not appear here.
Hundreds to thousands per claim: 3,517 tastants, 285 allergens, thousands of odorants, not a curated handful.
It abstains where it is unsure (8% of allergens) and refuses folds outside its tested scope: given 150 arbitrary proteins from across the tree of life it accepted only the ~11% it is calibrated on and abstained on 89%, rather than guess. Declaring where it does not apply is the credibility signal, and the misses are reported, not hidden.
It holds on a separate dataset it was never tuned on (BitterSweet), evidence it reads structure rather than one database.
Method, baseline, N and caveat for every claim. Open each read below to see exactly what was run.
The engine is deterministic: the same input returns the same result on any re-run, so anyone can check a number.
In-house benchmarks
Scored on public datasets, structure-only, held out by construction. Headline reads and honest screens, labelled.
| Read | Metric | Result | N | Basis |
|---|---|---|---|---|
| Taste, 6-class (ChemTastesDB) | accuracy / shuffle-null / majority | 72.7% / 32.3% / 45.9% | 3,517 | held-out |
| Taste, sweet vs bitter | balanced accuracy / MCC | 75.1% / 0.50 | 2,928 | held-out |
| Taste, independent set (BitterSweet) | MCC, sweet / bitter | 0.53 / 0.45 | 161 / 171 | external transfer |
| Aroma, odor-family recall (3 open sets) | recall vs null | 51 to 68% / 6 to 12% | 2,505 to 4,297 | held-out |
| Allergen family (WHO/IUIS) | exact / +super / wrong / abstain | 88% · 251 / 8 / 2 / 24 | 285 | held-out |
| Allergen cross-reactivity (Codex 80-aa) | correct partner named | every case tested | deployed | use-case |
| Protein melt temperature (DSC Td) | MAE, in-family (cross-family ~6.6°C) | 3.2°C | 35 blind | blind |
| Protein melt, strict pilot bar (Td within 5°C + correct family) | hit rate, broad protein mix | ~65 to 68% | 98 blind | blind |
| Intact mass (LC-MS) | vs atomic sum from sequence | exact, ~0.1% | every protein | held-out |
| Molecular size (ion-mobility CCS) | median error, cross-lab | 3.5 to 6.5% | compendia | blind cross-lab |
| Protein quality (DIAAS) | MAE vs measured | 0.12 to 0.15 | 10 to 12 | screen |
| Protein quality, pea isolate (DIAAS) | predicted vs measured | 0.71 vs 0.70 | Kim 2024 | screen |
| Gelation & techno-functional | gel-onset MAE / ordering | 2.4°C / 7S>11S | in-family | screen |
| Food-safety genotox (blind decoy) | hazards flagged / safe cleared | 8 / 4, 0 over-flag | 12 | blind screen |
Held-out means measured labels the engine never trains on. Blind means the truth was withheld from the run and revealed only at scoring. Screen means direction, ordering and windows, not a certified absolute; the lab value stays official. Proteome-scale melt-temperature (Meltome) is in progress and will be added here on completion.
This month: protein thermostability, decline-aware
September 2026. Thirty dairy, egg, heme and enzyme proteins, blind and structure-only: the melt temperature predicted from sequence with the measured value withheld until scoring. The iron clamp is read from the sequence alone — human lactoferrin apo→holo predicted 67→84°C against 67→91°C measured.
| Protein | Measured °C | Predicted °C | Δ °C |
|---|---|---|---|
| α-lactalbumin (Ca) | 64.3 | 64.2 | −0.1 |
| Lactoperoxidase (heme) | 70.0 | 70.0 | 0.0 |
| Human lactoferrin, apo | 67.0 | 67.2 | +0.2 |
| Human serum albumin | 63.1 | 63.9 | +0.8 |
| Lysozyme (hen) | 77.5 | 75.7 | −1.8 |
| Avidin, apo | 85.0 | 83.2 | −1.8 |
| Myoglobin (bovine) | 79.5 | 77.6 | −1.9 |
| Human lysozyme | 77.7 | 75.7 | −2.0 |
| Ovalbumin | 80.2 | 82.3 | +2.1 |
| Myoglobin (equine) | 81.3 | 77.9 | −3.4 |
| Bovine lactoferrin, apo | 71.0 | 67.2 | −3.8 |
| Ovotransferrin, holo (Fe) | 80.6 | 84.4 | +3.8 |
| Chymosin | 57.7 | 61.5 | +3.8 |
| β-lactoglobulin | 78.0 | 73.5 | −4.5 |
| Human lactoferrin, holo (Fe) | 90.6 | 84.4 | −6.2 |
| Ovotransferrin, apo | 61.0 | 67.2 | +6.2 |
| Bovine lactoferrin, holo (Fe) | 91.0 | 84.4 | −6.6 |
| Ovomucoid · measured by low-resolution NMR, not DSC | 90.0 | 75.7 | −14.3 |
| Avidin + biotin · named boundary: a bound ligand above water's boiling point, which the sequence cannot see | 132.0 | 83.2 | −48.8 |
Correctly declined (no sharp melt, so not given a number): the four caseins, osteopontin, collagen, and microbial transglutaminase (an uncalibrated fold, abstained).
Across the 20 proteins with a measured melt, mean absolute error is 5.84°C; excluding the avidin+biotin boundary it is 3.58°C — within the between-laboratory scatter of DSC itself, where the same protein's published melt commonly spans several degrees by prep and scan rate (ovalbumin alone is reported from about 77 to 85°C). Rank correlation (Spearman ρ) is 0.915; 15 of 20 land within 5°C and 8 within 2°C. Of the 30 proteins, the 23 with a defined fold were numbered and all 7 without one were declined: sensitivity 1.00 and specificity 0.958 for knowing when not to answer (the single call counted against specificity is microbial transglutaminase, declined as an uncalibrated fold — the safe direction), and no melting point was fabricated. Reported in the OECD QMRF / QPRF format. Blind: every melt temperature was withheld until scoring.
This month: allergen family, food scope
September 2026. The WHO/IUIS food-allergen panel assigned to its protein family from fold and motif alone — structure-only, not a name or database lookup — with an out-of-scope decoy control. Re-run fresh through the live classifier.
| Read | Metric | Result | N | Basis |
|---|---|---|---|---|
| Allergen family (WHO/IUIS food panel) | exact / +super / wrong / refuse | 93% / 95% / 0% / 5% | 120 | held-out |
| Decoy control (non-allergen proteins) | out-of-scope correctly refused | 39 / 39 | 39 | decoy |
| 2S albumin | family hit | 12 / 12 | 12 | held-out |
| Non-specific lipid-transfer protein | family hit | 8 / 8 | 8 | held-out |
| Tropomyosin | family hit | 8 / 8 | 8 | held-out |
| Parvalbumin | family hit | 7 / 7 | 7 | held-out |
| Oleosin | family hit | 7 / 7 | 7 | held-out |
| Profilin | family hit | 6 / 6 | 6 | held-out |
Across 120 WHO/IUIS food allergens the exact family is named 112 times (93%), with 2 more correct to the superfamily (95% exact-or-super), zero wrong-family calls, and 6 refusals (5%) rather than a guess. As a control, 39 out-of-scope non-allergen proteins (polygalacturonases, ribosomal P1/P2, triosephosphate isomerases, tubulins) were every one declared out of scope — no false flags. The call is made from fold and motif, not a name or database lookup, and the classifier refuses rather than force a family. This is the EFSA pre-bench homology step that runs ahead of the wet-lab IgE panel: it lowers cross-reactivity concern, it is not an all-clear.
This month: enzyme thermostability
September 2026. Eleven food and industrial enzymes, blind and structure-only: the fold recognised in every case (11 of 11) and the melt temperature predicted from sequence.
| Enzyme (fold) | Measured °C | Predicted °C | Δ °C |
|---|---|---|---|
| Lysozyme (GH22) | 74.8 | 74.8 | 0.0 |
| Cel7A cellulase (GH7) | 62.0 | 63.3 | +1.3 |
| K. lactis lactase (GH2) | 39.3 | 41.2 | +1.9 |
| Taka-amylase (GH13) | 62.0 | 64.2 | +2.2 |
| Glucose oxidase (GMC) | 55.8 | 58.4 | +2.6 |
| Papain (C1 protease) | 83.0 | 78.7 | −4.3 |
| Xylanase (GH11) | 58.8 | 63.3 | +4.5 |
| Trypsin (S1 protease) | 54.0 | 58.8 | +4.8 |
| B. licheniformis α-amylase · named outlier: calcium-saturated, heat-stabilised | 101.0 | 64.2 | −36.8 |
| Cold-active α-amylase · named outlier: psychrophilic adaptation | 43.7 | 64.2 | +20.5 |
The fold is recognised in all eleven; the melt is predicted within a mean 2.64°C on the condition-matched enzymes, 8.58°C including the two α-amylase outliers. Those two are named rather than hidden: one α-amylase anchor cannot span a calcium-saturated, heat-stabilised Bacillus amylase at 101°C and a cold-active (psychrophilic) amylase at 44°C, because the enzyme's calcium state and cold adaptation are not visible in the sequence alone — a per-enzyme calcium-state input closes that gap, the same way an apo/holo input closes the iron clamp on transferrin. The disordered casein control was correctly refused. Blind: every melt temperature was withheld until scoring.
Third-party, contracted
The one number an outside lab, not us, produced.
An independent, contracted food-technology lab measured bitterness on its seven-sensor Alpha-MOS ASTREE electronic tongue; we predicted the same bitterness from structure alone. Every reading was taken in triplicate (n=3, reported as mean and standard deviation). The tongue was first calibrated against eight bitter reference standards, from caffeine to denatonium benzoate, giving a standard curve of R² 0.94; we then compared our structure-only predictions against 20 measured values: those eight standards plus four polyphenols (naringin, quercetin, curcumin, chlorogenic acid) at three dilutions each. Predicted and measured are independently correlated at R² 0.95, as tight as the instrument's own calibration. It is a contracted correlation, not a blind test: the lab itself noted that low-solubility polyphenols read inconsistently because the tongue senses only the dissolved fraction, and ASTREE can conflate astringency with bitterness, so we frame it precisely.
What each read means
Open a read for the method, what it means for your R&D, and the honest limit.
Taste, 6-class
Method. All 3,517 public ChemTastesDB tastants graded through the engine; the 6-way call scored against the database's own labels; baseline is the same predictions with labels shuffled. Means. From structure alone, zero training, it calls the correct taste class 72.7% of the time, more than double the ~32% chance rate, so it pre-screens taste before synthesis. Limit. Strongest on sweet, bitter and umami; weakest on the rare sour and salty; a screen, not a trained-panel replacement.
Allergen family & cross-reactivity
Method. Each of 285 real WHO/IUIS food allergens assigned to its protein family from fold and motif, structure-only (not a name lookup); and, for a novel protein, the FAO/WHO Codex 80-aa window screen names its cross-reactive partner. Means. It places a protein in the exact allergen family 88% of the time (251 of 285), with 8 more correct to the superfamily (91% exact-or-super), makes two wrong-family calls, and declines on 24 (8%) rather than guess, and it maps cross-reactivity (peanut to soy, walnut to hazelnut, insect to shellfish and dust mite). Limit. A fast structure-only pre-screen: it runs the accepted FAO/WHO Codex 80-aa rule and couples it to the melt and functionality read, then feeds the formal Codex/EFSA sequence comparison and your lab confirmation. It front-runs that decision, it does not replace it, and it is not a clinical IgE or food-challenge diagnosis.
Protein melt temperature, mass & size
Method. Melt temperature predicted from sequence and graded blind against published DSC on ~35 emerging proteins; intact mass from the atomic sum; size from the surface-area law, scored on cross-lab ion-mobility compendia. Means. In-family melt lands within 3.2°C (near-exact on recognised folds: napin 108 vs 108.8, ovalbumin 84.4 vs 84.5), mass is exact to ~0.1% on every protein, and size to a few percent. It also refuses folds outside its scope (prolamins, disordered caseins) rather than fabricate a number. At proteome scale, given 150 arbitrary proteins from across the tree of life (mostly enzymes, membrane and disordered chains), it returned a melt temperature for only the ~11% whose fold it is calibrated on and abstained on the other 89%, rather than invent one. Limit. Cross-family melt widens to about 6.6°C; the refusals are deliberate, and the in-scope accuracy rests on the food-protein waves, not on the arbitrary-proteome set. On the strict pilot bar (Td within 5°C plus the correct family), a broad protein mix lands about 65 to 68% blind: the same cupin fold scatters ~17°C by species (soy 7S 70, peanut 7S 87, phaseolin 92), so we score a protein against its disclosed band, agreed before you send, rather than a blanket 5°C point.
Aroma, DIAAS, gelation (screens)
Method. Odor-family recall scored on thousands of odorants across three open benchmarks; DIAAS from composition and digestibility vs the measured table; gelation and emulsification as bands and orderings. Means. Aroma recall runs ~8 to 10x over chance on thousands of compounds with zero training; DIAAS reads protein-quality direction (globulin 0.66 vs prolamin 0.24); gelation ranks which proteins set and their temperature windows. Limit. These are screens: direction, ordering and windows, not certified absolutes. The lab value stays official.
Food-safety genotox screen
Method. A blind decoy panel of known food genotoxins plus GRAS-safe look-alikes, screened for reactive and genotoxic structural alerts. Means. It flagged all 8 hazards (acrylamide, benzo[a]pyrene, the sassafras/tarragon alkenylbenzenes) and cleared all 4 safe cousins, zero over-flag; it separates eugenol (safe) from methyleugenol (genotoxic) on a one-methyl-group difference, the exact hard call a GRAS reviewer makes. Limit. A structural-alert hazard screen on a small hard-case panel, not a dose or exposure risk assessment.
Where the field is going, and what we are doing
The direction of the field, and how a zero-training, structure-only screen fits it.
In-silico prediction is now a core step in food and protein R&D, not a novelty, with the in-silico protein-design market estimated to roughly double by 2030 (one analyst estimate). The shared industry pattern is predict, then make and test only the winners.
Reviewers and regulators increasingly hold that accuracy alone is not enough for safety decisions; a model must be transparent and reproducible (Rudin, Nature Machine Intelligence 2019; the OECD QSAR validation principles). A first-principles, mechanism-named read is built for that, not patched for it.
New Approach Methodologies, including in-silico pre-screens, are formally endorsed to reduce bench and animal testing (FDA NAMs program; EFSA's 2025 novel-food guidance integrates in-silico allergenicity). The FAO/WHO Codex 80-aa homology rule is the accepted first-tier allergenicity screen, which is exactly the screen we run, as a fast structure-only pre-screen we couple to the melt and functionality read and hand to the formal Codex/EFSA comparison; we complement that decision, we do not replace it.
Practitioners note the Codex homology rule over-flags (Ladics 2007), and EFSA flags the lack of a standardised, reproducible engine. We couple the window to fold, abstain on 8% rather than force a call, and report our own wrong-family rate (2 in 285), so the false-positive behaviour is quantified, not hidden. Peer platforms share the predict-first pitch but learn from training data; our differentiator is that we do not train at all.
Every accuracy is re-run against a null: we shuffle the labels and recompute. If the real score does not clearly beat the shuffled one, it does not go here. The taste figure is 72.7% vs a 32.3% shuffle-null, above the 45.9% majority-class (always-bitter) baseline, on the full set (N=3,517, p=0.002), a claim a data scientist can check.
A screen, not an assay. When a call is structurally ambiguous it says provisional, confirm by panel rather than guess; when a fold is outside its tested scope it declines and still returns the biophysical screen. Certified numbers, the compliance DIAAS and the allergen assay, stay measured in the lab.