Held out by construction · every number carries its baseline and its N

The benchmarks, in full.

Every result here is a live run on the current engine, scored against a public dataset, with its baseline, its sample size, and its misses shown. One reconciled set, reproduced this session. Structure in, deterministic out.

Analytical laboratory bench
The analytical bench a from-structure screen front-runs.
Why these numbers hold

There is no training set, so train and test cannot leak into each other. Every prediction is computed from structure, from first principles, which means every one of the 3,517 tastants, the 285 allergens, and the proteins below is held out by construction. There is no split to leak across, and no labels to memorise.

That is the first question a skeptical reviewer asks of any in-silico tool, and here it is answered by the method itself, not by a promise.

One honest distinction between the two lanes:

  • Small molecules (taste, aroma, off-note, safety, molecular size) are computed with no fitting of any kind, so a compound nobody has measured is read like a known one.
  • Proteins are calibrated per fold family, human-encoded rather than gradient-trained, so within the folds it knows it is accurate and outside them it refuses.

Neither lane trains on labels, so neither can leak; the novel-chemistry claim is strongest for the small molecules.

What a real benchmark needs

The checklist the field uses to tell a result from a demo. Each row below meets all five.

A stated null

Every accuracy is shown against its shuffle-label baseline (taste 32.3%, aroma 6 to 12%). A score that does not clearly beat its null does not appear here.

Real sample sizes

Hundreds to thousands per claim: 3,517 tastants, 285 allergens, thousands of odorants, not a curated handful.

A named boundary

It abstains where it is unsure (8% of allergens) and refuses folds outside its tested scope: given 150 arbitrary proteins from across the tree of life it accepted only the ~11% it is calibrated on and abstained on 89%, rather than guess. Declaring where it does not apply is the credibility signal, and the misses are reported, not hidden.

Transfer to another set

It holds on a separate dataset it was never tuned on (BitterSweet), evidence it reads structure rather than one database.

More than one number

Method, baseline, N and caveat for every claim. Open each read below to see exactly what was run.

Reproducible

The engine is deterministic: the same input returns the same result on any re-run, so anyone can check a number.

In-house benchmarks

Scored on public datasets, structure-only, held out by construction. Headline reads and honest screens, labelled.

ReadMetricResultNBasis
Taste, 6-class (ChemTastesDB)accuracy / shuffle-null / majority72.7% / 32.3% / 45.9%3,517held-out
Taste, sweet vs bitterbalanced accuracy / MCC75.1% / 0.502,928held-out
Taste, independent set (BitterSweet)MCC, sweet / bitter0.53 / 0.45161 / 171external transfer
Aroma, odor-family recall (3 open sets)recall vs null51 to 68% / 6 to 12%2,505 to 4,297held-out
Allergen family (WHO/IUIS)exact / +super / wrong / abstain88% · 251 / 8 / 2 / 24285held-out
Allergen cross-reactivity (Codex 80-aa)correct partner namedevery case testeddeployeduse-case
Protein melt temperature (DSC Td)MAE, in-family (cross-family ~6.6°C)3.2°C35 blindblind
Protein melt, strict pilot bar (Td within 5°C + correct family)hit rate, broad protein mix~65 to 68%98 blindblind
Intact mass (LC-MS)vs atomic sum from sequenceexact, ~0.1%every proteinheld-out
Molecular size (ion-mobility CCS)median error, cross-lab3.5 to 6.5%compendiablind cross-lab
Protein quality (DIAAS)MAE vs measured0.12 to 0.1510 to 12screen
Protein quality, pea isolate (DIAAS)predicted vs measured0.71 vs 0.70Kim 2024screen
Gelation & techno-functionalgel-onset MAE / ordering2.4°C / 7S>11Sin-familyscreen
Food-safety genotox (blind decoy)hazards flagged / safe cleared8 / 4, 0 over-flag12blind screen

Held-out means measured labels the engine never trains on. Blind means the truth was withheld from the run and revealed only at scoring. Screen means direction, ordering and windows, not a certified absolute; the lab value stays official. Proteome-scale melt-temperature (Meltome) is in progress and will be added here on completion.

This month: protein thermostability, decline-aware

September 2026. Thirty dairy, egg, heme and enzyme proteins, blind and structure-only: the melt temperature predicted from sequence with the measured value withheld until scoring. The iron clamp is read from the sequence alone — human lactoferrin apo→holo predicted 67→84°C against 67→91°C measured.

ProteinMeasured °CPredicted °CΔ °C
α-lactalbumin (Ca)64.364.2−0.1
Lactoperoxidase (heme)70.070.00.0
Human lactoferrin, apo67.067.2+0.2
Human serum albumin63.163.9+0.8
Lysozyme (hen)77.575.7−1.8
Avidin, apo85.083.2−1.8
Myoglobin (bovine)79.577.6−1.9
Human lysozyme77.775.7−2.0
Ovalbumin80.282.3+2.1
Myoglobin (equine)81.377.9−3.4
Bovine lactoferrin, apo71.067.2−3.8
Ovotransferrin, holo (Fe)80.684.4+3.8
Chymosin57.761.5+3.8
β-lactoglobulin78.073.5−4.5
Human lactoferrin, holo (Fe)90.684.4−6.2
Ovotransferrin, apo61.067.2+6.2
Bovine lactoferrin, holo (Fe)91.084.4−6.6
Ovomucoid · measured by low-resolution NMR, not DSC90.075.7−14.3
Avidin + biotin · named boundary: a bound ligand above water's boiling point, which the sequence cannot see132.083.2−48.8

Correctly declined (no sharp melt, so not given a number): the four caseins, osteopontin, collagen, and microbial transglutaminase (an uncalibrated fold, abstained).

Across the 20 proteins with a measured melt, mean absolute error is 5.84°C; excluding the avidin+biotin boundary it is 3.58°C — within the between-laboratory scatter of DSC itself, where the same protein's published melt commonly spans several degrees by prep and scan rate (ovalbumin alone is reported from about 77 to 85°C). Rank correlation (Spearman ρ) is 0.915; 15 of 20 land within 5°C and 8 within 2°C. Of the 30 proteins, the 23 with a defined fold were numbered and all 7 without one were declined: sensitivity 1.00 and specificity 0.958 for knowing when not to answer (the single call counted against specificity is microbial transglutaminase, declined as an uncalibrated fold — the safe direction), and no melting point was fabricated. Reported in the OECD QMRF / QPRF format. Blind: every melt temperature was withheld until scoring.

This month: allergen family, food scope

September 2026. The WHO/IUIS food-allergen panel assigned to its protein family from fold and motif alone — structure-only, not a name or database lookup — with an out-of-scope decoy control. Re-run fresh through the live classifier.

ReadMetricResultNBasis
Allergen family (WHO/IUIS food panel)exact / +super / wrong / refuse93% / 95% / 0% / 5%120held-out
Decoy control (non-allergen proteins)out-of-scope correctly refused39 / 3939decoy
2S albuminfamily hit12 / 1212held-out
Non-specific lipid-transfer proteinfamily hit8 / 88held-out
Tropomyosinfamily hit8 / 88held-out
Parvalbuminfamily hit7 / 77held-out
Oleosinfamily hit7 / 77held-out
Profilinfamily hit6 / 66held-out

Across 120 WHO/IUIS food allergens the exact family is named 112 times (93%), with 2 more correct to the superfamily (95% exact-or-super), zero wrong-family calls, and 6 refusals (5%) rather than a guess. As a control, 39 out-of-scope non-allergen proteins (polygalacturonases, ribosomal P1/P2, triosephosphate isomerases, tubulins) were every one declared out of scope — no false flags. The call is made from fold and motif, not a name or database lookup, and the classifier refuses rather than force a family. This is the EFSA pre-bench homology step that runs ahead of the wet-lab IgE panel: it lowers cross-reactivity concern, it is not an all-clear.

This month: enzyme thermostability

September 2026. Eleven food and industrial enzymes, blind and structure-only: the fold recognised in every case (11 of 11) and the melt temperature predicted from sequence.

Enzyme (fold)Measured °CPredicted °CΔ °C
Lysozyme (GH22)74.874.80.0
Cel7A cellulase (GH7)62.063.3+1.3
K. lactis lactase (GH2)39.341.2+1.9
Taka-amylase (GH13)62.064.2+2.2
Glucose oxidase (GMC)55.858.4+2.6
Papain (C1 protease)83.078.7−4.3
Xylanase (GH11)58.863.3+4.5
Trypsin (S1 protease)54.058.8+4.8
B. licheniformis α-amylase · named outlier: calcium-saturated, heat-stabilised101.064.2−36.8
Cold-active α-amylase · named outlier: psychrophilic adaptation43.764.2+20.5

The fold is recognised in all eleven; the melt is predicted within a mean 2.64°C on the condition-matched enzymes, 8.58°C including the two α-amylase outliers. Those two are named rather than hidden: one α-amylase anchor cannot span a calcium-saturated, heat-stabilised Bacillus amylase at 101°C and a cold-active (psychrophilic) amylase at 44°C, because the enzyme's calcium state and cold adaptation are not visible in the sequence alone — a per-enzyme calcium-state input closes that gap, the same way an apo/holo input closes the iron clamp on transferrin. The disordered casein control was correctly refused. Blind: every melt temperature was withheld until scoring.

Third-party, contracted

The one number an outside lab, not us, produced.

R² 0.95
Pearson r 0.97 · MAE 1.10 · N=20 in triplicate

An independent, contracted food-technology lab measured bitterness on its seven-sensor Alpha-MOS ASTREE electronic tongue; we predicted the same bitterness from structure alone. Every reading was taken in triplicate (n=3, reported as mean and standard deviation). The tongue was first calibrated against eight bitter reference standards, from caffeine to denatonium benzoate, giving a standard curve of R² 0.94; we then compared our structure-only predictions against 20 measured values: those eight standards plus four polyphenols (naringin, quercetin, curcumin, chlorogenic acid) at three dilutions each. Predicted and measured are independently correlated at R² 0.95, as tight as the instrument's own calibration. It is a contracted correlation, not a blind test: the lab itself noted that low-solubility polyphenols read inconsistently because the tongue senses only the dissolved fraction, and ASTREE can conflate astringency with bitterness, so we frame it precisely.

What each read means

Open a read for the method, what it means for your R&D, and the honest limit.

Taste, 6-class

Method. All 3,517 public ChemTastesDB tastants graded through the engine; the 6-way call scored against the database's own labels; baseline is the same predictions with labels shuffled. Means. From structure alone, zero training, it calls the correct taste class 72.7% of the time, more than double the ~32% chance rate, so it pre-screens taste before synthesis. Limit. Strongest on sweet, bitter and umami; weakest on the rare sour and salty; a screen, not a trained-panel replacement.

Allergen family & cross-reactivity

Method. Each of 285 real WHO/IUIS food allergens assigned to its protein family from fold and motif, structure-only (not a name lookup); and, for a novel protein, the FAO/WHO Codex 80-aa window screen names its cross-reactive partner. Means. It places a protein in the exact allergen family 88% of the time (251 of 285), with 8 more correct to the superfamily (91% exact-or-super), makes two wrong-family calls, and declines on 24 (8%) rather than guess, and it maps cross-reactivity (peanut to soy, walnut to hazelnut, insect to shellfish and dust mite). Limit. A fast structure-only pre-screen: it runs the accepted FAO/WHO Codex 80-aa rule and couples it to the melt and functionality read, then feeds the formal Codex/EFSA sequence comparison and your lab confirmation. It front-runs that decision, it does not replace it, and it is not a clinical IgE or food-challenge diagnosis.

Protein melt temperature, mass & size

Method. Melt temperature predicted from sequence and graded blind against published DSC on ~35 emerging proteins; intact mass from the atomic sum; size from the surface-area law, scored on cross-lab ion-mobility compendia. Means. In-family melt lands within 3.2°C (near-exact on recognised folds: napin 108 vs 108.8, ovalbumin 84.4 vs 84.5), mass is exact to ~0.1% on every protein, and size to a few percent. It also refuses folds outside its scope (prolamins, disordered caseins) rather than fabricate a number. At proteome scale, given 150 arbitrary proteins from across the tree of life (mostly enzymes, membrane and disordered chains), it returned a melt temperature for only the ~11% whose fold it is calibrated on and abstained on the other 89%, rather than invent one. Limit. Cross-family melt widens to about 6.6°C; the refusals are deliberate, and the in-scope accuracy rests on the food-protein waves, not on the arbitrary-proteome set. On the strict pilot bar (Td within 5°C plus the correct family), a broad protein mix lands about 65 to 68% blind: the same cupin fold scatters ~17°C by species (soy 7S 70, peanut 7S 87, phaseolin 92), so we score a protein against its disclosed band, agreed before you send, rather than a blanket 5°C point.

Aroma, DIAAS, gelation (screens)

Method. Odor-family recall scored on thousands of odorants across three open benchmarks; DIAAS from composition and digestibility vs the measured table; gelation and emulsification as bands and orderings. Means. Aroma recall runs ~8 to 10x over chance on thousands of compounds with zero training; DIAAS reads protein-quality direction (globulin 0.66 vs prolamin 0.24); gelation ranks which proteins set and their temperature windows. Limit. These are screens: direction, ordering and windows, not certified absolutes. The lab value stays official.

Food-safety genotox screen

Method. A blind decoy panel of known food genotoxins plus GRAS-safe look-alikes, screened for reactive and genotoxic structural alerts. Means. It flagged all 8 hazards (acrylamide, benzo[a]pyrene, the sassafras/tarragon alkenylbenzenes) and cleared all 4 safe cousins, zero over-flag; it separates eugenol (safe) from methyleugenol (genotoxic) on a one-methyl-group difference, the exact hard call a GRAS reviewer makes. Limit. A structural-alert hazard screen on a small hard-case panel, not a dose or exposure risk assessment.

Molecular structure
Every read is computed from the sequence, before the sample exists.

Where the field is going, and what we are doing

The direction of the field, and how a zero-training, structure-only screen fits it.

Computation moved to the front of the lab

In-silico prediction is now a core step in food and protein R&D, not a novelty, with the in-silico protein-design market estimated to roughly double by 2030 (one analyst estimate). The shared industry pattern is predict, then make and test only the winners.

Interpretable is the bar for high-stakes calls

Reviewers and regulators increasingly hold that accuracy alone is not enough for safety decisions; a model must be transparent and reproducible (Rudin, Nature Machine Intelligence 2019; the OECD QSAR validation principles). A first-principles, mechanism-named read is built for that, not patched for it.

Regulators are adopting in-silico screens

New Approach Methodologies, including in-silico pre-screens, are formally endorsed to reduce bench and animal testing (FDA NAMs program; EFSA's 2025 novel-food guidance integrates in-silico allergenicity). The FAO/WHO Codex 80-aa homology rule is the accepted first-tier allergenicity screen, which is exactly the screen we run, as a fast structure-only pre-screen we couple to the melt and functionality read and hand to the formal Codex/EFSA comparison; we complement that decision, we do not replace it.

The gap we answer

Practitioners note the Codex homology rule over-flags (Ladics 2007), and EFSA flags the lack of a standardised, reproducible engine. We couple the window to fold, abstain on 8% rather than force a call, and report our own wrong-family rate (2 in 285), so the false-positive behaviour is quantified, not hidden. Peer platforms share the predict-first pitch but learn from training data; our differentiator is that we do not train at all.

How we test

Every accuracy is re-run against a null: we shuffle the labels and recompute. If the real score does not clearly beat the shuffled one, it does not go here. The taste figure is 72.7% vs a 32.3% shuffle-null, above the 45.9% majority-class (always-bitter) baseline, on the full set (N=3,517, p=0.002), a claim a data scientist can check.

Where it defers

A screen, not an assay. When a call is structurally ambiguous it says provisional, confirm by panel rather than guess; when a fold is outside its tested scope it declines and still returns the biophysical screen. Certified numbers, the compliance DIAAS and the allergen assay, stay measured in the lab.

The food domain
One method, across the food-protein domain.
The set keeps growing. Each month we add held-out studies straight from the published literature, run structure-only, then checked against their own shuffle-null before they join the set. We are scaling onto the corpora the field respects: Meltome and ProThermDB for stability, ChemTastesDB and BitterSweet for taste, WHO/IUIS with COMPARE and AllergenOnline for allergens, Pyrfume for aroma.

Run it on your own proteins →