Tabular foundation models · 5×5 cross validation · two data sets

Swapping the head

NVIDIA's structured-data-models package1 puts several tabular foundation models behind one in-context interface: TabICLv22, Kumo Tabular in three sizes3 and Google's TabFM4. Each is the same kind of object as the TabPFN6 that reads Monroe's frozen molecular embedding5 in the strongest arm of the nine-method foundation-model comparison. So each one is dropped into that slot here, with the representation, the folds and the test sets held fixed, and scored against Monroe + TabPFN 3 on fifteen ADME endpoints. Kumo large is the best of them: on top of all 45 endpoint × metric combinations after Tukey's correction, and ahead of the reference on 28 of 45 in a paired test, behind on 6. It is also a small effect next to what the representation is worth.

5NVIDIA heads
1frozen encoder
15endpoints
5×5cross validation
1,875fold models fitted

What differs between the arms, and what does not

Every arm on this page reads the same 720-dimensional Monroe embedding of each molecule, cached once and never updated. Every arm is handed the same training molecules in every fold — the four fifths of the training set outside the held-out fifth that have a value for the endpoint — and is scored on the same untouched test set. Nothing is trained downstream in any of them: each head takes the labelled training rows as context and predicts the test rows in one forward pass. So a difference between two arms is the head and nothing else.

Nothing is tuned either. Each NVIDIA model runs at the ensemble size and precision NVIDIA's own TabArena adapter uses for it, with its own default preprocessing, and the point prediction is what that adapter returns: the mean of the predictive distribution, which is also what the reference head reports. LightGBM on Morgan counts is carried along as the floor every arm has to clear.

The scoreboard

Tukey HSD per endpoint and metric, counted the way every page on this site counts: a method is on top if the correction cannot separate it from the best mean, and it is only best alone if nothing else is on top with it. The tables show MAE, in the units of each endpoint.

ExpansionRx7

LightGBM + Morgan0 best alone · 1 tied for best · 26 worse
Monroe + TabPFN 3 (reference)0 best alone · 16 tied for best · 11 worse
Monroe + TabICL (tabicl)0 best alone · 16 tied for best · 11 worse
Monroe + TabICLv2 (sdm)0 best alone · 17 tied for best · 10 worse
Monroe + Kumo small0 best alone · 21 tied for best · 6 worse
Monroe + Kumo medium0 best alone · 22 tied for best · 5 worse
Monroe + Kumo large2 best alone · 25 tied for best · 0 worse
Monroe + TabFM0 best alone · 23 tied for best · 4 worse
EndpointLightGBMTabPFN 3 (ref)TabICL (tabicl)TabICLv2 (sdm)Kumo SKumo MKumo LTabFM
LogD0.545±0.0120.402±0.0150.397±0.0110.393±0.0090.385±0.012tied0.389±0.0120.375±0.011tied0.390±0.012
LogS0.454±0.0150.380±0.009tied0.382±0.013tied0.381±0.011tied0.393±0.0120.388±0.011tied0.382±0.013tied0.388±0.010tied
LOG_HLM0.332±0.0090.293±0.0090.289±0.008tied0.288±0.008tied0.285±0.010tied0.285±0.010tied0.284±0.008tied0.290±0.008tied
LOG_MLM0.410±0.0140.358±0.011tied0.371±0.0120.368±0.0110.370±0.0110.371±0.0120.366±0.012tied0.367±0.007tied
LOG_Caco_AB0.340±0.0220.248±0.011tied0.273±0.0140.274±0.0150.248±0.010tied0.245±0.013tied0.243±0.014tied0.241±0.009tied
LOG_Caco_Efflux0.432±0.0190.346±0.0120.340±0.0180.340±0.0150.329±0.014tied0.328±0.013tied0.327±0.014tied0.324±0.014tied
LOG_MPPB0.288±0.0200.259±0.013tied0.258±0.015tied0.259±0.015tied0.263±0.013tied0.267±0.014tied0.259±0.014tied0.267±0.013tied
LOG_MBPB0.223±0.0180.192±0.016tied0.184±0.017tied0.184±0.016tied0.192±0.011tied0.189±0.020tied0.181±0.016tied0.196±0.013
LOG_MGMB0.302±0.0340.204±0.033tied0.198±0.031tied0.198±0.033tied0.202±0.033tied0.206±0.032tied0.205±0.032tied0.196±0.033tied
Tukey HSD on MAE, ExpansionRx. Each bar is a method's mean over 25 folds with an interval widened to cover every pairwise comparison at once. Blue is the best mean, grey cannot be separated from it, red is worse.
Tukey HSD on MAE, ExpansionRx. Each bar is a method's mean over 25 folds with an interval widened to cover every pairwise comparison at once. Blue is the best mean, grey cannot be separated from it, red is worse.

Biogen ADME8

LightGBM + Morgan0 best alone · 0 tied for best · 18 worse
Monroe + TabPFN 3 (reference)0 best alone · 6 tied for best · 12 worse
Monroe + TabICL (tabicl)0 best alone · 6 tied for best · 12 worse
Monroe + TabICLv2 (sdm)0 best alone · 6 tied for best · 12 worse
Monroe + Kumo small0 best alone · 10 tied for best · 8 worse
Monroe + Kumo medium0 best alone · 17 tied for best · 1 worse
Monroe + Kumo large0 best alone · 18 tied for best · 0 worse
Monroe + TabFM0 best alone · 8 tied for best · 10 worse
EndpointLightGBMTabPFN 3 (ref)TabICL (tabicl)TabICLv2 (sdm)Kumo SKumo MKumo LTabFM
LOG_SOL0.413±0.0090.322±0.0040.326±0.0040.326±0.0040.319±0.004tied0.319±0.004tied0.315±0.004tied0.320±0.004
LOG_HLM0.395±0.0050.299±0.0030.313±0.0030.313±0.0030.298±0.0030.293±0.003tied0.293±0.003tied0.301±0.002
LOG_RLM0.470±0.0050.375±0.0030.383±0.0030.382±0.0030.370±0.0030.365±0.002tied0.365±0.002tied0.370±0.003
LOG_MDR1_ER0.385±0.0050.294±0.0020.300±0.0020.298±0.0020.291±0.0020.287±0.002tied0.287±0.003tied0.296±0.003
LOG_HPPB0.576±0.0330.438±0.017tied0.433±0.031tied0.434±0.027tied0.439±0.025tied0.449±0.030tied0.437±0.021tied0.429±0.019tied
LOG_RPPB0.554±0.0210.445±0.026tied0.456±0.032tied0.454±0.029tied0.449±0.028tied0.459±0.032tied0.453±0.026tied0.446±0.022tied
Tukey HSD on MAE, Biogen ADME.
Tukey HSD on MAE, Biogen ADME.

With seven heads on one representation, most combinations end in a tie, and one arm is never out of it. Kumo large is on top of 27 of 27 ExpansionRx combinations and 18 of 18 Biogen ones — no other arm on either data set is never significantly worse than the best. The TabPFN 3 reference is on top of 16 and 6, and the two TabICLv2 arms of the same number or one more. LightGBM on Morgan counts reaches the top of 1 of 45.

Against the reference

The scoreboard asks who is best. The question here is narrower: does swapping TabPFN 3 for one of NVIDIA's heads help? The folds are the pairing, which is legitimate because every arm saw identical training molecules in all 25. A win or a loss is a paired t-test at p < 0.05 on one endpoint and one metric, uncorrected, so read the counts as a direction rather than a verdict on any single cell.

Head, on Monroe's embeddingExpansionRx, 27Biogen, 18Mean over 15 endpoints
winsno calllosseswinsno calllossesΔ R²Δ MAE
Monroe + TabICL (tabicl)11970414-0.005+0.0031
Monroe + TabICLv2 (sdm)14670414-0.004+0.0024
Monroe + Kumo small101251260+0.006-0.0016
Monroe + Kumo medium12871215+0.003-0.0012
Monroe + Kumo large16831233+0.014-0.0055
Monroe + TabFM15391053+0.005-0.0022
Each head against Monroe + TabPFN 3 on ExpansionRx: mean paired MAE difference over 25 folds with a 95% interval, signed so that right of zero is better than the reference.
Each head against Monroe + TabPFN 3 on ExpansionRx: mean paired MAE difference over 25 folds with a 95% interval, signed so that right of zero is better than the reference.
The same on Biogen ADME.
The same on Biogen ADME.

All three Kumo Tabular sizes beat the reference more often than they lose to it, and the large one most clearly: 28 wins against 6 losses, a mean R² gain of +0.020 on ExpansionRx and +0.005 on Biogen. The gains are largest on LogD, human microsomal stability and the two Caco-2 endpoints. The reference holds on ExpansionRx mouse microsomal stability, where every head is behind it, and on Biogen rat plasma protein binding, where every head but Kumo small and TabFM is. Tissue binding is otherwise mixed: on ExpansionRx brain and muscle binding the two TabICLv2 heads are ahead of it, Kumo large on brain binding and TabFM on muscle binding.

TabICLv2 goes the other way: 14 wins against 21 losses, and the split is by data set. 14 of those losses are on Biogen, where it trails the reference on every endpoint but human plasma binding. On ExpansionRx it wins more than it loses, apart from mouse microsomal stability and Caco-2 A→B permeability, where it gives back 0.026 MAE — the widest gap to the reference anywhere on the page.

TabFM, at 1.6 billion parameters the largest model here, lands among the Kumo models rather than above them: 25 wins against 12 losses, more wins than Kumo small and more than twice the losses. Its best endpoint is ExpansionRx Caco-2 efflux, +0.085 R² over the reference. It loses to the reference on mouse microsomal stability, as every head does, and on mouse plasma binding, as the smaller Kumo models do. Against Kumo large directly it loses: Kumo large is significantly ahead on 28 of 45 endpoint × metric combinations and behind on 7, and TabFM takes 8.5 times the GPU time to get there.

The control: one checkpoint, two implementations

The source comparison already had a Monroe + TabICL arm, fitted with the TabICL authors' own tabicl package, version 2.1.1. That release loads the same checkpoint NVIDIA's port does, tabicl-regressor-v2-20260212.ckpt. Putting the two side by side tests the wrapper rather than the model. If NVIDIA's reimplementation were subtly wrong, the other heads' numbers would mean less.

MetricMedian |mean Δ|Largest |mean Δ|Mean |Δ| on one foldEndpoints p < 0.05
R²0.00300.00780.01833 / 15
Spearman ρ0.00160.00550.00934 / 15
MAE0.00090.00350.00572 / 15

They agree. The largest mean difference on any endpoint is 0.0078 in R² and 0.0035 in MAE, an order of magnitude below the gaps between heads. 9 of 45 endpoint × metric combinations do reach p < 0.05, and all of them favour NVIDIA's port by a hair. Paired folds are sensitive enough to see differences of a few thousandths, which is what two implementations with different preprocessing and random ensemble members produce. Against the reference the two land in exactly the same place, 21 losses each.

Does size help?

Kumo Tabular ships in three sizes: 28 million, 63 million and 216 million parameters, the largest also running twice the ensemble. Paired over the same folds, large against small:

Size helps, but not monotonically: the medium model buys nothing over the small one, losing to the reference 12 times to the small model's 5. The large model's edge over the small one is real but a few thousandths of MAE on most endpoints. It costs about seven times as much GPU time.

HeadEstimatorsMedian s / foldSlowest fold, sAll 375 folds, minWins / losses vs ref
Monroe + TabICLv2 (sdm)82.87.01814 / 21
Monroe + Kumo small81.65.31022 / 5
Monroe + Kumo medium85.113.03324 / 12
Monroe + Kumo large1610.425.96828 / 6
Monroe + TabFM3262.8379.657825 / 12

Wall-clock time on one RTX 5070 Ti, per fold, context and test set in one forward pass. The largest fold is ExpansionRx LogD, with 4,230 training rows and 2,270 test rows. Kumo small is the fastest head on the page and beats the reference more often than it loses to it.

TabFM is the exception to one forward pass. On the two largest ExpansionRx endpoints, LogD and LogS, 4,230 context rows and 2,270 test rows do not fit in 16 GB at once, so those 50 folds predict the test rows 1,024 at a time. A test row attends only to the context, never to another test row, and every chunk is seeded identically, so this is the single pass computed in parts; checked on a fold that does fit, the two agree to within 3×10−3 on a 6.9-unit range, which is half-precision rounding. Every other fold, TabFM and otherwise, is one pass and reproduces bit for bit.

The head is the small lever

The source comparison found that the representation is worth about ten times the head, from two representations and two heads. This page has seven heads on one representation, which prices the head more precisely.

On each endpoint, the spread in mean R² between the best and worst in-context head has a median of 0.042. Monroe + TabPFN 3's lead over LightGBM on Morgan counts — a change of representation and of model — has a median of 0.208. Choosing among modern tabular foundation models is worth about a fifth of what choosing the representation is worth. Kumo large is the best head here, but a better head does not turn a frozen encoder into a fine-tuned one: on the three ExpansionRx tissue-binding endpoints, CheMeleon's fine-tuned graph network is still ahead of every head on this page by 0.02 to 0.05 MAE.

What this does and does not say

Scope

One representation. Every head reads Monroe's embedding. Whether Kumo's lead holds on fingerprints, descriptors or another encoder is not tested here, and heads can rank differently on different inputs.

Defaults only. No head is tuned, and none is fine-tuned, although the package supports fine-tuning Kumo small and TabICLv2. The Kumo and TabFM recipes keep at most 500 of the 720 embedding dimensions per ensemble member, shuffled differently for each, so no single member sees the whole embedding. TabICLv2 reads all 720.

Two of the package's models are not arms. TimesFM 3 forecasts time series and Kumo Relational models linked tables. Neither has anything to say about one table of molecules with one numeric target.

Licences. TabICLv2's weights are BSD-3-Clause and Kumo Tabular's are OpenMDW. TabFM's are under Google's TabFM Non-Commercial License, which allows benchmarking like this and nothing commercial; it was accepted for this study on that basis.

Reported the way “Even More Thoughts on ML Method Comparisons” and the protocol paper it points at9 argue comparisons should be: distributions over folds, simultaneous confidence intervals, paired tests on the folds, and no bolded maxima. Where two methods cannot be separated, both are reported as tied.

Code, per-fold metrics and every figure: studies/sdm-tabular.

References

  1. NVIDIA. Structured Data Models. GitHub repository, 2026, Apache 2.0; run here at commit 5d66336. github.com/NVIDIA/structured-data-models
  2. Qu, J.; et al. TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation Model. ICML 2026. Weights BSD-3-Clause, at huggingface.co/jingang/TabICL. arXiv:2602.11139
  3. Qu, J.; et al. NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction. Hugging Face blog, 2026. Weights under OpenMDW 1.1, at huggingface.co/nvidia/Kumo-Tabular. huggingface.co/blog/nvidia/kumo-tabular
  4. Kong, W.; et al. Introducing TabFM: A Zero-shot Foundation Model for Tabular Data. Google Research blog, 2026. Weights under the TabFM Non-Commercial License v1.0. research.google/blog
  5. Banaszewski, B.; Fitzgibbon, A. W. Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference. Preprint, 2026. The frozen encoder every arm on this page reads. arXiv:2608.18982
  6. Hollmann, N.; Müller, S.; Purucker, L.; et al. Accurate Predictions on Small Data with a Tabular Foundation Model. Nature 2025, 637 (8045), 319–326. TabPFN 3 is the reference head. 10.1038/s41586-024-08328-6
  7. OpenADMET; Expansion Therapeutics. OpenADMET–ExpansionRx Blind Challenge Data. Hugging Face, 2026, CC BY 4.0. 10.57967/hf/9687
  8. Fang, C.; Wang, Y.; Grater, R.; et al. Prospective Validation of Machine Learning Algorithms for Absorption, Distribution, Metabolism, and Excretion Prediction: An Industrial Perspective. J. Chem. Inf. Model. 2023, 63 (11), 3263–3274. 10.1021/acs.jcim.3c00160
  9. Ash, J. R.; Wognum, C.; Rodríguez-Pérez, R.; et al. Practically Significant Method Comparison Protocols for Machine Learning in Small Molecule Drug Discovery. J. Chem. Inf. Model. 2025, 65 (18), 9398–9411. 10.1021/acs.jcim.5c01609