Tabular foundation models · 5×5 cross validation · two data sets
NVIDIA's structured-data-models package1 puts several tabular foundation models behind one in-context interface: TabICLv22, Kumo Tabular in three sizes3 and Google's TabFM4. Each is the same kind of object as the TabPFN6 that reads Monroe's frozen molecular embedding5 in the strongest arm of the nine-method foundation-model comparison. So each one is dropped into that slot here, with the representation, the folds and the test sets held fixed, and scored against Monroe + TabPFN 3 on fifteen ADME endpoints. Kumo large is the best of them: on top of all 45 endpoint × metric combinations after Tukey's correction, and ahead of the reference on 28 of 45 in a paired test, behind on 6. It is also a small effect next to what the representation is worth.
Every arm on this page reads the same 720-dimensional Monroe embedding of each molecule, cached once and never updated. Every arm is handed the same training molecules in every fold — the four fifths of the training set outside the held-out fifth that have a value for the endpoint — and is scored on the same untouched test set. Nothing is trained downstream in any of them: each head takes the labelled training rows as context and predicts the test rows in one forward pass. So a difference between two arms is the head and nothing else.
Nothing is tuned either. Each NVIDIA model runs at the ensemble size and precision NVIDIA's own TabArena adapter uses for it, with its own default preprocessing, and the point prediction is what that adapter returns: the mean of the predictive distribution, which is also what the reference head reports. LightGBM on Morgan counts is carried along as the floor every arm has to clear.
Tukey HSD per endpoint and metric, counted the way every page on this site counts: a method is on top if the correction cannot separate it from the best mean, and it is only best alone if nothing else is on top with it. The tables show MAE, in the units of each endpoint.
| Endpoint | LightGBM | TabPFN 3 (ref) | TabICL (tabicl) | TabICLv2 (sdm) | Kumo S | Kumo M | Kumo L | TabFM |
|---|---|---|---|---|---|---|---|---|
| LogD | 0.545±0.012 | 0.402±0.015 | 0.397±0.011 | 0.393±0.009 | 0.385±0.012tied | 0.389±0.012 | 0.375±0.011tied | 0.390±0.012 |
| LogS | 0.454±0.015 | 0.380±0.009tied | 0.382±0.013tied | 0.381±0.011tied | 0.393±0.012 | 0.388±0.011tied | 0.382±0.013tied | 0.388±0.010tied |
| LOG_HLM | 0.332±0.009 | 0.293±0.009 | 0.289±0.008tied | 0.288±0.008tied | 0.285±0.010tied | 0.285±0.010tied | 0.284±0.008tied | 0.290±0.008tied |
| LOG_MLM | 0.410±0.014 | 0.358±0.011tied | 0.371±0.012 | 0.368±0.011 | 0.370±0.011 | 0.371±0.012 | 0.366±0.012tied | 0.367±0.007tied |
| LOG_Caco_AB | 0.340±0.022 | 0.248±0.011tied | 0.273±0.014 | 0.274±0.015 | 0.248±0.010tied | 0.245±0.013tied | 0.243±0.014tied | 0.241±0.009tied |
| LOG_Caco_Efflux | 0.432±0.019 | 0.346±0.012 | 0.340±0.018 | 0.340±0.015 | 0.329±0.014tied | 0.328±0.013tied | 0.327±0.014tied | 0.324±0.014tied |
| LOG_MPPB | 0.288±0.020 | 0.259±0.013tied | 0.258±0.015tied | 0.259±0.015tied | 0.263±0.013tied | 0.267±0.014tied | 0.259±0.014tied | 0.267±0.013tied |
| LOG_MBPB | 0.223±0.018 | 0.192±0.016tied | 0.184±0.017tied | 0.184±0.016tied | 0.192±0.011tied | 0.189±0.020tied | 0.181±0.016tied | 0.196±0.013 |
| LOG_MGMB | 0.302±0.034 | 0.204±0.033tied | 0.198±0.031tied | 0.198±0.033tied | 0.202±0.033tied | 0.206±0.032tied | 0.205±0.032tied | 0.196±0.033tied |
| Endpoint | LightGBM | TabPFN 3 (ref) | TabICL (tabicl) | TabICLv2 (sdm) | Kumo S | Kumo M | Kumo L | TabFM |
|---|---|---|---|---|---|---|---|---|
| LOG_SOL | 0.413±0.009 | 0.322±0.004 | 0.326±0.004 | 0.326±0.004 | 0.319±0.004tied | 0.319±0.004tied | 0.315±0.004tied | 0.320±0.004 |
| LOG_HLM | 0.395±0.005 | 0.299±0.003 | 0.313±0.003 | 0.313±0.003 | 0.298±0.003 | 0.293±0.003tied | 0.293±0.003tied | 0.301±0.002 |
| LOG_RLM | 0.470±0.005 | 0.375±0.003 | 0.383±0.003 | 0.382±0.003 | 0.370±0.003 | 0.365±0.002tied | 0.365±0.002tied | 0.370±0.003 |
| LOG_MDR1_ER | 0.385±0.005 | 0.294±0.002 | 0.300±0.002 | 0.298±0.002 | 0.291±0.002 | 0.287±0.002tied | 0.287±0.003tied | 0.296±0.003 |
| LOG_HPPB | 0.576±0.033 | 0.438±0.017tied | 0.433±0.031tied | 0.434±0.027tied | 0.439±0.025tied | 0.449±0.030tied | 0.437±0.021tied | 0.429±0.019tied |
| LOG_RPPB | 0.554±0.021 | 0.445±0.026tied | 0.456±0.032tied | 0.454±0.029tied | 0.449±0.028tied | 0.459±0.032tied | 0.453±0.026tied | 0.446±0.022tied |
With seven heads on one representation, most combinations end in a tie, and one arm is never out of it. Kumo large is on top of 27 of 27 ExpansionRx combinations and 18 of 18 Biogen ones — no other arm on either data set is never significantly worse than the best. The TabPFN 3 reference is on top of 16 and 6, and the two TabICLv2 arms of the same number or one more. LightGBM on Morgan counts reaches the top of 1 of 45.
The scoreboard asks who is best. The question here is narrower: does swapping TabPFN 3 for one of NVIDIA's heads help? The folds are the pairing, which is legitimate because every arm saw identical training molecules in all 25. A win or a loss is a paired t-test at p < 0.05 on one endpoint and one metric, uncorrected, so read the counts as a direction rather than a verdict on any single cell.
| Head, on Monroe's embedding | ExpansionRx, 27 | Biogen, 18 | Mean over 15 endpoints | |||||
|---|---|---|---|---|---|---|---|---|
| wins | no call | losses | wins | no call | losses | Δ R² | Δ MAE | |
| Monroe + TabICL (tabicl) | 11 | 9 | 7 | 0 | 4 | 14 | -0.005 | +0.0031 |
| Monroe + TabICLv2 (sdm) | 14 | 6 | 7 | 0 | 4 | 14 | -0.004 | +0.0024 |
| Monroe + Kumo small | 10 | 12 | 5 | 12 | 6 | 0 | +0.006 | -0.0016 |
| Monroe + Kumo medium | 12 | 8 | 7 | 12 | 1 | 5 | +0.003 | -0.0012 |
| Monroe + Kumo large | 16 | 8 | 3 | 12 | 3 | 3 | +0.014 | -0.0055 |
| Monroe + TabFM | 15 | 3 | 9 | 10 | 5 | 3 | +0.005 | -0.0022 |
All three Kumo Tabular sizes beat the reference more often than they lose to it, and the large one most clearly: 28 wins against 6 losses, a mean R² gain of +0.020 on ExpansionRx and +0.005 on Biogen. The gains are largest on LogD, human microsomal stability and the two Caco-2 endpoints. The reference holds on ExpansionRx mouse microsomal stability, where every head is behind it, and on Biogen rat plasma protein binding, where every head but Kumo small and TabFM is. Tissue binding is otherwise mixed: on ExpansionRx brain and muscle binding the two TabICLv2 heads are ahead of it, Kumo large on brain binding and TabFM on muscle binding.
TabICLv2 goes the other way: 14 wins against 21 losses, and the split is by data set. 14 of those losses are on Biogen, where it trails the reference on every endpoint but human plasma binding. On ExpansionRx it wins more than it loses, apart from mouse microsomal stability and Caco-2 A→B permeability, where it gives back 0.026 MAE — the widest gap to the reference anywhere on the page.
TabFM, at 1.6 billion parameters the largest model here, lands among the Kumo models rather than above them: 25 wins against 12 losses, more wins than Kumo small and more than twice the losses. Its best endpoint is ExpansionRx Caco-2 efflux, +0.085 R² over the reference. It loses to the reference on mouse microsomal stability, as every head does, and on mouse plasma binding, as the smaller Kumo models do. Against Kumo large directly it loses: Kumo large is significantly ahead on 28 of 45 endpoint × metric combinations and behind on 7, and TabFM takes 8.5 times the GPU time to get there.
The source comparison already had a Monroe + TabICL arm, fitted with the TabICL authors' own tabicl package, version 2.1.1. That release loads the same checkpoint NVIDIA's port does, tabicl-regressor-v2-20260212.ckpt. Putting the two side by side tests the wrapper rather than the model. If NVIDIA's reimplementation were subtly wrong, the other heads' numbers would mean less.
| Metric | Median |mean Δ| | Largest |mean Δ| | Mean |Δ| on one fold | Endpoints p < 0.05 |
|---|---|---|---|---|
| R² | 0.0030 | 0.0078 | 0.0183 | 3 / 15 |
| Spearman ρ | 0.0016 | 0.0055 | 0.0093 | 4 / 15 |
| MAE | 0.0009 | 0.0035 | 0.0057 | 2 / 15 |
They agree. The largest mean difference on any endpoint is 0.0078 in R² and 0.0035 in MAE, an order of magnitude below the gaps between heads. 9 of 45 endpoint × metric combinations do reach p < 0.05, and all of them favour NVIDIA's port by a hair. Paired folds are sensitive enough to see differences of a few thousandths, which is what two implementations with different preprocessing and random ensemble members produce. Against the reference the two land in exactly the same place, 21 losses each.
Kumo Tabular ships in three sizes: 28 million, 63 million and 216 million parameters, the largest also running twice the ensemble. Paired over the same folds, large against small:
Size helps, but not monotonically: the medium model buys nothing over the small one, losing to the reference 12 times to the small model's 5. The large model's edge over the small one is real but a few thousandths of MAE on most endpoints. It costs about seven times as much GPU time.
| Head | Estimators | Median s / fold | Slowest fold, s | All 375 folds, min | Wins / losses vs ref |
|---|---|---|---|---|---|
| Monroe + TabICLv2 (sdm) | 8 | 2.8 | 7.0 | 18 | 14 / 21 |
| Monroe + Kumo small | 8 | 1.6 | 5.3 | 10 | 22 / 5 |
| Monroe + Kumo medium | 8 | 5.1 | 13.0 | 33 | 24 / 12 |
| Monroe + Kumo large | 16 | 10.4 | 25.9 | 68 | 28 / 6 |
| Monroe + TabFM | 32 | 62.8 | 379.6 | 578 | 25 / 12 |
Wall-clock time on one RTX 5070 Ti, per fold, context and test set in one forward pass. The largest fold is ExpansionRx LogD, with 4,230 training rows and 2,270 test rows. Kumo small is the fastest head on the page and beats the reference more often than it loses to it.
TabFM is the exception to one forward pass. On the two largest ExpansionRx endpoints, LogD and LogS, 4,230 context rows and 2,270 test rows do not fit in 16 GB at once, so those 50 folds predict the test rows 1,024 at a time. A test row attends only to the context, never to another test row, and every chunk is seeded identically, so this is the single pass computed in parts; checked on a fold that does fit, the two agree to within 3×10−3 on a 6.9-unit range, which is half-precision rounding. Every other fold, TabFM and otherwise, is one pass and reproduces bit for bit.
The source comparison found that the representation is worth about ten times the head, from two representations and two heads. This page has seven heads on one representation, which prices the head more precisely.
On each endpoint, the spread in mean R² between the best and worst in-context head has a median of 0.042. Monroe + TabPFN 3's lead over LightGBM on Morgan counts — a change of representation and of model — has a median of 0.208. Choosing among modern tabular foundation models is worth about a fifth of what choosing the representation is worth. Kumo large is the best head here, but a better head does not turn a frozen encoder into a fine-tuned one: on the three ExpansionRx tissue-binding endpoints, CheMeleon's fine-tuned graph network is still ahead of every head on this page by 0.02 to 0.05 MAE.
One representation. Every head reads Monroe's embedding. Whether Kumo's lead holds on fingerprints, descriptors or another encoder is not tested here, and heads can rank differently on different inputs.
Defaults only. No head is tuned, and none is fine-tuned, although the package supports fine-tuning Kumo small and TabICLv2. The Kumo and TabFM recipes keep at most 500 of the 720 embedding dimensions per ensemble member, shuffled differently for each, so no single member sees the whole embedding. TabICLv2 reads all 720.
Two of the package's models are not arms. TimesFM 3 forecasts time series and Kumo Relational models linked tables. Neither has anything to say about one table of molecules with one numeric target.
Licences. TabICLv2's weights are BSD-3-Clause and Kumo Tabular's are OpenMDW. TabFM's are under Google's TabFM Non-Commercial License, which allows benchmarking like this and nothing commercial; it was accepted for this study on that basis.
Reported the way “Even More Thoughts on ML Method Comparisons” and the protocol paper it points at9 argue comparisons should be: distributions over folds, simultaneous confidence intervals, paired tests on the folds, and no bolded maxima. Where two methods cannot be separated, both are reported as tied.
Code, per-fold metrics and every figure: studies/sdm-tabular.