NVIDIA releases Kumo Tabular, an open model for tabular prediction with no training

NVIDIA releases Kumo Tabular, an open model for tabular prediction with no training

NVIDIA has released Kumo Tabular, part of its NVIDIA Kumo Structured model collection, on Hugging Face. It is an open foundation model for tabular classification and regression. You hand it a table with labeled rows plus the rows you want predictions for, and it returns class probabilities or numeric predictions in a single forward pass, with no training, no tuning and no feature engineering. The model comes in three sizes, from 28M to 215M parameters, was pretrained only on artificial data, and is released under the OpenMDW-1.1 license for commercial use. It runs through NVIDIA's open-source structured-data-models library.

The pitch starts from the state of enterprise machine learning. Customer records, transactions, sensor logs, claims and orders live in tables, and for two decades the work has been done with gradient-boosted trees. NVIDIA's complaint is the lifecycle around them: every new question means collecting labels, engineering features, searching hyperparameters, validating and deploying a model that learns each task from scratch. Kumo Tabular borrows the in-context learning idea from large language models. A model pretrained on millions of tables reads a labeled table as its context and predicts labels for new rows directly, without updating any weights.

Architecturally, it is a Transformer built around the structure of a table, using column, row and in-context attention as introduced in TabICL and TabPFN. Cells are turned into tokens via Fourier features, with separate weights for numerical and categorical values; missing values need no imputation. Column attention learns what a value means within its column (whether a 42 is typical or extreme) using induced self-attention, so its cost grows linearly with the number of rows. Row attention learns how features interact, using rotary positions to tell columns apart. Four learnable [CLS] tokens join each row as its final readout, so the cost of the last stage no longer depends on the number of columns. A final Transformer then lets context rows attend to each other, while query rows attend to context rows only. Each prediction therefore depends only on the context and the row itself, and the context's keys and values are computed once and can be reused for follow-up predictions. Query rows use Test-GQA to shrink the cache each prediction reads. The head outputs class probabilities for classification and 999 quantiles for regression, from which a point prediction and an uncertainty estimate follow. A length-aware attention temperature, scaled with the logarithm of the number of keys and learned per attention head, is meant to keep attention sharp as tables grow longer or wider than typical training tables.

Training data is entirely synthetic. Each table is sampled from a Structural Causal Model: a configuration is drawn for the whole table, a random causal graph links hidden variables through randomly drawn functions (linear maps, small neural networks, trees or Gaussian processes), some nodes become numerical or categorical columns, one becomes the target and the rest stay hidden. Post-processing correlates column groups, clips outliers and injects missing values, and a quick tree-ensemble check discards any table without a learnable signal. NVIDIA added real-world messiness to the generator: several missingness patterns, coarsened features so duplicate rows may disagree on their label, categorical columns with many levels, and heavy-tailed regression targets. Classification and regression are trained as separate models, with cross-entropy loss and quantile loss respectively. As with TabICLv2, training runs in three stages: 1,024-row tables with up to 100 columns first and longest, then context lengths from 400 to 10,240 rows, then up to 60,000 rows, still with up to 100 columns. The Small, Medium and Large models saw about 35, 71 and 137 million artificial tables.

On results, NVIDIA says it ran all three sizes with default settings against the full TabArena leaderboard, which spans tuned gradient-boosted trees, AutoGluon and the latest tabular foundation models. Kumo Tabular ranks first overall with an ELO of 1950. The post also says it ran "17 faster" than LimiX-2 on a uniform single RTX 6000 Pro setup; the multiplication sign is missing in the text. Across the three sizes, NVIDIA claims a new state of the art on the accuracy-efficiency Pareto front. On BeyondArena it reaches an ELO of 1418 with an Improvability score of 7.78%, first on the leaderboard. On TALENT it takes the top overall ranking across classification accuracy, classification log-loss and regression RMSE, with average ranks of 6.67, 3.98 and 4.22. On ScoringBench, a benchmark for predictive distributions, Kumo Tabular-Large and Medium rank first and second on average rank.

Limits are stated plainly. The model handles numerical and categorical columns only; text, images or timestamps can be turned into features via built-in pre-processing recipes. One forward pass covers up to 10 classes, and the library extends this to any number of classes with error-correcting output codes. Accuracy may degrade on tables far beyond the training ranges or when query rows come from a different distribution than the context rows, so NVIDIA advises validating accuracy and calibration on your own held-out data before deployment. The training recipe and artificial data generators will be released soon.

Usage is a short snippet: load a pandas DataFrame into a sdm.TableTensor on CUDA, split rows with a known target (context) from rows with missing target (query), create sdm.models.KumoTabular(device="cuda"), and call it with x_context, y_context and x_query. The library downloads weights from the Hub on first use and provides the preprocessing, ensembling and many-class handling used in NVIDIA's evaluations. Code is at github.com/NVIDIA/structured-data-models and weights at huggingface.co/nvidia/Kumo-Tabular.

Key facts

  • Kumo Tabular is an open foundation model for tabular classification and regression that predicts labels of new rows in one forward pass, with no training, tuning or feature engineering.
  • Three sizes from 28M to 215M parameters, pretrained only on artificial tables sampled from Structural Causal Models; Small, Medium and Large saw about 35, 71 and 137 million tables.
  • NVIDIA reports first place on TabArena (ELO 1950), BeyondArena (ELO 1418, Improvability 7.78%), TALENT and ScoringBench, where Large and Medium rank first and second.
  • Released on Hugging Face under the OpenMDW-1.1 license for commercial use, with a GPU-native open-source library; the training recipe and data generators are promised soon.
  • Limits: numerical and categorical columns only, up to 10 classes per forward pass, and accuracy may degrade far beyond training ranges or under distribution shift.

Why it matters

Tabular data is the backbone of enterprise machine learning, and gradient-boosted trees have dominated it for two decades. NVIDIA argues the weak point is the workflow: labels, feature engineering, hyperparameter search, validation and deployment for every new question. Kumo Tabular applies the in-context learning idea from language models to tables, so a new task becomes a single forward pass over a labeled table. It builds on the TabICL and TabPFN line of attention designs, and NVIDIA says it sets a new state of the art on the accuracy-efficiency Pareto front across its three sizes.

Who it affects

Teams that predict things like churn, default, demand or price from customer records, transactions, sensor logs, claims and orders, and today build a gradient-boosted tree pipeline per task. It also touches researchers working on tabular foundation models, who now have open weights, a public library and, promised for later, the training recipe and synthetic data generators.

How to use it

Weights are at huggingface.co/nvidia/Kumo-Tabular and code at github.com/NVIDIA/structured-data-models. Install the structured-data-models library (imported as sdm), load a pandas DataFrame into a TableTensor on a CUDA device, treat rows with a known target as context and rows with a missing target as query, and call sdm.models.KumoTabular. Weights download from the Hub on first use. The library also handles preprocessing, ensembling and many-class problems beyond 10 classes via error-correcting output codes. The license is OpenMDW-1.1, stated to permit commercial use. Text, images and timestamps must first be turned into features with the built-in pre-processing recipes.

How solid is it

The results are NVIDIA's own evaluations across four benchmarks, and the post gives its own ELOs and ranks rather than competitor scores. No independent verification is mentioned. The TabArena run used default settings for all three sizes against tuned gradient-boosted trees, AutoGluon and recent tabular foundation models. The post does not say which size produced the 1950 ELO. The speed comparison with LimiX-2 is written as "17 faster" with no unit, so the exact factor is unclear, and no LimiX-2 accuracy comparison is given. On ScoringBench only Large and Medium are first and second, not every size.

Risks and caveats

The model covers numerical and categorical columns only. One forward pass handles up to 10 classes. Accuracy may degrade on tables far beyond the training ranges (the longest training context was 60,000 rows and up to 100 columns) or when query rows come from a different distribution than the context rows. NVIDIA tells users to validate accuracy and calibration on their own held-out data before deployment. Classification and regression are separate models. The training recipe and artificial data generators are not yet released, so the results cannot be reproduced from scratch today. Training compute and cost are not stated.