Activation alignment lets tabular models TabPFN-3 and TabFM close much of the gap with less context

Activation alignment lets tabular models TabPFN-3 and TabFM close much of the gap with less context

Tabular foundation models learn in context: they condition their predictions on labeled training examples supplied as context. Unlike traditional models, which separate training from inference, they must process all of those training examples in every forward pass, so each prediction is expensive. Restricting the number of examples cuts the cost, but the authors say it substantially degrades performance.

The paper proposes a third route called activation alignment. Instead of discarding context, it uses the full context to teach a model how to behave when it sees only a subset. The mechanism is a lightweight linear transformation, trained on synthetic unlabeled data. It maps the intermediate activations of a data-constrained "student" (which sees partial context) toward those of a full-context "teacher" (which sees all the data).

The authors say training the aligner needs no GPU and converges in seconds to minutes on commodity hardware.

The evaluation uses 38 classification datasets from the TabArena benchmark and the two leading tabular foundation models, TabPFN-3 and TabFM. Across all context budgets, the aligned student gives broad, statistically significant improvements over the unaligned baseline for both models. In low-data regimes, alignment recovers nearly half of the teacher's predictive advantage. The authors present the method as a practical, low-overhead way to get the inference speed of compact contexts while closing a significant fraction, not all, of the performance gap to the full-context teacher.

Key facts

  • Activation alignment trains a lightweight linear transformation on synthetic unlabeled data to map a partial-context student's intermediate activations toward those of a full-context teacher.
  • Tabular foundation models must process all training examples in every forward pass; cutting the context lowers cost but substantially degrades performance.
  • Evaluated on 38 TabArena classification datasets with TabPFN-3 and TabFM, the aligned student improves significantly over the unaligned baseline at all context budgets.
  • In low-data regimes, alignment recovers nearly half of the teacher's predictive advantage.
  • Training the aligner needs no GPU and converges in seconds to minutes on commodity hardware.

Why it matters

In-context tabular models pay for their accuracy at prediction time, because every forward pass reprocesses the whole training set. Trimming the context is the obvious fix, but it hurts accuracy. This work tries to keep the speed of a short context while recovering part of what is lost, and it does so with a small linear map rather than retraining the model.

Who it affects

Teams that use tabular foundation models such as TabPFN-3 or TabFM, and researchers working on efficient in-context learning for tables. The reported gains are for classification datasets from TabArena.

How to use it

The paper describes the recipe: train a linear aligner on synthetic unlabeled data so that a student with partial context mimics the intermediate activations of a full-context teacher, then run the student at inference. The authors say this needs no GPU and takes seconds to minutes on commodity hardware.

How solid is it

The authors report broad, statistically significant gains over the unaligned baseline for both models across all context budgets, tested on 38 TabArena classification datasets. The headline figure, nearly half of the teacher's advantage recovered in low-data regimes, is a share of the teacher-student gap, not an accuracy number. These are the authors' own results as stated in the abstract.

Risks and caveats

Alignment narrows the gap to the full-context teacher but does not close it; the authors claim a significant fraction. The evaluation covers classification only and two models, TabPFN-3 and TabFM. The source gives no absolute metric values, no speedup factor and no definition of "low-data regimes", so the practical size of the gains is hard to judge from the summary alone.