Latent protein languages PLL and SLL help transformers scale in protein generation

The paper starts from a problem it states plainly: autoregressive transformers remain comparatively weak for protein sequence and structure generation. Its focus is the target representation, meaning what the model is asked to predict. Amino acid tokens encode residue identities without explicit contextual semantics, and backbone coordinates need some discrete representation before an autoregressive model can handle them.
The authors introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet, with one token per residue, built on a frozen ESM-2 encoder. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision, while keeping the ability to decode back to backbone coordinates. They then pretrain autoregressive transformers separately on PLL and SLL tokens with next-token prediction, which gives two models: PLLM and SLLM.
The reported results are as follows. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038, against 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM cuts the fraction of samples falling below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model, across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training.
On speed, the authors say that for long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in their own measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. The authors also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample.
The authors conclude that learned latent protein languages are a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.
Key facts
- PLL maps sequences to a 4,096-state contextual alphabet, one token per residue, built on a frozen ESM-2 encoder; SLL adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision and still decodes to backbone coordinates.
- Under matched downstream sequence training, PLLM's fitted compute-scaling exponent is 0.038 versus 0.020 for the amino acid autoregressive model.
- PLLM reduces the fraction of unconditionally generated samples below a heuristic 1.5-bit entropy threshold by 54% relative to the amino acid model; SLL cuts best validation perplexity for sequence-to-structure prediction by 34% against the original GCP-VQVAE Lite tokenizer.
- For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in the authors' own measurements.
- SLLM compares favorably with other generative models on diversity and novelty in backbone generation; using its token confidence for inference-time sampling shows only early signs of helping.
Why it matters
The paper targets a specific weakness: autoregressive transformers, the architecture behind most language models, remain comparatively weak at generating protein sequences and structures. Its argument is that the choice of target representation matters. Instead of predicting raw amino acid identities, the model predicts tokens from a learned latent alphabet that carries context. The reported scaling exponent of 0.038 against 0.020, about 1.9 times higher, suggests such models may gain more from added compute. The claimed speed advantage over MSA-based AlphaFold2 for long proteins, about 1,000 times in the authors' measurements, is the other headline number.
Who it affects
Mainly researchers working on protein generative models, especially those using autoregressive transformers or discrete structure tokenizers such as GCP-VQVAE Lite. People who build on ESM-2 representations may also find the PLL design relevant, since it sits on a frozen ESM-2 encoder. The source names no products or users beyond these research settings.
How to use it
This is a research paper, not a product. Practically, it offers a recipe: encode sequences with a frozen ESM-2 encoder into a 4,096-state alphabet (PLL), or retrain a GCP-VQVAE Lite style tokenizer with auxiliary sequence and confidence supervision (SLL), then pretrain an autoregressive transformer with next-token prediction. No code or weights release is mentioned.
How solid is it
The numbers come from the authors' own experiments, as summarised in the paper's abstract. The 1,000 times speedup is stated only for long proteins, only against MSA-based AlphaFold2, and in the authors' own measurements; no hardware or protein lengths are given. The 54% and 34% figures are relative reductions, and the baseline absolute values are not given. Backbone generation results are described only as comparing favorably on diversity and novelty, with no figures and no named competitors. The source gives no model sizes, datasets or training compute.
Risks and caveats
The source does not say that structure prediction quality matches AlphaFold2, only that sampling is faster. The 1.5-bit entropy threshold is described as heuristic, and the source does not say it measures protein quality or foldability directly. The token-confidence sampling result is labelled early signs, with no numbers. The scaling-exponent comparison is a fitted value under matched downstream sequence training, so it should not be read as a guarantee at larger scales.
“Autoregressive transformers remain comparatively weak for protein sequence and structure generation.”
— Opening line of the paper's abstract