Multimodal Flow models text and images in one continuous space, MF-1 released

The paper presents Multimodal Flow, a fully continuous generative model of language and vision. Its starting point is a complaint about how most unified multimodal models are built. They either treat both language and quantized images as discrete tokens, or they pair discrete language prediction with continuous image generation. According to the authors, the first route introduces a visual quantization bottleneck, and the second requires modality-dependent objectives and sampling procedures. Fully continuous modeling, they say, avoids both problems and allows a shared generative process, but it remains underexplored for multimodal pretraining.
Multimodal Flow proposes a unified continuous architecture that combines multimodal continuous representations with a shared chunk-causal flow backbone. Text blocks and images are organized as ordered continuous hyperchunks, which preserves the order of text tokens and the spatial structure of images. The backbone learns a single vector field over these hyperchunks through Flow Matching. Joint attention handles cross-modal interaction, while separate feed-forward networks process each modality. During training the model predicts multiple target chunks in parallel; at inference it generates hyperchunks one after another.
The authors instantiate the design as MF-1 and pretrain it on multimodal data. Across 0.6B, 1.2B and 1.6B scales, they report that continued pretraining consistently improves multimodal modeling. With only 150B pretraining tokens, MF-1 reaches an average score of 82.8 across GenEval and DPG-Bench (image generation benchmarks) and 75.3 across VQAv2, MMBench and POPE (understanding benchmarks). They say this is competitive with unified models trained on substantially more data. Under matched data, optimization and parameter budgets, they also report that Multimodal Flow outperforms representative hybrid and discrete models.
The authors conclude that the results establish continuous chunk-based embedding flow modeling as a new fully continuous paradigm for unified multimodal modeling. Code and model are publicly released at https://github.com/hustvl/Multimodal-Flow.
Key facts
- Multimodal Flow is a fully continuous generative model of language and vision; it avoids the visual quantization bottleneck of discrete-token designs and the modality-dependent objectives of hybrid designs.
- Text blocks and images are arranged as ordered continuous hyperchunks, and one vector field over them is learned through Flow Matching, with joint attention and modality-specific feed-forward networks.
- With only 150B pretraining tokens, MF-1 scores 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench and POPE.
- Under matched data, optimization and parameter budgets, the authors report that it outperforms representative hybrid and discrete models; continued pretraining improves results across 0.6B, 1.2B and 1.6B scales.
- Code and model are publicly released on GitHub at hustvl/Multimodal-Flow.
Why it matters
Unified models that both understand and generate images usually pay a price. Discrete tokenization of images adds a quantization bottleneck, and mixing discrete text with continuous image generation needs separate objectives and sampling procedures for each modality. The authors argue that a fully continuous approach removes those trade-offs and gives one shared generative process, yet say it has been underexplored for multimodal pretraining. This paper is an attempt to fill that gap, and it frames its result as a new paradigm: continuous chunk-based embedding flow modeling.
Who it affects
Mainly researchers and engineers building unified multimodal models that must both read and generate text and images, and anyone comparing discrete, hybrid and continuous designs. The source is a research paper rather than a product announcement, so it speaks to people working on model architecture and pretraining.
How to use it
The authors say the related code and model are publicly released at https://github.com/hustvl/Multimodal-Flow. Anyone who wants to test the approach can start there. The abstract does not describe a license, inference speed, compute cost or hardware requirements.
How solid is it
The claims are the authors' own, drawn from the paper's abstract. The headline figures are specific: 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench and POPE, both with 150B pretraining tokens. The comparison under matched data, optimization and parameter budgets is the more informative claim, since it controls for the usual confounders. Still, the abstract gives no per-benchmark scores, no margin over the hybrid and discrete baselines, and does not say which unified models MF-1 is compared with or how many tokens they were trained on. It also does not say which of the 0.6B, 1.2B or 1.6B scales the headline scores belong to.
Risks and caveats
The statement that MF-1 is competitive with unified models trained on substantially more data cannot be checked from the abstract, because those models are not named. The abstract gives no limitations of the approach. The word paradigm is the authors' framing of their own results; independent replication is the thing to watch before treating continuous embedding flow as a settled alternative to discrete and hybrid designs.
“Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains underexplored for multimodal pretraining.”
— Multimodal Flow paper abstract