Framework lets frozen medical LLMs and VLMs learn from deployment, with gains up to 34.2%

A paper on Hugging Face Papers (2610.09146) starts from a familiar problem: large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. The authors say this is especially troubling in medicine, where new clinical evidence, updated guidelines and new therapies can change established practice.
They describe the two usual fixes and their weaknesses. Fine-tuning can update the model, but it requires access to model weights and additional training. Parameter-free methods avoid training, but they may overfit a fixed validation set, lack reliable domain knowledge, or lose visual details by saving experience only as text.
Their answer is a model-agnostic framework that lets a frozen LLM or VLM learn from deployment experience through three forms of external expertise. The first is a Skill that guides reasoning and tool use. The second is a Knowledge Memory that stores reliable facts supported by earlier cases or trusted external evidence. The third is a Multimodal Knowledge Base that keeps visual examples and guides the model to relate each retrieved case to the current image.
Instead of relying on a fixed validation set, a validation strategy keeps an update only if it helps on new cases without degrading performance on earlier ones.
The evaluation spans six benchmarks covering clinical diagnosis, clinical workflows, medical reasoning, and medical and non-medical visual reasoning, with four open-weight and closed-source base models. The authors report that the framework improves performance during online deployment by up to 34.2% over the base model on medical tasks. They also say it generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.
Key facts
- The framework lets frozen LLMs and VLMs learn from deployment experience without changing model weights or running additional training.
- It uses three forms of external expertise: a Skill for reasoning and tool use, a Knowledge Memory of reliable facts, and a Multimodal Knowledge Base of visual examples.
- A validation strategy keeps an update only if it helps on new cases without degrading performance on earlier ones, replacing a fixed validation set.
- Tested on six benchmarks and four open-weight and closed-source base models, it improves online performance by up to 34.2% over the base model on medical tasks.
- The authors say it generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.
Why it matters
Models deployed in medicine are normally frozen, yet clinical evidence, guidelines and therapies keep changing. The authors argue that neither fine-tuning (which needs model weights and more training) nor existing parameter-free methods (which can overfit a fixed validation set, lack reliable domain knowledge, or drop visual detail by storing experience only as text) solves this well. This framework targets that gap by keeping the model frozen and putting what it learns into external components, including one that stores visual examples.
Who it affects
The work is aimed at people building and running LLM and VLM systems for medical tasks such as clinical diagnosis, clinical workflows and medical reasoning, especially where model weights are not accessible, as with closed-source models. The authors also report that the approach works in non-medical domains and transfers to other models, so it is not limited to one model family.
How to use it
The paper describes a design rather than a product. A deployed model is paired with three external pieces: a Skill that guides reasoning and tool use, a Knowledge Memory that stores facts supported by earlier cases or trusted external evidence, and a Multimodal Knowledge Base that keeps visual examples and prompts the model to relate each retrieved case to the current image. Updates to these pieces are accepted only if they help on new cases without hurting performance on earlier ones. The source does not mention a code, dataset or model release.
How solid is it
The claims come from the authors' own abstract. They report results across six benchmarks and four base models, and the headline figure is a maximum: up to 34.2% over the base model on medical tasks. The source does not say whether that figure is in percentage points or a relative gain, nor which benchmark or base model it refers to. No per-benchmark results and no typical (average) improvement are stated, and the source does not name the specific benchmarks or base models.
Risks and caveats
The 34.2% is a best case, not an average, and without the benchmark and model behind it the size of the typical gain is unknown. The source describes benchmark evaluation and does not mention clinical deployment, regulatory approval or testing on real patients. Claims of generalization to unseen cases, transfer to other models and use in non-medical domains are the authors' own and are not backed here by any detail.
“Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve.”
— From the paper's abstract