Hugging Face ML Intern agent builds six small models from prompts, for USD 1.90 to USD 37 each

Hugging Face ML Intern agent builds six small models from prompts, for USD 1.90 to USD 37 each

A first-person post on the Hugging Face blog describes how its author built six small models in a few days by writing prompts for ML Intern, an agent available in HuggingChat. The agent plans the work, asks for a budget before it spends anything, runs a small test before the real job, then trains, evaluates and publishes on Hugging Face hardware. Each model started as a message in HuggingChat with ML-intern switched on and ended as a public model on the Hub, with its evaluation in the model card.

The trigger was a prompt rewriter. The official one that ships with Qwen-Image 2.1 is a 9B model that needs about 20 GB of memory and thinks for thousands of tokens before writing a paragraph. The author found only compressed copies of that same 9B model on the Hub, so described the small version wanted, and the next day had a 0.8B version that runs on a CPU. It returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. The whole project, including having the 9B model label 8,797 example requests, cost USD 16.

The six builds. (1) A citrus-disease model: using Claude, the author assembled citrus-disease-vlm-instruct from three Project-AgML sources on the Hub, 3,017 annotated images across 21 pests, illnesses, nutritional gaps and treatment approaches. ML-intern fine-tuned Qwen3.5-2B after benchmarking the base model first. On 335 test photos the base model named the right problem 14.9% of the time; after two epochs on one A10G the fine-tuned model reached 52.8%. Compute: about USD 1.90.

(2) A character LoRA: Huggy, drawn in the flat style of Hugging Face brand assets, on FLUX.2 klein base 4B, trained on 84 captioned drawings. The agent saved a checkpoint every 100 steps and drew the same prompts with each. Step 200 was the first where Huggy was fully on-model; from step 500 on, Huggy's style bled into unrelated prompts. The LoRA also works on the distilled klein model at 4 steps. Compute: about USD 7.60.

(3) Two new-trick LoRAs for Qwen-Image 2.1. The camera-angle LoRA (view an object from, say, 45 degrees left) had no existing version when the author checked a few days after the model's release. ML-intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each, giving 24,722 transparent images, on a CPU job costing a few cents. It finalised 461 training objects, 40 held out, and 1,844 before-and-after pairs spread evenly over 23 camera instructions. Training took 2,000 steps in about 90 minutes on one A100 (about USD 3.75). The project took about half a day and 48 jobs, counting failures from missing packages or wrong paths that the agent resubmitted, and cost about USD 16 in total.

The Doodle-in LoRA turns a magenta scribble on a photo, plus a short prompt naming an object, into that object with lighting and composition kept consistent. No dataset existed, so the prompt described how to make one: take an Open Images photo, remove an object with the LaMa inpainting model, draw a scribble where it was, and use the untouched photo as the target. ML-intern wrote and tested the scripts in a CPU sandbox, ran them as GPU jobs, and recorded the author and licence of each source photo. It built 6,042 training pairs and a 160-pair test set, 40 of whose pairs come from 23 object classes kept out of training. It measured the base model alone and with the Viggle turbo LoRA before training, then trained 2,000 steps in 1 hour 38 minutes on one A100 (about USD 4); a comparison on 48 test pairs picked step 500. Paired with the Viggle turbo LoRA at 6 steps, 67.5% of objects were detected where they were drawn, at 4.7 seconds per edit, and unseen classes did about as well as the rest (65.0% versus 64.2%). The project took a little over a day and 59 jobs, about USD 24 in total.

(4) Two models for small devices. The Pocket Rewriter: ML-intern generated 8,797 short image requests with a small instruct model through Inference Providers, following a mix set in the prompt (photos, posters, logos, infographics, about a third asking for exact quoted text, many in languages other than English). The 9B teacher rewrote them all on one A100 in 2 hours 37 minutes (about USD 6.50). After quality filtering, 1,840 examples went into training. The 0.8B and 2B students trained in 12 and 18 minutes on an A10G (USD 0.75 for both). The 0.8B also ships as an 812 MB GGUF file for CPUs. The project took about 11 hours and 24 jobs, about USD 16 in total.

Agate-Preview-002-4step: Logolabs' Agate Preview 002 is a 260M-parameter text-to-image model, small enough for a browser, but it needs 50 steps with guidance, which is 100 network passes per image. The author asked ML-intern to distill it to 4. The first run cached 155,000 training images as latents, baked the guidance into the model, then cut steps from 16 to 8 to 4 on A100s. The 4-step student beat the teacher run at the same 4 steps on GenEval and FID, and ML-intern exported it to ONNX for the browser; this run took about 13 hours and USD 22. A second run made 24,000 more image pairs with the teacher at 16 steps and fine-tuned the student for about an hour. GenEval rose from 0.509 to 0.536, against the teacher's 0.563 at 50 steps. The second run took about 8 hours; total across both runs, about USD 37.

The prompting method. The author spends effort on the first message: about 450 words for the citrus model, closer to 2,000 by the 6th project. A prompt opens with the idea and why, then names the dataset, base model and training script. Anything already checked goes under a heading reading "Verified facts, do not re-derive", so the agent does not spend budget rediscovering it. Two lines are called critical: a baseline before any training, and a smoke test with a check attached (for the image LoRAs, 50 training steps, then a check that the saved weights had changed, before paying for the full run). The prompt ends with the deliverables, what belongs in the model card, and a spending cap. ML-intern starts each task with a zero dollar budget and needs permission before paid jobs, so the cap holds; with no budget given, it suggests a couple of paths by project size and asks which to take. The author says the first attempt needs none of this: the citrus brief lacked the verified-facts section and still produced a model that more than tripled the base model's accuracy. All seven prompts are on GitHub at yvrjsharma/ml-intern-prompts.

To try it, switch on ML-intern mode in HuggingChat and paste a prompt. The author advises a small budget, a baseline and a smoke test, and reading the results before raising the cap.

Key facts

  • The author built six small models in a few days by prompting ML Intern in HuggingChat; each ended as a public Hub model with its evaluation in the model card, and compute per project ran from about USD 1.90 to about USD 37.
  • The Pocket Rewriter is a 0.8B student of the official 9B Qwen-Image 2.1 prompt rewriter. It runs on a CPU as an 812 MB GGUF file, returns valid output 99.7% of the time, uses about a quarter of the teacher's tokens, and cost USD 16.
  • The citrus-disease model went from 14.9% to 52.8% correct on 335 test photos after fine-tuning Qwen3.5-2B for two epochs on one A10G, at about USD 1.90.
  • The 4-step Agate student improved from 0.509 to 0.536 on GenEval after a second run, against the 50-step teacher's 0.563; total compute across both runs was about USD 37.
  • The author's method: a long first prompt, a "Verified facts, do not re-derive" section, a baseline before training, a smoke test, and a spending cap; ML-intern starts with a zero dollar budget and needs permission for paid jobs.

Why it matters

The post is a worked example of an agent doing the whole small-model pipeline: building a dataset, training, evaluating against a baseline and publishing. Its headline case is a 0.8B rewriter replacing a 9B model that needs about 20 GB of memory, for USD 16 of compute and about 11 hours. The Hub, the author says, offered only compressed copies of the big model, which is the gap the agent was asked to fill.

Who it affects

People who want a small, cheap model that does not exist yet: a domain classifier, a character or style LoRA, a new LoRA trick for a freshly released model, or a model that fits a CPU or a browser. The post is aimed at HuggingChat users who can write a detailed brief and have a dataset, or can describe how to make one. It also shows the Hub's role as the place where results are published.

How to use it

Switch on ML-intern mode in HuggingChat and paste a prompt. Start with a model you wish existed and a dataset you have, or one you can describe. State the idea and why, name the dataset, base model and training script, list checked facts under "Verified facts, do not re-derive", ask for a baseline and a smoke test, define the model card contents, and set a cap such as "Cap total spend at USD 12 and ask me before exceeding it." The agent asks for a budget before spending and suggests paths if none is given. The author's seven prompts are free to copy on GitHub at yvrjsharma/ml-intern-prompts. The text does not say whether ML-intern mode in HuggingChat is free or has its own pricing or limits beyond the GPU and CPU job charges.

How solid is it

This is one author's first-person account. Each model is said to have a public model card with its evaluation, and the post gives specific numbers, durations and job counts. The text does not say how the 99.7% valid-output rate or the quarter-of-the-tokens figure was measured, or on which test set, and no independent verification or third-party evaluation of any model is mentioned. The Huggy and camera-angle LoRAs have no numeric results in the text. The counts are loose too: the post says "five more models" after the first and refers to both "6 things" and "seven prompts", and the Agate project had two runs. The "What it cost" section appears only as a one-line caption, without the table.

Risks and caveats

Quality is modest in places: the fine-tuned citrus model names the right problem on 52.8% of test photos, so it is still wrong on about half. The Doodle-in LoRA detected 67.5% of objects where they were drawn. The Agate student at 4 steps still trails the 50-step teacher (0.536 against 0.563 on GenEval). Runs were not tidy: the camera-angle project took 48 jobs, some failing on missing packages or wrong paths. Costs are GPU and CPU job charges as reported for each session. The author notes the checkpoint choice matters: the Huggy LoRA's style bled into unrelated prompts from step 500 on, while step 200 was the first checkpoint where Huggy was fully on-model.

“Cap total spend at USD 12 and ask me before exceeding it.”

— The post's author, an example instruction from the prompts