🔍 Read the full analysis: How To Build The AI Model You Couldn't Find on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A Hugging Face contributor reports using the ML-intern agent to build and publish seven custom models over several days. Examples include a small prompt rewriter and a citrus-image classifier; performance and compute costs are self-reported, and the account is not an independent evaluation of the agent.
A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom models over several days, including a small prompt rewriter and a citrus disease image classifier, as described in the original analysis. The account offers examples of how an agent-assisted workflow can handle training and evaluation, but the reported performance and compute costs have not been independently verified.
The first project addressed a specific gap the contributor said they encountered: a smaller version of the prompt rewriter included with Qwen-Image 2.1. According to the contributor, that model has 9 billion parameters, needs about 20 GB of memory and can generate thousands of tokens before producing a paragraph. They said they found compressed versions but no smaller alternative, so they trained a 0.8-billion-parameter model using the larger model as a teacher. The contributor reports that the smaller model returned valid output 99.7% of the time and used about one-quarter as many tokens as its teacher. They put compute costs, including teacher-generated labels for 8,797 example requests, at about $16.
Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images. The contributor says the training set contained 3,017 annotated images across 21 categories. On a reported test set of 335 photos, the base model identified the correct problem 14.9% of the time, compared with 52.8% for the fine-tuned model after two training epochs on one A10G GPU. The stated compute cost was about $1.90. These figures describe the contributor’s evaluation and are not independently confirmed in the source account.
The other described work included a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the camera project, the contributor says the agent generated 24,722 transparent images of scanned household objects from 24 angles, then used selected objects for training and held others back for testing. Training took about 90 minutes on one A100; the contributor estimated total compute at about $16, including failed jobs that had to be resubmitted. The account says model cards and evaluations were published on Hugging Face, but gives detailed results for only some of the seven projects.
What Lower-Cost Model Building Could Change
The account illustrates a possible way for developers to try task-specific model customization without manually coordinating every training step. The contributor says the agent proposed plans, ran test jobs, requested spending approval and handled training, evaluation and publication. For people with a narrow use case, that could make it easier to test whether a smaller or fine-tuned model can meet a need.
The reported dollar amounts are compute costs for these projects, not a full accounting of the work. They do not include a complete estimate of time spent preparing or checking data, writing prompts or reviewing outputs. Nor do a few examples establish typical costs or results for other users. The citrus comparison is useful as a reported baseline-versus-fine-tuned result on the same test set, but its reach depends on how representative that test set is.
The account also describes a limitation: the contributor says later character-LoRA checkpoints began influencing prompts unrelated to the intended character. That observation suggests customization can produce unwanted effects as well as desired ones. It reinforces why testing on held-out examples and checking outputs beyond the target task matter before a model is relied on.
As an affiliate, we earn on qualifying purchases.
From Chat Prompt to Model Card
According to the contributor, each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and, when paid work was involved, requested a budget. The contributor says it then ran a small test before proceeding with training, evaluation and publication using Hugging Face hardware. When a prompt did not include a budget, the agent offered choices and asked the user to select one.
The contributor says their prompts became more detailed across the projects, growing from about 450 words for the first to nearly 2,000 by the sixth. They included information such as the dataset, base model and training script, as well as requests for a baseline, a small test run and a spending cap. The source says the prompts are available in a public GitHub repository. One instruction asked the agent to report the base model’s zero-shot score on the same metric before training, so results could be compared.
The account is a description of one contributor’s process, rather than an independent study of ML-intern or a guarantee that other users will see similar outcomes. It describes seven completed projects but provides detailed measurements for only some. The project examples therefore show what the author says they achieved, not a general benchmark for agent-assisted model development.
“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”
— The Hugging Face contributor, describing an instruction used in the project prompts
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Results
The figures and descriptions are self-reported. The source account does not provide independent replication or full evaluation protocols for every model, and it does not detail all seven projects. It is also unclear how the contributor measured the prompt rewriter’s 99.7% valid-output rate, whether test images were independently reviewed, or how consistently the agent checked data quality.
The reported compute costs do not establish the full cost of preparing datasets, composing prompts or examining results. Results on the cited test sets may not carry over to new images, prompts or other tasks. The account also does not show whether other users, budgets or hardware would produce comparable performance and costs. Those questions leave the examples informative as a case report, but not proof that the workflow will generalize.
machine learning model fine-tuning tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Published Models Await Wider Testing
The contributor says the models, evaluations and prompts are available through Hugging Face and GitHub. Interested readers can inspect the published materials, while independent tests on other datasets and tasks would help determine how well the reported gains hold up beyond the author’s examples.
The contributor’s stated process for future projects includes establishing a baseline before training, running a small test and setting a spending limit. Further reports could clarify the agent’s reliability, the amount of human review required and the full costs beyond compute. The source does not identify a scheduled independent assessment or a date for additional results.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did the contributor use ML-intern to build?
The contributor says the agent helped build and publish seven custom models over several days. Described examples include a 0.8-billion-parameter prompt rewriter, a citrus problem image classifier and two image-generation LoRAs.
How much did the reported projects cost?
The contributor reported about $1.90 in compute for the citrus classifier and about $16 for the prompt rewriter and the camera-angle LoRA project. These are project-specific compute estimates, not full costs or typical prices.
Are the reported model results independently verified?
No independent verification is described in the source account. The performance figures are reported by the contributor, and the account does not give complete evaluation details for all seven models.
What remains unknown about the results?
The source does not establish whether results generalize to other datasets, prompts or users. It also leaves unclear how some metrics were measured and how much time and review work the projects required beyond compute.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
