🔍 Read the full analysis: The Model That Didn't Exist, So You Made It Yourself on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
A Hugging Face contributor reports using the ML-intern agent to build and publish seven custom models over several days. The examples include a small prompt rewriter and a citrus-disease classifier, but performance and cost figures are self-reported and have not been independently verified.
A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom models over several days, including a 0.8-billion-parameter prompt rewriter and a citrus-problem classifier. The account offers a detailed example of an agent coordinating model work, but its performance and compute-cost figures are self-reported, not independently verified results.
The project began with a request for a smaller alternative to the prompt rewriter included with Qwen-Image 2.1. The contributor said that model has 9 billion parameters, needs about 20 GB of memory and can generate thousands of tokens before returning a paragraph. After finding compressed versions of the large model but no smaller alternative, they used it as a teacher to create a 0.8B-parameter model. They reported valid output 99.7% of the time and about one-quarter as many tokens as the teacher; compute, including labeling 8,797 example requests, cost about $16.
Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutrient deficiencies in images. The contributor said the dataset included 3,017 annotated images in 21 categories. On 335 test photos, they reported the base model correctly identified the problem in 14.9% of cases, compared with 52.8% after fine-tuning for two epochs on one A10G GPU. The reported compute cost was about $1.90. The account does not provide an independent assessment of the test set or results.
The contributor also described a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent reportedly generated 24,722 transparent images of household objects across 24 angles, then trained on selected objects while holding others back for testing. Training took about 90 minutes on one A100; the contributor put total compute at about $16, including failed jobs. Model cards and evaluations were published on Hugging Face, according to the account.
The described workflow began with a message in HuggingChat with ML-intern enabled. The agent proposed a plan, requested a spending limit before paid work, ran a small test and then handled training, evaluation and publication using Hugging Face hardware. The contributor said their prompts became more detailed over time, specifying datasets, base models, training scripts, baseline tests and spending caps. They said all seven prompts are available in a public GitHub repository.
Lowering the Barrier to Custom Models
The account matters because it shows how a developer might use an AI agent to coordinate tasks that otherwise require time and technical work: preparing or selecting data, running a baseline, testing a training job, evaluating results and publishing a model. The reported compute costs—about $1.90 to $16 for the projects described—suggest that some small, specialized experiments may be attempted without a large hardware budget. They are not a full estimate of the work’s cost, however, and do not include a complete accounting of data preparation, prompt writing or human review.
The examples also underline why comparison with a baseline is useful. The citrus result is presented against the untuned model on the same test set, giving readers a reported point of comparison rather than a score on its own. At the same time, the contributor said later character-LoRA checkpoints began affecting prompts unrelated to the character. That account illustrates a practical risk: fine-tuning can spread a learned style or behavior beyond the intended task. The results do not establish that the same workflow or gains will carry over to other users or applications.
As an affiliate, we earn on qualifying purchases.
From a Smaller Rewriter to LoRAs
The first project addressed a specific gap the contributor said they encountered on Hugging Face: compressed versions of the Qwen-Image 2.1 prompt rewriter were available, but they did not find a smaller model for the task. The reported response was to use the larger model as a teacher and train a compact alternative. This is one approach to customizing an existing model; it does not amount to a claim that the smaller system matches the teacher across all prompts or uses.
The later examples covered different forms of customization. The citrus project fine-tuned a base model for image classification, while the character and camera-angle projects used LoRA adapters, a method for adapting a model without retraining all of its parameters. The contributor said their prompts grew from about 450 words for the first project to nearly 2,000 by the sixth, with explicit requests for a baseline, a smoke test and a spending cap. The source says the account describes one contributor’s method, not an independent evaluation of ML-intern.
machine learning model fine-tuning kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Reliable Are the Reported Results?
The performance and cost figures come from the contributor’s account. The source does not provide independent replication, a full evaluation protocol for every model, or comparable results from other users. It also gives details for only some of the seven projects, so readers cannot assess the complete set from the supplied descriptions.
Several evaluation questions remain open. The account does not explain how the 99.7% valid-output rate was measured, whether test images were independently reviewed, or how performance would change on data beyond the reported test sets. It also does not give a full accounting of human time or other possible expenses. The contributor’s claims should be read as results from these particular projects, not as a general performance or cost guarantee for the agent.
As an affiliate, we earn on qualifying purchases.
Published Models Invite Further Checks
The contributor says the seven models, evaluations and prompts are available through Hugging Face and GitHub, where readers can inspect the published material. The source does not identify a scheduled follow-up or independent review. The next useful evidence would be repeated tests by other users, with clear evaluation methods and results across additional datasets and tasks.
For future projects, the contributor’s stated process calls for establishing a baseline before training, running a small smoke test and setting a spending limit. Whether those steps are adequate will depend on the model, data and intended use. Until outside checks are available, the reported examples show what one contributor says they achieved, while leaving the workflow’s broader reliability and typical costs unsettled.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is ML-intern in this account?
It is the agent the contributor says they used through HuggingChat to plan model projects, request spending approval, run tests, coordinate training and publish evaluations.
Which results did the contributor report?
Among the examples, the contributor reported that a 0.8B prompt rewriter produced valid output 99.7% of the time. For a citrus image classifier, they reported 52.8% correct identification on 335 test photos after fine-tuning, compared with 14.9% for the base model.
How much did the projects cost?
The contributor reported compute costs of about $1.90 for the citrus classifier and about $16 each for the prompt rewriter and camera-angle LoRA project. These are self-reported compute figures, not complete estimates of labor and all other expenses.
Have the results been independently verified?
The supplied account does not report independent replication. It also leaves details such as some evaluation procedures and how results generalize beyond the stated test data unclear.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
