3593 words
18 minutes
Your Fine-Tuned Model Is Tech Debt: The Checklist for When It's Still Worth It

Dan Bjornn is a senior data scientist at Lease End, an auto-financing company. He built a fine-tuned intent classifier for Lease End’s text-message app, which connects customers nearing the end of a car lease with the sales team to finance a buyout. A fine-tuned model is retrained on a fixed set of labeled examples, a process called supervised fine-tuning (SFT), so it learns to sort a message into one of a small number of categories. It worked. Bjornn says it helped bring in $12 million in revenue at a 50x return on investment within its first year, by his own account, with no audit or method disclosed [1]. Then a customer answered a first outreach text about their upcoming lease maturity with a simple “good morning” [1]. The system called them back immediately. It had read a pleasantry as a request to talk right now [1]. Fixing that bug meant gathering new examples, labeling them, retraining the model, and checking for new regressions (old bugs coming back). Bjornn says that cycle took about a week each time [1].

Figure 1 - Diagram of a fine-tuned model's value climbing early while its cost to change climbs afterward, with the $12M revenue point and the week-long fix cycle marked as company-stated

Figure 1 - The Win and the Debt Are the Same Event: Lease End’s fine-tuned classifier drove a company-stated $12 million in revenue at a 50x return in its first year. That early win is exactly what made the architecture around it hard to question, and every bug fix afterward cost about a week instead of an edit.


The classifier that paid for itself#

Before the fine-tune, Lease End’s app ran on a workflow built around retrieval-augmented generation, or RAG. RAG means looking up the most relevant stored material and handing it to the model along with the question, instead of relying on the model’s training alone. The app searched a vector database (a database that finds entries by similarity of meaning rather than exact match) of customer messages that had already been classified. It matched each new message against the closest labeled example to guess intent [1]. That handled the obvious cases but missed the nuance in real conversational messages. So the team moved to supervised fine-tuning, training a model to sort every incoming message into one of six intent categories [1].

Bjornn gives four reasons the move looked right. First, accuracy mattered on a decision the whole system hinged on: call now, schedule, or opt out. Second, fine-tuning made smaller, cheaper, faster models workable, and that mattered when replying in real time to thousands of messages a day. Third, six-category classification is a narrow, structured task, the textbook case for supervised fine-tuning. Fourth, the team expected the labeled data itself to be portable, on the theory that the same examples could go into any provider’s fine-tuning pipeline for comparable results [1]. For a while, that bet looked right. The classifier shipped, the app did well, and the revenue number followed [1].

Figure 2 - Diagram of four reasons Lease End chose to fine-tune, labeled accuracy, smaller and cheaper models, a narrow structured task, and expected vendor portability

Figure 2 - Why Fine-Tuning Looked Like the Obvious Move: Four reasons Bjornn names for the original decision. Better accuracy on a decision the system hinged on. Smaller and cheaper models at volume. A narrow six-category task that fits supervised fine-tuning’s sweet spot. An expectation that the labeled data would carry over to any provider.

What broke, and what it cost to fix#

The “good morning” [1] callback was one of two named production failures. Bjornn calls it the “overeager puppy” [1]. He calls the other one the “confused confirmer” [1]. A customer replied “Sounds good” to a scheduled-appointment confirmation, and the bot answered “Great, I’m calling you right now” [1] instead of waiting for the appointment. Every classifier ships with bugs like these. What made these two expensive was the work it took to fix either one.

Bjornn lays out the full cycle in order. Gather examples of the failure. Decide whether enough real data exists, or whether synthetic examples need to be generated and checked by hand. Label everything against the six-category scheme, then check the labels again. Run the fine-tuning job, then test the result against a holdout set (examples set aside and never trained on, used only for testing) to look for new regressions. The fine-tuning run itself was cheap. It took about an hour. Everything around it stretched the cycle to about a week. Bjornn says no fix landed clean on the first pass. Solving one failure mode routinely brought back another, a “whack-a-mole” [1] cycle the team had to walk back before anything could ship.

That cost forced triage. How often does this bug happen? How badly does it hurt the customer experience? Is there a stopgap that avoids a full cycle? [1] Triage kept the team from burning a week on every minor bug. It also meant bugs waited. The team held off retraining until enough of them piled up to justify the week-long cycle, and everything below that bar got a band-aid fix in the meantime, at best [1].

Figure 3 - Diagram of a loop labeled gather, label, fine-tune (about an hour), validate, walk back regressions, taking about a week, with the two named production failures entering the loop

Figure 3 - The Loop Around a One-Hour Training Run: The fine-tuning job itself took about an hour. Gathering data, labeling it, validating the result, and walking back the regressions it introduced stretched the full cycle to about a week, by Bjornn’s own account.

The tax nobody priced in#

Vendor portability never showed up. Fine-tuning data formats, the training volume required, and the training interface all differ between providers, and they even differ between versions from the same provider. Switching mid-flight would have meant redoing a week-long process from a different data shape. So the team stayed on one vendor and one model version, since it couldn’t absorb a migration cost on top of the routine retraining cost [1].

Bjornn calls this the calcification tax. It had a second face. The app was built in late 2024 on the workflow-based architecture that was standard at the time. The team’s engineering attention already went to keeping the fine-tuned pipeline running, which left no room to adopt newer approaches as the field moved on, even once better ones existed [1]. The team paid that tax on every later decision, whether or not they ever tried to migrate.

Figure 4 - Diagram of two locks, one labeled vendor lock-in from incompatible training formats, the other labeled architecture lock-in from engineering time consumed by retraining

Figure 4 - The Calcification Tax, Two Faces: Vendor lock-in came from training formats and volumes that do not transfer cleanly between providers or model versions. Architecture lock-in came from engineering time being consumed by the retraining cycle, leaving none to adopt newer approaches once they existed.

What made the team rethink it#

The push to rethink fine-tuning had nothing to do with the messaging app. After the team started using Claude Code, Anthropic’s coding agent, for coding tasks, Bjornn noticed the model never needed to be swapped when the task changed. What changed was “the skill, the resources that we passed it, the context” [1]. Better context gave better results, with nothing retrained. Bjornn calls it an aha moment, and an awkward one, since he had been the one championing fine-tuning [1]. That moment led the team to ask whether the intent classifier needed to be a fine-tuned model at all.

Moving the fix out of the weights#

The team was already building an agentic framework for other work, so it reused that. It moved the classifier from a fine-tuned model to “a series of skills, and tools, and resources” [1]. A frontier model loads those skills at call time. Frontier models are the largest, most capable general-purpose models from vendors such as OpenAI and Anthropic, called through an API instead of trained in-house. A skill, here, is a folder of instructions and worked examples the agent pulls in when a task needs it. Our skill-evals piece covers that mechanism, and how to prove a skill earns its place before it ships [2]. Deciding what a model sees at call time, and how to layer retrieval, compression, and memory, is its own discipline. We’ve covered it in detail elsewhere [3].

In the new fix cycle, you find a problem and edit the system prompt (the standing instructions sent with every request) or the affected skill. You check the change against a curated evaluation set, a fixed set of real past messages with known correct answers, built from the app’s own production history. You iterate a few times, then deploy by uploading markdown files to a cloud storage bucket, and by Bjornn’s account the whole cycle takes under an hour, against roughly a week for the old one [1]. Per-message API cost went up, because the rebuild called better frontier models instead of a cheap, small fine-tuned one [1]. Total cost still went down, because the team stopped spending a week of engineering time on every fix. Accuracy, the metric the whole system hinged on, was reported higher than it had ever been under fine-tuning, though no specific number was given [1].

Figure 5 - Side-by-side panels comparing the fix cycle before and after the rebuild, about a week versus under an hour, both labeled company-stated, with per-message cost up and total cost down

Figure 5 - What the Rebuild Changed, By Bjornn’s Own Account: The fix cycle fell from about a week to under an hour. Per-message API cost rose, because the rebuild called better frontier models, but total cost fell because the week of engineering time per fix disappeared. All figures are Lease End’s own, unaudited.

Figure 6 - Diagram contrasting decision logic baked into trained weights, changeable only by retraining, against a skill and context read at call time, changeable by editing a file

Figure 6 - Where the Decision Logic Actually Lives: In the fine-tuned version, the classification rule lived inside trained weights, changeable only by retraining. In the rebuild, it lives in a skill file and a system prompt a frontier model reads at call time, changeable by editing a file and redeploying.

A related piece asked a nearby question from the other direction. Once an agent can already reach a domain’s own patterns, does more prompt detail change whether the output is correct? The Prompt Specificity Ladder found that past that point, more detail mostly changed cost and code structure rather than the correctness score. This piece asks a question one step earlier, whether the decision belongs in trained weights at all, before anyone asks how much detail the prompt needs.

KEY INSIGHT: When a decision lives in the context a model reads at call time instead of in its weights, a fix becomes a text edit instead of a data-collection-and-retraining project. At Lease End the training run took an hour. The work around it took a week.

The reasons that didn’t hold up#

Bjornn ends his talk by checking the original four reasons against what actually happened. He splits the second one, smaller models, into its two promises: lower cost and lower latency. That makes five lines to check:

  • Accuracy. Fine-tuning was supposed to give the best accuracy on the decision the system hinged on. The context-engineered rebuild, which guides a frontier model with a system prompt and skills instead of retrained weights, beat the fine-tuned model on accuracy [1].
  • Lower cost at volume. The team expected smaller fine-tuned models to be cheaper at scale. Per-message cost did go up after the rebuild. Bjornn says he had been “looking at the wrong costs” [1]. Total cost, engineering time included, went down.
  • Lower latency. Smaller fine-tuned models were faster, but the gain was “so small… in practice it really didn’t make any difference” [1].
  • A narrow, structured task. Six-way intent classification is the case supervised fine-tuning is supposed to fit best. In Bjornn’s words, “our textbook case still became tech debt” [1].
  • Vendor control. The plan was to keep the labeled data and swap providers if needed, but that failed, because “it’s not as simple as just plugging the data in” [1].

Figure 7 - Checklist of the original reasons for fine-tuning, cost and latency checked separately, each marked as not holding up: accuracy, cost, latency, narrow task, vendor control

Figure 7 - Five Reasons, Checked Against What Happened: Bjornn’s own retrospective against his team’s original four reasons to fine-tune, with cost and latency checked separately. None held up the way the team expected going in.

When fine-tuning is still the right call#

Bjornn doesn’t say fine-tuning is always wrong. His closing rule sets a bar: “fine-tune only when you literally cannot call a frontier model. And even then, your decision still has to beat the tax” [1]. He names two situations where fine-tuning may still be justified. One is privacy and data-control requirements. The other is an offline-only deployment, where no frontier model call is possible at all. Even then, he says to confirm the fine-tune itself won’t cause problems in the long run [1].

Put Bjornn’s rule next to the cases below and we get our own list of four conditions for when fine-tuning still clears the bar:

  1. The decision is bounded and recurring, not an open-ended judgment call.
  2. The workflow already produces its own labels as a byproduct, so no separate labeling effort is needed.
  3. There is truly no way to call a frontier model at all.
  4. Someone owns the retraining pipeline, meaning the labeling loop and the regression checks, for as long as the model stays in production.

NVIDIA and Palantir, whose Foundry platform held the data, give a case that meets the first three conditions. Proprietary supply-chain data had to stay inside a single governed environment. So NVIDIA fine-tuned a 30-billion-parameter model with LoRA, a technique that trains a small set of added parameters instead of the model’s full weights [4]. It trained on allocation decisions (choices about which orders get scarce parts when supply is short) that the workflow already produced as a byproduct [4]. On NVIDIA’s own development benchmark, the tuned model reached 86.7% allocation-decision accuracy, against 55.5% for a much larger general-purpose model and 17.5% for its own untuned base [4]. The training run finished on two GPUs in minutes [4]. NVIDIA’s own caveat travels with the number every time it appears: “This doesn’t mean the smaller model is more capable overall. Its gains are concentrated in the domain it was post-trained on. Future production risk forecasting remained difficult despite fine tuning” [4]. Cheap training didn’t make the debt go away. NVIDIA still had to build a pipeline that turns every planner decision into training data. It also needed a backtest (replaying past decisions to check whether the new model would have done better) that gates each new version before it ships [4]. That’s the fourth condition. NVIDIA pays for it.

Fyxer shows what meeting that fourth condition looks like when it runs continuously, at product scale. Fyxer fine-tunes models to draft executive emails. Each time a user edits a draft, Fyxer turns the edit into a preference pair: the original draft against the version the user actually sent. It trains on those pairs with Direct Preference Optimization (DPO), a method that learns from pairs of a worse and a better output instead of single correct answers. A change ships only after it passes a statistically significant A/B test [5].

Even a fine-tune that clears the bar needs real curation effort. A quick export of logs won’t do. AWS’s guidance puts a typical task at roughly 2,000 high-quality samples, and more for complex multi-step reasoning [6]. AWS also relays research showing that small, clean sets can beat far larger, uncurated ones. LIMA used 1,000 carefully curated examples [7]. AlpaGasus kept only the highest-quality 17% of a larger set [8], which AWS describes as the top fifth [9].

One limit on all this. Everything above is about supervised fine-tuning, training a model to imitate a fixed set of labeled examples. That’s exactly what Lease End’s classifier was. Supervised fine-tuning can make a model lose skills it already had, which researchers call catastrophic forgetting, a close cousin of the whack-a-mole regressions Lease End kept hitting. Cameron Wolfe’s survey of recent research reports that reinforcement learning (RL) is comparatively resistant to that forgetting [10]. RL trains a model on outputs it already produces, instead of pushing it toward someone else’s fixed answers.

Figure 8 - Checklist diagram of four conditions for when fine-tuning still clears the bar: bounded recurring decision, workflow-produced labels, no frontier access, and an owned retraining pipeline

Figure 8 - The Bar, Not a Ban: Our own reading of when fine-tuning clears the bar, built from Bjornn’s rule and the NVIDIA and Fyxer cases above: four conditions. The decision is bounded and recurring. The workflow already produces its own labels. Calling a frontier model genuinely is not possible. Someone owns the retraining pipeline for as long as the model runs.

Figure 9 - Bar chart of NVIDIA's LoRA-tuned model reaching 86.7 percent accuracy against a general model at 55.5 percent and an untuned base at 17.5 percent, NVIDIA development benchmark

Figure 9 - A Case That Clears the Bar: NVIDIA’s LoRA-tuned 30B model reached 86.7% on NVIDIA’s own development benchmark, ahead of a larger general-purpose model at 55.5% and its own untuned base at 17.5%. NVIDIA’s own caveat: the gains are concentrated in the domain it was tuned on.

KEY INSIGHT: Before committing budget to a fine-tune, check all four conditions. Look for a bounded and recurring decision, a workflow that already produces its own labels, no real way to call a frontier model instead, and a named owner for the retraining pipeline once it ships.

Where the checklist runs out#

Every outcome number in this piece comes from one of two places. Either it’s a single speaker’s own, unaudited account of his own company’s system, or it’s a vendor’s own published case study of its own product. The $12 million, the 50x return, and the week-to-under-an-hour fix cycle are Bjornn’s, and he discloses no methodology or baseline anywhere in the talk [1]. He never gives an accuracy percentage, and he doesn’t name the model behind either the original classifier or the rebuilt system [1]. NVIDIA’s numbers come from its own development benchmark, not an audited production deployment [4], and Fyxer documents its mechanism in detail while its outcome numbers come with no method [5].

This piece is about the organizational cost of fine-tuning. A separate line of research raises a problem inside the model itself. A fine-tuned fact can apparently sit in a model’s weights without the model ever using it in multi-hop reasoning (answering a question that needs several facts chained together). The model can recite what it learned without being able to reason with it. That’s a different failure mode with a different cause. Nothing here should be read as evidence for or against it.

None of that makes the mechanism wrong. Fixing a bug in a prompt or a skill file is a smaller, faster, more reversible change than collecting new labeled data and running a retrain. That holds whoever’s numbers you trust. The question to ask is what a fix will cost for as long as the system stays in production. Lease End’s classifier worked. What grew was the price of changing it.

Figure 10 - Decision flow chart asking whether a frontier model with better context can do the task, routing to context engineering by default and fine-tuning only when access is genuinely blocked

Figure 10 - The Decision, In One Flow: Start by asking whether a frontier model with better context can do the task. If yes, that is the default path. If no, fine-tuning is worth considering only when a named owner exists for the retraining pipeline that comes with it.

Conclusion#

Bjornn puts the classifier’s contribution at $12 million [1]. A number that size makes an architecture hard to question. The debt was in what a fix cost once the decision lived in trained weights, instead of in a skill or prompt a frontier model reads at call time.

Before committing budget to a fine-tune, run it against Bjornn’s bar. Can you call a frontier model at all, and if you can, could it do this job with better context instead of retrained weights? If the honest answer to either is no, check one more thing. Does the plan already name who owns the retraining pipeline, the labeling loop, and the regression checks for as long as the model stays in production?


References#

[1] D. Bjornn, “Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards,” AI Engineer (YouTube channel), Aug 20, 2026. https://www.youtube.com/watch?v=4loPnxvWWhg

[2] G. Dotzlaw, “Don’t Ship Skills Without Evals: A Reproducible Way to Prove Your Claude Code Skills Work,” Dotzlaw Consulting, Sep 9, 2026. /insights/claude-code-18-skill-eval-methodology/

[3] G. Dotzlaw, K. Dotzlaw, and R. Dotzlaw, “The Context Engineering Stack: Compression, Retrieval, and Decision Memory,” Dotzlaw Consulting, Aug 12, 2026. /insights/ai-16-context-engineering-stack/

[4] N. Barber, R. Haber, and A. Jhunjhunwala, “From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry,” NVIDIA Technical Blog, Sep 10, 2026. https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/

[5] OpenAI, “How Fyxer built an AI executive assistant people trust,” OpenAI News, Sep 14, 2026. https://openai.com/index/fyxer

[6] AWS, “Preparing data for supervised fine-tuning Part 2: Advanced data strategies,” AWS Machine Learning Blog, Aug 26, 2026. https://aws.amazon.com/blogs/machine-learning/preparing-data-for-supervised-fine-tuning-part-2-advanced-data-strategies/

[7] C. Zhou et al., “LIMA: Less Is More for Alignment,” arXiv:2305.11206, 2023. https://arxiv.org/abs/2305.11206

[8] L. Chen et al., “AlpaGasus: Training A Better Alpaca with Fewer Data,” arXiv:2307.08701, 2023. https://arxiv.org/abs/2307.08701

[9] AWS, “Preparing data for supervised fine-tuning Part 1: Formatting and quality,” AWS Machine Learning Blog, Aug 26, 2026. https://aws.amazon.com/blogs/machine-learning/preparing-data-for-supervised-fine-tuning-part-1-formatting-and-quality/

[10] C. R. Wolfe, “Reinforcement Learning for LLMs: The Complete Guide,” Deep (Learning) Focus, Aug 24, 2026. https://cameronrwolfe.substack.com/p/llm-rl

Your Fine-Tuned Model Is Tech Debt: The Checklist for When It's Still Worth It
https://dotzlaw.com/insights/ai-49-finetuned-model-is-tech-debt/
Author
Gary Dotzlaw
Published at
2026-09-29
License
CC BY-NC-SA 4.0

Building production AI, or modernizing a legacy system?

That is the kind of work we do at Dotzlaw Consulting. Book a free 20-minute intro call and tell us what you are trying to build, or what is slowing you down.

Related reading

The Advisor Tool Is Real, and Anthropic Ranks It Last: The Cost Ladder and the One Number That Decides
Anthropic's advisor tool lets a cheap executor model consult a stronger advisor mid-task inside one API call, with no orchestration code. It is real, it is in beta, and Anthropic's own measured cost-lever guidance puts it dead last. A precise walk of the mechanism, the nine rungs you climb first, the consult rate that decides whether the pairing helps or hurts, and the four places a routing decision can live.
2026-09-03·AI & Modern Development
MCP Tool Design: The Two Ways Your Agent's Tools Fail (Bloat vs. Confusion)
Almost every MCP tool failure traces to one of two root causes, bloat or confusion, and the usual fix for one makes the other worse. A walk through AWS's six tool designs, Smartsheet's production token math, an independent eval where the arm with no tool catalog scored highest on correctness, and a checklist you can run against your own MCP server.
2026-09-02·AI & Modern Development
Prompt Architecture: Layer Your Prompts, Don't Bloat Them
One system prompt cannot be inviolable, situational, expressive, and self-checking at the same time. Split it into four stacked layers, make the last one code instead of text, and you get the only guarantee a prompt was never able to give you.
2026-08-24·AI & Modern Development
Pi + Obsidian CLI: The Agent That Never Forgets Because You Gave It a Place to Remember
Most agent-memory tools sell you a schema. This is the opposite bet: a portable, markdown-native, diffable second brain built from Obsidian, the Obsidian CLI, and Graphify, driven by the Pi coding agent. Boring plain text wins because it stays yours.
2026-07-27·AI & Modern Development
← Back to Insights