What actually happens when you train a model
Most people think training is just feeding data into something and hitting run. It isn't. The difference between a model that works in production and one that breaks on a Tuesday morning has almost nothing to do with the architecture. It has to do with how you structure the training loop, what you monitor, and where you stop.I spent about four years working on language model fine-tuning. The work was never glamorous. You sit in front of a terminal, watch loss curves, and figure out why your validation loss started climbing while your training loss kept dropping. That moment is called overfitting, but the word doesn't tell you what it feels like. It feels like you built a system that memorized the answer key instead of learning the material. Your model will pass every internal benchmark and then fail completely when someone asks it a question in slightly different wording than what it saw during training. The first thing most people get wrong is the data. They collect a bunch of examples and start training immediately. That's backwards. I spend about 40 to 60 percent of my total time on data curation before I even launch a single training run. The reason is simple: garbage quality in, garbage quality out, and you can't optimize your way out of bad data. In the dataset I described above, I found that roughly 15 percent of the responses were either incomplete, contained internal jargon that outsiders wouldn't understand, or were direct copies of documentation that didn't match the question being asked. I spent a full week cleaning, rewriting, and reformatting those entries. The model that came out of the cleaned dataset was noticeably better than what I got from the raw data, and it required fewer training epochs to reach the same performance level. The training duration is measured in epochs, not hours. One epoch means the model has seen every example in your dataset exactly once. For the technical documentation dataset I mentioned, I found that 2 to 3 epochs was the sweet spot. Going beyond that, say 5 epochs, would start producing noticeable degradation in the model's general language ability. It would become more accurate on the domain-specific questions but worse at answering everything else. This is the classic catastrophic forgetting problem. The model literally unlearns general capabilities to specialize in your data.
When I was working on that technical documentation project, I hit a situation where the validation loss spiked unpredictably. It wasn't a smooth climb. It would drop, spike up, drop again, then spike higher. After about two days of investigation, I traced it back to data ordering. The dataset had been sorted by topic, which meant the model would see 200 examples about the billing module in a row, then 200 examples about the API reference, then 200 about authentication. When the model finished cycling through billing, it had forgotten everything about the previous topic. The fix was shuffling the data properly before each epoch and adding a stratified sampler so that each batch contained examples from all topics roughly in proportion to their representation in the full dataset. Once I did that, the validation loss curve became smooth and predictable. The spike pattern disappeared entirely. The learning rate scheduler matters more than people give it credit for. A linear decay schedule works fine for small adjustments. But when you're doing more substantial fine-tuning, cosine decay gives you a gentler ramp down at the end, which helps the model settle into a good configuration rather than continuing to make aggressive weight updates. I switched from linear to cosine on a project once and saw the validation loss improve by about 8 percent over the same number of steps. The change took me five minutes to implement. Another thing that catches people off guard is the relationship between dataset size and model size. A small model with a lot of data generally outperforms a large model with very little data. I've seen this happen repeatedly. If you only have 2,000 training examples, fine-tuning a 70B parameter model will almost certainly overfit. A 3B parameter model trained on the same 2,000 examples will likely generalize better. The larger model has more capacity to memorize, and with insufficient data, that becomes a liability rather than an asset.
What this approach doesn't solve
This method works well for adapting a model to a specific domain or style. It does not fix fundamental problems with the base model. If the base model has poor reasoning capabilities, fine-tuning won't add them. If your data contains factual errors, the model will learn and reproduce those errors with high confidence. You can't fine-tune truthfulness into a model. Also, SFT doesn't help with hallucination reduction in any meaningful way. You'll get better format compliance and tone matching, but the model will still generate plausible-sounding but incorrect information if it encounters a question outside its training distribution. For that, you need retrieval-augmented generation or a separate evaluation and filtering pipeline, not more fine-tuning.The whole process - data cleaning, configuration, training, and basic evaluation - took me about three days for the technical documentation project. Training itself ran for roughly 14 hours on the 8x A100 cluster. The remaining time was spent on data work and monitoring. If your dataset is smaller and cleaner, you can cut that down to a day or so. If it's messier, plan on a week. There's no shortcut around the data quality work.
Get the Full Details
