Why Keeping Notes on ML Experiments Actually Matters

You train a model. It works. Then you come back two weeks later and have no idea which learning rate, optimizer tweak, or data preprocessing step made the difference. This happens constantly. I've lost count of how many times I've stared at a results spreadsheet and thought "what the hell was I doing here?" A journal in ML isn't some reflective writing exercise. It's a structured record of what you tried, what happened, and why you think it happened. The format matters less than the consistency. I started with simple text files organized by date, then moved to CSVs with fixed columns, and now I use a hybrid approach that combines structured metadata with free-text notes. Here's what I track for every experiment: timestamp, dataset version, model architecture (or modification to existing architecture), hyperparameters (learning rate, batch size, epochs, weight decay, momentum, optimizer type), training time, validation metrics (at least loss and accuracy, sometimes F1 or other domain-specific measures), and a brief note on unexpected observations. That's it. Nothing fancy.

I also log the seed value. You'd be surprised how many times two runs with identical hyperparameters produce different results due to randomness, and without recording seeds you can't reproduce them.

The System That Actually Sticks

Most people abandon journaling because they set up something too elaborate. I once tried maintaining a full database with relational tables for experiments, runs, parameters, and results. It took three hours to set up and two minutes to abandon. The system you use needs to require less effort than the act of not using it. My current setup is a folder structure like this: each experiment gets its own directory named with a date and short tag (2024-03-15-grad-manipulation). Inside that directory there's a config file in YAML format that captures all hyperparameters, a results.txt with raw output logs, and a notes.md for observations. I also run a one-line command at the start of training that dumps the git commit hash, Python version, PyTorch version, and CUDA version into a metadata.txt file. This takes about ten seconds and has saved me more times than I can count. The trick is making it automated where possible. I wrote a small Python wrapper around my training loop that automatically records the hyperparameter dictionary and appends to a master CSV. Everything that doesn't require the wrapper goes into the notes.md. No pressure to be perfect.

Get the Full Details

Lecture Slides for INTRODUCTION TO Machine Learning 2
Lecture Slides for INTRODUCTION TO Machine Learning 2

What Beginners Miss

One thing nobody tells you about experiment journals: the most valuable entries are the ones that document failure. A model that didn't converge, a learning rate that caused NaN values, a data augmentation pipeline that introduced bias — these are actually more useful than successful runs when you're debugging a similar problem six months later. I keep a separate section in each experiment directory called failures.txt where I write a few sentences on why something went wrong. This has probably helped me more than the success logs. Another common mistake is being too vague in your notes. "Tried a lower learning rate" tells you nothing two weeks from now. "Reduced learning rate from 0.001 to 0.0001 because training loss plateaued at epoch 12 with training loss 0.034 and validation loss 0.089" is worth gold. Also, don't forget to note what you didn't try. If you're running a grid search over batch size and learning rate but you skip momentum because you assumed it wouldn't matter for Adam, write that down. Assumptions you made and reasons you excluded certain variables become invisible to your future self very quickly.

A Specific Problem I Ran Into

Last year I was working on a segmentation model where I kept seeing mysterious performance drops after switching between two different GPU setups. The same code, the same hyperparameters, the same data. One machine gave me consistent results and the other gave me results that degraded randomly between runs. I spent about four hours digging through my journals before I noticed something: the config files had been saved with slightly different random seeds, and between the two machines the default seed generation was producing very different values. On one machine I was inadvertently comparing experiments that had different initialization conditions rather than just different hardware behavior. The fix was straightforward once I spotted it. I added an explicit seed logging step to my wrapper script and started requiring all future experiments to document the exact seed value. But it cost me two days of confused debugging. This is exactly the kind of thing that a proper journal prevents — not just by recording data, but by forcing you to notice patterns in your own documentation.

When Journaling Doesn't Help

I should be honest about where this approach breaks down. If you're running large-scale distributed training across multiple nodes, a simple folder-based system becomes unwieldy quickly. I've seen teams at that scale move to experiment tracking platforms like Weights & Biases or MLflow, which solve the coordination problem but introduce their own maintenance overhead. The journaling principle is the same, but the tooling shifts. Another limitation: if you change your experimental methodology mid-project without updating your journal format, your records become incomparable. I once had a month of training logs where I switched from logging validation loss to logging validation accuracy halfway through without realizing it. Cleaning that up took longer than keeping the journal in the first place. The best alternative if you find manual journaling unsustainable is to build automation into your training pipeline from day one. Even a simple decorator function that wraps your training call and outputs everything to a structured format is better than nothing. The goal isn't perfect documentation, it's useful documentation. There's a difference.

Journal of machine learning research
Journal of machine learning research

The Minimum Viable Journal

If all of the above feels like too much to start with, here's the bare minimum that will still save you time: create a CSV with columns for date, experiment name, key hyperparameters, validation metric, and a notes column. One row per experiment. Fill it out while the results are fresh in your memory, ideally within five minutes of the training run completing. Spend roughly as long on the journal entry as you would on a casual slack message. This usually cuts your debugging time down from several hours per confusing result to maybe twenty minutes. The habit itself matters more than the tool. I know people who use notebooks, spreadsheets, markdown files, databases, and nothing at all. The ones who consistently benefit from their journals are the ones who keep doing it even when they're busy. Because the alternative is always the same: spending three hours figuring out what you already documented in fifteen minutes.