Writing Execution Notes That Actually Help You Debug Later
Execution notes are documentation you leave behind after running something that failed, that ran too long, or that produced output you didn't expect. They are not post-mortems written a week later when you remember nothing clearly. They are written at the moment, while the terminal output is still on screen and your working memory hasn't degraded. The core format is simple. You record what you ran, what you expected, what actually happened, and the single most useful observation you can make about the gap between those two things. Everything else is secondary.I write them in plain text files with a predictable naming scheme. Something like exec_20240315_1430.md. The date matters because you'll be searching through months of them later. When a pipeline starts eating memory on a Friday afternoon, you don't want to be guessing which run triggered it.
Notes On An Execution as a Habit, Not a Ceremony
The biggest mistake people make is treating execution notes as optional paperwork. They aren't. They are a cognitive offload system. Your brain is terrible at holding environment details, exit codes, and partial error messages in working memory simultaneously. Writing them down frees up mental space for actually solving the problem. A typical note contains five sections. The command you ran. The environment state — versions, dependencies, OS patches, the weird variable you set three hours ago. What you expected to happen. What actually happened, quoted from the output. Your immediate hypothesis about the gap. This last section is the part most people skip, and it's the most valuable part.The Practical Template
I use a structure that looks like this: Purpose: One sentence. What were you trying to do? Command: The full command with flags, variables, and input paths included. Not a paraphrase. Environment: Key versions and state. Runtime version, dependency lock file status, recent config changes. Expected: What success looks like. A specific output, a line count, a completion message, an exit code. Actual: What happened. Copy-paste the relevant output. Don't summarize. Hypothesis: Your best guess about why the gap exists. Keep it short. You'll refine it or discard it later. Next steps: One or two things to try. If you've already tried them, note that too.This takes about ninety seconds to fill out during a failed run. That's it. You spend less time writing the note than you would spend trying to reconstruct the exact command from memory five minutes later.
Edge Case: The Intermittent Failure
The hardest thing to document isn't a consistent failure. It's the thing that fails one time out of ten, on a machine you don't have consistent access to, with a dataset that changes slightly between runs. I dealt with this with a data processing job last year. It would occasionally hang at the same point — a merge operation on a 40-gigabyte intermediate file. No crash. No error. Just a silent hang until I killed the process. Exit code was always zero, which made monitoring systems treat it as success. My workaround was to add periodic timestamped progress writes to stderr, not stdout. Stderr goes to the log file. Stdout goes to the next pipe stage. By watching stderr timestamps, I could see that the process was doing work, just very slowly and unevenly. The note I wrote captured the exact file state at each timestamp interval, and that data eventually pointed to a bloated index on a specific shard.The takeaway is that execution notes should capture observable state, not just outcomes. A hang is not the same as a crash. Document the difference.
Get the Full Details

What People Get Wrong
Writing too much. Execution notes are reference documents, not essays. If your note is longer than half a page, you're probably including information you already know or recording noise instead of signal. The command, the environment, the actual output, and your hypothesis. That's usually enough. Rushing the hypothesis section. This is where most notes become worthless. People write "unknown" or "need to investigate" and move on. Instead, write your actual best guess, even if it's wrong. A wrong hypothesis is better than no hypothesis because it gives you a concrete thing to test and potentially disprove. Disproving something is how you learn. Not linking to related runs. If you run the same command five times with different parameters and it fails in slightly different ways each time, cross-reference all five notes. The pattern across failures is often more useful than any single failure. Keeping notes in inconsistent locations. This sounds trivial but it's the single biggest factor in whether people actually use the system. Pick a location. A directory under your project root, a specific folder in your notes app, whatever. Put all execution notes there. If you can't find them in twenty seconds, the system failed.When This Approach Breaks Down
Execution notes don't help you when the problem is entirely environmental and transient — a flaky network mount that drops during a specific phase of the run, for example. In those cases, the note will just say "it failed this one time and worked the next" and you'll be right back where you started. They also don't scale well for high-frequency automated pipelines where you run hundreds of executions per day. At that point you need structured logging and automated failure analysis, not manual notes. The note-taking habit is most valuable for the failures that matter — the ones that take hours to diagnose and recur in ways that are hard to reproduce.If you're running hundreds of jobs daily with automated test suites and CI pipelines, stop reading this and look into structured log aggregation instead. This is for the cases where something breaks in a way that automated tests miss and someone needs to sit down and figure out why.
A Quick Worked Example
Let's say you're running a shell script that processes logs and generates a summary report. The script has been working for six months. Today it produces output but the line count is wrong — 847 lines instead of the expected 1,203. Your note: Purpose: Generate monthly summary from access logs. Command: ./generate_summary.sh /data/logs/2024-03 --format=csv --output=/tmp/report.csv Environment: Script v3.2, Python 3.11.6, input directory contains 47 log files totaling 12.3GB. No recent changes to the script or dependencies. Expected: CSV with 1,203 data rows plus header. Actual: CSV with 847 data rows plus header. Script exited 0. No warnings in stderr. Hypothesis: A filter condition in the aggregation step is silently dropping rows. Possibly a date range mismatch — the script expects ISO dates but some log entries use a different format. Next steps: Compare row IDs between expected and actual output. Check date format distribution in input logs.You now have a traceable record and a specific direction to investigate. Six months from now, when this same issue resurfaces with a different data source, you can find this note and remember exactly where to look. That's the entire point.
Notes On An Execution Best Practices
Write the note immediately after the run completes, before you switch to another task. Use the same template every time so your brain stops treating it as a novel problem. Include the full command with arguments — abbreviated commands are useless when you're debugging. Keep the hypothesis section honest and brief. Link to related notes when applicable. Store everything in one consistent location. Delete notes for runs that succeeded without incident after thirty days. The failures are what you need to remember.That's all there is to it. The system only works if you actually use it, and the system only works if you can find what you wrote when you need it. Everything else is decoration.