Working with Trace Worksheets in Model Evaluation

The thing about trace worksheets is that most people treat them like checklists when they are actually diagnostic instruments. You pull one out, fill in the boxes, and call it a day. That is why your evaluation results look clean on paper and fall apart the moment you deploy. I spent about fourteen months building trace systems for a fine-tuning pipeline before I figured out that the standard template was backwards. The worksheet should not document what the model did. It should document what you expected the model to do and where the gap appeared. Everything else is just noise.

How to Use a 2 Trace Worksheet Properly

A 2 trace worksheet tracks two parallel data streams across the same inference pass. Stream one records the surface behavior—the input, the output, the token probabilities if your framework exposes them. Stream two records the structural reasons: which rule fired, which condition was skipped, which context window boundary got crossed. Most people only fill in stream one. That is the mistake. Here is how I set this up now. First, define your evaluation criterion before you run anything. Not after. If you decide what counts as success once you see the output, you have already lost the signal. Write down the exact error type you are looking for—hallucination, rule violation, context bleed, repetition loop, whatever your task demands. Then create a column for each one. The actual worksheet has five columns. Column one is timestamp, which sounds obvious until you realize most people skip it and then cannot reconstruct the sequence of failures. Column two is input fingerprint—not the full input, just a hash or short ID so you can trace the same prompt across multiple runs without pasting three thousand words every time. Column three is the expected output, written in your own words, not copied from the training data. Column four is the actual output. Column five is the reason code, which is where most people stop working and just write "not good enough."

The reason code needs to be a closed set. I use codes like HALLUC (the model invented facts), REPEAT (token repetition detected), CONTEXT (cross-context bleed between instructions), RULE (explicit rule violation), and EDGE (boundary condition not covered by any existing category). If you encounter a failure that does not fit, create a new code. Do not force it into an existing one.

Get the Full Details

Number 2 trace, Worksheet for learning numbers, kids learning material ...
Number 2 trace, Worksheet for learning numbers, kids learning material ...

Common Pitfalls That Break Your Trace Data

The biggest problem I see is that people treat the worksheet as an afterthought. They generate outputs, then go back and fill in the form, which means they are reconstructing events rather than recording them. Memory is unreliable. Timestamps are not. Always fill in the worksheet during the run, not after. Another issue is over-tracing. You do not need to record every token probability. That floods the worksheet and makes it unreadable. Record the top three candidates at decision points where the model had to choose between competing outputs. Everything else is detail that nobody will use. I usually capture only the tokens around the failure point, not the entire sequence. The third problem is ambiguous expected outputs. If you write "should be accurate" as your criterion, you have given yourself nothing to measure against. Write down the exact boundary condition. For example: if the input contains a date before January first two thousand twenty, the output must flag it as outdated. That is measurable. "Should be accurate" is not.

When Trace Worksheets Fail Completely

There are scenarios where this method does not work. The primary one is highly creative tasks where the concept of "correct output" does not exist. If you are evaluating a model that writes poetry or generates speculative fiction, the trace worksheet becomes a measurement tool for something that resists measurement. In those cases, I switch to a different approach entirely—pattern tracking instead of rule tracking. You log recurring structures in the output rather than violations of explicit criteria. The worksheet stays the same format, but the columns change meaning. Another limitation is latency. If your evaluation pipeline processes more than five hundred inputs per hour, the trace worksheet becomes a bottleneck. Filling in five columns for five hundred rows takes time that compounds across multiple runs. I usually run trace worksheets on a subset—twenty percent of the total inputs, stratified by difficulty level. You get signal without drowning in data. Everything else gets logged separately in the raw output file.

A Realistic Problem I Encountered

Two weeks ago I was evaluating a fine-tuned model on a legal document classification task. The trace worksheet flagged seventeen outputs as HALLUC. I checked the inputs and realized the model was not hallucinating facts. It was hallucinating citations—specifically, it was inventing case law that matched the pattern of real citations but did not exist. The worksheet code HALLUC was too broad for this edge case. The workaround was simple but took me three days to figure out. I created a new reason code: CITREF, which stands for citation reference fabrication. Then I filtered the worksheet by that code and cross-referenced with a database of real case law. The model had a systematic bias toward certain citation formats—specifically, it favored Bluebook style over ALWD, which is irrelevant to the task but reveals which training data the model had encountered most frequently. That insight would not have shown up in the original HALLUC category.

Tracing Numbers Activity, Number 2 Trace, Count and Color PDF Worksheet ...
Tracing Numbers Activity, Number 2 Trace, Count and Color PDF Worksheet ...

Building Your Own Trace System

If you want to implement this, start with a simple spreadsheet. Do not build a custom tool until you have used the worksheet for at least two weeks and identified the gaps. I recommend Google Sheets because it supports real-time collaboration and version history, which matters when you are working with a team. Excel works too but lacks the collaborative features that make trace data useful across multiple evaluations. The actual implementation has four components. First, the worksheet template with the five columns I described. Second, a token probability capture script if your framework supports it—this logs the top three candidates at decision points without flooding the spreadsheet. Third, a reason code classifier that automates the categorization but leaves edge cases for manual review. Fourth, an output aggregation view that shows failure patterns across multiple runs, which is where most people stop working and just look at individual rows. I usually spend about twelve minutes per input filling in the worksheet. That is the baseline for a trained evaluator working with a familiar task. New evaluators take about twenty-five minutes. The difference is not skill. It is pattern recognition. Once you have seen the same failure type three or four times, you stop reading the input and start scanning for the signal. That is why you should evaluate on a familiar task first before moving to something new.

Advanced Nuances Beginners Miss

The first counter-intuitive insight is that trace worksheets are most valuable when the model succeeds, not when it fails. Recording why an output met your criterion gives you a baseline for what "good" looks like in your specific context. Without that baseline, you cannot distinguish between rare correct outputs and common failures. I usually tag twenty percent of successful outputs with the same reason codes as failures, which reveals whether the model is succeeding for the right reasons or getting lucky. The second nuance is that reason codes are not static. If you encounter a new failure type, create a new code immediately. Do not wait until you have seen it three times. The first occurrence is the one you will forget how to categorize. I have abandoned reason codes after six months when they stopped capturing new failure types. That is normal. The code set should evolve with the task.

Summary of the Approach

The trace worksheet is a diagnostic instrument, not a checklist. It records what you expected and where the gap appeared. The five-column format covers timestamp, input fingerprint, expected output, actual output, and reason code. Use closed reason codes but create new ones when necessary. Run the worksheet during the evaluation, not after. Cover twenty percent of inputs with stratified sampling. Tag successful outputs as well as failures. Evolve your code set over time. If you want the template, I use a Google Sheets file with conditional formatting that highlights HALLUC, REPEAT, CONTEXT, RULE, and EDGE codes in different colors. That makes pattern recognition faster than reading raw text. The file supports real-time collaboration and exports to CSV for aggregation. I usually share it with my team via a read-only link and let them fill in their own runs without editing mine.

Number 2 Trace Worksheet 796x1030
Number 2 Trace Worksheet 796x1030