Building a Dbt Problem Solving Worksheet That Actually Works
Most teams I've seen trying to use dbt end up spinning their wheels on the same errors over and over. The model compiles fine locally, runs great on dev, then fails in CI with some cryptic warehouse timeout or a macro expansion that makes zero sense. Instead of posting in Slack and waiting three hours for someone to notice, I built a worksheet. Not some fancy dashboard. Just a simple Google Sheet with a repeatable structure that forces you to capture the actual error before you start guessing. The core idea is straightforward. When something breaks, you stop the cycle of "have you tried restarting?" and fill out a structured form. I call it a Dbt Problem Solving Worksheet, and it's saved my team an uncountable number of hours because we stopped re-solving the same problems in different ways.
Dbt Problem Solving Worksheet
Here's what the sheet actually contains. First row is always the ticket or issue ID. Second is the model name being run. Third is the exact error message — copied directly from the terminal, not paraphrased. Fourth is the environment: dev, staging, prod, or CI. Fifth is the command that triggered it. Sixth is the hypothesis. Seventh is the fix applied. Eighth is whether it actually resolved the issue. Ninth is the root cause category. I've tried Notion databases, Confluence pages, Jira custom fields. They all fall apart because nobody fills them out when they're frustrated. A Google Sheet works because it takes twelve seconds to open and type into. I know, it's not glamorous. That's why it works. One thing people consistently mess up is the hypothesis row. They write "it didn't work" or "maybe the warehouse was busy." Neither of those are hypotheses. A real hypothesis needs to be testable. Something like "the join exploded because the staging table has nulls in the join key" is testable. You can verify or kill it in about twenty minutes. I force my team to write the hypothesis before they try any fixes. It's annoying at first. It saves time within a week.
Here's a specific edge case that took me way too long to figure out. We had a model that would fail intermittently in CI but never locally. The error was a warehouse credit exhaustion on what should have been a lightweight aggregation. The worksheet captured the pattern — it only happened when running the full materialization chain after midnight. The root cause turned out to be a cron job that ran a data export at 12:05 AM, competing for the same warehouse resources. The fix was staggering the schedule by ten minutes. Without the worksheet, I never would have connected those two things because I was looking at the wrong layer every time. The root cause category column is where most of the actual institutional knowledge lives. Over six months, you build a searchable history of what breaks and why. The categories I use are: data quality issue, warehouse constraint, macro bug, dependency ordering, permission issue, config mismatch, query complexity, and unknown. Unknown is not a failure state. It's a flag that says we need to come back to this one. I've seen teams skip this column because they think it's extra work. That column alone is worth setting up the sheet for. There's a version of this that people try where they auto-populate fields from dbt logs using a script. It sounds efficient until your engineer runs dbt in a weird terminal configuration or copies the error from a log file that's been rotated. Half the entries end up with garbled or missing error messages, and now you have worse data than if you'd just typed it manually. Manual entry forces you to read the error. That's the whole point.
Get the Full Details

Another counter-intuitive thing about this process: you should log the fixes even when they don't work. I used to only write down successful resolutions. That left a massive gap because the wrong turns contain just as much information. If you tried changing the warehouse size from X-Small to Medium and it still failed, that tells you the problem isn't compute-related. Not logging it means the next person hits the same dead end and spends another two hours going through the same motion. The biggest limitation of a worksheet like this is that it only works if people actually use it. I've watched well-structured troubleshooting frameworks die because someone introduced a new tool mid-process. If your team starts using this, enforce it for one month before you evaluate whether it's working. The early entries will be sloppy. The pattern recognition kicks in around entry fifteen or twenty when you're staring at the spreadsheet and notice "oh, we've seen this exact error pattern four times already this quarter." There's also a hard limit to what a worksheet can solve. If your team's fundamental problem is that nobody has dbt experience and they're flying blind, a troubleshooting form won't help. You need documentation and mentorship first. The worksheet is for teams that know the tool and just need a better way to track their failures. Don't use it as a substitute for basic competence.
If you want to grab the sheet, here's the template: Dbt Problem Solving Worksheet Template. It's set up with data validation dropdowns for environment, root cause category, and resolution status. The hypothesis column has a character minimum so you can't submit a one-word guess. Feel free to fork it and adjust the columns for your own stack. I've removed the CI timeout category because my setup doesn't hit those, but you might want to keep it. I don't track exactly how many hours we've saved since setting this up. The number is high enough that I stopped doing the math. What I can say is that we go from broken model to identified root cause in under an hour now instead of the two-day cycle we used to run through. That's not because the problems got easier. It's because we stopped treating each failure as a unique mystery.