Why most people pick the wrong problem to solve
When I first started working on production systems, I spent weeks building a caching layer for something that had no actual traffic. The system was fine under light load. Under heavy load, it was fine too. The real issue was a single database query that only ran once per user session and was responsible for maybe 0.3 seconds of latency. I fixed nothing useful while ignoring the thing that actually mattered. The difference between a small problem and a big problem usually comes down to impact surface area, not how loudly it complains. A bug that crashes the checkout page for two percent of users is a big problem. A memory leak that grows three megabytes per hour but won't OOM until four days out? That is a small problem wearing a costume.
Small Problem Vs Big Problem: How to tell which one you are actually dealing with
Here is the way I approach it now. I measure impact before I measure anything else. That means I look at how many users are affected, how often it happens, and how much revenue or trust is at stake. Everything else is secondary. There is a practical trick that most people miss. You should measure the problem during peak conditions. A formatting bug that only shows up when the server is under load, or a slow query that only becomes visible during a deployment window, is usually the big problem hiding behind a small one. I learned this the hard way when a client complained about a UI glitch that happened once per week. We reproduced it on staging and could not find it at all. It turned out the glitch only appeared when three other services were restarting at the same time. The root cause was a race condition in the shared session store. We fixed that and the UI glitch went away too. I keep a simple spreadsheet for this. The columns are impact score, frequency, and fix complexity. Impact score is one through five based on user reach. Frequency is daily, weekly, monthly, or rare. Fix complexity is low, medium, or high. The problem with the highest impact multiplied by frequency consistently rises to the top. Fix complexity matters, but only after you have already eliminated the low-impact noise.
One counter-intuitive thing I have noticed is that the scariest looking problem is rarely the biggest one. A loud error in production gets attention because humans are wired to react to visible stress. That does not make it important. The problem that silently costs money every hour is the one you should be solving. I once watched a team spend two sprints chasing a dashboard that displayed incorrect color coding while the actual conversion funnel had a broken tracking event that was costing them roughly twelve thousand dollars a month. The dashboard looked terrible. The tracking issue was invisible. The invisible one was the big problem. Another thing people get wrong is assuming a small problem will always stay small. It will not. A tiny consistency bug in a data pipeline can corrupt thousands of records over six months if nobody looks. I saw this happen with a log aggregation service where a timestamp parsing edge case only triggered on leap seconds. It seemed negligible for years. Then it started dropping entire days of data during Daylight Savings transitions and the compliance team flagged it during an audit. That was a small problem that became a big problem because nobody measured its growth over time. There are situations where this framework breaks down completely. If you are in a regulated industry and a minor issue has legal implications, it is automatically a big problem regardless of user count. If you are dealing with security, a single vulnerability that allows privilege escalation is a big problem even if the exploit chain is obscure. Never let the impact model override compliance or security requirements.
Get the Full Details

When you have a genuine small problem, sometimes the best move is to ignore it until it either grows or becomes irrelevant. I have deleted entire features that were causing headaches for three percent of users because those same users were not on a paid plan and the feature took up twenty percent of our maintenance capacity. That is not cowardice. It is resource allocation. Writing code to fix a low-impact problem is still spending engineering time, and that time cannot be used elsewhere. If you want a quick starting point, I use a modified version of the RICE scoring model but replace Reach with Impact and add a penalty factor for complexity. It is not perfect. It tends to undervalue problems that are hard to quantify, like developer frustration or technical debt that slows future work. For those, I add a separate category called maintainability tax and score it independently. That usually catches the problems that slip through the normal filters. The short version is this. Measure impact. Watch for peak conditions. Ignore the loud ones unless they are actually loud. And remember that a small problem today can be a big problem next quarter if you let it grow unchecked.