Management language is messy, and most people handle it wrong.
I spent years building performance review systems for mid-size tech companies, and the single most broken part was always how managers described their teams. Words carry weight in ways people don't expect. When a manager writes that someone is "struggling with communication," that phrase travels. It goes into a permanent file. It shows up in background checks. It frames someone's career for years. The vocabulary you pick matters more than you think. Most organizations try to fix this with a simple word bank. Pick an adjective, rate the person, move on. That approach collapses under actual use. "Dependable" means something completely different when two different managers write it. One person uses it for someone who hits deadlines. Another uses it for someone who never causes problems. The words exist on paper but not in practice, which is why I built what we called a descriptive management framework into our HR stack back in 2018.
Words To Describe Management
The core insight is that management descriptors need three anchors: a behavior, a frequency, and an outcome. Not just "poor performer" or "excellent leader." Those are verdicts, not descriptions. A verdict closes a conversation. A description opens one. When you say someone "initiates process improvements without being asked," you're describing something a hiring manager can verify. You're not handing them a label to interpret. The framework I used had four tiers. Tier one covered observable actions. Things like "documents workflows after completion" or "escalates blockers within 24 hours." Tier two added context about scope. Did they do it alone? With a team? Across departments? Tier three tracked consistency over time. Someone who does a good thing once isn't reliable. Someone who does it every quarter is. Tier four tied it to measurable impact. Time saved, revenue influenced, error rates reduced. This felt heavy at first. It took about twelve minutes per review cycle to fill out properly once the team got used to it. We cut evaluation time from forty minutes to roughly fifteen because nobody spent half an hour debating whether someone was "meh" or "good." Here is where it gets tricky and where most people abandon this approach. The descriptors need calibration against actual role expectations. A senior engineer and a junior developer get judged by different baselines. When I first rolled this out at a logistics company, we applied the same descriptor templates to everyone regardless of level. Within six months, we realized "takes initiative" meant something very different for a warehouse supervisor versus an entry-level coordinator. The fix was adding level-specific anchor phrases. Instead of just "proposes new ideas," it became "proposes and pilots process changes that affect their team." That small addition prevented half the conflicts we were seeing in calibration meetings.
Another thing nobody warns you about: words drift. Over eighteen months, the average manager in my experience shifted their usage by roughly two descriptor bands. Someone who rated people as "meets expectations" in January was routinely rating them "exceeds" by August. They hadn't changed their standards. They had just gotten comfortable with the words. The workaround was a quarterly calibration session where three managers independently described the same person using the framework, then compared notes. This exposed drift fast. One manager consistently used "collaborative" for anyone who didn't complain. Another used it for people who actively enabled other teams. Seeing both descriptions side by side made the difference obvious without having to lecture anyone. There is a real downside to this system that I should mention. It does not work well in companies with high turnover or where managers change frequently. The framework depends on institutional memory. When you have managers rotating every six months, the anchor phrases lose their shared meaning. People revert to shorthand. We saw this at a startup where the engineering lead switched every quarter, and the review documents became nearly useless within a year. In those cases, a lighter approach works better. A simple three-question template works: what did they do, how often, and what changed because of it. Skip the tier system. Keep it short enough that a new manager can actually use it without reading a handbook. Another pitfall is the false precision problem. Describing someone as "consistently resolves cross-team blockers within 48 hours" sounds more scientific than "good team player," but it only holds value if the 48-hour metric is real. If nobody tracks response times, the number is decorative. It gives the illusion of objectivity while adding nothing. The moment I stopped tying descriptors to things we could actually measure, the system degraded into corporate theater. People wrote elaborate descriptions that sounded impressive but verified nothing. The fix was ruthless: if you can't point to a record, a ticket, or a metric that supports the descriptor, you drop it. Not soften it. Drop it. This made the reviews shorter, not longer, and significantly more useful when someone actually needed to reference them later.
Get the Full Details

If you are looking to implement something like this, start small. Pick three roles in your organization. Build five descriptors per role using the behavior-frequency-impact structure. Train two managers to use them. Run it for one review cycle. See what breaks. Then expand. Do not attempt to roll this out across an entire company at once. The first implementation I tried that way failed because middle management interpreted the descriptors differently across three time zones, and there was no calibration mechanism in place yet. The second rollout, done with a smaller pilot group, stabilized within ninety days. The framework files and descriptor templates from my last implementation are available for reference. They include the four-tier structure with role-specific anchor phrases and the calibration session guide. You can download them here if you want to adapt them rather than build from scratch.