Reading toxicology slides is mostly about knowing which changes to ignore
Most pathologists I work with spend more time deciding what doesn't matter than what does. That might sound backwards, but it's how the work actually goes. You get a liver section from a rat that received the test article for 28 days and there's this tiny focus of hepatocyte necrosis. Is it treatment-related? Maybe. Or maybe the animal was slightly stressed during cage change and you got a periportal artifact. The distinction matters because regulatory reviewers will ask you about it either way. I've been doing this for about twelve years across three different contract organizations and one pharma internal lab. The work hasn't changed much in terms of what you look for, but the expectations around documentation have gotten heavier. Here's what I've learned about making sense of preclinical toxicity findings and actually using them for safety decisions.
What histopathology of preclinical toxicity studies interpretation and relevance in drug safety evaluation actually means in practice
The phrase itself sounds like something pulled from a guidance document, but the reality is simpler. You're looking at tissues from animals given a drug candidate at multiple doses, then deciding whether observed changes are relevant to human safety risk. That's it. Everything else is details. The key word is relevant. A change can be statistically significant across dose groups and still be irrelevant if it falls within the historical control range for that strain, that laboratory, and that age of animal. I once sat through a three-hour review meeting where someone wanted to flag mild renal tubular dilation as treatment-related simply because the p-value was under 0.05. The median change was half a micrometer and the historical control database for that colony showed 40 percent of animals had some degree of the same finding at the same anatomical location. We dropped it. The reviewer wasn't happy but the data was the data. Relevance also depends on context. A single focus of nephropathy in one animal out of ten in the high-dose group is almost never treated as a real finding unless there's corroborating clinical chemistry evidence. But diffuse cortical mineralization in eight out of ten animals at the same dose with elevated BUN and creatinine? That's a target organ effect and you write it up accordingly.
The interpretation workflow most people get wrong
There's a standard way to approach this that comes from the guidelines, but people tend to rush through steps they don't actually need to revisit. Here's what I do: Step one: Look at the raw data before you write anything. This means the clinical pathology results, the organ weights, the gross pathology notes, and yes, the survival data. I've seen pathologists grade slides in isolation and then produce interpretations that contradicted the clinical chemistry by a mile. The tissues don't lie but they also don't tell the whole story by themselves. Step two: Establish the baseline for that specific laboratory and colony. Historical control data isn't optional. If your lab's database for F344 rats shows spontaneous proliferative lesions at 12 percent incidence in the thyroid C-cells at 26 weeks, you don't flag a similar finding at 10 percent in the high-dose group without careful consideration. The NTP has databases you can reference, but they're not the same as your own colony's numbers. Different housing conditions, different diet, different vendor. These matter more than people admit.
Get the Full Details

Step three: Determine whether the change is adaptive or injurious. This distinction matters for relevance. Hepatocellular hypertrophy is often adaptive and reversible. True necrosis is not. The difference can be subtle in early lesions. I use a combination of cytoplasmic basophilia, nuclear enlargement, and mitotic activity to tell them apart. Mitoses in this context usually point toward regeneration rather than hypertrophy, but you need to see the surrounding architecture to be sure. Step four: Establish dose-response and temporal relationship. This is where most interpretations either hold up or fall apart. A finding that appears at all doses with increasing severity is straightforward. A finding that appears only at the highest dose with no severity gradient is harder to defend as treatment-related, especially if the animals were starting to lose weight or had reduced food consumption. Death is not a dose-response relationship.
Edge cases that will make your life difficult
Spontaneous lesions in aged animals are the biggest pain point. I ran a 6-month study in rats that started at 6 weeks of age and ended around 36 weeks. By that point, nearly every animal had some degree of chronic progressive nephropathy. Distinguishing drug-induced tubulointerstitial inflammation from age-related changes required comparing the pattern of infiltration, the presence of casts, and the degree of fibrosis against matched controls that were similarly aged. The controls also had CPN but the pattern was different. Drug-related inflammation tended to be more interstitial and lymphoplasmacytic while age-related changes were more tubular with hyaline casts. Another problem is species differences in susceptibility. Mouse lung is notoriously sensitive to irritants and you'll see inflammation in the nasal turbinates and lung even with compounds that are perfectly safe in rat. I once reviewed a study where a pharmaceutical company was concerned about pulmonary inflammation in mice treated with a compound that showed no corresponding finding in rats at comparable exposure levels. The mouse finding was attributed to species-specific hypersensitivity rather than a class effect. Human relevance was considered low based on the absence of any pulmonary endpoint in clinical trials up to that point. Here's something counter-intuitive that beginners rarely grasp: severity grading is less useful than you think for determining relevance. A mild (1+) change that is dose-related and has a plausible biological mechanism can be more clinically significant than a moderate (3+) change that occurs sporadically with no dose relationship. Don't let the grading scale fool you into downplaying consistent low-grade findings.
How to write interpretations that won't get thrown out by reviewers
Regulatory reviewers have seen every possible interpretation and they're not impressed by confidence without evidence. When you write up a finding, include: I keep a template for this and honestly it makes the writing process take about fifteen minutes instead of an hour. The first time I went through it without one it took me two hours and I still missed including the historical control comparison. Reviewers will catch that omission and ask for it anyway. One more thing that saves time: don't reinterpret slides after the clinical chemistry is available. This sounds obvious but I've watched pathologists second-guess their initial slide assessment once they saw elevated ALT and decided the liver findings must be more severe than they originally graded. Stick to what the tissue shows. If the clinical chemistry contradicts your morphological assessment, note the discrepancy rather than inflating the severity to make them match. Reviewers respect honesty more than forced consistency.

When histopathology alone can't save you
There are situations where you simply cannot determine relevance from tissue sections. A change in liver weight with no morphological correlate at the light microscopy level might indicate enzymatic induction without visible cellular change. You'd need special stains or electron microscopy to confirm. Conversely, a clear morphological lesion with normal organ weights and clinical chemistry might be incidental and not functionally significant. These mismatches happen more often than the literature suggests. For those cases, I recommend augmenting the standard workup with immunohistochemistry for proliferation markers like Ki-67 or PCNA. It takes about two extra days and costs roughly three hundred dollars per antibody, but it can resolve ambiguous proliferative lesions faster than waiting for a follow-up study. I've used this approach to distinguish reactive hyperplasia from neoplastic proliferation in thyroid follicular cell adenomas where the boundary was unclear on H&E alone. The bottom line is that histopathology interpretation in preclinical toxicology is as much about knowing what to exclude as what to include. Most findings are noise. Your job is to separate signal from it, document the signal properly, and not pretend the noise is meaningful just because a reviewer might expect you to find something. The data will tell you what it tells you. Trust it.