Getting Through Molecular Evolution Without Losing Your Mind
Molecular evolution courses tend to pile on a lot of mathematical theory while assuming you already know how to apply it. The problem isn't the material itself, it's that most students hit a wall somewhere around neutral theory and Kimura's two-parameter model, then spiral trying to catch up. I ran into this repeatedly when I was tutoring undergrads, and it usually comes down to one thing: people memorize formulas without understanding what the variables actually represent in a real phylogenetic context. The answer key for course 174 in molecular evolution circulates through a few reliable channels. Most commonly you'll see it shared on course-specific Discord servers, GitHub repositories maintained by former students, or academic file-sharing platforms like Course Hero. The most complete versions I've seen include worked solutions for problem sets covering maximum likelihood estimation, molecular clock calibration, and dN/dS ratio calculations. I tend to grab mine from GitHub repos because those tend to stay updated when the professor tweaks problem sets between semesters. One specific issue I ran into last semester was that the answer key for problem set 4 had a known error in the maximum parsimony calculation for the third question. The published answer listed a tree length of 47, but running it through PAUP with the same character matrix gave 52. I spent about forty minutes debugging before I realized the error was in the key itself, not my code. The workaround was to cross-reference with the in-class lecture notes where the professor had written out the stepwise character mapping on the board. That manual walkthrough caught the discrepancy immediately. My advice is always trust the lecture material over any circulated key if they conflict.
How the Material Actually Works in Practice
Most students approaching this course come from a biology background with limited quantitative training, which creates a specific bottleneck. Molecular evolution is fundamentally a statistical subject dressed in biological notation. When you're calculating transition/transversion ratios or fitting substitution models, you're doing maximum likelihood optimization, period. The biological framing just tells you what the parameters mean. One thing that catches people off guard is how sensitive likelihood calculations are to the starting tree. If you're running ML phylogenies and your bootstrap support values are all over the place, the issue is rarely your model choice. It's usually that your initial tree topology sent the optimizer toward a local optimum. I've seen students try twenty different substitution models before realizing they just needed to rerun with a different starting tree generated by neighbor-joining. The computational savings from fixing that one thing typically cut analysis time from several hours down to under thirty minutes on a standard laptop. Another area where the answer keys can mislead is in the molecular clock section. The standard problems assume a strict clock, but real data almost never fits that assumption. When I work with actual sequence data, I almost always test for clock-likeness using a likelihood ratio test before committing to a relaxed clock model. The answer keys rarely mention this step because textbook problems are sanitized. Skipping the clock test and jumping straight to dating analyses is probably the most common mistake I see, and it produces confidently wrong divergence time estimates that look perfectly reasonable to someone who doesn't know what to look for.
What the Answer Key Won't Tell You
The answer key gives you the right numbers, but it doesn't teach you how to interpret them in a research context. Here's what I wish someone had told me before starting this course: dN/dS ratios greater than one don't automatically mean positive selection. They can also indicate relaxed purifying selection, population bottlenecks, or even sequencing errors in low-coverage data. The standard approach of comparing a single dN/dS value across a branch is almost never sufficient for a real publication. You need branch-site models, and even then you need multiple lines of evidence. Similarly, the coalescent theory problems in this course make heavy use of the Wright-Fisher model assumptions. Those assumptions break down quickly with real population structure. If you're using these methods for actual research, you'll want to move toward structured coalescent implementations or approximate Bayesian computation rather than relying on the standard formula results. The textbook answers work beautifully for idealized scenarios. They're less useful when your samples come from five different geographic locations with asymmetric gene flow. I'd also note that the answer key tends to gloss over model selection entirely. In practice, picking the right substitution model using AIC or BIC should take five minutes and should be the first thing you do before any tree building. Skipping model selection and running a default GTR analysis is a habit I see from people who just want the answer key numbers to match. It produces results, but the confidence intervals are unreliable and reviewers will tear them apart. ModelFinder in IQ-TREE does this automatically and it's free, so there's really no excuse to skip it anymore.
Get the Full Details

A Few Practical Notes
If you're working through this course on your own, I'd recommend pairing the answer key with R packages like ape and phangorn for hands-on practice. Running the problems through actual software rather than just checking your handwritten answers against the key makes a noticeable difference in retention. The gap between knowing the formula and being able to implement it is wider than most students expect, and bridging that gap is what actually prepares you for research-level work in this field. The answer key itself is fine as a checkpoint tool, but treat it as a starting point for deeper investigation rather than the final word. Every problem in this course has at least one edge case that the standard solution sidesteps, and recognizing those edge cases is what separates people who can pass the exam from people who can actually do this work.