Why nobody actually uses the phylogenetic concept the way they claim to
Let's just get this out of the way. The Phylogenetic Concept Of Species defines a species as the smallest monophyletic group of organisms that share a common ancestor. That's the textbook version. In practice, it's much messier. Most people who talk about using this concept haven't actually built a species tree and then tried to draw a line around it. There's a gap between the definition and what happens when you open PAUP* or RAxML and realize your data doesn't cooperate. The actual workflow starts with sequence data. You pick markers appropriate to your taxonomic level. For recently diverged groups, nuclear genes or UCEs work better than mitochondrial DNA because mtDNA can mislead you through introgression or incomplete lineage sorting. I learned that the hard way with a set of amphibian populations I was working on. The mtDNA tree showed what looked like two clean species. The nuclear data told a different story. You align your sequences. Trim the alignment. I use Geneious or manual refinement in AliView depending on the dataset size. Then you select a model. ModelTest-NG or jModelTest will give you the best-fitting model based on AIC or BIC scores. Don't skip this step. Using the wrong substitution model is one of the most common errors I see in papers claiming to apply the phylogenetic species concept, and it completely undermines the results.
Build your tree. Maximum likelihood is the standard now. RAxML-NG or IQ-TREE handle most datasets. For larger phylogenomic datasets with hundreds or thousands of loci, ASTRAL or SVDquartets account for gene tree discordance rather than forcing everything into a single concatenated supermatrix. Concatenation sounds convenient but it can produce strongly supported wrong trees when there's significant incomplete lineage sorting. That's not a edge case. It's the norm for rapid radiations. Once you have the tree, you identify monophyletic clades. The species concept says each monophyletic clade with diagnosable differences represents a distinct species. Diagnosability usually means fixed or nearly fixed nucleotide differences at certain loci. This is where people get subjective about what "nearly fixed" actually means. I tend to use a threshold of 95 percent allele frequency across multiple individuals per population. Anything less and you're calling intraspecific variation a species boundary. I ran into a specific problem last year with a group of desert rodents where two populations were separated by a narrow valley. The phylogeny showed them as sister clades with about 2 percent divergence in the cytochrome b gene and fixed differences at three nuclear loci. On paper, they fit the phylogenetic species concept perfectly. But the morphological data showed continuous variation across the valley with no clear break. I spent about three weeks cross-referencing the genetic clusters against ecological niche models and realized the valley was a relatively recent barrier, maybe ten thousand years old. The populations hadn't been separated long enough for ecological divergence to catch up with the genetic split. I reported them as separate species but flagged the uncertainty explicitly in the manuscript. Peer review caught it too, which was annoying but fair.
Common mistakes that make phylogenetic species concepts look worse than they are
The biggest issue is treating monophyly as a binary yes-or-no question. Trees have support values. A clade with 68 percent bootstrap support isn't the same as one with 99 percent. Many papers I've reviewed gloss over this and just map species boundaries onto whatever topology their software produced without acknowledging that topological uncertainty exists. If your clade isn't well-supported, you don't have evidence for a species boundary. Period. Another problem is sampling density. If you sample two individuals from what you think is one widespread species, your tree might show them as monophyletic or paraphyletic purely due to sampling artifacts. Cryptic diversity is real, but so is sampling error. I always aim for at least five to ten individuals per putative species across their geographic range before making any taxonomic claims. It's more work but it prevents embarrassing reversals later when someone samples more thoroughly and splits your species again. Hybridization is the third major headache. The phylogenetic species concept assumes tree-like evolution. In many groups, especially plants and some fish, hybridization is common enough that reticulate evolution is the rule rather than the exception. A strict tree topology will misrepresent the relationships. Network methods like SplitsTree or PhyloNet can help visualize the conflict, but they don't neatly solve the species delimitation problem. In these cases, the phylogenetic concept breaks down and you need something like the general lineage concept or an integrative taxonomy approach that combines genetics, morphology, ecology, and behavior.
Get the Full Details
The operational definition of "smallest monophyletic group" also creates problems at the subspecies level. If every deeply nested monophyletic cluster qualifies as a species, you end up with a situation where geographically isolated populations with minor genetic differences get elevated to species status while their closer relatives don't. This is sometimes called "taxonomic inflation" and it's a real concern in groups with high levels of population structure. I've seen papers describe over a dozen species in what a field biologist would plainly recognize as a single highly structured species. The data technically supports the delimitation under a strict phylogenetic reading, but it doesn't reflect biological reality in any useful way.
When to use it and when to walk away
The phylogenetic concept works well for groups where deep divergences exist, gene flow is minimal, and monophyly is strongly supported across multiple independent loci. It's particularly useful for microbes and other organisms where the biological species concept is impossible to apply. For sexually reproducing animals with ongoing gene flow, it's best used as one line of evidence within an integrative framework rather than the sole criterion. If you're starting a project and want to apply this approach, begin with a comprehensive literature review to understand what taxonomic work already exists in your group. Check GenBank for existing sequences but don't rely on them exclusively because misidentified sequences are rampant. Build your own sampling plan. Prioritize type localities and geographic edges of known ranges. Sequence at least two nuclear and one mitochondrial marker minimum. Analyze with both concatenation and coalescent methods and compare the results. If they conflict, investigate why before drawing conclusions. The whole process from sample collection to a draft manuscript typically takes six to eighteen months depending on the organism and sequencing strategy. Budget extra time for the analysis phase because troubleshooting model selection and dealing with missing data in phylogenomic matrices is where most projects hit delays. A well-executed phylogenetic species delimitation study is solid work. A sloppy one just generates noise that makes the concept look unreliable.