So you need to classify mutations, let's get this straight
Gene mutation is one of those topics that gets taught with too much hand-waving in undergrad biology. You show up expecting a clean taxonomy, and what you actually get is a slide deck that conflates mechanism with consequence. I've spent more years than I want to admit sorting through sequencing data where the distinction between a point mutation and a frameshift wasn't just academic it determined whether we called a variant pathogenic or benign. The framework most people learn breaks mutations into broad buckets based on what happens at the DNA level. Substitution mutations swap one base pair for another. Insertions and deletions add or remove bases. Repeat expansions stretch out through repeated sequence units. Structural rearrangements move chunks of DNA around. Each category behaves differently in practice, and the reason matters when you're reading a VCF file at 11pm trying to figure out why your phenotype doesn't match the annotation.
Substitution mutations: the quiet ones that lie
Point substitutions are the simplest type to identify computationally but the hardest to interpret clinically. A single nucleotide change can be silent, missense, or nonsense depending on which codon position it hits and what the genetic code does at that spot. Synonymous substitutions don't change the amino acid because of codon degeneracy. Most variants get filed away as benign polymorphisms by automated pipelines, and that assumption costs people lives. Here's the part nobody emphasizes enough: synonymous mutations can absolutely disrupt splicing. The branch point, the splice donor, the acceptor sequence, these are all RNA-level signals encoded in what looks like coding DNA. I spent three weeks chasing a case where a variant annotated as benign missense turned out to be a cryptic splice-site disruptor after we ran RT-PCR. The SIFT and PolyPhen scores were both green. The ClinVar submission was empty. The patient had symptoms consistent with the gene's known disease association. Sequencing the cDNA resolved it within a day once we knew what we were looking for. Missense substitutions are where the real judgment call lives. Not every amino acid change is damaging. You need conservation data, structural modeling, functional assays, and population frequency all converging before you can make a confident call. The gnomAD database has changed this field dramatically. A missense variant showing up at 2 percent allele frequency in the general population is almost never a highly penetrant disease mutation, regardless of what your in silico tool predicts.
Insertions and deletions: the frameshift mess
Indels are where sequencing technology still trips up. Short insertions and deletions within homopolymer runs or repetitive regions are notoriously hard to call accurately with short-read platforms. You'll see false positives at rates that make no biological sense, especially in low-complexity regions where the aligner just gives up and places reads randomly. When an indel isn't a multiple of three bases, it causes a frameshift. The reading frame shifts downstream, translation runs into a premature stop codon, and the resulting protein is usually nonfunctional. This is the mechanism behind many loss-of-function variants in tumor suppressor genes. BRCA1 and BRCA2 have dozens of documented frameshift variants used as clinical test endpoints. The logic is straightforward, but the exception rate in practice is higher than the textbooks suggest. Some indels do maintain the reading frame. In-frame insertions or deletions remove or add one or more complete codons without shifting the frame. These can be deleterious when they disrupt a critical domain, like the CFTR delta F508 deletion that removes a single phenylalanine and causes cystic fibrosis, but they can also be tolerated in less constrained regions. You can't assume a frameshift is always worse than an in-frame event without looking at the specific protein architecture.
Get the Full Details

Repeat expansion mutations: the structural problem
Mutations involving tandem repeat expansions operate on a completely different principle than point mutations. Trinucleotide repeat disorders like Huntington's disease, fragile X syndrome, and myotonic dystrophy involve copy number changes that are essentially invisible to standard variant callers. A CAG repeat expanding from 15 copies to 45 copies looks like a normal variant to a GATK pipeline. It's a pathogenic expansion. The technical reason is that short-read sequencers can't span these repeats reliably. The reads map equally well to dozens of positions within the repeat tract, so the alignment score bottoms out and the variant caller discards the evidence. You need specialized assays like PCR with fragment analysis, long-read sequencing, or triplet-primed PCR to detect these. If your lab is only running standard whole-exome sequencing, repeat expansion disorders will slip through the net every single time.
Structural variants and chromosomal rearrangements
Large-scale mutations include deletions, duplications, inversions, translocations, and insertions that affect kilobases to megabases of genomic material. These are the mutations that show up on karyotypes and FISH assays, not on standard SNP arrays or exome sequencing. Copy number variants detected by array CGH or SNP arrays represent the intersection between structural variants and clinically actionable mutations, which is why clinical geneticists care about them so much. Translocations deserve a specific mention because they don't just disrupt genes they create fusion genes. The BCR-ABL1 fusion from the Philadelphia chromosome is the textbook example, but translocation-derived fusions appear in lymphomas, sarcomas, and leukemias routinely. The diagnostic workup for these requires break-apart FISH probes or RNA-seq fusion detection, not DNA-level variant calling. A DNA-only analysis will flag the breakpoint region as a structural variant but won't tell you whether the transcripts are being produced and whether they're oncogenic.
What the classification system doesn't tell you
The types of gene mutation framework is useful for organization, but it's not how the genome actually behaves. A single nucleotide variant can cause a frameshift if it sits at the right boundary of an exon. A "benign" structural variant might create a new regulatory element that dysregulates a distant gene. The mechanistic label you assign to a mutation tells you nothing about its clinical impact without additional context: zygosity, tissue-specific expression, compensation mechanisms, variable penetrance, and modifier genes all matter more than the mutation type itself. Also worth noting: the classification breaks down completely for somatic mutations in cancer. A tumor mutation isn't inherited, it accumulates under selective pressure, and the same substitution can be driver in one tissue and passenger in another. The COSMIC database exists precisely because the germline mutation framework doesn't apply to the vast majority of variants found in cancer genomes. There's also the issue of compound heterozygosity, which the basic taxonomy ignores entirely. Two different mutation types in trans at the same locus, one missense and one frameshift, can produce a recessive disease phenotype indistinguishable from homozygous loss-of-function. Your annotation pipeline will tag them separately and may miss the clinical significance if it evaluates each variant in isolation rather than in the context of the individual's diploid genotype.

If you're working through this clinically, I'd recommend starting with ACMG variant interpretation guidelines rather than mutation type classification. The guidelines force you to consider population data, computational predictions, functional evidence, segregation, de novo occurrence, and allelic data together. The mutation type is one data point in that framework, not the framework itself. And if you're doing any kind of diagnostic sequencing, budget extra time for orthogonal validation of indels and structural variants, because the primary caller will get some of them wrong and the false negative rate in repetitive regions is real.