Practical Protein Digestion for Mass Spectrometry
You don't need fancy software to start making sense of Amino Acids And Proteins in a lab setting. You need to understand what happens when you take a purified protein sample and decide whether you're doing bottom-up proteomics or running it intact. Most people I talk to jump straight to buying kits without thinking about their downstream application. That usually wastes money and gives you garbage data anyway. Tryptic digestion is the standard workflow, and here is what actually happens during it. You denature your protein with 8M urea or 0.1% SDS, reduce disulfide bonds with DTT at 56°C for 30 minutes, then alkylate with iodoacetamide in the dark for 20 minutes. After that you dilute the urea below 2M because trypsin dies in high urea, add calcium chloride to stabilize the enzyme, and digest at 37°C overnight. The result is peptides that are mostly 7 to 20 amino acids long, which is the sweet spot for most LC-MS/MS instruments.
Amino Acids And Proteins: What Actually Matters
Twenty standard amino acids form proteins through peptide bonds, yes, but the thing nobody tells you until it ruins your experiment is that two amino acids—cysteine and methionine—get easily modified in ways that will confuse your sequence search. Oxidized methionine adds 15.99 Da. Carbamidomethylated cysteine is your fixed modification. If you forget to alkylate properly, disulfide shuffling will happen during digestion and you will spend three days trying to identify a protein that appears as a dozen different charge states instead of one clean spectrum. The molar mass of a free amino acid ranges from 75 Da for glycine to 204 Da for tryptophan. When they link into a polypeptide chain, each peptide bond formation releases one water molecule, so the residue mass is always the free amino acid mass minus 18.015 Da. This matters because your mass spectrometer measures peptide mass, not protein mass directly, and if your software is using wrong residue masses your entire identification pipeline breaks. I ran into this exact problem once with a membrane protein prep. The cysteine content was unusually high and the alkylation step kept failing because the detergent I was using interfered with iodoacetamide access. Half the cysteines remained unmodified and formed random disulfide bridges during the tryptic digest. The resulting peptide mix had so many variant forms that the search engine flagged everything as low-confidence. I switched to -mercaptoethanol for reduction, removed the detergent by acetone precipitation before alkylation, and the next run gave clean spectra. Took about 20 extra minutes of sample handling but saved me a week of troubleshooting.
Collision Energy and Fragmentation Reality
HCD fragmentation is the standard for modern instruments, but the collision energy setting needs to match your peptide length and charge state. A +2 peptide at 800 Da needs different energy than a +3 peptide at 1200 Da. Most people set a fixed normalized collision energy and wonder why some peptides fragment completely while others sit there as precursor ions. I use an adaptive setting—typically 25 to 35% NCE depending on the predicted peptide properties—and it cuts down on missing sequences significantly. Here is something that surprises people: amino acid composition affects fragmentation pattern in ways that are not linear. Proline at the C-terminal side of a cleavage site promotes b-ion formation due to its rigid ring structure. That is why proline-rich peptides often give cleaner spectra. Conversely, sequences with multiple consecutive hydrophobic residues tend to form strong gas-phase ions that resist fragmentation and show up as intense precursor peaks with weak fragment signals.
Get the Full Details

Common Pitfalls in Sequence Interpretation
Missed cleavages happen. Trypsin does not cut after every arginine and lysine, especially when the next residue is proline. Expecting 100% cleavage efficiency means your identification rate drops because the search engine is looking for peptides that never exist. I allow up to two missed cleavages in my searches and it increases my confident identifications by roughly 15 to 20 percent on standard yeast digest samples. Another issue is contaminant peaks. Keratin from skin cells is everywhere in any lab environment. If your blank run shows abundant keratin peptides, your sample handling area needs attention more than your instrument does. I also see people mistake partial sequence matches for valid identifications. A peptide matching 6 out of 10 residues is not an identification, it is noise. The machine learning scoring in modern search engines like MaxQuant or Proteome Discoverer handles this, but only if you feed it reasonable parameters.
When Intact Protein Analysis Makes More Sense
Sometimes digestion is the wrong approach. If you are studying post-translational modifications on a specific lysine residue near a disulfide bond, bottom-up proteomics will lose that context because the peptide does not span both features. Native MS or top-down approaches preserve that information, though they require different instrumentation and more method development time. For routine identification and quantification, tryptic digestion remains the most efficient path. For mapping complex modification patterns on a single domain, consider keeping the protein intact and running it through an ETD or ECD fragmentation method on a capable instrument. The trade-off is speed versus information depth. Digestion plus LC-MS/MS gives you coverage across dozens of proteins in a single run. Top-down typically handles one or two proteins per run with deeper modification detail. There is no universal answer. It depends on what question you are actually trying to answer with your sample.