Getting Past the Basics

Protein is made of amino acids. That's the textbook answer, but the real story is messier than that. There are 20 standard amino acids that human biology string together into long chains, and how those chains fold determines everything about what the protein actually does. The ones I deal with in practice are never just simple chains on a page. I spent several years working with recombinant protein expression in a biotech lab, and one of the first things you learn is that the sequence alone doesn't tell you what's going to happen. You can have the perfect DNA construct, run it through the expression system, and end up with a protein that's aggregated into insoluble inclusion bodies because the folding kinetics didn't match your growth conditions. I once ran a batch where the target protein was producing at what looked like excellent levels on a gel, but when I tried to purify it, nothing came through the column. The problem wasn't contamination or degradation. The protein had misfolded into a form that looked normal by SDS-PAGE but was functionally useless. I ended up having to optimize the expression temperature down to 16°C and add arginine to the media to get it soluble. That single change cut my yield in half but actually gave me usable protein for the first time.

What Is Protein Composed Of at the Structural Level

Amino acids link via peptide bonds, which are covalent connections between the carboxyl group of one amino acid and the amino group of another. This creates the primary structure — the linear sequence. From there, local hydrogen bonding patterns create secondary structures like alpha helices and beta sheets. These elements then pack together through hydrophobic interactions, disulfide bridges, ionic bonds, and van der Waals forces into the tertiary structure. Some proteins, like hemoglobin, come together from multiple polypeptide subunits, which is quaternary structure. Here's something beginners routinely miss: the composition isn't just about which amino acids are present. It's about their post-translational modifications. Phosphorylation, glycosylation, acetylation — these can completely change a protein's behavior without altering the underlying sequence. I've seen people design binding assays and get confusing results only to later discover their protein was being glycosylated differently depending on whether it was expressed in E. coli or in mammalian cells. E. coli doesn't do glycosylation at all, which is great for simplicity but meaningless if you're trying to model a human therapeutic protein.

The Practical Breakdown

When you actually need to figure out what a protein is composed of — say you're characterizing a novel enzyme or validating a biosimilar — you're not going to get there from just the gene sequence alone. Here's the workflow that actually works in practice: Start with mass spectrometry. LC-MS/MS will give you the amino acid sequence by breaking the protein into peptides, ionizing them, and measuring their mass-to-charge ratios. You match those against a database. This is fast, usually takes a day or two if your instrument is running, and it catches unexpected modifications. I've used this to identify N-terminal methionine excision that hadn't been predicted, which matters because the presence or absence of that starter methionine can affect protein stability and immunogenicity. For determining the exact amino acid composition rather than just the sequence, you do acid hydrolysis followed by HPLC or LC-MS. You break every peptide bond with strong acid — typically 6N HCl at 110°C for 24 hours — and then quantify each free amino acid. This tells you the molar ratios. One caveat worth knowing: tryptophan gets destroyed in standard acid hydrolysis, so if you need to quantify it you have to use a different method like alkaline hydrolysis or directly measure it from the peptide map.

Get the Full Details

Protein - Wikipedia
Protein - Wikipedia

For secondary and tertiary structure, circular dichroism is the workhorse. It measures how proteins absorb polarized light differently depending on their folded state. An alpha-helix-rich protein will show a characteristic double minimum at 208 and 222 nanometers. I've used CD to confirm that a reformulated buffer preserved the native fold after a protein went through several lyophilization cycles. The yield numbers looked fine on a Bradford assay, but the CD spectrum told the real story — the protein was partially unfolded even though the concentration was correct.

Common Pitfalls and Where Things Fall Apart

The biggest mistake I see is treating protein composition as a static property. It isn't. The same gene can produce proteins with different compositions depending on the expression system, the cell line, the media, the induction timing, and even the scale. A batch made in a 5-liter bioreactor at commercial scale often has a different glycosylation profile than one made in a 2-liter shake flask during development. If you're comparing data across scales, you need to account for this. Another issue is the assumption that SDS-PAGE tells you about purity or composition. It doesn't. SDS-PAGE separates by molecular weight under denaturing conditions. Two proteins with the same apparent weight but completely different sequences will sit on the same band. I once spent weeks troubleshooting a purification that kept failing, only to realize the contaminant was a chaperone protein with nearly identical molecular weight to my target. The gel looked clean. The activity assay didn't lie, but I should have used something more definitive much earlier. If you're working with membrane proteins, most of the standard composition analysis methods struggle. These proteins need detergents to stay soluble, and detergents interfere with everything from MS ionization to CD measurements to HPLC separation. I've had good luck diluting the detergent below its critical micelle concentration before injecting onto the MS, but it's finicky and you lose some material in the process. It's one of those cases where the standard protocols need significant adjustment rather than straightforward application.

A Note on What You Can't Easily Determine

Solid-phase peptide synthesis can build custom sequences, but it's not scalable beyond roughly 50 amino acids without degradation becoming a serious problem. For larger proteins, you're stuck with biological expression or native isolation. Neither is perfect. Expression systems introduce artifacts like non-native disulfide bonds or missed cleavage sites. Native isolation gives you the real thing but in quantities that are often too small for comprehensive characterization. X-ray crystallography will tell you the precise atomic composition and arrangement, but you need crystals first, and not all proteins crystallize well. Cryo-EM has made large complexes much more tractable, but for small proteins under 50 kDa, the resolution still isn't where crystallography is. There's no single method that covers all cases cleanly.

67 going on 50… : HOW TO...CALCULATE DAILY PROTEIN NEEDS
67 going on 50… : HOW TO...CALCULATE DAILY PROTEIN NEEDS