Understanding What Proteins Actually Are

Proteins are long chains of amino acids folded into specific three-dimensional shapes, and that shape determines what they do. You've probably seen the standard list somewhere - enzymes, structural proteins, antibodies, hormones - but most people never get past memorizing those categories for a test. The real thing is messier than that. I remember working through a lab project back in grad school where we were trying to purify an enzyme, and the textbook description of how it "should" behave didn't match what was happening in the test tube at all. Turns out the protein was glycosylated in a way the expression system didn't support, so it folded wrong and precipitated out. Took me three weeks and a bunch of failed purification runs to figure that out. That's the kind of thing nobody tells you in an intro biology class. Let me give you some actual examples that cover the major functional categories without being dry about it. Hemoglobin carries oxygen in your blood - it's a tetrameric protein made of four subunits, two alpha and two beta chains. Each subunit holds a heme group with an iron atom that binds oxygen reversibly. That's why carbon monoxide poisoning is so dangerous; CO binds to that same iron site about 200 times more tightly than oxygen does, and it doesn't let go easily. Insulin is a peptide hormone, only 51 amino acids long, that regulates glucose uptake. It's made as proinsulin first, then cleaved to remove the C-peptide segment before it becomes active. You see a lot of people talking about insulin resistance without understanding that the problem isn't the protein itself breaking down - it's the receptor on the cell surface stopping responding properly. Keratin is one of the most abundant structural proteins in the body. It forms intermediate filaments that give mechanical strength to epithelial cells. Your hair and nails are basically dead cells filled with cross-linked keratin. Collagen is another structural protein, but it's outside cells in the extracellular matrix. It forms triple helices stabilized by hydroxyproline, which is why vitamin C deficiency causes scurvy - without vitamin C, prolyl hydroxylase can't modify the proline residues, the triple helix doesn't form correctly, and connective tissue falls apart. The connection between a vitamin and structural integrity is something most people don't make until they look at the biochemistry.

How Protein Structure Determines Function

There are four levels of structure to think about. Primary is just the linear sequence of amino acids in the polypeptide chain. Secondary involves local folding patterns - alpha helices and beta sheets held together by hydrogen bonds between the backbone amide and carbonyl groups. Tertiary is the overall three-dimensional shape of a single polypeptide. Quaternary structure only exists in proteins made of multiple subunits, like hemoglobin I mentioned earlier. The sequence determines the structure, and the structure determines the function. That's the central idea, but it's not as simple as it sounds in practice. One counter-intuitive thing that trips up a lot of people is intrinsically disordered proteins. These are regions or entire proteins that don't fold into a stable three-dimensional structure under physiological conditions. They were once considered experimental artifacts or misfolded junk. Now we know they're functionally important. Alpha-synuclein, which is involved in Parkinson's disease, is mostly disordered in its native state. Some viral proteins are largely disordered too. Disordered regions can bind to multiple different partners because they're flexible enough to adopt different conformations. A rigid lock-and-key model doesn't apply here. When I was studying protein-protein interactions, I kept trying to crystallize certain regions and failing repeatedly before I realized they weren't going to fold for X-ray crystallography no matter what I tried. That's when I switched to NMR and other methods better suited for studying dynamic systems. Another thing people miss is that post-translational modifications completely change how you think about proteins. The same gene can produce dozens of different protein products through alternative splicing, phosphorylation, acetylation, methylation, ubiquitination, glycosylation, and other modifications. Phosphorylation adds a phosphate group, usually to serine, threonine, or tyrosine residues, and it can switch a protein's activity on or off. Ubiquitination marks proteins for degradation by the proteasome. These modifications mean you can't just look at the gene sequence and predict everything a protein will do. The proteome is much larger and more dynamic than the genome suggests.

How Proteins Are Made and Broken Down

Translation happens on ribosomes in the cytoplasm or on the rough endoplasmic reticulum. mRNA carries the genetic code from DNA, and tRNA molecules bring the correct amino acids to the ribosome based on codon-anticodon pairing. The ribosome catalyzes peptide bond formation between adjacent amino acids, building the polypeptide chain from the N-terminus to the C-terminus. This is basic molecular biology that you'll find in any textbook, but the quality control steps afterward are where things get interesting and where experiments often go wrong. Chaperone proteins help newly synthesized polypeptides fold into their correct shapes. Hsp70 and the chaperonin complexes like GroEL-GroES are the main ones. They bind to exposed hydrophobic regions on unfolding or nascent chains and prevent aggregation. Without chaperones, most proteins would just clump together into insoluble masses. Misfolded proteins get tagged with ubiquitin and sent to the proteasome for degradation. This is the ubiquitin-proteasome system, and it's crucial for cellular housekeeping. When it fails, you get protein aggregation diseases like Alzheimer's, Parkinson's, and Huntington's. I worked with a researcher once who was studying protein aggregation in yeast models, and we spent months troubleshooting why our overexpressed proteins were forming inclusion bodies instead of folding properly. The fix was reducing the expression temperature and using chaperone co-expression vectors. That alone cut our aggregation problems by maybe 80 percent. Digestion of proteins starts in the stomach with pepsin, which is a protease that works best at low pH. It breaks proteins into smaller peptides. Then in the small intestine, trypsin, chymotrypsin, and carboxypeptidases from the pancreas do most of the work, chopping peptides into individual amino acids or very small dipeptides and tripeptides. Amino acid transporters in the intestinal epithelium move them into the bloodstream. If you're looking at this from a nutrition angle, the key takeaway is that your body doesn't really care about the source protein as long as you get all the essential amino acids. Complete proteins contain all nine essential amino acids in sufficient quantities. Most plant proteins are incomplete, which is why people who eat only plants need to combine different sources throughout the day.

Get the Full Details

Population vs. Sample | Definitions, Differences and Example
Population vs. Sample | Definitions, Differences and Example

Practical Considerations and Common Problems

If you're working with proteins in a lab setting, solubility is probably going to be your biggest headache. Recombinant proteins expressed in E. coli often form inclusion bodies - insoluble aggregates of misfolded protein. The workaround is usually to express at lower temperatures, use weaker promoters to slow down translation so the protein has time to fold properly, or tag the protein with something like MBP or GST to improve solubility. Sometimes you need to denature the inclusion bodies with urea or guanidine hydrochloride and then refold by gradual dialysis, but that refolding step is hit or miss and depends entirely on the protein. Purification strategies depend on what you need the protein for. If you just need it for Western blotting, a crude lysate might be fine. For enzymatic assays or structural studies, you need high purity and the right buffer conditions. Affinity chromatography using tags like His-tags binding to nickel or cobalt resin is the standard first step. Then you usually follow with size exclusion chromatography or ion exchange to polish it. The buffer matters a lot more than people expect. Salt concentration, pH, and the presence of reducing agents like DTT or beta-mercaptoethanol all affect stability. I once had a protein that was perfectly stable in one batch of buffer and completely degraded in the next because the supplier changed the glycerol content in their DTT solution. It took me two days to figure out that the variable was the reducing agent concentration, not the protein itself. Storage is another area where people make mistakes. Most proteins should be stored at -80 degrees Celsius for long-term storage, not -20. Freezer temperatures above -70 aren't cold enough to completely stop enzymatic degradation and ice crystal formation. Repeated freeze-thaw cycles destroy a lot of protein preparations. Aliquot your samples so you only thaw what you need. Adding glycerol to 50 percent can help with stability during freeze-thaw, but it interferes with some assays. You have to think about what you're going to do with the protein before you decide how to store it.

The limitations of common techniques are worth mentioning too. SDS-PAGE tells you about molecular weight and purity but destroys any information about native structure. Western blotting is sensitive but again, the denaturing conditions mean you're not looking at functional protein. X-ray crystallography gives high-resolution structures but requires crystals, and many proteins won't crystallize no matter what you try. Cryo-EM has become the alternative for difficult proteins, but it's expensive and needs specialized equipment. NMR works for smaller proteins up to maybe 50 kilodaltons before the spectra get too complex. No single method gives you the full picture, and most projects end up requiring at least two or three different approaches to get reliable results.