Love of Proteins
Every one of us was born into an age of technology and scientific enlightenment. Overwhelming as it is, the world we were plopped into resembles very little of what life was like a few thousand years ago, merely a tiny percentage of the age of our species. Ever since the Industrial Revolution, our world has exponentially grown in complexity.
For many contraptions, you could tear them apart and understand how they work. But cars, smartphones, rockets...? For the vast majority of us, these human innovations are so complex that we abstract them as wizardry: unnatural forces that coexist with our reality.
When I was an undergraduate, I took a class at MIT called 6.004: Computation Structures. Assuming you only knew classical electrodynamics, it bridged the distant worlds of circuits and the wizardry of a computer, giving rise to how transistors, ALUs, CPUs, and caching worked on a physical level. That and my human evolution class were by far the two most world-changing classes I took during my years there, and rekindled a spark of curiosity and fascination in the reality imposed on us.
I've been told many times that digging too deep can take away the magic. I disagree: a thorough understanding is the best way to appreciate and love the world we live in. I love cars more ever since I learned how engines and suspensions work. I love computers more ever since I learned how we fabricated and miniaturized compute and memory. And over the past couple of years, I've turned my attention toward the engineering feats that nature had fabricated long before the age of humanity: life itself.
One of the most mind-boggling things about a computer is how we could solve any problem we face in our daily lives: navigation, time management, communication, information resources. We are spoiled by how versatile computers can be. And yet, deep inside the cells of every life form on earth, we have a macromolecule that is just as insane and versatile: the protein.
When we talk about organic life, the building blocks primarily revolve around carbohydrates, lipids, proteins, and nucleic acids. Carbohydrates provide energy, some structure, and signaling. Lipids provide permeability. Nucleic acids provide information. And proteins provide... function? Like, any function? Sure, I can understand how cells use carbohydrates and lipids: it all comes down to chemical reactions built around fundamental physical principles. As for nucleic acids, if we create an analogy to the bits of a computer, it isn't too far-fetched to see how any information can be encoded by genetic code. But one thing I never understood since I was 14 is how a protein could do just about anything.
The computer is to the cell what the algorithm is to the protein. A protein could be programmed to carry out a very specific chemical reaction. It could be programmed to transport specific molecules from one specific area to another specific area. It could recognize versatile threats, while others can help eliminate them. A protein can push and pull things. It can be a motor, a rope, a storage container, a building foundation, a workshop table, an engine, a drill, a garbage disposal, a basket, a taxi driver, a police officer, an assassin, front-desk security... I'm not making this up as I go: I'm thinking of real examples of specific proteins (enzymes, transport proteins, channel proteins, fibrous proteins, etc.) that perform the roles in a biological system of what any machine or occupation does in a society. Given environmental conditions of constant threats, proteins provide order to maintain that system.
I use the word "specific" a lot because these proteins have a clearly defined role. An enzyme is built to recognize or bind to a very small set of molecules. It's not a chaotic, messy, incompetent worker. Proteins are efficient and don't mess around.
And the craziest thing about this: there is no intelligent mastermind behind these designs. These proteins were designed through billions of years of trial-and-error of gradual genetic mutations to see which were the most successful in surviving environmental threats and maintaining stable population levels. And because there were so many possible routes for these mutations to take, and many varying sets of environmental challenges to overcome, it brought about the grand speciation of life on earth today.
But just like the computer, I stare at a protein in fascination of how a macromolecule could possibly take on so many roles. I'd like to offer my interpretation of how this is possible: how one would transfer mere chemical and physical laws into a specialized function.
This post is a condensed re-telling of Chapters 3 through 6 of Albert Lehninger's Principles of Biochemistry, supplemented by research papers found through Google Scholar Labs, Gemini, and Wikipedia sources. The content is a lot, and at times I'll be too focused on the teachings, but I aim to interject with my own interpretations of how they work and why they exist the way they do.
The Amino Acid Building Block
At the heart of the protein lies the amino acid residue. Just like bits of a computer, the amino acid sequence is the code of the protein's algorithm. Given an alphabet of possible amino acids, they're all linked up the same way, but the magic lies in the macroscopic configuration of that long amino acid chain as it folds into the machine it's meant to be.
Each amino acid monomer consists of a carboxylic acid on one end and an amino group on the other. In the middle, you could technically have a carbon chain, but life on earth resorts to the simple alpha amino acid structure, where between the carboxylic acid and amino groups lies just a single alpha carbon, with one side group attached to it. This side group is the letter of the alphabet. Triads of nucleic acids in DNA and RNA encode exactly one of 20 amino acids (plus an extra one to designate when to stop building the protein) which all vary by this side group.
Lehninger Principles of Biochemistry 6th Edition, Page 79
These 20 different side groups have been specially selected by life as the most elite representations of how to abuse chemical and physical laws to their absolute limits with the use of carbon, hydrogen, oxygen, some nitrogen, and a little bit of sulfur. Some are small. Some are large and contribute to Van der Waals interactions. Some are polar and contribute to hydrogen bonding and dipole moments. Some are charged and contribute to electrostatic forces. Some have aromatic rings which can abuse the way electrons can delocalize and spread among conjugated systems. For anyone who has studied organic chemistry, the amino acids represent the elite group of functional groups that abuse every crazy principle you learn when you just limit yourself to those elements. By rearranging these interactions, it turns out you could achieve endless possibilities when it comes to specialized chemical and physical behaviors.
There are other amino acids out there, but those usually exist as modifications after the RNA has been encoded; usually, a functional group is swapped out or added. For example, 13% of collagen in muscle tissue is hydroxyproline, which is a proline with a hydroxyl group attached to one of the ring carbons.
Lehninger Principles of Biochemistry 6th Edition, Page 118
Lehninger Principles of Biochemistry 6th Edition, Page 118
The backbone itself is small and stiff. Each monomer bonds to one another through a peptide bond: a condensation reaction acylates the carboxylic group of one residue with the amino group of another residue to bond the carbonyl carbon and nitrogen together. What's really neat is that the carbonyl oxygen is so greedy to steal the carbon's electrons that the nitrogen ends up contributing some of its lone pair to the carbon-nitrogen bond in the pi orbitals. This ends up giving the peptide bond 40% double bond character. The pi orbital electrons restrict rotation around the bond length, stabilizing the peptide bond laterally; it can't bend. The amino acid backbone can only bend around its other two bonds: the carbonyl carbon to the alpha carbon bond, and the alpha carbon to the nitrogen bond.
The Conformations of Proteins
In each cell lies a ribosome that prints the RNA code into a raw amino acid sequence with all the side groups. Wherever proteins are synthesized, you're bound to be in a highly polar environment surrounded by water. The hydrophobic side groups hate this and will immediately want to clump together like oil droplets in water. What ends up happening is that your hydrophobic groups clump together into the inner core of the protein, and then the surface of the protein consists of a solvation layer of your hydrophilic side groups that just love to hydrogen bond with water.
Klein Organic Chemistry 3rd Edition, Page 1148
And here lies an astounding insight into one of the most perplexing questions in biology: by the 2nd law of thermodynamics, the universe naturally wants to become more disordered and chaotic, so why does the protein only want to fold in one of but a few configurations, giving rise to the high order and stability of life? Shouldn't it be constantly fluctuating between all sorts of folded configurations?
It turns out that we are obeying the 2nd law of thermodynamics. When you place hydrophobic molecules in water, water absolutely hates it. They can't hydrogen bond with each other and are locked in rigid conformations. So as the water molecules look for each other and the hydrophobic side groups look for each other, the entropy actually increases in the number of possible stable configurations the protein-water system could attain. The formation of this hydrophobic core is the primary instantaneous drive of protein folding.
Thermodynamics further describes that the change in a system over time is driven by a spontaneous search for minimizing the Gibbs free energy of your system, which is a function of your enthalpy (the internal energy), temperature, and entropy. If we consider a map of the countless configurations that a protein could take, we immediately notice that we can minimize this function by maximizing entropy by focusing first on clumping the hydrophobic side groups together and shoving the hydrophilic side groups outward. Beyond that, our other interactions take effect: the Van der Waals interactions, the salt bridges of our charged groups, hydrogen bonding (usually from the backbone or solvation layer), pi stacking, and disulfide bridges (if in an oxidizing environment like the endoplasmic reticulum).
You can imagine a vector space that maps protein conformations to the Gibbs free energy: there exist configurations where this Gibbs free energy reaches local minima. Once the protein folds in a way that it falls into one of these wells, it may struggle to get out of it. With enough heat and control, the protein will have enough energy to move between these potential wells, which is why denaturing a protein causes it to disorganize. But the protein is always jiggling around and breathing, bouncing between local minima wells that all lie in a much bigger potential well that represents ensemble configurations (think of it like the difference of wiggling your finger vs. wiggling your whole hand).
Lehninger Principles of Biochemistry 6th Edition, Page 146
These potential maps are a function of the environment: temperature, pH, charge, the presence of reductive or oxidizing agents. For example, disulfide bridges need an oxidizing environment to form, which is why proteins using cysteine have different purposes depending on whether they're built in free cytosol as opposed to the endoplasmic reticulum. An important environmental factor is the one artificially manufactured by chaperone proteins.
You often can't trust proteins to fold perfectly by themselves. If you're printing a protein, its hydrophobic regions are immediately naked to the world: anything else that's hydrophobic (such as other printing proteins) would want to clump together and mess up your whole system. And even then, the protein might find a configuration it's happy with, but it's misfolded in a way where it's left indecent: exposing a hydrophobic patch to the world. These chaperone proteins like Hsp70 supervise the amino acid sequences as they print and help ensure it only folds with itself. For the proteins that claim they finished folding but are indecent, chaperonins find them and lock them into a hydrophilic dressing room, which forces them to refold until they're decent. These chaperone proteins allow the protein to find its true, designated folded configuration where it falls into its intended potential well.
Lehninger Principles of Biochemistry 6th Edition, Pages 143, 148
Once the protein is finished, you may attach some prosthetic group which is some external molecule or functional group that acts as a cofactor to activate the protein for its specialized purpose. This is very common for enzymes, but there are many other proteins with prosthetic groups, such as the heme group in hemoglobin we'll touch on later.
Structure
So you have your amino acid sequence and it instantly folds in a way that maximizes internal stabilizing forces until it finds a potential well it doesn't want to leave. You might think this must result in disorganized messes. And in some cases, when the protein wants to be funky in certain areas, it can be, but with purpose (known as intrinsically disordered). But in most cases, it turns out that if you choose the right side groups, you can impose organized hierarchical structure.
So far, we've only talked about the primary structure of the protein: what an amino acid residue does in relation to any other given residue. But once you start organizing your amino acids to start acting as a team, you get secondary structure. Your secondary structures can further cooperate together into a tertiary structure through the ideas of "domains" and "motifs", turning our printed polypeptide into a folded protein subunit. Optionally, when different subunits find each other and stick to each other (through non-covalent forces), we get quaternary structure. We'll first start with secondary structure.
Secondary structure is what allows individual amino acids to cause larger localized regions of the protein to start having their own combined physical and chemical properties. In the case of alpha helices and beta sheets, life figured this out through periodicity. An alpha helix is a spiral of amino acids, whereas beta sheets are zig-zag lines that often bundle together side-by-side to create a "sheet".
When we looked at the amino acid backbone, we saw two bonds that you could twist. In a stable configuration, each of these bond angles is fixed and happy. By intentionally manipulating the bond angles of each residue, you could manipulate your protein to take on localized secondary structure. The secret of the alpha helix and beta sheet is to hack each residue to have about the same fixed bond angles in the structure.
Lehninger Principles of Biochemistry 6th Edition, Pages 120, 122
An alpha helix works by twisting each amino acid backbone such that they "orbit" around some axis, such that it takes exactly 3.6 residues to complete a revolution. The side groups stick out away from the spiral. The structure is held in place in two ways.
(1) Along the spiral, every carbonyl oxygen (being the greedy electron-hungry dude) internally hydrogen bonds with the naked amino hydrogen of the residue above it. Due to the 3.6 residue periodicity, all of these hydrogen bonds point in the same direction. That is, in an alpha helix, if there is a hydrogen bond with the amino hydrogen on top of the carbonyl oxygen, then you will never find one in that helix where the carbonyl oxygen is on top. This ends up creating a dipole moment, where the helix behaves like a magnet: it becomes gradually more negatively charged toward the carboxyl terminal, and gradually more positively charged toward the amino terminal. To compensate for this, the termini usually have charged side groups to stabilize the helix.
(2) The side groups are spaced out to perfectly match this 3.6 residue periodicity, such that side groups on one side of the helix will attract each other. For example, you might imagine alternating positive and negative charges about every 3.6 residues, making one very happy side of the helix. In reality, it's more common to have one side of the helix with hydrophobic side groups, and the other side with hydrophilic side groups, making the helix amphipathic.
To further stabilize the alpha helix structure, it tends to prefer certain side groups that don't throw a fuss. Two particularly fussy side groups are those belonging to glycine and proline. Glycine, having a side group of just a hydrogen, is too flimsy and wobbly. Proline, having its nitrogen group attached to a ring, can't bend the way a helix should, so it acts as a kink to disrupt the helix structure (sometimes intentionally). In general, alpha helices prefer small unbranched side chains like Alanine, Leucine, and Methionine, so there isn't much steric hindrance.
Lehninger Principles of Biochemistry 6th Edition, Page 123
Beta sheets are the closest you could have to a straight line of amino acids. It's not going to be perfectly straight, but you could get close. The trick to creating the beta sheets is to print your sequence in a way that the side groups alternate which direction they face orthogonally to your strand's direction: up, down, up, down, etc. This creates a periodicity of two residues, and thus you could do the same tricks we discussed above except for every two amino acids. So if you're trying to make a beta sheet, you could alternate your sequence between polar and non-polar groups to create an amphipathic strand. And in contrast to alpha helices, beta sheets actually prefer bulkier side groups that are branched or aromatic to fill up the space above and below the sheet.
The strands of beta sheets typically aren't found on their own and usually line up side-by-side to form rigid sheets through similar hydrogen bonding as before between the carbonyl oxygen and amino hydrogen. Because your strand has a periodicity of two residues, you can imagine that the geometry has reflectional asymmetry, and thus there are two possible ways to line up pairs of strands: parallel or anti-parallel.
In anti-parallel sheets, the carbonyl oxygen and amino hydrogen lie on the same plane, so this hydrogen bond is relaxed and strong, and the sheet ends up flatter. Anti-parallel sheets are much more common and are used to provide structure and rigidity to the protein. Anti-parallel sheets are flexible: they can be in the hydrophobic core or near the solvation layer; they could have two strands or many strands.
In contrast, in parallel sheets, the carbonyl oxygen and amino hydrogen are on opposite sides of the plane, so the hydrogen bond is more strained. To prevent the strands from clumping closer and leading to side group steric hindrance, the backbone angles instead take on a more stressed zig-zag configuration (where each zig-zag is almost a right-angle). As a result of this strain, parallel sheets are much less common, found only in the hydrophobic core, and have many more strands. They're most commonly used in active sites such as in the TIM barrel or Rossmann fold.
Lehninger Principles of Biochemistry 6th Edition, Page 124
Beyond alpha helices and beta sheets, there are also a few types of sharp turns common in proteins used to link between different structures, such as a beta sheet or an alpha helix, or one strand of a beta sheet to another strand. Beta turns consist of four residues where the first and fourth residues hydrogen bond, turning a sequence around in the reverse direction. Type I Beta Turns are more common and rely on proline as the second residue. Type II Beta Turns are less common and rely on glycine as the third residue. Gamma Turns are even less common and only consist of three residues, with the hydrogen bond between the first and third residues.
When we talk about tertiary structures, we consider these secondary structures as having their own macro properties. You can look at an alpha helix and see it as a magnetic cylinder. You can look at a beta sheet and see it as a wall or a string. The side groups along these geometric figures can work together to have a combined chemical or physical property: a bunch of aromatic side groups could allow electrons to be delocalized together, allowing for electron transport; a bunch of charged groups could make a massive charge density to pull or push substrates or other folds; a bunch of hydrophilic groups can allow the structure to be exposed to the surface; a bunch of hydrophobic groups can allow the structure to stick to the core. Alternatively, the side groups can be specific here-and-there, especially along binding sites or pockets to guide or bind directly to substrates.
The structure is tied to its function. A protein with more alpha helices may use them to create an induced fit around ligands, carry out allostery, or embed proteins in membranes. A protein with more beta sheets may use them for extracellular stability, strength, or precision. Because most enzymes need some level of precision and ability to use allostery or induced fits, both alpha helices and beta sheets are often needed.
Fibrous Proteins
Before we dive deep into tertiary and quaternary structures, there is a class of proteins that rely almost solely on secondary structures: the fibrous proteins. These include alpha-keratin (skin, fur, nails, horns), beta-keratin (scales, feathers), fibroin (silk), collagen (skin, bone, ligament), and elastin. This discussion aims to highlight the macroscopic properties of what happens when you mass-produce this secondary structure into a crystalline material, linking it with the clothing we wear.
Lehninger Principles of Biochemistry 6th Edition, Page 126
Alpha-keratin is a protein family found in all animals (but has versatile use by mammals) that consists of two alpha helices that wrap around each other like a coil. Think of it like coiling strands of rope together; when they coil, they gain tensile strength and resistance. These coils are held together mostly by non-covalent forces, but they also contain some number of disulfide bridges. The more disulfide bridges, the harder the protein. For example, the alpha-keratin of nails contains far more disulfide bridges than that of fur or hair. The coils can grow as dense fine threads, which can trap air to insulate the mammal. Mammals also produce sebum on alpha-keratin for water resistance.
Beta-keratin is a protein family found in birds and reptiles that consists of packed beta sheets. The beta sheets are spaced at the right distance to trap air in pockets for insulation while excluding water for water resistance. The sheets are held together by infrequent disulfide bridges, making them hard and stiff. For reptiles and birds in arid environments, the beta sheets are cross-linked with lipids to prevent dehydration.
Alpha-keratin and beta-keratin are nothing alike, but they actually gave rise to the historical terms "alpha helix" and "beta sheet". In the 1930s, William Astbury studied the molecular structure of biological fibers, calling them all "keratins", and fired X-rays to study their diffraction patterns. When he fired the X-rays at wool, he found a specific repeating pattern that he called the "alpha pattern", which was again found in other mammalian fibers. When he steamed the wool, he found a different pattern that he called the "beta pattern", which was found in reptilian and bird fibers. Initially, people thought this was two states of the same protein: alpha-keratin and beta-keratin. In 1951, when Linus Pauling deduced the 3D molecular structure of proteins, he discovered that amino acids could use hydrogen bonds to rearrange themselves as either helices or sheets. He found that the X-ray diffraction patterns of the helices experimentally matched that of alpha-keratin, and of the sheets matched that of beta-keratin, giving rise to the names "alpha helix" and "beta sheet". When it came to wool, alpha-keratin is normally structured as alpha helices, but the process of steaming denatures it to rearrange as beta sheets.
In reality, alpha-keratin and beta-keratin are not related at all. Beta-keratin evolved on its own in early sauropsids. Alpha-keratin evolved as part of the Intermediate Filaments (IF) family. When early eukaryotes were building the nucleus to house DNA, IFs first emerged as nuclear lamins to anchor chromatins and help reassemble the nucleus during cell division. When animals evolved multicellularity, they initially relied on actin-myosin networks to link cells, then later reused IFs to build tissues. When animals started walking on land, early alpha-keratin evolved as one type of IF that built up the outermost layer of skin (the stratum corneum) to prevent dehydration. Early mammals experienced another mutation that introduced a high amount of cysteine, which allowed them to play with cysteine content to form various fibers.
Fibroin is the protein found in insects as silk. Like beta-keratin, it also consists of packed beta sheets. But where beta-keratin is tough and brittle, fibroin is flexible. In fibroin, the beta sheets are packed tightly together, full of glycine, serine, and alanine, bonding the sheets together by van der Waals interactions and hydrogen bonding. This makes it poor for insulation and water resistance, but excellent for structure. This explains why a coat made of wool or fur is warm and tough, while a silk dress is designed for comfort and breathability. Arachnids have evolved by convergence a similar protein called spidroin where its crystalline beta sheets are interlaced by amorphous regions of amino acids that grant spider silk high extensibility and strength for catching insects.
Lehninger Principles of Biochemistry 6th Edition, Page 127
As a final example, collagen showcases an example of a protein that is built on neither alpha helices nor beta sheets. Instead, it develops a triple helix structure of three polypeptides looped around each other. Similar to alpha-keratin, this wrapping mechanism grants collagen incredible tensile strength and resistance. Because the amino acids need wider angles to twist and pack together, it relies on glycine, alanine, proline, and 4-hydroxyproline. 4-hydroxyproline is an example of an abundant amino acid that is formed by modifying proline after translation. Collagen is found in bone, inner skin, tendons, and cartilage. Because of its toughness and tight packing, wearing leather is highly durable and good for protection against abrasions, but because it lacks air pockets, it is not a very good insulator and needs to be tanned to resist water well.
Hemoglobin
To showcase the power of tertiary and quaternary structure in a protein, we will look at the first major globular protein humans have extensively studied: hemoglobin. Hemoglobin consists of four subunits: two alpha chains and two beta chains. We will first study the tertiary structure of each of these chains, and then look at the quaternary structure of how these chains interact, and how they perfect the art of oxygen storage and transport.
Lehninger Principles of Biochemistry 6th Edition, Pages 132, 159
Heme proteins are the class of proteins that include myoglobin and hemoglobin. Their structure is characterized by what is known as a globin fold, which only consists of alpha helices and no beta sheets. At the core of the globin fold lies the heme prosthetic group, which is a porphyrin ring with a ferrous iron held in the center. Ferrous iron can form six coordination bonds with its d-orbital electrons: four of them bond to nitrogen atoms found in the pyrrole corners of the porphyrin ring, one bonds to a proximal histidine found below it, and one bonds to the oxygen it aims to store and transport.
Lehninger Principles of Biochemistry 6th Edition, Page 162
Iron is particularly good at donating or accepting electrons without sticking (due to its d-orbital configuration), and because oxygen is so greedy to take electrons, ferrous iron turns out to be perfect for trapping oxygen molecules. However, iron is also really good at rusting. When exposed to both water and oxygen, the ferrous iron will oxidize from to the useless non-reactive . The globin fold hides the heme group behind two amphipathic alpha helices. It forms a hydrophobic barrier that scares water from entering. Furthermore, one of the alpha helices has a distal histidine sticking out as a "cap". When no oxygen is present, the distal histidine forms a distant weak coordination bond with iron that only allows tiny molecules like (and to an extent, ) to squeeze in. When oxygen does get in, it pulls the iron slightly out of the ring, forcing the distal histidine to close and trap the oxygen.
The globin fold highlights what makes alpha helices so powerful: they are massive cylinders with side group bristles that can slot into and slide past each other. The parts of a protein with alpha helices can breathe: jiggling and wobbling that massive cylinder all together. When the iron finds oxygen and gets warped, the proximal histidine gets pulled, which brings the entire helix with it like some kind of lever, which triggers a chain reaction in the rest of the protein. If that histidine were attached to a beta sheet, it would only locally distort the sheet rather than cause the chain reaction you need. In short, alpha helices allow a protein to act like an operating dynamic engine, whereas beta sheets usually form rigid walls and floors.
In a chemical system, thermodynamics dictates the equilibrium of a reaction, and this also applies to a ligand binding to a protein's binding site. Equilibrium is driven by the concentrations of your reactants and your products. If you're in an environment with a lot of oxygen and deoxyhemoglobin, such as your lungs, they will easily fill up. When the hemoglobin travels through your blood and encounters areas of cells with low , entropy drives the release of oxygen without reabsorption. This makes intuitive sense: if you bind really well to oxygen and you're surrounded by it, it's easy to fill up. But when you're no longer surrounded by a ton of oxygen, once you let go of the one you have and it floats away, it's much harder to find a replacement. This is the basics, but hemoglobin does several cool things to make this more efficient.
Allostery is a concept used to describe how a ligand binding to a protein can change its shape, affecting its ability to bind to another ligand. Hemoglobin experiences multiple types of allostery. The most significant way is related to why it consists of four subunits instead of just one (like myoglobin), and highlights how quaternary structure works.
Lehninger Principles of Biochemistry 6th Edition, Pages 141, 165
Hemoglobin has four globin subunits, each with their own heme group, and thus can bind to four . In cooperative binding, a modulator increases the affinity of a ligand. Each acts as a homotropic (i.e. its own) allosteric modulator on hemoglobin through cooperative binding. With each binding to hemoglobin, it is more likely to fill up its other slots for . Likewise, when any leaves hemoglobin, it is more likely to empty its other slots for . Think about it: if deoxyhemoglobin is in a high-oxygen environment like the lungs, entropy will drive a pretty likely chance for at least one oxygen to find the hemoglobin. Once one binds to a subunit, that subunit tries to "wake up" the rest of the protein, signaling it to be more responsive to pick up in the rest of its spots. Likewise, if hemoglobin is full of oxygen and is in a low-oxygen environment, entropy will drive at least one oxygen to leave and not come back, so that subunit will try to "wake up" the rest of the protein and signal that it is probably time to dump the rest of its oxygen.
Mathematically, you can graph non-cooperative ligand binding as a hyperbolic curve: saturation and affinity are highest at low concentrations but gradually flatten out to an asymptote. But cooperative binding is graphed as a sigmoidal curve: saturation and affinity grow faster as a polynomial before tapering off. This is modeled by what is known as the Hill equation, and the Hill plot is used to graph the degree of cooperativity, with the slopes of this plot labeled . If a protein has binding sites for a specific ligand (like ), then measures no cooperativity, measures positive cooperativity, and measures negative cooperativity (rare in proteins). Cooperativity is maxed out by , which represents the strange case where ligands cannot bind alone, and can only bind when all binding sites are ready to be bound simultaneously. This limit is physically impossible, but there are dimeric proteins and ion channels that do get close to this limit. Hemoglobin sits comfortably at , which fine-tunes hemoglobin to be saturated at 98% full in the lungs, drop 25% in resting tissues, and drop up to another 50% in active tissues.
Lehninger Principles of Biochemistry 6th Edition, Page 167
This explains the key distinction between myoglobin in muscles and hemoglobin in blood. Myoglobin is monomeric and thus experiences no cooperativity. It holds onto oxygen when there's a lot of it, and lets it go when there's less of it. This makes myoglobin perfect for storage. Meanwhile, hemoglobin fills up very quickly when there's a lot of it, lets go of it quickly when there's very little of it, and maintains stable behavior when oxygen levels are doing fine. This makes hemoglobin perfect for transport.
Chemically, this is achieved by the quaternary structure of the hemoglobin subunits. Many allosteric proteins, including hemoglobin, can have their stable conformations described as being in a T-state ("tense", low affinity) or R-state ("relaxed", high affinity). When deoxygenated, deoxyhemoglobin is most stable in the T-state conformation. In this conformation, the subunits are held together by salt bridges and hydrogen bonding. As mentioned earlier, when binds to a subunit, it tugs on the heme ring, which pulls on the heme's proximal histidine residue, dragging an alpha helix along with it like a lever, triggering a chain reaction that causes the subunit to slide and rotate against the other subunits. With some of these hydrogen bonds and salt bridges broken, the other subunits themselves are less stable and may rearrange. Eventually, this causes the whole protein to take on the R-state, which is optimally shaped to bind with higher affinity to oxygen.
Originally, there were two proposed models for how this state transformation works. On one hand, you have the concerted model, which proposes that the whole protein is either in the T-state or R-state, and that binding increases the probability of the transition. On the other hand, you have the sequential model, which proposes that each subunit undergoes its own state change. In reality, both models are needed to explain what's happening. Each subunit undergoes its own local conformation, and the whole protein itself has a concerted T-state or R-state conformation. When a subunit gets oxygenated, it undergoes its own conformation and causes strain to the rest of the protein, but it's still "tense": the rest of the protein is too "globally tense" to let that subunit fully "relax". But when the second or third oxygen binds to subunits, the strain is too much, and the protein is more thermodynamically stable in the R-state than the T-state.
Lehninger Principles of Biochemistry 6th Edition, Page 170
The subunits in hemoglobin are not all equal. The subunits are labeled , , , and . and are the same polypeptide, and likewise for and . When hemoglobin forms, an subunit and subunit get paired as a dimer (giving us and ), held together by hydrophobic forces. Then these dimers pair up into a tetramer, using hydrogen bonds and salt bridges to link and . The reason why the first oxygen doesn't trigger a state change is because it only adds strain to the subunit it's paired with, but not the other dimer. So if is oxygenated, it adds strain to , but doesn't relieve any stress from the dimer. When a second subunit gets oxygenated, it adds its own local strain. If were oxygenated, no state change happens because the dimer was already locally stressed, and holds together the global tension. But if or were oxygenated, suddenly the structural integrity of the T-state collapses, and the salt bridges shear apart, refolding into the R-state.
This superpower behind conformations is again rooted in alpha helices. Proteins are constantly jiggling and breathing, refolding between their thermodynamically stable conformations. Even in deoxyhemoglobin, the protein may spontaneously refold for a brief moment to the R-state. But the energy barrier is so high that it instantly caves back to the more stable T-state. When the first oxygen binds, it's slightly more stable when it accidentally refolds to the R-state, but not very much, so . When the second oxygen binds, it has a 2/3 chance of binding to one of the subunits on the other dimer, so . When the third oxygen hits, it's guaranteed to oxygenate a subunit on the other dimer, so .
Hemoglobin employs three other types of allostery. Two of them are linked to what is known as the Bohr effect: the binding of and to hemoglobin lowers its affinity for . This makes sense. In cellular respiration, mitochondria convert to . In red blood cells, which are mostly hemoglobin, about 20% of the protein content is carbonic anhydrase, which is an enzyme that converts into and . So if or levels are high, that indicates that surrounding cells are starved of oxygen. Carbonic anhydrase is valuable because isn't very soluble in blood plasma on its own, but is. So carbonic anhydrase converts to and when it picks it up, then converts them back into to be expelled from the lungs.
Lehninger Principles of Biochemistry 6th Edition, Page 170
Chemically, the Bohr effect describes how and modulate the activity of hemoglobin. is so small that it wedges itself among the hydrogen bonds and salt bridges and helps stabilize the T-state. actually binds to the amino end of the subunits, spitting out an . Hemoglobin is responsible for carrying up to 40% of and 20% of to the lungs and kidneys. The remaining is transported as soluble or dissolved in the blood plasma, and the remaining is used by the bicarbonate buffer.
The third allostery is by a molecule called BPG. BPG is a heterotropic allosteric modulator that also reduces hemoglobin's affinity to oxygen. It has a negative charge which interacts with a positively charged cavity in the subunits, stabilizing the T-state. It's primarily used to maintain homeostasis of the absorption and release of to cells when atmospheric varies, such as when you change elevation. It's also used locally within the body when it detects local tissue hypoxia.
Lehninger Principles of Biochemistry 6th Edition, Page 172
When acclimatizing to a higher elevation, the body immediately responds by hyperventilating to increase breath rate and heart rate. But this increases concentration, increasing blood alkalinity. After a few hours, the body uses BPG to modulate oxygen delivery and fix this issue. After a few days, the body begins making new red blood cells. After a few weeks or months, the body remodels itself by building more capillary networks, mitochondria, and myoglobin.
As a final aside regarding hemoglobin, there is one neat insight about proteins and genetics that can be learned from sickle-cell anemia. A gene mutation replaces a glutamic acid with a hydrophobic valine. In the T-state, this hydrophobic pocket is exposed. As mentioned earlier, these hydrophobic pockets can stick to each other, aggregating into a tubular fiber, distorting the red blood cell into a sickle shape. This is usually awful, rendering these sickle cells fairly useless, leaving the person deprived of oxygen. However, this gene is slightly favored in certain areas of Africa where malaria is common. It turns out to provide minor resistance to malaria. Malaria is caused by a parasitic eukaryote that feasts on hemoglobin. The immune system is inefficient when dealing with threats in the blood plasma itself, so it struggles to deal with malaria itself. But for those with sickle-cell anemia, when malaria feasts on hemoglobin, the sickle hemoglobin clumps together quickly. Those malaria-infected sickle cells get sent to the pancreas for destruction. This example highlights how a single amino acid mutation is enough to change the physical properties of hemoglobin that allows for an evolutionary trade-off.
Lehninger Principles of Biochemistry 6th Edition, Page 173
Ligand Binding
The example of hemoglobin bridges the gap between protein structure and function, and highlights how conformation state changes happen, and how modulators can regulate its behavior. Before we go into enzymes, which are proteins that carry out chemical reactions, I want to talk a little further about ligand binding. Each paragraph will briefly mention or review a different trait about how ligands bind to proteins.
Lehninger Principles of Biochemistry 6th Edition, Page 192
The binding sites of proteins are engineered to be specific and complementary to the intended ligand in size, shape, charge, and hydrophilicity. If your ligand is small, your binding site needs to be small. If your ligand has a negative charge, your binding site should have positive charges. If your ligand is hydrophobic, you may need to tuck your binding site away from the solvation layer, and use hydrophobic residues to stick to it via van der Waals interactions. The ligand is usually bulky enough that different areas of the ligand molecule will have different chemical properties, and the binding site can be shaped to complement each area. For example, if your ligand consists of an aromatic ring with two hydrophobic tails, your binding site may have two hydrophobic pockets to stabilize the tails, then use something like pi-stacking to stabilize the ring. This idea is known as the "lock-and-key" model. It allows the protein to use its residues to only bind to its intended ligands. Sometimes, it's just one. Sometimes, it's multiple, but it binds to some better than others (whether intentionally or not). We will see examples of the hyper-specificity of binding sites with chymotrypsin. In pharmacology, we hack this system by synthesizing unnatural ligands that can bind much better than the intended ligands, hijacking our biochemical pathways to treat symptoms.
One thing I was always confused about was why proteins had to be so large when most of them only serve to bind to a few ligands at local binding sites. It helps to think of proteins as engines that are engineered at micro-precision. In order for ligands to bind to binding sites with fine-tuned affinity and chemical reactivity, the proteins need to be sufficiently large enough to hold the binding sites in stable positions that allow them to interact with ligands as intended. If the protein were too small, it would be wobbly and janky. The ligands may not attach. They might fall out too easily. Other ligands might latch on that shouldn't. Enzymes might not be able to stabilize their intermediate states to carry out chemical reactions. Proteins need to be optimized for their environment to be most stable. If the pH is off, or interact with the binding site residues and can affect their charges, disrupting their ability to bind to their ligand.
The induced fit refers to a conformational adaptation by the protein in response to ligand binding. We saw this with hemoglobin where oxygenating the heme ring pulls in the distal histidine to serve as a "cap" and close the binding site, trapping the oxygen. A protein may undergo an induced fit to hold onto its ligand tightly. This is especially important for enzymes, which use their binding site to entice and ensnare their substrate. Once a substrate binds to an enzyme, its induced fit reshapes the binding site to prepare it for catalysis. When it undergoes a reaction, it may undergo another induced fit to carry out the next step of the reaction or to open the active site to release its products.
Allostery refers to the way that ligand binding affects properties of other binding sites. A homotropic modulator is one where the ligand modulates binding of itself in other binding sites. A heterotropic modulator is one where the ligand modulates binding of other types of ligands. In hemoglobin, we saw oxygen serve as a homotropic modulator in cooperative binding, and , , and BPG serve as heterotropic modulators of oxygen. Allosteric modulators generally work by either refolding the protein into an active or inactive state (as we saw with hemoglobin), or by closing or widening the binding site entrance.
Generated by Gemini
Binding sites are usually tucked away inside a tunnel in a protein. In order for the ligand to find its binding site, the protein uses residues to steer the ligand into and down the pocket. It may use charge from charged residues or the dipole of an alpha helix to attract and push the ligand. The tunnel may consist of alpha helices or beta sheets with hydrophobic or hydrophilic channels to slide the ligand along. The tunnel may also have gates that breathe open and close to allow ligands to pass through. When the ligand reaches its binding site, it's docked with the stabilizing forces we mentioned above, and an induced fit may latch onto it tighter.
When we study the thermodynamics and kinetics of ligand binding, we model the system as concentrations of ligands, proteins, and the ligand-protein complex. As such, we can model an energy curve that shows the activation energy required for this complex to form, or to separate into the ligand and protein. The association constant expresses the equilibrium of ligands binding to proteins, measuring ligand affinity. The dissociation constant measures the concentration at which half the ligand binding sites are saturated. When we discuss enzymes, we must consider this as a multi-step process, with different activation energies to jump between each step of the process.
These topics can be generalized to many types of globular proteins, not just enzymes. However, enzymes often take these topics much further due to the complications revolving around not just latching onto a ligand, but also chemically changing it.
Enzymes
An enzyme is a protein that carries out chemical reactions through catalysis. The binding site of an enzyme is called its active site, and its ligand is called its substrate. Enzymes often include a cofactor used to help catalyze the reaction. If the cofactor is tightly bound (potentially covalently) to the enzyme, it is known as a prosthetic group. For example, heme is a prosthetic group of hemoglobin. A cofactor can be either a metal ion or an organic coenzyme. Metal ion cofactors participate in catalysis, whereas organic coenzymes serve to carry ions or functional groups. Organic coenzymes are usually derived from vitamins, but there are also other metabolic coenzymes like ATP (to transfer a phosphate group), SAM (to transfer a methyl group), lipoic acid (to transfer an acetyl group), and glutathione (to neutralize free radicals).
To explain how an enzyme performs the magic of a chemical reaction, we have to talk about what they aim to solve. Chemical reactions happen when a delocalized electron of one molecule, in its many quantum states, crashes into another molecule, momentarily converts its heat and kinetic energy into a more energetic transition state, then falls back down to a different configuration than what it started with, but is happy and stable in that other configuration. This is far from the whole story and ignores a lot of scenarios but puts it very simply. In ordinary chemistry, reactants randomly collide in the perfect orientation that encourages this predictable reconfiguration to occur. In most cases, several electrons need to be reconfigured across reactants in sequential steps for the products to be revealed. Every reaction is theoretically reversible: if the reactants can be turned into products, products can be turned back into reactants. Therefore, an enzyme not only catalyzes the forward reaction but also the reverse reaction. If an enzyme's goal is to push a reaction in a direction, it must rely on a number of tricks to pull it off.
Klein Organic Chemistry 3rd Edition, Pages 278, 281
Chemical reactions can be described by both thermodynamics and kinetics. Given certain conditions, thermodynamics can be used to derive an energy curve for a reaction, showcasing how enthalpy and entropy can be used to show the direction of a reaction given initial concentrations of reactants and products, telling us the concentrations at equilibrium, whether the reaction is spontaneous, and how much free energy is gained or lost to the system following a reaction. Excluding some minor nuances, enzymes don't mess with any of this. Instead, enzymes work by shifting the kinetics of a reaction.
Lehninger Principles of Biochemistry 6th Edition, Page 192
Kinetics relate to the speed of a reaction. Without a catalyst, a reaction might be thermodynamically favorable, but it may take a very long time to occur by chance. This is dictated by the activation energy of a reaction. As mentioned above, an electron is happy and stable as a reactant or product. Momentarily though, an electron uses excess energy to attain some super energetic and unstable "transition state". This transition state allows the electron a moment to leave its happy stable bond to explore elsewhere, which allows it to discover a different bond it decides to fall into and form. It's similar to the energy funnel of protein folding we discussed much earlier, where the protein may use heat and kinetic energy to momentarily reach "transition states" between its stable conformations. In chemical reactions, this may look like certain bonds stretching further apart or closing together, eventually snapping apart old bonds or gluing together new ones between new pairs of atoms.
This transition state has an activation energy associated with it. The activation energy is associated with both the enthalpic and entropic conditions to get to that transition state. On enthalpy, the molecules need the ability for their electrons to convert excess energy into the potential energy needed to stretch or form those bonds. On entropy, the molecules need to position themselves at the correct orientation and distance for the intended electrons to interact. The secret trick to enzyme catalysis is how it minimizes both the enthalpy and entropy requirements to attain this transition state.
Lehninger Principles of Biochemistry 6th Edition, Page 193
When you lower the activation energy cost to attain this transition state, suddenly your reactions do not need that much energy to carry out their reactions. Imagine you have a solution of reactants. If the activation energy is too high, your two colliding reactants need a ton of excess energy to cross that gap. In a Boltzmann Distribution of molecules, only a certain number of molecules will pass this energy requirement. But if you lower your harsh energy requirement, suddenly a lot more reactants are passing the grade and are allowed to attain the transition state needed to become products. If you assume there is some number of collisions within a given interval, some percentage of them will pass and attain the transition state to return products. If you increase the likelihood of the collision turning reactants into products, you increase the number of reactions that occur in that fixed interval. Therefore, lowering the activation energy increases the rate of a reaction.
The Arrhenius Equation in statistical mechanics relates the rate constant of a reaction to its activation barrier, and is derived by taking the integral of the Boltzmann curve from to . The speed of a reaction at a given time scales linearly with the rate constant. So if you linearly lower the activation energy, you exponentially increase the rate constant and thus the speed of a reaction over time. This is why enzymes don't just increase a reaction by a small factor: they increase rates of reactions by many orders of magnitude. Most enzymes increase the rate of reactions by 10 to 17 orders of magnitude, with some enzymes as theoretically high as 26 orders of magnitude. This theoretical limit is actually attained by OMP decarboxylase for DNA synthesis: it can catalyze a reaction that would normally take 78 million years to spontaneously happen, to only take 18 milliseconds. Incredible.
Collision Theory and Maxwell-Boltzmann Distribution Curves
Now that we've shown it's possible, let's talk about how proteins pull this off. The secret ingredient lies at the heart of specificity. The specificity of a substrate to its active site goes beyond the simple lock-and-key model. The active site of an enzyme is not shaped to look like the substrate: it's shaped to look like its transition state! Imagine you wake up on a Monday morning and you have to think of what you're going to wear, what you're going to eat, and how you're going to drive to work. But one Monday morning, someone picks out your clothes for you, makes your coffee and breakfast, and drives you to work. Someone has done all that extra work for you to get your day started. When an enzyme's active site looks like the transition state, it makes that transition state more appealing to fall into without the high energy cost.
Lehninger's Biochemistry book offers a beautiful example of this with a metal stick analogy. Imagine you're trying to break a metal stick in half. Without an enzyme, you exert all your energy into bending the stick in half (its transition state) before it finally snaps. Now imagine you have an enzyme where the active site is lined with magnets. If the active site looked like the straight stick, the stick would bind really well to the enzyme, but is terrible for breaking it apart. But if the active site were in a bent shape to complement the shape of the transition state, the unbent stick still binds ok to the active site, but the important point is that suddenly you don't have to put so much energy into breaking the stick - the magnets will help bend it so you need much less energy to break it in half.
Lehninger Principles of Biochemistry 6th Edition, Page 196
In practice, the magnets are analogous to the stabilizing forces provided by the active site's residues. The residues use non-covalent weak forces to convert the heat and kinetic energy of the protein and reactant into stabilizing potential binding energy that serves as an "alternative route" for the substrate to attain its transition state without the crazy spontaneous stunts it would need to pull in the wild. In short, enzymes use some weak forces to attract the substrate to its binding site, use most of their weak forces to stabilize its transition state, then use repelling forces to push out products to prepare for the next reaction.
The weak interaction binding energy provides the free energy for electron redistribution, but there are other factors at play to reduce the activation energy. A large part of entropy is arranging the reactants at the right distance and orientation. If your binding sites force the reactants in the optimal arrangement, they won't have much trouble fitting together. Another factor is regarding de-solvation. In a normal solution, reactants roam in a solution full of water that loves to hydrogen bond and distract reactants. If an active site excludes water from entering, you eliminate those distractions, allowing the reactants to focus their attention on each other. Third, I have to bring up induced fitting again. When a substrate binds to an active site, the protein may undergo an induced fit that can provide the mechanical, chemical, or electric forces needed to further push the substrate into its transition state or push out the products.
I want to briefly return to kinetics again. Because catalyzed reactions include multiple steps, there are multiple activation energies. Because activation energy has an exponential relationship with the speed of the reaction, the step with the greatest activation energy is called the rate-limiting step, although there may be steps with near equal activation energies and thus are all partial bottlenecks. This rate-limiting step helps us approximate the rate of the overall reaction. The way to measure the maximum velocity of a catalyzed reaction is to measure the change in concentrations at the very start of a reaction: when you have only reactants and no products. Otherwise, the reverse catalyzed reaction would make your reaction appear slower over time as you build more products.
Michaelis-Menten kinetics is a common model for predicting the rate of many enzyme-catalyzed reactions and is based on the steady-state assumption: that an enzyme and substrate reversibly binding is much more common than the formation of reactants to products. This applies to many simple single-substrate non-allosteric enzymes. For a given reaction, we can experimentally attain the Michaelis constant to measure the affinity of an enzyme to its substrate, the turnover number to measure how many reactants convert to products when saturated, and the specificity constant to measure how efficient an enzyme is. This is capped by the diffusion-controlled limit of that represents enzymes that catalyze reactions as soon as substrates are able to diffuse to the active sites.
There are three main types of catalysis employed by enzymes. Enzymes may use some combination of the three.
- Acid-base Catalysis: The enzyme creates a charged intermediate in the substrate but uses proton transfers with an acid or base to stabilize the charge. When water is used as a weak acid or base, it is known as specific acid-base catalysis. When another substrate or one of its binding site residues is used as a weak acid or base, it is known as general acid-base catalysis.
- Covalent Catalysis: The binding site residues may act as nucleophiles to form covalent bonds with the substrate.
- Metal Ion Catalysis: Metal ions are used to orient substrates, stabilize charged states, or mediate oxidation-reduction reactions.
Metal ion catalysis is used in around 1/3 of all enzymes and relies on the properties of the special metal ion's atomic radius, redox state, stable positive charge, and its d-orbital coordination. For example, iron can swap oxidation states between and very easily with a low energy gap, and so can transport electrons easily and can easily carry out redox reactions. Copper is similar but is more electronegative and can undergo what is known as the Jahn-Teller effect to distort two of its coordination bonds. is really useful in hydrolytic enzymes because its full d-orbital shell means that it's redox-inert and can serve as a safe Lewis acid, allowing it to pull enough electron density from water for it to tear into and . Other metal ion cofactors include , , , , and .
Enzymes can be classified into one of the following categories:
- Oxidoreductase: Can transfer electrons from a donor to an acceptor.
- Transferase: Can transfer functional groups from one molecule to another.
- Hydrolase: Can perform hydrolysis to transfer functional groups to water. This includes all proteases which are used to break peptide bonds.
- Lyase: Can cleave , , and bonds by elimination to produce double bonds or rings.
- Isomerase: Can transfer functional groups to yield isomers.
- Ligase: Can form , , , and bonds by condensation reactions, usually coupled with ATP cleavage.
Now let's talk about methods of regulation. We refer to regulatory enzymes as those that exhibit changes in response to signals. A regulatory site on an enzyme is frequently found on a separate subunit of a protein, and there are some proteins that consist of many regulatory subunits. Enzymes may be allosteric just as ordinary proteins. Regulatory proteins exist as modulators themselves by wrapping around and binding to other proteins. Proteolytic cleavage refers to a process where a peptide segment of an enzyme is irreversibly cleaved for activation or deactivation.
Inhibition describes substrates that slow or halt an enzyme's activity. A competitive inhibitor competes with a substrate for the active site but doesn't itself react. An uncompetitive inhibitor binds to an allosteric site on the enzyme-substrate complex that pauses the enzyme from turning the reactants into products. A mixed inhibitor is a type of uncompetitive inhibitor that can halt action whether the enzyme is bound to a substrate or not. An irreversible inhibitor can inactivate an enzyme by either covalently bonding to the enzyme or resembling the transition state substrate so that it latches on tighter than the substrate itself.
Lehninger Principles of Biochemistry 6th Edition, Page 208
The most common type of regulation is through reversible covalent modifications. Examples include phosphorylation (from ATP), adenylylation (from ATP), acetylation of the amino end (from Acetyl-CoA), methylation, and ubiquitination (tags for proteolytic degradation). Phosphorylation is the most common of these. Protein kinases are transferases designed to transfer phosphoryl groups from ATP to an enzyme (usually on a serine, histidine, threonine, or tyrosine). Phosphoryl groups are bulky and negatively charged, which allows them to interact with multiple residues at once by attracting positive charges, repelling negative charges, or hydrogen bonding with polar residues. Kinases use regex-style residue sequences to recognize where to phosphorylate. Enzymes can be modified by several phosphoryl groups which can serve to activate or deactivate the enzyme.
Lehninger Principles of Biochemistry 6th Edition, Page 230
One of my favorite types of kinases is AMPK. AMPK measures how much free energy is in a cell and can turn on and off enzymes to control that energy usage. It has four binding sites for adenine molecules as allosteric modulators. When ATP is high, there is enough energy, so AMPK is auto-inhibited - it closes its active site. When AMP is high, there is low energy, so AMPK gets activated and gets ready to phosphorylate certain enzymes. AMPK binds to a specific residue sequence that tends to be associated with activation for energy-producing enzymes, and de-activation for energy-consuming enzymes. Energy-consuming enzymes have this sequence near their active sites, so phosphorylating here would close its active site. Energy-producing enzymes have this sequence near a regulatory hinge, so phosphorylating here would refold the protein into its active state. It's a great example of how a kinase is capable of both activating and de-activating multiple enzymes. It allows us to view certain sequences on proteins as special handshakes to communicate with other proteins, such as "uh oh, we're running low on energy."
Chymotrypsin
Just like how hemoglobin taught us valuable lessons about protein structure and role, chymotrypsin can teach us valuable lessons about how proteins can carry out lightning-speed chemical reactions like an assembly line one after another. We can also use this opportunity to compare chymotrypsin's structure to that of hemoglobin and how they're each optimized for their roles.
Proteases are a class of hydrolases specialized to cleave peptide bonds. Serine proteases, which include chymotrypsin and trypsin, and cysteine proteases work by using serine / cysteine as a nucleophile to covalently bond with the residue it aims to cleave. Aspartyl proteases, which include pepsin, cathepsin, and renin, use two aspartate residues to hydrolyze the peptide bond via acid-base catalysis. Metalloproteases use a metal ion to cleave the peptide bond.
Chymotrypsin offers an excellent example of not only how an enzyme operates, but how it selectively chooses its substrates. There are several proteases that cleave the peptide bond, but they have different requirements to make that happen. Chymotrypsin's binding site has a snug hydrophobic pocket that allows residues with aromatic rings (Trp, Phe, and Tyr) to fit tightly. When a denatured polypeptide rolls around the pancreas, it stumbles and slides across chymotrypsin. Residues along the peptide may try to fall into chymotrypsin's active site, but they'll either bounce off or fall out. But when one of Trp, Phe, or Tyr falls in, it gets locked in with high affinity, long enough for the minute-precision of the enzyme to carry out its reactions.
If we took a sneak preview inside the active site, we could spot a few landmarks that contribute to its binding and catalysis. First, we have that hydrophobic pocket for an aromatic ring to fit in to be deep, held by van der Waals interactions. Second, we can spot the oxyanion site, whose role it is to stabilize the negative charge on the residue's carbonyl oxygen during one of the intermediate steps of catalysis. This consists of glycine and histidine, which hydrogen bond with the negative oxygen to lower its reactivity so it only bonds with what it's supposed to. Third, we have the catalytic triad: a set of three residues that carry out a combination of both acid-base catalysis and covalent catalysis. In this case, our three residues are serine (the nucleophile), histidine (the base), and aspartate (the acid).
Lehninger Principles of Biochemistry 6th Edition, Page 216
The catalytic triad refers to a common motif found among many enzymes in biology by convergent evolution due to the strict constraints necessary for efficient catalysis. It involves three residues: one acts as a nucleophile, one acts as a base, and one acts as an acid. The goal is to create a super strong nucleophile that will attack the most electrophilic atom it's near, covalently bonding with the substrate. The active site allows a single water molecule to squeeze in and hydrolyze that covalent bond, releasing the product. This process is highly efficient at transferring functional groups, and so is found in many hydrolases and transferases. In the triad, the nucleophile attacks the substrate; the base takes protons away from the nucleophile to make it sufficiently reactive; the acid stabilizes the base's positive charge via hydrogen bonding or electrostatic attraction.
Let's tie this all together. The peptide snaps into chymotrypsin: the aromatic ring is in its pocket, the carbonyl oxygen is in the oxyanion hole, and the peptide bond is exposed near serine (the nucleophile) and histidine (the base), with an aspartate (the acid) behind the histidine's ring. The first half of the process is acylation. The peptide bond's carbonyl carbon is electrophilic, and so serine dumps a proton to histidine so it can use its oxygen tail to covalently bond with the carbonyl carbon, creating a carboxylic group, and pushing the extra negative charge to the oxyanion hole where it's stabilized. The adjacent residue's amino carbon is now in a close position where it's happy to easily take both the base's extra proton and that oxyanion hole's extra electron and break that peptide bond, snapping off that half of the peptide.
Now we need to release the other half of the peptide through deacylation. With the other peptide half gone, the active site opens up just enough for a single water molecule to squeeze through and find itself between the serine, basic histidine, and other peptide bond. The histidine wants to steal one of the water's hydrogen atoms, and so pulls on it with a strong hydrogen bond, making the water's oxygen highly nucleophilic to the point where water dumps its proton to histidine and uses its hydroxyl group to attack the carbon, pushing its extra electron again to the oxyanion hole. Now this carbon has three oxygens attached (one with a negative charge) and is unhappy. Serine is happy to take the histidine's extra proton and the oxyanion's extra electron to push away the other peptide half. At this point, the other peptide half flies away and the reaction is complete.
Lehninger Principles of Biochemistry 6th Edition, Pages 216-217
This reaction took place in two major steps which were nearly identical, known in organic chemistry as an addition-elimination mechanism. In both steps, a nucleophile was so desperate to bond that it dumped a proton somewhere stable and dumped an electron somewhere else stable so that it could form a covalent bond. This covalent bond made the central carbon unstable, and so one of the carbon's other two functional groups was more stable stealing that extra hydrogen and extra electron to leave. In the first step of the reaction, it was serine that was desperate to bond, and the amino group of the peptide bond that wanted to leave. In the second step of the reaction, it was water that was desperate to join the party, but the serine wanted to go home. (Or more figuratively, the serine wanted the party to end and kicked everyone else out.) If you look at the system this way, you might be able to see how you could generalize this to make it easy to transfer most functional groups from one substrate to another. In this case, water transferred its hydroxyl group to the broken peptide so it was happy to leave.
At its core, chymotrypsin showcases how enzymes can employ every trick in the field of organic chemistry to break down reactions into a series of inconsequential steps. It hacks into acid-base chemistry by positioning reactants such that it controls who can donate or receive protons, deciding which of the carbonyl carbon's bonding partners would be the better leaving group. The nucleophilic acyl substitution is a type of addition-elimination mechanism, which is fairly common.
If you flip through an organic chemistry textbook, you'll find biological equivalents of each reaction: redox, , , , radical, EAS, pericyclic, and condensation... Redox reactions are the most common, found in a third of all enzymes as oxidoreductases, where electron flow is essential for respiration, photosynthesis, and fatty acid oxidation. Addition-elimination reactions are primarily used by hydrolases, useful for breaking amide (including peptide) and ester bonds. are very common for transferases, such as by methyltransferases and kinases. It's much more common than because there is no carbocation intermediate to deal with, so is reserved for a few essential pathways (such as terpene synthesis). is the default elimination mechanism by enzymes used by lyases, favored over or which require strong bases, strong acids, or high heat. Condensation reactions (such as aldol and Claisen) are often seen in metabolism (glycolysis, Krebs cycle, fatty acid synthesis). Radical reactions are rare due to the danger of ROS, reserved for modifying un-reactive bonds. Aromatic reactions like EAS and pericyclic are very rare in biology because aromatic rings are inherently stable.
As a last aside, I want to compare chymotrypsin's structure to hemoglobin. About 3/4 of hemoglobin's residues are in alpha helices, and none in beta sheets. In comparison, about 1/3 of chymotrypsin's residues are in beta sheets, and only 1/9 in alpha helices. Hemoglobin is built to be flexible and resist strain as it holds onto and runs with : optimized for transport and storage. In comparison, chymotrypsin is a digestive enzyme found in the pancreas: surrounded by other proteases and acids that serve to break down proteins. Chymotrypsin consists of anti-parallel beta sheets wrapped around each other in a -barrel structure, strong and rigid enough to act as armor, hiding its essential backbone underneath while fending off surrounding threats. The beta sheets are also needed for high precision: the hydrophobic pocket, oxyanion hole, and catalytic triad need to be at the exact positions to carry out their reactions. Alpha helices are too bouncy and would displace these parts from each other. You also see alpha helices used to carry out regulation or induced fit, neither of which chymotrypsin really needs.
Lehninger Principles of Biochemistry 6th Edition, Page 214
So what are the few alpha helices in chymotrypsin doing? Alpha helices fold much faster than beta sheets so can act as nucleation sites to guide protein folding of the slower beta sheet formation. One massive alpha helix, in particular, slots itself near the C-terminal to hold that end together. We also see short alpha helices act like hinges between other domains of the protein. These alpha helices don't contribute to the reaction nor help with binding affinity - they serve to build and hold together the protein like the bolts and rods of an engine.
Conclusion
If you want to understand how life works, you have to understand the key traits: how to encode information, how to perform chemical and physical mechanisms, how to replicate, how to extract and use energy, how to protect against a hostile environment. If we abstract these questions philosophically to "what", "how", and "why", we find proteins are usually the answer to "how". The "what" and "why" can mostly be derived from mathematical and physical laws, and the environmental conditions of earth and our solar system. But a lot of the "how" is proteins.
If we are to understand "how" we got to this point as human beings, we must understand where proteins come from and why they're so efficient. No man will ever be able to tell the whole story of life and the universe, but each of us can grow to understand and accept it in our own way, leaving us satiated. I will continue to study and tell the story of our universe, piece by piece, and I hope that curiosity never dies in our species. The most beautiful aspect of life is that it can give rise to those capable of admiring the fruit of its work.
I wrote this post because I realized halfway through that the study of proteins was so extensive that I couldn't afford to forget it as soon as I finished these chapters, nor did I ever want to forget my interpretation, as it's so foundational to my goal in life to understand the universe and share that love with others. I hope this post has been informational to other readers. I also hope that this post will help me in the future build upon it in my future projects that aim to share knowledge and wisdom.