Key Takeaways
Proteins in natural biology are built from 20 canonical amino acids, which historically limited the chemical diversity available for protein engineering and biologic drug design.
Genetic code expansion enables the incorporation of noncanonical amino acids into proteins by using engineered orthogonal aminoacyl-tRNA synthetase and transfer RNA systems during translation.
Strategies such as stop codon reassignment, quadruplet codons, and sense-codon compression create additional coding capacity that allows new amino acids to be encoded in proteins.
Noncanonical amino acids can introduce bioorthogonal functional groups and other synthetic chemistries into proteins, enabling new molecular interactions and chemical reactivity.
These capabilities support emerging therapeutic applications, such as site-specific antibody–drug conjugates, cytokine engineering, and other chemically programmable protein therapeutics.
As production platforms mature, expanded genetic codes may transform proteins from molecules constrained by natural evolution into programmable chemical scaffolds for next-generation biologic drugs.
The Limits of the Natural Genetic Code
All proteins produced by living organisms are assembled according to the rules of the genetic code, which specifies how nucleotide triplets in messenger RNA correspond to amino acids during translation. In natural biology, this system encodes a fixed set of 20 canonical amino acids that serve as the building blocks of proteins. These amino acids provide the functional groups, structural features, and chemical interactions that determine how proteins fold, interact with other molecules, and carry out biological functions.
The chemistry available within these 20 residues is remarkably versatile. Through different combinations and arrangements, proteins can form stable structures, catalyze reactions, bind to specific molecular targets, and regulate cellular processes. This versatility has allowed natural evolution to produce an extraordinary diversity of biological functions from a relatively small chemical toolkit.
At the same time, the fixed nature of the genetic code imposes an inherent constraint on protein design. The side chains present in the canonical amino acids define the range of chemical interactions proteins can perform. Although protein engineers can alter sequences, redesign structures, and introduce mutations to improve stability or activity, these efforts remain confined to the same limited palette of chemical functionalities that evolution has used for billions of years.
This limitation has long shaped the field of protein engineering and biologics development. Whether designing enzymes, therapeutic antibodies, or other protein-based medicines, researchers have traditionally been restricted to modifying combinations of the same natural building blocks. Expanding the chemical diversity available to proteins presents a fundamental opportunity to broaden the capabilities of engineered biologics.
Advances in genetic code expansion (GCE) are beginning to address this constraint by enabling the incorporation of noncanonical amino acids into proteins during translation. These approaches introduce entirely new chemical groups into protein structures, allowing researchers to move beyond the chemistry defined by the natural genetic code. The result is the possibility of designing proteins with capabilities that extend beyond those accessible through the canonical amino acid set.
This raises a provocative question for the future of protein engineering and therapeutic development: What might become possible if proteins could be constructed from a far larger chemical vocabulary than the one provided by nature?
Expanding the Genetic Code
Efforts to move beyond the constraints of the natural amino acid repertoire have led to the development of GCE, a set of molecular engineering strategies that allow additional amino acids to be incorporated into proteins during translation. Rather than modifying proteins after they are produced, GCE alters the translation process itself so that synthetic amino acids can be inserted directly into a growing polypeptide chain at defined positions.1
The central challenge in this approach lies in integrating new amino acids into the highly conserved translation machinery of the cell without disrupting its normal operation. GCE addresses this problem through the introduction of orthogonal translation components, most notably engineered aminoacyl-tRNA synthetase and transfer RNA pairs that function independently of the host’s native protein synthesis system. These orthogonal aminoacyl-tRNA synthetase (aaRS) and tRNA pairs are designed to recognize a specific noncanonical amino acid and attach it to a corresponding tRNA without interacting with the host cell’s natural amino acids or translation components.2,3
Once charged with the noncanonical amino acid, the engineered tRNA can deliver that residue to the ribosome during translation when a designated codon is encountered. Because the orthogonal aaRS–tRNA system operates separately from the cell’s endogenous machinery, it allows the synthetic amino acid to be incorporated at specific sites in a protein while leaving the rest of the translation process unchanged.1
In practical terms, this approach introduces new amino acids into the genetic code by assigning them to codons that would otherwise be unused or reassigned. The resulting system effectively expands the number of chemical building blocks available to proteins beyond the canonical set encoded by natural biology.4
Through these strategies, researchers have demonstrated the incorporation of many different noncanonical amino acids into proteins across multiple expression systems. The resulting proteins can contain functional groups, chemical reactivity, and structural features not present in the natural amino acid repertoire, allowing engineered proteins to access forms of chemistry that biological evolution never encoded.5
Engineering New Codons for New Chemistry
Introducing new amino acids into proteins requires more than engineered translation machinery. It also requires additional coding capacity within the genetic code itself. Because the standard genetic code already assigns nearly all codons to the canonical amino acids or translation signals, GCE strategies must create space for new amino acid assignments. Researchers have developed several approaches to accomplish this, effectively rewriting parts of the genetic code so that specific codons can encode noncanonical amino acids.6
One widely used strategy is stop codon reassignment. In this approach, a codon that normally signals the termination of translation is repurposed to encode a synthetic amino acid. The engineered orthogonal aminoacyl-tRNA synthetase and transfer RNA pair recognize the reassigned codon and deliver the corresponding noncanonical amino acid during translation. Because stop codons occur relatively infrequently in coding sequences, they provide a convenient entry point for expanding the code without disrupting the translation of most proteins.
A second approach introduces quadruplet codons. Instead of the traditional three-nucleotide codons used by the natural genetic code, quadruplet codons consist of four nucleotides. These additional combinations dramatically increase the number of possible codons that can be assigned to amino acids. Engineered tRNAs capable of reading these four-base codons allow ribosomes to incorporate noncanonical amino acids at positions specified by these newly created codons.
Researchers have also explored sense-codon compression strategies. In this framework, certain codons that normally encode canonical amino acids are reassigned after redesigning the genome or expression system so that those codons are no longer required for their original function. Once freed, the codons can be repurposed to encode synthetic amino acids through orthogonal translation systems.
These approaches create new coding capacity within the genetic code. Rather than treating the genetic code as a fixed property of biology, GCE treats it as an engineering platform that can be modified and extended. By generating additional codons that encode synthetic amino acids, researchers can design proteins that incorporate chemical functionalities far beyond those available in the natural amino acid repertoire.
Expanding the Chemical Capabilities of Proteins
The primary value of noncanonical amino acids (ncAAs) lies in the new chemistry they introduce into proteins. While the canonical amino acids provide a diverse set of side-chain functionalities, their chemical properties remain constrained to those selected through biological evolution. GCE allows synthetic residues with entirely new functional groups to be embedded directly into proteins during translation, extending the range of chemical interactions proteins can perform.4
Many ncAAs are designed to carry reactive or bioorthogonal functional groups that are absent from natural amino acids. These may include azides, alkynes, photocrosslinkers, or other chemically reactive moieties that enable highly selective reactions under biological conditions. Because these groups can be incorporated at precise locations within a protein sequence, they create defined chemical handles that can participate in reactions unavailable to natural proteins.4,7
The presence of these synthetic functionalities expands the molecular toolkit available to protein engineers. Proteins can be constructed with specific sites for chemical modification, designed catalytic environments, or tailored interaction surfaces that depend on chemistries not found in the canonical amino acid set. In some cases, ncAAs also alter the physicochemical properties of proteins themselves, influencing stability, folding behavior, or molecular recognition through chemical interactions that natural proteins cannot easily access.7
These advances allow proteins to function as more than biological macromolecules assembled from a fixed set of building blocks. By encoding synthetic residues within their sequences, proteins can be designed with precisely positioned chemical functionalities, turning them into molecular frameworks whose reactivity and interactions are specified directly at the genetic level.
Implications for Therapeutic Protein Design
The chemical capabilities introduced by noncanonical amino acids are beginning to translate into new strategies for therapeutic protein design. Because ncAAs provide precisely positioned reactive groups within proteins, they offer a way to control chemical modification and molecular interactions with a level of precision that is difficult to achieve through conventional protein engineering approaches.2
One of the most prominent applications involves the development of site-specific conjugation strategies for antibody-based therapeutics. Traditional antibody–drug conjugates (ADCs) often rely on naturally occurring lysine or cysteine residues for payload attachment, which can lead to heterogeneous mixtures of products. Incorporating ncAAs into defined locations within an antibody creates a chemically addressable site that can be used to attach drug payloads in a controlled and reproducible manner, producing more uniform conjugate structures.2,5
Similar approaches are being explored in cytokine engineering and other protein-based therapeutics. By introducing synthetic amino acids at specific sites within a protein, researchers can enable targeted chemical modification or controlled conjugation strategies that alter pharmacological behavior without disrupting the underlying protein structure. This ability to introduce chemically programmable sites expands the range of design strategies available for modifying therapeutic proteins.5
More broadly, GCE provides a framework for designing biologics with molecular features that cannot be generated through sequence engineering alone. By combining synthetic amino acid incorporation with modern protein engineering and synthetic biology methods, researchers are beginning to develop protein therapeutics whose chemical properties are deliberately programmed into their genetic sequences.
Platforms and Production Systems
Although GCE was initially developed as a research tool, the technology has steadily progressed toward practical expression platforms capable of producing proteins containing ncAAs. Achieving this capability requires integrating orthogonal translation components with expression systems that can efficiently support both the host cell’s normal protein synthesis machinery and the engineered pathways needed to incorporate synthetic amino acids.
Researchers have demonstrated ncAA incorporation across several biological production systems, including bacteria, yeast, and mammalian cells. In each case, the core strategy involves introducing orthogonal aminoacyl-tRNA synthetase and transfer RNA pairs that selectively recognize a specific ncAA and direct its insertion during translation without interfering with endogenous protein synthesis. These systems have been successfully implemented in multiple host organisms, allowing the production of proteins containing synthetic residues in a variety of experimental and biotechnology contexts.1,3
Despite this progress, developing efficient translation platforms for ncAA incorporation remains technically demanding. The engineered aminoacyl-tRNA synthetase–tRNA pairs must function reliably within the host cell while avoiding cross-reactivity with native amino acids or translation components. At the same time, the expression system must maintain sufficient efficiency to produce useful quantities of the engineered protein while ensuring accurate incorporation of the noncanonical residue at the intended position.2,3
These requirements highlight one of the central engineering challenges of GCE. Orthogonal translation systems must be carefully optimized to balance fidelity, efficiency, and compatibility with the host cell’s metabolic environment. Improvements in synthetase engineering, codon reassignment strategies, and host strain development continue to refine these systems and broaden the range of ncAAs that can be incorporated into proteins.1,4
As these platform technologies mature, the field is gradually shifting from proof-of-concept demonstrations toward more systematic approaches to protein production with expanded genetic codes. What began as a powerful laboratory technique for studying protein function is increasingly being viewed as an engineering framework for designing and manufacturing proteins with precisely defined chemical capabilities.
Toward a Programmable Chemical Biology of Therapeutics
The continued development of GCE suggests a broader shift in how proteins are conceptualized in biotechnology. For decades, protein engineering focused primarily on modifying sequences within the constraints of the canonical amino acid set. Expanded genetic codes challenge that limitation by enabling proteins to incorporate synthetic amino acids that introduce new chemical capabilities directly into their structures. As these technologies mature, they create the possibility of designing proteins whose functions depend not only on sequence and structure but also on chemistry that lies outside the repertoire of natural biology.
One consequence of this shift is the potential emergence of proteins with functions that extend beyond those accessible through conventional mutagenesis. By embedding new functional groups or chemical reactivity into protein frameworks, expanded genetic codes allow researchers to explore new catalytic activities, novel interaction mechanisms, and engineered biological systems that operate with a broader chemical toolkit. In this sense, GCE represents an extension of synthetic biology, where the boundaries of biological function are expanded through deliberate engineering of molecular components.
The convergence of synthetic chemistry, molecular biology, and protein engineering is central to this vision. GCE connects these disciplines by linking chemical functionality directly to genetic information. Once a synthetic amino acid can be encoded within the translation machinery, its chemical behavior becomes programmable through DNA sequence in much the same way that natural protein structure and function are encoded today.
Viewed in this context, expanded genetic codes represent more than a technical advance in protein engineering. They point toward a future in which biological systems can be designed with chemical capabilities that evolution never explored. As researchers continue to develop new orthogonal translation systems and synthetic amino acids, the ability to program chemical functionality into proteins at the genetic level may become an increasingly important strategy for biotechnology and therapeutic development, and may unlock new therapeutic possibilities that could never be achieved using the canonical genetic code.
References
1. Kim, YouJin, et al. “tRNA engineering strategies for genetic code expansion.” Front. Genet. Sec. RNA. 6 Mar. 2024.
2. Huang, Yujia and Tao Liu. “Therapeutic applications of genetic code expansion.” Synth. Syst. Biotechnol. 3: 150–158 (2018).
3. Andrews, Joseph, Qinglei Gan, and Chenguag Fan. “’Not-so-popular’ orthogonal pairs in genetic code expansion.” Protein Sci. 32: e4559 (2023).
4. Pigula, Michael, and Peter G Schultz. “Recent Advances in the Expanding Genetic Code.” Curr. Opin. Chem. Biol. 83: 102537 (2024).
5. Huang, Yujia, et al. “Genetic Code Expansion: Recent Developments and Emerging Applications.” Chemical Reviews. 125: 523–598 (2024).
6. Kim, Donghyeon, Doeon Sung, and Jeong Wook Lee. “Expanding the genetic code: Strategies for noncanonical amino acid incorporation in biopolymer.” Bioresource Technology. 432: 132691 (2025).
7. Brouwer, Bart, et al. “Noncanonical Amino Acids: Bringing New-to-Nature Functionalities to Biocatalysis.” Chemical Reviews. 124: 10877–10923 (2024).












