Introduction to CRISPR What does CRISPR stand for. It stands for Clustered Regularly Interspaced Short Palindromic Repeats. This was discovered by scientists studying bacteria. Bacteria get infected all the time by viruses called Bacteriophages. They attach to the bacteria and inject their viral DNA. Scientists noticed that the bacteria genome had a sequence of repeats. It showed a small sequence of DNA then a repeat followed by another small sequence and another repeat. They named this area of the bacteria genome the CRISPR region. This is the bacteria's immune system. It works through type 1 CRISPR enzymes like CAS1 and CAS2, chopping up the bacteria DNA. Then a small sequence of the viral DNA called the protospacer is inserted into the CRISPR region of the bacteria genome. What happens next is very awesome. 1 The bacteria transcribes the entire CRISPR region into one long messenger RNA called the pre-cRNA transcript. This then gets chopped up so each DNA sequence that matches the virus' protospacer becomes the CRISPR RNA (crRNA) and the repeat becomes the trans activating RNA (tracrRNA). 2 The repeat segments become the tracrRNA that will get loaded into an enzyme called CAS9. CAS9 stands for CRISPR Associated Protein 9. The CAS9 will recognize the tracrRNA and bind to it. The crRNA will then bind to the tracrRNA to make a complete guide RNA. The two RNA segments bind together by complementary base pair binding by hydrogen bonding. Once the CAS9 is loaded with the complete guide RNA, it will go and find any DNA that has the matching protospacer for the guide. The guide is a complete complementary sequence of RNA that matches that protospacer in the Viral genome. The actual guide sequence is about 20 nucleotides long. Once it finds the matching viral DNA, the CAS9 enzyme will cut that DNA. This destroys the viral DNA. I know what you are thinking. If the DNA sequence copied into the CRISPR region of the Bacteria matches the viral DNA, how does the CAS9 not cut up that bacteria's DNA too? The answer to that is the PAM sequence of the CAS9 enzyme. This stands for Protospacer Adjacent Motif. This is a small segment of nucleotides in the DNA that is recognized by the CAS9 enzyme which is N-G-G. 3 The N stands for any nucleotide while the G stands for Guanine. This basically means the CAS9 enzyme not only needs to match the guide RNA sequence to the protospacer, but it also needs to match its PAM to the 2 Guanines near the protospacer. If both don't exist, the CAS9 will not cut. As you can guess, the CRISPR region of the bacteria genome would not contain the PAM sequence so the CAS9 would be selective for targeting the viral DNA, but not the bacteria's DNA. SgRNA The major change of CRISPR from the wild type was the creation of the single-stranded guide RNA. If you recall, in the wild type, the guide RNA came in 2 sections with the crRNA and the tracrRNA. With the single guide RNA (sgRNA), they use a small linker at the top of the complementary region to link them together. This makes things much easier. CAS9 4 The CAS9 enzyme is made up of a few basic domains. The first is the PI domain which stands for the PAM interacting domain. This part is about finding the PAM sequence. The other two major domains are the nucleases that cut the DNA. These are the RuvC domain and the HNH domain. They do the actual cutting. The rest of the CAS9 protein domains are mainly for structure. The CAS9 enzyme will open up the DNA. It will scan along the DNA with the 20 nucleotide guide until it finds its matching DNA sequence. Only if it finds both the right DNA sequence and has a matching PAM sequence will the Nucleases become active and cut the DNA. The CAS9 enzyme will cut both strands of the DNA at the exact same site called a Double Stranded Break (DSB). The cut will occur about 3 to 4 bases away from the PAM sequence. The CAS9 enzyme is about 4,100 bases in size and is about 1,368 amino acids long. This makes it very big for use in AAV vectors which have a max capacity of about 4,500 bases in size. CAS12 The CAS12 enzyme has one major difference over CAS9 in its structure. It only has one nuclease for cutting with the RuvC domain. It uses this nuclease to cut both strands. When the CAS12 enzyme cuts, it makes a staggered break in the DNA by about 5 nucleotides between cut sites. This is much preferred for DNA repair. It also cuts much further away from its PAM sequence. It cuts about 18 to 23 bases away from the PAM, which makes it capable of doing multiple edits. The CAS12 enzyme does not require the tracrRNA. It only uses the crRNA which makes it much smaller for the guide RNA. The CAS9 uses the full length gRNA which is about 100 nucleotides long. The crRNA used by the CAS12 enzyme will only be about 42 nucleotides long. The CAS12 enzyme is smaller in size than the CAS9 enzyme. It is only 3,800 bases long which is about 1,250 amino acids long. Its smaller size can allow it to be used in most AAV vectors. The last major difference between CAS12 and CAS9 is the PAM sequence. The CAS9 uses the N-G-G sequence, but CAS12 uses T-T-N which is Thymine rich. This allows for access to different areas of the genome for editing. 5 CRISPR MAD7 This is developed from a variant of the CAS12a enzyme found in the area of Madagascar. It looks and behaves much like CAS12a. It's about 1260 amino acids in length. It uses a T rich PAM sequence. Inscripta allows open use of the MAD7 for research purposes and licenses MAD7 to several companies. Recently, Fate did a partnership with them for MAD7 in editing iPSC cells. The Nickase The nickase is a modified CAS enzyme designed to only cut a single strand of the DNA. This is often done with CAS9 in which one of the two nucleases is inactivated. This allows for the cutting of just a single strand of the DNA. This becomes more important in Base Editing and Prime Editing, which we will look at next. 6 Double Stranded Breaks There is one major drawback to the first generation of CRISPR CAS enzymes. They make Double Stranded Breaks (DSB). When a double stranded break occurs, the cell can repair it one of two ways. The first is Homology Directed Repair (HDR) and the other is Non Homologous End Joining (NHEJ). The HDR system is only available when the cell is going through mitosis after it has copied its Chromosomes. It is only available when there is a sister Chromatid available. 7 Most cells do not turn over very frequently so this process is not available to those cells. The majority of the time, the process of DNA repair will be done using NHEJ. This is a very inaccurate way to repair the DNA. It basically just sticks the ends back together again. If the break isn't perfect, you can end up with Insertions and Deletions (Indels) of bases into the DNA. This can throw off the entire coding of the protein and lead to mutations. Translocations occur when multiple Double Stranded Breaks occur at one time. This can lead to different segments of the DNA being put back together that did not belong together. This occurred recently in an Allogene patient where they found some CAR-T cells that had the TCR locus translocated to another place in the genome. 8 Double Stranded Breaks will lead to p53 activation. When there is a low level of p53 activation the cell will attempt to undergo DNA repair by NHEJ or HDR. If you get a lot of p53 activation, this can lead to genomic instability which will cause the cell to undergo Programmed Cell Death (Apoptosis). This will cause the cell to die and lower the efficiency of the gene editing. Killing of the cells by Double Stranded Breaks is a concern with the efficiency of editing. This mostly becomes a problem when many edits are attempted at one time which can over-activate p53. Base Editing Before we get into base editing, we must first cover a bit about genetics. There are 4 different bases with Adenine, Guanine, Cytosine and Thymine. The Adenine and Guanine are called Purines, and they have 2 rings. The Cytosine and Thymine are called Pyrimidines and have only 1 ring. When a purine mutates into another purine or a pyrimidine mutates into another pyrimidine, this is called a transition mutation. When one of the purines is replaced by a pyrimidine or a pyrimidine is replaced with a purine, this is called a transversion mutation. 9 Why does this matter? When it comes to base editing, it can only do transitions, not transversions. The base editors come in 2 forms with an Adenine Base Editor (ABE) and a Cytosine Base Editor (CBE). The Base Editing comes with three major parts. The first is the CAS9 enzyme just like in normal CRISPR editing. The big difference with this CAS9 is that one of the nucleases is inactivated so that it will only cut a single strand of the DNA. We call this a nickase when it cuts only 1 strand. The second part is the guide RNA which is also the same as with normal CRISPR editing. The major difference is the Deaminase that is attached to the CAS9 enzyme. The system works by the guide RNA going along the DNA just like in other CAS9 editing until it matches both the guide and the PAM to the desired DNA. Then the deaminase will remove an amino group to an Adenine within the editing window. The editing window is about 4 to 5 bases wide. This converts the Adenine to Inosine which will be read as Guanine. Now you will have a mismatch of the two bases as the new inosine will be a Guanine and no longer match the Thymine that was originally paired with the Adenine. The nickase will cut the opposite strand that was not altered where the mismatch Thymine is located. This will kick in the DNA repair which will replace the Thymine with a new Cytosine to match the Inosine. This is the basic process of Base Editing. The first issue with base editing is that human cells have an enzyme called Uracil DNA Glycosylase (UNG) which is designed to detect and repair any Uracils that appear in the DNA from deamination. After all, deamination happens to our DNA all day long. This only happens to a Cytosine that has undergone deamination. You will notice that on the Cytosine Base Editor (CBE) there is another small enzyme called Uracil Glycosylase Inhibitor (UGI). This is there to block the activation of the UNG and repair of the Cytosine deamination before it can be converted to a Thymine. 10 The other issue with base editing is the editing window. It is 4 to 5 bases wide. More than one nucleotide of the same kind can appear inside that window. If you get 2 Adenines inside the window, there is no way to control which one of those Adenines gets changed. These are called bystander edits. This means every target in base editing must be extensively tested to see if the bystander edits will cause problems. Potential solutions can be to just shift the editing window in one direction or the other by a few bases if there is a bystander edit that causes an issue. Prime Editing The one big drawback to base editing is that it can not insert into the DNA because it only cuts a single strand. That brought about the idea of Prime Editing. This uses a CAS9 enzyme with one of the nucleases inactivated just like in base editing. It has a modified guide RNA called the Prime Editing Guide RNA (pegRNA). This includes the original guide RNA for finding the right location in the genome plus it includes a segment for the Reverse Transcriptase to copy into the DNA. This works by the guide RNA going along the DNA until it finds the correct site for the guide. Then it cuts one strand of the DNA. A section of the pegRNA will bind to the opened DNA strand to stabilize it using base pairing. Then the reverse transcriptase will use the rest of the pegRNA as a template to copy new bases into the DNA. There isn't a lot known about the details of how the system fully works or how the repair function of the DNA is handled. I know it has some challenges which is why they moved onto a Twin Prime editing version. This uses a prime editor on each of the strands of the DNA. To be honest, I don't know if this will ever be perfected. 11 CRISPR delivery There are three ways that the CRISPR machinery can be delivered into the cells. There are only two that are actually used so we will look at them. The most common is to encode the CAS enzyme into a messenger RNA (mRNA). Then the mRNA for the CAS and the RNA for the guide are delivered into the cells using a Lipid NanoParticle (LNP) or a viral vector. The most commonly used is the LNP delivery. The mRNA for the CAS9 is taken right into the Ribosomes where it gets translated into the fully functional CAS enzyme ready for use. Then it binds to the guide RNA and goes to work. The other method that is gaining more use is the fully functional CAS enzyme and guide RNA assembled into a LNP ready to go. This allows the machinery to be ready to edit the moment it gets into the cell. It also helps avoid any of the inefficiency issues with not enough of one or other components getting into cells in equal amounts. The drawback is it is far more expensive to make. This process is called the Ribonucleoprotein (RNP) delivery method. Delivery continues to be the major challenge to Genetic Editing of all kinds. There is a lot of amazing technology out there that is completely unable to reach the desired tissues of target. This is an issue of the vectors we currently have. The traditional AAV vectors are very small for CRISPR delivery and require multiple vectors to delivery different fragments. It is the least preferred method. The LNP technology is great, but lacks tropism to tissues outside of the normal blood circulations like Liver, Stem cells and Blood cells. 12 CRISPR applications All of the current applications of CRISPR technology have been using knockout of genes. This is seen in the SCD trials where they knock out the BCL11a gene which stops the production of fetal hemoglobin. Knocking out genes is very simple and very accurate. Many of these therapies will knock out multiple genes in T cells with 98% accuracy. When it comes to inserting a gene, it's a whole different story. The potential for Indels and mutations is far too high using NHEJ repair. We have seen in some of the CAR-T therapies that an insertion using CRISPR will be about 75% efficiency. 13 This has led to the targeted integration design. This is where they use the CAS enzyme to find and cut the DNA with a double stranded break. Then also insert a donor template of DNA which includes homology arms that match the DNA. This can drive HDR repair of the DNA, which is far more accurate. This method requires the delivery of the CAS and guide RNA in one vector while the donor DNA is carried by another vector, usually an AAV vector. So far the efficiency with this method is still very low. Some of the data from Graphite Bio has shown very low efficiency rates. The SLEEK platform for Editas showed upward of 87% targeted integration of a CAR into a T cell, which is better than the 75% seen previously. The SLEEK system uses CAS cutting along with a DNA donor template, but it only targets a specific exon of a protein vs trying to change an entire gene. The efficiency has been a lot better with 87% efficiency. 14