Identity-by-state analysis: a new method for PGT-M

Stephen Brown · Human Reproduction · 2020

The article by Ding et al. in the present issue of Human Reproduction builds on the authors’ previous innovations in the field of preimplantation genetic testing. In the present paper, they describe a new method of preimplantation testing that uses a sibling of an affected parent, rather than a prior affected child, to establish disease linkage. By so doing, this elegant work adds to a growing armamentarium of tools that can be used both for preimplantation genetic testing of Mendelian disorders (PGT-M) as well as preimplantation genetic testing of aneuploidy (PGT-A). Prior to a discussion of this paper and what it adds to the literature, it is useful to briefly trace the history of this rapidly evolving field. Early efforts at PGT-M relied on PCR-based methods for direct amplification and detection of parentally derived mutation(s) in DNA from single-cell biopsies. This approach was severely limited by the phenomenon of allele-drop-out, or ‘ADO’, a situation in which one of the two alleles at a given locus fails to amplify. ADO is critically important because it can, and not infrequently does, lead to disastrous false negative results. To address this problem, laboratories began to use tightly linked DNA polymorphisms (usually simple-sequence repeats) to follow the segregation of the chromosome segment containing the disease locus. If an embryo did not inherit polymorphic markers near the disease locus, one could be confident that it was unaffected. While this approach dramatically decreased the incidence of false negative results, it was patient/family-specific, time-consuming and not always possible to achieve. A major breakthrough came with the development of DNA amplification methods that allowed for array-based genotyping of single nucleotide polymorphisms (‘SNPs’) on single-cell or few-cell biopsies. Array-based genotyping of hundreds of thousands of SNPs on parents and a previously affected child made it possible to establish dense SNP haplotypes surrounding almost any disease locus in the genome. Once haplotypes were established, arrays could be used to genotype embryos and the inheritance of disease-causing variants could be robustly determined, without the need to directly test for the actual disease-causing variant (Handyside et al., 2010). This method, called karyomapping by the authors who described it, had a major impact on the field, since it allowed the same workflow to be used on almost every PGT-M case; however, technical challenges have prevented it from being used for aneuploidy detection. Another significant development came in 2015, when the present authors described a new method for combined PGT-M and PGT-A, which they called ‘haplarithmisis’ (meaning haplotype counting) (Zamani Esteki et al., 2015). This method allowed for the simultaneous establishment of SNP haplotypes as well as accurate determination of chromosome copy number on single-cell or few-cell biopsies. Although haplarithimsis is complicated, it fundamentally relies on haplotype-specific differences in SNP array intensity data to assess relative copy number of chromosomal segments. Because it is haplotype-specific, it yields not only genomic copy number but also parent of origin of duplicated or deleted segments. Several publications have demonstrated the utility of Haplarithmisis, including a recent paper in Human Reproduction where the method was used to assess embryonic ploidy (Destouni et al., 2018). Although karyomapping and haplarithmisis both provide a near universal method for performing PGT-M, both require either an affected child (sibling of the embryos) or parents (grandparents of the embryos) for determination of which haplotype(s) carry the disease alleles, a process called ‘phasing’. In practice, the relatives needed for phasing are frequently either lacking or difficult to access, which brings us back to this issue of Human Reproduction. The paper by Ding et al. addresses the situation in which a parent with a dominant or X-linked illness has a sibling (either affected or unaffected) available for testing. The authors show that, by analyzing the SNP genotypes of the affected parent and his or her sibling, the haplotype carrying the disease locus can be determined. This information can be used to assess the disease status of embryos. The easiest way to see how this works is to imagine a SNP located near a disease locus. If the affected parent is heterozygous (A/B) and the affected sibling is B/B, one can infer that the disease allele is linked to the haplotype carrying B. If an embryo is B/B, then the B (disease carrying) allele was inherited from the affected parent. This process, which is termed identity-by-state analysis, can only be applied when the two siblings have inherited the same parental haplotype from one parent and the alternate haplotype from the other parent, a situation that has a 50% chance to occur. If both siblings have inherited the same parental haplotypes surrounding the disease locus, then all SNPs will have the same genotype (A/B, A/A or B/B), making it impossible to determine which haplotype carries the disease allele. Likewise, the siblings may not have either parental haplotype in common, again making it impossible to determine which haplotype carries the disease. In practice, haplotype phasing with a sibling will only succeed in about half of the cases, and, in fact, the authors report that phasing of the disease locus was successful in about half the families in which it was attempted. Importantly, all of the cases reported involve either dominant or X-linked conditions. Presumably, this is because embryonic detection of recessive illness requires the phasing of both parental disease haplotypes, which would require both parents to have an informative sibling, which would occur in a maximum of 25% of the cases, assuming that each parent had one available sibling. So, to summarize, Ding et al. have expanded on their previous work by developing a method for PGT-M that uses siblings of parents to establish phasing of SNP haplotypes. At the same time, they provide a method for performing haplotype-specific copy number determination in embryos, using sibling-phased SNP data. While there is no doubt that this work improves the chances that a given couple/family will have the necessary relatives for successful PGT-M, there will still be many situations in which PGT-M is challenging. For instance, dominantly inherited illnesses are frequently caused by de novo mutations, making it impossible to phase haplotypes on the basis of existing family members. Also, couples at risk of recessive illness are increasingly being identified through expanded carrier screening, rather than by having an affected child. Thus, there is considerable demand for truly general methods for PGT-M and, preferably, methods that do not rely on family members. In this context, it should be mentioned that a widely used (but seldom reported) solution to the phasing problem is to use the embryos themselves to determine which parental haplotype(s) carry the disease variant(s). If the disease-causing variant can be detected in DNA amplified from an embryo biopsy, then phase is established. For this to provide reliable results, the disease variant must be one that can be reliably assessed by PCR and sequencing. It also requires a sufficient number of embryos, since each embryo has only a 50% chance to harbor a parental disease variant. What new PGT-M methods are on the horizon? Clearly, next-generation sequencing has completely changed the landscape of PGT-A. Aneuploidy assessment with shallow whole genome sequencing is fast, accurate, easy to multiplex and cost-effective. In contrast, SNP genotyping, whether for PGT-M, PGT-A or both, is relatively expensive, since both parents, other necessary relatives and every embryo require a SNP array. Therefore, from a work flow and cost perspective, sequencing-based methods for PGT-M would be highly desirable. In an ideal world, the same sequencing data could be used for both PGT-M and PGT-A. However, genotyping of SNPs through sequencing, while possible, requires much deeper sequencing than aneuploidy assessment and is consequently prohibitively expensive. Nonetheless, several PGT-M methods that rely on sequencing have been suggested. Yan et al. have described a method that they call ‘mutated allele revealed by sequencing with aneuploidy and linkage analysis’ or MARSALA (Yan et al., 2015). In their method, parents are first genotyped with SNP arrays, allowing the identification of a number of informative SNPs near the disease locus. Then, DNA from embryo biopsies first undergoes whole genome amplification, and PCR using an assortment of specific primers designed to amplify a number of SNPs that flank the disease locus is performed, using a small amount of embryo DNA as template. This PCR product is then added back to (‘spike’) the embryo DNA, and shallow whole genome sequencing is performed. This provides sufficient-read depth for aneuploidy assessment while, at the same time, provides plenty of read-depth to genotype multiple SNPs that flank the disease locus. This allows the SNP haplotype linked to the disease to be established, facilitating robust, accurate diagnosis. While this approach eliminates the need for an SNP array on each embryo, it has other problems: first, informative parental SNPs must be identified with either SNP arrays or deep whole genome sequencing. Second, each case requires individualized informatic analysis, primer design and synthesis. Therefore, it seems unlikely that MARSALA will become a widespread method for PGT-M. Haploseek is another promising effort to establish a universal work flow for both PGT-M and PGT-A (Backenroth et al., 2019). In this method, parents and an affected child are first genotyped with SNP arrays, allowing phasing of virtually any genomic locus, just as in karyomapping or haplarithmisis. Then, instead of genotyping embryos with SNP arrays, embryo biopsies are subjected to shallow genomic sequencing (<1× read depth). SNP genotypes of informative loci (previously identified using data from the family SNP arrays) are determined using a combination of prior information from the family’s SNP array data, coupled with sophisticated statistical analysis. In a proof-of-concept paper, the authors show that their method allows the accurate determination of the inheritance of parental haplotypes across the genome and hence disease prediction. The same sequencing data can be used to assess for aneuploidy using standard methods. While this system is both elegant and cost-effective, it still requires an affected child and at least three SNP arrays. In addition, a more extensive validation study will be needed before it can be used clinically. Yet a third approach to obtaining SNP genotypes with sequencing has been described by Masset et al. (2019). Their method makes use of an old trick to dramatically reduce the complexity of genomic DNA by restriction enzyme digest followed by linker ligation and PCR with a primer corresponding to the linker. The number of reads needed to achieve a sufficient depth for SNP genotyping is correspondingly reduced, making it reasonably inexpensive. The SNP genotypes obtained from parents, an affected relative and from embryos can be resolved into haplotypes, phased and used for PGT-M, using their previously described haplarithmisis method. At the same time, the same sequence data can be used for PGT-A, providing a cost-effective, unified work flow. All these developments show that the field is moving rapidly and that there is plenty of room for innovation and improvement. It seems quite likely that sequencing-based methods that allow for both PGT-A and PGT-M will quickly replace array methods. One can imagine that long-read sequencing on unamplified DNA from embryo biopsies might eliminate the need for family-based phasing of disease bearing haplotypes. Stay tuned! No relevant funding. No conflict of interest.

Read the paper · More papers on PaperTik