Papers1 provider · 1 record
November 17, 2015· Biotechnology and Bioengineering
article
Open access

Synthesis aided design: The biological design‐build‐test engineering paradigm?

Abstract

An open question in biotechnology concerns the extent to which rapid advancements in DNA reading and writing technologies will shift current paradigms in the engineering of biological systems. The prevailing paradigms involve an ad hoc combination of forward engineering via serial testing of specific hypotheses and reverse engineering via random exploration of phenotypic landscapes. Combining these approaches have proven successful in many cases; however, more concerted and comprehensive approaches that leverage computational resources for experimental design visualization, modeling, and optimization at systems-level are desired. Unfortunately, we have not yet entered an era in which biological simulations can accurately predict the behavior of designed systems. Against the backdrop of increasingly affordable synthetic DNA and high-throughput testing capabilities, it is reasonable to speculate that the biological design-build-test cycle may be optimized by directly synthesizing and testing thousands of designs iteratively (Fig. 1). That is, the promise of advances in DNA synthesis and sequencing is the ability to construct and test >10,000 variants of proteins, pathways, and ultimately genomes for low cost and on laboratory timescales. This capability enables the adoption of “weak” hypothesis approaches that allow the parallel testing of >10,000 genotype-phenotype hypotheses in machine learning driven strategies for searching combinatorial genome space. In this manner, systems biology datasets for thousands of variants can be leveraged to improve computational models, driving forward engineering via a “synthesis aided design” paradigm. The same strong to weak hypothesis paradigm shift began to happen in computer science almost 50 years ago. As described by Bradley Efron in a 1979 article regarding the “unthinkable” impact that automated computation would have on the state of statistical modeling, the “unthinkable” mentioned in the title is simply the thought that one might be willing to perform 500,000 numerical operations in the analysis of 16 data points. Or one might be willing to perform a billion operations to analyze 500 numbers. Such statements would have seemed insane 30 years ago, when a slow and noisy fifty pound desk calculator that added, subtracted, multiplied, and divided was the most sophisticated computational aid available to most scientists. Most of the statistical theory in common use was developed under the constraint of slow and expensive computation. Now computation is fast and cheap. It is not surprising that new theory is being developed, which takes advantage of the high-speed computer (Efron, 1979). This sentiment neatly describes the shift in experimental design strategies occurring today in biotechnology as we move from slow and expensive DNA synthesis to fast and cheap genome engineering. The first wave of reports demonstrating data-driven experimental design and testing of biological processes happened quite predictably at the level of short peptides (Hellberg et al., 1987). A range of 10–100 sequence variants were taken through at least two rounds of a design-build-test cycle using partial least squares and regression-based modeling, respectively, demonstrating the utility of active learning approaches in building predictive models for biological engineering problems (Mee et al., 1997; Norinder et al., 1997). Following the first wave of knowledge-based algorithm implementations in the 1990s, several reports demonstrated the feasibility of rapidly evolving proteins based on technological achievements in mutagenesis techniques, like DNA shuffling (Stemmer, 1994), as well as high-throughput screening methodologies, including phage display, automated sorting devices, and plate-based assays (Chaparro-Riggers et al., 2007; Crameri et al., 1996; Olsen et al., 2000; Zhang et al., 2002). Although many prior demonstrations of peptide to protein scale directed evolution exist, it was not until a seminal report by Fox et al. that machine learning concepts were extended to protein engineering in a way that allowed testing 1000s of “weak” sequence-function hypotheses on a laboratory timescale, ∼1 month per cycle (Fox et al., 2007). Briefly, several libraries of halohydrin dehalogenase, which plays a pivotal role in biosynthesis of the cholesterol-lowering drug Lipitor, were generated using a combination of rational and random mutagenesis strategies and quantitatively assessed on an individual basis. Importantly, the sequence-activity profiles for a broad range of activities were used to retrain a partial-least squares model correlating mutations with enzyme activity. Ultimately, the group discovered a sequence variant with 4,000-fold activity increase over wild-type. Continued advances in DNA synthesis now enable this same data-driven approach to be pursued in a completely rational manner, where at each stage the enzyme libraries can be synthesized according to whatever search algorithm is desired. More recently data-driven approaches have begun to surface at the scale of whole operons, inspiring the establishment of institutes like the Broad Foundry, where 1000s of pathway-scale constructs can be automatically built and arrayed for specific testing. Notably, constructing computational models with strong predictive power of pathway-level mutational effects can be hindered by the complexity of native regulatory networks as well as a lack of a priori knowledge about mutant-activity relationships. Alternatively, it is now possible to simply construct and test on the order of 10,000 alternative designs, and in this manner identify optimal designs in a “synthesis aided design” approach. Specifically, Smanski et al. refactored the Klebsiella oxytoca nitrogen fixation gene cluster without changing its function by (i) removing all non-coding DNA and regulatory elements, (ii) recoding each essential gene in the operon to remove any internal regulatory features including those yet to be discovered, and (iii) placing recoded genes into artificial operons whose expression levels are controlled by well characterized ribosome binding sites and spacer sequences (Smanski et al., 2014; Temme et al., 2012). Although the refactored cluster only retained 7% activity when expressed in E. coli with respect to the wild-type operon in its native host, this synthetic operon was able to serve as a foundation for applying active learning for iterative optimization. In fact, the most recent reports from this group have revealed the rapid assembly of tens of thousands of designs and optimal variants performing at close to 70% of wild-type activity. A series of reports suggest that we will soon see the “weak” hypothesis approach demonstrated at the genome-scale. Wang and coworkers reported the multiplex automated genome engineering (MAGE) approach as a rapid method for constructing billions of combinatorial mutants spanning a targeted set of genes (Wang et al., 2009). The application of MAGE in many ways parallels the application of ProSAR described above with two key caveats. First, the size of combinatorial genome space requires that the initial search strategy was limited to a small number of pre-selected genes (27 in the case of Wang et al.) relative to the size of genome. Second, the ability to specifically test large numbers of individual MAGE mutants was not possible in the absence of whole-genome (or extremely long-read length) sequencing. The result was a very sparse mapping of genotype to phenotype relationships relative to what could be accomplished at either the protein or pathway levels as described above. In this manner, MAGE and several excellent follow up studies provided a set of impressive examples of how to rapidly and comprehensively construct combinatorial genome libraries, but several additional technologies were required to fully prove out weak-hypothesis driven genome engineering. The Trackable Multiplex Recombineering (TRMR) method (Warner et al., 2010) from our own group was developed to address the first caveat above. In TRMR, barcoded promoter mutants spanning the entire genome were constructed and then applied to map the effect of changes to an individual gene expression level onto a trait of interest. This approach could then be applied to find the smaller set of target genes required for combinatorial library generation via MAGE. We demonstrated precisely such an approach (Sandoval et al., 2012) in the engineering of cellulosic hydrolysate tolerance into E. coli. Although tolerance was improved, the study highlighted key technology limitations, such as the unpredictable efficiency of ssDNA recombineering (Reynolds and Gill, 2015), and emphasized the need to be able to track combinatorial mutants at much greater depth as described above. We recently reported an approach for addressing the latter of these issues by employing emulsion linking PCR to deeply characterize MAGE libraries (Zeitoun et al., 2015). In particular, we characterized population diversity in 4 out of 27 RBS sites across the E. coli genome that were combinatorially varied using MAGE, searching on the order of 105–106 genotypes, representing four orders of magnitude greater tracking-depth than previously possible. With respect to efficiency, the advent and broad applicability of CRISPR technologies could not have come at a better time (Doudna and Charpentier, 2014). CRISPR allows selection for specific genome modifications simply via the inclusion of a guide RNA targeting the CRISPR nuclease machinery to the wild-type sequence (Jiang et al., 2013, 2015). CRISPR has been employed to increase the efficiency of ssDNA recombineering to 95% or greater (Findlay et al., 2014; Fu et al., 2014; Pines et al., 2015). In combination, these technologies provide the remaining pieces for realization of highly efficient genome engineering and optimization via a “weak” hypothesis driven strategy. What do these advances hold for the future of genome scale engineering? We are already seeing a rapid increase in the use of such technologies to demonstrate weak-hypothesis driven strain engineering (Cress et al., 2015; Li et al., 2015; Ronda et al., 2015; Salis et al., 2009) and the codification of this approach in the launching of several new bio-foundries (e.g., Zymergen, Synthetic Genomics, GingoBioworks, Copenhagen, Munich, Amyris, NYU, Edinburgh). These biofoundries operate at rates approximately 100× over the prior state of the art. Given the history of disruption of such machine-learning/weak-hypothesis approaches in parallel fields, we expect that the ship has sailed in terms of questioning this paradigm shift. Rather, the more relevant question is how we will most effectively take advantage of such a shift in the advancement of the field in general? Where are the most obvious application areas, both from a technology (protein or pathway or genome) and a product (antibodies, small molecules, etc.) perspective? How can we use the technology to develop the understanding of design rules required for construction of predictable models and what form will such models take (computational, biological, or both)? How do we restructure our workforce to address the increasing need for computational design and data analysis and reduced need for molecular cloning? Answers to these questions are not obvious, but the need to answer them is clear. We expect many answers will arise through large ongoing efforts from both the public and private sectors (see DARPA Living Foundries or recent financings of startups such as Twist, Gingko Bioworks, and Zymergen). Although centralized efforts are often effective, the community should continue to build upon decentralized efforts (e.g., iGEM, Foldit) that not only can provide a different perspective to such challenges but also help to disseminate technology and build a workforce. Doing so will require sustained support from funding agencies with missions tied to long-term impact as well as new ways of thinking about and judging the impact of innovations in this space. Ryan T. Gill, Andrea L. Halweg-Edwards Department of Chemical and Biological Engineering University of Colorado, Boulder Boulder, CO Aaron Clauset, Sam F. Way Department of Computer Science University of Colorado, Boulder Boulder, CO

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.