Projects for 2027-28

As part of your application, you will be asked to indicate your preferred research projects by selecting from the two project categories below: Teal and Fuchsia. You must select at least one project from each category and include your choices in the 'Proposed field and title of research project' section of the application form. Simply list the project colour and project number (for example, Teal 3 or Fuchsia 7).

The two project categories reflect the primary research environment for the DPhil:

  • Teal Projects: The primary scientific direction and project trajectory are defined by EIT. Most students on these projects will be primarily based at the Generative Biology Institute (GBI) within the Ellison Institute of Technology (EIT) at Oxford Science Park.
  • Fuchsia Projects: The scientific direction is either (i) developed jointly by the University of Oxford and EIT through shared scientific interests, including topics beyond the supervisors' principal research programmes, or (ii) led by the University of Oxford where the research is better conducted there or the necessary expertise is not available at EIT. Most students on these projects will be primarily based  in the laboratory of the student's University of Oxford supervisor. 

Regardless of project category, all students are fully integrated into the Generative Biology EIT DTP cohort and benefit from the same programme of research training, cohort activities, and professional development opportunities.

The projects available for 2027–28 entry are listed below.

Students walking, Christ Church

Copyright © University of Oxford Images / Public Affairs Directorate

Genome Synthesis and Encoded Polymer Synthesis

Lead supervisor: Jason Chin (Principal Investigator, Generative Biology Institute, EIT & Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) - chin@eit.org

Co-supervisor: Ben Davis (Professor of Chemical Biology, Department of Pharmacology, University of Oxford) – ben.davis@pharm.ox.ac.uk

The Chin group’s work pioneers: 1) the development and application of genome design and synthesis methods and 2) combines these approaches with cellular engineering for the encoded cellular synthesis of new polymers and materials.  Some recent examples are exemplified in the publications below.  Within these broad areas, or other areas within the scope of GBI, students are encouraged to propose their own project ideas as part of their application. 

  • Dunkelmann, D.L., Piedrafita, C., Dickson, A. et al. Adding α,α-disubstituted and β-linked monomers to the genetic code of an organism. Nature 625, 603–610 (2024). 
  • Robertson, W. E. et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057–1062 (2021).   
  • Robertson WE, Rehm FBH, Spinck M, et al. Escherichia coli with a 57-codon genetic code. Science. 2025;390(6771) 
  • Zürcher, J. F. et al. Continuous synthesis of E. coli genome sections and Mbscale human DNA assembly. Nature 619, 555–562 (2023).    
  • Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019).    
  • Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59–64 (2016). 

Pioneering the Human Genome Synthesis

Lead supervisors: Jason Chin (Principal Investigator, Generative Biology Institute, EIT & Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) - chin@eit.org 

Co-supervisor: KJ Patel (Director, Weatherall Institute of Molecular Medicine, University of Oxford) - ketan.patel@imm.ox.ac.uk 

Since the completion of the Human Genome Project at the start of the century, researchers have sought the ability to write our genome from scratch. Unlike genome editing, genome synthesis allows for changes at a greater scale and density, with more accuracy and efficiency, and will lead to the determination of causal relationships between the organisation of the human genome and how our body functions. To date, scientists have successfully developed synthetic genomes for microbes such as E. coli (Syn57). The genomes of higher organisms, such as mammals and plants, however, are on the scale of up to billions of bases (103 times larger than microbial genomes synthesised to date), and there are currently no technologies for writing these gigabase-scale genomes.  

We are developing  the foundational and scalable tools, technologies and methods needed to build human chromosomes and genomes, and other gigabase-scale genomes. Genome synthesis may allow us to understand the rules underpinning natural human genomes. This will open entire new fields of research in human health and have profound impacts on biotechnology, which may include the development of safe, targeted, cell-based therapies. Methods developed for the synthesis of human genomes will also provide a foundation for the synthesis of the gigabase-scale genomes of crops, animals and other organisms with important applications in food security and sustainability, such as climate-resistant crops.    

  • G Petris, S Grazioli, L van Bijsterveldt. Et al., High-fidelity human chromosome transfer and elimination. Science (2025).  
  • Zürcher, J. F. et al. Continuous synthesis of E. coli genome sections and Mbscale human DNA assembly. Nature 619, 555–562 (2023).     
  • Robertson WE, Rehm FBH, Spinck M, et al. Escherichia coli with a 57-codon genetic code. Science. 2025;390(6771) 

Microbial Synthesis

Lead supervisor: Jason Chin (Principal Investigator, Generative Biology Institute, EIT & Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) - chin@eit.org 

Co-supervisor: Frank Bürmann (Principal Investigator, Department of Biochemistry, University of Oxford) - frank.burmann@bioch.ox.ac.uk 

Our ability to write DNA has recently expanded to the genomic scale. The possibility of defining every single base in the genome of a cell enables manipulation of the most fundamental cellular properties, such as the genetic code.   

However, current genome synthesis methods are slow, narrow in scope, and limited in scale. To date, the few versions of only two bacteria have been successfully synthesized. This project aims to develop methodologies to make the synthesis of model organism genomes (i.e. E. coli) more rapid and enable the rapid testing of many genome designs.   

The ability to routinely synthesize designed genomes will enable uncovering of new to nature functionalities. Ultimately, the combination of microbial genome synthesis and artificial intelligence will enable biological design at the organism scale with implications in bioproduction, human health, agriculture, and beyond.   

  • Robertson WE, Rehm FBH, Spinck M, et al. Escherichia coli with a 57-codon genetic code. Science. 2025;390(6771)   
  • Zürcher, J. F. et al. Continuous synthesis of E. coli genome sections and Mbscale human DNA assembly. Nature 619, 555–562 (2023).      
  • Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019).      
  • Gibson, D. G. et al. Creation of a bacterial cell controlled by a chemically synthesized genome. Science 329, 52–56 (2010).     
  • Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59–64 (2016).  

Microbial Genome Synthesis

Lead supervisor: Jérôme Zürcher (Principal Investigator, Generative Biology Institute, EIT & Associate Professor of Biological Chemistry, Department of Chemistry, University of Oxford) - jeromez@eit.org   

Co-supervisor: Mathew Stracy (Group Leader, Dunn School of Pathology, University of Oxford) - mathew.stracy@path.ox.ac.uk 

Our ability to write DNA has recently expanded to the genomic scale. The possibility of defining every single base in the genome of a cell enables manipulation of the most fundamental cellular properties, such as the genetic code.   
 
However, current genome synthesis methods are slow, narrow in scope, and limited in scale. To date, the genomes of only two bacteria have been successfully synthesized. This project aims to develop methodologies to make the synthesis of model organism genomes (i.e. E. coli) more rapid and enable the synthesis of the genomes of non-model bacteria to broaden the scope of genome synthesis.     
 
The ability to routinely synthesize the genomes of a diverse set of organisms will not only allow reprogramming of the genetic code but also facilitate testing of generative genome designs. Ultimately, the combination of microbial genome synthesis and artificial intelligence will enable biological design at the organism scale with implications in bioproduction, human health, agriculture, and beyond. 

  • Robertson WE, Rehm FBH, Spinck M, et al. Escherichia coli with a 57-codon genetic code. Science. 2025;390(6771)   
  • Zürcher, J. F. et al. Continuous synthesis of E. coli genome sections and Mbscale human DNA assembly. Nature 619, 555–562 (2023).      
  • Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019).      
  • Gibson, D. G. et al. Creation of a bacterial cell controlled by a chemically synthesized genome. Science 329, 52–56 (2010).     
  • Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59–64 (2016).  

Engineering Universal Centromeres

Lead supervisor: Dr Linda van Bijsterveldt (Principal Investigator, Generative Biology Institute, EIT) - linda.vanbijsterveldt@eit.org 

Co-supervisor: Prof KJ Patel (Director, Weatherall Institute of Molecular Medicine, University of Oxford) - ketan.patel@imm.ox.ac.uk 

Synthetic chromosomes enable the engineering of entirely new biological functions into mammalian cells without disrupting host genomes. Their broadest applications depend on deploying a single modular chromosome across diverse organisms, yet synthetic or engineered chromosomes are frequently lost when transferred between species. 

The difficulty lies in maintaining centromere identity, which ensures a chromosome is correctly distributed into each new cell as they divide. Centromere function depends on species-specific molecular machinery, so a human-derived chromosome introduced into another organism is gradually lost over successive divisions, limiting its potential usefulness in an in vivo setting.

We recently generated mouse embryonic stem cells carrying intact human chromosomes [1]. This system provides an ideal platform to re-engineer centromeres so that they function independently of their host. This project aims to construct self-contained, universal centromeres that allow a single engineered chromosome to be transferred and stably maintained across many different species. 

Beyond improving stability, redesigning centromeres opens fascinating biological questions: What gives a centromere its identity? Could a single designed centromere work equally well in a plant, a mouse, and a human cell? Can centromere evolution be recreated in the laboratory, or even harnessed to bias a chromosome's own inheritance [2]?  

  • Petris, G, Grazioli, S, van Bijsterveldt, L, et al., High-fidelity human chromosome transfer and elimination. Science (2025).  
  • Henikoff, S., Ahmad K., Malik H.S. The centromere paradox: stable inheritance with rapidly evolving DNA. Science 293, 1098-1102 (2001). 

Something From Nothing

Lead supervisor: Chang Liu (Principal Investigator, Generative Biology Institute, EIT; Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) – chang.liu@eit.org 

Co-supervisor: Ben Davis (Professor of Chemical Biology, Department of Pharmacology, University of Oxford) – ben.davis@pharm.ox.ac.uk 

This project will aim to evolve new families of functional proteins starting from random sequence. We have two major motivations. The first is to witness de novo gene birth, a provocative evolutionary mechanism where new genes originate and evolve from non-coding sequences rather than through the canonical process of duplication and divergence of existing genes. 

The second is to sample (and populate) the true distribution of functional protein sequence space, uncontaminated by the (unknown) idiosyncrasies and historical contingencies that produced natural proteins. Our project will leverage orthogonal DNA replication (OrthoRep) systems to drive the rapid evolution of random sequence libraries to confer function in vivo. As we recently demonstrated, when libraries of known genes are encoded on OrthoRep to drive their accelerated evolution, novel gene functions can originate and evolve with surprising ease in cellular contexts. We now wish to observe the same from random sequences to originate and evolve a genuinely synthetic biology. 

  • Pisera et al. Continuous evolution of gene libraries towards arbitrary functions reveals the versatility of biomolecular evolution in vivo. bioRxiv (2026). https://doi.org/10.1101/2025.03.22.644768 
  • Rix et al. Continuous evolution of user-defined genes at 1-million-times the genomic mutation rate. Science 386, eadm9073 (2024).  
  • Tian et al., Establishing a synthetic orthogonal replication system enables accelerated evolution in E. coli. Science 383, 421-426 (2024). 
  • Rehm et al., Highly mutagenic continuous evolution in E. coli using a Φ29-based orthogonal replication system. Nature Biotechnology (2026). https://doi.org/10.1038/s41587-025-02944-x 
  • Ravikumar et al., Scalable, continuous evolution of genes at mutation rates above genomic error thresholds. Cell 175, 1946-1957 (2018). 

Continuous Evolutionary Data Generation and ML

Lead supervisor: Chang Liu (Principal Investigator, Generative Biology Institute, EIT; Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) – chang.liu@eit.org 

Co-supervisor: Harrison Steel (Associate Professor of Engineering Science, Department of Engineering Science, University of Oxford) – harrison.steel@eng.ox.ac.uk 

This project aims to advance protein and RNA design by experimentally generating large volumes of labeled protein evolutionary data whose distributions are strategically shaped for training ML-based design models. To do so, we will leverage orthogonal DNA replication (OrthoRep) systems that can continuously evolve protein and RNA sequences under selection at scale, yielding large and diverse sequence-fitness datasets. We will also establish a virtuous cycle between continuous evolution and design where the data outputs of the former are used to improve the latter, and the design outputs of the latter are used as starting points for further evolutionary data generation in the former. Biomolecular classes of interest include binders, protein complexes, and enzymes. This project may also admit the inclusion of genetically encoded unnatural amino acids as well as the involvement of automation. 

  • Rix et al. Continuous evolution of user-defined genes at 1-million-times the genomic mutation rate. Science 386, eadm9073 (2024). 
  • Alcantar et al. Mapping the evolution of computationally designed protein binders. bioRxiv (2025). https://doi.org/10.1101/2025.10.04.680454 
  • Hummer et al. Investigating the volume and diversity of data needed for generalizable antibody–antigen ΔΔG prediction. Nature Computational Science 5 635-647 (2025). 
  • Bennett et al. Atomically accurate de novo design of antibodies with RFdiffusion. Nature 649, 183-193 (2026). 
  • Tian et al. Establishing a synthetic orthogonal replication system enables accelerated evolution in E. coli. Science 383, 421-426 (2024). 

Reaction-Aware Generative Models for De Novo Design

Lead supervisor: Kiarash Jamali (Principal Investigator, Generative Biology Institute, EIT) - kiarash.jamali@eit.org 

Co-supervisor: Charlotte Deane, (Professor of Structural Bioinformatics, Department of Statistics, University of Oxford) - deane@stats.ox.ac.uk 

Protein design has advanced rapidly in recent years, progressing from the generation of novel folds to the design of new functions. Designing enzymes to catalyse desired chemical reactions is now a realistic goal. Current approaches begin from a predefined active site: either a set of catalytic residues taken from a natural enzyme, or a theozyme constructed by computation or first principles in chemistry. Protein design is then used to build a scaffold that holds these residues in the required geometry. While this approach has produced functional designs, requiring the active site coordinates to be specified in advance limits the chemical solutions that the model can explore. 

A more general strategy would encode the target reaction directly within the design process, allowing a generative model to design the active site and scaffold together. Rather than fixing one catalytic geometry in advance, the model could sample possible geometries conditioned on the chemical demands of the reaction. The central challenge is to represent those demands in a form that is both chemically meaningful and usable by a generative model.  

This PhD project will develop representations of target reactions based on their predicted mechanisms. These representations will be used to train a generative model that designs the enzyme directly. Thus, this project will extend enzyme design to geometries and reactions that are new to nature. 

Related work 

Watson et al., 2023. De novo design of protein structure and function with RFDiffusion.  

Jiang et al., 2008. De Novo Computational Design of Retro-Aldol Enzymes.  

Röthlisberger et al., 2008. Kemp elimination catalysts by computational enzyme design. 

Lauko et al., 2025. Computational design of serine hydrolases. 

Warshel, 1998. Electrostatic Origin of the Catalytic Power of Enzymes and the Role of Preorganized Active Sites. 

Unsupervised Joint Representation Learning for Protein Sequence and Structure

Lead supervisor: Kiarash Jamali (Principal Investigator, Generative Biology Institute, EIT) - kiarash.jamali@eit.org 

Co-supervisor: Charlotte Deane, (Professor of Structural Bioinformatics, Department of Statistics, University of Oxford) - deane@stats.ox.ac.uk 

Large generative models trained on internet-scale data have transformed machine learning in the past few years. One of the key ingredients in their success has been the availability of high-quality and diverse data sources. While machine learning has also recently transformed biology for structure prediction (AlphaFold2) and design (RFDiffusion), the lack of large-scale structural data has hindered further progress.   

So far, most structure prediction and design models have been trained on the Protein Data Bank, a database of experimentally resolved structures that contains a few hundred thousand structures, many of which are homologues of other entries in the database. Economically and therapeutically relevant classes of complexes such as protein-RNA and protein-ligand interactions are even a smaller subset of this data. This data sparsity has limited the application of generative models in areas such as enzyme design and drug discovery.  

In contrast, sequence data is abundant. Large sequence datasets have already been used for training protein and genome language models at scale. Furthermore, it has been shown that protein and genomic language models encode structures even though they are only trained to predict sequences, with a computationally efficient training run afterwards capable of extracting unsupervised structural features learned by these networks.  

The hypothesis of the PhD project is that training models by mixing large sequence data with sparse supervision from structure data will enable learning structure-aware representations that would improve generalisation compared to current structure-only approaches. Concretely, this PhD project will consist of designing semi-supervised and unsupervised approaches to jointly embed sequence and structure, focused on its applicability to modelling protein-ligand interactions. These representations will then be used for training new structure prediction and design models that will be more robust to overfitting on complexes with scarce structural data.  

Related work 

  • Jumper et al., 2021. Highly accurate protein structure prediction with AlphaFold. 
  • Lin et al., 2023. Evolutionary-scale prediction of atomic-level protein structure with a language model.  
  • Zhang et al., 2025. Predicting protein-protein interactions in the human proteome.  
  • Watson et al., 2023. De novo design of protein structure and function with RFdiffusion. 
  • He et al., 2021. Masked Autoencoders Are Scalable Vision Learners.  

Exploring Synthetic Endosymbiosis in Higher Eukaryotic

Lead supervisor: Rongzhen Tian (Principal Investigator, Generative Biology Institute, EIT & Associate Professor of Biological Chemistry, Department of Chemistry, University of Oxford) - rtian@eit.org  

Co-supervisor: Richard White (Professor of Genetics, Ludwig Institute for Cancer Research, Nuffield Department of Medicine, University of Oxford) - richard.white@ludwig.ox.ac.uk 

Endosymbiosis is one of the most important evolutionary events in the history of life, giving rise to mitochondria and chloroplasts and enabling the emergence of complex eukaryotic organisms. Although these events fundamentally reshaped cellular evolution, the molecular and evolutionary processes that convert a free-living bacterium into a stable intracellular endosymbiont and, ultimately, an organelle remain largely unknown. Recent advances in synthetic biology have demonstrated that natural and engineered bacterial endosymbionts can establish mutually beneficial relationships with eukaryotic cells, providing experimental models for studying the early stages of endosymbiosis. 

Building upon these advances, this project aims to establish a synthetic endosymbiosis platform in higher eukaryotic cells to experimentally explore and accelerate the evolution of stable intracellular symbionts under laboratory conditions. The project will first engineer mutually beneficial metabolic interactions between host and bacterium to create the selective pressure required for stable endosymbiotic maintenance. Subsequent integration of rational engineering, targeted and genome wide hypermutation, and AI-guided bioinformatics will drive and analyse the coevolution of both host and symbiont, enabling the rapid emergence of increasingly stable endosymbiotic relationships. 

This project will provide a unique opportunity to experimentally study one of the most fundamental evolutionary transitions in biology while laying the foundation for future applications including nitrogen fixing crop endosymbionts for sustainable agriculture, synthetic organelles for industrial biomanufacturing, and intracellular microbial therapeutics for human disease. 

  • Tian, R. et al. Establishing a synthetic orthogonal replication system enables accelerated evolution in E. coli. Science 383, 421–426 (2024).  
  • Giger, G. H. et al. Inducing novel endosymbioses by implanting bacteria in fungi. Nature 635, 415–422 (2024). 
  • Mehta, A. P. et al. Engineering yeast endosymbionts as a step toward the evolution of mitochondria. Proc. Natl. Acad. Sci. U.S.A. 115, 11796–11801 (2018) 

Ensemble Enzyme Design

Lead supervisor: Sam Pellock (Principal Investigator, GBI, EIT-Oxford, sam.pellock@eit.org

Co-supervisor: Peter Minary (Associate Professor of Computer Science, University of Oxford, peter.minary@cs.ox.ac.uk

Recent advances in deep learning have considerably progressed the de novo construction of enzymes, where it is now possible to generate new enzymes by simply specifying the coordinates of an active site. Despite this transformative advance, the resulting enzymes are often inefficient, and design methods typically accommodate only a single static state, limiting the design of enzymes and machines with complex multistep catalytic cycles. New approaches are needed that model the true ensemble nature of proteins to enable the construction of efficient and intricate enzymes currently out of reach to computation. 

To address this limitation, we will develop and utilize conformational sampling and optimisation algorithms, statistical potentials, and low-dimensional, manifold-based methods to generate and evaluate functional ensembles. By leveraging diverse structural representations—from atomistic to coarse-grained—we will efficiently sample protein conformational landscapes and couple this with probabilistic sequence design tools to yield a scalable design method for protein ensembles that accounts for the multistate nature inherent to enzymes. The resulting computational method will be used to design new-to-nature enzymes that proceed through multi-step reaction paths and expand what is possible with computational protein design.  

Evolutionary Engines for Generative Biology

Lead supervisor: Fabian Rehm (Generative Biology Institute, EIT) - fabian.rehm@eit.org

Co-supervisor: Samuel Sheppard (Professor of Microbial Genomics and Evolution, University of Oxford) – samuel.sheppard@biology.ox.ac.uk 

Evolution is biology’s engine for generating functional complexity. By creating variation, testing alternatives, and amplifying successful solutions, it can search vast design spaces that are difficult to navigate rationally. In nature, however, this process is slow and difficult to direct. This project aims to transform evolution into a programmable engineering platform that operates on experimental timescales. 

We will develop autonomous systems in which mutation, selection, feedback, and amplification are coupled within living cells. These systems will function as biological search engines, continuously generating and refining solutions toward user-defined goals with minimal human intervention. 

The central challenge is to extend continuous evolution beyond individual genes. Many valuable traits—including biosynthesis, environmental sensing, cell–cell communication, stress tolerance, motility, and spatial organisation—emerge from interactions among multiple genes. We will investigate how such complex phenotypes can be evolved continuously, reliably, and at scale.

Key questions include: 

  • How can evolutionary systems adapt mutation rates, selection pressures, and population structure during an experiment? 
  • How can continuous evolution be extended to pathways, regulatory networks, and multicellular behaviours? 
  • Which selection architectures reward complex phenotypes while limiting evolutionary escape? 
  • How can modular selection systems be made portable across genes, organisms, and applications? 

In addition to engineering evolutionary systems, the project will seek to uncover the general principles that govern adaptation. Genome sequencing of evolving populations, combined with genome-wide association/epistasis studies, will enable students to identify adaptive mutations, reconstruct evolutionary trajectories and define the genetic interactions that underpin complex phenotypes. These approaches will be complemented by analyses of natural populations, allowing experimental evolution to be interpreted in the broader context of naturally occurring adaptation, horizontal gene transfer, conserved gene modules and the evolutionary portability of biological functions. 

Students will work across synthetic biology, evolutionary engineering, molecular genetics, automation, and quantitative biology. They will design genetic circuits, build continuous evolution platforms, engineer selection schemes, analyse evolving populations, and apply these tools to challenging biological phenotypes. Students will also have the opportunity to develop and apply modern computational approaches and machine learning to understand the genetic basis of adaptation and improve the predictive design of evolutionary systems. 

We are seeking students who want to develop the next generation of evolution technologies: autonomous, modular, and adaptable systems capable of discovering biological solutions beyond the reach of conventional design. The project offers a unique opportunity to combine experimental evolution with quantitative genomics, providing training in both laboratory and computational approaches to engineering biological systems. 

Key papers 

  • Rehm et al. “Highly mutagenic continuous evolution in E. coli using a Φ29-based orthogonal replication system.” Nature Biotechnology (2026). 
  • Tian, Rehm et al. “A portable orthogonal replication system enables continuous gene evolution near the biological speed limit.” bioRxiv (2026). 
  • Tian, R. et al. “Establishing a synthetic orthogonal replication system enables accelerated evolution in E. coli.” Science (2024). 
  • Molina, R. S. et al. “In vivo hypermutation and continuous evolution.” Nature Reviews Methods Primers (2022). 

Create a Functional Human Chromosome Arm from First Principles

Lead supervisor: Leopold Parts (Principal Investigator, Generative Biology Institute, EIT) - leopold.parts@eit.org

Co-supervisor: Mira Kassouf (RDM Principal Investigator, Weatherall Institute of Molecular Medicine, University of Oxford) - mira.kassouf@imm.ox.ac.uk 

Our team’s goal is to write new chromosomes that implement beneficial functions in animal and plant cells. Since cellular identity and function are primarily determined by the genes that are expressed, and at what levels, faithfully recreating a desired gene expression programme provides a natural benchmark for the design of functional synthetic chromosomes. A natural way to decompose this challenge is to first define the target gene expression programme, then design the DNA sequence that implements this programme, and finally deliver the synthetic chromosome into cells. This Ph.D. project addresses the central challenge of designing regulatory DNA that implements a specified gene expression profile. 

The human genome is an evolutionary happenstance that has its present form due to millennia of innovations that have entrenched a complicated regulatory system. We hypothesise that without the constraints of the limited mutational processes that are available to evolution, much simpler genomes with equivalent function can be engineered. The target gene expression programme is modular, such that the individual regulatory units retain their expression properties regardless of their context, and can be composed without interfering with one another or the broader cell state via feedback loops.  

This combined computational and experimental project aims to generate a synthetic sequence that recapitulates the output of a chromosome arm in a defined cell type and condition. You will use generative and predictive AI models to design regulatory sequences for individual genes, experimentally test these designs in parallel, and hierarchically assemble validated modules into a complete synthetic chromosome arm for delivery into human cells. Working alongside experts in genome modelling, DNA synthesis and synthetic genomics, you will develop computational tools for gene expression design while gaining expertise in large-scale DNA assembly, genome engineering and functional genomic characterisation. 

Minimize a Human Chromosome Arm

Lead supervisor: Leopold Parts (Principal Investigator, Generative Biology Institute, EIT) - leopold.parts@eit.org

 Co-supervisor: Mira Kassouf (RDM Principal Investigator, Weatherall Institute of Molecular Medicine, University of Oxford) - mira.kassouf@imm.ox.ac.uk 

While we develop the ability to write entire chromosomes, we can already engineer those that exist. Modern CRISPR-Cas genome editing technologies enable radical changes to chromosome composition, providing a powerful means to explore the limits of genome function. For example, constructing minimal chromosomes through iterative deletion of non-essential regions can reveal the requirements for chromosome replication and segregation while testing whether much of the human genome is truly "junk", or instead required in some configuration. Beyond these fundamental questions, chromosome-scale engineering could enable applications such as constructing genomes carrying defined polygenic risk profiles to investigate common disease, or removing xenoantigens from donor species to facilitate safer xenotransplantation. 

We recently demonstrated that chromosome evolution can be directed towards sequence configurations unlikely to arise through natural evolution by repeatedly applying targeted mutational processes. Using iterative genome engineering, we created the most extensively engineered human genomes to date by installing recombinase recognition sites into repetitive DNA in HEK293T and HAP1 cells [1]. Over one year, this generated more than 1,600 targeted sequence insertions in a single cell line. We now seek to apply similar principles to engineer chromosomes with exceptional properties over multiple years. 

This combined computational and experimental Ph.D. project will develop an iterative deletion strategy to radically minimise one arm of a human chromosome. You will use AI-assisted tools to design CRISPR-Cas reagents, develop efficient iterative mutagenesis workflows, implement them using the automation infrastructure and technical expertise at the EIT, and comprehensively characterise the resulting genomes. This ability to steer genome evolution towards a desired outcome naturally complements efforts to write synthetic genomes from scratch and could become a robust platform for implementing new biological functions. The extent to which a chromosome can be minimised in practice will shed light on a fundamental open question: are our genomes mostly junk? 

Understanding and Engineering Minimal Signalling Modules that Perform Kinetic Proofreading for T Cell Antigen Discrimination

Lead supervisors: Omer Dushek (Professor of Molecular Immunology, Sir William Dunn School of Pathology, University of Oxford) - omer.dushek@path.ox.ac.uk

Co-supervisor: Linda van Bijsterveldt (Principal Investigator, Generative Biology Institute, EIT) - linda.vanbijsterveldt@eit.org 

T cells provide highly selective immune surveillance by continuously scanning cells for peptide–MHC (pMHC) antigens using their T cell receptors (TCRs). They must respond to high-affinity pathogen or tumour antigens while ignoring the vast excess of lower-affinity self-antigens, even when these are presented at much higher concentrations. We have previously developed quantitative platforms demonstrating that the TCR possesses substantially greater antigen discrimination than other receptor systems, including GPCRs, RTKs, cytokine receptors, and chimeric antigen receptors (CARs) [1]. This remarkable selectivity is thought to arise through kinetic proofreading, a mechanism originally proposed by Hopfield in the context of tRNA decoding, in which a temporal delay between ligand binding and signalling preferentially filters out short-lived, low-affinity interactions. However, despite strong evidence for kinetic proofreading as an operational mechanism, the molecular components and network architecture responsible remain unknown because more than 30 proteins participate in proximal TCR signalling. 

This project will define the minimal signalling module required for kinetic proofreading by reconstituting TCR signalling in heterologous cells. Building on established reconstitution systems, we will introduce the TCR together with candidate signalling molecules using conventional transfection approaches. In parallel, we will deploy megabase-scale DNA assembly platforms [3] to reconstruct increasingly complete signalling networks on a synthetic chromosome [4] and identify the minimal molecular circuitry sufficient to reproduce kinetic proofreading. 

Defining the molecular basis of kinetic proofreading will provide a rational framework for engineering T cells with enhanced antigen discrimination [2]. This has important translational implications for TCR-based cancer immunotherapy, where off-target recognition of low-affinity self-antigens can cause severe toxicity. Precisely tuning kinetic proofreading could enable safer, more selective T-cell therapies with improved therapeutic efficacy. 

  • Pettmann J, Huhn A, Abu Shah E, Kutuzov MA, Wilson DB, Dustin ML, Davis SJ, van der Merwe PA, Dushek O. The discriminatory power of the T cell receptor. Elife. 2021 May 25;10:e67092. doi: 10.7554/eLife.67092. PMID: 34030769; PMCID: PMC8219380. 
  • Cabezas-Caballero J, Huhn A, Kutuzov MA, Andre V, Shomuradova A, Peeters BWA, Gillespie GM, van der Merwe PA, Dushek O. Generation of T cells with reduced off-target cross-reactivities by engineering co-signalling receptors. Nat Biomed Eng. 2026 Apr;10(4):753-764. doi: 10.1038/s41551-025-01563-w. Epub 2026 Jan 2. PMID: 41482591; PMCID: PMC7618719. 
  • Zürcher JF, Kleefeldt AA, Funke LFH, Birnbaum J, Fredens J, Grazioli S, Liu KC, Spinck M, Petris G, Murat P, Rehm FBH, Sale JE, Chin JW. Continuous synthesis of E. coli genome sections and Mb-scale human DNA assembly. Nature. 2023 Jul;619(7970):555-562. doi: 10.1038/s41586-023-06268-1. Epub 2023 Jun 28. PMID: 37380776; PMCID: PMC7614783. 
  • Petris G, Grazioli S, van Bijsterveldt L, Murat P, Liu KC, Birnbaum J, Sale JE, Chin JW. High-fidelity human chromosome transfer and elimination. Science. 2025 Dec 4;390(6777):1038-1043. doi: 10.1126/science.adv9797. Epub 2025 Dec 4. PMID: 41343641. 

Decoding the Hidden Regulatory Grammar of Mammalian Non-Coding Genomes Through Synthetic Regulatory Domain Engineering  

Lead Supervisor: Mira Kassouf (RDM Principal Investigator, Weatherall Institute of Molecular Medicine, University of Oxford) - mira.kassouf@imm.ox.ac.uk

 Co-Supervisor: Leopold Parts (Principal Investigator, Generative Biology Institute, EIT) - leopold.parts@eit.org

Promoters, enhancers and insulators are the principal cis-regulatory elements controlling mammalian gene expression.

Recently, we identified facilitators1, a previously unrecognised class of regulatory sequence that modulates enhancer activity, demonstrating that the regulatory code remains incompletely understood. Despite extensive annotation, most non-coding DNA lies outside currently defined cis-regulatory elements and coding sequences, suggesting that additional regulatory information remains to be discovered. 

This project will use the α-globin locus as a tractable model to determine the contribution of such sequences to gene regulation2. Although the human and mouse α-globin loci share highly conserved coding regions and canonical regulatory elements, the intervening non-coding sequences (INCS) have diverged substantially during evolution and their role in gene regulation is unexplored. Replacement of the ~70kb mouse α-globin regulatory domain with the orthologous human sequence resulted in reduced α-globin expression despite preserving all known regulatory elements3. We previously suggested that evolution could have created a mismatch between the human regulatory elements and the heterologous mouse transcription factors. Alternatively, the INCS may also regulate gene expression3.   

To distinguish between these hypotheses, we will generate synthetic hybrid α-globin domains in which the mouse regulatory elements are retained while INCSs are systematically replaced with orthologous sequences from other species. Synthetic regulatory domains will be computationally designed and synthesised in the Parts laboratory4 before being integrated into the Kassouf laboratory's established α-globin experimental pipeline. The output gene expression and chromatin state measurements will inform which sequences outside mapped enhancers and facilitators have an independent impact on α-globin expression. 

The long-term goal is to determine the role, if any, of the rapidly evolving INCSs. Combining synthetic genomics with predictive modelling will advance our ability not only to understand the regulatory genome more fully but to write functional regulatory sequences from first principles, providing a foundation for predictive genome design and synthetic mammalian genomics. 

  • Blayney et al, Cell 2023 (DOI: 10.1016/j.cell.2023.11.030) 
  • Kassouf et al, 2023 (Review) (https://doi.org/10.1002/bies.202300047
  • Wallace et al, Cell 2007 (DOI: 10.1016/j.cell.2006.11.044) 
  • Koeppel et al, Nature Genetics 2024 (DOI: 10.1038/s41588-024-01981-7) 

Environmentally-Responsive Immunoparticles 

Lead Supervisor: Jason Davis (Professor of Chemistry, Department of Chemistry, University of Oxford) jason.davis@chem.ox.ac.uk

Co-Supervisor: Chang Liu (Principal Investigator, Generative Biology Institute, EIT; Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) – chang.liu@eit.org

 Controlling the spatiotemporal activity of drugs remains a longstanding challenge. While inherent drug properties—such as target specificity and biodistribution—provide some baseline control, precisely engineering these traits is difficult. Moreover, many cancer therapies target cell-surface proteins that are overexpressed in tumors but still present in healthy tissues. This overlap narrows the therapeutic window, as high systemic drug concentrations often trigger on-target, off-tumor toxicity in normal cells. 

A promising strategy to overcome this limitation is to engineer therapeutics or delivery vehicles that remain inert until activated by tissue-penetrating energy directed precisely at the target site. To this end, we aim to develop thermoresponsive drugs and delivery vehicles activated by focused ultrasound (FUS). FUS can non-invasively penetrate deep tissue, safely elevating temperatures in localized region with millimeter resolution. 

Our project entails two primary components. First, we will create general antibody engineering methods to create therapeutics that bind their targets only at elevated temperatures (e.g., 40 °C) while remaining inactive at normal physiological temperatures (e.g., 37 °C). While temperature-sensitive proteins are common, they typically unfold and lose activity at high temperatures—the exact opposite of our goal. Therefore, we will develop novel antibody design and evolution strategies to yield robust, heat-activated protein switches. Second, we will encapsulate these antibodies in activatable delivery vehicles. These will include thermally / hypoxically activateable liposomes and diffusively mobile particle surfaces that are pre-shrouded in hypoxically-responsive polymer coatings. 

Ultimately, integrating these two components should yield therapeutics that can be administered systemically but activated locally by FUS, maximizing efficacy at tumor sites while sparing healthy tissue. 

Further reading:  

  • Bennett, N.R, Watson, J.L, et al, (2026). “Atomically accurate de novo design of antibodies with RE diffusion”. Nature 649: 183. 
  • Duncan, A. M., C. M. Ellis, J. P. Smith, L. Leutloff, M. J. Langton and J. J. Davis (2025). "Organic functionality in responsive paramagnetic nanostructures." Front Chem 13: 1605538. 
  • Kang, S., D. Yuan, R. Barber and J. J. Davis (2024). "Antigen-Mimic Nanoparticles in Ultrasensitive on-Chip Integrated Anti-p53 Antibody Quantification." ACS Sens 9(3): 1475–1481. 
  • Smith, J. P., C. M. Ellis, A. M. Duncan and J. J. Davis (2026). "AND-logic MRI contrast by water flux modulation." Chemical Science, DOI: 10.1039/d6sc01039c 
  • Yuan, D., C. M. Ellis, F. E. Mozes and J. J. Davis (2023). "Ultrahigh magnetic resonance contrast switching with water gated polymer-silica nanoparticles." Chem Commun (Camb) 59(40): 6008–6011. 

Title: Environmentally-responsive immunoparticles 

 

Lead Supervisor: Jason Davis (Professor of Chemistry, Department of Chemistry, University of Oxford) jason.davis@chem.ox.ac.uk

Co-Supervisor: Chang Liu (Principal Investigator, Generative Biology Institute, EIT; Professor of Chemistry and Chemical Biology, Department of Chemistry, University of Oxford) – chang.liu@eit.org 
 

 

Controlling the spatiotemporal activity of drugs remains a longstanding challenge. While inherent drug properties—such as target specificity and biodistribution—provide some baseline control, precisely engineering these traits is difficult. Moreover, many cancer therapies target cell-surface proteins that are overexpressed in tumors but still present in healthy tissues. This overlap narrows the therapeutic window, as high systemic drug concentrations often trigger on-target, off-tumor toxicity in normal cells. 

 

A promising strategy to overcome this limitation is to engineer therapeutics or delivery vehicles that remain inert until activated by tissue-penetrating energy directed precisely at the target site. To this end, we aim to develop thermoresponsive drugs and delivery vehicles activated by focused ultrasound (FUS). FUS can non-invasively penetrate deep tissue, safely elevating temperatures in localized region with millimeter resolution. 

 

Our project entails two primary components. First, we will create general antibody engineering methods to create therapeutics that bind their targets only at elevated temperatures (e.g., 40 °C) while remaining inactive at normal physiological temperatures (e.g., 37 °C). While temperature-sensitive proteins are common, they typically unfold and lose activity at high temperatures—the exact opposite of our goal. Therefore, we will develop novel antibody design and evolution strategies to yield robust, heat-activated protein switches. Second, we will encapsulate these antibodies in activatable delivery vehicles. These will include thermally / hypoxically activateable liposomes and diffusively mobile particle surfaces that are pre-shrouded in hypoxically-responsive polymer coatings. 

 

Ultimately, integrating these two components should yield therapeutics that can be administered systemically but activated locally by FUS, maximizing efficacy at tumor sites while sparing healthy tissue. 

 

Further reading:  

  • Bennett, N.R, Watson, J.L, et al, (2026). “Atomically accurate de novo design of antibodies with RE diffusion”. Nature 649: 183. 

 

  • Duncan, A. M., C. M. Ellis, J. P. Smith, L. Leutloff, M. J. Langton and J. J. Davis (2025). "Organic functionality in responsive paramagnetic nanostructures." Front Chem 13: 1605538. 

 

  • Kang, S., D. Yuan, R. Barber and J. J. Davis (2024). "Antigen-Mimic Nanoparticles in Ultrasensitive on-Chip Integrated Anti-p53 Antibody Quantification." ACS Sens 9(3): 1475–1481. 

 

  • Smith, J. P., C. M. Ellis, A. M. Duncan and J. J. Davis (2026). "AND-logic MRI contrast by water flux modulation." Chemical Science, DOI: 10.1039/d6sc01039c 

 

  • Yuan, D., C. M. Ellis, F. E. Mozes and J. J. Davis (2023). "Ultrahigh magnetic resonance contrast switching with water gated polymer-silica nanoparticles." Chem Commun (Camb) 59(40): 6008–6011. 

 

Orthogonal N-Degron Pathways to Control Responses to Environmental Cues in Plant Cells 

Lead supervisor: Prof Francesco Licausi (Professor of Molecular Plant Physiology, Department of Biology, University of Oxford) – francesco.licausi@biology.ox.ac.uk  

Co-supervisor: Dr Sam Pellock (Principal Investigator, Generative Biology Institute, EIT) - sam.pellock@eit.org  

The N-degron pathways, formerly referred to a N-end rule pathways, are proteolytic systems dedicated to proteome surveillance1. They mediate protein degradation based on the identity of the amino acid exposed at the N-terminus of the polypeptide. They provide a solution for removing damaged and dysfunctional proteins, but these pathways also control the abundance of regulatory proteins in responses to (bio)chemical and physical cues. Well-characterised examples of such regulatory role include responses to oxygen (O₂) and nitric oxide (NO) in plant cells2, thereby acting as major regulator of responses to submergence, drought and high soil salinity. 

In this PhD project, we propose to expand the N-degron pathways through the conjugation and recognition of N-terminal non-canonical amino acids (ncAAs), enabling cue-directed proteolytic control of key regulators of environmental adaptation in plants. 

The PhD candidate will exploit a combination of rational design and directed protein evolution to evolve existing enzymes or engineer novel ones capable of using ncAAs that respond to chemical and physical stimuli as substrates. For example, redox- or light-sensitive amino acids could enable selective interactions with regulatory partners or inducible targeted degradation.

Engineered N-degron enzymes will initially be tested in vitro and in fast-replicating unicellular organisms. Subsequently, their stable expression in planta will be used to evaluate their potential for generating genetic circuits for environmental sensing. 

This project will contribute to the development of smart crops with enhanced resilience to the numerous environmental challenges associated with rapid climate change. Although the initial approach will rely on transgenesis, the resulting strategies could subsequently be implemented through genome editing to generate plants that comply with current legislation. 

  • 10.1042/BST20191094  
  • 10.1016/j.cub.2017.09.006 

Mix It Up: Deleting and Scrambling an Imprinting Control Region to Understand Gene Cluster Regulation and Hierarchical Signalling Cascades 

Lead supervisor: Dr Jacinta Kalisch-Smith (Group Leader and Principal Investigator, Institute of Developmental and Regenerative Medicine, Department of Physiology, Anatomy and Genetics, The University of Oxford) jacinta.kalisch-smith@idrm.ox.ac.uk    

Co-supervisors:  Leopold Parts (Principal Investigator, Generative Biology Institute, EIT) - leopold.parts@eit.org 

  • Dr Brent Ryan (Group leader and Principal Investigator, Department of Physiology, Anatomy and Genetics, The University of Oxford) brent.ryan@dpag.ox.ac.uk 
  • Prof Nicola Smart (Professor of Cardiovascular Science, Fellow and Tutor in Medicine, Christ Church College Oxford, Group Leader and Principal Investigator, Institute of Developmental and Regenerative Medicine, Department of Physiology, Anatomy and Genetics, The University of Oxford) nicola.smart@dpag.ox.ac.uk  

Expression of genes in a tissue and cell type specific manner is a finely controlled process. This specificity is conferred through the interaction of regulatory elements (enhancers, promoters, silencers) and transcription factor combinations which drive the expression of gene modules. These complex interactions together initiate signaling cascades which ultimately form a genome regulatory network (GRN). In the placenta, appropriate expression of gene clusters such as the DLK1-DIO3 imprinting control region (ICR) are necessary for placental development and pregnancy success [1].  

This project will use a genome engineering approach to understand the organisation of the DLK1-DIO3 ICR by using CRISPR-Cas9 gene editing to scramble and alter genes and regulatory elements controlling expression of this ICR. We will create iterative structural variants composed of deletions, inversions, translocations and duplications of whole genes, regulatory elements, or non-coding regions [2-4]. This will inform how the DLK1-DIO3 ICR is established and will give insight into function of this gene cluster, their regulatory element interactions, and how their spacing drives gene expression. In parallel, we will test whether the learnings from this exemplar locus generalize to designing synthetic imprinting programs in a different context. 

The goal of this combined wet lab and computational PhD project is to apply the genomic toolbox on the DLK1-DIO3 ICR to understand the genome architecture in this region, from a broad scale, down to the individual gene to understand specific elements that drive gene expression. Using novel next-generation sequencing datasets, you will learn bioinformatic analysis to predict enhancer interactions and will test these interactions using in vitro models. You will design and synthesise constructs for use in CRISPR prime editing (generating iterative mutations in the same cell line), as well as design candidate cis-acting epigenetic state maintenance sequences, and characterise the resulting outcomes. These mutant cell lines will map regulatory elements, dissect their interactions and construct a minimal circuit (genes/elements) required for building a functional genomic module in the placenta. This understanding can then be used to design an exemplar synthetic imprinting circuit in a cell line.    

Harnessing Organelles as DNA Delivery Systems 

Lead Supervisor: Richard White (Ludwig Institute, organelles and cell-cell communication, University of Oxford) - richard.white@ludwig.ox.ac.uk

Co-supervisor: Linda van Bijsterveldt (Principal Investigator, Generative Biology Institute, EIT) - linda.vanbijsterveldt@eit.org 

A major goal of synthetic biology is to write DNA sequences that enable cells and tissues to perform novel functions. Recent advances in DNA synthesis and assembly have made it possible to construct large genetic payloads, including entire chromosomes, but delivering this cargo into specific target cells remains a major obstacle.  

Organelles, because of their large size and loading capacity, represent a promising vehicle for such delivery. Recent work has demonstrated that certain organelles (i.e. mitochondria, melanosomes) can readily transfer from one cell to another, and carry cargo with them. Importantly, this organelle transfer occurs selectively between specific donor and recipient cell types, suggesting this natural mechanism could be co-opted as a programmable delivery route for synthetic DNA. 

In this project, we aim to explore how we can exploit organelles for the delivery of synthetic DNA cargoes. We envision 3 major long-term goals: 1) how to efficiently assemble large DNA fragments and package them into donor organelles, 2) how to enable transfer of these organelles into recipient cells, and 3) to define the rules governing donor-recipient compatibility, to enable targeted and cell-type-specific delivery. 

More information to come