Name of participant: Erik Kubaczka
Project’s name: DeeBOP: Deep Batch Bayesian Optimization for the design of sequences in biology
Project description:
In biology, sequences form the basis for a wide variety of functions. Whether it is the regulation of gene expression, the functions of proteins or the identification of diseased cells, all of these are controlled on the basis of DNA, RNA or amino acid sequences. By modifying existing sequences or designing new ones, we can programme new functions into cells or adapt existing mechanisms to new tasks.
In my DeeBOP project, we at TU Darmstadt are collaborating with Merck KGaA to design sequences using artificial intelligence. This is based on an approach that alternates between the design of new sequences and their experimental characterisation, thereby enabling our method to learn from previous designs. The new sequences are proposed using Bayesian optimisation. In doing so, we investigate how the inherently sequential method of Bayesian optimisation can be combined with modern high-throughput approaches in biology. These high-throughput approaches can characterise hundreds to thousands of constructs simultaneously. However, this places high demands on the selection of suitable sequences, as these must, on the one hand, perform the desired function optimally and, on the other hand, cover the sequence space as diversely as possible. As part of the DeeBOP project, we are developing new strategies for Bayesian optimisation in the high-throughput sector, thereby paving the way for its successful application in the biology of the future.
Software Campus Partner: Technical University of Darmstadt & Merck KGaA
Implementation period: 01.03.2025 – 28.02.2027




































![[KOM,BI]Co-citation-based machine learning to determine promising research projects](https://softwarecampus.de/wp-content/uploads/2022/02/KOMBI-768x768.png)




































































