Name of participant: Jan-Micha Bodensohn
Project’s name: ADEE: Automated Data Engineering in Enterprises with Large Language Models
Project description:
Large enterprises generate vast amounts of data. Every action, every process and every business transaction leaves a trace in the enterprise data, which often hides valuable insights. What costs have been incurred? How is demand developing? Where are bottlenecks or failures? To make use of this information, the relevant data must first be prepared for each application. The process which finds the required data and prepares it for an application is known as data engineering.
The fragmentation of enterprise data across a large number of systems makes data engineering a challenging task. While some data is stored in relational databases with well-designed schemes, much data resides only in so called data lakes, which store raw data in its original form without any uniform structure. Many projects therefore begin with a lengthy search for usable data, which must then be extracted, cleaned, and integrated. Even today, this process still involves a great deal of manual effort and requires technical skills as well as concrete knowledge of how the data is organized within the company.
The goal of the Software Campus project ADEE is therefore to fully automate data engineering in enterprise settings by automatically finding the required data for a given application and then preparing it for its downstream use. Rather than manually searching for the right data and implementing programs to process it, in our vision of automated data engineering, the user simply describes the data they need regardless of how it is actually stored within the company.
To find the required data and prepare it for its application, the project relies on Large Language Models (LLMs), which interact with the enterprise data as agents and create execution plans that prepare the raw data for its specific application. To achieve this, the project first investigates how LLM agents can efficiently interact with large amounts of data to solve isolated data problems. It then develops strategies that enable agents to solve complex data problems and thus automate the data engineering process end-to-end.
Software Campus Partner: TU Darmstadt & Celonis SE
Implementation period: 01.01.2026 – 30.06.2027



































![[KOM,BI]Co-citation-based machine learning to determine promising research projects](https://softwarecampus.de/wp-content/uploads/2022/02/KOMBI-768x768.png)





































































