Name of participant: Joshua Oehms
Project’s name: INDEX: INcorporating Domain-specific concepts and EXtracted knowledge for structured document representation in medical literature
Project description:
The INDEX research project addresses the fundamental challenge that much of the ever-growing body of medical knowledge is available only in unstructured natural language. Although knowledge has always been documented in writing, this content is often inaccessible or can only be utilised at an extremely high cost in terms of manpower. Structured searching and the aggregation of published content play a central role, particularly in systematic medical reviews and meta-studies. In this context, defined search criteria often have complex semantic meanings that cannot be adequately covered by simple keyword searches. Despite efforts at formalisation, much of this information remains difficult for researchers to access, which complicates evidence-based medical decision-making.
A major hurdle lies in the wide linguistic variation found in medical data, as technical terms, brand names and chemical designations are often used inconsistently. Digital search methods are generally unable to fully capture these semantic relationships without context sensitivity. Whilst modern NLP technologies assist with machine processing through vector embeddings, they frequently reach their limits when dealing with extensive scientific articles and complex semantic relationships. Crucial information is often obscured by a flood of irrelevant details in the full text, or ambiguous concepts are incorrectly assigned.
Against this backdrop, the INDEX project is investigating how the limitations of current vector- and keyword-based retrieval systems can be overcome by using domain-specific semantic structures. The project’s focus is on enhancing machine knowledge representation by directly incorporating structuring concepts into data models. To this end, the project is exploring various approaches, including hybrid strategies that combine vector-based similarity searches with the explicit extraction of biomedical features.
Software Campus Partner: TU München und Holtzbrinck
Implementation period: 01.01.2026 – 31.12.2027




































![[KOM,BI]Co-citation-based machine learning to determine promising research projects](https://softwarecampus.de/wp-content/uploads/2022/02/KOMBI-768x768.png)




































































