Explainability & trustworthiness of language models
I study why models say what they say, and whether we can trust the explanation.
My PhD research develops methods to test and improve the trustworthiness of large language models such as GPT and Llama with a particular focus on whether the explanations these models give for their own predictions actually hold up.
About
I'm a PhD student in Machine Learning and Natural Language Processing at the Department of Computer and Systems Sciences, Stockholm University. I'm part of the Data Science Research Group and the Natural Language Processing Research Group.
Before starting my PhD, I completed a Master's in Artificial Intelligence at Stockholm University and a Master's in Electrical Engineering and Information Technology at the Technical University of Munich. I also worked eight years as an electrical engineer specializing in automation technology.
Beyond my core research, I'm interested in the application of information technology to biology and in projects that serve the social good.
Research focus
Explainability & interpretability
Broader work on making machine learning predictions legible to the people who rely on them. One strand of this looks at self-explanations: when a language model is asked to explain its own output, is that explanation actually faithful to how the model arrived at its prediction, or just plausible-sounding? My findings suggest faithfulness is not guaranteed.
More trustworthy prompting methods
Work on making prompting pipelines principled rather than trial-and-error. RAG-E makes retrieval-augmented generation more transparent, using attribution methods to audit whether a generator's answer actually draws on the documents its retriever ranked as most relevant, exposing failure modes that per-component metrics miss. CICLe uses conformal prediction to decide, case by case, which classification queries a large language model can be trusted to answer directly and which need a fallback.
Applied ML & NLP
Applying these methods to real domains: food-safety hazard detection from incident reports, automotive fault detection from technician notes, and clinical data.
Publications
- Quantifying Retriever-Generator Alignment in RAG with Local Explanations
- Mind the gap: from plausible to valid self-explanations in large language models
- Evaluating the reliability of self-explanations in large language models
- CICLe: conformal in-context learning for large-scale multi-class food risk classification
- SemEval-2025 Task 9: the food hazard detection challenge
- Early prediction of the risk of ICU mortality with deep federated learning
See the full profile for an up-to-date list.
Teaching
- Machine Learning Lecturer & course assistant · since spring 2024
- Introduction to Programming Lecturer & course assistant · since fall 2024
- Project management & tools for health informatics Lecturer & course assistant · since fall 2024
- Principles and Foundations of AI Administration, lecturer & course assistant · fall 2023