Portrait of Maximilian Dreyer

Maximilian Dreyer

Interpretability for safe and reliable AI

EmailGoogle ScholarGitHubLinkedInCV (PDF)

I build methods that look inside AI models. They reveal which concepts a model has learned and relies on, uncover hidden flaws such as shortcuts and biases, and fix them directly inside the model. With layerwise, I bring this into practice by monitoring and steering AI systems while they run.

Co-founder · with Reduan Achtibat and Oleg Hein

Interpretability, in production.

layerwise is a Fraunhofer HHI spin-off building interpretable AI tooling for teams that need to understand model behaviour before it reaches real-world decisions. Instead of judging outputs after the fact, we read and steer a model’s internal signals at inference time to detect and correct failures such as hallucination, manipulation and prompt injection in language models and agents.

Visit layerwise.ai Built on AttnLRP, SemanticLens, CRP

news

research

AI models are trained, not engineered, so nobody specifies what their inner parts do. My research makes them understandable, checkable and fixable.

  1. 1.Understand
    Concept Relevance Propagation

    Understanding decisions through the concepts a model uses.

    89% vs. 54%

    of users spotted a flawed model with our explanations, vs. standard heatmaps

  2. 2.Audit
    SemanticLens & PCX

    Searching a model’s inner parts by text and checking what it has learned against expert knowledge.

    0 of 27

    image classes in a widely used model relied only on valid features

  3. 3.Correct
    RR-ClArC & Reveal2Revise

    Removing unwanted behaviour inside the model and checking the fix, without needing human annotations.

    up to +28%

    accuracy on biased test data after correction

work

Selected publications

  • Preprint 2026Senior author

    ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

    Cancellations in residual pathways cause attributions in ViTs to explode. Treating residual connections correctly yields fine-grained, faithful and stable explanations.

    Jim Berend, Reduan Achtibat, Daniel Schäffer, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer

  • Nature Machine Intelligence 2025

    Mechanistic understanding and validation of large AI models with SemanticLens

    A search engine for the inside of AI models: every component is mapped into a shared space of images and language, so a model can be searched, described, compared and audited.

    Maximilian Dreyer*, Jim Berend*, Tobias Labarta, Johanna Vielhaben, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek

  • ICLR 2026

    Dyslexify: A mechanistic defense against typographic attacks in CLIP

    Locates the attention-head circuit that reads text in images and ablates it, making CLIP robust to typographic attacks.

    Lorenz Hufe, Constantin Venhoff, Maximilian Dreyer, Erblina Purelku, Sebastian Lapuschkin, Wojciech Samek

  • Preprint 2025

    From What to How: Attributing CLIP’s latent components reveals unexpected semantic reliance

    Combines sparse autoencoders with attribution patching to reveal not only what CLIP’s components encode, but how they drive its predictions.

    Maximilian Dreyer, Lorenz Hufe, Jim Berend, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek

  • ICML 2024

    AttnLRP: Attention-aware layer-wise relevance propagation for transformers

    Extends LRP to attention layers, enabling faithful and efficient explanations of LLMs and Vision Transformers.

    Reduan Achtibat, Sayed M. V. Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek

  • Nature Machine Intelligence 2023

    From attribution maps to human-understandable explanations through Concept Relevance Propagation

    Explains not only where a model looks, but what it sees there, by attributing decisions to the concepts encoded in its hidden units.

    Reduan Achtibat*, Maximilian Dreyer*, Ilona Eisenbraun, Sebastian Bosse, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

  • AAAI 2024

    From hope to safety: Unlearning biases of deep models via gradient penalization in latent space

    RR-ClArC unlearns spurious concepts by penalizing the model’s sensitivity to them in latent space.

    Maximilian Dreyer*, Frederik Pahde*, Christopher J. Anders, Wojciech Samek, Sebastian Lapuschkin

  • MICCAI 2023

    Reveal to Revise: An explainable AI life cycle for iterative bias correction of deep models

    An iterative XAI life cycle that reveals, corrects and re-checks model biases, demonstrated on medical imaging.

    Frederik Pahde*, Maximilian Dreyer*, Wojciech Samek, Sebastian Lapuschkin

  • CVPR Workshops 2024

    Understanding the (extra-)ordinary: Validating deep model decisions with prototypical concept-based explanations

    Concept-based prototypes summarize a model’s strategies, so unusual behaviour and data-quality issues stand out automatically.

    Maximilian Dreyer, Reduan Achtibat, Wojciech Samek, Sebastian Lapuschkin

  • CVPR Workshops 2024Spotlight

    PURE: Turning polysemantic neurons into pure features by identifying relevant circuits

    Disentangles polysemantic units into monosemantic features via their circuits, increasing latent interpretability.

    Maximilian Dreyer, Erblina Purelku, Johanna Vielhaben, Wojciech Samek, Sebastian Lapuschkin

More publications

  • Contrastive Semantic Projection: Faithful neuron labeling with contrastive examples

    Oussama Bouanani, Jim Berend, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer

    World Conference on XAI 2026 · Senior author

  • From attribution to action: A human-centered application of activation steering

    Tobias Labarta, Maximilian Dreyer, Katharina Weitz, Wojciech Samek, Sebastian Lapuschkin

    CVPR Workshops 2026

  • Navigating neural space: Revisiting concept activation vectors to overcome directional divergence

    Frederik Pahde, Maximilian Dreyer, Moritz Weckbecker, Leander Weber, Christopher J. Anders, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

    ICLR 2025

  • Attribution-guided pruning for insight and control: Circuit discovery and targeted correction in small-scale LLMs

    Sayed M. V. Hatefi, Maximilian Dreyer, Reduan Achtibat, Patrick Kahardipraja, Thomas Wiegand, Wojciech Samek, Alexander Binder, Sebastian Lapuschkin

    Preprint 2025

  • Pruning by explaining revisited: Optimizing attribution methods to prune CNNs and transformers

    Sayed M. V. Hatefi, Maximilian Dreyer, Reduan Achtibat, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

    ECCV Workshops 2024

  • Reactive model correction: Mitigating harm to task-relevant features via conditional bias suppression

    Dilyara Bareeva, Maximilian Dreyer, Frederik Pahde, Wojciech Samek, Sebastian Lapuschkin

    CVPR Workshops 2024

  • Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification

    Christian Tinauer, Anna Damulina, Maximilian Sackl, Martin Soellradl, Reduan Achtibat, Maximilian Dreyer, Frederik Pahde, Sebastian Lapuschkin, et al.

    World Conference on XAI 2024

  • Revealing hidden context bias in segmentation and object detection through concept-specific explanations

    Maximilian Dreyer, Reduan Achtibat, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

    CVPR Workshops 2023

  • Explainability-driven quantization for low-bit and sparse DNNs

    Daniel Becking, Maximilian Dreyer, Wojciech Samek, Karsten Müller, Sebastian Lapuschkin

    xxAI – Beyond Explainable AI 2022

* equal contribution. Full list on Google Scholar.

talks

cv

Positions

  1. since 2026

    Co-founder

    layerwise

    Fraunhofer HHI spin-off reading and steering a model’s internal signals to detect and correct failures such as hallucinations and prompt injection.

  2. since 2026

    Postdoctoral Researcher

    Fraunhofer HHI

    Explainable AI Group (Dr. S. Lapuschkin), Department of Artificial Intelligence (Prof. W. Samek).

  3. 2022 – 2026

    Doctoral Researcher

    Fraunhofer HHI

    Concept-based explanations, model auditing and bias correction; medical imaging applications.

  4. 2020 – 2022

    Student Research Assistant

    Fraunhofer HHI

    XAI for segmentation and object detection, explainability-driven quantization.

  5. 2016 – 2020

    Student positions

    IAV Berlin, Baumer Hübner, DESY Zeuthen

    Autonomous-driving simulation, sensor metrology, teaching.

Education

  1. 2026

    Dr. rer. nat., summa cum laude

    Technische Universität Berlin

    From Local Explanations to Comprehensive Mechanistic Understanding of Deep Vision Models. Examiners: W. Samek, P. Biecek (Warsaw University of Technology), J. Choi (KAIST).

  2. 2022

    M.Sc. Computational Science

    University of Potsdam

    Thesis on concept-based explanations for segmentation and object detection, awarded by the Friends of Fraunhofer HHI.

  3. 2019

    B.Sc. Physics

    Humboldt-Universität zu Berlin

    Exchange semester at Uppsala University, 2018.

Patents
Inventor on three Fraunhofer patent families: SemanticLens (lead inventor), Concept Relevance Propagation and AttnLRP.
Software
Open-source implementations of all methods, including zennit-crp, LXT (LRP for transformers) and semanticlens, with about 530 GitHub stars.
Mentoring
Supervised nine Master’s and doctoral students, including as senior author on their work.
Award
Young Talent Award of the Friends of Fraunhofer HHI for the Master’s thesis, 2022.
Download full CV (PDF)

contact

I live in Berlin and work on making AI systems understandable, monitorable and correctable, in research at Fraunhofer HHI and in practice with layerwise. I’m always happy to talk about interpretability, AI safety or collaborations.