About Me
I am a Senior ML Engineer working on production GenAI systems, with a focus on evaluation, reliability, and LLMOps. At IABG, I lead and build systems such as chat assistants, RAG-based document search, meeting note-taking, and code-review automation. Much of my work is about making these systems measurable: defining evaluation criteria, monitoring product and model behavior, and deciding when a system is ready to move from test to broader rollout.
Before that, I completed a PhD at the Technical University of Munich on AI-based failure management for large-scale cloud systems. That research focused on failure prediction, root-cause analysis, and operational risk classification using logs, metrics, and other operational data. It still shapes how I think about ML systems: define the expected behavior, measure deviations, monitor the right signals, and use that evidence to improve reliability, safety, and deployment decisions.
Outside my day-to-day work, I like projects where ML meets a concrete system or interaction: reinforcement learning and games, IoT and home automation, computer vision for physical-world workflows, and NLP or language analysis. I tend to enjoy the full loop: modeling the problem, building the software around it, instrumenting it, and seeing where it breaks.
For contact, use email or LinkedIn.
Core Capabilities
Production-ready GenAI & LLMOps
LLM eval harnesses
LLM APIs
Guardrails
RAG
Hybrid retrieval
Embedding models
LoRA / PEFT
LangChain
HF Transformers
vLLM
Applied Machine Learning
PyTorch
sklearn
OpenCV
NLTK
PyG / GNNs
RL
Gym
Time-series analysis
Tabular foundation models
Pattern mining
MLOps, Platforms & Observability
MLflow
DVC
W&B
FastAPI
Docker
Kubernetes / Helm
ArgoCD
GitLab CI/CD
GitHub Actions
Prometheus
Grafana
Elasticsearch
Kibana
Software Engineering
Python
SQL
Bash
C++
Poetry / uv
Ruff
ty
Pylint
pytest
Security, Privacy & Governance
Confidential computing
Differential privacy
ML privacy attacks
Threat modeling
AI governance
Experience
Senior ML Engineer, IABG
October 2023 — present
Build and maintain an LLM evaluation platform/library for compliance review in regulated settings and continuous delivery of GenAI products, with a multi-dimensional taxonomy covering correctness, grounding and attribution, retrieval quality, safety and toxicity, privacy leakage, security, latency, and robustness. Build and ship production GenAI services including chat assistants, RAG-based document search, and code-review automation, with retrieval pipelines, prompt and tool orchestration, structured outputs, input/output guardrails, vLLM serving on Kubernetes, versioned CI/CD, and canary rollouts. Define KPI/SLA-driven monitoring for GenAI products with Grafana/Prometheus, and use post-incident reviews to close operational gaps. Project lead of SPARTA, an academic collaboration project on social-media monitoring and analysis, with work on political stance detection and malicious-account detection using vision-language embeddings and graph neural networks. Tech lead & core developer for RESISTANT, a zero-trust avionics research project, designing attack-graph modeling via partially-observable Markov decision processes, particle-filter-based risk assessment, and TabPFN-based tabular modeling for adaptive, security-critical aircraft systems. Engineer privacy and safety measures for sensitive AI deployments, including differential privacy, homomorphic encryption, trusted execution environments, and hardening against membership-inference, extraction, and inversion attacks. Mentor colleagues on research directions, clean-code practices, evaluation design, and architectural decisions for reliable and secure GenAI systems.
PhD Candidate, Huawei Munich Research Center + Technical University of Munich
January 2020 — September 2023
Researched AI-based failure management for large-scale cloud systems, covering online failure prediction, root-cause analysis, and operational risk classification. Published the first-ever survey on AIOps methods for failure management (ACM TIST), organizing the field by data requirements, intervention windows, and target problems. Built LogRule, a structured-log root-cause analysis method published in IEEE TNSM, achieving 37x faster runtime and improved explanation accuracy. Developed reliability and failure-prediction models for cloud infrastructure, including work on optical transceiver reliability recognized as an IEEE/ACM CCGRID Best Paper nominee. Developed an LLM-based command-line risk classifier for operational security workflows; lead inventor on the related filed patent. Mentored students through seminars, master theses, and guided research projects.
Education
TUM
PhD, Computer Science
Technical University of Munich · 2024
AI-based proactive failure management in large-scale cloud environments.
TUM
MSc, Computer Science
Technical University of Munich · 2019
Machine learning, deep learning, computer vision, and NLP.
PoliTo
BSc, Computer Engineering
Politecnico di Torino · 2017
Computer engineering foundations across software, systems, and applied ML.