
Two copies of one small AI model keep a remote machine in sync with the cloud, with no constant streaming. A beginner-friendly world model tutorial in PyTorch.

Computer Scientist (PhD, TUM) and Senior AI Researcher specializing in ML systems. Along with a portfolio of AI patents, my expertise spans hands-on LLM architecture, large-scale ML deployments, LLM inference optimization and several open-source contributions (incl. vLLM). I uniquely combine deep technical execution with IP strategy and bridge the gap between abstract mathematical models and physical silicon execution to build and protect the next generation of AI compute.

Two copies of one small AI model keep a remote machine in sync with the cloud, with no constant streaming. A beginner-friendly world model tutorial in PyTorch.

Turn a small open-source Qwen LLM into a fast, single-pass text classifier by swapping its language-modeling head

A beginner-friendly guide to building a world model in Python, letting it daydream its way through CartPole, and accurately measuring when the illusion collapses.

A hand-written CUDA inference runtime for Vision-Language-Action robots that decides what to remember, what to forget, and when it's simply too late to think.

Take Google's Open Knowledge Format (a Markdown+YAML skeleton for humans and agents), bolt one extra load-bearing field onto the frontmatter, and hand pre-tokenized integer arrays between three Qwen2.5-Coder agents (7B, 3B, 1.5B) through `/dev/shm`. 28–37% faster TTFT — and one runtime check that stands between "faster" and "confidently fluent nonsense."

In five to ten years, the sharpest manager in your company might not be human, might not sleep, and might exist entirely in shared GPU memory. Here is the systems engineering that has to land first.

Demystifying the LLM runtime: A from-scratch tutorial on building a custom C++/CUDA inference engine, understanding how bare-metal AI actually works, and meeting the synchronization bugs on the way.

Why the future of AI memory relies on persistent neural state, not vector databases.

Stop passing prompt strings between agents. How to use a β-VAE and a gated MLP to persist context across the hand-off boundary.

How a tiny C++ daemon uses 5G-style admission control and asynchronous layer pipelining to run three LLMs on an 8-year-old graphics card.