
Thomas FritschAI Engineer for RAG & self-hosted LLMs
Retrieval systems that cite their sources, and the infrastructure to run them on your own hardware.
Responsible for the whole chain: retrieval, model operations and the application around them — from the architecture decision to the point where the system is running in production.
Services & Expertise
RAG, LLMOps, self-hosted inference and agent workflows — built to run in production, not just in a notebook.
RAG Systems
Development of intelligent Retrieval-Augmented Generation systems for precise, context-based AI responses
- Custom Embedding Strategies
- Vector Database Integration
- Semantic Search Optimization
- Context-Aware Retrieval
Chunking and embedding strategy, vector store in ChromaDB or FAISS, retrieval evaluated against a fixed reference question set, and answers traceable back to the source passage.
LLMOps
End-to-end operationalization of Large Language Models for production environments
- Model Deployment & Serving
- Prompt Engineering & Versioning
- Fine-tuning Pipelines
- Cost Optimization
Versioned prompts, reproducible deployments via Docker, cost and latency tracing per request, and regression tests against a fixed evaluation set before every rollout.
Self-Hosted AI Infrastructures
Secure, privacy-compliant AI solutions on your own hardware or private cloud
- On-Premise Model Deployment
- Data Privacy & Compliance
- Custom Model Training
- GPU Optimization
Model serving on your own GPU hardware or EU hosting, access control, and monitoring — data never leaves your infrastructure, which is what makes GDPR compliance tractable.
AI Automation & Agents
Development of autonomous AI agents for complex workflow automation
- Multi-Agent Systems
- Tool & API Integration
- Autonomous Decision Making
- Workflow Orchestration
Agent workflows with tool and API integration, explicit error paths instead of silent failure, and human-in-the-loop approval at the steps where a wrong call gets expensive.
Source data is embedded and stored in the vector store; at runtime, retrieval passes only the relevant passages to the model.
Web apps end to end
Not just the AI part, but the whole product: frontend, backend, database, deployment. From the first line to running in production.
Frontend
Responsive web interfaces, component libraries, accessible interactionReactNext.jsTypeScriptBackend & API
REST and realtime APIs, authentication, background jobs, integrationsFastAPIPythonRedisDatabase
Data modeling, migrations, indexing — including vector search where AI is involvedPostgreSQLpgvectorSupabaseDeployment
Containers, CI/CD, monitoring and operations on EU infrastructure instead of hyperscalersDockerHetznerGrafanaAll layers from one hand — no handovers between frontend, backend and infrastructure teams. That keeps decisions consistent and the path from idea to deployment short.
Tournio is exactly this stack in real use: React frontend, FastAPI backend, PostgreSQL, Redis, Hetzner hosting — built solo, tested live.
Featured Project
Right now my focus is on a single product — built end to end, on my own.
Tournio
From idea to product
I'm building Tournio from the ground up — solo, no team, no investor.
Tournio is a SaaS platform for sports tournaments. Organizers create tournaments, manage schedules and results in real time — and embed the live view directly into their own website via iframe. Billing is per tournament, no subscription.
Its first real-world run was at the end of June 2026: the Busse-Kamine Beach Soccer Cup in Neubrandenburg. Live results on the big screen on site, the tournament table on the spectators' phones — all powered by Tournio.
First live test completed. Next up: payment integration, public beta launch, more sports.
About
From retrieval pipeline to running deployment
Thomas Fritsch
Building software since 2000, with a background in web and graphic design — that combination is where the engineering precision and the eye for form both come from.
Today that work centres on production-grade AI systems with a bias toward cost control, data residency in the EU, and architectures that stay debuggable once they are live — retrieval-augmented generation, autonomous agents, and the infrastructure underneath them.
Journey & Milestones
Building production RAG systems and self-hosted LLM stacks
Implementing autonomous AI agents and workflow orchestration
Combining traditional software engineering with emerging AI technologies
Building scalable web applications and cloud-native systems
Over two decades of practice in development and design — the foundation for everything since
Core Skills & Technologies
Building cost-effective RAG systems with advanced retrieval strategies, optimizing self-hosted LLM infrastructures, and exploring multi-agent architectures for complex automation workflows.
Get In Touch
Tell me about the system you want to build, and I will tell you what it takes.
Send a Message
Fill out the form below and I'll get back to you within 24 hours
Typical response time: 24-48 hours