thomasfritschtech
AI Engineering · Remote

Thomas FritschAI Engineer for RAG & self-hosted LLMs

Retrieval systems that cite their sources, and the infrastructure to run them on your own hardware.

Responsible for the whole chain: retrieval, model operations and the application around them — from the architecture decision to the point where the system is running in production.

Capabilities

Services & Expertise

RAG, LLMOps, self-hosted inference and agent workflows — built to run in production, not just in a notebook.

RAG Systems

Development of intelligent Retrieval-Augmented Generation systems for precise, context-based AI responses

Key Capabilities
  • Custom Embedding Strategies
  • Vector Database Integration
  • Semantic Search Optimization
  • Context-Aware Retrieval
Tech Stack
LangChainChromaDBPineconeFAISS
What this involves

Chunking and embedding strategy, vector store in ChromaDB or FAISS, retrieval evaluated against a fixed reference question set, and answers traceable back to the source passage.

LLMOps

End-to-end operationalization of Large Language Models for production environments

Key Capabilities
  • Model Deployment & Serving
  • Prompt Engineering & Versioning
  • Fine-tuning Pipelines
  • Cost Optimization
Tech Stack
LangSmithWeights & BiasesMLflowFastAPI
What this involves

Versioned prompts, reproducible deployments via Docker, cost and latency tracing per request, and regression tests against a fixed evaluation set before every rollout.

Self-Hosted AI Infrastructures

Secure, privacy-compliant AI solutions on your own hardware or private cloud

Key Capabilities
  • On-Premise Model Deployment
  • Data Privacy & Compliance
  • Custom Model Training
  • GPU Optimization
Tech Stack
OllamavLLMTGILocalAI
What this involves

Model serving on your own GPU hardware or EU hosting, access control, and monitoring — data never leaves your infrastructure, which is what makes GDPR compliance tractable.

AI Automation & Agents

Development of autonomous AI agents for complex workflow automation

Key Capabilities
  • Multi-Agent Systems
  • Tool & API Integration
  • Autonomous Decision Making
  • Workflow Orchestration
Tech Stack
LangGraphAutoGenCrewAIn8n
What this involves

Agent workflows with tool and API integration, explicit error paths instead of silent failure, and human-in-the-loop approval at the steps where a wrong call gets expensive.

Reference architecture · RAG
SOURCE DATAEMBEDDINGVECTOR STORERETRIEVALLLM

Source data is embedded and stored in the vector store; at runtime, retrieval passes only the relevant passages to the model.

Scope

Web apps end to end

Not just the AI part, but the whole product: frontend, backend, database, deployment. From the first line to running in production.

Frontend

Responsive web interfaces, component libraries, accessible interactionReactNext.jsTypeScript

Backend & API

REST and realtime APIs, authentication, background jobs, integrationsFastAPIPythonRedis

Database

Data modeling, migrations, indexing — including vector search where AI is involvedPostgreSQLpgvectorSupabase

Deployment

Containers, CI/CD, monitoring and operations on EU infrastructure instead of hyperscalersDockerHetznerGrafana
Approach

All layers from one hand — no handovers between frontend, backend and infrastructure teams. That keeps decisions consistent and the path from idea to deployment short.

In practice

Tournio is exactly this stack in real use: React frontend, FastAPI backend, PostgreSQL, Redis, Hetzner hosting — built solo, tested live.

Case study

Featured Project

Right now my focus is on a single product — built end to end, on my own.

SaaS

Tournio

From idea to product

Currently building

I'm building Tournio from the ground up — solo, no team, no investor.

Tournio is a SaaS platform for sports tournaments. Organizers create tournaments, manage schedules and results in real time — and embed the live view directly into their own website via iframe. Billing is per tournament, no subscription.

Its first real-world run was at the end of June 2026: the Busse-Kamine Beach Soccer Cup in Neubrandenburg. Live results on the big screen on site, the tournament table on the spectators' phones — all powered by Tournio.

model
pay-per-tournament
hosting
EU / Hetzner
team
solo
Tech stack
ReactFastAPIPostgreSQLRedisHetzner
Status

First live test completed. Next up: payment integration, public beta launch, more sports.

Background

About

From retrieval pipeline to running deployment

Thomas Fritsch — TFT monogram

Thomas Fritsch

Building software since 2000, with a background in web and graphic design — that combination is where the engineering precision and the eye for form both come from.

Today that work centres on production-grade AI systems with a bias toward cost control, data residency in the EU, and architectures that stay debuggable once they are live — retrieval-augmented generation, autonomous agents, and the infrastructure underneath them.

2000
Building software since
2022
Focused on AI since
Learning mode

Journey & Milestones

2024
AI Infrastructure Specialist

Building production RAG systems and self-hosted LLM stacks

2023
LLMOps & Automation

Implementing autonomous AI agents and workflow orchestration

2022
Full-Stack AI Development

Combining traditional software engineering with emerging AI technologies

2020–21
Software Engineering

Building scalable web applications and cloud-native systems

since 2000
Software Development, Web & Graphic Design

Over two decades of practice in development and design — the foundation for everything since

Core Skills & Technologies

Retrieval
ChromaDBFAISSpgvectorPinecone
Orchestration
LangChainLangGraphCrewAI
Backend
PythonFastAPIPostgreSQLRedis
Self-hosted inference
vLLMOllamaTGI
Operations
DockerKubernetesGrafana
Infrastructure
HetznerVercelSupabase
Current Focus

Building cost-effective RAG systems with advanced retrieval strategies, optimizing self-hosted LLM infrastructures, and exploring multi-agent architectures for complex automation workflows.

Availability

Get In Touch

Tell me about the system you want to build, and I will tell you what it takes.

Send a Message

Fill out the form below and I'll get back to you within 24 hours

Contact Information

GitHub@tomex74
Available for projects

Typical response time: 24-48 hours

Available
Thomas Fritsch — AI Engineer for RAG & Self-Hosted LLMs