>_ Tobias Sterbak
Freelance
Local AI · Private & open

Your AI, on your hardware.

I design and build local and private AI systems: on-premises LLM inference, retrieval over your own documents, fine-tuned open models and the infrastructure to run them — without sending data to third parties.

01 / Local AI

What I build.

Practical, self-hosted AI for teams that need privacy, control and predictable cost.

01

On-prem LLM inference

llama.cpp · vLLM · ollama

Server- and edge-side deployment of open models for chat, classification and structured extraction, sized and tuned for your hardware and latency budget.

02

Private RAG & search

pgvector · LanceDB · embeddings

Retrieval-augmented generation over your own documents and databases, with local embeddings so sensitive content never leaves your network.

03

Fine-tuning & eval

LoRA · QLoRA · PEFT

Adapting open models to your domain and tone with efficient fine-tuning, plus evaluation harnesses that measure what matters before you ship.

04

Local AI infrastructure

Docker · GPU · MLOps

Reproducible stacks for running and monitoring models on your own GPUs or workstations — reliable, observable and easy for your team to operate.

02 / Tooling

Open-source stack I work with.

The ecosystem that makes local AI practical — and my defaults for new projects.

03 / Approach

Why local AI makes sense.

A pragmatic look at when self-hosted models beat the cloud.

Privacy

Data stays in-house

Sensitive documents and queries never leave your network — no third-party API agreements or data retention questions.

Cost

Predictable infrastructure costs

Replace variable API bills with infrastructure you control and costs that become easier to forecast at sustained usage.

Control

Full control

Swap models, tune them and change the stack when you want — no lock-in to a vendor's roadmap.

Offline

Works offline

Reliable inference even without connectivity — ideal for field use, air-gapped environments and reliable automation.

Contact

Have something useful to build?

Need help selecting models and hardware, building local RAG or putting self-hosted inference into production? Tell me what you're working on.