§04 Services

AI integration

AI that answers from your data, with measured quality and predictable cost — instead of a flashy demo that never reaches production.

PoCin 2–3 weeks
evalquality measured by tests
€cost per query under control
Typical problems I solve

Sound familiar?

Company knowledge is scattered across PDFs, Confluence and email — nobody finds it.

The chatbot hallucinates and there’s no way to evaluate answer quality.

Model API costs grow without control.

Repetitive document work (classification, extraction) ties up people.

Reference architecture

What it usually looks like inside

ORCHESTRATIONMODELS & KNOWLEDGEUserchat · API · agentAPIauth · rate limitLLM routercost ↔ qualityRetrieverhybrid + rerankModelsClaude · OllamaSemantic cacheRedispgvectorembeddingsKnowledgePDF · Confluence · DBDWG · AI-01 · nightdev
  1. Userchat · API · agent
  2. APIauth · rate limit
  3. LLM routercost ↔ qualityRetrieverhybrid + rerank
  4. ModelsClaude · OllamaSemantic cacheRedispgvectorembeddingsKnowledgePDF · Confluence · DB
Reference architecture · AI integration
What you get

The scope you get

RAG on your data

Hybrid search + re-ranking, source citations, access control.

Tool-using agents

Agents calling your APIs and systems (MCP), with human-in-the-loop.

Model routing

A cheaper model where it’s enough, a stronger one where it matters.

Evaluation

Test sets and metrics (e.g. RAGAS) — quality in numbers.

Privacy

Local models (Ollama) or EU hosting when data can’t leave the company.

FinOps

Semantic cache, limits and a cost dashboard per feature.

Stack

Technologies in this area

  • Python
  • FastAPI
  • PostgreSQL
  • Redis
  • Claude
  • Ollama
  • LangChain
Questions

AI integration — FAQ

01Where do we start?

With one process that has a measurable outcome. We build a proof of concept on your data and evaluate quality and cost.

02Will my data go to OpenAI/Anthropic?

Only if you agree. Alternatively I use local or EU-hosted models.

03How do you measure quality?

We build a set of questions and expected answers, and every change is tested automatically (faithfulness, relevance, cost, latency).

Got a project that has to be fast and work at scale?

Describe it in 2 minutes. I'll reply within 24 hours with first insights and a proposed next step.