What I build: LLM apps, RAG, multi-agent workflows
Most AI projects I’m asked about fall into one of three shapes, and many products need more than one. I design them to share one data layer, one set of logs and one evaluation suite.
Enterprise documents are where RAG usually breaks. Scanned PDFs, tables, outdated versions and files that only some users may see all need handling first. I build permission-aware retrieval, so users only get answers from documents they are allowed to open, and I treat ingestion as production software, with tests, retries and monitoring.
- LLM applications: assistants, copilots and AI features inside your SaaS product, on OpenAI, Claude, Gemini or Vertex AI
- RAG over your documents and data, such as contracts, policies, manuals, tickets and records, with answers that cite their sources
- Production RAG pipelines for enterprise documents: ingestion, OCR for scanned files, chunking, embeddings, search and scheduled re-indexing
- Multi-agent workflow automation: agents that plan, call your tools and APIs, pass work between them and stop for human approval where a mistake would be costly
- Fine-tuning, where evaluation shows that prompting and retrieval have reached their limit
From demo to production: evaluation, guardrails, cost control
A demo answers the questions it was built for. Production gets the ones nobody planned, from users who paste in odd documents, ask the same thing five ways and sometimes try to break it. Evaluation, guardrails and cost control make the difference, so I build them in from the first week rather than bolting them on before launch.
Evaluation comes first. I build a test set from real examples of your users’ tasks, with expected answers agreed with your domain experts, and run it automatically on every change to prompts, models or retrieval. It is the same discipline I apply in evaluation and red-teaming work on frontier LLMs and autonomous coding agents through Mercor, Turing, Uber AI Solutions, DataAnnotation, Mindrift, AfterQuery and Terac, where part of the job is catching tests that pass when they shouldn’t.
Cost matters as much as accuracy, because inference spend grows with every user. I route simple steps to cheaper models, cache repeated work in Redis, set token and step budgets on agents and log the cost of each request, so you can see cost per customer rather than one monthly bill.
- Input and output checks for prompt injection, personal data and off-topic requests
- Least-privilege tools, so an agent can only do what the user in front of it is allowed to do
- Human confirmation before actions that send, delete, pay for or change records
- Tracing of every prompt, retrieval result and tool call, so failures can be found and reproduced
- An AI red-team assessment before launch, testing the whole system under attack
Case studies: Kaivo AI assistant, Point.Laz industrial AI, Carbon Giant OCR pipeline
Kaivo is an AI operating system for advertising agencies, which were juggling client ad accounts across Google, Meta, TikTok, LinkedIn, Microsoft, Reddit, Spotify and Shopify with no single view of performance. I designed a multi-service platform with an in-product AI assistant, “Kai”, and a signals engine that flags ROAS decline, budget pacing and creative fatigue, built with Next.js, FastAPI, microservices and LLM agents.
At Point.Laz, as Founding Principal Engineer, I built industrial AI for underground mine safety, where millimetric ground movement has to be detected through dust, steam and low light. LiDAR-to-cloud pipelines create high-precision 3D digital twins, suppress environmental noise and run automated convergence monitoring, with point clouds streamed to the browser. Mining engineers get remote inspection and automated movement alerts they can act on from anywhere.
For Carbon Giant, I built an OCR- and API-driven ingestion pipeline, using AWS Textract and Apideck, that turns invoices and Xero, QuickBooks and Sage data into Scope 1–3 emissions calculations with versioned DESNZ and EPA factors. It is the same problem most enterprise RAG projects face: getting messy real-world documents into clean, traceable data before anything reasons over it.
Stack
I choose tools to fit your existing systems and team, not the other way round. These are the ones I use most for production AI work.
The system runs in your own AWS, Azure or GCP account, under your access controls and data residency settings.
- Models: OpenAI, Claude, Gemini and Vertex AI, behind a model layer you can switch
- Retrieval: RAG with embeddings, vector search and source citations
- Agents: multi-agent and agentic workflows with tool calling and human-in-the-loop steps
- Backend: Python (FastAPI, Django), Node.js and TypeScript
- Data: Postgres, MongoDB and Redis, with Celery for ingestion and background jobs
- Evaluation: automated evaluation pipelines and regression test sets that run in CI
What you get
- A written architecture and scope covering models, retrieval, tools and data flows
- A production LLM application, RAG pipeline or multi-agent workflow, deployed in your cloud
- An evaluation test set and automated pipeline that runs on every change
- Guardrails, tool permissions and human-approval steps matched to the risk of each action
- Tracing, cost tracking and alerts for model spend, latency and failures
- Weekly demos throughout the build
- Documentation and a handover session, so your team can run and extend the system
How it works
Step 01 · Before the build
Discovery and scope
NDA first, then a working session on the problem, your data, your users and what a wrong answer would cost. You get a written architecture, scope and a fixed or capped price before any work starts.
Step 02 · Weeks 1–2
Evaluation set and prototype
I build the evaluation set with your domain experts, then a working prototype on real data, so quality is measured from the start rather than judged by eye.
Step 03 · Agreed milestones
Build
Hands-on delivery with weekly demos: retrieval, tools, guardrails, integrations and the interface your users see. The evaluation suite runs on every change.
Step 04 · Before go-live
Harden and launch
Cost and load testing, a red-team pass for prompt injection, data leakage and unsafe tool use, then a staged rollout to real users.
Step 05 · At launch
Hand over
Documentation, runbooks and a handover session, so your team can run, change and extend the system without me.
Proof
“Steady progress, predictable delivery, and code that’s easy to review and integrate.”
Sean Metcalf
Founder, getKaivo
“Solid technical knowledge in automations using n8n and Meta Cloud API — he solved complex integrations efficiently.”
Manuel Garcia Fuentes
CEO, BuzzBlend
“He was able to deliver many parts of our system, communicate efficiently, and was fun to work with throughout the engagement.”
Chen Atlas
CTO & Founder, Optery
Price
From £12,000. For a scoped build from architecture to launch, with weekly demos and a full handover. Every engagement starts with a free 1-hour intro call, and you get a written scope with a fixed or capped price before any work starts.
FAQ
How much does it cost to build an AI agent in the UK?
My Build Projects start from £12,000 for a scoped AI feature, from architecture to launch. The price depends on the data sources, tools and integrations involved, and how much hardening the risk calls for. Model and cloud usage is billed to your own accounts, so you see those costs directly.
Can you take our AI prototype into production?
Yes, and it is a common starting point. I measure where the prototype fails using an evaluation set built from real user questions, then add the retrieval fixes, guardrails, logging and cost controls it needs. If parts are worth rebuilding, I’ll say which and why before we agree the scope.
Do we need RAG, fine-tuning or both?
Most products should start with RAG, because answers stay grounded in your current documents and the data can change without retraining. Fine-tuning helps with a fixed format, tone or a narrow, repetitive task, once evaluation shows prompting and retrieval have reached their limit. Many production systems never need it.
Should we build on OpenAI, Claude or Gemini?
It depends on the task, your data residency needs and your budget, so I test candidate models against your own evaluation set rather than public leaderboards. A model layer between your product and the provider lets you switch, or use different models for different steps, without a rewrite.
Is our data safe when it goes to an LLM provider?
It can be, with the right setup. I check the provider’s data-use and retention terms for the API you use, keep personal data out of prompts unless the task needs it, enforce document permissions in retrieval and log access. If you need SOC 2 or ISO 27001 controls around the feature, I build them in from the start.
How do you stop an AI agent taking the wrong action?
By limiting what it can do, not only what it is told. Each tool gets the narrowest permissions it needs, risky actions need human confirmation and every step is traced. Before launch, I test the agent for prompt injection and unsafe tool use with the methods from my AI red teaming service.