I'm an applied Automations engineer working out of Nairobi (UTC+3, overlapping European and US-East mornings). Most of my work sits in one narrow place: taking an LLM system from "it seems to work" to "here is the number."
That means retrieval pipelines that are measured instead of vibed, agent workflows with a tool layer that can be traced, and the unglamorous glue around both — auth, payments, rate limits, multi-tenancy — that decides whether any of it survives contact with a real business.
Retrieval & RAG · hybrid search (BM25 + vector), reranking, chunking strategy, Pinecone / Supabase
Agents & tooling · MCP tool layers, Claude Agent SDK, LangGraph, spec-driven build systems
Evaluation · fixed question sets, faithfulness, tool-call correctness, latency & cost per request
Automation · n8n / Make orchestration, M-Pesa Daraja + Stripe rails, event-driven integrations
Product surface · React 19 / TypeScript / Next, Flutter, FastAPI / Laravel, Postgres, Docker
| What I built | The measurable outcome |
|---|---|
| Production RAG assistant (Pinecone + Supabase) | +30% retrieval accuracy, measured against a fixed set of real user questions — not eyeballed |
| Document-mapping automation (n8n / Make) | 60% faster end-to-end mapping cycle |
| Internal workflow automation | 20% efficiency gain on the manual process it replaced |
| Forge Lite — Claude Code build system | 31 capability playbooks + 25 system docs; spec → running React app with a human in the loop |
The middle column is the point. Most LLM portfolios say "it works." I hand over the measurement that shows by how much.
llm-eval-harness — a small, honest evaluation suite for RAG and agent systems: retrieval accuracy, answer faithfulness, tool-call correctness, latency, and cost per request. Public eval report included, methodology and all. Shipping in slices — follow along.
Write-up on the eval design: Medium →
- forge-lite — spec → React app via Claude Code. No queues, no deploy infra, no custom UI. A library of context, not a platform.
- team-manager — project and task management for small teams. React 19 + TypeScript + Vite + Tailwind v4.
- patient-tracker — Afya Yangu, a Flutter patient-records app built for a Kenyan clinic workflow.
- Plant-Leaf-Disease-Detector — CNN classifier for crop leaf disease from images.
Remote roles — applied AI / AI engineering / full-stack / Backend / Automations Engineering. I work UTC+3 and overlap comfortably with European hours and US-East mornings. I contract through an Employer of Record (Deel, Remote.com, Oyster, Multiplier all cover Kenya) so hiring me is a one-form problem, not a compliance project.
Fixed-scope automation builds — three packages, priced up front, no hourly guessing:
| Package | What you get | Turnaround |
|---|---|---|
| Automation Audit | 90-minute session + written report naming your 3 highest-value automatable workflows, with hours saved per month | 3 days |
| Single Workflow Build | One n8n/Make workflow, live, documented, Loom walkthrough, 14 days support | 7–10 days |
| Document Assistant | RAG assistant over your internal docs — with an eval set of 20 real questions and a measured accuracy number at handover | 3–4 weeks |
📩 Start here (90-second intro) · or message me on LinkedIn.



