I spent fifteen years as a therapist. Then I shipped production software. Now I build the AI that has to be trusted in front of real people.

People, and now machines, produce confident, coherent output that's sometimes completely wrong. The real work is building the systems that catch it. That's what I do.

LoanSlam. The AI proposes; the code enforces.

My own fail-closed conversation engine for a regulated lending domain, built after a 30-day Loans by MAL contract, as my vision of how it should work. An untrusted LLM proposes; deterministic code decides; every turn is audit-traced. Prototypal, over synthetic data.

live · details

The Pit. If you can't see the agent, you can't trust it.

Multi-agent AI evaluation platform. Structured contests between agent configurations with observable traces, scoring, failure tagging, and cost visibility.

live · repo · details

Sortie. Different models have different blind spots. Make them check each other.

Async adversarial multi-model code review system. Parallel LLM fan-out, debrief synthesis with convergence analysis, severity-gated merge blocking. Python.

repo · details

Becoming Diamond. Shipped, paid for, and edited by the client themselves.

Paid client build, production and customer-facing: a marketing site plus a gated member portal delivering a 30-day video course, with AI chat, Stripe membership, and a git-based CMS for non-technical editing. Next.js, React 19, TypeScript, Stripe, Decap CMS.

live · details

Halo. A human and an agent, driving the same tools the same way.

Agent/tool-layer infrastructure: CLI modules with isolated stores and NATS event sourcing, so humans and agents drive the same tools through the same surface. Python, Docker, Kubernetes.

repo · details

all work

If you have an AI idea stuck in pilot, I take one workflow to a shipped, measured system. Fixed fee. A baseline before a line is built.

kai@oceanheart.ai →