Advisory

A second set of eyes on the serving stack.

For teams already running production inference and paying a real GPU bill. I take the system apart, break down what is actually happening, and sit on the calls that move cost and latency. You keep the implementation.

I take only 2 clients a month. Not a hire. Not a Joule service.

Start here

Architecture & SLO Audit

$5,000 USD

A focused live review and written findings memo, scoped to 5 specific questions or traces you name.

One week. 100% credited toward Month 1 if a retainer starts within 30 days.

Base scope up to 5 traces. Complex multi-cluster setups scoped and billed separately.

  • Single-session live architecture review
  • Async memo on up to 5 agreed questions or traces
  • 60-minute debrief and remediation findings
  • Written roadmap delivered within 7 calendar days

Then the seat

Strategic Retainer

$10,000 / month

A standing seat on the serving stack: compute economics, agentic systems, or both.

Three-month minimum. Scoped to 2 concurrent engineering workstreams.

Secures your seat. We kick off within 5 business days. If intake shows an irreconcilable scope mismatch, the payment is refunded immediately.

  • 2 working sessions a month with engineering leadership
  • Async RFC and architecture review (24 to 48 hour SLA)
  • Monthly compute economics and SLO governance memo
  • Formal TAB seat for investor and technical diligence

A decade building production ML. Author of LLMOps (O'Reilly). GPU Engineering: AI Inference and System Design (Packt) is almost finished, launch is end of 2026 or early 2027. Invited Faculty for AI Inference Engineering at Andreessen Horowitz Academy (The Academy SF).

Checkout

Pay. Then we schedule.

Payment unlocks the calendar and the intake. I take only 2 clients a month.