Advisory

A second set of eyes on the serving stack.

For teams with scaled inference workloads where every millisecond and dollar per million tokens counts, I audit what is actually happening across your cluster, identify the recoverable spend hiding behind gross utilization, and advise the technical decisions that change your unit economics. Your team owns the execution.

I take only 2 clients a month. Not a hire. Not a Joule service.

Start here

Architecture & SLO Audit

$5,000 USD

A focused live review and written findings memo, scoped to 5 specific questions or traces you name.

One week. 100% credited toward Month 1 if a retainer starts within 30 days.

Base scope up to 5 traces. Complex multi-cluster setups scoped and billed separately.

  • Single-session live architecture review
  • Async memo on up to 5 agreed questions or traces
  • 60-minute debrief and remediation findings
  • Written roadmap delivered within 7 calendar days

Then the seat

Strategic Retainer

$10,000 / month

A standing seat on the serving stack: compute economics, agentic systems, or both.

Three-month minimum. Scoped to 2 concurrent engineering workstreams.

Secures your seat. We kick off within 5 business days. If intake shows an irreconcilable scope mismatch, the payment is refunded immediately.

  • 2 working sessions a month with engineering leadership
  • Async RFC and architecture review (24 to 48 hour SLA)
  • Monthly compute economics and SLO governance memo
  • Formal TAB seat for investor and technical diligence

A decade building production ML. Author of LLMOps (O'Reilly). GPU Engineering: AI Inference and System Design (Packt) is almost finished, launch is end of 2026 or early 2027. Invited Expert and Faculty Member on AI Inferencing for Andreessen Horowitz Academy (The Academy SF).

Checkout

Pay. Then we schedule.

Payment unlocks the calendar and the intake. I take only 2 clients a month.