Eliot Leu is your partnerfor building ambitious AI systems.
刘俊 · Vakkie
Build an Exposé-style window manager, then wire parallel agents to a PMP acceptance list.
Software creation is changing.
I work at the edge of research, engineering, product, and finance: turning models into shippable systems, systems into manageable spend, and spend into decisions you can explain.
There is much to learn, try, and build.
- 2026Multi-agent delivery OSActive
- 2026Prompt eval and regression harnessPublished
- 2025Cost-object model for AI programsPublished
- 2025Product-grade autonomy sliderPublished
- 2024Full-stack agent workbenchPublished
- 2024Research reproduction pipelinePublished
- 2023PMP methods for AI deliveryCertified
- 2023CMA management accountingCertified
- 2022Prompts as product interfacesPublished
Agents turn ideas into running systems
Research prototypes, product specs, and cost constraints share one delivery chain. Agents write the code; you keep the decisions: interfaces, data, evals, and release.
Build an Exposé-style window manager, then wire parallel agents to a PMP acceptance list.
Works in parallel, lands in scope
PMP discipline on AI programs: scope, risk, stakeholders, and milestones become executable plans. Agents explore in parallel. The project manager decides, accepts, and reviews.
In every tool, at every step
A prompt is not a spell. It is a versioned, evaluated, reversible interface. Prompt engineering belongs in the CLI, the review, the doc, and the product surface — so model behavior stays predictable.
You are an executor in the program, not the author of scope. - Read the acceptance list before touching code - Scope changes return to a human - Failures enter the eval set; they are not retried in chat
Let research constrain imagination
Measure first, then scale. Eval sets, failure cases, and reproductions come before the demo. The question is not whether the model can talk. It is whether it holds on your task distribution.
Give autonomous agents a product boundary
Autonomy is not abandonment. Product work is permissions, feedback, and interrupts: the user holds the autonomy slider, from inline complete to full agent, with a clear consequence at every notch.
The user always knows what this notch changes, who is accountable, and how to interrupt.
Manage AI spend in the language of finance
Tokens, labor, rework, and risk can all be booked. CMA methods give AI programs cost objects, contribution margins, and rolling forecasts — not an unexplained cloud bill.
The new way to build with AI.
“He turns a vague model demo into something you can schedule, accept, and review. Researchers follow it. Finance follows it too.”
“Prompts in his hands have tests. A change gets a regression. Failures enter the set — they do not die as another chat retry.”
“We booked agent tokens, rework, and human review on one contribution-margin sheet. The conversation moved from vibe to tradeoff.”
“He does not write feature lists. He writes the consequence of every notch on the autonomy slider. That is product work. Few people actually do it.”
“Full-stack, for him, is not one person who knows every framework. It is research, code, release, and cost staying on one chain.”
“PMP is often decoration on AI programs. He uses it as a real language of scope and risk, so a crowd of agents cannot walk the goal off the map.”
Recent log
- Multi-agent research boardSep 10, 2026
- Prompt regression suite v3Sep 2, 2026
- Greenfield product experiments, no legacy repoAug 27, 2026
- Rolling forecasts and contribution margin for AI spendAug 19, 2026
Research & notes
- Mar 27, 2026·ResearchAn evaluation report on Composer-style workflowsEliot Leu·8 min
- Aug 14, 2026·ProductThe autonomy slider: giving control backEliot Leu·6 min
- Aug 12, 2026·EngineeringPrompts are interfaces, not spellsEliot Leu·5 min
- Aug 18, 2026·FinanceBooking tokens to cost objectsEliot Leu·9 min
Selected slices
Parallel agents, scope baseline, risk register, and acceptance list share one view. The PM decides; agents explore.
Prompts as interfaces: versions, goldens, failure clusters, regression gates. Built for prompts that have to live.
Experiment cards, data slices, metric boards, and paper-grade reproduction scripts. Evals before demos.
A delivery chain from PRD to preview: code, data, permissions, release. Humans appear at decision points.
Complete, targeted edit, plan, and full agent. Each notch has permissions, audit, and interrupts.
Tokens, labor, rework, and risk reserves booked to cost objects. Contribution margin and rolling forecasts, weekly.
