Skip to content
Eliot Leu

Eliot Leu is your partnerfor building ambitious AI systems.

刘俊 · Vakkie

Eliot Leu · Vakkie · Studio
mission-control.md·app-manager.ts
Reading scopeGrepping risksDrafting planEval gateReady

Build an Exposé-style window manager, then wire parallel agents to a PMP acceptance list.

## Trigger
Menu, hotkey F3, or double-click a program on the board.
## Tasks
Create MissionControlView.tsx
Map agent states onto the risk register
Accept: agents cannot rewrite the baseline

Software creation is changing.

I work at the edge of research, engineering, product, and finance: turning models into shippable systems, systems into manageable spend, and spend into decisions you can explain.

There is much to learn, try, and build.

Public track
  • 2026Multi-agent delivery OSActive
  • 2026Prompt eval and regression harnessPublished
  • 2025Cost-object model for AI programsPublished
  • 2025Product-grade autonomy sliderPublished
  • 2024Full-stack agent workbenchPublished
  • 2024Research reproduction pipelinePublished
  • 2023PMP methods for AI deliveryCertified
  • 2023CMA management accountingCertified
  • 2022Prompts as product interfacesPublished

Agents turn ideas into running systems

Research prototypes, product specs, and cost constraints share one delivery chain. Agents write the code; you keep the decisions: interfaces, data, evals, and release.

Full-stack studio
mission-control.md·app-manager.ts
Reading scopeGrepping risksDrafting planEval gateReady

Build an Exposé-style window manager, then wire parallel agents to a PMP acceptance list.

## Trigger
Menu, hotkey F3, or double-click a program on the board.
## Tasks
Create MissionControlView.tsx
Map agent states onto the risk register
Accept: agents cannot rewrite the baseline

Works in parallel, lands in scope

PMP discipline on AI programs: scope, risk, stakeholders, and milestones become executable plans. Agents explore in parallel. The project manager decides, accepts, and reviews.

PMP board
Agents do not rewrite the baseline
Scope
Board v1: grid, states, acceptance
Risk
Model writes exploration as requirements
Accept
Baseline changes need a human signature

In every tool, at every step

A prompt is not a spell. It is a versioned, evaluated, reversible interface. Prompt engineering belongs in the CLI, the review, the doc, and the product surface — so model behavior stays predictable.

Prompt interface
system.md · v3.4
You are an executor in the program, not the author of scope.
- Read the acceptance list before touching code
- Scope changes return to a human
- Failures enter the eval set; they are not retried in chat
128 goldensRegressionReversible

Let research constrain imagination

Measure first, then scale. Eval sets, failure cases, and reproductions come before the demo. The question is not whether the model can talk. It is whether it holds on your task distribution.

Research lab
Lab · reproducible
eval.yaml
iddistributionscoregate
MNIST-12Slice drift0.81Fail
Prompt-88Tool consistency0.94Pass
Agent-04Scope adherence0.89Watch

Give autonomous agents a product boundary

Autonomy is not abandonment. Product work is permissions, feedback, and interrupts: the user holds the autonomy slider, from inline complete to full agent, with a clear consequence at every notch.

Product boundary
Autonomy slider

The user always knows what this notch changes, who is accountable, and how to interrupt.

CompleteEditPlanAgent
Current: Complete. Cursor-local only. You still commit.

Manage AI spend in the language of finance

Tokens, labor, rework, and risk can all be booked. CMA methods give AI programs cost objects, contribution margins, and rolling forecasts — not an unexplained cloud bill.

Management accounting
AI spend ledger
FY26 · W37
41%
contribution
Board agents$12.4k38%Margin up
Human review$6.1k22%On plan
Rework reserve$2.8k9%Watch

The new way to build with AI.

He turns a vague model demo into something you can schedule, accept, and review. Researchers follow it. Finance follows it too.
Ke Lin·Head of Product, consumer tech
Prompts in his hands have tests. A change gets a regression. Failures enter the set — they do not die as another chat retry.
Heng Zhao·Head of quantitative research
We booked agent tokens, rework, and human review on one contribution-margin sheet. The conversation moved from vibe to tradeoff.
Wan Su·Finance business partner
He does not write feature lists. He writes the consequence of every notch on the autonomy slider. That is product work. Few people actually do it.
Cheng Jiang·B2B product advisor
Full-stack, for him, is not one person who knows every framework. It is research, code, release, and cost staying on one chain.
Mu Han·Director of engineering
PMP is often decoration on AI programs. He uses it as a real language of scope and risk, so a crowd of agents cannot walk the goal off the map.
Lan Pei·Head of PMO

Start building together.