Seats present: claude, codex, mistral, gemini, groq, nvidia Seats missing: cerebras (benched: billing/credits needed on the provider account - fix it, then run a self-check to re-seat it); glm (no key configured); openrouter (benched: key invalid); sambanova (no key configured); cohere (no key configured); deepseek (no key configured); kimi (no key configured) Seats with failed calls: gemini (2 call(s) failed), groq (4 call(s) failed), nvidia (2 call(s) failed)
Synthesis
Board Synthesis — The Four Forks Verdict
Seats: Present: claude, codex, mistral, gemini, groq, nvidia. Missing: cerebras (benched: billing/credits needed on the provider account — fix it, then run a self-check to re-seat it); glm (its API key not set); openrouter (benched: invalid key — replace its API key locally, then self-check to re-seat); sambanova (its API key not set); cohere (its API key not set); deepseek (its API key not set); kimi (its API key not set). Failed calls: gemini (2 — absent from operator and red-team), groq (4 — capital-allocator only), nvidia (2 — absent from capital-allocator and red-team). Six of thirteen seats, three of them partial.
Consensus
- The forks mostly resolved. Canvas: ship, but as an experiment on the founder, not a moat — its completion criterion is behavioral (≥3 real Decisions in 7 days), which answers nvidia's bait charge with data. MCP: slot fourth as a thin one-tool
run_councilwrapper aimed at the Claude-locked friend, not a 48h ecosystem play. Corpus: defer the product, instrument now — passively log and dedupe every source touched (near-zero cost) and measure cross-run reuse. Citations: defer; attach raw links mechanically first. - Consensus 14-day order: Days 1–2 canvas + interview Decision capture; Day 3 cost card + token ledger (no seat knows actual run cost — "expensive" is a feeling); Days 4–5 blind head-to-head; then MCP-for-one-friend, then survivors only.
- The binding constraint is the founder's habit, not engineering. Multiple seats independently prescribed fixed Tue/Thu 2-hour blocks to engineer out motivation dependence.
- The silent risk is seat ops — seven benched seats and mid-round call failures; claude-operator says cap repair at 3 hours or ship with six seats.
Live disagreements
- MCP timing: gemini (distribution) still wants a 48h public ship as a zero-friction acquisition surface; groq (capital) says kill it outright; claude, codex, mistral, nvidia hold the middle — thin slice, fourth, gated on the blind test.
- Citations: groq alone says citations survive as the anti-copycat wedge ("the only AI council that shows its work"); claude (red-team) counters that the council opines from training data — there are no sources to cite until real web research exists.
- Test sequencing: gemini and nvidia (first-principles) want the blind test Day 1, before any building; the majority builds canvas first.
Kill criteria
- Founder writes fewer than 3 Decisions within 7 days of canvas shipping — the interface was an alibi; the canvas-as-moat claim dies.
- Council fails to win at least 2 of 3 blind judgments (or costs >$6/run) — the multi-model premise dies, and everything downstream of it.
- friend doesn't run the MCP tool twice unprompted within 7 days — MCP dies as a channel.
- Cross-run source reuse under ~20% after 10 runs — the corpus wedge dies even in theory.
Next actions
- Ship canvas + interview-style Decision capture (Days 1–2); the founder writes 3 real Decisions within 7 days.
- Add token-per-seat logging and the pre-run cost card (Day 3); estimate lands within 25% of actual across 5 runs.
- Run the $5 blind head-to-head (Days 4–5) with rubric and blinding written before running; 3 independent judges.
The way through
The biggest objection: the whole plan queues behind a founder who hasn't used his own product. The realistic route is already in the build order — the interview capture removes his stated friction, and fixed weekly blocks remove the motivation dependence; his behavior becomes the cheapest instrument in the portfolio. If the blind test fires the kill: the pivot worth testing before walking away is collapsing the council into a single-model Decision-capture tool shipped inside Claude via MCP — the founder's friend's exact belief, served instead of fought.
Dissent
By: codex
- “Consensus” is overstated. Only six of thirteen seats participated, three partially; the reported majority may be an artifact of correlated models and failed adversarial calls.
- The binding constraint is asserted, not established. Founder non-use could reflect poor output, unclear stakes, latency, or absent product value—not deficient habit. Calendar blocks do not “engineer out” motivation; they subsidize weak demand.
- Canvas-first is backwards. Building UI before testing whether the council beats a strong single model risks polishing an invalid premise and contaminating the founder’s judgment through investment bias. Run the blind test first.
- The blind test is badly specified. Three judgments and “2 of 3” provide almost no statistical confidence. There is no baseline model, task sampling method, quality rubric, latency constraint, or control for verbosity/style preference. Failure would indict the test or orchestration—not “the multi-model premise.”
- Every kill criterion is numerology. Three Decisions, $6/run, two unprompted uses, 20% reuse, and ten runs have no customer economics or observed baseline behind them. They invite threshold gaming.
- The founder and one friend are convenience samples. Founder behavior cannot validate a market; one already-Claude-committed friend cannot validate MCP as a distribution channel. “Unprompted” is especially meaningless when the tool was built specifically for him.
- Seat ops is not a three-hour side issue. Reliability is integral to the advertised product. Partial councils silently change output quality, cost, and reproducibility; graceful degradation and minimum-quorum semantics are missing.
- “Passively log every source” is not near-zero cost. It creates provenance, privacy, retention, canonicalization, and access-control obligations. Raw links also provide false credibility when claims are not mapped to evidence.
- The council skipped the commercial question: named buyer, painful recurring workflow, incumbent alternative, willingness to pay, required turnaround time, and why disagreement synthesis beats one capable model.
- The fallback compounds platform risk. A single-model tool inside Claude abandons differentiation while becoming dependent on the platform that can absorb it.
Survives: instrumenting real costs before making economic claims is sound.
Theory tests
Tier-2 sub-agents: research lenses chosen and deployed by the chairman across the strongest seats.
Theory tests
lens-behavioral — CONDITIONAL (both sub-agents; AGREED). claude and codex converge on two things: Decision capture must be a zero-friction default (claude: an un-closeable one-tap KILL/PROCEED/PIVOT; codex: interview-as-landing-state, ≤90 seconds, ≤3 taps), and visible per-run metering amplifies pain-of-paying — Buildpad went $20→$39 flat precisely because subscription anesthetizes the meter. Strongest point (claude): the blind test measures System 2 while buying is System 1 — judges can crown the council while nobody pays.
lens-distribution — CONDITIONAL (both; SPLIT on the channel). Both condemn the same hole — a 14-day plan with zero distribution days — then prescribe opposite channels. claude: MCP directories are an empty app store; a 48h run_council listing is page-one placement for $0, with the founder's friend demoted to user #2. codex: the channel is short-form video (30 "Claude vs eight rivals" clips at $39/mo self-serve); MCP is delivery, not distribution — fake-door it. Strongest point (claude): the MCP CAC-arbitrage window closes in months, and this product is already API-shaped.
lens-unit-econ — CONDITIONAL (both; AGREED). Both observe the plan contains zero revenue events and every kill criterion is behavioral, none denominated in dollars. Both prescribe a prepaid, guaranteed flat offer (claude: $200 council audit with full refund; codex: $299 Decision Sprint, gated on 2-of-10 prospects prepaying). Strongest point (claude): at $39/mo, a 15-run subscriber is gross-margin negative before infra.
Where the theories collide: Price and shape. distribution-codex wants strangers converting at $39/mo self-serve within 14 days; unit-econ-claude's arithmetic says $39/mo loses money on every active subscriber; behavioral-claude says nothing recurs, so subscription is the wrong shape entirely. Secondary collision: MCP is distribution-claude's entire acquisition strategy, while unit-econ-codex forbids writing a line of it before a $99 prepayment.
Strongest objection still standing: behavioral-claude's "every sale is a cold start." No lens or seat named what brings a paying customer back for run #2 — yet unit-econ-codex's LTV ≈ $1,250 silently assumes a 10%-churn subscription. Until something recurs, every LTV figure on this board is fiction.
Decision
(Yours to write by hand. The council never decides for you and never posts anything anywhere.)