MCP Hub
Back to servers

@velvetmonkey/flywheel-ideas

The local-first falsifiable decision ledger — MCP server that turns an Obsidian vault into a compounding decision system with multi-model AI council dissent and outcome-driven refutation propagation.

npm251/wk
Updated
Apr 23, 2026

Quick Install

npx -y @velvetmonkey/flywheel-ideas

flywheel-ideas

Track why you believed a bet, challenge it, and learn if it was right.

A local-first falsifiable decision ledger for your Obsidian vault. Every idea becomes a tracked object with a lifecycle state machine, a declared-assumption ledger, a multi-model AI council that stress-tests the assumptions, and an outcome log that refutes broken assumptions and automatically flags every dependent idea for re-review. The vault compounds. Six months from now, it catches the assumption that failed.

npm version license: Apache 2.0 status: stable

Status: 0.1.0 GA. v0.1 closed loop shipped across M1–M14 + alpha.4 hardening + alpha.5 consolidation. M13 (real claude -p e2e in CI), claude-auth error classification, and the v0.2 product surface carry to v0.1.1 / v0.2.

Who this is for

  • Solo founders making product, pricing, GTM, and hiring bets
  • Staff+ / principal engineers making architecture and platform bets
  • Consultants, investors, and researchers who revisit why they believed something

If you've ever made a consequential decision and six months later couldn't articulate the assumption that turned out wrong — this is for you.

The loop

idea.create                          # "Should we switch to event-driven architecture?"
   ↓
assumption.declare                   # "In the context of X, facing Y, we assume Z, accepting W"
   ↓
council.run(pre_mortem)              # Real multi-model dissent attacks the assumptions
   ↓
idea.transition → committed          # You made the call
   ↓
outcome.log(refutes: [A1, A3])       # Later: reality arrived
   ↓
Every idea in your vault that relied on A1 or A3 is flagged for re-review.

That's the product. Everything else supports this loop.

What makes it different

What it really doesThe actual gap we fill
Notion / Jira Product DiscoveryManual idea capture + AI featuresCloud-only; no dissent; no outcome loop
Log4brains / adr-toolsADRs in gitPost-hoc records; no live dissent
LangGraph / CrewAIStateful agent orchestration frameworks (with memory)Developer libraries, not products; no vault-native decision ontology
autodecisionPersona deliberation + persistent runs in ~/.autodecision/Siloed from your knowledge graph; single-model; no outcome loop into the runs
obsidian-agent-client / EchoVault / PKM AssistantAI agents in ObsidianPlumbing; no decision ledger

The gap nobody fills: a local, vault-native, human-readable decision ledger with falsifiable assumptions and outcome-driven refutation propagation.

How it works

Ideas are markdown notes. Frontmatter tracks id, state, declared assumptions, lineage. The vault is source of truth; ideas.db is an index.

Writes go through flywheel's shared orchestration. Artifacts are indexed immediately — visible in flywheel-memory's search with no watcher delay. If the preferred write path isn't available, we fall back cleanly (MCP-subprocess to flywheel-memory; direct fs with a warning).

The council is a subprocess dispatcher. council.run(idea_id, depth, mode) spawns parallel child processes — one per (model, persona) cell — shelling out to your installed CLIs with explicit consent. Concurrency capped. Failures classified (timeout / auth / rate-limit / parse / exit-nonzero); one cell failing never aborts the session. Each cell runs a mandatory two-pass self-critique. Views persist as markdown; template synthesis distills them with evidence citations.

Outcomes refute assumptions. outcome.log(refutes: [...]) marks assumptions refuted and flags every dependent idea. outcome.log.undo reverses if you logged by mistake. This is the compounding mechanism.

Every response teaches you the next move. {result, next_steps} — the tool surface is the onboarding flow.

Quickstart

Prerequisites

  • Node.js 22+
  • Obsidian vault with flywheel-memory initialized (recommended — degrades gracefully without it, but the compounding mechanism works best when writes are fully indexed)
  • One or more of claude, codex, gemini CLIs on $PATH
  • VAULT_PATH env var pointing at your vault (no hardcoded paths)

Install

The flywheel-ideas MCP server runs as a subprocess launched by your MCP client (Claude Desktop, Claude Code, Cursor, etc.). You don't need to npm install it globally — npx will fetch + cache it on first use.

Grant approval for council dispatch out-of-band (the LLM cannot grant it itself):

# Session-scoped (recommended while getting started)
export FLYWHEEL_IDEAS_APPROVE=session
# Persistent across restarts — edit <vault>/.flywheel/ideas-approvals.json manually

Wire into your MCP client (example mcp.json entry):

{
  "mcpServers": {
    "flywheel-ideas": {
      "command": "npx",
      "args": ["-y", "@velvetmonkey/flywheel-ideas"],
      "env": {
        "VAULT_PATH": "/path/to/your/vault",
        "FLYWHEEL_IDEAS_APPROVE": "session"
      }
    }
  }
}

The -y flag auto-accepts the initial npx download prompt. After the first run the server is cached locally; subsequent launches are instant.

Prefer a global install? npm install -g @velvetmonkey/flywheel-ideas, then point the MCP client at flywheel-ideas-mcp directly.

When a v0.2 pre-release train starts, opt in via the @alpha dist-tag (e.g. npx -y @velvetmonkey/flywheel-ideas@alpha). The latest tag always tracks the most recent stable.

See CHANGELOG.md for what's new in each release.

First flow

> idea.create({title: "Move subsystem X to event-driven architecture"})
> assumption.declare({
    idea_id: "idea-abc",
    context: "current synchronous design",
    challenge: "ops burden of cross-service debugging",
    decision: "events worth the tradeoff",
    tradeoff: "observability cost",
    load_bearing: true
  })
> council.run({id: "idea-abc", mode: "pre_mortem", confirm: true})
> idea.transition({id: "idea-abc", to: "committed", reason: "council confirmed; A1 still load-bearing"})

# ... 9 months later ...
> outcome.log({
    idea_id: "idea-abc",
    text: "Migration completed. Observability cost was 3x forecast; debugging time dropped 40%. A2 refuted — cost model was wrong. A1 validated.",
    refutes: ["asm-2"], validates: ["asm-1"]
  })
# → automatically flags 3 other ideas in the vault that cited asm-2.

Tool surface (v0.1)

idea

  • create({title, body?}) / read(id) / list({state?, limit?}) / transition({id, to, reason?})

assumption

  • declare({idea_id, text? | structured, signpost_at?, signpost_reason?, load_bearing?}) / list({idea_id})
  • lock({idea_id}) / unlock({idea_id}) — OSF-style pre-registration
  • signposts_due({window?}) — sweep for elapsed signposts (fed into next_steps)

council

  • run({id, depth: "light"|"full", mode: "standard"|"pre_mortem", confirm}) / view(view_id) / list({idea_id})

outcome

  • log({idea_id, text, refutes?: [asm_ids], validates?: [asm_ids]}) / list({idea_id?}) / undo(outcome_id)

Roadmap

v0.1 — the closed loop (in progress)

The core product: four MCP tools forming idea → assumption → council → outcome → propagation.

  • ideacreate · read · list · transition · forget. Lifecycle state machine (8 states), atomic DB + frontmatter transitions with rollback-on-failure, stale-row filtering, next_steps guidance for every state.
  • assumptiondeclare · list · lock · unlock · signposts_due · forget. Y-statement structured input OR free text, OSF-style pre-registration lock, load-bearing tagging, signpost-based re-evaluation surfacing.
  • council — multi-model subprocess dispatcher + deterministic synthesis.
    • ✅ M6: approval / dispatch-log plane (out-of-band consent; LLM cannot self-grant)
    • ✅ M7: CLI characterization + error classifier (claude/codex/gemini quirks catalogued)
    • ✅ M8: real claude dispatcher, 2 personas, single-pass, deterministic SYNTHESIS.md, pre_mortem mode
    • ✅ M9: two-pass metacognitive + codex dispatch + concurrency limiter
    • ✅ M10: gemini dispatch + full matrix (3 × 5 = 15 cells) + CLI-interleaved concurrency + strict benign-stderr filter
    • ✅ M11: evidence-aware synthesis with sentence-level Jaccard agreement/disagreement sections
  • outcome — refutation propagation (the compounding mechanism), reversible via undo. Shipped at M12.
  • memory-bridge — flywheel-memory custom-category registration on startup (M14, alpha.3). Path-security + maxBuffer + frontmatter-sync hardening shipped in alpha.4. Consolidation (needs_review filter, model_version capture, comment sweep) shipped in alpha.5.

v0.1.0 GA shipped 2026-04-23. What's still upcoming, carrying to v0.1.1 / v0.2:

  • ⏳ M13 — real claude -p end-to-end test in CI with flake-aware demotion.
  • ⏳ CLI-error classifier auth + rate_limit patterns (real failure samples now captured during the GA dogfood).
  • clis arg passthrough on council.run (gap surfaced in dogfood — orchestrator dispatched all 3 CLIs even when subset was requested).

v0.2 — depth on the loop + closing the feedback loop

idea.freeze / council.freeze (OSF snapshot) · metric & guardrail fields on assumptions · RAND ABP shaping + hedging actions · Assumption Radar (semantic vault-wide signal detection) · daily-note outcome capture · automated signpost surfacing · agent-driven outcome detection · lifecycle enforcement · decision_delta view · digest · lineage queries · steelman council mode · Anti-Portfolio pass memos · Ollama / LM Studio local models

v0.3+ — operator calibration

Obsidian plugin · personal calibration dashboard (Brier trend) · persona effectiveness A/B · exportable decision portfolios · state-of-mind context capture (Farnam Street) · Brier-scored assumption updating

Long-term

Your vault becomes an empirical record of your predictions. Personal calibration emerges over years. Cross-idea reasoning flags assumption conflicts across your roadmap. A community benchmark — seeded with the author's own historical decisions — lets anyone compare council quality.

Design principles

  • Vault is source of truth. ideas.db is an index.
  • Every response carries next_steps. The tool surface teaches itself.
  • No auto-transitions, no auto-decisions. User writes the final rationale; council provides dissent, never verdict.
  • Explicit consent per MCP spec. Subprocess spawns + vault writes require approval; dispatches are audited.
  • Reversibility. Outcomes can be undone; refutation propagation unwinds.
  • Single write path across the flywheel family — writes go through vault-core orchestration when available.
  • Graceful degradation. If the preferred write path isn't available, we fall back cleanly.
  • No LLM SDK lock-in. Uses whatever CLIs you have on $PATH.
  • Apache 2.0. No viral reach.
  • Tested hard. Unit · property-based · integration · real claude -p e2e in CI with flake-aware demotion.

Ecosystem

  • vault-core — shared core library (authoritative vault-write orchestration)
  • flywheel-memory — vault indexing MCP
  • flywheel-ideas (this repo) — decision ledger

Packages

This repo publishes two npm packages. Most users only interact with the first:

PackagePurposeUsers
@velvetmonkey/flywheel-ideasThe MCP server users add to their client configEnd users
@velvetmonkey/flywheel-ideas-coreInternal domain library the server depends onTransitive

Release + publish workflow: see RELEASE.md.

License

Apache 2.0. See LICENSE.

Reviews

No reviews yet

Sign in to write a review