Unwind design Rewind · Spec · Play draft for review · built 2026-10-08 · ca711ce

Unwind design: Rewind · Spec · Play

In short: This folder describes where Unwind is heading and how we get there. Teams work on a large codebase through a shared, self-hosted Unwind Server, splitting it into slices. Unwind splits into Rewind, which understands a legacy system and compiles it into a stack-neutral Spec, and Play, which rebuilds the Spec deterministically from a client's Target Kit of recipes, with an LLM filling the explicit holes. The CLI does the work next to the code and pushes artifacts (never code) to the server, which stores them, converges slice Specs into one project Spec and shows progress. Both halves are verified by computation, not by assertion. The ideas borrow heavily from OpenRewrite/Moderne, but Unwind takes no dependency on their stack.

Status: design, for review. Nothing here is built yet. Live HTML version: https://unwind-design.cliftonc.nl

Unwind destination: Server · Rewind → Spec → PlayUnwind destination: Server · Rewind → Spec → Playdeterministic (engine / CLI)LLM agentartifacthuman / externalnew in destinationREWIND · understand the sourceper slice · runs locally via the CLISource repoany language · gitSemantic Modeltyped facts · ids · edgestree-sitter + compilersLayer docs + grilltagged [MUST] / [SHOULD] / [DON'T]coverage = manifest − docsrw-spec (compile)model + docs → Spec fragmentpush artifacts per sliceUNWIND SERVER · shared backbone from day 0slices + owners · artifacts in git · state + index in SQLite · basic UI · token authconverges slice fragments into one project Spec · tracks progress per sliceartifacts only: source code never leaves the developer's machineconvergedSPEC · stack-neutral typed IRentities · endpoints · events · operationsscenarios · priorities · provenanceContext gapsagents find what code can't say→ interview briefs → stakeholders→ answers enrich the SpecBehaviour parityscenarios run vs legacy → goldensreplay vs target → parity %observations enrich the SpecPLAY · rebuild in the target stackper slice · ordered by seamsTarget Kitclient recipe book(own git repo)Planchoose / tailor Kitphasing · riskGeneraterecipes → code + holescorrect by constructionFill holesbusiness logic onlyVerifyre-scan target − SpecMUST completeness %gaps → regenerate / refill until verified (loop mode)push verify resultsTarget repo · new stackgenerated structure + filled holes + native parity testsSurfaces: unwind CLI (does the work, next to the code) · Unwind Server UI · MCP adapter later
The big picture: Unwind Server as the shared backbone · Rewind → Spec → Play → Target, verified per slice

One-page summary

Where Unwind is today. Unwind is a Claude Code plugin with this pipeline:

  1. deterministic scripts (@unwind/core, tree-sitter) inventory a codebase;
  2. LLM specialists write [MUST]/[SHOULD]/[DON'T]-tagged layer docs;
  3. coverage is proven by set arithmetic (manifest − docs);
  4. uw-grill attacks the business logic;
  5. uw-plan interviews the user;
  6. uw-build has LLM subagents write the target code, which verify-rebuild re-scans and diffs.

Every line of target code is still written by an LLM. The extracted facts are names only: no types, no call graph.

What we learned from OpenRewrite/Moderne (01):

  • Their Lossless Semantic Tree (type-attributed, built by the real compiler) and their recipe model (scan → generate → edit, declarative composition, before/after golden tests, data tables) are excellent ideas.
  • But OpenRewrite performs no cross-language translation. Its non-JVM languages and most recipe packs are under a source-available licence (MSAL) or are proprietary, and they run through a JVM host and the Moderne CLI.
  • Decision: we copy the concepts and depend on none of the code.

The flow, end to end (02 §2.1):

  1. A team stands up the Unwind Server (03), creates a project, and developers unwind login with a simple token.
  2. The codebase is scanned once locally; the server proposes slices (units of work) and people or agents claim them.
  3. Each slice is analysed locally with Rewind (scan → tagged docs → grill → context gaps → parity observations); the CLI pushes artifacts per slice. Source code never leaves the developer's machine.
  4. The server converges slice Spec fragments into one project Spec, reporting uncovered items, overlaps, conflicts and seams.
  5. Play rebuilds slice by slice from the Spec plus the client's Target Kit; verification results are pushed back, so the server shows completeness and parity per slice.

The destination (02):

  • Unwind Server (from day 0): the team's shared system of record and UI. Hono + React, git + SQLite, one Docker image; token login; artifacts only; slices are first-class and converge into one project Spec (03).
  • @unwind/model is the shared contract: ids, Semantic Model, Spec, Kit and verification schemas.
  • Rewind (understand) covers scan, typed Semantic Model, tagged docs, grill and context gaps, and produces the Spec: a typed, stack-neutral IR of entities, endpoints, events, config and operations, each carrying priority and provenance.
  • Play (rebuild) takes a Spec plus a Target Kit and runs recipes and blueprints to generate code. The LLM fills the @unwind-holes, then verification runs: structural, typed and behavioural.
  • Target Kits are versioned per-client git repos encoding their golden path. They are mined primarily from the client's own reference app (04).
  • Behaviour parity (06): scenarios generated from the Spec run against the legacy app (to observe it and enrich the Spec) and against the rebuild (to give a parity verdict).
  • Context gaps (07): an agent round finds what the code can't tell us and produces interview briefs for stakeholders, end users, developers and ops. The answers flow back into the Spec with provenance.
  • Surfaces: one engine with a CLI first (the skills shell out to it; it works offline and pushes to the server when logged in), the server for shared state and UI, and an MCP adapter later.

How we get there (08): phases 0–8 (with 5b–5d in parallel), each shippable on its own, each keeping the graceful fallback to today's pure-LLM flow.

Reading order

#DocRead it if you want…
01Review: OpenRewrite / Moderne vs Unwindwhat we looked at, what we borrow, and why there is no dependency
02Destination architecturethe whole picture: server, Rewind / Spec / Play / Kits / surfaces
03Unwind Server and slicesthe shared day-0 server, auth, push/pull, slices and convergence
04Target Kits, recipes, blueprints, holeshow the target becomes deterministic
05Semantic Model and storethe LST-inspired typed model behind Rewind
06Behaviour paritytests that run against both the legacy app and the rebuild
07Context gaps and interviewsfinding and closing gaps the code can't answer
08Roadmapthe phased plan with exit criteria
09Open questionsdecisions still to make, each with a recommendation

How to review

  • Sections are numbered (§4.4, §7.2). Cite them in offline comments.

  • The markdown here is the source of truth. The HTML site is generated from it.

  • Conventions every proposal must respect come from the repo's CLAUDE.md:

    • the manifest schema is additive-only;
    • AST/real parsers over regex;
    • graceful fallback when Node/core is unavailable;
    • candidate ids are the join key.

    A proposal that bends one of these says so explicitly.

01 · Review: OpenRewrite / Moderne compared with Unwind

In short: OpenRewrite's Lossless Semantic Tree and its recipe model are the best existing answer to "deterministic, verifiable code change at scale", and almost every concept is worth adopting. But OpenRewrite only transforms code within one language, and its non-JVM languages and most recipe packs are source-available or proprietary. Unwind therefore copies the concepts and takes no dependency on the code.

Research was done on 2026-10-07 from primary sources: the openrewrite/rewrite monorepo, docs.openrewrite.org, docs.moderne.io and moderne.ai. Items we could not confirm are marked [unverified].

1.1 What the LST is

A Lossless Semantic Tree (LST) is "a tree representation of code" that adds two things a classic AST lacks: type attribution and format preservation (docs: LSTs, moderne.ai/platform/lst).

  • Type attribution. Every typed node carries a JavaType (docs: type attribution). The variants are Class, Parameterized, GenericTypeVariable, Method, Variable, Array, Primitive and Unknown. Between them they hold the fully-qualified names, supertypes, method signatures, generics and annotations. The same type model is reused for Kotlin, TypeScript, Python, C# and Go.
  • Built by the real compiler: javac for Java, the TypeScript compiler API (ts.createProgram + getTypeChecker()) for JS/TS, Roslyn for C#. That is why OpenRewrite needs the classpath or dependencies. When it can't resolve them, types fall back to JavaType.Unknown and type-gated recipes silently match nothing (FAQ).
  • Format preserving. Whitespace and comments are stored on the nodes, so printing the tree reproduces the source byte-for-byte. Edits pick up the local style.
  • Markers are immutable metadata attached to any node: search results, provenance, build-tool facts (docs: markers).
  • Storage.
    • The OSS Maven/Gradle plugins hold the whole LST in memory for one run.
    • The Moderne CLI (mod build) uses the build tool only for discovery, parses in its own JVM, and serialises LSTs into JAR artifacts published to Artifactory/Nexus (how LST artifacts are produced). The on-disk format is proprietary [unverified detail].
  • Languages and bridging. Non-JVM languages run as peers over JSON-RPC (RewriteRpc). They are implemented in their own language: TypeScript for JS/TS, Python, C# and Go peers.

1.2 What a recipe is

ConceptWhat it doesSource
Visitor recipeRecipe.getVisitor() walks the LST and edits nodesrecipes
ScanningRecipe<Acc>Three phases: scan every file into an accumulator, then generate new files, then edit. Cross-file knowledge only flows through the accumulatorscanning recipes
Declarative recipeYAML recipeList that composes other recipes with optionsdocs: types of recipes
Refaster / JavaTemplateBefore→after templates, compiled into type-attributed snippetsrefaster recipes
PreconditionsPreconditions.check(...) gates a recipe on facts (e.g. "repo has dependency X")recipes docs
Search recipes + data tablesRecipes that report instead of edit. They emit SearchResult markers and typed rows (CSV)data tables
TestingRewriteTest: rewriteRun(java(before, after)), with type validation after the run and repeated cycles to catch non-idempotent recipesrecipe testing
GenerationOnly ScanningRecipe.generate() creates files (e.g. CreateTextFile). All the documented examples are small artifacts, not applicationsscanning recipes docs

Recipes can be authored in TypeScript for JS/TS targets: @openrewrite/rewrite provides Recipe, JavaScriptVisitor, and pattern/template rewrite rules. They still run through the Moderne CLI (writing a JS recipe).

1.3 The Moderne platform around it

  • Moderne CLI (mod). Multi-repo builds and runs. Free for public repositories; private repositories need a licence (CLI licence).
  • Platform / DX. Multi-repo LST store, Trigrep search (trigram plus structural), data-table analytics, and coordinated pull requests.
  • AI.
    • An MCP server inside the CLI, with tools like find_types, find_methods, run_recipe, query_datatable and learn_recipe (MCP overview).
    • The Moddy agent, which uses recipes as deterministic tools [current status unverified].
  • Prethink (docs). About 140 recipes that write agent context into .moderne/context/: service endpoints, DB connections, DTO schemas, test gaps, code smells and a FINOS CALM architecture model. This is the nearest analogue to Rewind.

1.4 The two facts that decide our strategy

1. No cross-language translation. We found no recipe that translates code between languages: not Java→Kotlin, not JS→TS, not COBOL→Java. Every visitor edits one language's tree and prints it back in that same language. Cross-language work stops at three things:

  • multi-file-type recipes (a POM condition driving a YAML edit);
  • reading one language as evidence for another;
  • framework migrations within one language (Spring→Quarkus).

Unwind's core job is to rebuild in a different stack. OpenRewrite does not do that.

2. Licensing and runtime.

The licensing below is our reading of OpenRewrite licensing. It is not legal advice.

LayerLicence
rewrite-core, Java, Kotlin, Groovy, XML/YAML/JSON/HCL/TOML/Protobuf, Maven/Gradle, rewrite-templating, rewrite-analysis, RewriteRpcApache-2.0
JS/TS (@openrewrite/rewrite), Python, C#, Go, Ruby, Scala; rewrite-static-analysis, rewrite-spring, rewrite-migrate-java, rewrite-prethink …MSAL: may not be commercialised or provided as a managed service
io.moderne.* recipes, Prethink at scale, data-flow / impact analysisProprietary

In practice, non-JVM languages only run through the JVM host plus the Moderne CLI. That conflicts with Unwind's lightweight Node runtime, its "degrade gracefully" rule, and its OSS (MIT) distribution.

Decision: no dependency on the Moderne stack, not even interoperability. We adopt the concepts.

1.5 Known limits we must not reproduce

  • Missing types fail silently. Type-gated logic silently matches nothing when resolution fails. Unwind must instead report the tier it reached (§5.2).
  • The whole LST in memory causes out-of-memory failures on large repos. Unwind keeps facts, not trees, and writes them to disk (§5.5).
  • Noisy whole-file rewrites and huge pull requests were found too unwieldy to review (Adyen case study). Play generates into a new target repo slice by slice, with holes clearly marked.
  • Stateful recipes are bug-prone (cronn). Unwind recipes are pure functions with golden fixtures (§4.3).

1.6 Concept transfer

Concept transfer: OpenRewrite / Moderne → UnwindConcept transfer: OpenRewrite / Moderne → Unwinddeterministic (engine / CLI)LLM agentartifacthuman / externalnew in destinationOpenRewrite / ModerneUnwind destination conceptlives inLST (typed, lossless)Semantic Model (typed + bound, not lossless)RewindLST artifacts · mass ingestModel Store (per-commit, multi-repo)StoreSearch recipes + data tablesDetector recipes → typed fact tablesRewindPrethink context filesSpec + tagged layer docsRewind → SpecVisitor / ScanningRecipeTarget recipe (scan → generate → edit)PlayDeclarative YAML recipeBlueprint (whole service / module)KitRecipe marketplace / BOMRecipe Book · Target Kit (per client)Kit repoPreconditionsappliesTo(spec node, kit profile)PlayRewriteTest before / afterGolden fixtures + idempotence runKitJavaTemplate / RefasterTemplates mined from exemplar codeKitModdy · learn_recipeExemplar → recipe promotionPlay / KitModerne MCP serverunwind mcp (adapter over the CLI)MCPMarkers (SearchResult)Holes (@unwind-hole) + provenancePlayCopied as concepts only: no Moderne / OpenRewrite code or runtime dependency (MSAL & proprietary licensing).
OpenRewrite/Moderne concepts mapped to Unwind equivalents
OpenRewrite / ModerneUnwind destination conceptLives inDoc
LST (typed, lossless)Semantic Model: typed and bound, but not lossless. Unwind never edits the source, so format preservation is not neededRewind05
LST artifacts, mass ingestModel Store: per-commit, versioned, multi-repo queriesStore05
Search recipe + data tablesDetector recipes emitting typed fact tablesRewind05
Prethink contextSpec plus tagged layer docsRewind → Spec02
Visitor / ScanningRecipeTarget recipe: scan Spec → generate → editPlay04
Declarative YAML recipeBlueprint: a whole service or moduleKit04
Recipe marketplace / BOMTarget Kit / Recipe Book: versioned per clientKit repo04
PreconditionsappliesTo(specNode, profile)Play04
RewriteTest before/afterGolden fixtures (Spec fragment → expected files) plus an idempotence runKit04
JavaTemplate / RefasterParameterised templates mined from exemplar codeKit04
Moddy / learn_recipeExemplar → recipe promotionPlay04
Moderne MCPunwind mcp: a thin adapter over the CLISurfaces02
SearchResult markersHoles (@unwind-hole) and provenance markersPlay04

1.7 What stays uniquely Unwind

  • [MUST] / [SHOULD] / [DON'T] tagging of every documented item, with rationale.
  • Completeness proven by set arithmetic (scan − docs, Spec − target), never asserted.
  • Grilling plus domain-expert verdicts, which decide whether a behaviour deserves to be reproduced.
  • Cross-stack rebuild with verification, which OpenRewrite does not attempt.
  • New in this design: behaviour parity against the legacy app (06) and context-gap interviews (07).

02 · Destination architecture

In short: A self-hosted Unwind Server is the shared backbone from day 0: the team's system of record and UI, git- and SQLite-backed, receiving artifacts only, with slices as the first-class unit of work (03). Around it sit one shared model package, two plugins and one contract between them. Rewind understands a legacy system, slice by slice, and compiles it into a stack-neutral Spec. Play rebuilds the Spec from a client's Target Kit. A single engine sits behind a CLI-first surface that does the work next to the code and pushes results to the server. MCP is a thin adapter. Deterministic code owns the facts and the structure; the LLM owns semantics and holes; completeness is always computed.

Unwind destination: Server · Rewind → Spec → PlayUnwind destination: Server · Rewind → Spec → Playdeterministic (engine / CLI)LLM agentartifacthuman / externalnew in destinationREWIND · understand the sourceper slice · runs locally via the CLISource repoany language · gitSemantic Modeltyped facts · ids · edgestree-sitter + compilersLayer docs + grilltagged [MUST] / [SHOULD] / [DON'T]coverage = manifest − docsrw-spec (compile)model + docs → Spec fragmentpush artifacts per sliceUNWIND SERVER · shared backbone from day 0slices + owners · artifacts in git · state + index in SQLite · basic UI · token authconverges slice fragments into one project Spec · tracks progress per sliceartifacts only: source code never leaves the developer's machineconvergedSPEC · stack-neutral typed IRentities · endpoints · events · operationsscenarios · priorities · provenanceContext gapsagents find what code can't say→ interview briefs → stakeholders→ answers enrich the SpecBehaviour parityscenarios run vs legacy → goldensreplay vs target → parity %observations enrich the SpecPLAY · rebuild in the target stackper slice · ordered by seamsTarget Kitclient recipe book(own git repo)Planchoose / tailor Kitphasing · riskGeneraterecipes → code + holescorrect by constructionFill holesbusiness logic onlyVerifyre-scan target − SpecMUST completeness %gaps → regenerate / refill until verified (loop mode)push verify resultsTarget repo · new stackgenerated structure + filled holes + native parity testsSurfaces: unwind CLI (does the work, next to the code) · Unwind Server UI · MCP adapter later
The big picture: Unwind Server as the shared backbone · Rewind → Spec → Play per slice

2.1 The shape

The flow. Developers and agents run the unwind CLI locally, next to the code → they push artifacts per slice to the Unwind Server (the shared record and UI) → the server converges slice fragments into one project Spec → Play rebuilds per slice from the Spec plus a Target Kit → verification results are pushed back, so progress, completeness and parity are visible per slice. Code never leaves the developer's machine; everything works offline and syncs when logged in.

Components, in flow order:

  1. Unwind Server (03): projects, slices and owners, artifacts in git, state and index in SQLite, basic UI, token auth. Present from day 0.
  2. @unwind/model (§2.3): the shared ids and schemas every other part speaks.
  3. Rewind (§2.4): understands the source, per slice, and produces Spec fragments.
  4. The Spec (§2.5): converged on the server into one project Spec; the only Rewind → Play contract.
  5. Play (§2.6) with Target Kits (§2.7): rebuilds per slice and verifies.
  6. Surfaces (§2.8): CLI first, server UI, MCP later.
   ┌──────────────── UNWIND SERVER (shared backbone, day 0) ────────────────┐
   │ projects · slices + owners · artifacts (git) · state/index (SQLite) ·   │
   │ convergence → project Spec · metrics · UI · token auth                  │
   └───────▲───────────────────────▲─────────────────────────▲───────────────┘
      push │ per slice        push │ fragments          push │ verify results
            ┌──────────── @unwind/model (shared contract) ────────────┐
            │ ids · Semantic Model · Spec (IR) · Kit schema ·          │
            │ verification + state schemas · versioning                │
            └──────────────────────────────────────────────────────────┘
 SOURCE REPO ─► REWIND ─► Semantic Model ─► tagged docs ─► SPEC ──────┐
              (understand)   (facts)        (LLM + grill)  (neutral)  │
                                                                      ▼
 CLIENT KIT REPO ─► TARGET KIT (profile · conventions · recipes ·    PLAY ─► TARGET REPO
 (mined from ref app)  blueprints · golden fixtures)                 (rebuild)   │
                                                                      ▲          │
                                 VERIFY (target re-scanned by Rewind ◄───────────┘
                                         → diffed against the Spec)
 SPEC ─► CONTEXT GAPS ─► interview briefs ─► (external interview tool) ─► answers ─► enrich SPEC
 SPEC ─► PARITY SCENARIOS ─► run vs LEGACY (observe → goldens → enrich SPEC)
                        └──► run vs TARGET (same scenarios → behavioural parity)

2.2 Today compared with the destination

Today vs destination: Server + the Rewind / Play splitToday vs destination: Server + the Rewind / Play splitdeterministic (engine / CLI)LLM agentartifactnew in destinationTODAY · one plugin (uw-*), single userREWIND plugin (rw-*) · per slicePLAY plugin (pl-*) · per sliceSHARED · server + surfaces (day 0)uw-scanscan.mjs · seed-layers.mjsrw-scan+ Semantic Model tiers (types, calls)uw-analyze-*layer specialistsrw-analyze-*unchanged · tagged docsuw-verify · uw-completemanifest − docsrw-verify · rw-completeunchangeduw-grillhotspots · questionnairesrw-grillquestionnaires → context gapsrw-speccompile the stack-neutral Specrw-context-gapsinterview briefs · ingest answersrw-observescenarios vs legacy → goldenspl-kitmine · test · version Target Kitsuw-planfree-text stack decisionspl-planchoose / tailor a Kit (typed profile)pl-generaterecipes + blueprints → code + holesuw-build · uw-build-layerLLM writes every linepl-build-layerfills holes onlymerge-rebuild-map · verify-rebuildnames + fieldspl-verify · pl-parity+ types · holes · behaviourlocal docs/unwind/ onlyno shared state · no slicesUnwind Server (unwind serve)slices · git + SQLite · UI · token auth · convergenceuw-graph · uw-dashboard · uw-publishscripts in skills/scriptsunwind CLI · MCP adapterone engine; skills shell out; CLI pushesSPEC — the only Rewind → Play contractrenamed /evolved
Today's uw-* pipeline vs the Rewind/Play destination
Today (uw-*)DestinationChange
uw-scan → scan-manifest.jsonrw-scan → Semantic Model (typed, bound)Additive manifest fields; extra tiers (05)
uw-analyze-*, uw-verify, uw-completeSame, under rw-*Rename plus aliases
uw-grill → questions/rw-grill, folded into context gapsOne questionnaire mechanism (07)
(none)rw-spec → SpecNew: the Rewind→Play contract
(none)rw-observe / pl-parityNew: behaviour parity (06)
uw-plan, free-text stack decisionspl-plan: choose or tailor a KitTyped profile replaces free text
uw-build-layer writes every linepl-build: generate from Kit, LLM fills holesDeterministic first (04)
verify-rebuild: names, method+path, field-name JaccardPlus field types, holes, behavioural parityStronger verdicts
skills/scripts/*.mjsunwind CLI (scripts become shims)Consolidation (§2.8)
Single user, local docs/unwind/Unwind Server: shared, git + SQLite, slices, UINew from day 0 (03)

2.3 @unwind/model: the shared contract

The only package both plugins import. It holds:

  • The id scheme. Today this is manifest/candidates.ts and symbolId in manifest/manifest-schema.ts (kind:path:name). The destination version adds class-qualified, overload-safe ids (e.g. method:src/a.ts:Orders.create(2)), because methods and overloads collide today. Ids remain the join key across the model, Spec, coverage, graph, kit output and verification.
  • Schemas: Semantic Model (additive extension of ScanManifest), Spec, Kit manifest, rebuild state and verification (today's graph/rebuild-state-schema.ts and graph/rebuild-verification.ts types), the gap register and Scenario.
  • Versioning. Every artifact carries schemaVersion. Changes are additive-only, as with the manifest today; breaking changes need a migration and a major bump.

2.4 Rewind (understand)

Today's pipeline (scan → seed → analyze → verify-coverage → complete → grill), extended with:

  1. Semantic Model: a typed, bound fact model in tiers T0/T1/T2 (05).
  2. Detector recipes: layers/contract-detectors.ts split into a fixture-tested registry.
  3. rw-spec: compiles manifest, graph and tagged docs into the Spec.
  4. rw-observe: runs Spec scenarios against the legacy app and enriches the Spec (06).
  5. rw-context-gaps: finds what the code can't answer and produces interview briefs (07).

Every step runs locally through the CLI, scoped to a slice when one is claimed (--slice), and unwind push sends the resulting artifacts to the server, which converges the fragments into the project Spec (03 §3.8).

Rewind is useful on its own for documentation, onboarding, audits and due diligence, without ever running Play.

2.5 The Spec: the only Rewind → Play contract

A stack-neutral, typed IR of everything the target must preserve. Where a standard exists, the Spec embeds it rather than inventing a format: JSON Schema for shapes, OpenAPI-style operations for endpoints.

Node kindCarries
entityTyped fields (neutral types), keys, FKs/relations, enums, physical name
endpointFull composed path, method, params, request/response schema refs, auth, status codes, handler ref
event / jobPayload schema, producer/consumer, schedule
configEnv/config keys, defaults, secrets flag
integrationExternal system, protocol, calls made
operationA named business rule: description, doc ref, behavioural assertions, linked scenarios
scenarioGiven/when/then at a boundary (06)

Every node carries id, priority (MUST/SHOULD/DONT plus rationale), docRef, sourceRef and provenance: scanned, inferred, observed, interview:<role>:<date> or expert.

{
  "schemaVersion": "1.0",
  "entities": [{
    "id": "table:src/db/schema.ts:orders",
    "name": "orders", "priority": "MUST",
    "fields": [
      { "name": "id", "type": "uuid", "pk": true },
      { "name": "customerId", "type": "ref(customers)", "nullable": false },
      { "name": "status", "type": "enum(pending,paid,shipped,cancelled)" },
      { "name": "totalCents", "type": "int" },
      { "name": "createdAt", "type": "datetime", "default": "now" }
    ],
    "docRef": "layers/database/tables.md#orders",
    "provenance": ["scanned", "inferred"]
  }],
  "endpoints": [{
    "id": "endpoint:src/routes/orders.ts:POST /api/orders",
    "method": "POST", "path": "/api/orders", "priority": "MUST",
    "request": { "$ref": "#/schemas/CreateOrder" },
    "responses": { "201": { "$ref": "#/schemas/Order" }, "422": { "$ref": "#/schemas/ValidationError" } },
    "auth": "session", "handlerRef": "function:src/services/orders.ts:createOrder",
    "provenance": ["scanned", "observed"]
  }],
  "operations": [{
    "id": "operation:src/services/orders.ts:applyDiscount",
    "priority": "MUST",
    "rule": "Orders over 10000 cents get 5% off; never stack with coupon discounts.",
    "rationale": "Contractual with wholesale customers (interview:stakeholder:2026-11-02)",
    "docRef": "layers/service/orders.md#applydiscount",
    "scenarios": ["scenario:orders:discount-threshold"],
    "provenance": ["inferred", "interview:stakeholder:2026-11-02"]
  }]
}

A Spec can also be hand-written, which makes Play usable for greenfield work.

2.6 Play (rebuild)

  1. pl-plan. The interview narrows to choosing or tailoring a Kit: phasing, re-use and risk. It writes a typed profile in place of today's free-text rebuild-decisions.json (which stays as a record).
  2. Resolve the Kit. Pin kit@version in rebuild-state.json.
  3. Generate. Recipes and blueprints run over the Spec and write target files, map entries (the existing rebuild-map/*.json format) and holes (04). The scaffold recipe finally sets the dormant config.scaffolded flag in rebuild-state-schema.ts.
  4. Fill. pl-build-layer subagents fill only holes plus any unmapped [MUST] items.
  5. Verify. Structural and typed diff (an extended graph/rebuild-verification.ts), hole accounting, and behavioural parity (pl-parity).
  6. Loop. The existing loop-until-verified mechanics, with completeness % and parity % as termination signals.
  7. Push. Rebuild maps and verification results are pushed per slice; the server tracks each slice through its Play states (planned → generating → filling → verified → cut-over) and orders slices by their seams (03 §3.7–3.8).

Fallback. If there is no Kit match, or no Node/core, today's pure-LLM uw-build flow runs unchanged, and the skill says so.

2.7 Target Kits

Per-client, versioned git repos that encode the client's golden path: stack profile, conventions, type map, recipes, blueprints and golden fixtures. They are mined primarily from the client's reference app. Starter kits (first: hono-drizzle-zod) ship with Play. See 04.

2.8 Surfaces: one engine, CLI first, a shared server from day 0

Surfaces: one engine, CLI-first, shared server from day 0Surfaces: one engine, CLI-first, shared server from day 0deterministic (engine / CLI)new in destinationartifactskills (rw-*, pl-*) · CI · humansany agent via Bashunwind CLI · PRIMARYdoes the work · --json · offline-firstteam · reviewers · leadsbrowser (basic UI)unwind serve · DAY 0system of record · slices · UIMCP clientsnon-CLI agentsunwind mcp · laterstdio adapter, 1:1 with CLI / APIpushpull@unwind/enginemodel · rewind (scan, detectors, spec, observe, context-gaps) · play (kits, recipes, generate, verify, parity) · sliceslocal docs/unwind/working copy · works offlineproject git reposerver · artifacts · full historynode:sqliteauth · slices · metrics · indexThe CLI does the work next to the code; the server stores, merges, indexes and shows it. MCP and HTTP add no logic of their own.
One engine; CLI primary; server as the shared system of record; MCP as adapter
@unwind/engine (library: model, rewind, play, kits, parity, gaps, slices)
     ├── unwind CLI      ← PRIMARY execution surface. Skills shell out to it. CI / humans / any agent.
     │                     works fully offline; when logged in, pushes artifacts to the server
     ├── unwind serve    ← DAY 0: shared system of record + basic UI (Hono API, git + SQLite, slices)
     └── unwind mcp      ← later: thin stdio MCP adapter; tools map 1:1 to CLI commands / API routes

Division of labour:

  • The CLI does the work (scan, analyze, generate, verify) next to the code.
  • The server stores, merges, indexes and shows it: projects, slices, history, convergence and metrics.
  • Source code never goes to the server; only artifacts do. See 03 for storage, auth (simple bearer tokens via unwind login), push/pull, slices, convergence, UI and the API.

Why the CLI rather than MCP for the skills:

  • The skills already shell out to node skills/scripts/*.mjs via _resolve-plugin-root.sh/ensure_unwind_core. That is a CLI in all but name, so this is a consolidation.
  • A CLI works in CI, for humans and from any agent, with no daemon or connection state. MCP availability varies by client and would weaken the graceful fallback.
  • --json output and exit codes make it testable. Long-running steps (generate, parity) fit processes better than tool calls.

Command map:

unwind rewind  scan | seed | coverage | grill-brief | spec | observe | context-gaps | context-ingest
unwind play    plan-brief | kit mine | kit test | generate | merge | verify | parity
unwind slices  propose | list | claim | release | status
unwind         login | logout | whoami | project link|create | push | pull | status
unwind         graph | publish | serve | mcp
common flags   --project <src> --target <dir> --kit <repo@ver> --slice <id> --json --plan (dry run)

Distribution is an npm package (npx @unwind/cli) plus the server Docker image. ensure_unwind_core becomes ensure_unwind_cli. The current .mjs scripts become one-line shims for one release.

Server tech stack (detail in 03 §3.3):

  • Hono on Node (@hono/node-server), with zod validation (@hono/zod-validator) and hono/client RPC types shared with the CLI and UI;
  • React + Vite UI with TanStack Query (TanStack Router recommended) and Tailwind + daisyUI, dark by default and mapped onto the dashboard tokens, reusing the React Flow/ELK graph and DocsViewer from packages/dashboard;
  • node:sqlite (Drizzle recommended for schema and migrations) plus system git;
  • one Docker image serving API and UI from the same Hono app.
  • This is the stack of the pilot starter kit, so Unwind dogfoods its own target kit.

The MCP adapter and the HTTP API add no logic of their own, so behaviour is identical across surfaces.

2.9 Packaging and naming

  • One repo and two plugins: Rewind (rw-* skills) and Play (pl-* skills). Unwind stays the umbrella brand.
  • Packages: @unwind/model, @unwind/engine (initially today's @unwind/core, renamed and grown), @unwind/cli, @unwind/server (Hono API + storage) and @unwind/app (today's @unwind/dashboard, grown into the server UI).
  • The uw-* skills stay as deprecated aliases for one release.

2.10 Invariants

  • Hybrid rule. Deterministic code owns every verifiable fact and every structural artifact. The LLM owns semantics and holes. Completeness is computed by set arithmetic over ids.
  • Additive schemas only. No reshaping FileSymbols; add optional fields.
  • AST and real parsers over regex, on both sides: detectors read ASTs, and recipes edit target files with ts-morph or tree-sitter.
  • Graceful fallback at every step, announced to the user.
  • Files are the source of truth. Locally, that is docs/unwind/. On the server, it is the project's git repo. SQLite holds auth and operational state plus a rebuildable index.
  • Code stays local. Only artifacts are pushed to the server (03 §3.11).
  • Slices are first-class. Candidate ids define slice membership, and convergence is set arithmetic over them (03 §3.7–3.8).

03 · Unwind Server and slices

In short: From day 0, a self-hosted Unwind Server (one Node process in one Docker image) is the team's shared system of record. It provides a small web UI. The CLI keeps doing the work locally and pushes artifacts only: source code never leaves developers' machines. Storage is git (reviewable artifacts, full history) plus SQLite (auth, slices, metrics, query index). Slices are first-class units of work that run through both halves. Teams analyze them in parallel, the server converges their Spec fragments into one project Spec, and the same slices later become Play's strangler-style rebuild units.

Unwind Server: shared system of record, code stays localUnwind Server: shared system of record, code stays localdeterministic (engine / CLI)LLM agentartifacthuman / externalnew in destinationDEVELOPER MACHINES · CISource repocode never leaves this boundarySkills + agents (rw-* / pl-*)analyze · grill · fill holescall the CLI via Bashunwind CLIscan · spec · generate · verifylogin · push · pull · statusdocs/unwind/ (working copy)~/.config/unwind/credentials.json (0600)secret scrub + allow-list before pushartifacts onlyno source codepushpullBearer uwt_…UNWIND SERVER · one Docker image · unwind serveHono APIzod-validated · token authhono/client types → CLI + UIUI (React + Vite)TanStack Query · daisyUIslice board · docs · graphEngine: convergence · metrics · indexingmerge slice Spec fragments · uncovered / overlaps / conflicts / seamsgit: repos/<project>.gitdocs · Spec · findings · briefscommit per push, by owneroptional mirror to GitHubnode:sqlite (Drizzle)users · hashed tokens · slicesruns · metrics · FTS5 searchindex rebuildable from gitTeam in the browser: leads, reviewers, experts (token login)TLS / SSO via a reverse proxy later. Backup = copy the /data volume. The server never needs repository access.
Unwind Server: CLI pushes artifacts, code stays local

3.1 Why a server from day 0

On a large codebase, one person running Unwind in one session does not scale.

  • Analysis is spread across people and agents over weeks.
  • Reviewers who are not developers need to read the docs and answer questions.
  • Leads need to see progress, and the Spec has to be one shared thing, not N local copies.

Files in each developer's docs/unwind/ are fine for one person. For a team they need a home with history, ownership, progress and a UI. The server provides that home without changing how the work is done: the CLI stays the execution surface (doc 02 §2.8), and the server stores, merges, indexes and shows.

3.2 Shape: unwind serve

  • A single Node process (unwind serve), shipped as one Docker image (ghcr.io/nearform/unwind-server).
  • One Hono app serves the HTTP JSON API and the built UI as static assets.
  • Everything lives under one data volume:
/data
  unwind.db                 node:sqlite: auth, projects, slices, runs, metrics, index (WAL mode)
  repos/<project>.git       bare git repo per project: the reviewable artifacts
  tmp/                      push staging
docker run -d --name unwind -p 8080:8080 \
  -v unwind-data:/data \
  -e UNWIND_ADMIN_TOKEN="$(openssl rand -hex 32)" \
  ghcr.io/nearform/unwind-server:latest
  • Backup is a copy of /data. Run the SQLite online backup, or stop the container first.
  • TLS and SSO come from a reverse proxy in front (Caddy, nginx, Cloudflare Tunnel). They are out of scope for day 0 (§3.12).
  • Running without Docker works too: npx @unwind/cli serve --data ./unwind-data.

3.3 Tech stack

The server is built on the same stack as the pilot starter kit (hono-drizzle-zod, doc 04). Unwind therefore dogfoods its own target kit: the server is a living reference app that the kit can be mined from and tested against.

ConcernChoiceNotes
HTTP APIHono on Node (@hono/node-server)Typed routes. The one app also serves the UI's static assets.
Validationzod via @hono/zod-validatorThe schemas live in @unwind/model, shared with the CLI.
Typed clienthono/client RPC typesThe CLI and UI call the API through the same inferred types, with no hand-written client.
UIReact + ViteIt grows out of packages/dashboard.
Server state in UITanStack QueryCaching, refetch and optimistic slice claims.
Routing in UITanStack Router (recommended)Type-safe routes and search params. It replaces today's hand-rolled urlState.ts over time.
Components and themingTailwind + daisyUIDark default. daisyUI themes are mapped onto the existing dashboard tokens (--color-* → --c-*).
Reused UIReact Flow + ELK graph, DocsViewer/MarkdownViewTaken as they are from packages/dashboard.
Databasenode:sqlite (built into Node ≥22.5)WAL mode. FTS5 for search.
ORM / migrationsDrizzle over node:sqlite (recommended)Typed schema plus generated migrations. Matches the starter kit.
GitSystem git, called from the server process (recommended), installed in the imageFully compatible with real git (packs, receive, mirror push). isomorphic-git is a fallback for git-less environments.
PackagingOne Docker image (node:22-slim + git)unwind serve is the entrypoint.

3.4 Storage: git for artifacts, SQLite for state and index

Git is the source of truth for reviewable artifacts. Each project has one bare repo, with the same layout as a local docs/unwind/ plus slice folders:

architecture.md
layers/**                        tagged layer docs (shared)
slices/<slice-id>/
  slice.json                     scope, owners, state
  docs/**                        slice-specific layer docs
  spec.fragment.json             this slice's Spec fragment
  findings.json · gaps.json      grill findings, context gaps
spec/project.spec.json           converged Spec (written by the server, §3.8)
questions/** · interviews/**     questionnaires, briefs, responses
rebuild/**                       rebuild-map, state, verification (Play)
.cache/scan-manifest.json        manifest (ids are the join key)
  • Every push is a commit authored by the token's owner: Author: Dana Lee <dana@client.com>, with a trailer Unwind-Slice: orders.
  • That gives history, diff, blame and revert for free. It also allows an optional mirror push to a GitHub/GitLab repo for clients who want the artifacts next to their code.

SQLite holds operational state plus a query index:

Table groupContentsRebuildable from git?
users, tokensIdentity and hashed tokens (§3.5)No (back them up)
projects, slices, claimsOwnership, state, claimsPartly: slice state is mirrored to slice.json
runsWho ran which CLI command, when and against which commit, with result metricsNo (audit)
metricsTime series: doc coverage, context coverage, parity %, completeness %, convergence % per slice and per projectYes, recomputable per commit
spec_nodes, gaps, findingsParsed Spec, gap register and grill findings, for queries and the UIYes
search (FTS5)Full-text search over docs and the SpecYes

unwind serve --reindex rebuilds every "Yes" table by walking git history, so the index can always be thrown away and rebuilt.

3.5 Auth: simple bearer tokens

The model is deliberately minimal.

  • Bootstrap. UNWIND_ADMIN_TOKEN (env) is an admin token on first start. The admin creates users, for example in the UI, by name and email.
  • Personal tokens. Users create tokens in the UI (Settings → Tokens). Each has a name, scopes and an optional expiry. A token is shown once as uwt_<32 random bytes base62> and stored sha256-hashed.
  • Scopes: read (view and pull), write (push, claim slices, answer questions) and admin (users, projects, tokens). Project membership is a simple allow-list per project.
  • Requests use Authorization: Bearer uwt_…. The UI uses the same token, kept in an HttpOnly cookie after a token-paste login.

CLI:

unwind login https://unwind.internal.client.com   # prompts for a token, verifies via /api/me
# → ~/.config/unwind/credentials.json (mode 0600): { "servers": { "<url>": { "token": "uwt_…", "user": "dana" } }, "default": "<url>" }
unwind whoami                                      # user, server, scopes, projects
unwind logout [--server <url>]

UNWIND_TOKEN / UNWIND_SERVER env vars override the file, for CI.

3.6 CLI sync: local working copy, push and pull

  • The local docs/unwind/ stays the working copy. Every command still works with no server, in line with the graceful-fallback principle. The server is additive.
  • The project is linked to a server project in docs/unwind/.unwind.json: { "server": "<url>", "project": "acme-billing" }.
unwind project link acme-billing            # or: unwind project create acme-billing
unwind status                               # local vs server: ahead/behind, changed files, slice claims
unwind push [--slice orders] [-m "message"] # changed files → one commit on the server
unwind pull [--slice orders]                # fast-forward the working copy

How push works:

  1. The CLI sends a bundle of changed files (paths, content and hashes) plus the base revision it last pulled.
  2. The server writes a commit if base == HEAD, or if none of the touched paths changed since base (a non-conflicting path-level merge).
  3. Otherwise it returns 409 with the conflicting paths. The CLI pulls, replays and retries.
  4. Because each slice owns its own folder (§3.4), concurrent pushes from different slices almost never conflict.

Before a push, the CLI:

  • scrubs secrets (token, key and connection-string detectors; the push is refused if anything is found, with --allow to override per file);
  • enforces the artifact allow-list (§3.11).

Skills call the CLI as they do today. When logged in and linked, long-running skills (rw-analyze, rw-grill, pl-build) push at checkpoints, so progress shows up live in the UI.

3.7 Slices: the first-class unit of work

Slices: parallel analysis, convergence, Play by sliceSlices: parallel analysis, convergence, Play by slicedeterministic (engine / CLI)LLM agentartifactnew in destinationCodebaseorders · dana, sambilling · leecatalog · agent + kimauth · raviproposed from import graphPARALLEL REWINDorders: analyze → grill → fragbilling: analyze → grill → fragcatalog: analyze → grill → fragauth: analyze → grill → fragConvergencemerge fragments by candidate iduncovered: ids in no sliceoverlaps: id in 2+ slicesconflicts: same id, differsseams: cross-slice edgesconvergence % → 100Project Specspec/project.spec.jsonseams = interface contractsconverged at 100%PLAY BY SLICE · strangler order from the seam graphauthcut-overno upstream seamscatalogverifiedseams: adapters → legacyordersfillingseams: adapters → legacybillingplannedseams: adapters → legacyRewind states: proposed → claimed → analyzing → covered → grilled → spec-ready → accepted. Play: planned → generating → filling → verified → cut-over.
Slices: parallel analysis, convergence, Play by slice

A slice is a bounded area of the codebase that one owner or team carries from analysis through rebuild:

{
  "id": "orders",
  "name": "Orders & checkout",
  "scope": {
    "paths": ["src/orders/**", "src/checkout/**", "db/migrations/*orders*"],
    "layers": ["service", "api", "database"],
    "capabilities": ["ordering", "payments"]
  },
  "owners": ["dana", "sam"],
  "rewind": { "state": "covered" },
  "play": { "state": "planned" },
  "metrics": { "coverage": 0.94, "contextCoverage": 0.61, "parity": null, "completeness": null }
}

Proposal. After the first local full scan, unwind slices propose (or the UI) suggests slices deterministically:

  • clustering the import graph (today's importMap, later call edges);
  • seeded by top-level directories and layers;
  • balanced by candidate count.

Humans then rename, merge, split and assign. The scope globs resolve to a set of candidate ids, and that set is the slice's membership.

States run through both halves:

HalfStates
Rewindproposed → claimed → analyzing → covered → grilled → spec-ready → accepted
Playplanned → generating → filling → verified → cut-over
  • Transitions are mostly computed. For example, covered means slice coverage is 100% per verify-coverage over the slice's ids, and verified means [MUST] completeness is at its target.
  • claimed, accepted and cut-over are explicit human actions.

Per-slice artifacts: docs, coverage and gaps, Spec fragment, grill findings, context gaps, interview briefs and, later, rebuild map, verification and parity.

3.8 Convergence: fragments into one project Spec

The server merges the slices' spec.fragment.json files into spec/project.spec.json, using candidate ids as the join key. This is the same set arithmetic Unwind already uses for coverage. On every push it reports:

SignalDefinitionResolution
UncoveredManifest ids in no slice scopeWiden a scope, or create a slice
OverlapsAn id in two or more slice scopesAssign one owner; the other slice references it
ConflictsThe same id with a different priority or content across fragmentsOwners resolve in the UI; the decision is recorded with rationale
SeamsImports, calls or reads/writes crossing slice boundariesBecome interface contracts in the Spec that both slices must honour
  • Convergence % = ids that are in exactly one slice, conflict-free and spec-ready, divided by all manifest ids (excluding [DON'T]). The project Spec is converged at 100%.
  • Seams drive Play. They form a slice dependency graph, which gives the strangler-style rebuild order. A slice can be rebuilt and cut over once its upstream seams are satisfied, either by already-rebuilt slices or by an adapter onto the legacy app. Seams are also first-class scenarios for parity (doc 06).

3.9 UI (basic, day 0)

The UI is built by evolving packages/dashboard into the app served by unwind serve, reusing its graph and docs components.

ViewShows
ProjectsList, with last push and headline metrics
Project overviewCoverage, context, convergence, parity and completeness over time (from metrics); recent activity
Slice boardKanban by state (Rewind and Play lanes), owners and a Claim button; filter by owner or capability
Slice detailRendered docs (existing DocsViewer), gaps, Spec fragment, findings, open questions and its metrics
GraphExisting React Flow + ELK view, coloured by slice, with seams highlighted
ConvergenceUncovered, overlaps and conflicts (with resolve actions), and seams
Questions & interviewsGrill questionnaires and interview briefs; answer in place (writes back via a commit)
ActivityGit log per project and slice, with diffs
SettingsTokens, users and project members (admin)

3.10 API sketch

All routes are JSON and bearer-authenticated, typed through hono/client. The future MCP adapter maps onto these routes.

Method and routePurposeScope
GET /api/meCurrent user, scopes, projectsread
GET/POST/DELETE /api/tokensPersonal token managementread / write
GET/POST /api/projectsList or create projectsread / admin
GET /api/projects/:pOverview and headline metricsread
GET/POST /api/projects/:p/slicesList slices, or create/propose themread / write
PATCH /api/projects/:p/slices/:sRename, rescope or change statewrite
POST /api/projects/:p/slices/:s/claimClaim or release a slicewrite
POST /api/projects/:p/pushFile bundle plus base revision → commit (409 on conflict)write
GET /api/projects/:p/pull?since=<rev>Changed files since a revisionread
GET /api/projects/:p/artifacts/*pathRead any artifact (optionally ?rev=)read
GET /api/projects/:p/specThe converged project Specread
GET /api/projects/:p/graphrebuild-graph.json plus slice colouringread
GET /api/projects/:p/convergenceUncovered, overlaps, conflicts, seams, %read
GET/POST /api/projects/:p/runsRecord or list CLI runs and metricsread / write
GET /api/projects/:p/search?q=FTS5 over docs and Specread

3.11 Security and data policy

  • Code stays local. The server never needs repository access or credentials. The CLI scans and analyzes locally and pushes artifacts only:
    • manifests (paths, symbol names, line numbers);
    • docs;
    • Spec;
    • findings and questions;
    • rebuild maps and verification.
  • Code appears on the server only where docs already quote it, such as grill evidence quotes. Clients can disable quotes per project (policy.quotes: false), in which case the CLI strips fenced code from docs on push.
  • Allow-list on push. Only known artifact paths are accepted, on both the server and the CLI.
  • Secret scrubbing before push (§3.6).
  • Audit: every write is a git commit plus a runs row, tied to a named user.
  • Self-hosted by default. The server runs inside the client's network next to the code, so artifacts never leave it unless the client mirrors them.

3.12 Scale and scope

Large codebases:

  • Run one full scan locally, push the manifest, and propose slices.
  • People and agents then analyze slices in parallel: each rw-analyze --slice orders run sees only that slice's seeds.
  • The server's convergence keeps the whole picture honest.
  • Incremental refresh (detect-changes) re-opens only the slices whose ids changed.

Out of scope for day 0 (possible later options):

  • SSO/OIDC (put a reverse proxy in front for now);
  • multi-tenant SaaS;
  • server-side repo cloning and scheduled re-scans;
  • a Cloudflare-native variant;
  • fine-grained RBAC beyond read/write/admin plus project allow-lists;
  • real-time collaborative editing (a git commit per push is enough).

04 · Target Kits, recipes, blueprints and holes

In short: A Target Kit is a client's golden path packaged as a versioned git repo: a typed stack profile, conventions, a type map, recipes (pure Spec→code generators) and blueprints (recipes composed into whole services). Play runs the Kit over the Spec to generate code that is correct by construction. Anything it can't derive becomes an explicit hole for the LLM. Kits are mined mainly from the client's own reference app.

Target Kit anatomy: recipes and blueprintsTarget Kit anatomy: recipes and blueprintsdeterministic (engine / CLI)artifacthuman / externalclient-kit/ (own git repo)kit.yaml name · version · extendsprofile.yaml runtime · framework · orm …conventions.yaml naming · layout · errorstype-map.yaml neutral → target typesrecipes/ entity-drizzle-table/ recipe.ts templates/ fixtures/spec.json fixtures/expected/**blueprints/ crud-service.yamlexemplars/ mined-from provenanceRecipe contract (ScanningRecipe-style)appliesTopreconditionscanaccumulate over Specgenerateper node → fileseditshared files (AST)outputs: target files · rebuild-map entries · holespure · deterministic · idempotent · never clobbers filled holesBlueprint = declarative composition of recipesblueprints/crud-service.yaml (crud-service@hono-drizzle-zod)scaffoldentity → drizzle-tableentity → zod-schemaentity → repositoryendpoint → hono-routeroute-registry (edit)endpoint → testoptions: pagination=cursorids=uuidv7loadsKits inherit (extends: starter/hono-drizzle-zod) · each rebuild pins kit@version in rebuild-state.json
Anatomy of a Target Kit

4.1 Why kits

Today uw-build-layer writes every line with an LLM. The trouble with that:

  • Forty tables come out forty slightly different ways.
  • Conventions drift between slices.
  • Every re-run is a fresh roll of the dice.

Most of a rebuilt service is structural: schema, routes, validators, wiring, repositories, test skeletons. Structure is fully determined by the Spec plus the client's conventions. Kits make that part deterministic, reviewable and repeatable, and leave the LLM to do what only it can: business logic.

Kits are per client because "the target" is never just "Hono + Drizzle". It is this client's Hono + Drizzle: their error envelope, logging, id strategy, folder layout and internal SDKs.

4.2 Kit repo layout

acme-kit/
├── kit.yaml            # name, version, extends, compatible spec schemaVersion
├── profile.yaml        # typed stack profile
├── conventions.yaml    # naming, layout/module map, error model, logging, ids, pagination
├── type-map.yaml       # Spec neutral types → target types
├── recipes/
│   └── entity-drizzle-table/
│       ├── recipe.ts           # or recipe.yaml for simple declarative recipes
│       ├── templates/table.ts.tmpl
│       └── fixtures/{spec.json, expected/**}
├── blueprints/
│   ├── crud-service.yaml
│   └── full-app.yaml
└── exemplars/          # provenance: ref-app files each recipe was mined from (commit-pinned)
# kit.yaml
name: acme-kit
version: 1.4.0
extends: unwind/hono-drizzle-zod@^1     # inherit the starter kit, override what differs
spec: ">=1.0 <2"
# profile.yaml
runtime: cloudflare-workers
language: typescript
framework: hono
apiStyle: rest
orm: drizzle
db: d1
validation: zod
auth: acme-session-sdk
testing: vitest
infra: wrangler
# conventions.yaml (excerpt)
naming: { tables: snake_plural, columns: snake, files: kebab }
layout: { routes: "src/routes/{module}.ts", schema: "src/db/schema/{entity}.ts" }
errors: { envelope: "{ error: { code, message, details } }", validationStatus: 422 }
ids: uuidv7
pagination: cursor
# type-map.yaml (excerpt)
string:   { drizzle: "text", zod: "z.string()" }
int:      { drizzle: "integer", zod: "z.number().int()" }
datetime: { drizzle: "integer({ mode: 'timestamp' })", zod: "z.coerce.date()" }
uuid:     { drizzle: "text", zod: "z.string().uuid()" }

4.3 The recipe contract

Modelled on OpenRewrite's ScanningRecipe: scan → generate → edit.

export interface TargetRecipe<Acc = unknown> {
  name: string;
  description: string;
  /** Precondition: does this recipe apply to this Spec node under this profile? */
  appliesTo(node: SpecNode, profile: StackProfile): boolean;
  /** Optional cross-node pass (e.g. collect all entities for a schema barrel). */
  scan?(spec: Spec, acc: Acc): void;
  /** Pure: Spec node + kit → files, source→target map entries, holes. */
  generate(node: SpecNode, kit: ResolvedKit, acc: Acc): {
    files: GeneratedFile[];
    mappings: MapEntry[];     // same shape as rebuild-map/*.json mappings
    holes: Hole[];
  };
  /** Optional: AST edits to shared files (router registry, DI, migrations index). */
  edit?(project: TargetProject, acc: Acc): void;
}

Rules:

  • Pure and deterministic. The same Spec plus the same Kit version always produce byte-identical output.
  • Idempotent. Re-running gives no diff. Generator-owned regions are rewritten, and filled holes are never clobbered.
  • AST edits only for shared files (ts-morph for TS targets, tree-sitter otherwise). No regex, per repo convention.
  • Equivalent by construction. The output must verify as equivalent in graph/rebuild-verification.ts: same method plus normalised path, and the same field names and (now) types.
  • Golden-tested. fixtures/spec.json → expected/** is checked by unwind play kit test, and every recipe is run twice to assert idempotence. This is our version of RewriteTest.

Recipes are TS modules plus template files. Recipes for non-TS targets still emit text and are finished by a target-language formatter. Simple recipes can be declarative (recipe.yaml: template plus appliesTo selector).

4.4 Blueprints

A blueprint is a declarative composition of recipes, the analogue of a declarative recipe list. Each one covers a whole service, module or application skeleton.

# blueprints/crud-service.yaml
name: crud-service
description: REST CRUD module per entity, with validation, repository and tests
recipes:
  - scaffold                       # package.json, tsconfig, app entry, wrangler, vitest (once)
  - entity-drizzle-table:   { for: entity }
  - entity-zod-schema:      { for: entity }
  - entity-repository:      { for: entity, options: { softDelete: true } }
  - endpoint-hono-route:    { for: endpoint }
  - route-registry                 # edit phase: mount routes in src/app.ts
  - endpoint-parity-test:   { for: scenario }   # native tests from Spec scenarios (§6.6)

During pl-plan, each Spec slice is assigned a blueprint: crud-service for the orders module, event-consumer for webhooks, and so on.

4.5 Holes: the deterministic/LLM boundary

A recipe emits a hole wherever the Spec does not determine the code. Typical cases are business-rule bodies, non-trivial mappings and bespoke validation.

// src/routes/orders.ts (generated by endpoint-hono-route@1.4.0; do not edit outside holes)
import { Hono } from "hono";
import { zValidator } from "@hono/zod-validator";
import { CreateOrder } from "../schemas/order";
import { ordersRepo } from "../db/repos/orders";

export const orders = new Hono();

// @unwind-id endpoint:src/routes/orders.ts:POST /api/orders
orders.post("/api/orders", zValidator("json", CreateOrder), async (c) => {
  const input = c.req.valid("json");
  // @unwind-hole id=operation:src/services/orders.ts:applyDiscount kind=operation doc=layers/service/orders.md#applydiscount
  throw new Error("unwind-hole: applyDiscount not implemented");
  // @unwind-hole-end
  const order = await ordersRepo.create(c.env.DB, input);
  return c.json(order, 201);
});
  • The LLM builder fills only holes. It works from the linked doc and Spec node, and keeps the @unwind-id provenance marker.
  • Verification. A Spec node whose target still contains an unfilled hole is claimed, not present, so it does not count toward completeness.
  • Regeneration is safe. The generator rewrites everything outside @unwind-hole … @unwind-hole-end, and preserves the hole's contents once filled.

4.6 The Play build loop

Play build loop: generate → holes → fill → verifyPlay build loop: generate → holes → fill → verifydeterministic (engine / CLI)LLM agentartifacthuman / externalSpecMUST / SHOULD nodesTarget Kitpinned kit@versionGenerate (blueprint)recipes: scan → generate → editTarget code with holesstructure correct by constructionLLM fills holesbusiness logic only · keeps @unwind-idVerifystructure + types · unfilled hole = claimedMUST completeness ≥ target ?yes → done · no → looploopsrc/routes/orders.ts// generated: endpoint→hono-route (hono-drizzle-zod@1.0.0)app.post('/orders', zValidator('json', OrderCreate), async (c) => { const input = c.req.valid('json') // @unwind-hole id=operation:src/orders/service.ts:placeOrder // kind=operation doc=layers/service/orders.md#placeorder throw new NotImplemented('placeOrder') // @unwind-end })method · path · schema exact → verifies as equivalentonly this region is written by the LLMRegeneration rewrites generator-owned code but preservesfilled holes, so the loop is safe to re-run.
Play build loop: generate → holes → fill → verify
  1. unwind play generate --slice database --plan: a dry run that lists the files, map entries and holes.
  2. unwind play generate --slice database writes:
    • the target files;
    • docs/unwind/.cache/rebuild-map/<slice>.generated.json, which merge-rebuild-map folds in unchanged;
    • holes.json.
  3. pl-build-layer is dispatched with holes plus the unmapped [MUST] nodes from rebuild-graph.json. This also fixes today's mismatch, where the untagged seed file is pasted in.
  4. unwind play merge, then verify, then parity. The loop continues until completeness % and parity % reach their targets, or until two dry rounds pass (the existing LoopState).

4.7 Kit mining (primary authoring path)

Kit mining: from a reference app to a versioned recipe bookKit mining: from a reference app to a versioned recipe bookdeterministic (engine / CLI)LLM agenthuman / externalReference appclient golden-path serviceRewindSemantic Model + Specof the reference apppl-kit mineprofile · conventionstype-map (from facts)Exemplar selectionbest entity / route / testper blueprint slotVersioned Kitgit tag kit@1.2.0pinned per rebuildGolden testsspec fragment → expectedidempotence runRegenerate gaterecipe(fragment) ≡ exemplar ?via the verifierParameterizeexemplar → template+ recipe.tsfail → re-parameterizepassSecondary authoring pathshand-authored recipes (TS + templates) · in-flight promotion: the LLM builds the first instanceduring a rebuild, which is promoted to a recipe that generates the remaining N
Mining a kit from a reference app

Clients rarely want "a generic Hono app". They want their service template. So we mine it:

  1. Rewind the reference app. Run the normal scan, Semantic Model and Spec on the client's golden-path service.
  2. Infer the deterministic parts with unwind play kit mine --from <ref-repo>:
    • profile.yaml, from dependency manifests and detected frameworks;
    • conventions.yaml, from naming and layout statistics;
    • type-map.yaml, from observed Spec type → target type pairs.
  3. Pick an exemplar per recipe slot that the chosen blueprint needs (entity, route, repository, test, …). The selection is ranked by Spec completeness and conformity with conventions.
  4. Parameterise. The LLM turns each exemplar into a template plus recipe.ts. The exemplar's own Spec fragment becomes the golden fixture.
  5. Regenerate gate. A recipe is accepted only if it regenerates its exemplar from that Spec fragment, structurally equivalent per the verifier (and textually close after formatting). Golden tests then lock it in.
  6. Version. Tag the kit (v1.0.0). Each rebuild pins kit@version in rebuild-state.json.

Secondary paths:

  • Hand-authored recipes, written by platform engineers like any other TS module.
  • In-flight promotion. During a rebuild, the LLM hand-builds the first instance of a pattern. Play offers to promote it into a recipe (through the same regenerate gate), and the recipe then generates the remaining N.

4.8 Starter kits

Play ships a starter kit as the reference implementation of the format. The first is hono-drizzle-zod (TypeScript, Workers/Node; recipes: scaffold, entity → Drizzle table, entity → Zod schema, entity → repository, endpoint → Hono route, route registry, scenario → vitest test). Client kits usually extends a starter and override conventions and templates. Later starters are listed in 08 (Spring/JPA, FastAPI/SQLAlchemy).

4.9 What kits do not do

  • They do not translate code. Kits generate from the Spec, never from the source code. That is the whole point of the Spec boundary.
  • They do not guess. An unknown type or rule becomes a hole and is reported, never invented.
  • They do not replace review. Generated slices arrive as small, mapped diffs in the target repo, and the holes are listed for reviewers.

05 · Semantic Model and store

In short: Borrow the semantics of the LST (types and symbol binding), not its losslessness, because Unwind never edits the source. Build the model in tiers (compiler-accurate T2, index-based T1, tree-sitter T0), each falling back gracefully to the one below. Store it as additive manifest facts plus fact tables, and add a local index later for multi-repo queries.

Semantic Model tiers (LST-inspired, graceful fallback)Semantic Model tiers (LST-inspired, graceful fallback)deterministic (engine / CLI)artifactT2 · compiler-accurate (where it runs in Node)TS compiler API: ts.createProgram + TypeCheckerpyright (npm) for Python, optionaltypes · symbol binding · calls · handler bindingT1 · index-based (optional external tools)SCIP indexers: scip-java · scip-dotnet · scip-pythonexternal binaries, Apache-2.0definitions · references · signaturesT0 · tree-sitter (always available)today's tier: ts/js · python · rust · java · c#+ syntactic types · DDL column types+ route-prefix composition from the ASTnot available → fall backnot available → fall backSemantic Modelscan-manifest.json + additive typed fieldsNew facts (all additive)field · param · return typesinheritance (extends / implements)decorators / annotations as datacalls · reads · writes · derives_from edgesendpoint → handler binding + full pathcolumn types · keys · FKs · relationsconfig / env surfacemessaging producers / consumersclass-qualified, overload-safe idsbody-aware fingerprints (incremental)tier provenance on every factNot copied from the LST: lossless formatting. Unwind never edits the source, so whitespace fidelity buys nothing.
Semantic Model tiers with fallback

5.1 What we have today

The analysis is in packages/core/src (see structure/, imports/, layers/, graph/). Its limits:

AreaTodayGap
SymbolsFunctions, classes, method/property names, param names (TS/JS, Python, Rust, Java, C#)No types, no return types, no generics, no inheritance, no interfaces/type aliases/enums (TS)
BindingFile→file importMap only (relative TS/JS, Java FQN from the path, Python dotted)No symbol→symbol binding; Rust/C#/TS path aliases unresolved
CallsNoneNo call graph, no endpoint→handler, no function→table reads/writes. derives_from is declared but never emitted
Data modelTable/entity field names (Drizzle and JPA via AST, others via regex)No column types, keys, FKs, relations or indexes; EF Core has no fields
EndpointsMethod + path (Express-style regex; Spring/Nest/FastAPI/ASP.NET via AST)Class/router prefixes not composed; no request/response shapes, auth or status codes
LifetimeThe tree is deleted after per-file extraction (structure/tree-sitter-plugin.ts)Nothing survives for later queries
Idskind:path:nameOverloads and same-named methods collide

These are exactly the facts a rebuild needs to be typed and checkable, and exactly what the verifier lacks (graph/rebuild-verification.ts notes that field types are not in the manifest).

5.2 Tiers

TierHowLanguages (initially)Gives
T2 compiler-accurateThe language's own type checker, in-process in NodeTS/JS via the TypeScript compiler API (ts.createProgram + getTypeChecker(), the same route OpenRewrite takes for JS/TS); Python via pyright (an npm package) laterResolved types, symbol binding, references, calls, inheritance
T1 index-basedOptional external SCIP indexers (Apache-2.0: scip-typescript, scip-python, scip-java, scip-dotnet)Java, C#, Python, othersDefinitions, references, signatures
T0 syntacticToday's tree-sitter extractors, extendedAll six grammarsSyntactic types (annotations, DDL column types), decorators as data, composed route prefixes

Rules:

  • Fall back, and report. T2 runs only when the prerequisites exist (e.g. a tsconfig.json and a resolvable typescript). Otherwise the extractor drops to T1 or T0 and records the tier reached per file: semanticTier: 0|1|2. Unlike OpenRewrite's silent Unknown, a tier shortfall is visible in coverage reports and in the dashboard.
  • No hand-rolled regex for new facts. Tree-sitter, compiler APIs or real parsers only. SQL uses a real SQL parser, and ORM schemas use the ORM's own snapshot or schema (Drizzle meta/*_snapshot.json, schema.prisma).
  • Facts, not trees. We extract and persist facts. We never hold or serialise whole trees, which avoids the LST memory ceiling.

5.3 New facts (all additive)

Additions to manifest/manifest-schema.ts are optional fields only. FileSymbols is never reshaped.

FactWhere it lands
Field / param / return typesSymbolDefinition.fieldTypes?, SymbolFunction.paramTypes?, returnType?
Inheritance, interfaces, type aliases, enumsSymbolClass.extends?, implements?; new optional types?[] on FileSymbols
Decorators / annotations as datadecorators?: {name, args}[] on symbols
Symbol references`symbolRefs?: {from, to, kind: call
Endpoint handler + full pathSymbolEndpoint.handler?, fullPath?
Column types, keys, FKs, relations, indexesSymbolDefinition.columns?: {name, type, nullable, pk, fk?, default?}[]
Config/env surfaceNew fact table config
Messaging producers/consumersNew fact table messaging (finally setting hasMessaging in the layer map)
Tier and provenancesemanticTier? per file; provenance on facts

The graph gains real edges. build-graph.ts emits calls, reads and writes, plus derives_from (already declared in rebuild-graph-schema.ts). This enables:

  • endpoint → service → repository → table slicing;
  • impact analysis for uw-refresh;
  • better ordering for Play.

Ids gain class-qualified and arity-qualified forms (method:path:Class.name(n)). The old forms are kept as aliases, so existing anchor-id docs keep matching.

5.4 Detector recipes and fact tables

layers/contract-detectors.ts (about 1,160 lines, mixed AST and regex) becomes a registry of small detectors, the analogue of OpenRewrite's search recipes and data tables:

export interface Detector<Row> {
  name: string;                       // "drizzle-tables", "spring-endpoints", "prisma-models"
  appliesTo(file: ManifestFile): boolean;
  detect(ctx: { tree?: Tree; source: string; model: SemanticModelView }): Row[];
  table: FactTableName;               // "entities" | "endpoints" | "config" | "messaging" | ...
}
  • One detector per framework or ORM. Each has fixtures (input file → expected rows), like the existing layers/*.test.ts.
  • Each fact table is typed and becomes a queryable rows file (.cache/facts/<table>.json).
  • Community-extensible: adding a framework means adding a detector and its fixtures, not editing a monolith.
  • The regex fallback stays only for files with no grammar or parser (graceful degradation), as today.

5.5 Store

  • Phase 1 (now to the App). Files stay authoritative under docs/unwind/.cache/: scan-manifest.json, facts/*.json, spec.json. They are git-friendly and diffable, as today.
  • Later. A local index in node:sqlite over models, Specs, fact tables and verification results, keyed by repo@commit. It enables:
    • multi-repo portfolio queries ("every endpoint touching orders across 40 services");
    • the Recipe Book (which kits and recipes were used where);
    • the App's views.
  • Rebuildable. The index can always be rebuilt from the files. It is never the only copy.
  • Incremental. Today's fingerprints (fingerprint/fingerprint.ts) drive re-extraction. Once call edges exist, a body change that alters calls, reads or writes is no longer "cosmetic".

5.6 Spec compilation (rw-spec)

unwind rewind spec builds the Spec (02 §2.5) from three sources:

  • the Semantic Model's facts (types, shapes, bindings);
  • rebuild-graph.json (priorities, coverage, grill verdicts);
  • fenced blocks in the tagged layer docs: DDL, JSON Schema and OpenAPI, which analysis-principles.md §2/§8 already asks specialists to write. These are parsed with real parsers.

Precedence: T2 facts first, then T1, then doc blocks, then T0. Disagreements are recorded as conflicts, not silently resolved. Each unresolved type becomes unknown and is flagged; Play will turn it into a hole.

5.7 What we deliberately don't copy

  • Losslessness and format preservation. Unwind never prints the source back.
  • A JVM host, RPC peers and proprietary artifacts. Everything runs in Node with optional external indexers.
  • In-place source edits. Play generates into a new target repo.

06 · Behaviour parity: Spec-derived tests against legacy and rebuild

In short: Structural verification proves the rebuild has the right endpoints and tables. It never proves it behaves the same. We generate stack-neutral scenarios from the Spec and run them twice. Against the legacy app, the results observe and enrich the Spec; against the rebuild, the same scenarios give a behavioural parity verdict. This is characterization (golden-master) testing, driven by the Spec.

Behaviour parity: the same scenarios against legacy and rebuildBehaviour parity: the same scenarios against legacy and rebuilddeterministic (engine / CLI)LLM agentartifactSpecendpoints · entities · rulesScenario generationSpec · MUST rules · grill · legacy testsrecorded traffic (HAR / proxy / UI)Scenariosgiven · when · then (or record) · covers: [ids]Boundary drivers (pluggable; each degrades to "manual / not runnable")HTTP / APIDB state snapshotsBrowser: Playwright + agent recordingDesktop: computer-use (best-effort)CLI / batchMessagingLegacy appsandbox only, never prodGoldensresponses · DB side-effectsrw-observe: enrich the Specprovenance "observed" · disagreements → grill questionsTarget apprebuilt serviceNormalize + mapscrub ids / time · vs goldensParity reportparity % over [MUST] scenarios → loop termination signalKit recipes also emit the scenarios as native target tests (e.g. vitest + supertest), so parity becomes the rebuilt repo's regression suite.Legacy can't run? Scenarios stay as Spec acceptance criteria; goldens come from domain experts via the questions/ flow.
Behaviour parity: scenarios run against legacy and target

6.1 Why

  • rebuild-principles.md §8 already says present ≠ correct. Today's verifier (graph/rebuild-verification.ts) checks method plus path and field names, and every other kind of node can only ever reach present.
  • VerificationDepth already declares run-tests (graph/rebuild-state-schema.ts:31), but nothing implements it.
  • Running tests against the legacy app has a second payoff. It turns inferred behaviour into observed behaviour:
    • real response shapes;
    • validation messages;
    • edge cases;
    • confirmation or refutation of grill hypotheses.

6.2 Scenarios

A scenario is a new Spec node kind. It is stack-neutral and lives in the Spec.

id: scenario:orders:discount-threshold
covers: [endpoint:src/routes/orders.ts:POST /api/orders, operation:src/services/orders.ts:applyDiscount]
priority: MUST
source: generated            # generated | legacy-test | grill | traffic | interview | expert
given:
  db:
    customers: [{ id: c1, tier: wholesale }]
  auth: { as: customer, id: c1 }
when:
  http: { method: POST, path: /api/orders, json: { customerId: c1, items: [{ sku: A, qty: 3, priceCents: 4000 }] } }
then:
  status: 201
  json:
    totalCents: 11400          # 12000 − 5%
    status: pending
  db:
    orders: { count: +1 }
record: false                # true = expected unknown; capture from legacy as golden
  • covers links scenarios to Spec node ids. Behavioural coverage is computed with the same set arithmetic as documentation coverage: [MUST] nodes minus nodes covered by a passing scenario.
  • Sources:
    1. Generated from the Spec: per endpoint and entity (happy path, validation failure, auth failure, not-found, pagination), and per [MUST] operation.
    2. Mined from legacy tests. The uw-analyze-*-tests layers already catalogue them; we translate their intent into scenarios.
    3. Grill findings: each suspected bug or edge case becomes a probe.
    4. Captured traffic: HAR files, proxy logs and recorded UI sessions, scrubbed (§6.4).
    5. Interviews and experts: "users rely on X" (07).

6.3 Boundary drivers

Drivers are pluggable. When a driver cannot run, the scenario is still kept and marked manual / not runnable.

DriverMechanismDifficultyOrder
HTTP / APIfetch-based runner; JSON/body/status/header assertionsLow1st
DB stateSeed fixtures plus before/after table snapshots; assert side effects, not just responsesLow–Med1st
CLI / batch / filesInputs → stdout, exit code and output files vs goldensLow2nd
MessagingPublish/consume against a local broker; assert emitted eventsMed2nd
Web UIPlaywright (Apache-2.0) for replay. Agent-driven browser use (Chrome DevTools MCP / Claude in Chrome) to explore and record, then frozen into a deterministic Playwright scriptMed–High3rd
Desktop / legacy GUIAgent computer-use to explore and record, plus OS accessibility automation (e.g. Windows UI Automation) for replayHigh (best-effort)4th

Exploration by agents is non-deterministic, but replay must be deterministic. Recorded sessions are therefore converted into scripted scenarios before they count toward parity.

6.4 Normalisation and intentional differences

  • Scrubbers for nondeterminism: generated ids, timestamps, ordering of unordered collections, tokens and nonces. Each scrubber is declared per scenario or globally.
  • A mapping layer for intentional differences, reusing the structural verifier's equivalence rules:
    • path-parameter normalisation (normalizeEndpointPath);
    • field-name normalisation (norm);
    • target conventions from the Kit (e.g. error envelope, status 422 vs 400).
  • Verdict-aware. [DON'T] items are excluded. fix-in-rebuild grill verdicts expect the legacy result to differ; the scenario holds the corrected expectation, and legacy is recorded only as a reference.

6.5 Lifecycle

  1. rw-observe (Rewind):

    • generate or collect scenarios;
    • run them against a sandboxed legacy instance;
    • record goldens where record: true;
    • write observations back into the Spec with provenance observed.

    Where an observation disagrees with the docs, a grill question or context gap is raised (07).

  2. pl-parity (Play): run the same scenarios against the target, apply scrubbers and the mapping layer, and write parity-report.json plus parity-gaps.md.

  3. Verification depth behavioural, which implements today's run-tests depth. Parity % over [MUST] scenarios becomes a second termination signal for loop mode, alongside structural completeness %.

6.6 Parity tests live on in the target

Kit recipes (e.g. endpoint-parity-test in 04 §4.4) can emit scenarios as native tests in the target stack, for example vitest plus Hono's test client. The parity suite then stays in the rebuilt repo as its permanent regression suite, owned by the team, and needs no Unwind at runtime.

6.7 Safety and limits

  • Never run against production. rw-observe requires an explicit legacy base URL or connection, plus a sandbox confirmation.
  • Read-only first. Scenarios that change state run only against disposable or seeded environments (a Docker recipe, a snapshot restore).
  • Secrets and personal data. Captured traffic is scrubbed before it is stored. Goldens never contain credentials.
  • When the legacy app can't run at all:
    • scenarios still exist as Spec-level acceptance criteria;
    • expected results come from domain experts via the questionnaire flow (docs/unwind/questions/, see 07);
    • they run against the target only.
  • Not a proof. Parity covers the scenarios that exist. Behavioural coverage (§6.2) makes the uncovered remainder visible rather than implied.

07 · Context gaps: find what the code can't tell us, and ask the people who know

In short: Once a Spec exists, an agent round sweeps it for context gaps: intent, usage, non-functional requirements and tribal knowledge that no amount of code reading can recover. Each gap is routed to who can answer it, packaged into interview briefs for stakeholders, end users, existing developers and ops, and the answers are ingested back into the Spec with provenance. How the interviews are conducted is deliberately left open; an external AI-interview tool plugs in through a file contract.

Context gaps: ask the people who know what the code can't sayContext gaps: ask the people who know what the code can't saydeterministic (engine / CLI)LLM agentartifacthuman / externalSignalsSpec + layer docsgrill findingsparity observationsgit: churn · authors · TODOrw-context-gapsagents classify gapswith evidenceSettled in-runanswerable from codeGap registerwho-can-answer · prioritylinked Spec idsRouted interview briefsbusiness stakeholdersend users (per persona)existing developersops / supportcompliance— each brief: context · goals ·questions + probes · who to askInterview toolyour AI-interview toolor humans: method openResponsesinterviews/responses/*gap id → answer · rolecontext-ingestmap answers → gapsprovenance interview:role:dateSpec, enrichedrationale · retags · newscenarios · conflicts shownContext coverage = share of [MUST] Spec nodes with no open high-priority gap, reported next to doc coverage and parity %.The grill's checkbox questionnaires (questions/) become one audience-specific output of this round: one ingest path, not two.
Context gaps: Spec → gap register → briefs → interviews → enriched Spec

7.1 Why

Each existing mechanism has a blind spot:

  • Coverage proves every item is documented.
  • Parity proves what the system does (06).
  • Grilling challenges whether behaviour should be kept.

None of them capture intent and lived context, which is where rebuilds fail quietly:

  • why a rule exists, and whether the reason still holds;
  • which features are actually used, and which are dead weight;
  • workarounds users rely on, including "bugs" that became features;
  • volumes, SLAs, peaks and retention: non-functional requirements that never appear in code;
  • regulatory and contractual drivers;
  • planned changes the rebuild should anticipate;
  • what developers know is fragile, and what ops does by hand.

7.2 Where it sits

It runs after rw-spec, ideally after rw-grill and a first rw-observe pass, and before pl-plan. It can be re-run whenever the Spec changes; uw-refresh can trigger it for affected slices.

7.3 rw-context-gaps: the agent round

Specialist agents sweep these inputs:

  • the Spec;
  • the layer docs;
  • grill findings;
  • parity observations;
  • git signals: churn, authorship, age, dead code, TODO/FIXME/HACK comments, reverted commits.

Each gap they find is classified:

Gap typeExample signal
Unexplained [MUST] ruleOperation with no rationale; magic threshold 10000
Usage unknownEndpoint with no tests, no observed traffic, no frontend caller
Observed ≠ documentedParity recorded 400 where the docs say 422
Implicit workflowFive endpoints always called in sequence from one screen
Missing NFRNo timeout/retry config; no retention policy for audit_log
Integration ownershipExternal API called with no owner and no contract
Manual opsRunbook-shaped scripts; cron entries outside the repo
Persona / permissionAuth roles referenced in code but not explained anywhere

Each gap records:

  • evidence (quoted code and line, like grill findings);
  • the Spec node ids it would resolve;
  • priority, derived from the [MUST] impact;
  • the who-can-answer route: business stakeholder, end user (by persona), existing developer, ops/support, or compliance.

As in the grill, gaps the code can answer are settled in-run. The rest go into the gap register (docs/unwind/.cache/gaps/register.json).

7.4 Interview briefs: the outbound contract

The briefs are written to docs/unwind/interviews/briefs/<audience>/<capability>.{md,json}: one per audience and business capability, in both a human-readable and a machine-readable form.

# Brief · Stakeholder · Order pricing
**Context.** The current system applies a 5% discount to orders over €100 and never
combines it with coupons. (Diagram: order flow.) We are rebuilding this service.
**Goals.** Confirm which pricing rules must carry over, and why.

## Questions
1. **Why does the 5% wholesale discount exist?** *(gap G-014 · operation:…:applyDiscount)*
   - Intent: is this contractual, promotional, or historical?
   - Probe: is €100 still the right threshold? Who could change it?
2. **Should discounts ever stack with coupons?** *(gap G-015)*
   - Intent: the code forbids it, but support tickets suggest customers expect it.

**Suggested interviewees:** Head of Sales; wholesale account manager.

The JSON form carries the same content plus:

  • gapIds and specNodeIds;
  • the question intent and follow-up probes;
  • suggested interviewees: derived from git authorship for developers and from personas for end users, but never contacting anyone automatically.

The format is designed so an external AI-interview tool can run a rich, adaptive interview from it, while staying simple enough for a human interviewer.

7.5 Responses and ingest: the inbound contract

  • In: docs/unwind/interviews/responses/*, holding transcripts or structured answers in whatever format the interview tool produces.
  • Adapter contract (small and tool-specific), mapping each answer to { gapId, answer, confidence, intervieweeRole, date, quote? }.
  • Ingest (unwind rewind context-ingest): an agent applies the answers. It can:
    • add a rationale to rules;
    • retag priorities, with a mandatory rationale, exactly like grill verdicts (drop → [DON'T]);
    • write a fix-in-rebuild correction into the doc body;
    • create parity scenarios from answers like "users rely on X" (06 §6.2);
    • raise follow-up gaps.
  • Provenance. Every change is stamped interview:<role>:<date>.
  • Conflicts are surfaced, never resolved silently. Disagreements between the code, observed behaviour and interviews become new gaps or grill questions.

7.6 One mechanism, not two

Today uw-grill writes checkbox questionnaires for domain experts into docs/unwind/questions/, and grill-answers.mjs ingests the ticks. In the destination design, those questionnaires become one audience-specific output of the context-gap round: a "domain expert, checkbox format" brief. They use the same register, the same ingest path and the same provenance. The grill keeps its job of finding hotspots; the context-gap round owns asking people.

7.7 Context coverage

Context coverage is the share of [MUST] Spec nodes with no open high-priority gap. It is reported alongside:

  • documentation coverage (verify-coverage);
  • structural completeness and behavioural parity (verify / parity).

All three together make readiness for Play measured, not asserted. pl-plan shows them up front and warns before building slices that still have open high-priority gaps.

7.8 Open by design

How interviews are scheduled and conducted is deliberately left out of scope: human, AI-led, async survey, or a workshop. The brief and response formats are the contract. Integrating a specific tool (for example the user's own AI-interview product) is an adapter: either a file hand-off or an API push, to be decided (09).

08 · Roadmap

In short: Nine phases from today's uw-* plugin to the destination. Each phase is shippable on its own, keeps the graceful fallback to today's flow, and has a testable exit criterion, proven on the drizzle-cube example where possible. The order delivers value early: Spec first, then the CLI, the shared server MVP (team use on large codebases, slices as first-class) and the split, then deterministic generation, then richer semantics, mining, behaviour, context and slice convergence, then surfaces.

Roadmap: each phase shippable, graceful fallback keptRoadmap: each phase shippable, graceful fallback keptdeterministic (engine / CLI)new in destinationartifact0Phase 0 · Designbuilds: design doc set + HTML siteexit: reviewed and agreed1Phase 1 · Shared model + Spec v1builds: @unwind/model · rw-spec · typed stack profileexit: drizzle-cube → valid Spec, typed entities + endpoints2Phase 2 · CLI + Server MVP + splitbuilds: unwind CLI · serve (Hono, git, sqlite) · tokens · push/pull · slicesexit: team pushes ≥3 slices to one server; board shows them; 409 path works3Phase 3 · Kits + recipe engine + starter kitbuilds: kit schema · scan/generate/edit runtime · holes · hono-drizzle-zodexit: db + api slices generate, compile, verify equivalent; re-run = no diff4Phase 4 · Semantic Model T0 / T2builds: TS compiler tier · calls/handler/reads/writes · detector registryexit: TS sources fully typed; verifier diffs field types5Phase 5 · Kit miningbuilds: pl-kit mine · exemplar → recipe · regenerate gate · versioningexit: kit mined from repo A rebuilds repo B in house style5bPhase 5b · Behaviour paritybuilds: scenarios · HTTP + DB drivers · rw-observe · pl-parityexit: API scenarios recorded on legacy, replayed on rebuild → parity %5cPhase 5c · Context gapsbuilds: gap register · routed briefs · response adapter · context-ingestexit: briefs for ≥3 audiences; answers ingested with provenance5dPhase 5d · Slice convergence + Play by slicebuilds: merge fragments · uncovered/overlaps/conflicts/seams · strangler orderexit: 100% convergence; one slice rebuilt + verified via seam adapters6Phase 6 · MCP adapterbuilds: unwind mcp, 1:1 with the CLI / APIexit: non-Claude agents get identical results7Phase 7 · App depth + portfoliobuilds: metrics over time · conflict UX · answers in UI · Recipe Bookexit: portfolio across N projects; kits browsable / editable8Phase 8 · Breadthbuilds: SCIP tier (Java / C# / Python) · Spring + FastAPI kits · blueprintsexit: ≥3 typed source languages · ≥3 starter kits
Roadmap timeline

8.1 Phases

PhaseStepsExit criteria
0. DesignThis document set and the HTML site.Reviewed and agreed.
1. Shared model + Spec v1Extract @unwind/model from packages/core (ids, schemas). Define Spec v1 (02 §2.5). Add rw-spec, compiling the Spec from manifest + graph + tagged docs, with fenced DDL/JSON-Schema/OpenAPI parsed by real parsers. Add a typed stack profile written by the plan interview, replacing the API-style regex in skills/scripts/verify-rebuild.mjs.drizzle-cube produces a valid Spec with typed entities and endpoints.
2. CLI + Server MVP + Rewind/Play splitConsolidate skills/scripts/*.mjs into the unwind CLI (@unwind/engine + @unwind/cli, --json everywhere); turn the scripts into shims. Server MVP (03): unwind serve (Hono + node:sqlite + system git, one Docker image); bearer-token auth (login / whoami / logout); projects; push / pull / status with optimistic concurrency and secret scrubbing; slices (propose from the import graph, claim, state machine, per-slice coverage); basic UI (projects, slice board, slice detail with DocsViewer, graph coloured by slice, activity, tokens). Two plugins in one repo; rw-* / pl-* skills; uw-* aliases; manifests and marketplace updated. Play reads the Spec, plus docs for semantics.Both plugins install independently, and the pipeline passes on drizzle-cube offline. Two users push the drizzle-cube analysis as ≥ 3 slices to one Docker server and see it on the slice board; a concurrent push to a different slice does not conflict, and a push touching the same path returns 409 and succeeds after pull.
3. Kit format + recipe engine + starter kitKit schema and loader with extends; recipe runtime (scan / generate / edit, idempotent apply, hole protection); blueprint composer; golden-fixture runner (kit test); hono-drizzle-zod starter kit; pl-build runs generate → holes → LLM → verify; the verifier counts holes and checks types; the scaffold recipe sets config.scaffolded.drizzle-cube's database and API slices generate, compile and verify equivalent. Re-runs give no diff.
4. Semantic Model T0/T2TS compiler-API tier; tree-sitter type and route-prefix extraction; calls/reads/writes/derives_from edges; handler binding; per-file semanticTier; detector-recipe registry refactor of contract-detectors.ts.Spec entities and endpoints are fully typed for TS sources. The verifier diffs field types.
5. Kit miningplay kit mine (profile, conventions, type map); exemplar selection; LLM parameterisation behind the regenerate gate; kit versioning and pinning; in-flight promotion.A kit mined from one reference repo rebuilds another repo's slice in that house style.
5b. Behaviour parityScenario schema in the Spec; generation from Spec + legacy tests + grill findings; HTTP + DB-state drivers first; scrubbers and mapping layer; rw-observe (record goldens, enrich the Spec with provenance); pl-parity and the behavioural verification depth; kit recipes emit native target tests. Then browser (Playwright plus agent recording), CLI/batch and messaging drivers; desktop is best-effort.drizzle-cube's API scenarios are recorded against the legacy app and replayed against the rebuild, with a parity % over [MUST] scenarios.
5d. Slice convergence + Play by sliceSlice Spec fragments merged into spec/project.spec.json on push; uncovered / overlaps / conflicts / seams report and resolve UI; convergence % metric; seams as interface contracts plus parity scenarios; Play slice states (planned → generating → filling → verified → cut-over); seam graph drives the strangler-style build order; legacy adapters for unsatisfied seams.drizzle-cube reaches 100% convergence across its slices; one slice is rebuilt and verified on its own, with its seams to unrebuilt slices satisfied by adapters.
5c. Context gapsGap taxonomy and register schema; rw-context-gaps agent round; audience routing; brief format (md + json); response adapter contract plus context-ingest; provenance on Spec nodes; grill questionnaires folded in; context-coverage metric. The external interview tool is an adapter, not core.drizzle-cube produces routed briefs for ≥ 3 audiences, and a sample response set ingests back into the Spec with provenance.
6. MCP adapterunwind mcp: a stdio MCP server whose tools map 1:1 onto CLI commands and server API routes, with no separate logic. The skills keep using the CLI.Non-Claude-Code agents can drive Rewind/Play with results identical to the CLI.
7. App depth + portfolioServer UI grows: metrics over time, convergence and conflict resolution UX, questions and interview briefs answered in place, FTS search, multi-project portfolio, Recipe Book and Kit browser, kit editor, optional mirror push of artifacts to GitHub/GitLab.A portfolio view across N projects; kits browsable and editable in the App; questions answered in the UI land as commits.
8. BreadthSCIP tier for Java/C#/Python; starter kits for Spring/JPA and FastAPI/SQLAlchemy; blueprints for event consumers, scheduled jobs and BFFs; more parity drivers.≥ 3 source languages typed, and ≥ 3 starter kits.

8.2 Dependencies

0 ─► 1 ─► 2 ─► 3 ─► 5 (mining needs the recipe engine)
          │    ├──► 5b (parity uses the Spec; native tests use kits)
          │    └──► 5d (convergence needs server slices (2) + Spec; Play-by-slice needs 3)
          ├──► 4 (can start after 1; strengthens 3's verification and seam detection)
          └──► 5c (needs the Spec; better after 5b observations)
2 ─► 6 ─► 7 ─► 8

Phases 4, 5b, 5c and 5d can run in parallel with 3 once the Spec exists. Phase 8 is open-ended.

8.3 Cross-cutting rules for every phase

  • Graceful fallback. If the new path is unavailable (no kit match, no TS compiler, no runnable legacy app, no server), the previous behaviour runs and the skill says so.
  • Additive schemas. Optional fields only. The schemaVersion bumps and migrations are documented in @unwind/model.
  • AST over regex for every new detector and every target edit.
  • Tests next to the source (node --test, as in packages/core/src/**/*.test.ts). Every recipe and detector ships with fixtures.
  • Docs move with the code. CLAUDE.md, the README and the principle files (skills/analysis-principles.md, skills/rebuild-principles.md) are updated in the same phase. In particular, rebuild-principles.md gains sections on holes, the Spec and kits.

8.4 Files most affected (for orientation)

TodayBecomes
packages/core/src/manifest/{manifest-schema.ts, candidates.ts}@unwind/model
packages/core/src/layers/contract-detectors.tsDetector-recipe registry
packages/core/src/structure/tree-sitter-plugin.tsT0 tier plus the T2 hook
packages/core/src/graph/{build-graph.ts, rebuild-verification.ts, rebuild-state-schema.ts}New edges; hole, type and behavioural verdicts; kit pin; scaffolded used
skills/uw-plan, skills/uw-build, skills/uw-build-layerpl-plan, pl-build, pl-build-layer (Kit-aware)
skills/scripts/*.mjsunwind CLI subcommands (with shims)
skills/uw-grill questionnairesOne output of rw-context-gaps
packages/dashboard@unwind/app: the server UI (served by unwind serve)
(new)@unwind/server: Hono API, git + node:sqlite storage, auth, slices, convergence

09 · Open questions

In short: These decisions are deliberately deferred. Each has a recommendation, so we can move forward by default and revisit when the evidence arrives. Comment on them by number.

9.1 Naming and prefixes

  • Question: Are "Rewind" (understand) and "Play" (rebuild) the plugin names, with Unwind as the umbrella brand? Are the skill prefixes rw- and pl-?
  • Recommendation: Yes. Keep the uw-* names as deprecated aliases for one release.

9.2 Recipe language

  • Question: Should recipes be TS modules plus templates, fully declarative YAML, or both?
  • Recommendation: TS modules plus template files as the main form, because the engine and the first starter target are TS. Offer declarative recipe.yaml for simple template-only recipes. Recipes for non-TS targets emit text and are finished by a target-language formatter.

9.3 Kit licensing and sharing

  • Question: How are kits licensed and shared?
  • Recommendation: Client kits are private repos owned by the client. Starter kits use the same licence as Unwind (MIT). A future public "Recipe Book" index lists starter and community kits only.

9.4 Spec and external standards

  • Question: Should the Spec define its own formats, or embed existing standards?
  • Recommendation: Embed. Use JSON Schema for shapes and OpenAPI-style operations for endpoints, wrapped in Spec nodes that carry id, priority and provenance. This gives import/export to existing tooling for free.

9.5 Store timing

  • Question: When does the node:sqlite index arrive?
  • Recommendation: Day 0, inside the server (phase 2, 03 §3.4). Locally, the CLI stays file-only. On the server, git holds the artifacts and SQLite holds auth/ops state plus a rebuildable index.

9.6 Interview tool integration

  • Question: Do we push briefs to the external AI-interview tool via its API, or hand off files?
  • Recommendation: Start with file hand-off (briefs out, responses in), because it needs no credentials and is easy to audit. Add an API adapter once the tool's interface is stable. The brief and response formats (07) are the contract either way.

9.7 Parity environments

  • Question: Who provides a runnable legacy sandbox?
  • Recommendation: Support all three options: a client-provided environment, a Docker recipe produced from the infrastructure layer, and recorded traffic as the zero-setup fallback. Never production.

9.8 Scenario format

  • Question: Should scenarios use a custom YAML schema, or reuse an existing one (e.g. OpenAPI Arazzo workflows for API flows)?
  • Recommendation: Custom and minimal (given / when / then / covers), with import/export adapters for Arazzo and HAR. Arazzo is API-only, and scenarios must also cover UI, desktop, CLI and messaging.

9.9 Behavioural verification depth

  • Question: Should run-tests (graph/rebuild-state-schema.ts:31) be renamed behavioural, and how should it combine with structural completeness?
  • Recommendation: Keep run-tests as the stored enum value for compatibility and label it "behavioural" in the UI. Loop termination requires both structural completeness % and [MUST] parity % to reach their targets.

9.10 Id migration

  • Question: How do we introduce class-qualified, overload-safe ids without breaking existing anchor-id docs?
  • Recommendation: Emit the new ids alongside the old ones as aliases. Coverage matches on either. A one-time rw-complete pass can rewrite anchors.

9.11 Optional external tools

  • Question: Should the plugin bundle SCIP indexers, pyright or Playwright, or discover them?
  • Recommendation: Discover, never bundle. Each tool is optional, its tier or driver is reported when it is missing, and a helper suggests the install command.

9.12 Generated code ownership

  • Question: After hand-off, who owns the generated files, and can a team "eject" from regeneration?
  • Recommendation: Yes. unwind play eject <slice> strips the generator markers and keeps the @unwind-id provenance comments, so verification still works. From then on the slice is hand-maintained.

9.13 Slice auto-proposal algorithm

  • Question: How should unwind slices propose cut a large codebase into slices?

  • Recommendation: Deterministic and explainable first:

    • community detection (e.g. Louvain/Leiden) over the import graph, later the call graph;
    • seeded by top-level directories and layers;
    • balanced by candidate count (target ~200–800 candidates per slice);
    • each proposal shows its cohesion and seam count.

    An LLM pass may then suggest business-capability names. Humans always confirm.

9.14 Conflict resolution UX

  • Question: When two slice fragments disagree on the same id (priority or content), how is it resolved?
  • Recommendation:
    • The server never auto-picks. The convergence view shows both versions side by side with their provenance, and one owner resolves.
    • The resolution is a normal commit with a rationale, like grill verdicts.
    • Push still succeeds, but the slice can't reach accepted while it has open conflicts.

9.15 Git layout: per project, or a branch per slice

  • Question: Should each slice work on its own git branch in the project repo, or should everything live on main with per-slice folders?
  • Recommendation: One repo per project and one main branch, with per-slice folders (03 §3.4). Path-level optimistic concurrency makes slice conflicts rare, convergence always reads one tree, and history stays linear. Revisit "review branches" (push to a branch, approve into main) if teams want PR-style review of analysis.

9.16 Cloudflare / hosted variant

  • Question: Should the server also run Cloudflare-native (Workers + Durable Objects SQLite + R2) or as hosted SaaS?
  • Recommendation: Later, behind a storage interface. Day 0 is self-hosted Node + Docker only, because clients want artifacts inside their network and real git is simplest there. Keep the git and SQLite access behind a small ArtifactStore / StateStore interface so a Cloudflare adapter is possible without touching the API.

9.17 Server auth beyond tokens

  • Question: When do we need SSO/OIDC and finer-grained roles?
  • Recommendation: Not on day 0. Use bearer tokens with read/write/admin scopes plus project allow-lists, and put a reverse proxy (or Cloudflare Access) in front for SSO. Add native OIDC only when a client requires it.