Generation Pipeline
Every stage of every build type, in the order it runs — generated from the pipeline the builder itself executes.
Reading this page
Stages marked Code are deterministic and free; AI call stages call a model; Images generate photographs; Browser stages drive a real headless browser. A condition after a stage name says when it runs — Lite, for example, skips repair, copy rewrites, photographs and visual review.
Websites · 17 stages
Site Spec engine.
- 1FeatureCode
Reads the request to work out which of the nine build types it asks for, and how decisively.
- Scores phrases per feature, weighting multi-word phrases above bare words — 'pitch deck' outranks an incidental 'site'
- Matches on word boundaries, so 'app' does not fire inside 'happen' nor 'deck' inside 'decking'
- Reports the words that decided it, so a surprising answer can be argued with
- Reports confidence and a runner-up instead of resolving a genuinely ambiguous request
- Never overrides the build type the user picked — it names a disagreement rather than acting on one
- 2Blueprint / plannerCode
Detects the industry from the brief and produces an ordered page outline the generator must follow.
- Ranks industry cues by specificity, so 'a boutique physiotherapy clinic' resolves to clinic and not ecommerce
- Orders the page along a spine — open, make the case, show proof, answer the objection, ask
- Is computed ONCE from the user's brief and read by both the generator and the audit, so a page cannot be written to one playbook and marked against another
- 3Page plannerCode
Decides how many pages the site has and what each one is for.
- Gives a page only to content someone arrives at directly — services, menu, work, team, pricing
- Leaves supporting material (stats, testimonials, FAQ) on the page it supports
- Caps at five pages, because a thin page damages a site more than its absence
- Honours a brief that asked for a single landing page
- 4KnowledgeCode· a knowledge base is configured for this deployment
Asks the design knowledge base what it knows about a brief like this one — palette, type pairing, section order, motion and legibility rules drawn from 14,701 indexed documents.
- Fills a quota per category, so the answer is not five paraphrases of whichever row matched the words best
- Runs once per build and goes into the prompt, rather than being a tool the model must remember to call
- Returns nothing when the cluster is unreachable, so a build is degraded by its absence and never blocked
- Reports the documents used and their licences, so what informed a build is auditable afterwards
- 5Industry intelligenceCode
15 industry playbooks: required, expected and forbidden sections, the primary action, the angle.
- Names the sections a real business in that field cannot ship without
- Names the ones that are wrong for it, so a restaurant does not get a pricing table
- Supplies the one action every call-to-action on the page should aim at
- 6GenerateAI call
One forced tool call returns the whole validated Site Spec.
- Forced tool call, so the model cannot answer in prose
- Retries once on a validation failure, logging the field SHAPE and never the content
- Accepts reference images as visual direction
- 7AuditCode
Scores structure, conversion, copy, SEO, accessibility and mobile, and emits a repair instruction per issue.
- Deterministic — the same spec scores the same twice, which a vision model cannot promise
- Weights structure and conversion above polish, because they decide whether the page works at all
- Every issue carries the sentence that would fix it, so a failing audit feeds straight into a targeted edit
- Deliberately has no 'Design %' — there is no defensible way to compute one, and a made-up number devalues the real ones
- 8RepairAI call· the audit scores below 85 and found something repairable
One targeted edit against the audit's findings. Discarded unless it scores better.
- 9CopywriterAI call· Power mode, or the audit's copy dimension is below 85
A second pass that changes only the words. Discarded unless the page's shape and images survive.
- Sees a finished page and changes nothing but the language
- Shape is enforced in CODE, not asked for in the prompt — same sections, same order, same image URLs, same theme, or the rewrite is thrown away
- A rewrite that changed nothing is discarded too
- 10Copy re-checkCode
Re-scores the rewritten copy and keeps it only if it preserved the structure and scored no worse.
- Rejects a rewrite that changed the page structure
- Rejects a rewrite that changed nothing
- Rejects a rewrite that scored lower than the copy it replaced
- Reports which of those happened, so a discarded rewrite is visible rather than silent
- 11ImagesImages
Real photography for up to 8 slots in priority order, converted to WebP.
- Eight slots maximum, filled in priority order so the hero is never the one that gets dropped
- Writes alt text describing the photograph, not the business
- 12Design systemCode
12 palette presets; the button-label and brand-text colours are computed from the brand colour's luminance.
- Twelve palette presets, each with its own card, border, shadow and button tokens
- Derives the button label colour from the brand's WCAG luminance — a pale brand used to get a white label at 1.5:1
- Derives a separate readable colour for the brand used AS TEXT, keeping the hue while buying the contrast
- 13Layout rhythmCode
Bands the supporting sections and varies a repeated section variant.
- Bands the supporting sections so the page has boundaries the eye can find
- Semantic rather than positional, so the rule survives the mobile reorder
- Picks a grid's column count from its card count, so six cards do not render four-then-two
- 14Mobile planCode
Reading order, swipe rails and the sticky action bar for narrow screens.
- Promotes the ask above the long tail of the page, but never above the first substantive section
- Turns a collection into a swipe rail only when there are enough cards for the peek to mean anything
- Pins the primary action and a tap-to-call button to the bottom of the screen
- Changes only the PAINTED order — the DOM order a screen reader and a crawler follow is untouched
- 15SEOCode
Meta and Open Graph tags, JSON-LD, FAQ schema, LCP preload, per-site robots.txt and sitemap.xml.
- Meta description, Open Graph, Twitter card, Organization and FAQ JSON-LD
- Serves a robots.txt and sitemap.xml scoped to each published site — customer domains used to serve the app’s own
- Preloads the LCP image with the SAME srcset and sizes the <img> uses, so the preload is not a second download
- 16Visual reviewBrowser· a browser binary is present; skipped silently otherwise
Photographs the page at two widths and has a vision model review what it can see.
- Photographs the page at a desktop width and a real phone, with touch emulation on
- Freezes scroll animations first, or it reports washed-out text no visitor will ever see
- Caps the capture at the vision API's own 1568px limit, so body copy stays legible to the reviewer
- Told which page it is looking at — reading the wrong page's section list once produced two invented critical findings
- 17Improvement loopBrowser· Power mode, website builds only — at most 3 rounds and 150 seconds
Up to 3 rounds of review, edit and re-review. Every round is discarded unless it survives the audit, keeps every image and scores better.
- Applies its own findings to a COPY, never the live spec
- Stops on a flat round, because a loop that continues walks a good page downhill one improvement at a time
Designs · 7 stages
Site Spec engine.
- 1Artefact planCode
Decides what is being designed and at what distance it is read, which fixes the aspect, the copy budget and the hierarchy.
- Read distance is the constraint: a poster is read from four metres and a story from thirty centimetres, and that decides everything else
- Sets the headline and body word budgets the auditor was previously checking against its own assumptions
- Names the safe area per format, so a feed crop or a story interface does not eat the message
- 2FeatureCode
Reads the request to work out which build type it asks for, and how decisively.
- Reports a disagreement with the picked build type rather than acting on one
- Never overrides the user's explicit choice
- 3KnowledgeCode· a knowledge base is configured for this deployment
Asks the knowledge base what it knows about a brief like this one, filling a quota per category.
- Runs once per build and goes into the prompt, rather than being a tool the model must remember to call
- Records every retrieval against this feature and job, so which knowledge helped is answerable afterwards
- Returns nothing when the cluster is unreachable, so a build is degraded by its absence and never blocked
- 4GenerateAI call
One tool call returning the complete specification, validated against the schema before anything else runs.
- 5AuditCode
Checks each poster reads at a glance on its own canvas, and that its layout carries what it promises.
- Holds the headline to a per-ASPECT budget — a 9:16 story cannot set what a banner can, and the schema's single 160-character cap could not express that
- Catches a statement layout carrying bullets, which is two designs in one record
- Catches a split layout with nothing for its second half, which renders half empty
- Checks a multi-poster set shares one aspect, or it cannot be delivered as a set
- 6RepairAI call· the audit found something worth an edit, and the tier allows a repair pass
Applies one targeted edit carrying every repairable design finding, and keeps it only if it scores better.
- At most ONE pass — a build the user is waiting on cannot loop toward a target
- A repair that scores no better is discarded rather than shipped
- 7ImagesImages· the tier allows image generation and the spec has empty slots
Images
Slides · 7 stages
Site Spec engine.
- 1Deck planCode
Decides which kind of deck this is, then fixes its slide order, its length and its bullet budget before anything is written.
- Six deck kinds, each with the spine its argument has to arrive in — a pitch that opens with the ask is broken however good the slides are
- Caps bullets per slide, because a slide past that is a document nobody reads
- Names the layouts that are wrong for the kind, so a training deck gets no pricing slide
- 2FeatureCode
Reads the request to work out which build type it asks for, and how decisively.
- Reports a disagreement with the picked build type rather than acting on one
- Never overrides the user's explicit choice
- 3KnowledgeCode· a knowledge base is configured for this deployment
Asks the knowledge base what it knows about a brief like this one, filling a quota per category.
- Runs once per build and goes into the prompt, rather than being a tool the model must remember to call
- Records every retrieval against this feature and job, so which knowledge helped is answerable afterwards
- Returns nothing when the cluster is unreachable, so a build is degraded by its absence and never blocked
- 4GenerateAI call
One tool call returning the complete specification, validated against the schema before anything else runs.
- 5AuditCode
Checks every slide can be read from the back of a room and carries what its layout needs.
- Catches a stat slide with no metrics and a chart slide with no chart — both render as an empty slide while the build reports success
- Holds slides to six bullets and 120 characters each, past which a slide is a document
- Catches the all-content deck the layout enum exists to prevent
- 6RepairAI call· the audit found something worth an edit, and the tier allows a repair pass
Applies one targeted edit carrying every repairable slides finding, and keeps it only if it scores better.
- At most ONE pass — a build the user is waiting on cannot loop toward a target
- A repair that scores no better is discarded rather than shipped
- 7ImagesImages· the tier allows image generation and the spec has empty slots
Images
Documents · 6 stages
Site Spec engine.
- 1Document planCode
Decides which kind of document this is and which sections it cannot ship without — and which would be wrong for it.
- Six kinds: a proposal needs scope and pricing, a manual needs verification and troubleshooting, and neither wants the other's sections
- Sets the heading depth, because past three levels nobody follows the tree
- Names the forbidden sections — the ones a generic model adds anyway and which reveal a template
- 2FeatureCode
Reads the request to work out which build type it asks for, and how decisively.
- Reports a disagreement with the picked build type rather than acting on one
- Never overrides the user's explicit choice
- 3KnowledgeCode· a knowledge base is configured for this deployment
Asks the knowledge base what it knows about a brief like this one, filling a quota per category.
- Runs once per build and goes into the prompt, rather than being a tool the model must remember to call
- Records every retrieval against this feature and job, so which knowledge helped is answerable afterwards
- Returns nothing when the cluster is unreachable, so a build is degraded by its absence and never blocked
- 4GenerateAI call
One tool call returning the complete specification, validated against the schema before anything else runs.
- 5AuditCode
Checks the document's structure, tables and readability before it is rendered.
- Catches a contents page with no headings after it, which renders empty
- Catches a heading followed straight by another heading, which is an empty section
- 6RepairAI call· the audit found something worth an edit, and the tier allows a repair pass
Applies one targeted edit carrying every repairable document finding, and keeps it only if it scores better.
- At most ONE pass — a build the user is waiting on cannot loop toward a target
- A repair that scores no better is discarded rather than shipped
Data Visualizations · 7 stages
Site Spec engine.
- 1FeatureCode
Reads the request to work out which build type it asks for, and how decisively.
- Reports a disagreement with the picked build type rather than acting on one
- Never overrides the user's explicit choice
- 2KnowledgeCode· a knowledge base is configured for this deployment
Asks the knowledge base what it knows about a brief like this one, filling a quota per category.
- Runs once per build and goes into the prompt, rather than being a tool the model must remember to call
- Records every retrieval against this feature and job, so which knowledge helped is answerable afterwards
- Returns nothing when the cluster is unreachable, so a build is degraded by its absence and never blocked
- 3GenerateAI call
One tool call returning the complete specification, validated against the schema before anything else runs.
- 4DatasetCode· the prompt contains a table
Extracts a table pasted into the prompt and reports its shape.
- 5Data qualityCode· a dataset was extracted
Reports issues found in the supplied data rather than silently correcting them.
- Reports problems instead of fixing them, because a silently corrected figure is worse than a flagged one
- Separates serious issues from minor ones, so a build is not failed by a stray blank cell
- 6Chart auditCode· a dataset was extracted
Reconciles every plotted value against the source data, so no chart can invent a number.
- Compares by position rather than label, so a renamed chart is still the same chart
- Reports the charts that did not reconcile rather than quietly redrawing them to match
- 7ImagesImages· the tier allows image generation and the spec has empty slots
Images
Instant web apps · 2 stages
Sandbox engine.
- 1GenerateAI call
One forced tool call returns a self-contained HTML document (inline CSS and JS, no external resources).
- 2SandboxCode
Renders into a sandbox="allow-scripts" iframe with an opaque origin — no cookies, no same-origin, no reach into our APIs.
Cross-platform apps · 7 stages
App Spec engine.
- 1App plannerCode
11 app-type playbooks. Screens are earned from the type and its capabilities, never from a fixed template.
- 2GenerateAI call
One forced tool call returns the whole validated App Spec.
- 3Defect checkCode
Walks the navigation graph for dangling targets, duplicate ids and unreachable screens. Correctness, not quality.
- 4AuditCode
Scores structure, navigation, flow, content and accessibility.
- 5RepairAI call· the audit scores below 85
One targeted edit. Discarded unless it scores better.
- 6EmitCode
Turns the spec into an Expo + React Native + TypeScript project as a file map. Screens are editable TSX; the components are fixed.
- 7EditorAI call· NOT a build stage — this runs when an existing app is edited, on a different request entirely
Applies an instruction to an existing app. Rejected only if it introduces a structural defect nobody asked for.
2D Games · 14 stages
Sandbox engine.
- 1BriefCode
Picks the genre from the words in the prompt and fixes the core loop, the goal and both endings before any inference runs.
- Eight genres, each with a real win and lose condition — an unrecognised prompt gets a playable arcade shape, not an empty plan
- Word-boundary matching, so "collectible" in a shooter prompt does not select the collector genre
- 2ArchitectureCode
Names the states, the entities and the systems the runtime has to own, plus the delta-time rule the loop must obey.
- 3Level designCode
Fixes the world shape, the spawn cadence and the difficulty curve — what rises, by how much, and against what floor.
- A curve with numbers in it, so "it gets faster" cannot be the whole specification
- Authored levels where randomness would make a game uncompletable, spawners where it would not
- 4Gameplay systemsCode
Decides scoring, progression and the feedback that makes both legible during play.
- 5AssetsCode
Plans what is drawn and what moves while it is drawn, and synthesized audio — nothing fetched, since the sandbox blocks it.
- 6PhysicsCode
Chooses the collision model the genre actually needs — boxes, squared-distance circles or grid cells — and names the tuning constants.
- 7GenerateAI call
One tool call, briefed with the plan above, returning a complete self-contained HTML document containing the whole game.
- 8Runtime testCode· a browser can be launched — reported as skipped when it cannot
Opens the game in a real browser, taps and presses it, and watches for two seconds.
- Wraps requestAnimationFrame and addEventListener before the document loads, so it counts what the game DID rather than what its source says
- Catches the throw on line 40 that leaves every keyword the source audit looks for in a script that never finished running
- Nudges with a tap and a keypress first, so a title screen is not mistaken for a dead loop
- Samples the canvas for a flat single colour — the blank rectangle a build otherwise reports as a success
- Answers "unknown" for a WebGL or DOM game it cannot sample, rather than failing one that works
- 9PlaythroughCode· the game labels its state — reported as skipped when it does not
Starts the game, plays it for six seconds, watches a run end without input, then presses Restart and checks it comes back.
- Everything the runtime test checks is also true of a screensaver; this checks the things that make it a game
- Catches a score that never moves under six seconds of continuous input — nothing to achieve
- Catches a run that cannot end: twelve idle seconds where nothing kills you, in the genres where something should
- Catches the opposite too — a death inside 1.2s, which lands before a person could react to it
- Catches a game over with no way back except reloading the page
- Reads the score and phase off the screen, so a counter with nothing rendered behind it cannot satisfy it
- Answers unknown for an unlabelled game — unmeasured is never reported as unwinnable
- 10QA auditCode
Reads the emitted document for the things that would make the game unplayable or dead on arrival.
- Catches external scripts and audio files the sandbox CSP blocks — the game renders blank while the build reports success
- Catches localStorage and cookies, which throw on the sandbox's opaque origin and stop the script
- Catches a game with no requestAnimationFrame loop, and one that listens for no input at all
- Catches a keyboard-only game, which cannot be played on a phone
- Catches movement that is not multiplied by delta time — the frame-rate bug the brief calls out in capitals
- Catches a canvas sized from window.innerWidth instead of its own clientWidth, which the brief warns leaves it blank
- Catches a game with no way to play again after losing
- 11RepairAI call· the runtime test or the QA audit found something worth an edit
One targeted pass carrying every failing runtime check and every repairable audit finding together.
- 12Re-testCode· a repair ran
Audits the repaired document, so the repair is measured rather than assumed.
- 13Final validationCode· a repair ran
Plays the repaired game in a browser, because a repair that satisfies every source rule can still throw.
- 14Keep or discardCode· a repair ran
Scores both versions on the audit and the runtime failures together, and keeps the repair only if it measured better.
- A runtime failure outweighs source findings — swapping a game that runs for one that reads better is the worst trade available
- A repair that measures no better is discarded, and which way it went is reported rather than silent
- A skipped browser contributes nothing to the comparison instead of marking both sides down
3D Games · 12 stages
Game Spec engine.
- 1PlanCode
Reads the prompt and fixes the genre, the goal, the world size and the win condition before any inference runs.
- Picks a genre from the verbs in the prompt rather than asking the model to choose
- Sets a win condition the playtest can later check is reachable
- 2ReferenceCode· the prompt names an existing product
Turns a named product — "like Minecraft" — into mechanics, structure and an originality boundary.
- Detects the reference without an inference call, so the same prompt always resolves the same way
- Carries the list of names and assets that must not appear in the result
- 3ResearchCode· the prompt names a technology, or the feature has a known stack
Looks up the current facts a build depends on — package versions, peer ranges, licences — from an allowlist of public sources.
- Fetches rather than recalls, so a version is what npm says today and not what a model was trained on
- A failed lookup produces nothing at all — it never falls back to a guess
- 4KnowledgeCode· a knowledge base is configured for this deployment
Asks the design knowledge base what it knows about a brief like this one — palette, mood, legibility, and the Three.js guidance that applies to a browser game.
- Fills a quota per category, so the answer is not five paraphrases of one row
- Returns nothing when the cluster is unreachable, so a build is never blocked by it
- 5VarietyCode· the workspace has built games before
Reads what this workspace has already been given, so the next build is not a near-copy of the last.
- Names the previous builds rather than seeding randomness, so the instruction is checkable
- Varies the look and the name, never the genre the user asked for
- 6GenerateAI call
One forced tool call emits the complete GameSpec: world, player, entities, objectives and theme.
- 7LegibilityCode
Brightens a scene that would render as an unreadable black screen, in code rather than by asking again.
- 8AuditCode
Scores the spec for playability, genre fit, world and challenge, and lists what is wrong with it.
- 9RepairAI call· the audit found something worth fixing
Sends the audit's complaints back for one revision, and keeps the result only if it scores better.
- 10PlaytestCode
Walks the objective graph to prove the game can actually be finished, and reports how long it takes.
- Walks the objective graph rather than sampling, so an unreachable goal is proven rather than guessed
- Refuses to publish a game whose goal cannot be reached
- 11PerformanceCode
Counts draw calls, triangles and shadow casters, and says which of them will cost frames.
- 12PublishCode
Emits the runtime and puts it at a URL that runs in any browser.
Animations · 6 stages
Sandbox engine.
- 1BriefCode
Chooses the technique, the loop shape and the timing before any motion is written.
- Seven piece kinds, each with the technique that suits it — canvas for a thousand particles, SVG geometry for a logo, CSS alone for a loader
- Decides whether the piece resolves and holds, repeats without a visible join, or keeps evolving — the one fact that makes a reveal wrong as an ambient field
- Gives real timings and easings rather than "smooth" — 1.6–2.4s for a reveal, 260–420ms for a transition
- 2GenerateAI call
One tool call returning a complete self-contained HTML document that animates on load.
- 3Motion auditCode
Reads the emitted document for the things that would make the piece dead on arrival.
- Catches external libraries the sandbox CSP blocks — the piece renders blank while the build reports success
- Catches localStorage and cookies, which throw on the sandbox's opaque origin and stop the script
- Catches a piece with no motion at all: a still image sold as an animation
- Catches a missing play/pause control, and a piece that ignores prefers-reduced-motion
- 4RepairAI call· the motion audit found something worth an edit
One targeted pass carrying every repairable motion finding.
- 5Re-testCode· a repair ran
Audits the repaired document, so the repair is measured rather than assumed.
- 6Keep or discardCode· a repair ran
Keeps the repair only if it scored better than the original.
- A repair that scores no better is discarded — shipping it would be a regression
- Reports which way it went, so a discarded repair is visible rather than silent
Spreadsheets · 7 stages
Sandbox engine.
- 1GenerateAI call
One forced tool call returns a self-contained HTML document (inline CSS and JS, no external resources).
- 2Plan checksCode· Spreadsheets only.
Derives the checks from the same prompt the build was briefed with, so the test exercises the rules that were actually asked for rather than a generic guess.
- 3Browser testBrowser· Spreadsheets only.
Opens the document in a headless browser and drives it — types into cells and reads the results back, because a sheet with hardcoded totals is pixel-identical to one that recalculates.
- 4RepairAI call· Only when the browser test found a failure.
One more pass, briefed with the exact checks that failed.
- 5Re-testBrowser· Only after a repair.
Runs the same checks against the repair, to find out whether it helped.
- 6Keep or discardCode· Only after a repair.
Keeps the repair only if it fixed more than it broke. A second opinion that scores worse is a regression, and shipping it would be worse than shipping the original.
- 7SandboxCode
Renders into a sandbox="allow-scripts" iframe with an opaque origin — no cookies, no same-origin, no reach into our APIs.