stylus nexus field manual · consolidated 2026

The startup plan for
shipping with agents.

A short guide for anyone starting to build with AI coding agents: the working loop, the four mistakes that cost us the most time, and what to do this week.

think→ plan→ build→ verify→ review→ ship

Why this page

One workflow, not two guides.

This used to be two long guides — one for development, one for testing. Both got read once and bookmarked. This is what's left after cutting everything that wasn't a decision you'd actually make: the loop, the four failures that cost the most time, and the concrete first moves. Then it points you at the repo you fork and the skills you install, so you're building instead of reading.

It is not software to install, a full course, or research. It is one team's experience, kept short on purpose. Some examples use Claude Code file names; the loop and the failures apply to any coding agent.

The loop

Six steps, in order, every time.

01

Think

Diverge before you commit. What are we actually solving, what's the unstated assumption, what could go wrong three different ways? No plan yet — just surfacing the options and edge cases a rushed first instinct would skip. This is the step deadline pressure cuts first, which is exactly the wrong one to cut.

02

Plan

Converge and commit. Turn what thinking surfaced into a specific, written plan — files to touch, sequence, test cases in plain English, a definition of done. Specialist agents, one per area you work in, weigh in from their own lane on that plan, not one generalist agent holding every domain in its head at once. The full-starter template ships backend, docs, test-review, and browser-testing agents plus templates for adding your own.

03

Build

Agents write against a CLAUDE.md that stays small on purpose, domain guidances that load only when relevant, and hooks that inject context automatically — instead of a growing wall of always-loaded instructions.

04

Verify

Does it run? Build, typecheck, test, lint — mechanical pass/fail, no judgment involved. If this fails, review never starts. AI mock fixtures keep the loop fast; real calls are for local spot-checks, never CI.

One check earns its own rule, because we shipped the same bug three times before adding it: any value threaded through two or more hops needs a test at its final destination proving it arrived. The bug class survives ordinary review because every hop is locally correct — each function does receive the value and pass it on — but a destructure omits it, a mapping drops it, or a spread overwrites it. Nothing throws. A test at hop three proves nothing about hop five.

05

Review

Should it run this way? Everything that passed verify but is still wrong — a bug, a security hole, a convention violated, an edge case nobody wrote a test for. Verify is a machine's job; review needs judgment, which is why solo developers get it from a second opinion instead of a teammate.

A second opinion covers code. It does not cover the things only you can judge — whether the draft says what you meant, whether the plan is the one you want built. There the bottleneck is not judgment but the channel: describing a wording change in chat costs more than making it, and “the third paragraph feels off” throws away the phrase that caused the feeling. Mark the artifact up directly instead, and hand back edits and comments anchored to exact quotes.

06

Ship

What made it this far gets measured, not assumed. A prompt or model change that "feels better" is guessed at 40% wrong — ten generations and a pass/fail count catches what manual testing misses.

Lessons, paid for already

Four ways this goes wrong — and the one thing they had in common.

Each of these shipped, was used for weeks, and had to be unwound. Skip straight to the fix.

The mega-agent

One agent definition covering frontend, backend, database, content, security, analytics. Generic output that missed domain nuance — backend patterns applied to frontend code, UI suggestions during infrastructure work. Fix: split by domain. One agent, one lane.

The mega-CLAUDE.md

Every rule, pattern, gotcha, and convention in one file. Felt thorough. Started burning thousands of tokens loading AI-safety docs to edit CSS, and rules buried mid-file stopped getting followed. Fix: a small root file, domain guidances loaded on demand.

Ephemeral knowledge

An agent discovers a workaround, it lives in that conversation, and vanishes when the conversation ends. Same discovery, same 30-minute debugging session, the following week. Fix: persistent per-agent memory, one line per insight, linked for depth.

Gut-feel prompting

Prompt changes judged by whether they "felt better." When actually measured, intuition was wrong about 40% of the time — changes assumed to help sometimes made things worse. Fix: ten generations, structured pass/fail, before trusting a change.

The meta-lesson

All four came from the same instinct: keep things simple by centralizing. One agent, one file, one conversation, one gut feel. The fix was always the same — decompose, specialize, measure. More infrastructure upfront. It scales where the simple version doesn't.

First moves

What to actually do this week.

Before any of this: a validation gate

▸ expand

The most expensive mistake isn't a bad pattern — it's building the wrong thing well. Paste this into your agent before you scaffold anything.

Before we start a new project, run me through a validation gate. Don't
let me skip a step because I'm excited, and push back if I try to write
code before step 3 is done.

1. Ask who the target user is and what's the most annoying part of
   their day related to this idea. Help me write one sentence: "When
   ___ happens, I want to ___, so I can ___."

2. Classify it: painkiller (costs real hours or dollars every week) or
   vitamin (nice to have). Tell me plainly if it's a vitamin -- that
   doesn't clear this gate.

3. Help me design the cheapest possible demand test I can run in 7
   days -- a waitlist, a preorder, or a manual pilot offer -- and define
   what "someone actually commits" looks like for this specific idea.
   A soft "yes that's cool" doesn't count. Money or a hard commitment does.

4. Only once I report back real evidence from step 3, help me scope a
   24-hour MVP: one core loop, end to end. Ship it only if all three are
   true -- a new user gets value within 60 seconds, that value is
   immediate, and I can manually fix any bad output myself.

Adapted from the validation discipline in The Vibe Coding Playbook (Murugappan) — own the pain data, then prove willingness-to-pay before you open the code editor.

Start your CLAUDE.md at 20 lines, not 500

Dev commands and 3-4 gotchas. It grows on its own — every entry earns its place because you hit that exact issue twice.

Keep memory files under 200 lines

They load into context every run. Use topic files for depth, one line per insight in the index, and link out rather than inline.

Start with three skills, not eleven

Brainstorm, TDD, verify. Those three alone change how you work. Adopt the rest — plan, execute, debug, review, ship, measure — once you feel the gap.

Run agents sequentially before parallel

Master the chain — brainstorm, plan, execute, verify, ship — one agent at a time. Parallel dispatch is a coordination problem you don't need on day one.

One MCP server, not seven

Start with context7 or GitHub. Add the next one only when an agent keeps needing the same external system and doesn't have it.

Test at more than one viewport

A layout that holds at 1920px can break at 768px. Separate desktop and mobile projects in the Playwright config, not an afterthought.

Split memory by scope, not by size

What every agent needs (deploy target, workflow rules) goes in shared project memory. What only one agent needs (a specialist's domain quirks) goes in that agent's own memory. Get this backwards and you either bloat every agent's context or lose knowledge that should've been shared.

Default CI to manual, tier up from there

Every workflow starts as a manual trigger — zero cost until you actually want a check. Add smoke tests on PRs once local checks aren't enough. Save nightly regression and visual diffing for when you have minutes to spare.

Screenshot it, don't eyeball it

Feed a fresh screenshot to a vision-capable model instead of manually scanning the UI. It catches spacing, alignment, and contrast reliably — and it's the check that gets skipped first when you're in a hurry. It won't catch a bad design decision, so it's a floor, not a replacement for review.

Parallel dispatch, when you're ready for it

Only once two tasks share no state and don't depend on each other's output — a UI pass and a copy pass on the same feature, say. Dispatch both, let them work independently, synthesize the results yourself. No message bus, no orchestration layer. Just tasks that were never sequential to begin with.

Apply these principles agentically

▸ expand

The gate above is a one-time check before you start. This is different — paste it into a project's instructions (CLAUDE.md, AGENTS.md, or the start of a session) so your agent treats these as standing rules, not things it read once and forgot.

Adopt these as working rules for this project, not just advice to
remember. Push back on me if I ask you to skip one under deadline
pressure -- that's exactly when they matter most.

THINK BEFORE YOU PLAN
When I bring you a task, diverge before you converge: surface the
unstated assumption and at least one alternative before proposing a
plan. Don't skip straight to a plan because I seem rushed -- say so
if that's what's happening.

VERIFY, THEN REVIEW -- NOT THE SAME GATE
"It builds and passes tests" and "it's actually correct" are two
different checks. Never report something done on the first alone.
Name the specific bugs, security issues, or convention violations you
checked for in review, even if you found none.

MEMORY BY SCOPE, NOT CONVENIENCE
Before writing anything to persistent memory, ask whether every
future session needs it or only one specialty does. Shared memory
gets facts every agent needs. Anything domain-specific goes in that
domain's own memory, not the shared file.

DECOMPOSE INSTEAD OF CENTRALIZING
If you're holding multiple unrelated domains in one agent
definition, one long instructions file, or one sprawling
conversation, stop and split it. That instinct -- keep it simple by
centralizing -- is the most common way this goes wrong.

MEASURE BEFORE CLAIMING BETTER
Before calling a prompt, model, or approach change an improvement,
run it against a handful of real cases and count pass/fail. "This
feels better" is not evidence, and you should say so if that's all
either of us has.

SCREENSHOT VISUAL CHANGES, DON'T ASSERT THEM
After any UI change, take a screenshot and describe what you
actually see before claiming it's correct. Correct-looking code is
not the same claim as a correct-looking screen.

CI STAYS MANUAL UNTIL IT EARNS AUTOMATION
Default new workflows to a manual trigger. Only propose automatic
triggers once there's a concrete reason -- a bug that would've been
caught, a teammate who needs the signal.

PARALLEL ONLY WHEN GENUINELY INDEPENDENT
Before running two things in parallel, state why they share no state
and don't depend on each other's output. If you can't state that
plainly, run them sequentially instead.

BLAST RADIUS BEFORE STATE CHANGES
Before any delete, force operation, bulk edit, or production change,
print the exact scope it will touch and confirm the evidence supports
that action. "It should only affect X" is a guess until you've listed
what it affects.

WHEN THE DOCS AND REALITY DISAGREE, SAY SO
If live code, data, or config contradicts the spec, the issue, or
these instructions, report the contradiction and stop that thread.
Don't quietly reconcile the two -- that's how a wrong assumption ships.

STOP AFTER TWO FAILED FIXES
If verification fails twice after a reasonable fix, stop. Report the
exact commands, the output, and your best hypothesis. Don't loop a
third time, and don't switch to a different approach without saying so.

Before you forget

Three ways to lose a week.

Security

Auth bypass for E2E tests must never be reachable in production. Gate it behind an env var that only exists in test/dev — middleware checks process.env.E2E_BYPASS_AUTH === '1', and in production that variable simply doesn't exist.

Cost

GitHub Actions' free tier is 2,000 minutes a month. Each E2E run costs 2–10 minutes. Setting all your workflows to run automatically on every push can burn 500+ minutes a month on an active repo — start with smoke-on-PR, promote to nightly and visual only once you need them.

False confidence

Tests pass, types check, lint is clean, and the layout is still broken — overlapping elements, wrong spacing, invisible text on a same-colored background. None of that is visible to a programmatic assertion. Only a pixel-diff against a saved baseline catches it, and it's the check most setups skip.

What's actually ours

The loop, and who covers each step.

We don't own every step, and we're not pretending to. Where we have a published skill, it's linked below. Where we don't, that's someone else's tool we depend on — linked instead of reinvented.

StepCoverageWhere
Think not ours brainstorming is part of obra/superpowers, a plugin we run and don't try to replace.
Plan not ours writing-plans and test-driven-development, also obra/superpowers.
Build ours db-truth and db-migration-safety (schema work), read-the-damn-docs (third-party APIs) — agent-plugins.
Verify ours prove-it — agent-plugins, ship-pipeline pack.
Review ours review-merge-pipeline, the whole second-opinion pack, and redline for the judgment calls code review doesn't cover — agent-plugins.
Ship ours deploy plus the release-ops pack — agent-plugins.

Get moving

Fork the repo, or install just the skills.

Both get you the same patterns this page describes. The difference is whether you want a starting repo or a set of skills that drop into the project you already have.

Starter repo

Fork full‑starter

CLAUDE.md, agents, skills, hooks, Playwright E2E, and CI workflows — pre-merged, one repo, setup docs for Claude Code, Codex, and Cursor. Fork it on GitHub, clone your fork, and run npm install. Then open it in your AI tool and paste Read CLAUDE-SETUP.md and set up my project. (or the Codex or Cursor equivalent). The README has the five steps.

git clone \
  https://github.com/YOUR-USERNAME/full-starter
github.com/stylusnexus/full-starter →
Marketplace

Install agent‑plugins

The review-merge-ship loop, evidence-before-done verification, and more as installable skills — the same ones run in production, into whatever project you already have. Claude Code adds the marketplace once, then installs whichever packs you want; every other agent uses npx skills add instead.

/plugin marketplace add \
  stylusnexus/agent-plugins
github.com/stylusnexus/agent-plugins →