A short guide for anyone starting to build with AI coding agents: the working loop, the four mistakes that cost us the most time, and what to do this week.
Why this page
This used to be two long guides — one for development, one for testing. Both got read once and bookmarked. This is what's left after cutting everything that wasn't a decision you'd actually make: the loop, the four failures that cost the most time, and the concrete first moves. Then it points you at the repo you fork and the skills you install, so you're building instead of reading.
It is not software to install, a full course, or research. It is one team's experience, kept short on purpose. Some examples use Claude Code file names; the loop and the failures apply to any coding agent.
The loop
Diverge before you commit. What are we actually solving, what's the unstated assumption, what could go wrong three different ways? No plan yet — just surfacing the options and edge cases a rushed first instinct would skip. This is the step deadline pressure cuts first, which is exactly the wrong one to cut.
Converge and commit. Turn what thinking surfaced into a specific, written plan — files to touch, sequence, test cases in plain English, a definition of done. Specialist agents, one per area you work in, weigh in from their own lane on that plan, not one generalist agent holding every domain in its head at once. The full-starter template ships backend, docs, test-review, and browser-testing agents plus templates for adding your own.
Agents write against a CLAUDE.md that stays small on purpose, domain guidances that load only when relevant, and hooks that inject context automatically — instead of a growing wall of always-loaded instructions.
Does it run? Build, typecheck, test, lint — mechanical pass/fail, no judgment involved. If this fails, review never starts. AI mock fixtures keep the loop fast; real calls are for local spot-checks, never CI.
One check earns its own rule, because we shipped the same bug three times before adding it: any value threaded through two or more hops needs a test at its final destination proving it arrived. The bug class survives ordinary review because every hop is locally correct — each function does receive the value and pass it on — but a destructure omits it, a mapping drops it, or a spread overwrites it. Nothing throws. A test at hop three proves nothing about hop five.
Should it run this way? Everything that passed verify but is still wrong — a bug, a security hole, a convention violated, an edge case nobody wrote a test for. Verify is a machine's job; review needs judgment, which is why solo developers get it from a second opinion instead of a teammate.
A second opinion covers code. It does not cover the things only you can judge — whether the draft says what you meant, whether the plan is the one you want built. There the bottleneck is not judgment but the channel: describing a wording change in chat costs more than making it, and “the third paragraph feels off” throws away the phrase that caused the feeling. Mark the artifact up directly instead, and hand back edits and comments anchored to exact quotes.
What made it this far gets measured, not assumed. A prompt or model change that "feels better" is guessed at 40% wrong — ten generations and a pass/fail count catches what manual testing misses.
Lessons, paid for already
Each of these shipped, was used for weeks, and had to be unwound. Skip straight to the fix.
One agent definition covering frontend, backend, database, content, security, analytics. Generic output that missed domain nuance — backend patterns applied to frontend code, UI suggestions during infrastructure work. Fix: split by domain. One agent, one lane.
Every rule, pattern, gotcha, and convention in one file. Felt thorough. Started burning thousands of tokens loading AI-safety docs to edit CSS, and rules buried mid-file stopped getting followed. Fix: a small root file, domain guidances loaded on demand.
An agent discovers a workaround, it lives in that conversation, and vanishes when the conversation ends. Same discovery, same 30-minute debugging session, the following week. Fix: persistent per-agent memory, one line per insight, linked for depth.
Prompt changes judged by whether they "felt better." When actually measured, intuition was wrong about 40% of the time — changes assumed to help sometimes made things worse. Fix: ten generations, structured pass/fail, before trusting a change.
First moves
The most expensive mistake isn't a bad pattern — it's building the wrong thing well. Paste this into your agent before you scaffold anything.
Before we start a new project, run me through a validation gate. Don't let me skip a step because I'm excited, and push back if I try to write code before step 3 is done. 1. Ask who the target user is and what's the most annoying part of their day related to this idea. Help me write one sentence: "When ___ happens, I want to ___, so I can ___." 2. Classify it: painkiller (costs real hours or dollars every week) or vitamin (nice to have). Tell me plainly if it's a vitamin -- that doesn't clear this gate. 3. Help me design the cheapest possible demand test I can run in 7 days -- a waitlist, a preorder, or a manual pilot offer -- and define what "someone actually commits" looks like for this specific idea. A soft "yes that's cool" doesn't count. Money or a hard commitment does. 4. Only once I report back real evidence from step 3, help me scope a 24-hour MVP: one core loop, end to end. Ship it only if all three are true -- a new user gets value within 60 seconds, that value is immediate, and I can manually fix any bad output myself.
Adapted from the validation discipline in The Vibe Coding Playbook (Murugappan) — own the pain data, then prove willingness-to-pay before you open the code editor.
Dev commands and 3-4 gotchas. It grows on its own — every entry earns its place because you hit that exact issue twice.
They load into context every run. Use topic files for depth, one line per insight in the index, and link out rather than inline.
Brainstorm, TDD, verify. Those three alone change how you work. Adopt the rest — plan, execute, debug, review, ship, measure — once you feel the gap.
Master the chain — brainstorm, plan, execute, verify, ship — one agent at a time. Parallel dispatch is a coordination problem you don't need on day one.
Start with context7 or GitHub. Add the next one only when an agent keeps needing the same external system and doesn't have it.
A layout that holds at 1920px can break at 768px. Separate desktop and mobile projects in the Playwright config, not an afterthought.
What every agent needs (deploy target, workflow rules) goes in shared project memory. What only one agent needs (a specialist's domain quirks) goes in that agent's own memory. Get this backwards and you either bloat every agent's context or lose knowledge that should've been shared.
Every workflow starts as a manual trigger — zero cost until you actually want a check. Add smoke tests on PRs once local checks aren't enough. Save nightly regression and visual diffing for when you have minutes to spare.
Feed a fresh screenshot to a vision-capable model instead of manually scanning the UI. It catches spacing, alignment, and contrast reliably — and it's the check that gets skipped first when you're in a hurry. It won't catch a bad design decision, so it's a floor, not a replacement for review.
Only once two tasks share no state and don't depend on each other's output — a UI pass and a copy pass on the same feature, say. Dispatch both, let them work independently, synthesize the results yourself. No message bus, no orchestration layer. Just tasks that were never sequential to begin with.
The gate above is a one-time check before you start. This is different — paste it into a project's instructions (CLAUDE.md, AGENTS.md, or the start of a session) so your agent treats these as standing rules, not things it read once and forgot.
Adopt these as working rules for this project, not just advice to remember. Push back on me if I ask you to skip one under deadline pressure -- that's exactly when they matter most. THINK BEFORE YOU PLAN When I bring you a task, diverge before you converge: surface the unstated assumption and at least one alternative before proposing a plan. Don't skip straight to a plan because I seem rushed -- say so if that's what's happening. VERIFY, THEN REVIEW -- NOT THE SAME GATE "It builds and passes tests" and "it's actually correct" are two different checks. Never report something done on the first alone. Name the specific bugs, security issues, or convention violations you checked for in review, even if you found none. MEMORY BY SCOPE, NOT CONVENIENCE Before writing anything to persistent memory, ask whether every future session needs it or only one specialty does. Shared memory gets facts every agent needs. Anything domain-specific goes in that domain's own memory, not the shared file. DECOMPOSE INSTEAD OF CENTRALIZING If you're holding multiple unrelated domains in one agent definition, one long instructions file, or one sprawling conversation, stop and split it. That instinct -- keep it simple by centralizing -- is the most common way this goes wrong. MEASURE BEFORE CLAIMING BETTER Before calling a prompt, model, or approach change an improvement, run it against a handful of real cases and count pass/fail. "This feels better" is not evidence, and you should say so if that's all either of us has. SCREENSHOT VISUAL CHANGES, DON'T ASSERT THEM After any UI change, take a screenshot and describe what you actually see before claiming it's correct. Correct-looking code is not the same claim as a correct-looking screen. CI STAYS MANUAL UNTIL IT EARNS AUTOMATION Default new workflows to a manual trigger. Only propose automatic triggers once there's a concrete reason -- a bug that would've been caught, a teammate who needs the signal. PARALLEL ONLY WHEN GENUINELY INDEPENDENT Before running two things in parallel, state why they share no state and don't depend on each other's output. If you can't state that plainly, run them sequentially instead. BLAST RADIUS BEFORE STATE CHANGES Before any delete, force operation, bulk edit, or production change, print the exact scope it will touch and confirm the evidence supports that action. "It should only affect X" is a guess until you've listed what it affects. WHEN THE DOCS AND REALITY DISAGREE, SAY SO If live code, data, or config contradicts the spec, the issue, or these instructions, report the contradiction and stop that thread. Don't quietly reconcile the two -- that's how a wrong assumption ships. STOP AFTER TWO FAILED FIXES If verification fails twice after a reasonable fix, stop. Report the exact commands, the output, and your best hypothesis. Don't loop a third time, and don't switch to a different approach without saying so.
Before you forget
Auth bypass for E2E tests must never be reachable in production. Gate it behind an env var that only exists in test/dev — middleware checks process.env.E2E_BYPASS_AUTH === '1', and in production that variable simply doesn't exist.
GitHub Actions' free tier is 2,000 minutes a month. Each E2E run costs 2–10 minutes. Setting all your workflows to run automatically on every push can burn 500+ minutes a month on an active repo — start with smoke-on-PR, promote to nightly and visual only once you need them.
Tests pass, types check, lint is clean, and the layout is still broken — overlapping elements, wrong spacing, invisible text on a same-colored background. None of that is visible to a programmatic assertion. Only a pixel-diff against a saved baseline catches it, and it's the check most setups skip.
What's actually ours
We don't own every step, and we're not pretending to. Where we have a published skill, it's linked below. Where we don't, that's someone else's tool we depend on — linked instead of reinvented.
db-truth and db-migration-safety (schema work), read-the-damn-docs (third-party APIs) — agent-plugins.
review-merge-pipeline, the whole second-opinion pack, and redline for the judgment calls code review doesn't cover — agent-plugins.
Get moving
Both get you the same patterns this page describes. The difference is whether you want a starting repo or a set of skills that drop into the project you already have.
CLAUDE.md, agents, skills, hooks, Playwright E2E, and CI workflows — pre-merged, one repo, setup docs for Claude Code, Codex, and Cursor. Fork it on GitHub, clone your fork, and run npm install. Then open it in your AI tool and paste Read CLAUDE-SETUP.md and set up my project. (or the Codex or Cursor equivalent). The README has the five steps.
git clone \ https://github.com/YOUR-USERNAME/full-startergithub.com/stylusnexus/full-starter →
The review-merge-ship loop, evidence-before-done verification, and more as installable skills — the same ones run in production, into whatever project you already have. Claude Code adds the marketplace once, then installs whichever packs you want; every other agent uses npx skills add instead.
/plugin marketplace add \ stylusnexus/agent-pluginsgithub.com/stylusnexus/agent-plugins →