The Describe-Refresh-Despair Loop (A Love Story Gone Wrong)
Let’s be real about what’s happening here.
You have a vision in your head. A feeling of what the page should look like. Something about the proportions, the spacing, the way elements breathe together.
But translating that feeling into words?
That’s where the wheels come off the wagon.
The loop goes something like this:
Describe a change in words (that you hope Claude interprets correctly)
Wait for Claude to apply it
Refresh the page
Realize it’s not quite what you meant
Try different words—maybe “airy” instead of “spacious”?
Repeat 8-12 times
Eventually settle for “close enough” while muttering under your breath
I burned way too many hours in this loop last week, redesigning a page. Every iteration felt like playing telephone with my own design instincts.
(Spoiler: I was the one garbling the message.)
The problem isn’t Claude. The problem is that words are a terrible interface for visual decisions.
.
.
.
Enter the Claude Code Playground Skill: Your New Best Friend
Anthropic built an official plugin—the Playground skill—that adds something radical to Claude Code:
A visual layer between your brain and your codebase.
Here’s how it works: Claude analyzes your existing page, then generates a self-contained HTML file with sliders, dropdowns, and presets that let you see different design directions instantly.
No code changes. No refreshing. No “what I said vs. what I meant” shenanigans.
Just a live preview that updates as you click.
Once you’ve dialed in the exact design you want—with your actual eyeballs—the playground generates a natural language prompt describing your choices.
Copy. Paste. Execute.
One pass. Done.
👉 The key insight: You’re no longer translating visual intuition into words. The playground does that translation for you.
.
.
.
The Real-World Test: Redesigning a WooCommerce Product Page
Enough theory. Let me show you exactly how this worked on my LicenseWP product page.
The existing page was… fine. Functional. The kind of “fine” that makes you wince slightly every time you look at it.
The image area felt cramped. The pricing cards looked squished together. The “Add to cart” button was doing its best impression of a wallflower at a party.
I wanted to experiment with different proportions, card styles, and CTA treatments.
The old me would have spent an hour going back-and-forth with Claude in text prompts.
The new me? Installed the Playground skill.
.
.
.
Step 0: Install the Plugin (30 Seconds, Tops)
First things first—you need the Claude Code Playground skill installed.
Run /plugin in Claude Code, switch to the Discover tab, and search for “playground.”
Hit space to toggle it on.
That’s it. Claude now knows how to build interactive design playgrounds. (11.4K installs can’t be wrong, right?)
.
.
.
Step 1: Send Claude to Do Reconnaissance
Here’s where the magic starts.
I gave Claude a prompt telling it to visit my product page, study the layout like an art critic at a museum, and then build a Design Layout Playground based on what it found.
Claude opens Chrome and starts analyzing—like a very thorough house inspector, but for web pages.
It reads the DOM, extracts text content, and runs JavaScript to pull the full HTML structure.
Screenshots. Scrolling. Full page analysis from header to footer.
(Claude is nothing if not thorough.)
.
.
.
Step 2: Claude Builds You a Custom Design Tool
Once Claude understands your page, it loads the Playground skill and reads the design template.
Then—and here’s the part that made me do a little happy dance—it creates a single self-contained HTML file with everything baked in.
1,253 lines. A complete interactive design tool, built specifically for your page, in about 30 seconds.
Here’s what Claude built for me:
5 presets — each one a cohesive design direction
7 control groups — layout, gallery, typography, cards, buttons, tabs, colors
Live preview — using my actual page content (real product names, real prices, real structure)
Ferpetesake, this is exactly what I needed.
.
.
.
Step 3: Play With Designs Like a Kid With New LEGOs
Open the HTML file in your browser.
And then—I’m not going to lie—prepare to lose 20 minutes just playing.
The left panel has all the controls. The right panel shows a live preview that updates instantly as you change anything.
(I may have spent an unreasonable amount of time just clicking the “Gallery Style” options back and forth. Don’t judge me.)
I started by clicking through the presets to find a direction:
Bold SaaS — too aggressive for this product
Compact & Dense — too cramped (we’re selling premium themes, not packing a suitcase)
Clean Editorial — closer! But needed tweaks
Then I fine-tuned individual controls:
Grid split: 50/50 (equal width for image and details)
Gallery style: Framed (light background with border instead of that gradient)
Every single change reflected instantly in the preview.
👉 Here’s what hit me: I wasn’t describing what I wanted anymore. I was seeing it. And clicking until it looked right.
.
.
.
Step 4: Copy the Magic Prompt
Once the design felt right, I scrolled to the bottom of the playground.
And there it was.
The Prompt Output panel had already written a clear instruction describing exactly what I chose—and only what differed from the defaults.
Redesign the product single page at http://localhost:8107/product/theme-pro-tc011/
with the following design changes:
- product grid split: 50/50
- section spacing: 64px
- gallery style: framed
- gallery border radius: 24px
- product title size: 36px
- tier card border radius: 8px
- selected tier highlight: border-only
- CTA button style: pill-shaped
- CTA button width: full-width
- CTA button size: large
- tab style: pill tabs
- tab alignment: stretch
No ambiguity. No “make it more modern” nonsense.
Just precise specifications that Claude can execute without guessing.
One click on Copy Prompt.
.
.
.
Step 5: Let Claude Do Its Thing
Back in Claude Code. Paste.
Claude immediately recognizes the design instructions and enters plan mode to explore the codebase.
33 tool uses. 101k tokens. 2 minutes of thinking.
Claude reads every relevant CSS file, understands the variable system, and maps out exactly what needs to change.
Then it presents the plan:
Every change mapped out. Exact line numbers. Before/after CSS.
(Is it weird that I find this deeply satisfying? Don’t answer that.)
Verification steps included. Dark mode considerations. Mobile responsiveness checks.
I approved. Claude started executing.
9 CSS edits. All applied.
And… done.
.
.
.
The Result (AKA: The Part Where I Do a Victory Lap)
Here’s the final redesigned product page:
Equal-width layout. The image and product details now have balanced visual weight.
Framed gallery. Clean border instead of that dated gradient background.
Full-width CTA. The “Add to cart” button finally commands the attention it deserves.
Pill tabs. Stretched across the full width with a modern, cohesive feel.
The design matches what I saw in the playground preview—applied in a single pass.
HECK YES.
.
.
.
The Before & After (Because We All Love a Good Transformation)
Before
After
40/60 grid split (cramped image)
50/50 split (balanced)
Gradient gallery background
Framed with border
Small inline CTA button
Full-width pill CTA
Underline tabs
Stretched pill tabs
10+ prompt iterations
1 pass
The old workflow would’ve taken an hour of back-and-forth (and probably some mild frustration-snacking).
This took 15 minutes—and honestly, most of that was me playing with the controls because it was genuinely fun.
.
.
.
When the Claude Code Playground Skill Really Shines
👉 Redesigning existing pages. You already have something. You want to explore variations without breaking it.
👉 Client projects. Preview before you commit. Show options before you build. (Clients LOVE this, by the way.)
👉 Design indecision. When you don’t know what you want—and let’s be honest, that’s more often than we’d like to admit—clicking through presets beats describing in words.
👉 Reducing the prompt iteration loop. One visual session replaces 10+ text-based rounds of “no, more like… actually less like that… wait, go back.”
The playground acts as a translation layer between your visual intuition and Claude’s execution capabilities. You figure out what “right” looks like with your eyes, then communicate that with precision.
The Prompt Template (Steal This)
Here’s the full prompt I use. Copy it. Adapt it. Make it yours.
First, use the browser to visit and read this page: [YOUR_PAGE_URL]
Study the page's current layout structure, section hierarchy, component patterns,
and overall visual design. Take note of how content is organized and what
elements are present.
Then, use the "playground" skill (design-playground template) to build an
interactive Design Layout Playground based on what you found on that page.
The playground should let me visually explore different layout and component
style combinations for that page.
## Presets
Include 3–5 named presets that snap all controls to a cohesive combination,
inspired by what would work well for the page's content. For example:
- "Clean Editorial" — airy spacing, narrow content width, minimal components
- "Bold & Modern" — full-width hero, elevated cards, bold CTAs
- "Compact Dashboard" — tight spacing, grid cards, minimal chrome
- Adapt these to fit the actual content and purpose of the page
## Preview
- Single live preview panel that updates instantly on every control change
- The preview should use a simplified but recognizable representation of the
actual page content (use real section names, headings, and placeholder text
that matches the page structure)
- Use raw CSS (no Tailwind or frameworks)
## Output Location
- Save the playground HTML file to `notes/playground/` folder (create it if
it doesn't exist)
## Prompt Output
- Generate a natural language instruction at the bottom that I can copy and
paste back into Claude to implement the chosen design
- The prompt should describe the layout and component decisions in enough
detail to be actionable without the playground
- Only mention choices that differ from the defaults
- Frame it as a direction, e.g.: "Redesign the page with a full-width hero
section, 3-column card grid with elevated shadows and 16px gap, airy
section spacing (64px), pill-shaped CTAs positioned inline..."
- Include the source page URL in the generated prompt for context
Replace [YOUR_PAGE_URL] with whatever page you want to redesign.
.
.
.
Your Turn
Next time you’re about to type “make it more modern” or “adjust the spacing” or “try a different card style”—stop.
Build a playground first.
Let your eyes make the decisions. Let the Claude Code Playground skill translate those decisions into words. Let Claude execute them precisely.
What page are you going to redesign with this workflow?
Go install the plugin. Run /plugin, search “Playground”, toggle it on.
Now.
(And maybe clear your schedule. Because once you start playing with those sliders, you might lose track of time.)
Well, not entirely wrong. The tests ran. Things got verified. But what I showed you? That wasn’t a true Ralph loop—not the way Geoffrey Huntley originally designed it. And the difference matters more than I realized at the time.
(Stay with me here. This confession has a happy ending.)
.
.
.
The Problem I Didn’t See Coming
The real Ralph loop is supposed to wipe memory clean at the start of each iteration. No leftover context. No accumulated baggage. Just a fresh, focused agent tackling one task at a time.
The Ralph loop plugin from Claude Code’s official marketplace? It preserves context from the previous loop. The plugin relies on a stop hook to end and restart each iteration—but the conversation history tags along for the ride.
And that’s where everything quietly falls apart.
Here’s what this actually looks like in practice:
Imagine you’re setting out on a multi-day hiking trip. Every morning, you pack your backpack for that day’s trail.
Now imagine that instead of emptying your pack each night, you just… keep adding to it. Day one’s water bottles. Day two’s snacks. Day three’s rain gear (even though it’s sunny now). By day five, you’re hauling 40 pounds of stuff you don’t need, and you can barely focus on the trail in front of you.
That’s context rot.
It happens when an AI model’s performance degrades because its context window gets bloated with accumulated information from previous tasks. The more history your agent carries forward, the harder it becomes for the model to stay sharp on what actually matters right now.
👉 The takeaway: Fresh context isn’t a nice-to-have. It’s the whole point.
.
.
.
What Context Rot Actually Looks Like
Let me make this concrete with Claude Code testing:
Iteration 1: Claude runs test TC-001. Context is clean. Performance is sharp. The backpack is light.
Iteration 5: Claude runs test TC-005. But it’s also dragging along memories of TC-001 through TC-004. The pack is getting heavy.
Iteration 15: Claude runs test TC-015. The model is now swimming through accumulated history, trying to find what actually matters among all the gear from previous days.
Iteration 25: Claude runs test TC-025. Performance has degraded. The model makes weird mistakes. It forgets what it was supposed to verify—because it’s exhausted from carrying everyone else’s context.
Same trail. Same agent. Completely different performance.
And here’s the frustrating part: you might not even notice it happening. The tests still run. They just run… worse. Slower. Less reliably. With occasional bizarre failures that make you question your own test plan.
.
.
.
The Solution That Was Already There
So I went looking for a better approach to Claude Code testing—something that would give me the clean-slate benefits of a proper Ralph loop without the context accumulation problem.
And I found it in a tool I’d been using for something else entirely: Claude Code’s task management system.
Here’s where it gets interesting.
The task management system gives you the same effect as a properly implemented Ralph loop—but with something the Ralph loop never had: dependency management.
Think back to the hiking metaphor.
Each sub-agent is like a fresh hiker starting a new day with an empty pack. They get their assignment, they complete their section of trail, they report back. Then the next hiker takes over with their own empty pack.
No accumulated gear.
No context rot.
No performance degradation over time.
But here’s the bonus: the task management system also handles situations where “Day 3’s trail can’t start until Day 2’s bridge gets built.” Dependencies get tracked automatically. Tests that need prerequisites don’t run until those prerequisites pass.
(Is that two features in one? Well, is a package of two Reese’s Peanut Butter Cups two candy bars? I say it counts as one delicious solution.)
.
.
.
How Claude Code Testing Actually Works With Task Management
Let me show you exactly how to set this up.
Fair warning: there are a lot of screenshots coming. But I promise each one shows something important about the workflow.
Step 1: Put the Prompt as a Command
First, store the entire testing prompt as a command file. This makes triggering your Claude Code testing workflow trivially easy—just a slash command away.
The full prompt (I’ll include it at the end—it’s long but worth having) tells Claude exactly how to read your test plan, create tasks, set dependencies, execute tests sequentially, and track results.
Step 2: Trigger the Command
With the command saved, execution is just:
That’s it. Type /run_test_plan and let the system take over.
Step 3: Claude Reads Your Specs
Since we’re starting fresh—no memory of previous execution—Claude first reads your original specs, test plan, and implementation plan to understand the context.
(Remember: empty backpack. The agent needs to load up on just what it needs for this journey.)
Step 4: Claude Creates the Tasks
After understanding the context, Claude creates one task per test case. Watch how it automatically detects dependencies:
See that dependency analysis?
TC-004, TC-005, TC-006, TC-008 depend on TC-003 (password field must exist first)
TC-016, TC-017 depend on TC-014 (categories must exist first)
TC-019 through TC-023 depend on TC-018 (priority dropdown must exist first)
TC-029 depends on TC-027 (accent color must be saved first)
The system figured this out by reading the test plan. No manual configuration required.
Step 5: Dependencies Get Locked In
All 30 tasks created.
Now Claude sets up the dependencies and verifies everything:
Step 6: Test Status File Created
Claude creates a test-status.json file to track everything—machine-readable, resumable, and audit-friendly:
Tasks blocked by TC-003: TC-004, TC-005, TC-006, TC-008
Tasks blocked by TC-014: TC-016, TC-017
Tasks blocked by TC-018: TC-019, TC-020, TC-021, TC-022, TC-023
Tasks blocked by TC-027: TC-029
Step 7: First Task Begins
Here’s where the magic happens.
Claude spawns a sub-agent—with fresh context—to execute TC-001:
“You are a test execution sub-agent. You have ONE job: execute and verify ONE test case.”
That’s the key instruction. Fresh hiker. Empty backpack. Single trail.
Step 8: Browser Automation for Testing
The sub-agent uses Claude Code’s browser automation to test like a real user would:
It navigates to URLs, clicks buttons, fills forms, takes screenshots at verification points, and checks the DOM state against expected outcomes.
Real browser. Real interactions. Real Claude Code testing.
Step 9: Test Status Gets Updated
After completing a test, the sub-agent updates the status file:
Step 10: Human-Readable Results Too
The results also get appended to a markdown log for human review:
Every test gets logged in two places:
test-status.json for machine parsing
test-results.md for human review
(Because sometimes you want to query the data programmatically, and sometimes you just want to read what happened over coffee. Both are valid.)
Step 11: Automatic Progression to Next Task
Once TC-001 completes, Claude automatically moves to TC-002:
Look at those stats: 36 tool uses, 48.1k tokens, 4 minutes 19 seconds for TC-001.
Then a completely fresh sub-agent spawns for TC-002. New hiker. New backpack. No accumulated context from TC-001.
Step 12: Bugs Found? Claude Fixes Them.
TC-002 found a bug. Here’s what happened:
“TC-002 passed after 1 fix (Quick Setup wasn’t saving color options — now fixed in WizardAjax.php).”
The sub-agent detected the failure, analyzed the root cause, implemented a fix, and re-ran the test. All autonomously. All within the same fresh context.
Step 13: Dependencies Unlock Automatically
Now watch the dependency system in action.
Once TC-003 passes:
“TC-003 passed. Now TC-004, TC-005, TC-006, TC-008 are unblocked.”
The password field exists now. All the tests that depend on it can finally run.
👉 This is why dependencies matter: They prevent tests from running before their preconditions are met—avoiding the exact conflicts where one agent messes with something another agent needs.
Steps 14-15: The Marathon Continues
It keeps going. Test after test. Each sub-agent fresh and focused:
Every test runs sequentially. Every sub-agent gets clean context. Every dependency is respected. No context rot in sight.
Step 16: All Tests Complete
After 2 hours and 12 minutes:
30 tests. All passed. Zero known issues.
Step 17: The Full Summary
The orchestrator writes a comprehensive summary:
Here’s what got verified:
All 6 Critical tests passed (password handling, priority validation, accent color persistence)
Server-side validation confirmed working (urgent priority rejected, invalid hex rejected, password never stored in wp_options)
UI behaviors verified (notice dismiss, auto-hide, field error clearing, color sync)
Accessibility attributes verified on priority dropdown
And that bug that got fixed? handleQuickSetup() in WizardAjax.php wasn’t saving desq_primary_color or desq_accent_color options. Found during TC-002. Fixed autonomously.
.
.
.
Why This Actually Works Better
Let me be direct about the comparison:
Aspect
Ralph Loop Plugin
Task Management System
Context
Preserves across iterations
Fresh per sub-agent
Dependencies
None
Built-in blocking
Parallel Safety
Risky
Sequential by default
State Tracking
Basic stop hook
JSON + Markdown logs
Bug Fixing
Manual
Automatic (up to 3 attempts)
Resumability
Limited
Full state recovery
The Ralph loop was supposed to start each iteration with a clean slate. The task management system actually delivers on that promise—and adds dependency management that prevents tests from stepping on each other.
.
.
.
The Full Prompt (Copy This)
Here’s the complete command file to drop into .claude/commands/run_test_plan.md:
PROMPT: Execute Test Plan Using Claude Code Task Management System
Loading longform...
Couldn't load content. Please try again.
We are executing the test plan. All implementation is complete. Now we verify it works.
## Reference Documents
- **Test Plan:** `notes/test_plan.md`
- **Implementation Plan:** `notes/impl_plan.md`
- **Specs:** `notes/specs.md`
- **Test Status JSON:** `notes/test-status.json`
- **Test Results Log:** `notes/test-results.md`
---
## Phase 1: Initialize
### Step 1: Check for Existing Run (Resumption)
Before creating anything, check if a previous test run exists:
1. Check if `notes/test-status.json` exists
2. Check if there are existing tasks via `TaskList`
**If both exist and tasks have results:**
- This is a **resumed run** — skip to Phase 2 (Step 7)
- Announce: "Resuming previous test run. Skipping already-passed tests."
- Only execute tasks that are still `pending` or `fail` (with fixAttempts < 3)
**If no previous run exists (or files are missing):**
- Continue with fresh initialization below
### Step 2: Read the Test Plan
Read `notes/test_plan.md` and extract ALL test cases. Auto-detect the TC-ID pattern used (e.g., `TC-001`, `TC-101`, `TC-5A`, etc.).
For each test case, note:
- TC ID
- Name
- Priority (Critical / High / Medium / Low — default to Medium if not stated)
- Preconditions
- Test steps and expected outcomes
- Test data (if any)
- Dependencies on other test cases (if any)
### Step 3: Analyze Test Dependencies
Determine which test cases depend on others. Common dependency patterns:
- A "saves data" test may depend on a "displays default" test
- A "form submission" test may depend on "form validation" tests
- An "end-to-end" test may depend on individual component tests
If no clear dependencies exist between test cases, treat them all as independent.
### Step 4: Create Tasks
Use `TaskCreate` to create one task per test case. Set `blocked_by` based on the dependency analysis.
**Task description format:**
```
Test [TC-ID]: [Test Name]
Priority: [Priority]
Preconditions:
- [Required state before test]
Steps:
| Step | Action | Expected Result |
|------|--------|-----------------|
| 1 | [Action] | [Result] |
| 2 | [Action] | [Result] |
Test Data:
- [Field]: [Value]
Expected Outcome: [Final verification]
Environment:
- Refer to CLAUDE.md for wp-env details, URLs, and credentials
- WordPress site: http://localhost:8105
- Admin: http://localhost:8105/wp-admin (admin/password)
---
fixAttempts: 0
result: pending
lastTestedAt: null
notes:
```
### Step 5: Generate Test Status JSON
Create `notes/test-status.json`:
```json
{
"metadata": {
"testPlanSource": "notes/test_plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": "<count>",
"pending": "<count>",
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"<TC-ID>": {
"name": "Test case name",
"priority": "Critical|High|Medium|Low",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
### Step 6: Initialize Test Results Log
Create `notes/test-results.md`:
```markdown
# Test Results
**Test Plan:** notes/test_plan.md
**Started:** [CURRENT_TIMESTAMP]
## Execution Log
```
### Verify Initialization
Use `TaskList` to confirm:
- All TC-IDs from the test plan have a corresponding task
- Dependencies are correctly set via `blocked_by`
- All tasks show `result: pending`
Cross-check task count matches `summary.total` in `notes/test-status.json`.
---
## Phase 2: Execute Tests
### Step 7: Determine Execution Order
Use `TaskList` to read all tasks and their `blocked_by` fields. Determine sequential execution order:
1. Tasks with no `blocked_by` (or all dependencies resolved) come first
2. Tasks whose dependencies are resolved come next
3. Continue until all tasks are ordered
**For resumed runs:** Skip tasks where `result` is already `pass` or `known_issue`.
### Step 8: Execute One Task at a Time
For the next eligible task, spawn ONE sub-agent with the instructions below.
**One sub-agent at a time. Do NOT spawn multiple sub-agents in parallel.**
---
#### Sub-Agent Instructions
**You are a test execution sub-agent. You have ONE job: execute and verify ONE test case.**
1. **Read your task** using `TaskGet` to get the full description
2. **Parse the test steps** from the description (everything above the `---` separator)
3. **Parse the metadata** from below the `---` separator
4. **Read CLAUDE.md** for environment details, URLs, and credentials
5. **Execute the test:**
Using browser automation:
- Navigate to URLs specified in the test steps
- Click buttons/links as described
- Fill form inputs with the test data provided
- Take screenshots at key verification points
- Read console logs for errors
- Verify DOM state matches expected outcomes
Follow the test plan steps EXACTLY. Do not skip steps.
6. **Determine the result:**
**PASS** if:
- All expected outcomes verified
- No unexpected console errors
- UI state matches test plan
**FAIL** if:
- Any expected outcome not met
- Unexpected errors
- UI state doesn't match
7. **If PASS:** Update the task description metadata via `TaskUpdate`:
```
---
fixAttempts: 0
result: pass
lastTestedAt: [CURRENT_TIMESTAMP]
notes: [Brief description of what was verified]
```
Mark the task as `completed`.
8. **If FAIL and fixAttempts < 3:**
a. Analyze the root cause
b. Implement a fix in the codebase
c. Increment fixAttempts and update via `TaskUpdate`:
```
---
fixAttempts: [previous + 1]
result: fail
lastTestedAt: [CURRENT_TIMESTAMP]
notes: [What failed, root cause, what fix was applied]
```
d. Re-run the test steps to verify the fix
e. If now passing, set `result: pass` and mark task as `completed`
f. If still failing and fixAttempts < 3, repeat from (a)
9. **If FAIL and fixAttempts >= 3:** Mark as known issue via `TaskUpdate`:
```
---
fixAttempts: 3
result: known_issue
lastTestedAt: [CURRENT_TIMESTAMP]
notes: KI — [Description of the issue, steps to reproduce, severity, suggested fix]
```
Mark the task as `completed`.
10. **Update Test Status JSON** — Read `notes/test-status.json`, update the test case entry and recalculate summary counts, then write back:
- Set `status` to `pass`, `fail`, or `known_issue`
- Update `fixAttempts`, `notes`, `lastTestedAt`
- Increment `metadata.totalIterations`
- Update `metadata.lastUpdatedAt`
- Recalculate `metadata.summary` counts
- If known_issue, add entry to `knownIssues` array
11. **Append to test results log** (`notes/test-results.md`):
```markdown
## [TC-ID] — [Test Name]
**Result:** PASS | FAIL | KNOWN ISSUE
**Tested At:** [TIMESTAMP]
**Fix Attempts:** [N]
**What happened:**
[Brief description of test execution]
**Notes:**
[Observations, errors, or fixes attempted]
---
```
**CRITICAL: Before finishing, verify you have updated ALL THREE locations:**
1. Task description (metadata below `---` separator) via `TaskUpdate`
2. `notes/test-status.json` (test case entry + summary counts)
3. `notes/test-results.md` (appended human-readable entry)
Missing ANY of these = incomplete iteration.
---
### Step 9: Verify and Continue
After each sub-agent finishes, the orchestrator:
1. Uses `TaskGet` to verify the task description metadata was updated
2. Reads `notes/test-status.json` to confirm JSON was updated and summary counts are correct
3. Reads `notes/test-results.md` to confirm a new entry was appended
4. **If any location was NOT updated**, update it before proceeding
5. Determines the next eligible task (unresolved, dependencies met)
6. Spawns the next sub-agent (back to Step 8)
### Step 10: Repeat Until All Resolved
Continue until ALL tasks have `result: pass` or `result: known_issue`.
```
Completion check:
- result: pass → resolved
- result: known_issue → resolved
- result: fail → needs re-test (if fixAttempts < 3)
- result: pending → not yet tested
ALL resolved? → Phase 3 (Summary)
Otherwise? → Next task
```
---
## Phase 3: Summary
### Step 11: Generate Final Summary
When all tasks are resolved, append a final summary to `notes/test-results.md`:
```markdown
# Final Summary
**Completed:** [TIMESTAMP]
**Total Test Cases:** [N]
**Passed:** [N]
**Known Issues:** [N]
## Results
| TC | Name | Priority | Result | Fix Attempts |
|----|------|----------|--------|--------------|
| TC-XXX | [Name] | High | PASS | 0 |
| TC-YYY | [Name] | Medium | KNOWN ISSUE | 3 |
## Known Issues Detail
### KI-001: [TC-ID] — [Issue Title]
**Severity:** [low|medium|high|critical]
**Steps to Reproduce:** [How to see the bug]
**Suggested Fix:** [Potential solution if known]
## Recommendations
[Any follow-up actions needed]
```
---
## Rules Summary
| Rule | Description |
|------|-------------|
| 1:1 Mapping | One task per test case — no grouping |
| Dependencies | Use `blocked_by` to enforce test execution order |
| Sequential | One sub-agent at a time — do NOT spawn multiple in parallel |
| Sub-Agents | One sub-agent per task — fresh context, focused execution |
| Max 3 Attempts | After 3 fix attempts → mark as `known_issue` |
| Metadata in Description | Track `fixAttempts`, `result`, `lastTestedAt`, `notes` below `---` separator |
| Test Status JSON | Always update `notes/test-status.json` after each test |
| Log Everything | Append results to `notes/test-results.md` for human review |
| Resumable | Detect existing run state and continue from where it left off |
| Completion | All tasks resolved = all results are `pass` or `known_issue` |
## Do NOT
- Spawn multiple sub-agents in parallel — execute ONE at a time
- Leave tasks in `fail` state without either retrying or escalating to `known_issue`
- Modify test plan steps — execute them exactly as written
- Forget to update `notes/test-status.json` after each test
- Forget to append to the test results log after each test
- Skip the dependency analysis
- Use `alert()` or `confirm()` in any fix (see CLAUDE.md)
We are executing the test plan. All implementation is complete. Now we verify it works.
## Reference Documents
- **Test Plan:** `notes/test_plan.md`
- **Implementation Plan:** `notes/impl_plan.md`
- **Specs:** `notes/specs.md`
- **Test Status JSON:** `notes/test-status.json`
- **Test Results Log:** `notes/test-results.md`
---
## Phase 1: Initialize
### Step 1: Check for Existing Run (Resumption)
Before creating anything, check if a previous test run exists:
1. Check if `notes/test-status.json` exists
2. Check if there are existing tasks via `TaskList`
**If both exist and tasks have results:**
- This is a **resumed run** — skip to Phase 2 (Step 7)
- Announce: "Resuming previous test run. Skipping already-passed tests."
- Only execute tasks that are still `pending` or `fail` (with fixAttempts < 3)
**If no previous run exists (or files are missing):**
- Continue with fresh initialization below
### Step 2: Read the Test Plan
Read `notes/test_plan.md` and extract ALL test cases. Auto-detect the TC-ID pattern used (e.g., `TC-001`, `TC-101`, `TC-5A`, etc.).
For each test case, note:
- TC ID
- Name
- Priority (Critical / High / Medium / Low — default to Medium if not stated)
- Preconditions
- Test steps and expected outcomes
- Test data (if any)
- Dependencies on other test cases (if any)
### Step 3: Analyze Test Dependencies
Determine which test cases depend on others. Common dependency patterns:
- A "saves data" test may depend on a "displays default" test
- A "form submission" test may depend on "form validation" tests
- An "end-to-end" test may depend on individual component tests
If no clear dependencies exist between test cases, treat them all as independent.
### Step 4: Create Tasks
Use `TaskCreate` to create one task per test case. Set `blocked_by` based on the dependency analysis.
**Task description format:**
```
Test [TC-ID]: [Test Name]
Priority: [Priority]
Preconditions:
- [Required state before test]
Steps:
| Step | Action | Expected Result |
|------|--------|-----------------|
| 1 | [Action] | [Result] |
| 2 | [Action] | [Result] |
Test Data:
- [Field]: [Value]
Expected Outcome: [Final verification]
Environment:
- Refer to CLAUDE.md for wp-env details, URLs, and credentials
- WordPress site: http://localhost:8105
- Admin: http://localhost:8105/wp-admin (admin/password)
---
fixAttempts: 0
result: pending
lastTestedAt: null
notes:
```
### Step 5: Generate Test Status JSON
Create `notes/test-status.json`:
```json
{
"metadata": {
"testPlanSource": "notes/test_plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": "<count>",
"pending": "<count>",
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"<TC-ID>": {
"name": "Test case name",
"priority": "Critical|High|Medium|Low",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
### Step 6: Initialize Test Results Log
Create `notes/test-results.md`:
```markdown
# Test Results
**Test Plan:** notes/test_plan.md
**Started:** [CURRENT_TIMESTAMP]
## Execution Log
```
### Verify Initialization
Use `TaskList` to confirm:
- All TC-IDs from the test plan have a corresponding task
- Dependencies are correctly set via `blocked_by`
- All tasks show `result: pending`
Cross-check task count matches `summary.total` in `notes/test-status.json`.
---
## Phase 2: Execute Tests
### Step 7: Determine Execution Order
Use `TaskList` to read all tasks and their `blocked_by` fields. Determine sequential execution order:
1. Tasks with no `blocked_by` (or all dependencies resolved) come first
2. Tasks whose dependencies are resolved come next
3. Continue until all tasks are ordered
**For resumed runs:** Skip tasks where `result` is already `pass` or `known_issue`.
### Step 8: Execute One Task at a Time
For the next eligible task, spawn ONE sub-agent with the instructions below.
**One sub-agent at a time. Do NOT spawn multiple sub-agents in parallel.**
---
#### Sub-Agent Instructions
**You are a test execution sub-agent. You have ONE job: execute and verify ONE test case.**
1. **Read your task** using `TaskGet` to get the full description
2. **Parse the test steps** from the description (everything above the `---` separator)
3. **Parse the metadata** from below the `---` separator
4. **Read CLAUDE.md** for environment details, URLs, and credentials
5. **Execute the test:**
Using browser automation:
- Navigate to URLs specified in the test steps
- Click buttons/links as described
- Fill form inputs with the test data provided
- Take screenshots at key verification points
- Read console logs for errors
- Verify DOM state matches expected outcomes
Follow the test plan steps EXACTLY. Do not skip steps.
6. **Determine the result:**
**PASS** if:
- All expected outcomes verified
- No unexpected console errors
- UI state matches test plan
**FAIL** if:
- Any expected outcome not met
- Unexpected errors
- UI state doesn't match
7. **If PASS:** Update the task description metadata via `TaskUpdate`:
```
---
fixAttempts: 0
result: pass
lastTestedAt: [CURRENT_TIMESTAMP]
notes: [Brief description of what was verified]
```
Mark the task as `completed`.
8. **If FAIL and fixAttempts < 3:**
a. Analyze the root cause
b. Implement a fix in the codebase
c. Increment fixAttempts and update via `TaskUpdate`:
```
---
fixAttempts: [previous + 1]
result: fail
lastTestedAt: [CURRENT_TIMESTAMP]
notes: [What failed, root cause, what fix was applied]
```
d. Re-run the test steps to verify the fix
e. If now passing, set `result: pass` and mark task as `completed`
f. If still failing and fixAttempts < 3, repeat from (a)
9. **If FAIL and fixAttempts >= 3:** Mark as known issue via `TaskUpdate`:
```
---
fixAttempts: 3
result: known_issue
lastTestedAt: [CURRENT_TIMESTAMP]
notes: KI — [Description of the issue, steps to reproduce, severity, suggested fix]
```
Mark the task as `completed`.
10. **Update Test Status JSON** — Read `notes/test-status.json`, update the test case entry and recalculate summary counts, then write back:
- Set `status` to `pass`, `fail`, or `known_issue`
- Update `fixAttempts`, `notes`, `lastTestedAt`
- Increment `metadata.totalIterations`
- Update `metadata.lastUpdatedAt`
- Recalculate `metadata.summary` counts
- If known_issue, add entry to `knownIssues` array
11. **Append to test results log** (`notes/test-results.md`):
```markdown
## [TC-ID] — [Test Name]
**Result:** PASS | FAIL | KNOWN ISSUE
**Tested At:** [TIMESTAMP]
**Fix Attempts:** [N]
**What happened:**
[Brief description of test execution]
**Notes:**
[Observations, errors, or fixes attempted]
---
```
**CRITICAL: Before finishing, verify you have updated ALL THREE locations:**
1. Task description (metadata below `---` separator) via `TaskUpdate`
2. `notes/test-status.json` (test case entry + summary counts)
3. `notes/test-results.md` (appended human-readable entry)
Missing ANY of these = incomplete iteration.
---
### Step 9: Verify and Continue
After each sub-agent finishes, the orchestrator:
1. Uses `TaskGet` to verify the task description metadata was updated
2. Reads `notes/test-status.json` to confirm JSON was updated and summary counts are correct
3. Reads `notes/test-results.md` to confirm a new entry was appended
4. **If any location was NOT updated**, update it before proceeding
5. Determines the next eligible task (unresolved, dependencies met)
6. Spawns the next sub-agent (back to Step 8)
### Step 10: Repeat Until All Resolved
Continue until ALL tasks have `result: pass` or `result: known_issue`.
```
Completion check:
- result: pass → resolved
- result: known_issue → resolved
- result: fail → needs re-test (if fixAttempts < 3)
- result: pending → not yet tested
ALL resolved? → Phase 3 (Summary)
Otherwise? → Next task
```
---
## Phase 3: Summary
### Step 11: Generate Final Summary
When all tasks are resolved, append a final summary to `notes/test-results.md`:
```markdown
# Final Summary
**Completed:** [TIMESTAMP]
**Total Test Cases:** [N]
**Passed:** [N]
**Known Issues:** [N]
## Results
| TC | Name | Priority | Result | Fix Attempts |
|----|------|----------|--------|--------------|
| TC-XXX | [Name] | High | PASS | 0 |
| TC-YYY | [Name] | Medium | KNOWN ISSUE | 3 |
## Known Issues Detail
### KI-001: [TC-ID] — [Issue Title]
**Severity:** [low|medium|high|critical]
**Steps to Reproduce:** [How to see the bug]
**Suggested Fix:** [Potential solution if known]
## Recommendations
[Any follow-up actions needed]
```
---
## Rules Summary
| Rule | Description |
|------|-------------|
| 1:1 Mapping | One task per test case — no grouping |
| Dependencies | Use `blocked_by` to enforce test execution order |
| Sequential | One sub-agent at a time — do NOT spawn multiple in parallel |
| Sub-Agents | One sub-agent per task — fresh context, focused execution |
| Max 3 Attempts | After 3 fix attempts → mark as `known_issue` |
| Metadata in Description | Track `fixAttempts`, `result`, `lastTestedAt`, `notes` below `---` separator |
| Test Status JSON | Always update `notes/test-status.json` after each test |
| Log Everything | Append results to `notes/test-results.md` for human review |
| Resumable | Detect existing run state and continue from where it left off |
| Completion | All tasks resolved = all results are `pass` or `known_issue` |
## Do NOT
- Spawn multiple sub-agents in parallel — execute ONE at a time
- Leave tasks in `fail` state without either retrying or escalating to `known_issue`
- Modify test plan steps — execute them exactly as written
- Forget to update `notes/test-status.json` after each test
- Forget to append to the test results log after each test
- Skip the dependency analysis
- Use `alert()` or `confirm()` in any fix (see CLAUDE.md)
.
.
.
Your Turn
If you’ve been frustrated with AI-generated code that “works” but doesn’t actually work, give this a shot.
Define your success criteria upfront with a solid test plan. Let Claude Code testing handle the execution and verification through task management. Walk away while it iterates.
The test-fix-retest loop is boring. Tedious. The kind of thing every developer has always done manually.
Now you don’t have to.
What feature are you going to test with this workflow?
Set it up. Let it run. Come back to green checkmarks.
(And maybe grab a coffee while you wait. Your backpack is empty now—you’ve earned the rest.)
52 minutes. 13 tasks. 38 test cases worth of functionality. All built by sub-agents running in parallel.
Here’s what I didn’t tell you.
Half of it didn’t work.
(I know. I KNOW.)
.
.
.
The Part Where I Discover My “Complete” Implementation Is… Not
Let me show you what happened when I actually tested the WooCommerce integration Claude built for me.
Quick context: I have a WordPress theme for coworking spaces. Originally, it used direct Stripe integration for payments. But here’s the thing—not everyone wants Stripe. Some coworking spaces prefer PayPal. Others need local payment gateways. (And some, bless their hearts, are still figuring out what a payment gateway even is.)
The solution? Let WooCommerce handle payments. Hundreds of gateway integrations, tax calculations, order management—all built-in.
Claude followed my implementation workflow perfectly.
PERFECTLY.
The settings page looked gorgeous:
There’s even a Product Sync panel showing 3 published plans synced to WooCommerce at 100% progress. One hundred percent!
My plans:
Hot Desk ($199/month),
Dedicated Desk ($399/month),
Private Office ($799/month)
—all published and ready to go:
And look!
They synced perfectly to WooCommerce products:
Everything looked GREAT.
So I clicked “Get Started” on the Hot Desk plan to test the checkout flow. You know, like a responsible developer would do. (Stop laughing.)
And here’s what I saw:
The old Stripe checkout.
The direct integration I was trying to REPLACE.
I switched the payment mode to WooCommerce. I synced the products. Everything in the admin looked correct.
But the frontend? Still using the old Stripe integration.
Ferpetesake.
.
.
.
Why Claude Thinks “Done” When It’s Really “Done-ish”
Here’s where I went full detective mode.
I checked the codebase. The WooCommerce checkout code exists. Functions written. Hooks registered. File paths correct. All present and accounted for.
So why wasn’t it working?
The code was never connected to the rest of the system.
(Stay with me here.)
Claude wrote the WooCommerce checkout handler. Beautiful code. But the pricing page? Still calling the old Stripe checkout function. The new code sat there—perfectly written, completely unused—like a fancy espresso machine you forgot to plug in.
And here’s the thing: this happens ALL THE TIME with AI-generated code.
Claude writes features.
It creates files. It generates functions. And in its summary, it reports “Task complete.”
But “code exists” and “code works”?
Two very different things.
You’ve probably experienced this.
Claude builds a feature. You test it. Something’s broken. You point out the bug. Claude apologizes (so polite!), fixes that specific issue, and introduces two new ones.
The Reddit crowd calls this “nerfed” or “lazy.”
They’re wrong.
👉 Claude lacks visibility into whether its code actually runs correctly in your system.
It can’t see the browser. It can’t watch a user click through your checkout flow. It can’t verify that function A actually calls function B in production.
The fix? Give Claude the ability to test its own work.
(This is where it gets good.)
.
.
.
The Most Important Testing? Not What You Think
You might be thinking: “Just write unit tests. Problem solved.”
And look—unit tests help. Integration tests help more.
But here’s what nobody talks about:
Perfect code doesn’t mean a working product.
The WooCommerce checkout code passed every logical check. Functions syntactically correct. Hooks properly registered. A unit test would have given it a gold star and a pat on the head.
But the pricing page template still imported the old Stripe checkout URL.
That’s a wiring problem. Not a code problem.
The test that catches this? User acceptance testing.
Actual users (or something simulating actual users) verifying the end product meets their needs. Clicking buttons. Filling forms. Going through the whole dang flow.
This is exactly why my implementation workflow generates a test plan BEFORE the implementation plan. The test plan represents success criteria from the user’s perspective:
Can a user switch payment modes?
Does the checkout redirect to WooCommerce?
Does the order confirmation show correct details?
These questions can’t be answered by reading code. They require clicking through the actual interface.
Which brings us to Ralph Loop.
.
.
.
Meet Ralph Loop: Your Autonomous Claude Code Testing Loop
Here’s the workflow I use to make Claude test its own work:
This is an autonomous loop that picks up a test case, executes it in an actual browser, checks against acceptance criteria, logs results, and repeats. If a test fails? Claude fixes the code and retests.
(Yes, really. It fixes its own bugs. I’ll show you.)
The core insight: you can’t throw a vague prompt at an autonomous loop and expect magic. The loop needs structure.
Specifically, it needs:
A test plan defining every test case upfront
A status.json tracking pass/fail for each case
A results.md where Claude logs learnings after each iteration
Let me show you exactly how I set this up for Claude Code testing.
.
.
.
1. Create the Ralph Test Folder
First, create a folder to store all your Ralph loop files:
Four files. That’s it.
prepare.md — Instructions for generating the status.json from your test plan
prompt.md — The loop instructions Claude follows each iteration
status.json — Tracks the state of all test cases (starts empty)
results.md — Human-readable log of each iteration (starts empty)
2. The Prepare Prompt
The prepare prompt tells Claude how to read your test plan and initialize the status file:
PROMPT: Ralph Loop Testing Agent (Prepare prompt)
Read the test plan file and generate a `status.json` file with all test cases initialized.
## Input
- **Test Plan:** `notes/test_plan.md`
## Output
- **Status File:** `notes/ralph_test/status.json`
## Instructions
1. Read the test plan markdown file
2. Extract ALL test cases (format: TC-XXX)
3. For each test case, extract:
- TC ID (e.g., "TC-501")
- Name (the test case title after the TC ID)
- Priority (Critical/High/Medium/Low)
4. Generate a JSON file with this exact structure:
```json
{
"metadata": {
"testPlanSource": "notes/test_plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": <count>,
"pending": <count>,
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"TC-XXX": {
"name": "Test case name from plan",
"priority": "Critical|High|Medium|Low",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
````
5. Save the file to the output path
## Extraction Rules
- Test case IDs follow pattern: `TC-NNN` (e.g., TC-501, TC-522)
- Test case names are in headers like: `#### TC-501: Checkout Header Display`
- Priority is usually listed in the test case details or status tracker table
- If priority not found, default to "Medium"
## Example
Input (from test plan):
```markdown
#### TC-501: Checkout Header Display
**Priority:** High
...
#### TC-502: Checkout Elements - Step Progress Display
**Priority:** High
...
```
Output (status.json):
```json
{
"metadata": {
"testPlanSource": "./docs/test-plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": 2,
"pending": 2,
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"TC-501": {
"name": "Checkout Header Display",
"priority": "High",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
},
"TC-502": {
"name": "Checkout Elements - Step Progress Display",
"priority": "High",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
## Validation
After generating, verify:
- [ ] All TC-XXX IDs from the test plan are included
- [ ] No duplicate TC IDs
- [ ] Summary.total matches count of testCases
- [ ] JSON is valid (no syntax errors)
- [ ] File saved to correct path
````
Read the test plan file and generate a `status.json` file with all test cases initialized.
## Input
- **Test Plan:** `notes/test_plan.md`
## Output
- **Status File:** `notes/ralph_test/status.json`
## Instructions
1. Read the test plan markdown file
2. Extract ALL test cases (format: TC-XXX)
3. For each test case, extract:
- TC ID (e.g., "TC-501")
- Name (the test case title after the TC ID)
- Priority (Critical/High/Medium/Low)
4. Generate a JSON file with this exact structure:
```json
{
"metadata": {
"testPlanSource": "notes/test_plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": <count>,
"pending": <count>,
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"TC-XXX": {
"name": "Test case name from plan",
"priority": "Critical|High|Medium|Low",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
````
5. Save the file to the output path
## Extraction Rules
- Test case IDs follow pattern: `TC-NNN` (e.g., TC-501, TC-522)
- Test case names are in headers like: `#### TC-501: Checkout Header Display`
- Priority is usually listed in the test case details or status tracker table
- If priority not found, default to "Medium"
## Example
Input (from test plan):
```markdown
#### TC-501: Checkout Header Display
**Priority:** High
...
#### TC-502: Checkout Elements - Step Progress Display
**Priority:** High
...
```
Output (status.json):
```json
{
"metadata": {
"testPlanSource": "./docs/test-plan.md",
"totalIterations": 0,
"maxIterations": 50,
"startedAt": null,
"lastUpdatedAt": null,
"summary": {
"total": 2,
"pending": 2,
"pass": 0,
"fail": 0,
"knownIssue": 0
}
},
"testCases": {
"TC-501": {
"name": "Checkout Header Display",
"priority": "High",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
},
"TC-502": {
"name": "Checkout Elements - Step Progress Display",
"priority": "High",
"status": "pending",
"fixAttempts": 0,
"notes": "",
"lastTestedAt": null
}
},
"knownIssues": []
}
```
## Validation
After generating, verify:
- [ ] All TC-XXX IDs from the test plan are included
- [ ] No duplicate TC IDs
- [ ] Summary.total matches count of testCases
- [ ] JSON is valid (no syntax errors)
- [ ] File saved to correct path
````
Key elements:
Points to your test plan location (notes/test_plan.md)
Specifies the output file (notes/ralph_test/status.json)
Defines the JSON structure with metadata and test case tracking
Nothing fancy. Just clear instructions.
3. The Loop Prompt
The prompt.md file contains the instructions Claude follows every single iteration:
PROMPT: Ralph Loop Testing Agent (Execution Prompt)
You are a testing agent in an iterative loop. Each iteration:
1. Read state from files
2. Pick ONE test case
3. Execute or fix it
4. Update state files
5. Check if done → output completion promise OR continue
**You have NO memory of previous iterations.** Files are your memory.
---
## Files to Read FIRST
| File | Purpose |
| ------------------------------ | ----------------------------------------------- |
| `notes/ralph_test/status.json` | Current state of all test cases (JSON) |
| `notes/test_plan.md` | Full test plan with steps and expected outcomes |
| `notes/ralph_test/results.md` | Human-readable log (append results here) |
**Optional context:**
- `notes/impl_plan.md` — Implementation details
- `notes/specs.md` — Specifications details
---
## Environment
wp-env is running:
- Dev site: http://localhost:8101
- Test site: http://localhost:8102
- Admin: http://localhost:8101/wp-admin
Commands:
- Run inside sandbox: Standard commands
- Run outside sandbox: npm, docker, wp-env commands
## Test Credentials
### Admin
- URL: http://localhost:8101/wp-admin
- Username: admin
- Email: wordpress@example.com
- Password: password
### Reset Password (if needed)
```bash
wp user update admin --user_pass=password
```
---
## This Iteration
### Step 1: Read State
Read the test status JSON file. Understand:
- Which test cases exist
- Status of each: `pending`, `testing`, `pass`, `fail`, `known_issue`
- Fix attempts for failing tests
### Step 2: Check Completion
**If ALL test cases are `pass` or `known_issue`:**
Output completion promise and final summary:
<promise>ALL_TESTS_RESOLVED</promise>
Summary:
- Total passed: X
- Known issues: Y
- Recommendations: ...
**Otherwise, continue to Step 3.**
### Step 3: Pick ONE Test Case
Priority order:
1. `testing` — Continue mid-test
2. `fail` with `fixAttempts < 3` — Needs fix
3. `pending` — Fresh test
### Step 4: Execute Test
Using Chrome browser automation (natural language):
- Navigate to URLs
- Click buttons/links
- Fill form inputs
- Take screenshots
- Read console logs
- Verify DOM state
**Follow the test plan click-path EXACTLY.**
### Step 5: Record Result
Update test status JSON:
**PASS:**
```json
{ "status": "pass", "notes": "What was verified", "lastTestedAt": "ISO timestamp" }
```
**FAIL:**
```json
{ "status": "fail", "fixAttempts": <increment>, "notes": "What failed", "lastTestedAt": "ISO timestamp" }
```
Update `metadata.totalIterations` and `metadata.lastUpdatedAt`.
### Step 6: Handle Failures
**If FAIL and fixAttempts < 3:**
- Analyze root cause
- Implement fix in codebase
- Next iteration will re-test
**If FAIL and fixAttempts >= 3:**
- Set status to `known_issue`
- Add to `knownIssues` array with: id, description, steps, severity
### Step 7: Update Human Log
Append to test results markdown:
```markdown
## Iteration [N] — [TIMESTAMP]
**TC:** TC-XXX — [Name]
**Status:** ✅/❌/⚠️
**Notes:** [What happened]
---
```
### Step 8: Continue or Complete
- If all TCs resolved → Output `<promise>ALL_TESTS_RESOLVED</promise>`
- Otherwise → Continue working (loop will restart)
---
## Rules
1. ONE test case per iteration
2. Update files BEFORE finishing
3. Follow test steps EXACTLY
4. Screenshot key verification points
5. Max 3 fix attempts → then known_issue
6. Output promise ONLY when truly complete
You are a testing agent in an iterative loop. Each iteration:
1. Read state from files
2. Pick ONE test case
3. Execute or fix it
4. Update state files
5. Check if done → output completion promise OR continue
**You have NO memory of previous iterations.** Files are your memory.
---
## Files to Read FIRST
| File | Purpose |
| ------------------------------ | ----------------------------------------------- |
| `notes/ralph_test/status.json` | Current state of all test cases (JSON) |
| `notes/test_plan.md` | Full test plan with steps and expected outcomes |
| `notes/ralph_test/results.md` | Human-readable log (append results here) |
**Optional context:**
- `notes/impl_plan.md` — Implementation details
- `notes/specs.md` — Specifications details
---
## Environment
wp-env is running:
- Dev site: http://localhost:8101
- Test site: http://localhost:8102
- Admin: http://localhost:8101/wp-admin
Commands:
- Run inside sandbox: Standard commands
- Run outside sandbox: npm, docker, wp-env commands
## Test Credentials
### Admin
- URL: http://localhost:8101/wp-admin
- Username: admin
- Email: wordpress@example.com
- Password: password
### Reset Password (if needed)
```bash
wp user update admin --user_pass=password
```
---
## This Iteration
### Step 1: Read State
Read the test status JSON file. Understand:
- Which test cases exist
- Status of each: `pending`, `testing`, `pass`, `fail`, `known_issue`
- Fix attempts for failing tests
### Step 2: Check Completion
**If ALL test cases are `pass` or `known_issue`:**
Output completion promise and final summary:
<promise>ALL_TESTS_RESOLVED</promise>
Summary:
- Total passed: X
- Known issues: Y
- Recommendations: ...
**Otherwise, continue to Step 3.**
### Step 3: Pick ONE Test Case
Priority order:
1. `testing` — Continue mid-test
2. `fail` with `fixAttempts < 3` — Needs fix
3. `pending` — Fresh test
### Step 4: Execute Test
Using Chrome browser automation (natural language):
- Navigate to URLs
- Click buttons/links
- Fill form inputs
- Take screenshots
- Read console logs
- Verify DOM state
**Follow the test plan click-path EXACTLY.**
### Step 5: Record Result
Update test status JSON:
**PASS:**
```json
{ "status": "pass", "notes": "What was verified", "lastTestedAt": "ISO timestamp" }
```
**FAIL:**
```json
{ "status": "fail", "fixAttempts": <increment>, "notes": "What failed", "lastTestedAt": "ISO timestamp" }
```
Update `metadata.totalIterations` and `metadata.lastUpdatedAt`.
### Step 6: Handle Failures
**If FAIL and fixAttempts < 3:**
- Analyze root cause
- Implement fix in codebase
- Next iteration will re-test
**If FAIL and fixAttempts >= 3:**
- Set status to `known_issue`
- Add to `knownIssues` array with: id, description, steps, severity
### Step 7: Update Human Log
Append to test results markdown:
```markdown
## Iteration [N] — [TIMESTAMP]
**TC:** TC-XXX — [Name]
**Status:** ✅/❌/⚠️
**Notes:** [What happened]
---
```
### Step 8: Continue or Complete
- If all TCs resolved → Output `<promise>ALL_TESTS_RESOLVED</promise>`
- Otherwise → Continue working (loop will restart)
---
## Rules
1. ONE test case per iteration
2. Update files BEFORE finishing
3. Follow test steps EXACTLY
4. Screenshot key verification points
5. Max 3 fix attempts → then known_issue
6. Output promise ONLY when truly complete
This is crucial for Claude Code testing to work properly.
Each iteration, Claude:
Reads status.json to understand current state
Picks the next pending test case
Executes the test in an actual browser
Updates status.json and results.md
Ends the iteration (which triggers the next loop)
Rinse. Repeat. Until done.
4. Initialize the Status File
Run the prepare prompt to generate your starting state:
Claude reads your test plan and creates status.json with all 38 test cases initialized:
The generated status file looks like this:
Every test case has:
status: “pending”, “pass”, “fail”, or “knownIssue”
fixAttempts: How many times Claude tried to fix this case
notes: What Claude observed during testing
lastTestedAt: Timestamp of the last test
All 38 tests. Ready to go. Pending status across the board.
The --completion-promise tells Ralph to keep looping until Claude outputs “ALL_TESTS_RESOLVED.”
The --max-iterations prevents infinite loops. (Because nobody wants that.)
6. Watch Claude Test Its Own Work
Claude starts by reading the state files to understand the current status:
It picks TC-001: Display Payment Mode Settings.
Then it launches a browser—an actual browser!—navigates to the settings page, and verifies each requirement:
All checks pass. TC-001: PASS ✅
(Look at all those green checkmarks. Gorgeous.)
Claude updates the status file:
Then updates results.md with a human-readable log:
Notice the stop hook at the bottom: “Ralph iteration 2.”
The loop automatically triggers the next iteration.
No manual intervention.
No babysitting.
Just… Claude Code testing itself.
.
.
.
7. The Loop Continues (Without You)
Iteration 2 starts.
Claude reads the state (1 pass, 37 pending), picks TC-002:
TC-002 requires WooCommerce to be deactivated.
So what does Claude do? Runs wp plugin deactivate woocommerce, then tests the settings page behavior.
The test passes—the WooCommerce option correctly shows “Setup Required” when the plugin is inactive:
Claude reactivates WooCommerce and updates the status:
And appends to results.md:
Iteration 2 complete.
Stop hook triggers iteration 3:
This continues automatically. Test after test after test.
You could go make coffee. Take a walk. Do your taxes.
(Okay, maybe not taxes.)
.
.
.
8. When Tests Fail, Claude Fixes Them
HERE’S where Ralph Loop really shines.
During testing, Claude encounters a failing test. The pricing page isn’t displaying plan prices correctly.
Does it give up? Does it log “FAIL” and move on?
Nope.
Claude investigates, finds the issue—the template is using old meta keys instead of the Plan model—and fixes it:
Then Claude retests to verify the fix worked:
The pricing page now shows correct prices. Claude clicks “Get Started” to continue testing the checkout flow.
Test. Find bug. Fix bug. Retest. Confirm fix.
All automatic.
.
.
.
9. All Tests Pass
After 3 hours and 32 minutes, all 38 test cases resolve:
Summary of Test Results:
Payment Mode Configuration: 5 tests ✅
Product & Plan Sync: 2 tests ✅
Checkout Flow: 6 tests ✅
Subscription Lifecycle: 6 tests ✅
Renewal Processing: 2 tests ✅
Plan Changes: 6 tests ✅
Cancellation: 2 tests ✅
Member Portal: 4 tests ✅
Admin Features: 2 tests ✅
Emails: 2 tests ✅
Security: 1 test ✅
Total: 38 tests. All passing.
The critical P0/P1 tests that Claude fixed during the loop:
TC-004: Mode Switch Blocking ✅
TC-009: Guest Checkout Prevention ✅ (with fix)
TC-016: Race Condition Prevention ✅
TC-018: Pre-Renewal Token Validation ✅
TC-020: 3D Secure Handling ✅
TC-038: Token Ownership Validation ✅
HECK YES.
.
.
.
The Proof: It Actually Works Now
Remember the checkout problem from the beginning? The one that made me question my life choices?
Let’s see what happens now.
The pricing page displays correctly:
Click “Get Started” on Hot Desk, and you’re redirected to the WooCommerce checkout:
See the difference?
This is the WooCommerce checkout page.
The order summary shows “Hot Desk” with “Billing Cycle: Monthly.” The account creation notice appears because subscriptions require accounts.
(This is the moment I did a small victory dance. Don’t judge.)
Scroll down to payment options—Stripe through WooCommerce:
The Stripe integration now runs through WooCommerce. Same payment processor, but managed by WooCommerce’s subscription system. I can swap in PayPal, Square, or any other gateway without touching theme code.
Complete the purchase, and you land on the welcome page:
Everything works.
The flow connects end-to-end.
The WooCommerce integration that Claude “completed” previously?
Now it’s actually complete.
.
.
.
The Complete Journey: From Idea to Working Product
Let me zoom out and show you how all three parts of this series connect:
Phase 1: Bulletproof Specs
We started by brainstorming comprehensive specifications.
Using the AskUserQuestion tool, Claude asked 12 clarifying questions covering everything from subscription handling to checkout experience to refund policies. Then Claude critiqued its own specs, finding 14 potential issues before we wrote any code.
Phase 2: Test Plan
Before implementation, we generated a test plan.
38 test cases defining exactly what success looks like—from a user’s perspective. These became our acceptance criteria.
Phase 3: Implementation Plan + Sub-Agents
We created an implementation plan mapping tasks to test cases. Then executed with sub-agents running in parallel waves, keeping context usage low while building everything in 52 minutes.
Phase 4: Claude Code Testing + Fixing with Ralph Loop
Finally, we let Ralph loose. The autonomous loop tested each case in an actual browser, found the bugs Claude missed during implementation, fixed them, and verified the fixes.
3 hours 32 minutes later: 38/38 tests passing.
.
.
.
What I’ve Learned About Building With AI
Here’s what this whole journey taught me.
We all want AI to one-shot solutions on the first try. To type a prompt, hit enter, and watch magic happen. And when it doesn’t work perfectly? We blame the AI. Call it nerfed. Call it lazy. Move on to the next shiny tool.
But here’s the thing I keep coming back to:
Even the most experienced developer can’t one-shot a complex feature.
We write code. Test it. Find bugs. Fix them. Test again. That’s just how building software works. Always has been. Probably always will be.
AI is no different.
The breakthrough—the real breakthrough—comes from giving AI the ability to verify its own work. The same way any developer does. Write the code. Test it against real user scenarios. See what breaks. Fix it. Test again.
Ralph Loop makes this autonomous.
You don’t have to manually test 38 scenarios. You don’t have to spot the bugs yourself. You don’t have to describe each fix.
You define success criteria upfront (test plan), give Claude the ability to test against those criteria (browser automation), and let it iterate until everything passes.
👉 That’s the entire secret: structured iteration with clear success criteria.
Not smarter prompts. Not better models. Not more tokens.
Just… iteration.
The same boring, unsexy process that’s always made software work.
Being the organized person I am, I created the most detailed inventory list you’ve ever seen. Every box labeled. Every item cataloged. “Kitchen – Plates (12), Bowls (8), That Weird Garlic Press I Never Use But Can’t Throw Away (1).”
I handed this masterpiece to the movers and said, “Here you go!”
They looked at me like I’d handed them a grocery list in Klingon.
Because here’s what my beautiful inventory didn’t tell them: Which boxes go to which rooms. What order to load things. Which items are fragile. What depends on what. The fact that the bookshelf needs to go in before the desk, or nothing fits.
They weren’t wrong to be confused. I’d given them a comprehensive what without any how.
This is exactly what happens when you hand Claude Code your bulletproof specs and say “implement this.”
You followed the process. You answered every clarifying question. You had Claude critique its own work. Your specs are comprehensive—2,000+ lines of detailed requirements, edge cases, and architectural decisions.
Now you’re ready to build.
So you fire up Claude Code and type: “Read the specs and implement it.”
Claude starts working. Files appear. Code flows.
And thirty minutes later? Half your edge cases are missing. The checkout flow doesn’t match what you specified. That critical race condition prevention you spent three rounds of Q&A perfecting?
Nowhere to be found.
Here’s the thing: Comprehensive specs don’t automatically translate to comprehensive implementation.
Your specs might be 2,000 lines. Claude’s context window is limited. As implementation progresses, early requirements fade from memory. The AI starts making shortcuts. Details slip through the cracks like sand through fingers.
Sound familiar?
(If you’re nodding right now, stay with me.)
The issue isn’t Claude’s capability. It’s the gap between what’s documented and what gets built. Even human developers working from perfect documentation miss things. They get tired. They make assumptions. They interpret requirements their own way.
Claude faces the same challenges—plus context limits that force it to work with only a subset of information at any given moment.
👉 The solution isn’t better prompting. It’s better process.
And that process? It’s what I’ve been calling the Claude Code implementation workflow. Let me show you what I mean.
.
.
.
The Missing Middle Layer
Here’s what most people do:
I call this the “hope-based development methodology.”
(That’s a joke. Please don’t actually call it that.)
Here’s what actually works:
The test plan answers: “How will we know if each requirement is implemented correctly?”
The implementation plan answers: “What specific tasks need to happen, and in what order?”
The task management answers: “How do we keep Claude focused and prevent context overload?”
Think of it like this: your specs are the inventory list. The test plan is the “here’s how we’ll know each box arrived safely” checklist. The implementation plan is the “which room, which order, what depends on what” instruction sheet.
The movers—er, Claude—can actually do their job now.
Let me walk you through exactly how this Claude Code implementation workflow… well, works.
.
.
.
Step 1: Create a Test Plan From Your Specs
Before writing any code, Claude needs to understand what success looks like.
Now, I know what you’re thinking: “Wait—isn’t a test plan for after implementation?”
Traditionally, yes. But for AI-driven development, creating the test plan first serves a completely different purpose.
👉 It forces Claude to deeply analyze every requirement and translate it into verifiable outcomes.
When Claude creates test cases for “handle race conditions during renewal processing,” it has to think through exactly what that means. What are the preconditions? What actions trigger the behavior? What should the expected results be?
It’s like asking someone to write the exam questions before teaching the class. Suddenly, they understand the material much more deeply.
Here’s the prompt I use:
PROMPT: Create a Comprehensive Test Plan based on Specs
Create a comprehensive test plan that will verify the implementation matches the specs at `notes/specs.md`
### Step 1: Identify Test Scenarios
Based on the specs:
- Happy path flows
- Error conditions
- Edge cases
- State transitions
- Responsive behavior
- Accessibility requirements
### Step 2: Create Test Cases
For each scenario, create detailed test cases:
```markdown
### TC-NNN: [Test Name]
**Description:** [What this test verifies]
**Preconditions:**
- [Required state before test]
- [Required data]
**Steps:**
| Step | Action | Expected Result |
|------|--------|-----------------|
| 1 | [Action to take] | [What should happen] |
| 2 | [Action to take] | [What should happen] |
**Test Data:**
- Field 1: `value`
- Field 2: `value`
**Expected Outcome:** [Final verification]
**Priority:** Critical / High / Medium / Low
````
### Step 3: Organize by Category
Group test cases:
- Functional tests
- UI/UX tests
- Validation tests
- Integration tests (if applicable)
- Edge case tests
### Step 4: Create Status Tracker
```markdown
## Status Tracker
| TC | Test Case | Priority | Status | Remarks |
| ------ | --------- | -------- | ------ | ------- |
| TC-001 | [Name] | High | [ ] | |
| TC-002 | [Name] | Medium | [ ] | |
```
### Step 5: Add Known Issues Section
```markdown
## Known Issues
| Issue | Description | TC Affected | Steps to Reproduce | Severity |
| ----- | ----------- | ----------- | ------------------ | -------- |
| | | | | |
```
## Output
Save the test plan to: `notes/test_plan.md`
Include:
1. Overview and objectives
2. Prerequisites
3. Reference wireframe (if applicable)
4. Test cases (10-20 typically)
5. Status tracker
6. Known issues section
## Test Case Guidelines
- Each test should be independent
- Use specific, concrete test data
- Include both positive and negative tests
- Cover all screens from wireframe (if applicable)
- Test all states from prototype
- Consider mobile/responsive
## Do NOT
- Over-test obvious functionality
- Skip error handling tests
- Forget accessibility basics
Create a comprehensive test plan that will verify the implementation matches the specs at `notes/specs.md`
### Step 1: Identify Test Scenarios
Based on the specs:
- Happy path flows
- Error conditions
- Edge cases
- State transitions
- Responsive behavior
- Accessibility requirements
### Step 2: Create Test Cases
For each scenario, create detailed test cases:
```markdown
### TC-NNN: [Test Name]
**Description:** [What this test verifies]
**Preconditions:**
- [Required state before test]
- [Required data]
**Steps:**
| Step | Action | Expected Result |
|------|--------|-----------------|
| 1 | [Action to take] | [What should happen] |
| 2 | [Action to take] | [What should happen] |
**Test Data:**
- Field 1: `value`
- Field 2: `value`
**Expected Outcome:** [Final verification]
**Priority:** Critical / High / Medium / Low
````
### Step 3: Organize by Category
Group test cases:
- Functional tests
- UI/UX tests
- Validation tests
- Integration tests (if applicable)
- Edge case tests
### Step 4: Create Status Tracker
```markdown
## Status Tracker
| TC | Test Case | Priority | Status | Remarks |
| ------ | --------- | -------- | ------ | ------- |
| TC-001 | [Name] | High | [ ] | |
| TC-002 | [Name] | Medium | [ ] | |
```
### Step 5: Add Known Issues Section
```markdown
## Known Issues
| Issue | Description | TC Affected | Steps to Reproduce | Severity |
| ----- | ----------- | ----------- | ------------------ | -------- |
| | | | | |
```
## Output
Save the test plan to: `notes/test_plan.md`
Include:
1. Overview and objectives
2. Prerequisites
3. Reference wireframe (if applicable)
4. Test cases (10-20 typically)
5. Status tracker
6. Known issues section
## Test Case Guidelines
- Each test should be independent
- Use specific, concrete test data
- Include both positive and negative tests
- Cover all screens from wireframe (if applicable)
- Test all states from prototype
- Consider mobile/responsive
## Do NOT
- Over-test obvious functionality
- Skip error handling tests
- Forget accessibility basics
Claude reads through all the specs, identifies what needs to be tested, and generates a structured test plan.
For my WooCommerce integration, Claude created 38 test cases organized into 12 sections:
Notice the priority distribution:
Critical (P0): 7 tests — Must pass before deployment
High (P1): 20 tests — Essential functionality
Medium: 11 tests — Important but not blocking
Each test case maps directly to a requirement in my specs. Nothing ambiguous. Nothing assumed. Nothing left to interpretation.
(This is the part where past-me would have skipped ahead to coding. Don’t be past-me.)
.
.
.
Step 2: Create an Implementation Plan That Maps to Test Cases
Now Claude knows what success looks like. Next question: how do we get there?
The implementation plan bridges test cases to actual tasks. Every task links back to specific test cases it will satisfy. It’s the “what depends on what” instruction sheet for our movers.
PROMPT: Create an Implementation Plan That Maps to Test Cases
Specs is approved, test plan is ready. Now we need an implementation plan.
- **Specs**: `notes/specs.md`
- **Test Plan:** `notes/test_plan.md`
## Your Task
Create a detailed implementation plan that maps to the test cases.
### Step 1: Analyze Test Cases
For each test case (TC-NNN):
- What functionality must exist?
- What files need to be created/modified?
- What dependencies are needed?
### Step 2: Create Task Breakdown
Group test cases into implementation tasks:
```markdown
## Implementation Plan: [PHRASE_NAME]
### Overview
[Brief description]
### Files to Create/Modify
[List all files]
### Implementation Tasks
#### Task 1: [Name]
**Mapped Test Cases:** TC-001, TC-002, TC-003
**Files:**
- `path/to/file1.php` - [description]
- `path/to/file2.js` - [description]
**Implementation Notes:**
- [Key detail 1]
- [Key detail 2]
**Acceptance Criteria:**
- [ ] TC-001 passes
- [ ] TC-002 passes
- [ ] TC-003 passes
#### Task 2: [Name]
...
````
### Step 3: Identify Dependencies
- What from previous phrases is needed?
- What order should tasks be implemented?
- Any external dependencies?
### Step 4: Estimate Complexity
- Simple: 1-2 tasks, straightforward
- Medium: 3-5 tasks, some complexity
- Complex: 6+ tasks, significant work
## Output
Save the implementation plan to: `notes/impl_plan.md`
Include:
1. Overview
2. Files to create/modify
3. Tasks with TC mappings
4. Dependencies
5. Complexity estimate
## Guidelines
- Every test case must map to a task
- Tasks should be completable in one session
- Include enough detail to guide implementation
- Reference design system patterns
## Do NOT
- Include actual code (next step)
- Over-engineer simple features
Specs is approved, test plan is ready. Now we need an implementation plan.
- **Specs**: `notes/specs.md`
- **Test Plan:** `notes/test_plan.md`
## Your Task
Create a detailed implementation plan that maps to the test cases.
### Step 1: Analyze Test Cases
For each test case (TC-NNN):
- What functionality must exist?
- What files need to be created/modified?
- What dependencies are needed?
### Step 2: Create Task Breakdown
Group test cases into implementation tasks:
```markdown
## Implementation Plan: [PHRASE_NAME]
### Overview
[Brief description]
### Files to Create/Modify
[List all files]
### Implementation Tasks
#### Task 1: [Name]
**Mapped Test Cases:** TC-001, TC-002, TC-003
**Files:**
- `path/to/file1.php` - [description]
- `path/to/file2.js` - [description]
**Implementation Notes:**
- [Key detail 1]
- [Key detail 2]
**Acceptance Criteria:**
- [ ] TC-001 passes
- [ ] TC-002 passes
- [ ] TC-003 passes
#### Task 2: [Name]
...
````
### Step 3: Identify Dependencies
- What from previous phrases is needed?
- What order should tasks be implemented?
- Any external dependencies?
### Step 4: Estimate Complexity
- Simple: 1-2 tasks, straightforward
- Medium: 3-5 tasks, some complexity
- Complex: 6+ tasks, significant work
## Output
Save the implementation plan to: `notes/impl_plan.md`
Include:
1. Overview
2. Files to create/modify
3. Tasks with TC mappings
4. Dependencies
5. Complexity estimate
## Guidelines
- Every test case must map to a task
- Tasks should be completable in one session
- Include enough detail to guide implementation
- Reference design system patterns
## Do NOT
- Include actual code (next step)
- Over-engineer simple features
Claude analyzes both the specs and test plan, then generates a phased implementation plan:
The result: 4 phases, 12 implementation tasks, each explicitly linked to the test cases that will verify them.
Phase 1: Foundation & Configuration (2 tasks → TC-001 to TC-007)
Plus a critical path of P0/P1 items that must work before deployment:
Mode switch blocking when active subscriptions exist
Race condition prevention (double-charge protection—ferpetesake, the payments!)
Pre-renewal token validation
Token ownership security
Now we have specs, a test plan, AND an implementation plan. Three documents that all reference each other. A complete picture.
But here’s where most people (including past-me, again) would stumble.
.
.
.
Step 3: Execute With Sub-Agents (This Is Where It Gets Fun)
Here’s something I learned the hard way.
There’s research showing that LLM performance degrades as context size increases. When Claude’s context fills up with implementation details from Task 1, it starts losing precision on Task 8. It’s like asking someone to remember the first item on a grocery list after they’ve been shopping for an hour.
The fix: run each task in its own sub-agent.
Each sub-agent gets fresh context. It focuses on one task, implements it, and reports back. The orchestrating agent manages dependencies and progress. No context pollution. No forgotten requirements.
It’s like having a team of movers where each person is responsible for exactly one room—and they all have fresh energy because they haven’t been carrying boxes all day.
Here’s the prompt that kicks off the Claude Code implementation workflow execution:
PROMPT: Claude Code implementation workflow execution
We are executing the implementation plan. All design and planning is complete.
## Reference Documents
- **Specs:** @notes/specs.md
- **Implementation Plan:** @notes/impl_plan.md
- **Test Plan:** @notes/test_plan.md
## Phase 1: Task Creation
### Before Creating Tasks
1. Review @notes/impl_plan.md completely
2. Understand test case expectations from @notes/test_plan.md
3. Reference wireframe/prototype for UI (if applicable)
4. Check design system for patterns (if available)
### Create Tasks from Implementation Plan
Parse the implementation plan and use `TaskCreate` to create a task for each implementation item:
1. **Extract all tasks** from @notes/impl_plan.md
2. **Identify dependencies** between tasks (what must be done before what)
3. **Create each task** with:
- Clear description including the specific files to create/modify
- Mapped test cases (TCs) that verify the task
- `blocked_by`: tasks that must complete first
- `blocks`: tasks that depend on this one
Tasks should be granular enough to run independently but logical enough to represent complete units of work.
---
## Phase 2: Task Execution
### Execution Strategy
Execute tasks using sub-agents for parallel processing:
1. **Group tasks into waves** based on dependencies
2. **Run each task in its own sub-agent** - This keeps context usage low (~18% vs ~56%)
3. **Process waves sequentially** - Wave N+1 starts only after Wave N completes
### For Each Task (Sub-Agent Instructions)
1. Use `TaskGet` to read full task details
2. Create/modify specified files
3. Implement functionality to pass mapped TCs
4. **Self-verify the implementation:**
- Check that code compiles/runs without errors
- Verify the functionality matches test expectations
- Ensure design consistency with existing patterns
5. Use `TaskUpdate` to mark task complete with a brief summary of what was done
6. Note any deviations, concerns, or discovered issues
### If Issues Are Discovered
- Use `TaskCreate` to add new fix/bug tasks
- Set appropriate dependencies so fixes run in correct order
- Continue with other independent tasks
---
## Phase 3: Completion Summary
After all tasks are complete, provide:
### 1. Summary of Changes
- Files created
- Files modified
- Key functionality added
### 2. Self-Verification Results
- What works as expected
- Any concerns or edge cases noted
- Tasks that required fixes (if any)
### 3. Ready for Testing
- Confirm all tasks marked complete
- List any setup needed for testing
- Note any known limitations
---
## Important Notes
- **Do NOT run full test suite** - that's the next step
- **Use `TaskList`** periodically to check overall progress
- **Dependencies are critical** - ensure tasks don't start before their blockers complete
- **Keep sub-agent context focused** - each sub-agent only needs info for its specific task
We are executing the implementation plan. All design and planning is complete.
## Reference Documents
- **Specs:** @notes/specs.md
- **Implementation Plan:** @notes/impl_plan.md
- **Test Plan:** @notes/test_plan.md
## Phase 1: Task Creation
### Before Creating Tasks
1. Review @notes/impl_plan.md completely
2. Understand test case expectations from @notes/test_plan.md
3. Reference wireframe/prototype for UI (if applicable)
4. Check design system for patterns (if available)
### Create Tasks from Implementation Plan
Parse the implementation plan and use `TaskCreate` to create a task for each implementation item:
1. **Extract all tasks** from @notes/impl_plan.md
2. **Identify dependencies** between tasks (what must be done before what)
3. **Create each task** with:
- Clear description including the specific files to create/modify
- Mapped test cases (TCs) that verify the task
- `blocked_by`: tasks that must complete first
- `blocks`: tasks that depend on this one
Tasks should be granular enough to run independently but logical enough to represent complete units of work.
---
## Phase 2: Task Execution
### Execution Strategy
Execute tasks using sub-agents for parallel processing:
1. **Group tasks into waves** based on dependencies
2. **Run each task in its own sub-agent** - This keeps context usage low (~18% vs ~56%)
3. **Process waves sequentially** - Wave N+1 starts only after Wave N completes
### For Each Task (Sub-Agent Instructions)
1. Use `TaskGet` to read full task details
2. Create/modify specified files
3. Implement functionality to pass mapped TCs
4. **Self-verify the implementation:**
- Check that code compiles/runs without errors
- Verify the functionality matches test expectations
- Ensure design consistency with existing patterns
5. Use `TaskUpdate` to mark task complete with a brief summary of what was done
6. Note any deviations, concerns, or discovered issues
### If Issues Are Discovered
- Use `TaskCreate` to add new fix/bug tasks
- Set appropriate dependencies so fixes run in correct order
- Continue with other independent tasks
---
## Phase 3: Completion Summary
After all tasks are complete, provide:
### 1. Summary of Changes
- Files created
- Files modified
- Key functionality added
### 2. Self-Verification Results
- What works as expected
- Any concerns or edge cases noted
- Tasks that required fixes (if any)
### 3. Ready for Testing
- Confirm all tasks marked complete
- List any setup needed for testing
- Note any known limitations
---
## Important Notes
- **Do NOT run full test suite** - that's the next step
- **Use `TaskList`** periodically to check overall progress
- **Dependencies are critical** - ensure tasks don't start before their blockers complete
- **Keep sub-agent context focused** - each sub-agent only needs info for its specific task
I know. That’s a lot of prompt. But here’s the thing—you set this up once, and then you watch the magic happen.
Task Creation
First, Claude creates all 13 tasks from the implementation plan:
Then sets up dependencies between them. (This is the “bookshelf before desk” part.)
Wave-Based Execution
Before diving into implementation, Claude explores the existing codebase to understand patterns:
Wave 1 starts with 2 sub-agents running Tasks 1.1 and 1.2 in parallel:
Once Wave 1 completes, Wave 2 begins:
Wave 2 completes, Wave 3 (the critical phase) starts with 3 sub-agents:
Wave 3 completes, and Wave 4 launches with 6 sub-agents in parallel:
(Six sub-agents. Running in parallel. Each with fresh context. This is the future we were promised.)
Implementation Complete
All 13 tasks across 4 waves:
Total time: 52 minutes 22 seconds.
That’s 13 tasks. 38 test cases worth of functionality. Parallel execution keeping context usage low throughout.
Remember my WooCommerce integration specs? The 2,000+ lines that would have turned into a context-overloaded mess if I’d just said “implement this”?
Every requirement addressed. Every edge case accounted for. Every critical feature implemented.
.
.
.
The Complete Claude Code Implementation Workflow
Let me put this all together:
Step 1: Create a test plan from your specs
Claude analyzes every requirement
Generates verifiable test cases
Establishes priority levels
Step 2: Create an implementation plan mapped to test cases
Groups requirements into logical tasks
Links each task to specific test cases
Identifies dependencies between tasks
Step 3: Execute with sub-agents
Create all tasks with dependency tracking
Process in waves based on blocking relationships
Each sub-agent gets fresh context
Parallel execution where dependencies allow
It’s the complete Claude Code implementation workflow—from inventory list to fully unpacked apartment. (Okay, I’ll stop with the moving metaphor now.)
.
.
.
But Wait. There’s a Catch.
You followed the workflow. Claude built everything in 52 minutes. Your implementation summary shows all 38 test cases “implemented.”
Here’s the uncomfortable truth I need to share with you:
“Implemented” and “working correctly” aren’t the same thing.
Claude sometimes takes shortcuts. Features get partially built. Edge cases get acknowledged in comments but not actually handled. The sub-agent said “Done”—but did it really do what the test case required?
You need a verification step. A way to systematically check that every test case actually passes. A process for catching the gaps before you discover them in production.
You hire a contractor to “update the bathroom.” You wave your hand vaguely at the space, mention something about “modern vibes,” and leave for work.
You come home to discover they’ve ripped out your vintage clawfoot tub—the one you specifically loved—and replaced it with a walk-in shower. They’ve chosen gray tiles. (Gray! When you’re clearly a warm terracotta person.) And oh, they’ve moved the toilet to a spot that now requires you to shimmy sideways to close the door.
“But you said you wanted it updated,” they say, genuinely confused.
Here’s the thing: they weren’t wrong. You did say that.
You just forgot to mention all the things living rent-free in your head—the constraints, the preferences, the “obviously don’t touch THAT” items you assumed were, well, obvious.
This is exactly what happens when you ask Claude Code to add a feature to your existing app.
.
.
.
The “Just Add This Feature” Trap
You’ve got a working application.
It’s not perfect, but it works. Now you want to add something—a new feature, a refactor, maybe squash a bug that’s been giving you side-eye for weeks.
So you fire up Claude Code and describe what you want in two or three sentences.
Claude gets to work. Files change. Functions appear. Code flows.
And then—record scratch—you realize it completely missed the point.
Not because Claude is bad at its job. But because you gave it the software equivalent of “update the bathroom.”
Your codebase has history.
Conventions. Edge cases you’ve already wrestled to the ground. Payment integrations that absolutely, positively cannot break (ferpetesake, the payments!). Claude needs to understand ALL of this before writing a single line of code.
This is where Claude Code spec-driven development comes in.
And honestly?
It’s transformed how I approach every feature request.
.
.
.
Why Spec-Driven Development Changes Everything
Here’s what most people do: describe the feature → let Claude code → discover problems → fix problems → discover MORE problems → question life choices.
Here’s what works: describe the feature → trigger clarifying questions → answer EVERY question → generate comprehensive specs → review specs → THEN code.
The difference isn’t just efficiency (though yes, that too).
👉 The real magic is that Claude surfaces edge cases YOU haven’t thought about.
All those assumptions living in your head? Claude will ask about them. All those “obviously we’d handle it THIS way” decisions? Claude will make you choose explicitly.
It’s like having a brilliant—if occasionally pedantic—architect who refuses to pick up a hammer until every single detail is documented.
Let me show you exactly how this works.
.
.
.
The 3-Phase Method for Claude Code Spec-Driven Development
I’ve been refining this workflow for weeks, and it breaks down into three distinct phases.
(Stay with me—I promise this is more exciting than it sounds. Okay, maybe “exciting” is strong. But genuinely useful? HECK YES.)
.
.
.
Phase 1: Build Your Specs
This is where you go from “vague idea in your head” to “comprehensive document that leaves nothing to chance.”
Step 1: Describe the Task + Trigger the Magic Words
Here’s the prompt structure that makes everything work:
[Describe your task/feature in detail]
Based on the tasks above, write the specs at @notes/specs.md.
Use ASCII diagrams where necessary to illustrate the UI/UX.
Use the AskUserQuestion tool to Ask me clarifying questions until you are 95% confident you can complete this task successfully. For each question, add your recommendation (with reason why) below the options. This would help me in making a better decision.
That last paragraph? That’s the secret sauce.
“Ask me clarifying questions until you are 95% confident” tells Claude not to make assumptions. Not to guess. Not to fill in blanks with whatever seems reasonable.
Instead, Claude will ASK.
The initial prompt triggering Claude’s question-asking mode
Step 2: Let Claude Explore First
Here’s something beautiful: Claude doesn’t immediately pepper you with questions.
First, it explores your codebase. It reads your existing implementation. It studies your patterns, your database schema, your webhook handlers.
In my WooCommerce integration project, Claude spent about 2 minutes analyzing the existing Stripe integration—36 tool uses, 97.5k tokens of context-gathering—before asking its first question.
Then the questions began:
Claude explores codebase first, then asks its first clarifying question about subscription handling approach
Notice what’s happening here (and this is the part that made me unreasonably happy):
Claude presents clear options—not vague open-ended questions
Each option includes an explanation of implications
Claude adds its own recommendation with reasoning
You can select an option OR type something custom
That recommendation is EVERYTHING. You’re not making decisions in a vacuum. You’re evaluating Claude’s expert analysis against your own context.
Step 3: Keep Answering Until Claude Hits 95%
After your first answer, Claude asks another question. Then another.
Second question about renewal payment flows with three implementation options
Each question builds on your previous answers. Claude’s understanding deepens with every response.
Third question about checkout UX with three different approaches
“But wait,” you might be thinking. “Isn’t this just Plan Mode with extra steps?”
Ah. Good question. (I asked it too.)
Plan Mode asks one batch of questions, then starts planning.
This approach keeps questioning until Claude reaches genuine confidence. If something’s unclear—or conflicts with an earlier answer—Claude asks follow-up questions. It’s more thorough. Sometimes annoyingly so. But annoyingly thorough beats catastrophically incomplete every time.
Step 4: Review the Complete Q&A
By the end of my WooCommerce project, Claude had asked 15 different questions covering:
Subscription plugin choice
Renewal flow mechanics
Checkout experience
Payment mode switching
Migration strategy
Product mapping
Refund handling
Member portal UI
Day pass handling
Coupon systems
Settings location
Email notifications
First half of the complete Q&A session—eight questions answeredSecond half of Q&A plus the requirements summary Claude compiled
Every edge case I hadn’t considered? Surfaced as a question.
Every architectural decision that could bite me later? Addressed before a single line of code.
👉 This is the core of Claude Code spec-driven development: front-loading decisions instead of discovering them mid-implementation.
Step 5: Watch Your Specs Materialize
Once Claude hit 95% confidence, it summarized everything and started writing:
Claude’s confidence summary showing 15 requirements before writing specs
The result? A 2,068-line specification document with:
Architecture decisions (documented!)
Database schema changes (planned!)
State machines for subscriptions (visualized!)
ASCII diagrams for UI mockups (actually helpful!)
Detailed flow documentation (comprehensive!)
Edge case handling (the stuff that usually bites you)
Specs created—2068 lines covering architecture, checkout flow, renewals, and moreSpecs include ASCII diagrams for architecture, checkout flow, state machines, and UI mockups
Two thousand lines might sound excessive. It’s not. It’s every decision you’d otherwise make on the fly—except now they’re documented, reviewable, and consistent.
Bonus Step: Split Large Specs
Quick practical note: if your specs exceed ~1,500 lines, split them.
Claude’s file reading has a 25k token limit. Large files cause problems during implementation. The fix is simple:
This specs file is too big. Please split it into 3 parts.
Keep the original file as the index file that links to the 3 parts.
Asking Claude to split the large specs file into manageable parts
Claude reorganizes everything into digestible chunks:
Result—four files total with specs.md as navigation hub linking to three focused documents
Your specs might be comprehensive, but they could still contain conflicts, gaps, or requirements that are—how do I put this diplomatically—completely impossible to implement.
Step 6: Start Fresh and Ask Claude to Critique
Start a new session. Fresh context. Then ask Claude to review its own work with skeptical eyes:
Read the `notes/specs.md`, analyze the codebase and tell me what could be
the potential problems with this specs.
Let's look at it from different angles and try to consider every edge cases.
Use ASCII diagrams to illustrate if needed. Don't edit anything yet.
Let's focus on analysis.
The review prompt—fresh session, analysis-only mode
That last line—”Don’t edit anything yet”—is important. You want analysis first. Changes second. Mixing them leads to chaos.
Step 7: Receive Your Brutally Honest Analysis
Claude reads the specs, re-explores your codebase, and delivers feedback that might sting a little:
Claude identifies Problem #1—dual source of truth conflicts during mode transitions
The analysis includes ASCII diagrams showing exactly where problems would occur:
Problem #2—payment token lifecycle gaps with three failure scenarios illustrated
In my project, Claude found 14 potential issues, organized by severity:
The full prioritized issue list—P0 Critical down to P3 Lo
P0 – Critical (the “oh no” tier):
Race condition in renewal processing → Double charges 😱
Token deletion not detected → Failed renewals
Mode switch with active Stripe subs → Billing confusion
3D Secure on unattended renewals → Payment failures
Would I have caught all these myself?
Honestly? Maybe half. And probably not until they caused real problems with real users and real money.
👉 Having Claude review its own specs is like getting a code review before the code exists.
.
.
.
Phase 3: Refine Your Specs
Now you have a prioritized list of issues. Time to fix them—but not by guessing.
Step 8: Address Issues With Another Round of Questions
Ask Claude to fix the critical issues, but trigger the questioning mode AGAIN:
Let's address all the P0 and P1 issues.
Update the specs file.
Use the AskUserQuestion tool to Ask me clarifying questions until you are 95% confident you can complete this task successfully. For each question, add your recommendation (with reason why) below the options.
Requesting fixes for critical and high-priority issues with question-asking enabled
Fixing these issues requires MORE decisions. Should you block mode switches entirely? Allow mixed modes temporarily? Auto-cancel subscriptions?
Claude asks:
Question about mode switch behavior—block, auto-cancel, or allow mixed modeQuestion about proration and coupon math during plan changes
Each answer shapes the solution. After this round, Claude updates the specs with concrete, specific fixes:
All P0 and P1 issues addressed with specific solutions and file locations
P0 Fixes:
Renewal Race Condition → Database row locking with SELECT FOR UPDATE
Token Deletion → Pre-renewal validation 7 and 3 days before renewal
Mode Switch Danger → Block switching if ANY active Stripe subscriptions exist
P1 Fixes:
Proration + Coupon → Calculate from actual paid amounts, keep coupon on upgraded plan
Guest Checkout → Force account creation for subscription products
3D Secure → Request off-session exemption, fallback to email payment link
You can repeat Phase 2 and Phase 3 as many times as needed. Review, refine, review, refine—until you’re genuinely confident in your specs.
.
.
.
The Method, Summarized
Because I know you’re going to want to reference this later:
Phase 1: Build Specs
Describe what you want + add the AskUserQuestion trigger
Let Claude explore your codebase first
Answer every clarifying question
Review the generated specs
Split large specs into parts (if needed)
Phase 2: Review Specs 6. Start a new session 7. Ask Claude to find problems (analysis only—no edits)
Phase 3: Refine Specs 8. Fix critical issues with another round of questions 9. Repeat Phases 2-3 until satisfied
YOU decide when specs are ready.
Not Claude.
You.
.
.
.
What About Implementation?
Ah. Yes.
You have these beautiful, bulletproof specs. Now what?
You could tell Claude to implement everything in one go. But for specs this comprehensive, that approach has problems:
Claude might skip features (context limits are real)
Long implementations drift from the plan
Details get forgotten mid-stream
The solution?
Claude Code’s task management system—which breaks large specs into trackable tasks, maintains progress across sessions, and ensures nothing slips through the cracks.
When I turned ten, my parents gifted me a battleship model kit.
You know the ones—those intricate naval vessels with approximately 1,000 tiny plastic pieces, decals thinner than a whisper, and instructions that assume you already have the steady hands of a neurosurgeon.
The box art was magnificent. A mighty warship cutting through ocean waves, every gun turret perfectly positioned, every railing impossibly detailed.
I tore it open that morning. Snapped together the first few pieces. The hull took shape. Then the deck. This was going to be amazing.
And then I hit the superstructure.
Tiny railings that snapped if you looked at them wrong. Parts that looked identical but absolutely weren’t. Decals that crumpled the moment I breathed near them. Suddenly, my mighty warship looked less “naval destroyer” and more “sad boat that lost a fight with a bathtub drain.”
I shoved it in my closet. It sat there for two years.
Here’s the thing: the kit wasn’t defective. I’d just hit The Dip.
.
.
.
The Honeymoon Is Real (Enjoy It While It Lasts)
When Anthropic dropped Claude Opus 4.5 in November, it felt like someone handed us the instruction manual to the universe.
You typed a prompt. Claude built exactly what you described. You shipped in 30 minutes what used to take an entire afternoon of Stack Overflow rabbit holes and frustrated Googling.
The dopamine hit was real.
“I’ll never code the old way again,” you probably said. (I certainly did. Out loud. To no one in particular.)
But here’s what nobody tells you when you’re in the honeymoon phase:
A smarter model doesn’t eliminate the wall you’re about to hit. It just moves it further down the road.
.
.
.
Then the Fog Rolls In
Three weeks in. Same prompt that worked beautifully yesterday. Completely different result today.
Claude starts “forgetting” context. Hallucinating function names that don’t exist. Proposing fixes that miss the point so entirely you wonder if you’re even having the same conversation.
You rephrase. Clarify. Add more detail.
Still broken.
It’s like driving in thick fog. You know the road is there. You know your car works fine. But you can’t see three feet in front of you, and suddenly every turn feels dangerous.
This is the dip in vibe coding. And if you’ve been working with Opus 4.5 for more than a few weeks, you’ve been here.
The symptoms look something like this:
“Claude won’t listen to me anymore”
“Works for simple cases, breaks on anything real”
“My whole process is falling apart”
“Maybe I actually need to understand this stuff…”
Sound familiar? (Be honest. I won’t tell anyone.)
.
.
.
The Spiral We Don’t Talk About
Let me tell you what happened to me.
It was 11 PM. Third cup of coffee—the one that crosses the line from “productive fuel” into “questionable life choices.” I’d been debugging with Claude for over an hour.
The bug was obvious. At least to me. I could see exactly where the logic broke down.
Claude kept proposing fixes that missed the point entirely.
I rephrased. Added context. Tried twelve different angles.
Nothing.
“You’re supposed to be the smartest model on the planet,” I muttered at my screen. (Yes, I talk to my computer. Don’t pretend you don’t.)
“Why can’t you just see this?”
.
.
.
And then came the spiral.
Maybe Anthropic is serving me a weaker model. They probably quantized it to save on compute—everyone knows they throttle heavy users. The conspiracy theories started writing themselves.
I opened Reddit. Started typing.
“Anyone else notice Claude got way worse lately?”
Found others who agreed. Upvotes. Comments. Validation.
Felt good.
The bug was still there.
.
.
.
The Uncomfortable Question
I closed Reddit and stared at my screen.
What if the model didn’t change?
What if I was the variable?
(Stay with me here—this is where it gets good.)
I went back to my conversation with Claude. Read through it again. And this time, I noticed something that made me want to crawl under my desk.
Claude wasn’t seeing what I was seeing.
I had context in my head—error logs I’d scanned, behavior I’d observed, patterns I’d noticed across multiple test runs. None of that was in the conversation. I was asking Claude to diagnose a problem while hiding half the symptoms.
Remember that fog metaphor? I was the fog.
I’d been yelling at my car for not driving properly while I was the one smearing mud on the windshield.
.
.
.
The Insight That Changed Everything
Here’s what I learned—and I want you to write this down somewhere:
👉 LLMs don’t fail because they’re not smart enough. They fail because they can’t see what you see.
Opus 4.5 is brilliant. But brilliance without visibility is just confident guessing. (And confident guessing is somehow more frustrating than obvious failure, isn’t it?)
The moment I understood this, my entire approach changed.
Instead of crafting cleverer prompts, I started providing better context. Added logging. Pasted actual error output—not my interpretation of it, but the raw, ugly, unfiltered thing. Described exactly what I observed.
Suddenly the “dumb” Claude became brilliant again.
The model didn’t change. What I gave it changed.
The fog cleared. The road was there all along.
.
.
.
What Seth Godin Knew All Along
Quick detour. (I promise it’s relevant.)
Seth Godin wrote about The Dip years ago. His insight: The Dip is the long, hard slog between starting something and mastering it. The point where initial excitement fades and real work begins.
Most people quit here. That’s exactly what makes pushing through valuable.
My battleship model kit? I eventually finished it. Two years later, with steadier hands and more patience. It’s not perfect—one gun turret sits slightly crooked—but it’s done. And I learned something about myself in the finishing that I never would have learned in the quitting.
The dip in vibe coding works the same way.
You’ve moved past the toy examples. You’re building real things now. And real things have edge cases, complexity, and problems that require iteration.
That’s always been true. AI doesn’t change the fundamental nature of engineering—it just accelerates everything. The breakthroughs come faster.
So do the walls.
.
.
.
The Dip Survival Kit
Okay. Enough philosophy. Let’s get tactical.
Here are five things I do now when Claude stops cooperating—and more importantly, when I feel the urge to open Reddit instead of solving the actual problem.
1. Visibility First (Clear the Fog)
The problem: Claude fails because it lacks context, not intelligence. You’re driving in fog and blaming the car.
The fix:
Add logging before asking Claude to debug. Paste actual error messages—not your summary of them, the real thing. Describe exactly what you observed.
Opus 4.5 has a massive context window. Use it. Don’t summarize when you can show.
A prompt pattern that actually works:
“Here’s what I’m seeing: [actual output]. Here’s what I expected: [expected output]. Here’s the relevant code: [code]. What’s the disconnect?”
More context = better reasoning. That’s literally why Opus outperforms previous models.
2. Tool Pivot (Different Hammers for Different Nails)
The problem: One tool can’t do everything brilliantly—even Opus 4.5.
The fix:
I’ve started using different AI tools for different phases:
Task
Best Tool
Brainstorming & Planning
Codex / GPT-5
Implementation
Claude Code (Opus 4.5)
Review & QA
Codex reviewing diffs
Claude excels at implementation and extended reasoning. Using it for everything creates blind spots. Match the tool to the task—not the task to the tool.
3. Blueprint Over Code (Clarity Before Creativity)
The problem: Vague prompts produce vague results. Even from Opus 4.5.
The fix:
Write detailed specs before asking Claude to build anything. Use ASCII wireframes for UI. Define success criteria upfront. Get to 95% clarity before implementation begins.
Bad prompt:
“Build me a login form”
Better prompt:
“Build a login form with email and password fields. Email validation on blur. Password requires 8+ characters, one uppercase, one number. Show inline errors below each field. On success, redirect to /dashboard. On failure, show error banner at top. Use existing Button and Input components.”
Here’s the thing: smarter models reward precision. Opus 4.5 can follow complex specifications better than any previous model. Don’t waste that intelligence on ambiguity.
4. The Strategic Retreat (Walking Away Is a Skill)
The problem: Tunnel vision makes everything harder. You’ve been staring at the same bug for two hours and you’re no longer thinking—you’re just reacting.
The fix:
Walk away. Literally.
Work on something else. Go for a walk. Sleep on it. (Yes, actual sleep. Revolutionary concept, I know.)
Your brain keeps processing in the background. You’ll return without the emotional charge. Often—annoyingly often—the solution becomes obvious after distance.
The anti-pattern to avoid:
Don’t keep hammering the same prompt hoping Opus 4.5 will suddenly “get it.” Ten variations of a bad approach is still a bad approach. Step back. Reassess.
5. Experiment Deliberately (Failure as Data)
The problem: Random attempts waste time and tokens. Frustration makes you sloppy.
The fix:
Change one variable at a time. Document what you tried. Look for patterns in what’s failing. Treat debugging as research, not combat.
The mindset shift:
“I’m not failing. I’m eliminating approaches that don’t work.”
A Claude-specific tip:
Start a fresh conversation when context gets polluted. Sometimes Opus 4.5 locks onto a wrong approach and keeps circling back to it like a dog with a favorite squeaky toy. New conversation = clean slate. Sometimes that’s genuinely all you need.
.
.
.
The Other Side of the Dip
Let me be honest with you.
The dip in vibe coding is real. It’s coming for every single person reading this—if it hasn’t arrived already.
Even with Claude Opus 4.5. Even with the most capable AI model we’ve ever had.
But here’s what the dip actually means: you’ve moved past the tutorials. You’re building something real. You’re pushing against the edges of what’s possible.
That’s engineering. That’s the craft.
Every new Claude model will have a honeymoon phase. And every one will eventually lead to a dip. The skill isn’t avoiding it.
The skill is navigating it.
Push through. Tinker. Experiment. Clear the fog. Give Claude what it can’t see on its own.
The developers who learn this with Opus 4.5 will be ready for whatever Anthropic ships next.
New project in mind—a WordPress theme for coworking spaces. And I want to do this right, you know? Start with requirements. Be professional about it. Act like someone who has their act together.
So I type:
I want to build a WordPress theme for Coworking space website
for the coworking space owner to use it to promote and manage
their coworking space.
I want you to help me brainstorm for the product requirement
document (PRD) for this theme.
Put the PRD at @requirements/coworkpress.md
I hit enter, feeling like a real PM.
Claude responds: “I’d be happy to help you brainstorm a PRD for a coworking space WordPress theme. Let me create a comprehensive document covering the key aspects.”
Then it writes. And writes. And writes.
313 lines.
[PRD] CoWorkPress – WordPress Theme
Loading longform...
Couldn't load content. Please try again.
Product Overview
Product Name: CoWorkPress Type: WordPress Theme Target Market: Coworking space owners and operators Purpose: Enable coworking space owners to create professional websites to promote their spaces, showcase amenities, display pricing, and manage bookings.
Problem Statement
Coworking space owners need an affordable, easy-to-use solution to:
Establish an online presence without hiring developers
Showcase their space professionally to attract members
Display pricing and membership options clearly
Allow potential members to book tours or reserve spaces
Build community and promote events
Target Users
Primary Users (Site Owners)
Small coworking space owners - Single location, 10-50 desks
Pricing Model - One-time purchase vs. subscription?
Free vs. Premium - Freemium model with premium add-ons?
Booking Depth - Build custom booking or rely on plugin integrations?
Multi-location - Core feature or premium add-on?
Community Features - Include member directory/forum?
E-commerce - Direct payment processing for memberships?
Notes & Ideas
Consider integration with popular coworking management platforms
Virtual tour integration could be a key differentiator
Mobile app companion for members (future consideration)
Community marketplace for member services
Analytics for space utilization (premium feature)
Product Overview
Product Name: CoWorkPress Type: WordPress Theme Target Market: Coworking space owners and operators Purpose: Enable coworking space owners to create professional websites to promote their spaces, showcase amenities, display pricing, and manage bookings.
Problem Statement
Coworking space owners need an affordable, easy-to-use solution to:
Establish an online presence without hiring developers
Showcase their space professionally to attract members
Display pricing and membership options clearly
Allow potential members to book tours or reserve spaces
Build community and promote events
Target Users
Primary Users (Site Owners)
Small coworking space owners - Single location, 10-50 desks
Pricing Model - One-time purchase vs. subscription?
Free vs. Premium - Freemium model with premium add-ons?
Booking Depth - Build custom booking or rely on plugin integrations?
Multi-location - Core feature or premium add-on?
Community Features - Include member directory/forum?
E-commerce - Direct payment processing for memberships?
Notes & Ideas
Consider integration with popular coworking management platforms
Virtual tour integration could be a key differentiator
Mobile app companion for members (future consideration)
Community marketplace for member services
Analytics for space utilization (premium feature)
At first glance, it looks thorough. Professional, even. Like something a senior engineer would put together after three cups of coffee and a really productive morning.
But here’s where things get squirrely.
Look closer at what Claude produced:
WordPress 6.0+ and PHP 8.0+ version requirements
Full Site Editing (FSE) block-based architecture decisions
Plugin recommendations: WPForms, Amelia, Yoast SEO
A three-phase development roadmap with version numbers
Wait.
Hold on.
I haven’t even started coding yet—haven’t written a single line—and I’ve already “decided” on performance budgets, responsive breakpoints, and plugin integrations?
Huh.
.
.
.
The Problem: Requirements That Aren’t Actually Requirements
Here’s what happened (and stay with me, because this is the part that changes everything): I asked Claude for requirements. Claude gave me implementation decisions.
Those 313 lines contain:
Technical stack choices I haven’t made yet
Performance constraints that might change
Architecture patterns I haven’t validated
Integration approaches I haven’t researched
And now this entire document sits in my context window—taking up space, influencing every future conversation. Every time I ask Claude to help me build something, it references these premature decisions like they’re gospel.
The PRD became a constraint instead of a guide.
Here’s the thing. Claude conflates two very different things:
What you want to build (requirements)
How to build it (implementation)
And honestly? That’s not Claude’s fault. It’s mine. I didn’t tell it otherwise.
But here’s the real insight—the one that made me completely rethink how I approach Claude Code requirements:
User perspective yields better AI output.
(I know, I know. That sounds like something you’d find on a motivational poster in a Silicon Valley bathroom. But hear me out.)
When you describe what users experience, Claude understands context. It grasps the why behind features. It catches edge cases you’d miss while you’re busy thinking about database schemas.
When you describe technical implementation? Claude follows a checklist. No context. No judgment. Just execution.
So what if I asked differently?
.
.
.
The Reframe: Making Claude Interview You First
I tried again. Same project. Different approach.
This time, I asked for a Business Requirements Document—focused on user stories and business logic, not technical specs.
But here’s the key change (and this is the part you’ll want to screenshot): I told Claude to interview me first.
The exact prompt:
I want to build a WordPress theme for Coworking space website
for the coworking space owner to use it to promote and manage
their coworking space.
I want you to help me brainstorm for the business requirement
document (BRD) for this theme. Focus on business logic and
rules, user stories, and acceptance criteria.
Use the "wp-theme-dev" skill to understand the development
patterns and guidance for building WordPress themes.
IMPORTANT: DO NOT INCLUDE ANY TECHNICAL IMPLEMENTATION DETAILS.
Use the AskUserQuestionTool to Ask me clarifying questions
until you are 95% confident you can complete this task
successfully. For each question, add your recommendation
(with reason why) below the options. This would help me
in making a better decision.
Put the BRD at @requirements/coworkpress.md
Three things matter here:
“Focus on business logic, user stories, and acceptance criteria” — This sets explicit scope. No wandering into implementation land.
“DO NOT INCLUDE ANY TECHNICAL IMPLEMENTATION DETAILS” — A hard boundary. (Yes, I’m yelling. Sometimes Claude needs yelling.)
“Ask me clarifying questions until you are 95% confident” — This is the magic ingredient. The secret sauce. The thing that changes everything.
That last part triggers AskUserQuestionTool—a built-in Claude Code feature that lets Claude pause and interview you before generating anything.
What happened next?
It changed how I start every single project.
.
.
.
The Interview Begins
Claude didn’t start writing.
It started asking.
Look at that interface. (Isn’t it beautiful?) Claude organized its questions into themes:
Target Users
Business Model
Space Types
And each question comes with options—plus Claude’s recommendation with reasoning. It’s like having a really thoughtful product manager who actually wants to understand what you’re building before they start building it.
For the business model question, Claude recommended Membership-based because “Monthly/annual subscriptions with different tiers (hot desk, dedicated desk, private office)” represents the most common coworking model.
I selected the recommendation. Next question.
.
.
.
Round 2: Booking and Member Management
After answering the first batch, Claude moved deeper.
Now it’s asking about:
Booking systems (meeting rooms, event spaces)
Member portal requirements
Community features
Notice how Claude builds on previous answers. Because I said “membership-based” earlier, it now asks about member-specific features like portals and directories.
The questions aren’t random. They follow a logical thread—like a conversation with someone who’s actually paying attention.
I selected Both Directory + Events for community features. Claude recommended it because “Full community engagement with directory and events” maximizes the value proposition for members.
Claude recommended Multiple Locations because “Two or more locations with shared or separate memberships” covers both current single-location spaces and future expansion.
It’s thinking about growth—not just current state.
That’s… actually really thoughtful?
.
.
.
Round 4: Sales Funnel and Operations
Now Claude explored:
Lead capture and nurturing
Notification preferences
Reporting needs
For reporting, I went with Comprehensive Dashboard—occupancy rates, revenue, member growth, booking analytics, space utilization.
Claude’s reasoning: space owners need visibility into their business health, and a comprehensive dashboard reduces the need for external tools.
(I mean, yes. Exactly. Thank you for understanding.)
.
.
.
Round 5: Policies and Edge Cases
The final rounds covered operational details:
Amenities handling
Access control methods
Guest policies
And here’s where it gets interesting.
These questions catch edge cases I wouldn’t have thought about until I was knee-deep in development—cursing at my screen at 11 PM, wondering why I didn’t think of this earlier.
Guest policy? I hadn’t considered it. But it affects:
Pricing structure
Access control logic
Liability considerations
Member value perception
Claude surfaced this decision before I started building.
Before I wrote a single line of code.
Before I had to refactor anything.
.
.
.
21 Questions Later: The Full Picture
After 7 rounds, Claude had asked 21 questions across multiple categories.
Space types: Full-service (hot desks, dedicated, private offices, meeting rooms)
Features & Functionality
Booking: Real-time online system
Member portal: Full dashboard with profile, bookings, billing
Community: Directory + events calendar
Operations & Scale
Multiple locations with location-based time zones
Two admin roles (Admin + Staff)
Comprehensive reporting dashboard
Policies & Rules
Flexible billing (monthly/quarterly/annual)
30-day cancellation notice
Plan-based meeting room quotas
No guest access
Corporate accounts supported
Claude now has context. Real context. Not assumptions—actual decisions I made.
“I now have comprehensive understanding of your requirements. Let me create the Business Requirements Document.”
Yes. Yes you do. Finally.
.
.
.
The Output: A Document I Can Actually Build From
[BRD] CoworkPress – WordPress theme
Loading longform...
Couldn't load content. Please try again.
Document Information
Field
Value
Project Name
CoworkPress
Document Type
Business Requirements Document (BRD)
Version
1.0
Created
2026-01-13
Status
Draft
1. Executive Summary
CoworkPress is a WordPress theme designed for coworking space owners to promote and manage their coworking business. The theme enables owners to showcase their spaces, manage memberships, handle bookings, and build community—all through a highly customizable, professional website.
1.1 Business Objectives
Enable coworking space owners to establish a professional online presence
Streamline membership management and billing operations
Provide real-time booking capabilities for spaces and resources
Support multi-location operations under a single platform
Deliver comprehensive business analytics and reporting
Foster community engagement through member directories and events
1.2 Target Market
Primary User (Theme Buyer): Coworking space owners and operators managing one or more locations
End Users (Website Visitors):
Freelancers and remote workers seeking flexible workspace
Small teams and startups needing dedicated desks or private offices
Enterprise clients looking for satellite offices or meeting spaces
2. Business Context
2.1 Business Model
The coworking space operates on a membership-based model with flexible billing cycles (monthly, quarterly, annual). Revenue is generated through:
Recurring membership subscriptions
Additional hour purchases when quotas are exhausted
Larger venues for workshops, seminars, and gatherings
2.3 Multi-Location Operations
The business supports multiple physical locations, each operating in its own time zone. Memberships may be location-specific or provide access across locations depending on the plan.
3. User Personas
3.1 Coworking Space Owner (Admin)
Goals:
Manage all aspects of the coworking business from a single dashboard
Track business performance through analytics and reporting
Customize the website to match brand identity
Oversee memberships, bookings, and staff operations
Permissions: Full system access including financial data, member management, settings, and reporting
3.2 Front Desk Staff
Goals:
Assist members with day-to-day inquiries
Manage bookings and check-ins
Handle basic member account updates
Process tour requests and follow-ups
Permissions: Limited access—cannot access financial reports, billing settings, or system configuration
3.3 Prospective Member (Lead)
Goals:
Explore available spaces and amenities
Understand pricing and membership options
Schedule a tour of the facility
Compare plans to find the right fit
Journey: Website visitor → Tour booking → Lead nurturing → Member conversion
3.4 Individual Member
Goals:
Book meeting rooms and resources within quota
View upcoming bookings and membership status
Access invoices and billing history
Connect with other members through directory
Attend community events
Permissions: Access to member portal, personal bookings, own profile, community features
3.5 Corporate Account Manager
Goals:
Manage employee memberships under company account
View consolidated billing for all employees
Add or remove team members from the plan
Track company-wide space usage
Permissions: Access to company dashboard, employee management, corporate billing
4. Functional Requirements
4.1 Public Website
4.1.1 Homepage
User Stories:
As a visitor, I want to immediately understand what the coworking space offers so I can decide if it meets my needs
As a visitor, I want to see the different locations available so I can find one near me
As a visitor, I want clear calls-to-action so I know how to take the next step
Acceptance Criteria:
[ ] Hero section communicates value proposition clearly
[ ] Location selector/showcase displays all available spaces
[ ] Membership tiers are summarized with pricing
[ ] Prominent CTAs for tour booking and sign-up
[ ] Testimonials or social proof elements are visible
[ ] Mobile-responsive layout
4.1.2 Pricing Page
User Stories:
As a visitor, I want to compare membership plans side-by-side so I can choose the right one
As a visitor, I want to understand what's included in each tier so there are no surprises
As a corporate buyer, I want to see corporate account options so I can evaluate for my team
Acceptance Criteria:
[ ] All membership tiers displayed with clear pricing
[ ] Monthly, quarterly, and annual pricing shown with any discounts
[ ] Amenities included in each tier clearly listed
[ ] Meeting room hour quotas displayed per plan
[ ] Corporate account option highlighted
[ ] CTA to sign up or contact sales for each tier
4.1.3 Locations Page
User Stories:
As a visitor, I want to see all locations so I can find the most convenient one
As a visitor, I want to view details about each location so I can assess the facilities
Acceptance Criteria:
[ ] All locations listed with address and contact information
[ ] Each location shows available space types
[ ] Operating hours displayed per location
[ ] Photo gallery or visual representation of each space
[ ] Link to book a tour at specific location
4.1.4 Virtual Tour Page
User Stories:
As a remote visitor, I want to experience the space virtually so I can evaluate without visiting in person
As a decision-maker, I want to share a virtual tour with my team so we can discuss together
Acceptance Criteria:
[ ] 360-degree tour or video walkthrough available
[ ] Key amenities and features highlighted
[ ] Easy to navigate between different areas
[ ] Works on mobile devices
[ ] Option to schedule in-person tour after viewing
4.1.5 About Page
User Stories:
As a visitor, I want to learn about the company and team so I can trust the brand
As a visitor, I want to understand the coworking philosophy so I know if it aligns with my values
Acceptance Criteria:
[ ] Company story and mission clearly presented
[ ] Team members or leadership showcased (optional)
[ ] Community values or culture highlighted
[ ] Contact information accessible
4.1.6 Contact Page
User Stories:
As a visitor, I want to easily reach the team with questions so I can get answers before committing
As a visitor, I want location-specific contact options so I reach the right team
Acceptance Criteria:
[ ] Contact form with required fields (name, email, message)
[ ] Phone number and email displayed
[ ] Location selector if multiple locations exist
[ ] Map integration showing physical location(s)
[ ] Expected response time communicated
4.2 Tour Booking & Lead Management
4.2.1 Tour Scheduling
User Stories:
As a prospective member, I want to book a tour online so I can visit at a convenient time
As a staff member, I want to see scheduled tours so I can prepare for visitors
As an admin, I want tour data captured so I can follow up with leads
Acceptance Criteria:
[ ] Calendar interface showing available tour slots
[ ] Location selection for multi-location operations
[ ] Lead information captured (name, email, phone, company, needs)
[ ] Confirmation email sent to prospect after booking
[ ] Tour appears in staff/admin dashboard
[ ] Reminder sent 24 hours before scheduled tour
4.2.2 Lead Nurturing
User Stories:
As a staff member, I want to track lead status so I know who to follow up with
As an admin, I want to see conversion metrics so I can improve sales process
Acceptance Criteria:
[ ] Lead status tracking (new, toured, follow-up, converted, lost)
[ ] Notes field for staff to record interactions
[ ] Follow-up reminders for staff
[ ] Lead source tracking (how they found us)
[ ] Basic lead-to-member conversion reporting
4.3 Membership Management
4.3.1 Membership Plans
User Stories:
As an admin, I want to create and manage membership plans so I can offer different options
As a member, I want to clearly understand my plan benefits so I know what I'm paying for
Business Rules:
Each plan has a name, description, price, and billing cycle options
Plans include specific amenities (bundled, not à la carte)
Plans define meeting room hour quotas per billing period
Plans may be location-specific or cross-location
Pricing supports monthly, quarterly, and annual cycles
Acceptance Criteria:
[ ] Admin can create, edit, and deactivate plans
[ ] Each plan specifies included amenities
[ ] Meeting room quotas defined per plan
[ ] Billing cycles with pricing for each cycle length
[ ] Plans can be restricted to specific locations or allow all
4.3.2 Member Sign-up
User Stories:
As a prospective member, I want to sign up for a membership online so I can start using the space
As an admin, I want to manually add members so I can onboard people who signed up offline
Members must provide valid payment method to activate membership
MEM-002
Free trials limited to one per person based on email
MEM-003
Upgrades take effect immediately with prorated billing
MEM-004
Downgrades take effect at next billing cycle
MEM-005
Cancellations require 30-day notice
MEM-006
Cancelled memberships remain active until end of billing cycle
MEM-007
Paused memberships do not accrue quota or charges
5.2 Booking Rules
Rule
Description
BKG-001
Bookings deduct from quota at time of confirmation
BKG-002
Cancellations refund quota only if within allowed window
BKG-003
Members cannot book more than their remaining quota unless purchasing extra
BKG-004
Quotas reset at start of each billing cycle
BKG-005
Unused quota does not roll over
BKG-006
All bookings display in location's local time zone
5.3 Corporate Account Rules
Rule
Description
CORP-001
Corporate accounts receive single consolidated invoice
CORP-002
Account manager can add/remove employees
CORP-003
Employee count limited by corporate plan
CORP-004
Employees share company quota or have individual quotas (configurable)
5.4 Billing Rules
Rule
Description
BILL-001
Billing cycles: monthly, quarterly, annual
BILL-002
Annual plans may include discount (configurable)
BILL-003
Failed payments trigger retry and member notification
BILL-004
Membership suspended after X failed payment attempts
BILL-005
Additional hour purchases charged immediately
6. Non-Functional Requirements
6.1 Usability
Website must be mobile-responsive
Member portal accessible on mobile devices
Booking process completable in under 2 minutes
Admin interface intuitive without extensive training
6.2 Accessibility
Theme should follow WCAG 2.1 AA guidelines
All interactive elements keyboard accessible
Proper color contrast for readability
Alt text for images
6.3 Performance
Public pages load within 3 seconds
Booking availability updates in real-time
Dashboard loads within 5 seconds
6.4 Security
Member data protected with secure authentication
Payment information handled by PCI-compliant processor
Admin actions logged for audit trail
Personal data export capability (GDPR compliance)
6.5 Scalability
Support for multiple locations
Handle hundreds of concurrent members
Booking system supports high-demand periods
7. Out of Scope
The following are explicitly NOT included in this theme:
Digital access control integration (key cards, QR codes, mobile app)
Guest pass management
À la carte amenity purchases
Franchise/white-label multi-tenant operations
Native mobile applications
Third-party calendar sync (Google, Outlook)
Chat/messaging between members
Automated marketing campaigns
Inventory management for physical supplies
8. Glossary
Term
Definition
Hot Desk
Shared workspace with no assigned seating
Dedicated Desk
Reserved desk assigned to a specific member
Private Office
Enclosed office space for individual or team use
Quota
Allocated meeting room hours included in membership
Billing Cycle
Period between recurring payments (monthly, quarterly, annual)
Lead
Prospective member who has shown interest but not signed up
Conversion
When a lead becomes a paying member
Churn
When a member cancels their membership
Utilization
Percentage of available time a resource is booked
Document Information
Field
Value
Project Name
CoworkPress
Document Type
Business Requirements Document (BRD)
Version
1.0
Created
2026-01-13
Status
Draft
1. Executive Summary
CoworkPress is a WordPress theme designed for coworking space owners to promote and manage their coworking business. The theme enables owners to showcase their spaces, manage memberships, handle bookings, and build community—all through a highly customizable, professional website.
1.1 Business Objectives
Enable coworking space owners to establish a professional online presence
Streamline membership management and billing operations
Provide real-time booking capabilities for spaces and resources
Support multi-location operations under a single platform
Deliver comprehensive business analytics and reporting
Foster community engagement through member directories and events
1.2 Target Market
Primary User (Theme Buyer): Coworking space owners and operators managing one or more locations
End Users (Website Visitors):
Freelancers and remote workers seeking flexible workspace
Small teams and startups needing dedicated desks or private offices
Enterprise clients looking for satellite offices or meeting spaces
2. Business Context
2.1 Business Model
The coworking space operates on a membership-based model with flexible billing cycles (monthly, quarterly, annual). Revenue is generated through:
Recurring membership subscriptions
Additional hour purchases when quotas are exhausted
Larger venues for workshops, seminars, and gatherings
2.3 Multi-Location Operations
The business supports multiple physical locations, each operating in its own time zone. Memberships may be location-specific or provide access across locations depending on the plan.
3. User Personas
3.1 Coworking Space Owner (Admin)
Goals:
Manage all aspects of the coworking business from a single dashboard
Track business performance through analytics and reporting
Customize the website to match brand identity
Oversee memberships, bookings, and staff operations
Permissions: Full system access including financial data, member management, settings, and reporting
3.2 Front Desk Staff
Goals:
Assist members with day-to-day inquiries
Manage bookings and check-ins
Handle basic member account updates
Process tour requests and follow-ups
Permissions: Limited access—cannot access financial reports, billing settings, or system configuration
3.3 Prospective Member (Lead)
Goals:
Explore available spaces and amenities
Understand pricing and membership options
Schedule a tour of the facility
Compare plans to find the right fit
Journey: Website visitor → Tour booking → Lead nurturing → Member conversion
3.4 Individual Member
Goals:
Book meeting rooms and resources within quota
View upcoming bookings and membership status
Access invoices and billing history
Connect with other members through directory
Attend community events
Permissions: Access to member portal, personal bookings, own profile, community features
3.5 Corporate Account Manager
Goals:
Manage employee memberships under company account
View consolidated billing for all employees
Add or remove team members from the plan
Track company-wide space usage
Permissions: Access to company dashboard, employee management, corporate billing
4. Functional Requirements
4.1 Public Website
4.1.1 Homepage
User Stories:
As a visitor, I want to immediately understand what the coworking space offers so I can decide if it meets my needs
As a visitor, I want to see the different locations available so I can find one near me
As a visitor, I want clear calls-to-action so I know how to take the next step
Acceptance Criteria:
[ ] Hero section communicates value proposition clearly
[ ] Location selector/showcase displays all available spaces
[ ] Membership tiers are summarized with pricing
[ ] Prominent CTAs for tour booking and sign-up
[ ] Testimonials or social proof elements are visible
[ ] Mobile-responsive layout
4.1.2 Pricing Page
User Stories:
As a visitor, I want to compare membership plans side-by-side so I can choose the right one
As a visitor, I want to understand what's included in each tier so there are no surprises
As a corporate buyer, I want to see corporate account options so I can evaluate for my team
Acceptance Criteria:
[ ] All membership tiers displayed with clear pricing
[ ] Monthly, quarterly, and annual pricing shown with any discounts
[ ] Amenities included in each tier clearly listed
[ ] Meeting room hour quotas displayed per plan
[ ] Corporate account option highlighted
[ ] CTA to sign up or contact sales for each tier
4.1.3 Locations Page
User Stories:
As a visitor, I want to see all locations so I can find the most convenient one
As a visitor, I want to view details about each location so I can assess the facilities
Acceptance Criteria:
[ ] All locations listed with address and contact information
[ ] Each location shows available space types
[ ] Operating hours displayed per location
[ ] Photo gallery or visual representation of each space
[ ] Link to book a tour at specific location
4.1.4 Virtual Tour Page
User Stories:
As a remote visitor, I want to experience the space virtually so I can evaluate without visiting in person
As a decision-maker, I want to share a virtual tour with my team so we can discuss together
Acceptance Criteria:
[ ] 360-degree tour or video walkthrough available
[ ] Key amenities and features highlighted
[ ] Easy to navigate between different areas
[ ] Works on mobile devices
[ ] Option to schedule in-person tour after viewing
4.1.5 About Page
User Stories:
As a visitor, I want to learn about the company and team so I can trust the brand
As a visitor, I want to understand the coworking philosophy so I know if it aligns with my values
Acceptance Criteria:
[ ] Company story and mission clearly presented
[ ] Team members or leadership showcased (optional)
[ ] Community values or culture highlighted
[ ] Contact information accessible
4.1.6 Contact Page
User Stories:
As a visitor, I want to easily reach the team with questions so I can get answers before committing
As a visitor, I want location-specific contact options so I reach the right team
Acceptance Criteria:
[ ] Contact form with required fields (name, email, message)
[ ] Phone number and email displayed
[ ] Location selector if multiple locations exist
[ ] Map integration showing physical location(s)
[ ] Expected response time communicated
4.2 Tour Booking & Lead Management
4.2.1 Tour Scheduling
User Stories:
As a prospective member, I want to book a tour online so I can visit at a convenient time
As a staff member, I want to see scheduled tours so I can prepare for visitors
As an admin, I want tour data captured so I can follow up with leads
Acceptance Criteria:
[ ] Calendar interface showing available tour slots
[ ] Location selection for multi-location operations
[ ] Lead information captured (name, email, phone, company, needs)
[ ] Confirmation email sent to prospect after booking
[ ] Tour appears in staff/admin dashboard
[ ] Reminder sent 24 hours before scheduled tour
4.2.2 Lead Nurturing
User Stories:
As a staff member, I want to track lead status so I know who to follow up with
As an admin, I want to see conversion metrics so I can improve sales process
Acceptance Criteria:
[ ] Lead status tracking (new, toured, follow-up, converted, lost)
[ ] Notes field for staff to record interactions
[ ] Follow-up reminders for staff
[ ] Lead source tracking (how they found us)
[ ] Basic lead-to-member conversion reporting
4.3 Membership Management
4.3.1 Membership Plans
User Stories:
As an admin, I want to create and manage membership plans so I can offer different options
As a member, I want to clearly understand my plan benefits so I know what I'm paying for
Business Rules:
Each plan has a name, description, price, and billing cycle options
Plans include specific amenities (bundled, not à la carte)
Plans define meeting room hour quotas per billing period
Plans may be location-specific or cross-location
Pricing supports monthly, quarterly, and annual cycles
Acceptance Criteria:
[ ] Admin can create, edit, and deactivate plans
[ ] Each plan specifies included amenities
[ ] Meeting room quotas defined per plan
[ ] Billing cycles with pricing for each cycle length
[ ] Plans can be restricted to specific locations or allow all
4.3.2 Member Sign-up
User Stories:
As a prospective member, I want to sign up for a membership online so I can start using the space
As an admin, I want to manually add members so I can onboard people who signed up offline
No premature implementation decisions. Just clear requirements I can build against.
And because Claude interviewed me first, these requirements reflect my decisions—not Claude’s assumptions.
That’s the difference.
.
.
.
👉 The Prompt Template (Copy-Paste This)
Here’s the exact prompt structure you can steal (please steal it—that’s why I’m sharing it):
I want to build [YOUR PROJECT DESCRIPTION].
I want you to help me brainstorm for the business requirement
document (BRD) for this [project type]. Focus on business logic
and rules, user stories, and acceptance criteria.
[Optional: Use the "[skill-name]" skill to understand the
development patterns and guidance for building [technology].]
IMPORTANT: DO NOT INCLUDE ANY TECHNICAL IMPLEMENTATION DETAILS.
Use the AskUserQuestionTool to Ask me clarifying questions
until you are 95% confident you can complete this task
successfully. For each question, add your recommendation
(with reason why) below the options. This would help me
in making a better decision.
Put the BRD at @[your-path]/[filename].md
Why each piece matters:
“Focus on business logic, user stories, and acceptance criteria” Forces Claude to think from user perspective, not implementation. This is where the magic happens.
“DO NOT INCLUDE ANY TECHNICAL IMPLEMENTATION DETAILS” Hard boundary that keeps the document clean. (Again with the yelling. It works.)
“95% confident” Gives Claude a threshold—it keeps asking until it has enough information. No more guessing.
“Add your recommendation (with reason why)” Claude explains its thinking, helping you make informed decisions instead of random choices.
.
.
.
Where Do Technical Details Go?
Okay, so if technical specs don’t belong in your requirements, where do they live?
Two places:
Project rules — Your .claude/rules/ directory
Agent skills — Reusable patterns Claude loads on-demand
I covered this in detail in my previous issue on self-evolving Claude Code rules. (If you missed it, go read it. I’ll wait.)
The key insight: Claude loads skills when it needs them.
Instead of burning context on every technical decision upfront, Claude reads skill descriptions at session start. When it encounters a database task, then it loads your database rules. When it’s building an API, then it loads your API patterns.
Your requirements stay clean. Your context window stays focused. Technical standards still get applied—just at the right moment.
Think of it as separation of concerns for your AI workflow:
Requirements (BRD) = What you want to build (the what and why)
Project rules & skills = How to build it (the how)
Simple. Elegant. Actually works.
.
.
.
👉 The Real Lesson
The difference between mediocre AI output and genuinely great output often comes down to one thing:
Did you let the AI assume, or did you make it ask?
When you hand Claude a prompt and let it run, it fills gaps with assumptions. Sometimes those assumptions are good. Often? They’re not. And you won’t know until you’re three hours into building the wrong thing.
When you force Claude to interview you first—to surface decisions, explore edge cases, and validate understanding—you get output that reflects your vision. Not Claude’s best guess.
21 questions took about 5 minutes to answer.
Those 5 minutes saved me from 313 lines of premature decisions and gave me requirements I can actually build against.
That’s a pretty good trade.
Your turn:
Next time you start a project, don’t ask Claude for a PRD.
Last week, I watched Claude make the exact same mistake I’d corrected three days earlier.
Same project. Same codebase. Same error.
I’d spent fifteen minutes explaining why we use prepared statements for database queries—not string concatenation. Claude understood. Claude apologized. Claude fixed it beautifully.
And then, in a fresh session, Claude did it again. Like our conversation never happened.
Here’s the thing: it wasn’t Claude’s fault.
(Stay with me.)
The correction I’d made? It lived and died in that single session. I never added it to my Claude Code rules. Never updated my project guidelines. Never captured what we’d learned together.
So Claude forgot. Because Claude had to forget.
And honestly? I’ve done this more times than I’d like to admit.
.
.
.
The Uncomfortable Truth About Your Claude Code Rules
You’ve probably got project rules somewhere.
Maybe in a CLAUDE.md file. Maybe in a markdown doc you include at the start of sessions. Maybe scrawled on a Post-it note stuck to your monitor.
(No judgment.)
These rules matter.
They’re the guardrails that keep Claude from hallucinating, making mistakes, or generating code that looks like it was written by someone who’s never seen your codebase before.
But here’s what nobody talks about: most Claude Code rules are frozen in time.
You wrote them once—probably when you were optimistic and caffeinated and full of good intentions. Maybe you updated them once or twice when something broke spectacularly. And then… they fossilized.
Meanwhile, you’re out there learning. Every session teaches you something. Better patterns. Sneaky edge cases. Production bugs that made you question your life choices at 2am.
But none of that learning makes it back into your rules.
Your rules stay stuck at whatever understanding you had on Day One. A Level 5 Charmander trying to fight Level 50 battles.
(More on Pokémon in a minute. I promise this metaphor is going somewhere.)
.
.
.
The Three Problems That Keep Your Rules Stuck
Let me break down why this happens—because understanding the problem is half the battle.
Problem #1: The Inclusion Tax
You have to remember to add your rules file at the start of every session. Miss it once—maybe you’re rushing, maybe you’re excited about a feature, maybe you just forgot—and suddenly you’re debugging code that violates your own standards.
It’s like having a gym membership you forget to use. The potential is there. The execution… less so.
Problem #2: Context Decay
Even when you do include your rules, long sessions dilute them. By message #50, your carefully crafted “always use prepared statements” guideline has degraded into… whatever Claude feels like doing.
The rules are still technically in the context window. They’re just competing with 47 other messages for Claude’s attention. And losing.
Here’s what that looks like:
Problem #3: The Maintenance Black Hole
Rules should evolve. You know this. I know this. We all know this.
But who actually updates their rules file regularly?
(I’m raising my hand here too. We’re in this together.)
.
.
.
The CLAUDE.md Trap
Here’s what most developers do—and I get why it seems logical:
They cram everything into CLAUDE.md.
Security rules. Database patterns. API conventions. Coding standards. Error handling preferences. That one weird edge case from six months ago. All of it, stuffed into one giant file that gets loaded at the start of every session.
The thinking makes sense: “Claude will always have my rules!”
The reality?
A 2,000-line CLAUDE.md file burns through your token limits before you even start working.
You’re paying the context tax on every single message—whether you need those rules or not. Working on a simple UI tweak? Still loading all your database migration patterns. Fixing a typo? Still burning tokens on your entire security playbook.
It’s like packing your entire wardrobe for a weekend trip. Sure, you’ll have options. But you’ll also be exhausted before you get anywhere.
.
.
.
What If Your Rules Could Level Up?
Okay, here’s where the Pokémon thing comes in. (Told you I’d get there.)
Think about how evolution works in Pokémon. A Charmander doesn’t stay a Charmander forever. It battles. It gains experience. It evolves into Charmeleon, then Charizard. Each evolution makes it stronger, better suited for tougher challenges.
What if your Claude Code rules could work the same way?
Every correction you make.
Every “actually, do it this way instead” moment. Every hard-won insight from debugging at 2am. What if all of that could strengthen your rules automatically?
Not stuffed into a bloated CLAUDE.md. Not forgotten when sessions end. Actually captured and integrated into a system that gets smarter over time.
This is possible now. And it’s not even that complicated.
Enter Agent Skills (And Why They Change Everything)
If you haven’t explored Agent Skills yet, here’s the short version: Skills let Claude activate knowledge only when it needs it.
Instead of loading everything upfront—all 2,000 lines of your rules, whether relevant or not—Claude starts by reading just the skill descriptions. A few lines each. When a task triggers a specific skill, then it reads the full details.
Working on database queries? Claude loads the database skill. Building an API endpoint? It loads the API patterns. Simple UI change? It loads… just what it needs for that.
The token savings are significant. Instead of paying for your entire rulebook on every message, you pay for maybe 50 lines of descriptions, plus whatever specific reference you actually need.
But here’s where it gets interesting for our “evolving rules” problem…
The Card Catalog Architecture
Here’s the approach that solves all three problems we talked about earlier:
Structure your skill like a card catalog, not a library.
Your SKILL.md file becomes an index—brief descriptions pointing to reference files. The actual rules live in a references/ folder:
Instead of cramming everything into SKILL.md like this:
## Database Schema Rules
Always use snake_case for table names...
[50 more lines of database rules]
## Firebase Security Rules
Auth patterns must follow...
[40 more lines of Firebase rules]
Your SKILL.md becomes a simple directory:
## Database Schema Rules
See: references/database.md
Guidelines for table naming, migrations, and schema patterns.
## Firebase Security Rules
See: references/security.md
Authentication patterns and Firestore rules structure.
Why does this matter for evolution?
When you capture new insights, they go into the specific reference file—not bloating the main index. Your skill grows by expanding its reference library, not by inflating one massive file.
And when you pair this structure with the build-insights-logger skill? Your rules start updating themselves based on real learnings.
Your Charmander starts evolving.
.
.
.
Building Your Self-Evolving Skill: The Complete Walkthrough
Alright, let’s get practical. Here’s exactly how to set this up—step by step, with real screenshots from when I did this myself.
Step 1: Convert Your Existing Rules to an Indexed Skill
If you already have Claude Code rules somewhere (in CLAUDE.md, a rules file, or scattered notes), this is your starting point.
Here’s the prompt I used:
Convert the following project rules at @notes/rules.md into a live documentation skill using the "skill-creator" skill.
## Requirements
### 1. Skill Structure
Create a skill folder with this structure:
```
.claude/skills/[skill-name]/
├── SKILL.md # Index file (card catalog, NOT the full library)
├── references/ # Detailed rule files
│ ├── [category-1].md
│ ├── [category-2].md
│ └── ...
└── [README.md](http://README.md) # Optional: skill usage guide
```
### 2. SKILL.md Design (The Index)
The SKILL.md should act as a **card catalog**, not contain the full rules.
Example:
```
[...content of SKILL.md...]
## Quick Reference
Brief overview of what this skill covers (2-3 sentences max).
## Rule Categories
### [Category 1 Name]
Brief description (1 line). See: `references/[category-1].md`
### [Category 2 Name]
Brief description (1 line). See: `references/[category-2].md`
[Continue for all categories...]
## When to Load References
- Load `references/[category-1].md` when: [specific trigger]
- Load `references/[category-2].md` when: [specific trigger]
```
### 3. Reference File Design
Each reference file in `references/` should:
- Focus on ONE category/domain of rules
- Include code examples where helpful
- Be self-contained (can be understood without reading other files)
- End with a "Quick Checklist" for that category
### 4. Evolvability
Structure the skill so new learnings can be easily added:
- Each reference file should have a `## Lessons Learned` section at the end (initially empty)
- The SKILL.md index should be easy to extend with new categories
- Use consistent formatting so automated updates are possible
## Notes
- Prioritize token efficiency: Claude should only load what it needs
- Keep SKILL.md under 100 lines if possible
- Each reference file should be focused (ideally under 200 lines)
- Use the indexed structure so the build-insights-logger can update specific reference files later
Two prerequisites before you run this:
Install the “skill-creator” skill first
Turn on Plan Mode (Shift+Tab)
Claude initiates the skill-creator and starts working:
It explores your codebase for existing patterns to follow. (I love watching this part—it’s like Claude doing research before diving in.)
In my case, Claude asked clarifying questions before finalizing the plan. This is Plan Mode doing its job—thinking before coding:
After answering, Claude proposed a complete plan:
The plan includes the SKILL.md design—notice how it acts as an index, not a container for all the rules:
I agreed to proceed with auto-accept edits. And here’s what Claude created:
The result:
SKILL.md: 83 lines (the index/card catalog)
15 reference files, each under 200 lines
Every reference file includes code examples, a Quick Checklist, and an empty Lessons Learned section
Those empty Lessons Learned sections? They’re intentional. That’s where the evolution happens.
👉 Don’t have existing rules to convert? You can ask Claude to analyze your codebase and extract patterns into a skill. Or check out my previous post on how to compile your own project rules first.
Step 2: Install the Build-Insights-Logger
This is the skill that captures learnings during your sessions and routes them to the right place.
(Yes, I built this. Yes, I’m biased. But it solves a real problem.)
Step 3: Work Like You Normally Would
Here’s the beautiful part: you don’t need to change how you work.
With your skill installed and the insights logger ready, just… build. Code. Debug. Do your thing.
I’ll show you what this looks like. I asked Claude to audit my WordPress plugin for security vulnerabilities:
Claude found several issues, fixed them, and documented the changes. Standard stuff.
But here’s where it gets good.
After the implementation, I triggered the insights logger:
Please jot down what you have learned so far using the "build-insights-logger" skill.
Claude activated the skill:
And logged 6 insights from the session:
Six insights. Automatically categorized. Saved to .claude/insights/session-2026-01-05-143600.md.
No manual documentation. No “I should write this down” that never happens. Just… captured.
Step 4: Review and Integrate (The Curation Step)
You can review immediately after one session, or batch multiple sessions together. I usually wait until I have a few sessions worth of insights—but that’s personal preference.
Here’s how to review:
Please use the "build-insights-logger" skill to review the insights logged so far.
Claude reads all session files and presents a summary organized by category:
Now here’s the critical part: don’t add everything.
This is where human judgment matters. Review each insight and ask: “Will this apply to future projects, or is it specific to this one-off feature?”
Some insights are gold. Some are situational. You’re the curator here.
I selected insights 1, 2, 5, and 6—the ones that generalize across WordPress projects:
Important detail: Notice my instruction. By default, the build-insights-logger updates CLAUDE.md. I explicitly redirected it to update the skill instead, with a reminder to keep SKILL.md clean—insights should go to the relevant reference files, not the index.
Claude reads the relevant reference files and adds the insights where they belong:
The session file gets archived. And your skill is now smarter than it was an hour ago.
.
.
.
What You Actually Built Here
Let’s step back for a second.
Four steps. That’s all it took:
Convert rules to an indexed skill
Install the insights logger
Build like normal
Review and integrate learnings
But what actually changed?
You’re no longer maintaining documentation. You’re growing a knowledge base.
Every correction you make—every “actually, do it this way instead” moment, every insight from debugging at 2am—the system captures it. You review it. The good stuff evolves your skill. Your next session starts with rules that reflect what you actually learned, not what you thought you knew when you started.
The gap between your experience and your documentation? It closes.
That mistake I mentioned at the beginning—Claude repeating an error I’d already corrected? It doesn’t happen anymore. Because the correction made it into my Claude Code rules. Automatically. As part of my normal workflow.
Your Charmander evolved into Charizard.
And it keeps evolving.
.
.
.
Your Turn
If you’ve been cramming rules into CLAUDE.md—or worse, keeping them in your head and hoping for the best—try this process on your next project.
Start with whatever rules you have. Even messy ones. (Especially messy ones, honestly.)
Convert them to an indexed skill. Install the insights logger. Build for a few sessions. Then review your insights folder.
You’ll be surprised at what Claude captured. And you’ll be amazed at how much smarter your rules become when they’re allowed to learn alongside you.
Imagine trying to teach someone to cook over the phone.
You’re walking them through your grandmother’s pasta recipe—the one with the garlic that needs to be just golden, not brown. You describe every step perfectly. The timing. The technique. The little flip of the wrist when you toss the noodles.
And then they say: “It’s burning. What do I do?”
Here’s the thing: you can’t help them. Not really. Because you can’t see the pan. You can’t see how high the flame is. You can’t see that they accidentally grabbed the chili flakes instead of oregano. All you have is their panicked description and your best guess about what might be going wrong.
This, my friend, is exactly what happens when you ask Claude Code to fix a bug.
(Stay with me here.)
.
.
.
The Merry-Go-Round Nobody Enjoys
You’ve been on this ride before. I know you have.
You describe the bug to Claude. Carefully. Thoroughly. You even add screenshots and error messages because you’re a good communicator, dammit.
Claude proposes a fix.
You try it.
It doesn’t work.
So you describe the bug again—this time with more adjectives and maybe a few capitalized words for emphasis. Claude proposes a slightly different fix. Still broken. You rephrase. Claude tries another angle. Round and round we go.
This is the debugging merry-go-round, and nobody buys tickets to this ride on purpose.
The instinct—the very human instinct—is to blame the AI.
“Claude isn’t smart enough for this.”
“Maybe I need a different model.”
“Why can’t it just SEE what’s happening?”
That last one?
That’s actually the right question.
Just not in the way you think.
Here’s what I’ve learned after spending more time than I’d like to admit arguing with AI about bugs: Claude almost never fails because it lacks intelligence. It fails because it lacks visibility.
Think about what you have access to when you’re debugging. Browser dev tools. Console logs scrolling in real-time. Network requests you can inspect. Elements that highlight when you hover. The actual, living, breathing behavior playing out on your screen.
What does Claude have?
The code. Just the code.
That’s it.
You’re asking a brilliant chef to fix your burning pasta—but they can only read the recipe card. They can’t see the flame. They can’t smell the smoke. They’re working with incomplete information and filling in the gaps with educated guesses.
Sometimes those guesses are right. (Claude is genuinely brilliant at guessing.)
Most of the time? Merry-go-round.
.
.
.
The Two Bugs That Break AI Every Time
After countless Claude Code debugging sessions—some triumphant, many humbling—I’ve noticed two categories that consistently send AI spinning:
The Invisible State Bugs
React’s useEffect dependencies.
Race conditions. Stale closures. Data that shapeshifts mid-lifecycle like some kind of JavaScript werewolf. These bugs are invisible in the code itself. You can stare at the component for hours (ask me how I know) and see nothing wrong. The bug only reveals itself at runtime—in the sequence of events, the timing of updates, the order of renders.
It’s happening in dimensions Claude can’t perceive.
The “Wrong Address” Bugs
CSS being overridden by inline JavaScript. WordPress functions receiving unexpected null values from somewhere upstream. Error messages that point to line 7374 of a core file—not your code, but code three function calls removed from the actual problem.
The error exists.
But the source? Hidden in cascading calls, plugin interactions, systems talking to systems.
Claude can’t solve either category by reading code alone.
So what do we do?
We give Claude eyes.
(I told you to stay with me. Here’s where it gets good.)
.
.
.
Method 1: Turn Invisible Data Into Evidence Claude Can Actually See
Let me walk you through a real example.
Because theory is nice, but showing you what this looks like in practice? That’s the good stuff.
I had a Products Browser component. Simple filtering and search functionality—the kind of thing you build in an afternoon and then spend three days debugging because life is like that sometimes.
Each control worked beautifully in isolation:
Search for “apple” → Three results. Beautiful.
Filter by “laptops” → Five results. Chef’s kiss.
But combine them?
Search “apple” + category “laptops” → Broken. The filter gets completely ignored, like I never selected it at all.
Classic React hook dependency bug.
If you’re experienced with React, you spot this pattern in your sleep. But if you’re newer to the framework—or if you vibe-coded this component and touched a dozen files before realizing something broke—you’re stuck waiting for Claude to get lucky.
I spent three rounds asking Claude to fix it. Each fix addressed a different theoretical cause. None worked.
That’s when I stopped arguing and started instrumenting.
Step 1: Ask Claude to Add Logging (Not Fixes)
Instead of another “please fix this” prompt, I asked Claude to help me see what was happening:
Notice what I didn’t say: “Fix this bug.”
What I said: “Add logging to track data changes.”
This is the mindset shift that changes everything.
Claude added console.log statements to every useEffect that touched the view state:
Each log captured which effect triggered, what the current values were, and what got computed. Basically, Claude created a running transcript of everything happening inside my component’s brain.
Step 2: Run the Test and Capture What You See
I opened the browser, selected “laptops” from the category filter, then typed “apple” in the search box.
The console lit up like a Christmas tree of evidence.
Step 3: Feed the Logs Back to Claude
Here’s where the magic happens. I copied that console output—all of it—and pasted it directly into Claude:
And Claude? Claude saw everything:
Claude found the bug immediately.
The logs revealed the whole story: when I selected a category, useEffect:filters fired and correctly filtered the products. But then when I typed in the search box, useEffect:search fired—and it ran against the full product list, completely ignoring the category filter.
The search effect was overwriting the filter results.
Last effect wins. (JavaScript, you beautiful chaos gremlin.)
Claude proposed the fix: replace multiple competing useEffect hooks with a single useMemo that applies all transforms together:
The difference between “Claude guessing for 20 minutes” and “Claude solving it instantly” was 30 seconds of logging.
That’s not hyperbole. That’s just… math.
.
.
.
Method 2: Map the Problem Before Anyone Tries to Solve It
The second method works for a different beast entirely—the kind of bug where even the error message is lying to you.
Here’s a WordPress error that haunted me for hours:
Deprecated: strpos(): Passing null to parameter #1 ($haystack) of type string
is deprecated in /var/www/html/wp-includes/functions.php on line 7374
Warning: Cannot modify header information - headers already sent by
(output started at /var/www/html/wp-includes/functions.php:7374)
in /var/www/html/wp-includes/option.php on line 1740
If you’ve done any WordPress development, you recognize this particular flavor of suffering.
The error points to core WordPress files—not your code. Something, somewhere, is passing null to a function that expects a string. But where? The error message is about as helpful as a fortune cookie that just says “bad things happened.”
I’d made changes to several theme files.
Any one of them could be the culprit.
And the cascading nature of WordPress hooks meant the error could originate three or four function calls before the actual crash.
After a few rounds of Claude trying random fixes (bless its heart), I tried something completely different.
The Brainstorming Prompt That Changes Everything
Instead of “fix this,” I asked Claude to brainstorm debugging approaches—and to visualize them with ASCII diagrams.
(I know. ASCII diagrams. In 2025. But stay with me, because this is where Claude Code debugging gets genuinely interesting.)
Claude Maps the Error Chain
Claude started by analyzing the flow of the problem:
The diagram showed exactly what was happening: some theme code was passing null to WordPress core functions, which then passed that null to PHP string functions, which threw the deprecation warning.
But which theme code? Claude identified the suspect locations:
Four possible sources.
Each with code examples showing what the problematic pattern might look like.
This is Claude thinking out loud, visually. And it’s incredibly useful for Claude Code debugging because now we’re not guessing—we’re investigating.
Multiple Debugging Strategies (Not Just One)
Rather than jumping to a single fix and hoping, Claude laid out several approaches:
Option A: Search all filter callbacks for missing return statements.
Option B: Find which WordPress functions use strpos internally.
Option C: Add debug_backtrace() at the error point to trace the caller.
Option D: Search for common patterns like wp_redirect with variables.
Four different angles of attack.
This is what systematic debugging looks like—and it’s exactly what you need when you’re stuck in the merry-go-round.
Claude Does Its Homework
Here’s where Opus 4.5 surprised me.
Instead of settling on the first approach, it validated its theories by actually searching the codebase:
It searched for wp_redirect calls, add_filter patterns, get_option usages—systematically eliminating possibilities like a detective working through a suspect list.
Then it updated its diagnosis based on what it found:
The investigation narrowed.
The error was coming from path-handling functions—something was returning a null path where a string was expected.
The Summary That Actually Leads Somewhere
Claude concluded with a clear summary of everything we now knew:
And multiple approaches to fix it, ranked by how surgical they’d be:
Did it work?
First attempt. Approach A—adding a debug backtrace—immediately revealed a function in FluentCartBridge.php that was returning null when $screen->id was empty.
One additional null check.
Bug gone.
All those rounds of failed attempts? They were doomed from the start because Claude was guessing blindly. Once it could see the error chain visually—once it had a map instead of just a destination—the solution was obvious.
.
.
.
Why This Actually Works (The Part Where I Get a Little Philosophical)
Both of these methods work because they address the same fundamental gap in Claude Code debugging: AI doesn’t fail because it’s not smart enough. It fails because it can’t see what you see.
When you’re debugging, you have browser dev tools, console logs, network requests, and actual behavior unfolding on your screen. Claude has code files.
That’s it.
It’s working with incomplete information and filling the gaps with educated guesses.
Here’s the mindset shift that changed everything for me:
👉 Stop expecting AI to figure it out. Start helping AI see what you see.
You become the eyes. AI becomes the analytical brain that processes patterns and proposes solutions based on the evidence you feed it.
It’s a collaboration. A partnership. Not a vending machine where you insert a problem and expect a solution to drop out.
When to Use Logging
Add logs when the bug involves:
Data flow and state management
Timing issues and race conditions
Lifecycle problems in React, Vue, or similar frameworks
Anything where the sequence of events matters
The logs transform invisible runtime behavior into visible evidence.
React’s useEffect, state updates, and re-renders happen in milliseconds—too fast to trace mentally, but perfectly captured by console.log. Feed those logs to Claude, and suddenly it can see the movie instead of just reading the script.
When to Use ASCII Brainstorming
Use the brainstorming approach when:
Error messages point to the wrong location
The bug could originate from multiple places
You’ve already tried the obvious fixes (twice)
The problem involves cascading effects across systems
Asking Claude to brainstorm with diagrams forces it to slow down and map the problem systematically. It prevents the merry-go-round where AI keeps trying variations of the same failed approach. By exploring multiple angles first, you often find the root cause on the very first real attempt.
.
.
.
The Line Worth Tattooing Somewhere (Metaphorically)
Here’s what I want you to take away from all of this:
Don’t argue with AI about what it can’t see. Show it.
The next time Claude can’t solve a bug after a few rounds, resist the urge to rephrase your complaint. Don’t add more adjectives. Don’t type in all caps. (I know. I KNOW. But still.)
Instead, ask yourself: “What am I seeing that Claude isn’t?”
Then find a way to bridge that gap—through logs, through diagrams, through screenshots, through any method that gives AI the visibility it needs to actually help you.
.
.
.
Your Next Steps (The Warm and Actionable Version)
For state and timing bugs:
Pause. Take a breath. Step off the merry-go-round.
Ask Claude to add logging that tracks the data flow.
Run your test, copy the console output, paste it back to Claude.
Watch Claude solve in one shot what it couldn’t guess in twenty.
For complex, cascading bugs:
Paste the error message (yes, the whole confusing thing).
Add: “Let’s brainstorm ways to debug this. Use ASCII diagrams.”
Let Claude map the problem before it tries to solve it.
Pick the most surgical approach from the options it generates.
That bug that’s been driving you up the wall? The one Claude keeps missing?
I was bouncing between ChatGPT Pro, Claude web, and Cursor like a pinball with a deadline. Copy from o1 pro. Paste into my editor. Fix the bug it introduced. Pray it works. Try Cursor for a second opinion. Watch it rewrite my entire file when I asked for one measly line.
Rinse. Repeat. Question your life choices.
(We’ve all been there. And if you say you haven’t, well, I’m not sure I believe you.)
Then May hit. Anthropic added Claude Code to their Max plan—same $200/month I was already burning on ChatGPT Pro, but now I could stop copy-pasting and start orchestrating.
That shift changed everything.
Here’s the thing: I wrote 30+ articles this year documenting every breakthrough, every spectacular failure, every “wait, that’s how it’s supposed to work?” moment. If you only read one piece from me in 2025—make it this one.
What follows are the 4 immutable laws of Vibe Coding I discovered this year. They turned chaotic AI sessions into systematic, predictable wins. Once you see them, you can’t unsee them.
Ready? Let’s go.
.
.
.
Rule #1: The Blueprint Is More Important Than The Code
Let me tell you about the single biggest mistake I see developers make.
They type “build me a task management app” and hit Enter. Claude generates code. Components. Database schemas. Authentication logic.
And then… it’s nothing like what they imagined.
They blame the AI. “It hallucinated again.”
But here’s what I’ve learned after shipping dozens of projects with Claude Code: hallucinations are usually just ambiguity in your prompt. That’s it. That’s the secret nobody wants to admit.
AI is a terrible architect. Give it vague instructions, and it fills in the blanks with whatever patterns it’s seen most often. (Which, spoiler alert, aren’t YOUR patterns.)
But AI is an amazing contractor.
Give it clear blueprints—specific requirements, explicit constraints, visual references—and it executes with surgical precision. Like a really talented carpenter who just needs you to stop saying “make it nice” and start handing over actual measurements.
The technique: Interview yourself first
Instead of asking Claude to “build me an app,” I use a brainstorming prompt (inspired by McKay Wrigley and Sabrina Ramonov) that flips the entire script.
The AI interviews me.
“What’s the core problem this solves?”
“Who uses it?”
“What does the main screen look like?”
“What happens when the user clicks X?”
By the time I’ve answered those questions, I’ve got a Product Requirements Document. Not AI-generated slop—my vision, clarified.
Claude becomes the junior dev who asks great questions before writing a single line of code. I stay the architect who actually understands what we’re building.
(This is the way it should be.)
The secret weapon: ASCII wireframes
Text descriptions get misinterpreted. Every. Single. Time.
You say “a sidebar with navigation.” Claude hears “full-width hamburger menu.”
So I started including ASCII art wireframes in my prompts:
Sounds primitive, right? Almost embarrassingly low-tech.
The results say otherwise.
When I started including visual plans, my first-try success rate hit 97%. Claude understood layout and hierarchy immediately. No more “that’s not what I meant” rewrites. No more three rounds of “closer, but still wrong.”
👉 The takeaway: Stop typing code and start drawing maps. The blueprint is where the real work happens.
Rule #2: Separate The “Thinker” From The “Builder”
At the beginning, I was using Claude Code for everything.
Planning. Building. Reviewing. Debugging.
One model to rule them all.
And it almost worked.
Almost.
But I kept running into the same problems. Claude would rewrite perfectly good code. Add complex abstractions I never asked for. Solve a simple bug by restructuring half my app.
I asked for email OTP login. I got a 12-file authentication framework.
I asked to fix a type error. Claude decided my entire architecture was wrong.
(It wasn’t. I promise you, it wasn’t.)
The discovery: Specialized roles
Then I stumbled onto a workflow that changed everything—and honestly, I felt a little silly for not seeing it sooner.
Use one model to think. Use another to build.
For me, that’s GPT-5/Codex (The Thinker) and Claude Code (The Builder).
Codex asks clarifying questions. It creates comprehensive plans. It reviews code like a senior engineer who’s seen every possible edge case and still remembers them all.
Claude Code executes. Fast. Reliably. It handles files, terminal commands, and edits without wandering off into philosophical debates about code architecture.
Together? Magic.
The review loop
The workflow looks like this:
Plan (Codex): Describe what I want to build. Codex asks questions, creates a detailed implementation plan.
Build (Claude Code): Feed the plan to Claude. Let it execute.
Review (Codex): Paste the implementation back to Codex. It checks against the original plan, catches bugs, finds edge cases.
That third step—the review loop—catches issues that single-model workflows miss every time. EVERY time.
Taming the overengineering monster
Claude has a tendency to overcomplicate. It’s well-documented at this point. (If you’ve used it for more than a week, you know exactly what I’m talking about.)
My fix? The Surgical Coding Prompt.
Instead of “add this feature,” I tell Claude:
“Analyze the existing patterns in this codebase. Implement this change using the minimal number of edits. Do not refactor unless explicitly asked. Show me the surgical changes—nothing more.”
From 15 files to 3 files. From 1000+ lines to 120 lines.
Same functionality. 90% less complexity.
👉 The takeaway: Treat your AI models like a team, not a swiss-army knife. Specialized roles produce specialized results.
“Why do I keep explaining the same patterns over and over?”
Every new project, I’d spell out my authentication approach. My database schema conventions. My error handling patterns. Every. Single. Time.
Claude would forget by the next session. Sometimes by the next prompt.
I was treating AI like a goldfish with a keyboard.
(No offense to goldfish. They’re trying their best.)
The “I know kung fu” moment
Then Claude launched Skills—and everything clicked.
Skills let you package your coding patterns into reusable modules. Instead of explaining “here’s how I do authentication” for the 47th time, you create an auth-skill. Enable it, and Claude instantly knows your entire implementation.
The exact patterns. The exact folder structure. The exact error messages.
Every project uses the same battle-tested approach. Zero drift. Zero “well, last time I used a different library.”
It’s like downloading knowledge directly into Claude’s brain.
Matrix-style. (Hence the name.)
Building your first skill
The process is stupidly simple:
Take code that already works in production
Document the patterns using GPT-5 (it’s better at documentation than execution)
Transform that documentation into a Claude Skill using the skill-creator tool
Deploy to any future project
The documentation step matters. GPT-5 creates clean, structured explanations of your existing implementations. Claude Skills uses those explanations to replicate them perfectly.
The compound learning effect
Here’s where it gets really interesting.
I built an Insights Logger skill that captures lessons while Claude “code”. Every architectural decision, every weird bug fix, every “oh that’s why it works that way” moment—automatically logged.
At the end of each session, I review those insights. The good ones get promoted to my CLAUDE.md file—the permanent knowledge base Claude reads at the start of every project.
Each coding session builds on the last. Compound learning, automated.
👉 The takeaway: Prompting is temporary. Skills are permanent. If you’re explaining something twice, you’re doing it wrong.
Rule #4: Friction Is The Enemy (So Automate It Away)
Let me describe a scene you’ll recognize.
You’re deep in flow state. Claude Code is humming along. Building components, wiring up APIs, making real progress.
And then:
Allow Claude to run `npm install`? [y/n]
You press Enter.
Allow Claude to run `git status`? [y/n]
Enter.
Allow Claude to run `ls src/`? [y/n]
Enter. Enter. Enter. Enter. Enter.
By prompt #47, you’re not reading anymore. You’re a very tired seal at a circus act nobody asked for.
(Stay with me on this metaphor—it’s going somewhere.)
Anthropic calls this approval fatigue. Their testing showed developers hit it within the first hour of use.
And here’s the terrifying part: the safety mechanism designed to protect you actually makes you less safe. You start approving everything blindly. Including the stuff you should actually read.
The sandbox solution
Claude Code’s sandbox flips the entire model.
Instead of asking permission for every tiny action, the sandbox draws clear boundaries upfront. Work freely inside them. Get blocked immediately outside them.
On Linux, it uses Bubblewrap—the same tech powering Flatpak. On macOS, it’s Seatbelt—the same tech restricting iOS apps.
These boundaries are OS-enforced. Prompt injection can’t bypass them.
Claude can only read/write inside your project directory. Your SSH keys, AWS credentials, shell config? Invisible. Network traffic routes through a proxy allowing only approved domains.
You run /sandbox, enable auto-allow mode, and suddenly every sandboxed command executes automatically. No prompts. No friction. No approval fatigue.
The 84% reduction in permission prompts? Nice. The kernel-level protection that actually works? Essential.
Parallel experimentation with Git Worktrees
Here’s another friction point that kills vibe coding: fear of breaking the main branch.
My fix: Git Worktrees with full isolation.
Standard worktrees share your database. They share your ports. Three AI agents working on three features leads to chaos. (Ask me how I know.)
I built a tool that gives each worktree its own universe. Own working directory. Own PostgreSQL database clone. Own port assignment. Own .env configuration. Now I run three experimental branches simultaneously. Let three Claude instances explore three different approaches. Pick the winner. Delete the losers.
No conflicts. No fear. No “let me save my work before trying this crazy idea.”
👉 The takeaway: Safe environments allow for dangerous speed. Eliminate friction, and experimentation becomes free.
Ready to set it up?
Claude Code Sandbox Explained walks through the complete configuration—including battle-tested configs for Next.js, WordPress, and maximum paranoia mode.
The Synthesis: What Separates Hobbyists From Shippers
These 4 rules are what separate “people who play with AI” from “people who ship software with AI.”
Rule #1: The blueprint is more important than the code.
Rule #2: Separate the thinker from the builder.
Rule #3: Don’t just prompt—teach skills.
Rule #4: Friction is the enemy.
Each rule builds on the last.
Clear blueprints feed into specialized models. Specialized models benefit from reusable skills. Reusable skills only matter if friction doesn’t kill your flow.
It’s a system. Not a collection of random tips.
Where to start
Don’t try to implement all four at once.
That’s a recipe for burnout.
Start with Rule #4. Enable the sandbox. Regain your sanity. Stop being a tired circus seal.
Then move to Rule #1. Before your next feature, write the PRD first. Interview yourself. Draw the ASCII wireframe.
Rule #2 and Rule #3 come naturally after that. You’ll feel the pain of overengineering (and want specialized roles). You’ll get tired of repeating yourself (and want skills).
The system reveals itself when you need it.
Your challenge for 2026
Pick one project you’ve been putting off. Something that felt too complex for AI assistance.
Apply Rule #1: Write the blueprint first. ASCII wireframes and all.
Apply Rule #4: Set up the sandbox before you start.
Then let Claude execute.
Watch what happens when AI has clear boundaries and clear instructions. Watch how different it feels when you’re orchestrating instead of babysitting.
What will you build first?
Here’s to an even faster 2026.
Now go ship something.
This post synthesizes a year’s worth of vibe coding experimentation. Browse the full archive to dive deeper into any technique—from CLAUDE.md setup to sub-agent patterns to WordPress automation.