What Went Wrong With My First AI Coding Workflow

By Rakshit Yadav (@yadavrakshit60)•Aug 2026•10 min read

The first workflow felt magical for a week and expensive for a month. I was generating more code than I could read, in chats that remembered the wrong things, with no file that said what "done" meant. I did not lose the repo. I lost time undoing confident mistakes.

This is a post-mortem, not a confession contest. Six failures. Six files or habits that exist because of them. I am not inventing an outage. I am explaining why AgenticKit's guardrails look the way they do.

The takeaway

Most AI-coding failures are process bugs: unbounded chats, no definition of done, no separation of write and review. Models make that cheaper to hit. They do not invent the shape.

Pick one failure you hit this month and add one guardrail. Not a new tool.

Failure 1: One chat to rule them all

I kept yesterday's thread alive because it "already knew the app." After enough turns the agent started editing files I had not named. Official Cursor guidance now says to start a new conversation when you change tasks or the agent repeats itself. I learned that by watching a billing change collide with a rename from three prompts ago.

What exists because of this: a habit, then a written rule. New feature, new chat. Point at the plan file. Do not paste the transcript. The shipping workflow is that habit spelled out.

Failure 2: Rules as a novel

I pasted a style guide into one always-on file and wondered why tenant isolation still did not stick. The model kept the nearby code and dropped bullet 47. Cursor's own docs cap a rule around 500 lines and tell you not to dump a style guide. I had done both.

What exists because of this: a small always-on core and glob-scoped API / DB / UI rules. The rules article is the corrected stack.

Failure 3: "Add billing" with no contract

I typed the product idea and accepted a Checkout session, a webhook, and a new settings IA. There was no list of events we handle, no idempotency table, no "not in this slice." Half of Stripe landed. The other half was implied.

What exists because of this: /feature-spec before /ship on anything M or L. Complexity labels. Skip flags in CURSOR.md so a billing slice can check Skip UI if you are only doing the webhook.

Failure 4: The agent reviewed its own work

I asked the same chat "any security issues?" and it said it looked good. Of course it did. It had just written the code.

What exists because of this: code-reviewer and security-auditor as separate desks. Security returns PASS or FAIL. FAIL stops /ship. The tech-lead playbook forbids that role from writing production code, for the same reason a planner who implements will skip the plan. The error-handling skill is the written version of “do not swallow the exception.”

Failure 5: Tenant and webhook landmines

Two classes of bug showed up more than once:

  • A list endpoint that trusted organizationId from the client
  • A Stripe webhook that parsed JSON before verifying the signature, or processed the same event.id twice

Those are not exotic. They are the default if you do not write them down.

What exists because of this: required { error, code }, tenant-from-session in the core rule, webhook rules in the API file (raw body, signature, persist event.id, return 200). A required tenant isolation test in the tests rule. /audit before a release.

HN has a version of this story where an agent nearly ran a destructive query against production. I did not need that version. The IDOR was enough.

Failure 6: Marketing in the same context as SQL

I asked for a launch thread in the chat that had just written a migration. The thread used the migration's nouns. The tweet read like a changelog. The next engineering turn picked up "punchy" and put it in a commit message.

What exists because of this: two kits. Engineering commands and marketing commands do not share a context window unless you install both and then still start a new chat. How the 34 agents are structured is the org chart behind that split.

The corrected process

Short version, with the file that each step is for:

  1. Fill CURSOR.md. Human-owned product truth, scan-owned stack. Memory setup.
  2. Always-on rule: read that file, keep the error shape, tenant from session.
  3. Spec S/M/L. Split L.
  4. /ship (or the hand pipeline) for one slice. Writer roles in order.
  5. Tests including tenant isolation. Reviewer is not the writer. Security is PASS/FAIL.
  6. Scripts from CURSOR.md. Red means HOLD.
  7. Marketing in a different chat, preferably a different kit.

The rebuild story is what a week looks like after this. I am not claiming the first month was clean. I am claiming each guardrail maps to a repeatable failure, not to a blog trend.

What I still will not automate

  • Deciding the product is for someone
  • Merging a diff I have not read
  • Production data changes
  • Anything I would not give a new hire on day one without a supervisor

If you only add one thing after reading this, add the tenant test and the new-chat habit. They catch more than a new model switch.

Related Tools & Agents

🛠️ Free Tool: rules-grader🤖 Agent: security-auditor🤖 Agent: code-reviewer🤖 Agent: debugger⚡ Command: /audit⚡ Command: /fixSkill: security-basicsSkill: error-handling

Audit the routes you already have

The /audit command is the guardrail from failure 5, written down.

Open /audit