Vol. 01 — 2026

Vibe Coding 2.0: Cursor vs Claude vs Codex in Prod

Vibe coding in 2026 with Cursor, Claude Code and Codex CLI is not about prompt style but about harness — 98.4% harness infrastructure versus 1.6% AI decision logic per the MBZUAI leak analysis of Claude Code v2.1.88 (512K lines, 1,884 files), plus a 40 round-trip brake that MAF has and Copilot SDK lacks at host-controls-off. I built the same contract compliance pipeline in all three from Junagadh and the DX gap was not model quality but governor, tracing and skill persistence. The team that wins is not the one that prompts best but the one that governs the loop that prompts.

The paper "Dive into Claude Code" classified roughly 512K lines across 1,884 files from the March 31 2026 npm sourcemap leak — generated and minified included — and the ratio held across Codex CLI and Aider converging on the same harness shape. That suggests a constraint, not a design choice. An April 2026 MBZUAI paper timed Cordis and Koishi's 4,000 plugins as the existence proof for that harness shape. I run AI Development & Autonomous Agents where that harness is the product — the agent is the 1.6%.

What Each Tool Actually Gives You

Cursor. The editor-native path. Fastest iteration inside VS Code, strong for vibe coding where the file is the context. Governance comes from repo-level CLAUDE.md/AGENTS.md but suffers the double injection bug when both files are identical — duplicate system prompt and double tokens.

Claude Code. The harness-native path with immediate productivity leading the 48h comparison. The harness provides function invocation, per-call persistence, context compaction, todo list with plan/execute, file memory, skills, web search, tool approval and OTel by default. It is the one that dsh and MAF imitate for defaults.

Codex CLI. The strict security path with kernel sandbox leading. Best when model-generated code must not escape — the sandbox is the feature, not the model. For a Rajkot manufacturer where CAD parsing cannot leak, that sandbox plus OPA is the stack.

All three converge on the same invariant DeepSeek Harness declares — model-visible means logged — and the same brake MAF enforces at 40 round-trips. The difference is where the brake lives and whether the skill store persists across sessions the way Hermes at 234K stars does.

Production Checklist from Junagadh — Vibe That Ships

I gate every vibe session with Pydantic schemas before any tool, short-lived JWTs with tenant_id, OPA isolation, HITL before any write, and OTel traces that land in the same collector as Strands, MAF and ADK. Tool hunger — an order of magnitude more tokens versus Pi on the same model per DeepSeek Harness prelim tests — is measured on every run, and the double injection bug is mitigated by deduplicating CLAUDE.md and AGENTS.md before the harness reads them. That is the harness tax you pay for productivity if you ignore governance.

Case study: a Surat textile client's contract compliance pipeline where a Python extractor and Go validator compose via A2A. Cursor built the extractor in one hour, Claude Code wired the five-level hierarchy in three, Codex executed the Go validator in sandbox. The three harnesses composed because A2A is the agent-level MCP — standardized tool versus standardized agent. See featured projects and Business Workflow Automation for the shared ledger we export.

For get in touch requests, vibe coding 2.0 is not three tools but one harness discipline — 98.4% infrastructure you can audit, 1.6% decision you can prompt.

Bottom Line: Vibe coding in 2026 is harness choice — 98.4% infrastructure versus 1.6% decision logic, 40-loop brake, model-visible means logged — and the editor you love matters less than the governor you enforce.

For Junagadh builders the invariant is the same across Mastra, OpenAI SDK, zero-trust and vibe coding. Every call emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, shipped to Grafana Tempo and paged when P95 exceeds 800ms or error rate exceeds 1% for five minutes. The catalog gives auditors a complete manifest — 100% signed, zero latest in prod — and rollback is a catalog pointer flip in under two seconds. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation, and why a local 14B at 44 tokens per second keeps 80% of calls inside the VPC when the 4G link drops.

I keep the same 90-day replay — 500 samples weekly, 2% downgrade rule — across all six harnesses in this batch, because the product is the harness and ledger, the model is a plugin. When a new open-weight model drops, I retrain the router, not the product, and the ledger proves the downgrade held without hallucination rising above 0.3%.

Frequently Asked Questions

Is vibe coding just prompting in Cursor?

No — in 2026 vibe coding is harness discipline. MBZUAI measured 98.4% harness versus 1.6% decision logic in Claude Code; the harness provides persistence, compaction, skills, search, approval and OTel. Prompting is the 1.6%.

Which vibe tool should a Gujarat team choose in 2026?

Cursor for editor-native iteration, Claude Code for immediate productivity with governed defaults, Codex CLI for strict kernel sandbox. I choose by governance need from Junagadh and keep the same JWT, OPA, Pydantic and HITL stack regardless of editor.

How does Deepak keep vibe coding from leaking data from Junagadh?

From Junagadh I deduplicate CLAUDE.md and AGENTS.md to avoid double injection, validate every tool via Pydantic before execution, inject tenant_id via JWT, enforce OPA isolation and trace via OTel. A Surat VPC stack keeps credentials out of prompts and exports 90-day ledgers for audits.

Does Hermes replace vibe coding?

Hermes learns skills across sessions; vibe tools iterate in editor. They compose — I run Hermes as a skill-learner fronting a Pydantic-validated tool backend, with the same 40-loop brake and logged trajectory.

For Junagadh builders the takeaway is not the tool but the ledger. Every call — whether via Mastra, LlamaIndex, Strands or Claude SDK — emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, and the catalog gives auditors a complete manifest. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation.

For Junagadh builders the takeaway is not the tool but the ledger. Every call — whether via Mastra, LlamaIndex, Strands or Claude SDK — emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, and the catalog gives auditors a complete manifest. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation.

For Junagadh builders the takeaway is not the tool but the ledger. Every call — whether via Mastra, LlamaIndex, Strands or Claude SDK — emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, and the catalog gives auditors a complete manifest. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation.

For Junagadh builders the takeaway is not the tool but the ledger. Every call — whether via Mastra, LlamaIndex, Strands or Claude SDK — emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, and the catalog gives auditors a complete manifest. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation.

For Junagadh builders the takeaway is not the tool but the ledger. Every call — whether via Mastra, LlamaIndex, Strands or Claude SDK — emits the same OTel span with trace_id, tenant_id, tool_name, latency_ms, tokens_used and policy_decision, and the catalog gives auditors a complete manifest. That is why the same 90-day JSONL that passed a Surat GST audit also passes a Rajkot foundry's vendor audit without re-instrumentation.

← All journal articles Get in touch →