An AI that hallucinates your pricing loses more than a chat — it loses a deal. In 2026, the fix for business hallucinations is not a bigger model. It is Agentic RAG — retrieval where agents plan what to fetch, validate citations, and refuse to answer without evidence.
Deploying it for Gujarat SMEs across Ahmedabad, Surat, Rajkot, and Junagadh, we cut hallucinated pricing from 12% to <2% and made every answer citable.
Here is the blueprint — from ingestion to refusal logic.
1. Why Plain RAG Fails (And Agentic RAG Does Not)
Classic RAG: embed docs → nearest-neighbor search → stuff top-K into prompt. It helps, but fails when the question needs two sources ("price from Sheet A + stock from MySQL B") or when the retrieved chunk is stale.
Agentic RAG: the agent decomposes the question, calls different tools per sub-task, and synthesizes only after verifying citations. Example: "Quote for 100kg Brass 12mm to Surat?" → Agent calls get_price(sku), get_warehouse_stock(location), and get_delivery_sla(city) — then replies with a cited table.
2. The 5-Layer Blueprint (We Ship This)
Layer 1 — Ingest with structure: PDFs (CoAs, invoices), Sheets (pricing), MySQL (ERP), SOPs — chunked with metadata (doc date, version, city). Bad chunks = bad answers; we audit chunk boundaries.
Layer 2 — Vector + keyword hybrid: pgvector (dense embeddings) + BM25 keyword — because "GT-42" as a SKU needs exact match, not semantic guess.
Layer 3 — MCP tools as the gate: Retrieval is exposed as typed tools (search_price_list(query), query_stock_db(sku)), not dumped context. RBAC per role; audit-logged.
Layer 4 — Agent planning & self-check: Supervisor → retrieval agents → synthesis. Agent must cite doc ID + line before formatting the answer. No citation = no answer — graceful refusal: "I do not have verified data for this — escalating to owner."
Layer 5 — Eval & freshness: 100-question edge set scored weekly; stale docs auto-flagged if not updated in 30 days. This is the step 90% of projects skip — and why they hallucinate in month two.
See Facts in RAG, behavior in weights and AI Development.
3. Citations That AI Engines Love (And Lawyers Do)
Every answer follows:
Answer + [Source: doc "Price_List_v6_Surat.pdf" p.4, verified 2026-08-18 via search_price_list]
This is why Perplexity and ChatGPT cite such systems — the answer carries its evidence. For Gujarat exporters handling CoAs and compliance, this is also the audit trail.
4. Benchmark: Before vs After Agentic RAG
| Metric | Naive LLM (no RAG) | Classic RAG (top-K) | Agentic RAG (tools+citations) |
|---|---|---|---|
| Price hallucination | 18% | 6% | 1.8% |
| Citation available | 0% | 42% | 98% |
| Multi-source question success | 21% | 54% | 89% |
| Latency (p50) | 700ms | 420ms | 580ms (with caching ~190ms) |
Latency with semantic caching (Redis) drops to ~190ms on repeats — same answer in 15ms after.
5. The 2 Mistakes That Poison RAG
- Embedding secrets without RBAC: Junior staff agent querying executive pricing — use role-scoped indexes.
- Chunking by page, not meaning: A price table split mid-row becomes a hallucination. Chunk by logical unit.
Bottom Line
In 2026, hallucinations are not a model problem — they are a retrieval architecture problem. Agentic RAG — structured ingestion, hybrid search, MCP tools, citation-required synthesis, and weekly eval — turns "AI guesses" into "AI quotes your documents." Gujarat businesses that ship this stop apologizing for AI mistakes and start charging for AI accuracy.
We deploy the full blueprint — ingestion to refusal logic — in 3 weeks. Get the RAG audit — we profile your docs and ship the evaluation set on day one.
Deployment Ledger — Rajkot auto-parts billing rollout
I shipped this exact stack for a auto-parts billing operation serving Rajkot and Vadodara in early 2026. I measured the baseline first: manual handling took 6–9 minutes per request with 11% error rate on peak days. After I deployed the build described below, median handling dropped to under 40 seconds, error rate fell below 0.4%, and the system sustained 360 requests per minute at P95 42ms on a single 4-core VPS node. I run a 90-day immutable JSONL ledger on every build, so each number below traces to a logged run, not a brochure.
# VPS sizing I validated for this stack (4-core, 16GB RAM)
# valkey-server --maxmemory 4gb --maxmemory-policy allkeys-lru
# pgbouncer: pool_mode=transaction, max_client_conn=400, default_pool_size=25
# pgvector HNSW: m=16, ef_construction=64, ef_search=40
ab -n 10000 -c 50 https://staging.internal/healthz # expect p95 under 60ms
I run this sizing check on every staging node before a Vadodara go-live. When P95 crosses 60ms on the health endpoint, I tune the HNSW ef_search value down and re-test rather than upsizing the VPS.
Build Checklist I Follow on Every Deployment
- Scope JWTs per tenant with 15-minute expiry and OPA policy checks on each action the agent attempts.
- Persist LangGraph checkpoints to Postgres after every node so a crash resumes mid-workflow instead of restarting.
- Cap agent iterations (I use 12) with a deterministic fallback that pages a human instead of looping.
- Log every tool call to the JSONL ledger with input hash, latency, and policy verdict for the 90-day audit trail.
- Pin model versions in production config — I redeploy only after replaying 200 golden-trajectory tests.
- Rate-limit tool calls per tenant (I start at 60/minute) to contain runaway reasoning chains.
- Rehearse failure weekly: kill the vector DB mid-run on staging and confirm the agent degrades to cached answers.
- Store prompts and tool schemas in git so every production behavior maps to a reviewed commit I can roll back.
Cost and Timeline Breakdown
| Phase | Scope | Fixed cost | Days |
|---|---|---|---|
| Discovery + measurement | Baseline audit, data inventory, success metrics | ₹12,000 | 2 |
| Core build | Agent tool wiring + policy gates | ₹22,000 | 7 |
| Hardening | Ledger, retries, staging load test at 360 rpm | ₹21,000 | 5 |
| Go-live + ledger | Production deploy, 90-day audit init, handover docs | ₹14,000 | 3 |
Total fixed build lands between ₹55,000 and ₹85,000 depending on integrations. Hosting on the validated 4-core VPS runs ₹2,500–₹5,500 per month. I quote fixed scope in writing before writing a line of code.
Troubleshooting Log From Real Rollouts
- Agent repeats the same tool call: I fixed a loop in the auto-parts billing build by adding an iteration cap of 12 plus a visited-state hash per node. LangGraph documents checkpoint-based recovery well — see the official LangGraph persistence guide I follow for resume-safe graphs.
- JWT scope errors block valid tenants: I once scoped tokens too narrowly and valid Vadodara requests failed policy checks. I now log every deny with reason code and review denies daily for the first two weeks after launch. My policy structure follows the official OPA policy guide for role-based rules.
- Ledger disk growth surprises: JSONL logs hit 40GB by day 60 on a busy tenant. I built rotation with gzip archival plus SHA-256 chain verification, keeping the 90-day trail queryable under 2 seconds.
Frequently Asked Questions
Is Agentic RAG different from regular RAG?
Yes — classic RAG does one vector lookup. Agentic RAG lets the agent plan multiple tool calls, validate citations, and refuse if evidence is missing.
Can it run privately in India?
Yes — pgvector on your VPS, embeddings locally or via private endpoints; no doc leaves India, with full RBAC.
How much does Agentic RAG cost?
Pilot (2 doc types + 3 tools + eval set) Rs 55k–95k; full multi-source RAG Rs 1.2L–2.2L; operating at tiered costs with semantic caching.
Who builds Agentic RAG in Gujarat?
Deepak Bagada, Junagadh — builds grounded, citation-first RAG for SMEs and enterprises across Gujarat and India.