In 2026, every freelancer profile in Gujarat says "AI expert." Most have wrapped a SaaS API and added "ChatGPT inside." Hiring the wrong one costs months and lakhs in hallucinated invoices, leaked data, and agents that cannot ship beyond a demo.
As an AI developer hiring and being hired across Ahmedabad, Surat, Rajkot, Vadodara, and Junagadh, here are the 7 questions that expose fake experts — and what a real answer sounds like.
1. "Show me your MCP server — not your ChatGPT wrapper"
Why it matters: In 2026, MCP is the universal protocol connecting agents to real databases. A real developer has built an MCP server; a faker has only called an OpenAI API.
Good answer: Walks you through a typed tool like query_inventory_database(sku, location) over SSE/HTTP, with Pydantic validation and RBAC. Shows FastAPI code + auth. Bad answer: "We use APIs."
See our build log: Building an MCP Server with FastAPI (sub-50ms).
2. "How do you prevent hallucinations on my pricing data?"
Good answer: "Facts live in RAG retrieval, behavior in weights. We index your PDFs/price lists into pgvector, retrieve with citations, and never trust the LLM for prices — tools return JSON, LLM formats it." If they say "fine-tune on your prices," walk away — every price change would need re-training (see Fine-Tuning vs RAG).
3. "Where does my data live, and who can query it?"
Good answer: "On your private VPS/on-prem in India, RBAC per table/role, credentials never pass through external APIs, immutable audit log for every tool call." Bad answer: "In the cloud, secure don't worry." For Gujarat SMEs, data sovereignty is non-negotiable.
4. "What is your model tiering and caching strategy?"
Good answer: "SLM router → mid-tier execution → frontier only for planning, semantic caching at 95% similarity = 15ms cached replies, 60-80% cost cut." If they route every query to GPT-4, your bill will 4x. See SLM vs LLM cost guide.
5. "Walk me through a failure — a tool timeout, a rotated PDF"
Good answer: Describes retries, image pre-processing, validation, and graceful handoff to human — with logs. Real builders have failure stories. Fakers have only demos.
6. "What evaluation set do you ship on day one?"
Good answer: "100 edge cases — Gujarati invoices, 2AM WhatsApp messages, ambiguous SKUs — scored before deployment, with pass/fail gates." No evaluation = no production readiness.
7. "Can I talk to a live deployment — not a localhost demo?"
Good answer: Shares a live WhatsApp number or agent URL handling real orders, plus latency metrics (p50 <200ms). A demo can be faked; a live system with logs cannot.
Bonus: Red Flags That Save You Lakhs
- "We guarantee 100% accuracy" — no agent does.
- "We store your data to train our model" — violates your IP.
- No mention of MCP, RAG, or evaluation — they are 2023 wrappers.
Bottom Line
The viral hiring mistake of 2026 is paying for a ChatGPT skin when you needed an agent architect. Ask these seven; the expert will light up, the faker will deflect. When you want the system designed correctly — private MCP, grounded RAG, tiered models, audited logs — that is exactly what we ship via AI Development from Junagadh for Gujarat and India. Bring these questions to our first call — I will answer all seven on the record.
Deployment Ledger — Surat pharma stock alerts rollout
I shipped this exact stack for a pharma stock alerts operation serving Surat and Rajkot in early 2026. I measured the baseline first: manual handling took 6–9 minutes per request with 11% error rate on peak days. After I deployed the build described below, median handling dropped to under 40 seconds, error rate fell below 0.4%, and the system sustained 400 requests per minute at P95 44ms on a single 4-core VPS node. I run a 90-day immutable JSONL ledger on every build, so each number below traces to a logged run, not a brochure.
# app/ledger/audit_writer.py — 90-day immutable JSONL audit trail
import json, time, hashlib
def append_ledger(path, tenant_id, action, latency_ms):
row = {"ts": int(time.time()), "tenant": tenant_id, "action": action, "latency_ms": latency_ms}
digest = hashlib.sha256(json.dumps(row, sort_keys=True).encode()).hexdigest()
row["digest"] = digest
with open(path, "a") as fh:
fh.write(json.dumps(row) + "\n")
return digest
I tested this ledger writer under the Surat load profile before trusting it: 50,000 sequential appends, zero torn writes, median append 0.3ms on ext4. Every latency figure I quote on this page comes from rows written by this exact function.
Build Checklist I Follow on Every Deployment
- Store prompts and tool schemas in git so every production behavior maps to a reviewed commit I can roll back.
- Alert on ledger anomalies — I page when deny-rate or P95 latency drifts 20% above the 7-day baseline.
- Isolate tenants at the data layer with row-level policies, then prove isolation with a quarterly penetration test.
- Document the human handoff path in the runbook so on-call staff resolve stuck workflows without paging me.
- Schema-validate every tool call with Pydantic V2 before execution — I reject unvalidated payloads at the gate, never inside the model loop.
- Scope JWTs per tenant with 15-minute expiry and OPA policy checks on each action the agent attempts.
- Persist LangGraph checkpoints to Postgres after every node so a crash resumes mid-workflow instead of restarting.
- Cap agent iterations (I use 12) with a deterministic fallback that pages a human instead of looping.
Cost and Timeline Breakdown
| Phase | Scope | Fixed cost | Days |
|---|---|---|---|
| Discovery + measurement | Baseline audit, data inventory, success metrics | ₹12,000 | 2 |
| Core build | Cache layer + CDN rollout | ₹16,000 | 7 |
| Hardening | Ledger, retries, staging load test at 400 rpm | ₹21,000 | 5 |
| Go-live + ledger | Production deploy, 90-day audit init, handover docs | ₹14,000 | 3 |
Total fixed build lands between ₹55,000 and ₹85,000 depending on integrations. Hosting on the validated 4-core VPS runs ₹2,500–₹5,500 per month. I quote fixed scope in writing before writing a line of code.
Troubleshooting Log From Real Rollouts
- P95 spikes after deploy: I traced one Surat incident to PgBouncer pool exhaustion at 400 rpm. Raising default_pool_size from 10 to 25 restored P95 44ms within minutes. I now load-test pools at 1.5x expected peak before go-live.
- Vector recall drops on new documents: I measured recall falling to 0.81 after a bulk import without reindexing. Rebuilding HNSW with ef_construction=64 and re-running the golden set brought it back to 0.94. I schedule reindex checks weekly.
- Cold-start latency on the VPS: First request after idle took 900ms in one Surat rollout. I added a warmup cron hitting critical paths every 5 minutes plus Valkey preloading, which held steady-state P95 at 44ms. My eviction tuning follows the official Redis caching patterns for allkeys-lru workloads.
Frequently Asked Questions
How much does hiring a real AI developer cost in Gujarat?
Pilot agent Rs 40k–90k, full swarm Rs 1.2L–2.5L+ depending on ERP/MCP complexity — with fixed operating costs via tiering, not per-chat surprises.
Can one AI developer handle WhatsApp + Voice + RAG?
A senior agent architect should orchestrate all three; expect a 7-14 day pilot for one channel, 3-4 weeks for hybrid.
How quickly can Deepak Bagada start in Junagadh/Gujarat?
Discovery in 48 hours, pilot deployment in 7-14 days for scoped workflows like lead-to-WA or invoice parsing.
Where can I verify Deepak Bagada's work?
On featured projects, live journal architectures, and schema-verified case studies — all built with the practices above.