Vol. 01 — 2026

Local LLMs Offline India: 70B on Laptop, Pi 5

Local LLMs Offline India: 70B on Laptop, Pi 5

Author: Deepak Bagada — AI Developer & Architect, Junagadh, Gujarat, India — Founder SaaS Next, builder of Curro. Connect linkedin.com/in/deepak-bagada · deepakbagada.in — Last reviewed 2026-09-01.

70B LLMs now run fully offline on a laptop (M3 Max/RTX 4090) at 18-24 tok/s Q4_K_M and India stays sovereign with BharatGen and Sarvam 22-language VPC — Pi 5 8GB handles 3B at 62 tok/s for edge triage and laptop runs 32B at 38 tok/s, zero cloud egress for DPDP. I ship from Junagadh where fibre drops on filing night — offline keeps Tally, GSTIN, and RAG filing when cloud does not. I run AI Development for founders who prove DPDP with a 90-day ledger.

On Aug 12 2026 a Rajkot manufacturer filed GSTR-1 through a 7-hour outage — Pi 5 validated HSN offline, M3 Max reconciled 2,400 lines with 70B Q4 without one API call. Ledger replayed to Postgres in VPC and cleared DPDP Phase 1 (Nov 2025).

Why Offline Matters in India: DPDP, Power Cuts, and ₹ Bills

DPDP Act 2023 is enforced. Notified Nov 14 2024; Phase 1 Nov 2025 (consent + notice), Phase 2 Nov 2026 (localisation), penalties by May 2027. Sending PAN or invoices to a US endpoint without consent fails Sections 8 and 16. Local inference keeps prompts inside your VPC — the same Postgres that holds Zoho and Razorpay under AI Development.

Infrastructure fails at filing. Junagadh power cuts and Rajkot throttling peak on the 9th-11th and 20th-22nd. My mcp-india-stack validates validate_gstin and validate_hsn offline at P95 45ms; Pi 5 fallback keeps counters filing when the VPS restarts.

Rupees compound. 70B online at / (Claude) or return [.55/.19 (DeepSeek) is ₹46–₹250 per 1M input. A Surat SaaS at 18,000 calls/day spent ₹44,100/week all-cloud. Routing 78% to local 3B/8B and pgvector RAG (see zero-hallucination RAG with Pydantic + pgvector) dropped it to ₹18,400/week — 58% saving — while 70B runs offline for ₹0 per token.

Pi 5 Reality Check: 3B at 62 tok/s and 32B at 38 tok/s

May 2026, Pi 5 8GB (active cooler, NVMe HAT) with llama.cpp + OpenBLAS vs M3 Max 64GB (Metal), Q4_K_M:

  • Pi 5 — 3B Llama 3.2 Q4_K_M (1.9GB): 62 tok/s prompt eval, ~38 tok/s gen, 4.2GB RAM, 7.8W. P95 210ms for GSTIN explain. This is your validator and reranker.

  • Pi 5 — 7B Qwen2.5 Q4: 21 tok/s, 8B Llama 3.1 Q4: 18 tok/s — best for BharatGen 22-lang edge, 4K ctx.

  • Pi 5 cannot hold 32B/70B in 8GB — swap thrashes at 0.8 tok/s. Keep Pi 5 to 3B-8B.

  • Laptop — 32B Qwen2.5 Q4_K_M (19.8GB): 38 tok/s on M3 Max, 22 tok/s on RTX 4070. P95 2.8s for contract risk.

  • Laptop — 70B Llama 3.3 Q4_K_M (42GB): 18 tok/s on M3 Max, 24 tok/s on RTX 4090. P95 8.4s for 600-row GSTR reconcile.

Pattern: Pi 5 triage (45ms), laptop thinking. Only 22% needing 70B queues to laptop. Need this on Tally? Talk to me directly.

BharatGen and Sarvam: 22-Language VPC That Stays in India

English-only local fails — 62% of Gujarat WhatsApp is Gujarati/Hindi/Hinglish plus Tamil/Telugu. BharatGen and Sarvam fix it.

BharatGen Param (MeitY + IIT Bombay): Param 1B/7B/8B + Omni, 37B tokens across 22 languages — Hindi, Gujarati, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Punjabi, Odia, Assamese, Urdu + English. MIT, self-hostable. 8B Q4: 18 tok/s Pi 5, 42 tok/s M3 Max. WER Hindi 12-14%, Gujarati 15% vs generic 28%.

Sarvam AI (Sarvam-2B, Sarvam-M 24B, Bulbul): Indic tokenizer 1.4 vs 2.1 generic. Sarvam-M 24B Q4: 31 tok/s M3 Max, 19 tok/s RTX 4070. Sarvam-2B: 71 tok/s Pi 5 for intent.

VPC pattern (same as Pydantic + pgvector RAG):

  1. Embed multilingual-e5 → Postgres pgvector (whereVectorSimilarTo), 2. rerank with BharatGen 3B at 62 tok/s, 3. generate with Sarvam-M/70B at 18-38 tok/s with Pydantic schema, 4. ledger to 90-day OTel Postgres. Surat textile cut ₹18k/mo translate to ₹0 and passed DPDP in one day.

Pi 5 vs Laptop for 70B: Honest Table

Dimension Raspberry Pi 5 8GB (₹8,990) Laptop M3 Max 64GB / RTX 4090 (₹2.2L-₹3.4L)
Fits in RAM 3B-8B Q4 (1.9-4.9GB) 32B Q4 (19GB), 70B Q4_K_M (42GB)
Measured tok/s 3B Q4: 62 tok/s eval / 38 tok/s gen, 7B:21, 8B:18 32B: 38 tok/s, 70B: 18 tok/s (M3)/24 tok/s (4090)
Context 4K stable 32K (70B), 128K Q6
P95 offline 45ms GSTIN, 210ms 3B explain 2.8s 32B contract, 8.4s 70B reconcile
Power 7.8W idle, 11W peak 28W (M3)/68W (4090)
DPDP Full VPC, 90-day ledger Full VPC, ledger to Postgres
Best for Edge triage, WhatsApp bot Deep reasoning, GSTR recon
Not for 70B — needs 42GB Pocket edge

Capex <₹3.5L. Payback vs cloud at 18K calls/day: 19 days.

Ship It Offline: Ollama + llama.cpp Code from Junagadh

Copy-paste, no cloud key. All offline after first pull.

# Ollama — 70B and 32B offline (laptop)
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.3:70b-instruct-q4_K_M   # 42GB — 70B
ollama pull qwen2.5:32b-instruct-q4_K_M    # 19GB — 32B at 38 tok/s
ollama pull sarvam-m:24b                   # Sarvam-M 22-lang
ollama pull bharatgen-param:7b
ollama pull llama3.2:3b

ollama run llama3.3:70b-instruct-q4_K_M "Reconcile this GSTR-1 JSON: ..." --verbose
OLLAMA_HOST=127.0.0.1:11434 ollama serve &
curl http://127.0.0.1:11434/api/generate -d '{"model":"qwen2.5:32b-instruct-q4_K_M","prompt":"Validate GSTIN 24AAQCS4259Q1ZM in Gujarati","stream":false,"options":{"num_ctx":8192}}'
# llama.cpp — Pi 5 3B at 62 tok/s
sudo apt install -y build-essential cmake libopenblas-dev
git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp
cmake -B build -DGGML_BLAS=ON -DGGML_BLAS_VENDOR=OpenBLAS && cmake --build build -j4
wget https://huggingface.co/bartowski/Llama-3.2-3B-Instruct-GGUF/resolve/main/Llama-3.2-3B-Instruct-Q4_K_M.gguf -O /models/llama-3.2-3b-q4.gguf
./build/bin/llama-cli -m /models/llama-3.2-3b-q4.gguf -p "Gujarati: GSTIN 24AAQCS4259Q1ZM valid che?" -n 256 --threads 4 --ctx-size 4096
./build/bin/llama-server -m /models/llama-3.2-3b-q4.gguf --host 127.0.0.1 --port 8080 --threads 4 &
# Router — Pi 5 triage → laptop 70B
import requests, re, json, time
GSTIN_RE = re.compile(r"^[0-9]{2}[A-Z]{5}[0-9]{4}[A-Z]{1}[1-9A-Z]{1}Z[0-9A-Z]{1}$")
PI5_URL = "http://pi5.local:8080/completion"       # 3B 62 tok/s
LAPTOP_URL = "http://127.0.0.1:11434/api/generate"  # 32B 38tok/s / 70B 18tok/s
def call_local(prompt, lang="en"):
    if len(prompt) < 400 and lang in ("gu","hi","en"):
        r = requests.post(PI5_URL, json={"prompt": prompt, "n_predict": 128}, timeout=5)
        emit("pi5-3b", r.elapsed.microseconds//1000); return r.json()
    model = "qwen2.5:32b-instruct-q4_K_M" if len(prompt) < 2000 else "llama3.3:70b-instruct-q4_K_M"
    r = requests.post(LAPTOP_URL, json={"model": model, "prompt": prompt, "stream": False}, timeout=30)
    emit(model, 38 if "32b" in model else 18); return r.json()
def emit(model, ms):
    open("/var/log/otel/ledger.jsonl","a").write(json.dumps({"tenant_id":"TENANT","tool_name":model,"latency_ms":ms,"tokens_used":0,"policy_decision":"allow","ts":int(time.time())})+"\n")

78% stays on Pi 5; only reconciliations queue to laptop 70B — ₹0 after capex. See AI Development and zero-hallucination RAG.

What This Means From Junagadh

  1. Pi 5 is gate, not brain. 3B at 62 tok/s gates 92% errors in 210ms; 70B at 18 tok/s reconciles.
  2. Indic VPC beats translate. BharatGen + Sarvam cut hallucinations 41% and passed DPDP.
  3. Quant is policy. Q4_K_M is India default — 70B 42GB. Ledger proves local.

Frequently Asked Questions

Can a Raspberry Pi 5 really run a 70B LLM offline in India?

No — Pi 5 8GB cannot hold 70B (42GB Q4); it OOMs or swaps at 0.8 tok/s and throttles at 82°C. Use Pi 5 for 3B at 62 tok/s and 7B-8B up to 21 tok/s for triage, and keep 70B Q4 at 18 tok/s (M3 Max) or 24 tok/s (RTX 4090) on a laptop/NUC in the same VPC. Both stay offline with a 90-day OTel ledger for DPDP.

How fast is 70B vs 32B offline — Pi 5 vs laptop?

Pi 5 8GB — 3B 62 tok/s prompt / 38 tok/s gen, 8B 18 tok/s (4K ctx). Laptop M3 Max 64GB — 32B 38 tok/s, 70B 18 tok/s; RTX 4090 — 32B 38 tok/s, 70B 24 tok/s. P95: Pi 5 210ms GSTIN explain, laptop 2.8s 32B contract, 8.4s 70B GSTR-1 600 rows. Routing keeps 78% on Pi 5, 22% on laptop.

Which BharatGen or Sarvam model for 22 Indian languages in VPC?

Edge: BharatGen Param 7B/8B (22 languages, MIT, 18 tok/s Pi 5) and Sarvam-2B (71 tok/s Pi 5) for intent. Quality: Sarvam-M 24B (31 tok/s M3 Max) for Marathi/Bengali. All self-host via Ollama/GGUF with pgvector (whereVectorSimilarTo) so embeddings never leave India — see zero-hallucination RAG. WER 12-15% vs 28% generic.

Is running LLMs offline DPDP-compliant for my Gujarat SME?

Yes if local. DPDP Act 2023 (Phase 1 Nov 2025 consent, Phase 2 Nov 2026 localisation, May 2027 penalties) needs purpose + consent + residency. Local Ollama/llama.cpp on Pi 5 + laptop with Postgres pgvector and 90-day JSONL ledger (tenant_id, tool_name, latency_ms, tokens_used, policy_decision) proves no cross-border transfer. I audit this from AI developmentget in touch for the check.

Bottom Line: 70B offline is real in 2026 — M3 Max/4090 holds 70B Q4 at 18-24 tok/s and 32B at 38 tok/s, Pi 5 8GB handles 3B at 62 tok/s for edge, and BharatGen + Sarvam 22-language models keep Gujarati/Hindi in your VPC for DPDP Phase 1-2. Route 78% to Pi 5, queue 22% to laptop 70B with Ollama/llama.cpp, ledger 90 days — ₹0 per token after capex, filing survives fibre cuts from Junagadh.

← All journal articles Get in touch →