Over the last few months I built dalimb.com — a Marathi-first service for pomegranate farmers in Maharashtra. Daily mandi prices, local weather with agronomic alerts, a disease and pest guide, and short how-to articles with Marathi audio. I don’t monetise it and I don’t plan to; it’s social work, not a product.

But it is a real production site with real users, and the way it’s built is worth writing up. Claude Code did the building. Cloudflare runs it. There is not a single LLM call in production.

That last part is the interesting bit, so let’s start there.

Build-time AI, not runtime AI

Most “AI-powered site” write-ups mean the same thing: the application calls an LLM on every request. That’s a legitimate architecture, but it is also the slowest, most expensive and least predictable part of any stack you bolt it onto.

This project inverts it. AI is used heavily — but entirely at build time:

  • Claude Code wrote the Cloudflare Worker, the static frontend, and the Python build scripts.
  • Claude drafted the Marathi content, which was then reviewed by a human before publishing.
  • Claude diagnosed production incidents and designed the fixes.
  • Text-to-speech runs in the build pipeline, not per request. The MP3s are static files.

The deployed site calls no model at all. A farmer opening the prices page on a ₹7,000 Android phone over a 3G connection gets static HTML and a JSON fetch from the edge — no inference, no cold start, no token bill.

The tradeoff is worth stating plainly: build-time AI can’t personalise per user, and it can’t answer a question you didn’t anticipate. For a content and data site, you don’t need either. What you get instead is latency measured in milliseconds, a cost of roughly zero, and — most importantly — every AI output is reviewable before it ships. For agricultural advice, that last property isn’t a nice-to-have.

The architecture

  Farmer's phone (Marathi, PWA)
            │
            ▼
  ┌───────────────────────────┐
  │ Cloudflare Pages          │  HTML (SEO)
  │ hand-written HTML + PWA   │  + app JSON
  └─────────────┬─────────────┘  + audio
                │ fetch() JSON (CORS)
                ▼
  ┌───────────────────────────┐
  │ Cloudflare Worker         │  REST/JSON API
  │ "dalimb-api"              │  + daily cron
  └──┬──────────┬─────────┬───┘
     │          │         │
     ▼          ▼         ▼
   D1         KV       external APIs
  (SQLite)   (cache)   MSAMB · Agmarknet
  markets    prices    Open-Meteo · Sarvam
  prices     weather   Resend · Turnstile
  subscribers

Two planes, deliberately separated: public read-only data (prices, weather, articles) and private transactional data (signups, newsletter). They share a Worker but nothing else.

The Cloudflare half

Pages: static HTML, no framework

The frontend is hand-written HTML with vanilla JS and CSS tokens for light/dark. No React, no build step for the pages themselves.

This wasn’t nostalgia. The target device is a cheap Android phone on patchy rural data. Shipping a framework runtime to parse and hydrate before anyone sees a price is a tax paid by exactly the people least able to afford it. Plain HTML plus a service worker gives you an installable PWA with offline fallback and last-known prices cached, at a fraction of the payload.

Workers: one Worker, two entry points

Everything server-side is a single Worker. The config is small:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
{
  "name": "dalimb-api",
  "main": "src/index.ts",
  "compatibility_date": "2026-09-01",
  "d1_databases": [
    { "binding": "DB", "database_name": "dalimb", "database_id": "<id>" }
  ],
  "kv_namespaces": [
    { "binding": "CACHE", "id": "<id>" }
  ],
  "triggers": {
    "crons": ["0 1,7,13 * * *"]
  }
}

A Worker can export both an HTTP handler and a scheduled handler, which means your API and your data refresh job are the same deployment, sharing the same code:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
export interface Env {
  DB: D1Database;
  CACHE: KVNamespace;
  TURNSTILE_SECRET: string;
}

export default {
  async fetch(req: Request, env: Env): Promise<Response> {
    const url = new URL(req.url);
    switch (url.pathname) {
      case "/api/prices": return getPrices(env);
      case "/api/health": return getHealth(env);
      default:            return new Response("Not found", { status: 404 });
    }
  },

  // Cron hits exactly the same refresh path as the manual admin endpoint,
  // so there is only ever one code path to debug.
  async scheduled(_event: ScheduledController, env: Env, ctx: ExecutionContext) {
    ctx.waitUntil(refreshPrices(env));
  },
} satisfies ExportedHandler<Env>;

D1: history you have to build yourself

Mandi prices are the daily-habit hook, and neither upstream source offers historical data. So the Worker persists one snapshot per market per day and trends accrue forward from launch:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
CREATE UNIQUE INDEX IF NOT EXISTS idx_daily_prices_market_date
  ON daily_prices (market_id, price_date);

INSERT INTO daily_prices
  (market_id, price_date, min_price, max_price, modal_price, arrivals)
VALUES (?1, ?2, ?3, ?4, ?5, ?6)
ON CONFLICT (market_id, price_date) DO UPDATE SET
  min_price   = excluded.min_price,
  max_price   = excluded.max_price,
  modal_price = excluded.modal_price,
  arrivals    = excluded.arrivals;

The unique index plus ON CONFLICT ... DO UPDATE makes the refresh idempotent — running it three times a day overwrites today’s row instead of duplicating it. Writes go through env.DB.batch() so the whole day lands in one round trip.

KV: cache, with a version in the key

KV holds the hot payloads — latest prices, weather per taluka. The one trick worth copying is putting a version prefix in the cache key:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
const CACHE_VERSION = "v2";

async function cached<T>(
  env: Env, key: string, ttlSeconds: number, build: () => Promise<T>,
): Promise<T> {
  const k = `${CACHE_VERSION}:${key}`;
  const hit = await env.CACHE.get<T>(k, "json");
  if (hit) return hit;

  const fresh = await build();
  await env.CACHE.put(k, JSON.stringify(fresh), { expirationTtl: ttlSeconds });
  return fresh;
}

KV has no “delete everything matching a prefix” operation. When the shape of a cached payload changes, bumping v2 to v3 invalidates every stale entry on the next deploy, instantly and for free. I learned this the way everyone does — by shipping a payload change and watching old objects keep serving.

Turnstile: bot protection without Google

The newsletter signup is protected by Cloudflare Turnstile rather than reCAPTCHA — same platform, free, no third-party tracking on a page aimed at farmers. The frontend renders the widget and sends the token; the Worker verifies it before writing anything:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
async function verifyTurnstile(token: string, ip: string, secret: string) {
  const body = new FormData();
  body.append("secret", secret);
  body.append("response", token);
  body.append("remoteip", ip);

  const res = await fetch(
    "https://challenges.cloudflare.com/turnstile/v0/siteverify",
    { method: "POST", body },
  );
  const data = (await res.json()) as { success: boolean };
  return data.success === true;
}

One design note to think about before copying: if the secret is unset, this implementation fails open so that a misconfigured deploy can’t silently break signups. That’s a deliberate availability-over-security call for a free newsletter with nothing sensitive behind it. If you adopt the pattern, make sure something alerts you when it’s in that state — otherwise “temporarily open” quietly becomes permanent.

Edge geolocation: local weather, no permission prompt

Asking for GPS permission on first visit is a great way to lose a user. Cloudflare already knows roughly where the request came from:

1
2
3
4
5
6
// request.cf is populated at the edge. latitude/longitude arrive as
// strings, not numbers — coerce them.
const cf = request.cf;
const lat  = cf?.latitude  ? Number(cf.latitude)  : null;
const lon  = cf?.longitude ? Number(cf.longitude) : null;
const city = cf?.city ?? null;

The home page uses this to show local weather immediately, with nothing stored. Precise location stays opt-in behind a taluka picker and browser GPS on the dedicated weather page. IP geolocation is coarse — fine for “which district’s weather”, useless for anything needing accuracy.

The Claude half

Pair-programming the thing

The Worker, the frontend and the Python build scripts were written and iterated with Claude Code. That part is unremarkable now and I won’t dwell on it.

Generating content at scale without generating slop

This is the transferable technique. The site needed a library of Marathi how-to articles. Naively prompting for forty articles produces forty pieces of mush.

What worked:

  1. Hand-write one article to the quality bar. One, properly, reviewed by someone who knows pomegranate agronomy.
  2. Turn that article into an explicit spec — not vibes. The real constraints were: Marathi throughout, Devanagari numerals, no pesticide dosages under any circumstances, and comma-paced sentences because the text would be read aloud by a TTS engine.
  3. Fan out parallel Claude subagents against the gold-standard article plus the spec, one per topic.
  4. Human review for agronomic accuracy before anything publishes.

Step 2 is where the quality comes from. The “no dosages” rule is a good example of what belongs in a spec rather than left to the model’s judgement: wrong dosage advice can damage a crop or injure the person spraying it. That’s a capability boundary, and boundaries should be written down, not hoped for.

Step 4 is non-negotiable and it’s the step people skip. A fluent, confident, wrong article about blight treatment is worse than no article.

Debugging and ops design

The most useful thing Claude did wasn’t writing code — it was working through production incidents. The government price API started returning HTTP 521 and blocking cloud IPs; separately, a stale KV cache kept serving an old payload shape after a deploy. Both got diagnosed and then designed around — source failover for the first, cache-key versioning for the second. The fixes below came out of those sessions.

The pattern worth stealing: one source, three outputs

Each article is authored once, as JSON. A build script renders it three ways:

  1. An HTML page with the full text inline, plus JSON-LD (Article, FAQPage, BreadcrumbList, AudioObject) and OpenGraph — the SEO and AI-citation surface.
  2. App JSON where the body is pre-parsed into typed blocks (heading, paragraph-with-spans, list), so a future native app renders it without shipping a markdown parser.
  3. Audio — a cleaned spoken script sent to a TTS API, chunked to the per-request limit and concatenated.

The audio step is idempotent: a manifest stores a hash of each spoken script, so only changed articles get re-synthesized. Without that you re-pay for the whole library every time you fix a typo.

The payoff is that the HTML never has to change to serve the app. SEO is locked in on day one, and the app is a consumer of the same build output rather than a second content system to keep in sync.

1
2
3
4
5
# adding an article is a four-step pipeline
python3 scripts/build_articles.py
python3 scripts/build_audio.py --only <slug>
npm run deploy:frontend
python3 scripts/seo.py sitemap && python3 scripts/seo.py indexnow

Designing for a data source that breaks

The hardest part of this project wasn’t the stack. It was that the authoritative price data lives on government websites that go down, change shape and block cloud IPs.

What ended up working:

  • A primary and a fallback source, with the fallback demoted after it started returning 521s to datacenter ranges.
  • Retries on transient 5xx, because a lot of these failures are momentary.
  • Cron three times a day, not once. A single daily job means one bad minute costs you the whole day’s data.
  • A health endpoint that reports staleness, not just liveness:
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
async function getHealth(env: Env): Promise<Response> {
  const row = await env.DB
    .prepare("SELECT MAX(price_date) AS latest_date FROM daily_prices")
    .first<{ latest_date: string | null }>();

  const latest = row?.latest_date ?? null;
  const daysStale = latest
    ? Math.floor((Date.now() - Date.parse(latest)) / 86_400_000)
    : null;

  return Response.json({
    ok: true,
    latest_date: latest,
    days_stale: daysStale,
    stale: daysStale === null || daysStale > 1,
  });
}
  • Honest UI when data is old. If today’s prices didn’t arrive, the page says so and shows the last known date. Silently displaying stale prices as current is the one failure mode that actually costs someone money.

/api/health returning a stale boolean means uptime monitoring catches “the site is up but the data is three days old” — which is the failure that matters here, and the one a plain 200-check misses entirely.

What it costs

Effectively nothing, which is the point for an unfunded project. Cloudflare’s free tiers at the time of writing: Pages serves unlimited static requests with 500 builds a month; Workers allows 100,000 requests a day; KV gives 100,000 reads but only 1,000 writes a day; D1 allows 5 million rows read and 100,000 written daily. Check the current limits before planning around these — they move.

The ceiling that bites first is KV writes. 100,000 reads a day is generous; 1,000 writes is not. If your instinct is to write to KV on every request — a counter, a session touch, a rate-limit bucket — you’ll blow through it long before you approach any other limit. Cache reads from KV; put writes in D1 or batch them in the cron.

The only paid-ish pieces are the TTS calls and transactional email, both small and both build-time or low-volume.

When this stack fits — and when it doesn’t

It fits when you’re building a read-heavy content or data site, your working set is small, you want global latency without managing regions, and the team is one person with no budget and no appetite for on-call.

It doesn’t fit when you need heavy relational queries or large joins (D1 is SQLite at the edge, and it will tell you so), when you need long-running compute (Workers are CPU-time limited per request), or when you genuinely need per-request personalisation that only a model can produce. If you need an LLM in the request path, you need a different cost model and a different latency budget — be honest about that up front rather than discovering it after launch.

For everything else, the combination is hard to beat right now: an agentic coding tool that collapses the build time, and an edge platform whose free tier is a genuinely production-grade target rather than a trial.

You can see the result at dalimb.com. It’s in Marathi — but the prices page, the weather alerts and the article player are all readable enough from the structure if you want to poke at what the stack actually produces.