charity-site-ai-assistant-article
I Run a Charity Website With Its Own AI Assistant. It Costs Me $11 a Year.
How a static site, a folder of Markdown files, and ~250 lines of hand-written search code gave a small foundation in Odisha a question-answering bot — with no vector database, no API keys, and no monthly bill.
The budget was the design document
Service to Humanity Foundation is a small charity. Small charities do not have infrastructure budgets. They have volunteers, a bank account, and a lot of questions coming in from people who want to know what the organisation actually does, where the money goes, and how to help.
When I started building their website, I gave myself one constraint and let everything else follow from it:
The only recurring cost allowed is the domain name. $11 per year. Nothing else.
No hosting bill. No OpenAI key. No Pinecone instance. No Vercel Pro. If a design decision needed a credit card, it was the wrong design decision.
That constraint turned out to be the most useful thing about the project. It killed every lazy answer before I could reach for it.
The site is live at sthfjagda.com. Here is where the constraint landed:
25,000 visits · 1,960 unique visitors · $11 a year
No hosting bill. No inference bill. No database bill. One domain renewal.
Those three numbers are the whole article. The rest of it is how.
A word about whose website this is. Service to Humanity Foundation was established on January 30, 1981 by Dr. Sashikanta Acharya (PhD, TU Berlin). It runs six children's homes in Odisha, India, currently caring for more than fifty orphaned and deprived children — food, shelter, school enrolment, books, uniforms, healthcare, and a continuing hand into higher education and first jobs. It is registered under the Juvenile Justice Act of 2000 and holds 80G tax-exempt status. The office and home are in Jagda, PO–Jhirpani, Rourkela.
Forty-five years of work, and until recently, no good way for a stranger on the internet to ask it a question.
The part everyone expects, and why I didn't build it
The default 2026 answer to "I want a chatbot that knows about my organisation" is well known:
- Chunk your documents.
- Embed them with an embeddings API.
- Store the vectors in a managed vector database.
- On each question, embed the query, do a similarity search, stuff the results into GPT-class model, stream the answer.
It works. It also assumes you have an API key with a billing relationship behind it, a database that bills by the hour, and a server that stays awake. For a corpus of 35 sections across 7 documents, that entire stack is a rounding error of value on top of a permanent monthly cost.
There is a second tempting answer: run the model in the browser. WebLLM, WebGPU, a quantised 3B model. I ruled that out too. Our visitors are overwhelmingly on mid-range Android phones in India. Asking them to download gigabytes of weights so they can find out how to sponsor a child is not a feature, it's an insult.
So I inverted the problem. The browser does search, never inference. The language model — when there is one — only rewrites what search already found.
Architecture
Two pipelines: one runs at build time, one runs per question.
Build time: Markdown in, JSON index out
┌──────────────┐ node scripts/build-knowledge-index.mjs ┌──────────────────────────────┐
│ knowledge- │ ───────────────────────────────────────▶ │ public/knowledge-base/ │
│ base/*.md │ runs as the pre-step of `pnpm dev` │ index.json │
│ (human-edit) │ and `pnpm build` │ {version, generatedAt, │
└──────────────┘ │ chunks[35]} ≈ 18 KB │
└──────────────┬───────────────┘
│ next build
▼
┌──────────────────────────────┐
│ out/ (fully static export) │
│ out/knowledge-base/index.json│
└──────────────────────────────┘
Runtime: one question, three rungs, one guarantee
Visitor types a question in <AiChatWidget/>
│
▼
[1] Browser-side BM25 retrieval lib/knowledge-base/client.ts
fetch /knowledge-base/index.json (cached, single-flight per session)
tokenise → synonym-expand → BM25 rank → top 5 sources
│
▼
[2] Phrasing attempt A — Cloudflare functions/api/chat.ts
POST /api/chat {question, sources[4]} (Pages Function — production)
→ Workers AI: env.AI.run(model, …)
← 200 {answer, model} mode = "ai" (✨)
│ any failure / local dev
▼
[3] Phrasing attempt B — local Ollama lib/ai/ollama.ts
GET 127.0.0.1:11434/api/tags (auto-discover model, 700 ms ping)
POST 127.0.0.1:11434/v1/chat/completions mode = "ai" (✨)
│ unavailable / any error
▼
[4] Extractive answer — ALWAYS succeeds lib/knowledge-base/answer.ts
heading re-rank → sentence selection mode = "knowledge" (♥)
│
▼
Bubble renders answer + top-2 source labels ("Sources: …")
Rungs 2 and 3 are pure progressive enhancement. Every failure falls silently to the rung below. Rung 4 has no dependencies and therefore no failure mode. The visitor never sees an error message, because there is no error to show.
Deployment topology
Cloudflare edge (per request)
├── Static assets (CDN cache) out/** — HTML/JS/CSS, index.json, images
├── /api/chat → Pages Function functions/api/chat.ts (env.AI, env.AI_MODEL)
└── everything else → single-page app shell, anchor navigation
The stack
| Layer | Technology | Notes |
|---|---|---|
| Framework | Next.js 16 (App Router) | output: 'export' — fully static |
| UI | React 19, Tailwind CSS 4, shadcn/ui, lucide-react | |
| Language | TypeScript 5.7, strict: true |
|
| Package manager | pnpm 9.15 | pinned via packageManager |
| Retrieval | Custom BM25 engine | lib/knowledge-base/client.ts — zero dependencies |
| AI phrasing (prod) | Cloudflare Workers AI | @cf/meta/llama-3.3-70b-instruct-fp8-fast, free tier |
| AI phrasing (dev) | Ollama on localhost | auto model discovery, OpenAI-compatible endpoint |
| Edge compute | Cloudflare Pages Functions | Node 20 |
| Hosting | Cloudflare Pages | Git integration, global CDN |
Runtime dependency budget for the entire AI subsystem: zero npm packages. The only runtime I/O is one static JSON fetch, one same-origin POST /api/chat, and — on my laptop only — localhost calls to Ollama.
The content layer: publishing is editing a Markdown file
The hardest requirement was not technical. It was this: a non-engineer volunteer has to be able to change what the bot says.
So the corpus is a folder of Markdown files with optional YAML front matter:
---
title: "Donations and Volunteering"
tags: [donate, donation, volunteer]
---
## How to donate ← every heading starts a new searchable chunk
Body text…
## Volunteering
Body text…
The rules are deliberately boring:
- Every
##–######heading starts a new chunk. One heading, one answerable unit. - Front matter is optional.
tagsboost retrieval. README.mdis documentation, not corpus — the builder skips it.- Fenced code blocks are ignored.
- Facts must be self-contained per section. A chunk that only makes sense if you read the one above it will retrieve badly.
A dependency-free Node script walks the folder, parses front matter, splits on headings, strips Markdown emphasis to display-ready plain text, and emits stable readable IDs like 01-about-foundation::where-we-work. The output:
{
"version": 1,
"generatedAt": "2026-09-12T20:41:07.284Z",
"chunks": [
{
"id": "01-about-foundation::where-we-work",
"file": "01-about-foundation.md",
"title": "About the Foundation",
"heading": "Where we work",
"text": "We are primarily active in the state of Odisha, India, …",
"tags": ["about", "history", "mission", "odisha", "founder"]
}
]
}
Eighteen kilobytes. Smaller than most hero images.
Two details I'd defend to anyone building something similar:
The builder is a pre-step of pnpm dev and pnpm build, not a separate command. There is no "did you rebuild the index?" failure mode, because there is no index command to forget.
A broken corpus exits with code 1 and aborts the deploy. Content errors fail the build, not production.
Retrieval: BM25, by hand, in the browser
The search engine is about 250 lines of TypeScript with no dependencies: a tokeniser, a light stemmer, a curated synonym map, and BM25 scoring over heading, title, tags and body.
Why BM25 and not embeddings? At 35 chunks, lexical search with a hand-tuned synonym table beats a vector pipeline on every axis that matters here — cost, privacy, latency, and the number of things that can break at 2am. Embeddings earn their complexity when the corpus is large enough that vocabulary mismatch dominates. Ours isn't. The synonym map does that job for a tenth of the effort, and the front-matter tags reinforce it — the about document carries [about, history, mission, odisha, founder, established, address, location, located, rourkela, visit], which is a volunteer's honest list of the ways a real person might phrase the question, written by the person who answers those questions on the phone.
The whole index is fetched once per session, cached, and single-flighted so a fast typist doesn't trigger five parallel fetches. Ranking a question against 35 chunks takes microseconds. No question a visitor asks ever leaves their browser unless a language model is actually going to rephrase the answer.
The AI layer, and the thing it is not allowed to do
Here is the part I care most about.
Look at what this corpus actually contains: a founding date, a founder's name and doctorate, a registration act, an 80G tax-exemption status, a Form 10AB approval order, a street address in Rourkela, a count of six homes and fifty children, and instructions on how to sponsor a child.
Every one of those is a fact a language model would be delighted to approximate. A hallucinated answer here is not an amusing failure — it is a wrong address given to someone who drove to Rourkela, an invented legal claim about tax exemption, or a confidently fabricated sponsorship process. A forty-five-year-old organisation's credibility is on the line, and it isn't mine to spend.
There is one more decision worth naming, and it isn't code. The site's Donate section publishes bank transfer details — account name, branch, account number, IFSC. None of that is in the knowledge base. The corpus says only that the Donate section shows the current options, so the bot answers "how do I donate?" by pointing at the page rather than by reciting an account number from memory.
That's belt-and-braces on top of an already-constrained model, and I'd do it again. Retrieval is good, not perfect; a mis-ranked chunk that puts the wrong four excerpts in front of the model is a normal Tuesday. The cost of that failure should never be a digit wrong in a bank account number. Some facts should simply not be in the system that phrases things.
So the model is architecturally prevented from knowing anything it shouldn't:
It is fed only retrieved excerpts. The top four BM25 results and nothing else. No web access. No repository access. No memory between questions.
The system prompt forbids invention, in as few words as possible:
You are "Humanity Assistant", the friendly AI helper for the Service to Humanity Foundation website. Answer ONLY using the knowledge-base excerpts provided. Keep answers to 2-4 sentences. If the excerpts don't contain the answer, say so and suggest contacting the foundation via the website's Donate section. Never invent facts. Do not mention "knowledge base" or "excerpts" in your reply.Temperature 0.2, max 400 tokens. The instruction to not mention "knowledge base" matters more than it looks — a visitor asking about a child's schooling should not be told about the retrieval architecture.
If the model fails, produces nothing usable, or the endpoint errors, the extractive engine answers instead — by selecting real sentences from the real corpus. Rung 4 is not a fallback in the apologetic sense. It is the floor, and the floor is always made of published content.
The model's only job is phrasing. It turns retrieved sentences into a fluent paragraph. It is a stylist, not a source. That distinction is the whole security model, and it's why I can put this on a real charity's site without losing sleep.
The UI is honest about which rung answered: a ✨ icon when a model phrased it, a ♥ icon when it's extractive. Both cite their top two sources underneath. A visitor can always trace an answer back to a document a human wrote.
The request from browser to /api/chat is validated and clamped server-side with boring constants: question capped at 1,000 characters, at most 4 sources, each excerpt truncated to 600 characters, title 120, heading 200. Anything with no usable excerpt is rejected with a 422 before the model is ever invoked. A public endpoint pointed at a free-tier AI binding is exactly the kind of thing people probe.
There's a smaller piece of defensiveness I'm fond of: the Function tries the messages shape first and falls back to the prompt shape if the model doesn't accept it, because Workers AI models differ on this and I'd rather the code survive a model swap than have the site depend on me noticing.
And the honest answer to "what happens when you exhaust the free tier?" — the Function returns a 502, the browser falls to rung 4, and visitors get extractive answers with a ♥ instead of a ✨ until the quota resets. That is the system working as designed rather than an outage, which is the entire reason the ladder exists.
What it actually costs
| Item | Cost |
|---|---|
| Domain | $11 / year |
| Hosting (Cloudflare Pages, static + Functions) | $0 |
| AI inference (Workers AI free tier, Llama 3.3 70B open weights) | $0 |
| Vector database | $0 — there isn't one |
| Embeddings API | $0 — there isn't one |
| Local dev inference (Ollama) | $0 |
| Total recurring | $11 / year |
And what that $11 has carried:
| Metric | Figure |
|---|---|
| Total visits | ~25,000 |
| Unique visitors | ~1,960 |
| Visits per visitor | ~12.8 |
| Cost per visit | $0.00044 |
| Index shipped to each visitor | 16 KB |
| Runtime npm dependencies in the AI subsystem | 0 |
That return-visit ratio is the number I find most interesting. Twelve-odd visits per person is not drive-by traffic — it's people coming back, which is what you'd hope for from donors, families, and the alumni the foundation still supports. It also means the static index is being served from cache almost every time, which is exactly why the bill stays flat.
What I'd tell you to steal
- Let the constraint do the architecture. "No recurring cost" eliminated more bad designs in an afternoon than a week of deliberation would have.
- Separate retrieval from generation completely. Retrieval is cheap, deterministic, testable, and private. Generation is expensive, probabilistic, and occasionally unavailable. Building them as independent layers is what makes the free-tier LLM optional rather than load-bearing.
- Design the degradation ladder before the happy path. I wrote the extractive engine first. Everything above it is an upgrade, which means no rung can take the site down.
- Small corpora do not need vector search. Be honest about your scale. Mine is 35 chunks; I'll revisit at around 500.
- Make publishing boring. The measure of success is whether a volunteer who has never opened a terminal can correct a phone number. Anything that requires me to be in the loop is a design failure with a long tail.
What's still on the list
Conversational follow-ups (each question is currently independent), an answer-quality eval set so corpus edits can be regression-tested, retrieval telemetry on which questions return nothing useful, and the vector-search revisit if the corpus grows an order of magnitude.
The part that actually matters
Everything above is scaffolding. This is the building.
Service to Humanity Foundation has been raising orphaned and destitute children in Rourkela since 1981, on public donations, under a founder who could have built a fortune with a TU Berlin doctorate and chose this instead — and a superintendent the children simply call Ma. Six homes. Fifty-plus children today. Forty-five years.
₹2,150 covers one child's basic needs for a month. Donations are 80G tax-exempt, the Form 10AB and CSR-1 certificates are published on the site, and the foundation reports monthly to the district judge.
If the engineering held your attention, the work it points at deserves more of it: sthfjagda.com — the Donate section carries the bank details, and they need teachers, tutors, mentors and doctors as much as they need money.
The repo is public, including the entire corpus the bot is allowed to speak from — you can read every sentence it can possibly say: github.com/Biswajit107927/service-humanity-website · the knowledge base itself
I'm a Senior Data Engineer at AWS. This was a weekend project, which is rather the point — the interesting engineering constraint wasn't scale, it was that there was no money at all.