Case file
W2M AI
A Claude-powered assistant that answers staff questions straight from the production database, with guardrails.
In one paragraphFor engineersFor hiring managersThe short version
Operations staff ask questions like "how many orders shipped last week?" in English, Burmese, or Thai and get an answer from live data in seconds, without a developer writing a report. Every answer is logged, budgeted, and limited to what that person is allowed to see.An agentic tool-use loop over the live MongoDB with read-only query, aggregation, and schema tools, plus an ask-user tool for clarification. Role-scoped collection allowlists, sensitive-field stripping, hard caps on documents, pipeline stages, timeouts, and iterations, per-user daily token budgets, rate limiting, and a full audit log. Streams over SSE into a Tauri desktop app.Operations staff ask questions like "how many orders shipped last week?" in English, Burmese, or Thai and get an answer from live data in seconds, without a developer writing a report. Every answer is logged, budgeted, and limited to what that person is allowed to see.I taught an AI to answer questions about a shipping company's data, and spent most of the effort making sure it could not do anything else.
Choose a lens in the header to reorder this page for your reading.Engineer lens: the engineering story comes before the outcome.Hiring lens: the outcome comes right after the problem; the engineering story follows.Curious lens: the story first, the engineering detail last.
The problem
Every operations team has questions that are not on any report. How many orders shipped last week? Which customers still have deposits outstanding? What did we move for this franchisee in March? In a small company, each of those becomes a message to the developer, who writes a query, and the answer arrives when the developer is free.
The idea was to let staff ask the database directly, in their own language, and get a real answer from live data. The risk is obvious: an AI with access to the production database can leak what it should not, run something expensive, or confidently answer wrong. The project was mostly about that risk.
What I built
A desktop assistant, delivered as a Tauri 2 app, backed by an Express service that runs an agentic loop on Anthropic's Claude models.
- Tools. Query, aggregation, and schema tools over MongoDB, so the model can discover structure and then read data. An ask-user tool lets it stop and ask a clarifying question instead of guessing.
- Models. Claude Sonnet 4.5 and Haiku 4.5, used where each fits.
- Guardrails. Role-scoped collection allowlists, sensitive-field stripping, and hard caps on documents returned, pipeline stages, timeouts, and loop iterations.
- Budgets. Per-user daily token budgets and rate limiting, with a full AI audit log of every request, tool call, and answer.
- Delivery. Answers stream over SSE into a React UI that renders tables, in whichever of English, Burmese, or Thai the question was asked.
The engineering story
The loop. A request becomes a conversation between the model and a small set of tools. The model can inspect the schema, run a query or an aggregation, or ask the user a question, and it keeps going until it has an answer or hits a limit. Every step is logged.
What the model is allowed to see. Each user's role maps to an allowlist of collections, and sensitive fields are stripped from results before the model sees them. The model never gets a raw connection; it gets the tools, and the tools enforce the rules.
What the model is allowed to spend. Hard caps on documents per result, aggregation pipeline stages, query timeouts, and loop iterations keep one question from turning into a load problem. Per-user daily token budgets and rate limiting keep the bill predictable, and the audit log makes any surprise explainable.
Asking instead of guessing. The ask-user tool exists because ambiguous questions are the normal case. "Last week" can mean two things; the assistant is allowed to ask which.
Streaming. Answers stream over SSE so staff see progress and partial tables instead of a spinner, which matters when a question takes several tool calls.
What it changed
Staff get answers from live data without waiting for a developer, in the language they think in. Every answer is traceable to the tool calls that produced it, limited to what that person is allowed to see, and capped in cost.
For me it changed what "supporting the business" means: fewer one-off queries, and a log of what people actually want to know, which is a better product roadmap than any meeting.