The receipts behind the numbers we quote. Short, technical write-ups of what we're shipping, including the parts that took us longer than they should have.
Latency budgets, barge-in, code-switching across Hindi, Tamil, and English, and why sound design is trust engineering. Written from the reference build that runs live on this site.
The autonomy ladder we use in every engagement — suggest, act with approval, act with audit — and why promotion between levels is earned by eval scores, never by demos.
An honest break-even analysis. Utilization decides everything, and sometimes the right answer is to keep paying the API. Here is the arithmetic we run at scoping.
How we wire retrieval, evals, and human review so agents stay tethered to source data, and how we know when they're drifting.
A walkthrough of the routing, batching, and quantization decisions we use to run Qwen3.5 fleets on client GPUs without blowing budget.
How a small pipeline of classifiers, clustering, and LLM summarization turned a year-end HR survey into a 6-page brief leadership actually read.
Every note above comes from a build we ran ourselves. If one reads like a problem on your desk, bring it to a call.
Free, 45 minutes. A founder replies within two days.