Cost per useful output
Not cost per token. A cheap model that needs three attempts and a human correction is more expensive than the one that answers once. The unit that matters is the finished, accepted result.
Stack / model-agnostic by design
Model-agnostic by design. The right model is the one that holds up in production at the price the workload can bear — which is rarely the one at the top of a benchmark table.
Our default stack. Direct API, not a wrapper and not a reseller.
Claude runs across our own platforms today: astrology interpretation in VedicPupil, content generation in Keywrd Scribe, and transaction classification in Arthastra’s reconciliation path — where it sits as the paid fallback tier behind a free model and a deterministic rules engine, and costs a fraction of a cent per call precisely because it is asked so rarely.
Chosen per workload, not per preference. Every one of these is in production somewhere in the portfolio.
GPT models where the ecosystem, tooling or multimodal fit is stronger. Some workloads are simply better served by what has already been built around a provider, and pretending otherwise costs the client months.
Indian-language models. Hindi and regional inference built for Indian users, with data staying in India. In Arthastra it is the tier-one model for exactly this reason: a translation round trip to an English-first model loses the accounting nouns that carry the meaning.
Inference infrastructure for sub-second response, used where latency is the product rather than a metric on a dashboard. A model that answers correctly in nine seconds has failed an interactive flow no matter how good the answer was.
Four questions, asked in this order, on every workload.
Not cost per token. A cheap model that needs three attempts and a human correction is more expensive than the one that answers once. The unit that matters is the finished, accepted result.
Measured against what the user is doing, not against a benchmark. A batch reconciliation can wait eight seconds. A confirm modal cannot wait one.
What happens when the model is wrong. If a wrong answer posts money, the model does not get to post — it proposes, and something deterministic decides. This question eliminates more candidate designs than the other three combined.
Where the inference physically happens and what the provider retains. For Indian personal data this is now a compliance question with a deadline attached, not a preference — see DPDP compliance engineering.
We don’t sell you a model. We build the system around it.
Model choice is 10% of the work
Swapping a model is a configuration change. Everything that makes the model dependable is engineering, and it is where the entire budget of a serious AI build actually goes.
Capacity: 2–3 engagements per year
Start a conversation →Tell us the workload, the latency budget and what a wrong answer costs. The model follows from those three answers.