Inference cost is outrunning revenue
Per-token pricing is fine in beta. At scale it quietly turns a 80% gross margin into a 40% one.
Home / Industries / SaaS
SaaS
AI features for software products, built so that gross margin survives adoption. We size the serving stack for your tenth thousand user, not your demo.
If none of these sound like you, we are probably not the right call yet — and we will say so.
Per-token pricing is fine in beta. At scale it quietly turns a 80% gross margin into a 40% one.
Their security questionnaire asks where the data goes, and 'a third-party API' is the wrong answer.
The demo worked. Making it reliable for every tenant is a different engineering problem.
Turning a working prototype into a feature that every tenant can use, that support can debug and that finance can forecast is where most AI roadmaps stall. That is the part we do.
Predictable unit economics.
We model cost per active user before the build and pick the serving approach that holds at your projected scale.
Tenant isolation that survives a security review.
Per-tenant data boundaries, retention controls and the documentation your buyers ask for.
Self-hosted where it pays.
Open-weight models on dedicated GPUs often beat API pricing past a threshold. We tell you where your threshold is.
Built into your stack.
Your repo, your CI, your observability. We leave behind code your team owns, not a black box.
Different teams, different first project. The platform underneath is the same.
From prototype to a feature every tenant can switch on, with evaluation and guardrails.
Serving infrastructure, autoscaling, caching and the cost dashboard to go with it.
The architecture answers for SOC 2, ISO 27001 and enterprise questionnaires.
An honest read on which AI ideas on the roadmap are worth funding this year.
Four things, in this order. Each one is useful on its own.
Two weeks to rank the roadmap by value, effort and true running cost.
Retrieval, agents or fine-tuned models, shipped inside your product and your release process.
Dedicated GPU capacity with autoscaling, so inference cost tracks usage instead of leading it.
Regression tests for model behaviour, so the next model upgrade is a decision, not a gamble.
No surprises about sequence, and no invoice before there is something to look at.
We take the AI items on your roadmap and put a real running cost against each one.
Serving approach, tenancy model, data boundaries and the cost ceiling, agreed in writing.
Our engineers work in your repo and your sprint cadence, with your team reviewing.
Your team takes it, or we run the serving layer under an SLA. Both are fine.
The questions we work through before recommending anything — data readiness, hosting constraints, review process and the running cost at year two. Use it with any vendor.
Short answers. Longer ones are a conversation.
It depends on your volume and how much variance you can tolerate. Below a certain steady-state throughput, an API is cheaper and simpler. Above it, dedicated GPUs win — often substantially. We run the numbers on your actual traffic during discovery and show you the crossover point rather than pushing a preference.
Yes, that is the default. Our engineers work in your repository, your branching model and your review process. You own everything we write.
Data boundaries are designed before anything is built: per-tenant retrieval scopes, isolated storage, and controls that prevent one tenant's content reaching another's context window. This is the part enterprise buyers probe hardest, so it gets documented properly.
We provide the architecture documentation, data-flow diagrams and control descriptions that those questionnaires ask for. Your team still owns the response, but they are not writing it from scratch.
Good — that shortens discovery. We assess what is there, tell you what carries over to production and what needs rebuilding, and price from that.
Send us the feature and your rough traffic numbers. We will come back with an approach and a running cost.