A big thank you to those of you who have been building on our Conversational AI (a.k.a. agents) and Conversation Intelligence (a.k.a. in-call analysis) beta releases. We’re seeing some very interesting use cases and huge enthusiasm, which we love. We can’t wait to see your products in the hands of end-users and making a difference in their lives and business - it is why we do what we do.
To that end, we have one humongous project ongoing in relation to Carrier Services billing. Our call billing is mature and untouched but when it comes to other things, like billing number rentals, trunk charges etc. it is overly simplistic. Not the kind of simplistic that gets marketed as ground-breaking, but simplistic enough to have been a major constraint over the years. It just isn’t right that we can solve technical problems only to be constrained by billing agility and while the early temptation is always to deal with exceptions one-by-one, over time that becomes an unmanageable knot. So, for perhaps the last decade, we’ve just accepted the constraint. No more - newer services have so many moving parts and we need to be able to bill like a software company, not be wedded to the very deep but very narrow paradigm of telephony billing. That means bundles, allowances, multiple meters, bespoke subscriptions and all manner of good stuff. That is an ongoing project whose first outing is AI and WhatsApp (before taking over billing existing services), meaning we have a blank canvas. Yay!
Billing this stuff is complicated and when we first started on this journey a few years ago, we wanted to simplify it. It was so insanely complicated it was nigh on impossible to tell what a service was going to actually cost. That’s fair enough for a hobby project, not for building a business on. Along the way we’ve seen multiple providers dramatically change pricing and either row back or publicly apologise. We’re not even into enshitification yet, they seem to be just trying to optimise what they can grab. That isn’t our style but we’ve come to appreciate that simplification dramatically raises risk and either leads to all customers funding the edge cases, or us taking massive arbitrage risk. Best value comes from us exposing some complexity, offset by that which is eliminated by Simwood’s unique position (the only actual carrier globally offering these services on-net), and pricing in a way that enables our customers to price in a genuinely competitive way against those seeking to underpin a bubble!
So, our pricing has 5 moving parts, well, 6 if you consider the big bit that isn’t there - telephony. You don’t need to deal with a reseller of a carrier and pay 2c a minute for the privilege, because all our services are on-net on calls you already pass, on numbers you already have. You can use BYoC if you’re unlucky enough to have numbers locked into a dinocarrier or the mustard morons, but porting to Simwood is within your gift there.
As to the remaining, we’re offering 4 tiers of service and these are available to any production Carrier Services account and are unaffected by service level (Virtual Interconnect, Managed Interconnect, etc.) as the economic benefits of those are exposed through telephony and we want to encourage innovation and growth here. Each tier trades a monthly subscription against consumption economics, but even at the headline lowest commit level, you should be able to compete with any of the bubble-operators. At the top-end, we have completely bespoke pricing for those achieving scale or looking to bring on a 1,000-seat call centre or two. The per minute rate covers all agent and intelligence functions and there is a generous allowance baked into the subscription, the subscription being the cheapest way to buy those minutes. This covers the core functions of our agents: listening, speaking.
The notable exception here is ‘thinking’, i.e. the LLM costs and this is where the real complexity lies. Different models have dramatically different economics, and different customers or end-user profiles have dramatically different consumption patterns - we’ve proved this in the beta. Trying to blend these leads to us overcharging you or taking on significant risk, with fairness being an edge case. So instead, just as you can choose your model, we’re going to be passing on LLM costs at cost. That includes the substantial benefit from LLM-caching, where it is achievable. BYoLLM is on our list but you get the economic benefits of that from day one. The remaining benefit is really operational, e.g. charges at the same rate on your Anthropic bill rather than ours.
There is another critical function which is a huge Simwood USP and that’s ‘remembering’ - storing your knowledgebases for agents or indexing past conversations for state. Other providers are severely limited here, both in terms of functionality and price. ElevenLabs for example will (crudely in our opinion) apply a limit across the entire account - pretty useless - from 2MB up to 1GB on published plans. To access 1GB you’d be spending $990 a month and $0.08 per minute. Twilio is more generous - they’ll charge you $0.018 per GB per hour for as much as you like - so $1,296 for 100GB, before you retrieve any of it at $0.005 per time! We’re storing and indexing on-net, and are not grotesquely greedy, so we’re allowing as much storage as you need with charging storage at a sensible rate, i.e. £0.0018/GB/hour. Furthermore, unlike others, we’ll only charge based on the original document, not the multiple of that after vectorisation!
The final moving part is concurrency and this one is hard. Nobody in the market is overt about their limits and true constraints; it is all fuzzy language and obfuscation. Concurrency risks being incredibly expensive, particularly as we progress to full on-net processing, so we need to find a balance. Accordingly, every tier comes with a small number of included concurrent channels and additional channels can be purchased on top. These are account level limits which you will be able to divide and contend by trunks as you do with other limits on Simwood accounts. They will all be best-effort and have a hard limit, in the interests of fairness. Until POA territory these will be the same rate regardless of subscription level.
So, that’s a subscription which includes minutes, overage at a rate which reflects the subscription, included levels of concurrency increasing with subscription, and the ability to scale that as you need. Oh, and LLM usage at cost - you pick the model and entirely control your own cost profile.
| Plan | Monthly | Concurrent channels | Included minutes(Conversational AI / Conversation Intelligence) | Minutes above that |
|---|---|---|---|---|
| Try | £0 | 5 | 0 | 6.00p |
| Build | £50 | 10 | 1,250 | 5.00p |
| Run | £300 | 20 | 10,000 | 4.00p |
| Bespoke | POA | POA | POA | POA |
Additional concurrent channels £5 a month each, same price on every plan. LLM tokens passed through at provider cost on every plan.
While LLM costs will be passed through at cost it is helpful to record what those costs are (at the time of writing) for anyone wanting to benchmark; the tables are at the end of this post. These are in USD so will vary for billing in GBP. For reference, if you’ve been using our services in beta, you will have used Gemini 2.5 Flash, GPT 4.1 or 5 Mini for Conversational AI and Haiku and Sonnet for Conversation Intelligence. Frontier models are available but absolutely not required. E&OE of course, and these will change, not least due to promotions, new models etc. One worth knowing today: Google’s current Gemini 3.x Flash pricing is promotional and doubles on 1 January 2027, and because we pass cost through, that increase reaches you.
Last but by no means least, we need to talk about data residency. Our ultimate aim is for all of this processing to occur on-net and we’re quietly moving towards that. We don’t intend to narrate every element as in all likelihood the transition will require hybrid working, but where this stuff is processed matters to customers and we know some are using it as a selling point. In every case possible, we are electing for UK processing first, EU if the UK is not available. Google don’t allow nomination but their services are anycast so logically we will naturally be using UK and EU instances first. When it comes to T&Cs however, we think it is more important that services work, so we will permit a fall-back to other instances (e.g. the USA) in the event of the preferred instance not being available.
We hope you’ll share our excitement at this and, naturally, look forward to feedback in our Community Slack or via account managers. When comparing, please do look at others like Twilio, ElevenLabs, Vapi, Retell etc. In every case they either can’t provide, or charge separately for, the connectivity, and the basic costs are substantially higher than ours - you guys should be able to compete with any of them on this pricing before we get into bespoke pricing for legitimate scale.
Finally, as regards timing, the billing is the remaining constraint here. As soon as that is done, these products will come out of beta, and we estimate that’ll be this quarter as billing is a Q3 objective and Charles doesn’t miss! Pete on the other hand…
**
LLM pass-through rates
Tokens are passed through at provider cost, so these are costs, not prices. Per 1,000,000 tokens, from each provider’s own pricing page, retrieved 11 September 2026. FX GBP 0.740295 to USD 1.00, ECB reference rates, 11 September 2026, one rate in every row.
Fast, the live turn loop
| Model | Provider | In $ | Out $ | In £ | Out £ | Cache write $, full cost | Cache read $ | Min tokens to cache |
|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | 0.20 | 1.20 | 0.15 | 0.89 | 0.25 | 0.02 | 1,024 |
| Gemini 3.1 Flash-Lite | 0.25 | 1.50 | 0.19 | 1.11 | not available | not available | not available | |
| Gemini 3.5 Flash-Lite | 0.30 | 2.50 | 0.22 | 1.85 | not available | not available | not available | |
| Claude Haiku 4.5 | Anthropic | 1.00 | 5.00 | 0.74 | 3.70 | 1.25 | 0.10 | 4,096 |
Standard, a capable agent turn and post-call analysis
| Model | Provider | In $ | Out $ | In £ | Out £ | Cache write $, full cost | Cache read $ | Min tokens to cache |
|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 Flash | 0.30 | 2.50 | 0.22 | 1.85 | not available | not available | not available | |
| Gemini 3.8 Flash | 0.75 | 3.75 | 0.56 | 2.78 | 0.75 | 0.075 | 4,096 | |
| Gemini 3.5 Flash | 1.50 | 9.00 | 1.11 | 6.66 | 1.50 | 0.15 | 4,096 | |
| Claude Sonnet 5 | Anthropic | 2.00 | 10.00 | 1.48 | 7.40 | 2.50 | 0.20 | 1,024 |
| GPT-5.6 Terra | OpenAI | 2.00 | 12.00 | 1.48 | 8.88 | 2.50 | 0.20 | 1,024 |
| Claude Sonnet 4.6 | Anthropic | 3.00 | 15.00 | 2.22 | 11.10 | 3.75 | 0.30 | 1,024 |
Frontier
| Model | Provider | In $ | Out $ | In £ | Out £ | Cache write $, full cost | Cache read $ | Min tokens to cache |
|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 Pro, to 200k | 1.25 | 10.00 | 0.93 | 7.40 | 1.25 | 0.125 | 2,048 | |
| Gemini 3.1 Pro Preview, to 200k | 2.00 | 12.00 | 1.48 | 8.88 | 2.00 | 0.20 | 4,096 | |
| Claude Opus 5 | Anthropic | 5.00 | 25.00 | 3.70 | 18.51 | 6.25 | 0.50 | 512 |
| GPT-5.6 Sol | OpenAI | 5.00 | 30.00 | 3.70 | 22.21 | 6.25 | 0.50 | 1,024 |
| Claude Fable 5.1 | Anthropic | 10.00 | 50.00 | 7.40 | 37.01 | 12.50 | 0.25 | 512 |
A note on caching. Google charge no premium to write a cache, so their write figure is simply the standard input rate, and the storage charge they publish applies to caches you choose to store rather than the automatic caching in use here, so there is nothing extra to pass on. Where a model is marked not available, Google publish no cached rate for it at all, so there is no caching benefit to pass through at any price. The minimum matters as much as the rate: below it a request is simply processed uncached, with no error and no warning.