Monetizing AI: why publishers are stuck
The BYOLLM trap: when the publisher gives up control
Anonymous testimonial, product director at a business software publisher (300+ customers)
The PoC was magic. One week of testing, a few tokens burned on a proprietary LLM, and our customers were blown away: the AI worked.
Then industrialization derailed. Our healthcare and public-sector accounts refused to depend on a non-EU LLM, which is understandable. Our SMB customers, meanwhile, were mostly afraid of missing out on AI. So we figured the simple solution was to let them bring their own API key and just handle the integration ourselves.
It was a slow-motion disaster. SMBs have no data team, no GPUs, no visibility into what’s running on their side. Our AI consumes, sometimes fails: who pays, who debugs, is it the model, our API, or their infrastructure? Nobody could say.
We recently spent four days on a single customer just tracing each step. Multiply that by fifty customers, and that’s two hundred days of support a year, pure waste.
We were looking for a way out. Token-based created too much commercial friction. Premium had the same problem. BYOLLM was the worst of the three.
What we didn’t know is that there was a fourth path: taking back technical control without becoming an infrastructure company ourselves. We came across hosted inference platforms like Agora, and it changed everything. We were completely blind to that option, and clearly we weren’t the only ones.

The token/volume paradox
Token prices keep falling. Every new model generation costs less per call than the last. On paper, AI should get cheaper to integrate every year.
It doesn’t, for a simple reason: a simple conversational exchange burns a few hundred tokens; an agentic workflow (plan, call tools, verify the result, retry if needed) can burn ten to fifty times more to handle a single user task. Volume climbs faster than price falls.
The result: your inference bill can double year over year even as each individual token costs less than it did last year. The intuition that AI gets cheaper over time doesn’t hold at the scale of a full agentic task: we break down this gap between sticker price and real cost in a dedicated article.
This paradox changes the original question. It’s no longer just about who pays the bill, you or the customer, it’s about picking a model that survives usage that can grow tenfold overnight, on an architecture you don’t control.
Four models facing the paradox
Each monetization model absorbs this paradox in its own way. None of them cancels it out.
Token-based: you charge for actual usage
Every token costs money. You pass it on to the customer.
Upside: predictable revenue for you, aligned with real usage. The customer sees exactly what they’re paying for.
Trap: the customer absorbs the volume explosion directly on their bill. They don’t see “token” as a business metric, just that their invoice doubled this month. They’ll want a cap, and you’ll get the calls.
Premium: AI as an optional paid module
You add a line item to the bill, say $2,000/year for AI.
Upside: predictable revenue, protected margin. Customers who pay are engaged.
Trap: the flat fee doesn’t move when volume climbs. You collect the same amount while inference costs you more: margin erodes quietly. And on the sales side, across the pricing engagements we run, 40 to 60% of customers refuse an extra line item at every renewal.
Embedded: AI is bundled into the existing subscription
No separate line item, AI becomes just another feature.
Upside: mass adoption, zero commercial friction. Customers use it without debate.
Trap: no revenue to absorb the shock. Every spike in agentic usage on the customer’s side costs you margin directly, with nothing in return. It’s a retention tactic, not a revenue source.
Outcome-based: billing follows a delivered result
Billing follows an observed business outcome (a ticket resolved, an invoice processed, a candidate qualified), not the underlying technology.
Upside: price tracks perceived value, and the volume paradox becomes painless for the customer: they’re not paying for your tokens, they’re paying for a result.
Trap: you still have to measure that result unambiguously, a heavy contractual exercise. And if the agentic workflow needs several attempts to get there, you’re the one absorbing that invisible extra cost.

The real question: who’s actually in control?
The token/premium/embedded/outcome-based debate distracts from a blunter question: who controls the architecture?
Without control over your inference infrastructure, you depend on three things:
- Token volume, not unit price. You just saw it with the volume paradox: this trajectory escapes you as long as you’re calling a proprietary LLM without controlling the architecture around it.
- Latency and availability. If the API goes down, your customers go down with it.
- Your customers’ compliance requirements. Those who refuse any dependency on non-EU infrastructure force you into fragmented, case-by-case support.
BYOLLM seems to solve the third problem, since the customer brings their own infrastructure. It creates two new ones:
- Fragmented support. You have to debug N different configurations, each with its own provider, its own quotas, its own outages.
- Invisible dependency. The customer’s technical team becomes your critical path. If they don’t know how to configure it, you’re stuck.
BYOLLM isn’t a solution, it’s an abdication.
Taking back control without carrying the infrastructure
There’s a fourth path: a hosted, controlled inference platform. The principle: you control the architecture, the customer keeps their sovereignty.
The publisher from the testimonial above discovered Agora, a multi-agent platform with integrated inference, on-premise, in a sovereign cloud, or in a private cloud. No more relying on an external proprietary LLM, no more letting each customer manage their own model.
Concrete results:
- The volume paradox disappears. On-premise, every extra agentic request costs you at the margin, not at an external provider’s rate: more volume no longer means a bigger bill.
- Unified support. One architecture, one stack. No more “it’s your model, my API, their infrastructure.”
- Customer sovereignty. Data never leaves, which brings your offering into compliance with healthcare, public-sector, and defense requirements, a market that was previously out of reach for you.
- Product autonomy. You iterate and test new models freely, without depending on a proprietary API.

Decision table: which model for what?
| Model | When to use it | Predictable revenue? | Absorbs the volume explosion? | Lost customers? |
|---|---|---|---|---|
| Token-based | Highly specialized AI (OCR, fraud detection) | No | No, the customer absorbs it | Few |
| Premium | General-purpose AI, customers willing to pay for modular add-ons | High | No, the publisher absorbs it | 40 to 60% |
| Embedded | Retention tool, marketing argument | Indirect | No, even worse | Few |
| Outcome-based | Unambiguously measurable business outcome | Medium (delayed) | Partly, decoupled from displayed volume | Few (heavy negotiation) |
| Controlled inference (Agora) | General-purpose AI with sovereignty constraints | High | Yes, near-zero marginal cost | Nearly none |
The final question
Three questions help you decide.
Do you have customers who refuse to depend on a non-EU LLM? Token-based or BYOLLM will cause churn among them. Controlled inference settles the question.
Does your AI do several things (general-purpose) or a single measurable task (specialized)? General-purpose: embedded or premium can work, but controlled inference remains the ideal option. Specialized and measurable: token-based or outcome-based become tolerable.
Is your token volume growing faster than you expected? If so, that’s the volume paradox already at work in your business, the one controlled inference neutralizes.
Two “yes” answers out of these three signal that BYOLLM isn’t a solution but a band-aid, whose bill arrives through support.
Choosing the right pricing model is one thing; making sure the AI you're charging for absorbs a volume explosion without derailing is another. Agora Software helps publishers secure that industrialization, from the business model all the way to deployment.
Bring AI into your software with Agora Software.
Let's talk