PortaAIM: The Inference Service Delivery Platform

PortaAIM brings real-time billing to AI infrastructure, combining PortaBilling with an AI gateway for authorization, usage charging, balance control, and settlement.

Metering is only the easy part

Once customers start using AI, billing logic quickly spreads into code, pricing, balances, and settlement.

PortaAIM handles the billing logic

Set the commercial rules once and apply them consistently around every AI request.

You don’t have to build AI billing yourself

PortaAIM pairs an AI gateway with PortaBilling’s proven charging engine.

Authorize before it runs

Check the balance and reserve funds before a request starts, not after.

Bill in any unit

Rate tokens, images, audio, tool calls, or other units without schema changes.

Catch abuse before it costs you

Block suspicious usage, free-tier abuse, and spend spikes before they consume capacity.

Bill every response correctly

Rate cached, completed, and blocked requests correctly instead of losing billable usage.

Route to a capable model with the lowest cost

Choose the lowest-cost model that can handle each request automatically.

White-label the customer experience

Give customers your branded billing portal for usage, charges, invoices, and payment collection.

Built on 25 years of real-time billing

PortaAIM brings PortaBilling’s carrier-grade charging experience to AI infrastructure.

PortaBilling has spent more than 25 years rating, authorizing, and billing high-volume communications for nearly 500 operators in over 100 countries. PortaAIM applies the same proven flexible approach to AI-enablement.

Full API and MCP access

Connect external applications, services, and agentic AI workflows directly to PortaAIM.

Choose how you pay

  • Perpetual license – Pay upfront, with integration support, proactive monitoring, and 24/7 coverage backed by response commitments.
  • Pay as you go – Pay monthly based on usage, with the same support commitments.

Choose where it runs

  • Your own infrastructure – Run in-country or in your own environment when deployment requirements demand it.
  • Hosted by PortaOne – Or start hosted, with setup and support included if that is simpler to begin with.

How it works

Every AI request moves through the same billing flow, from authorization to final charge.

An AI request is only as safe as what happens before it runs.  PortaAIM already knows the model and maximum response size, so it can calculate the worst-case cost before a single token is used. It reserves that amount, then charges for actual usage and releases the rest when the request finishes.

1

Authorize

PortaAIM checks the customer, service, pricing rules, and available balance before execution.

2

Reserve

Funds are reserved against the maximum expected cost before the request runs.

3

Execute

The approved request passes through the AI gateway to the selected model or provider.

4

Meter

PortaAIM records the actual billable usage in tokens, images, audio, tool calls, or other units.

5

Reconcile

The final charge is booked, and any unused reserved balance is released automatically.

Built for reliable AI infrastructure

Billing without adding delay

PortaAIM checks pricing and balance in real time without adding noticeable delay to the request.

Keep running if a region goes down

PortaAIM can run across multiple sites, so service continues if one region goes down.

Extend it without coding

Add local billable units or provider connections through documented API endpoints

AI governance guardrails

Configure guardrail modules to support risk-based governance frameworks such as the EU AI Act.

PortaAIMGW v4 illustration

See PortaAIM running

Book a demo and see the complete flow in action.

FAQ

Billing and enforcement

Most usage-based billing tools calculate a balance after a request finishes, then expect your own code to enforce it. PortaAIM’s gateway is built in, and it does the enforcement itself by authorizing and reserving the cost before the request runs, not after.

PortaAIM reserves the worst-case cost before the request starts, based on the model and the requested output size. If the balance can’t cover it, the request is stopped before it runs, not after.

Actual, per request. The worst-case cost is reserved before the request runs, then the real amount used is measured and reconciled once it completes, not estimated at a session or monthly level.

We’re following the new agent payment protocols closely. The account and hierarchy layer underneath them, who’s allowed to spend how much, checked before it happens, reconciled at month end, is close to what PortaBilling already does for telecom. If you’re building on x402, AP2, or MPP and need a real ledger behind it, get in touch to talk about a pilot.

Cached responses and requests blocked by a guardrail are rated on their own terms, not billed the same as a completed request, so invoices transparently reflect actual activity.

Yes. PortaAIM monitors for compromised API keys, free-tier farming, and unusual spend velocity, using the same fraud-detection patterns built for prepaid telecom accounts.

Yes. PortaAIM can route each request to the lowest-cost model that still meets your quality and provider requirements, the same way carrier billing has done least-cost routing for voice traffic for decades.

Deployment and infrastructure

Yes. PortaAIM can be deployed in your own environment, including in-country for sovereign or data-protection requirements, with source code supplied as standard.

PortaBilling runs inside your environment too, not just the gateway. Usage is tracked and billed locally, without sending data to an external system.

Yes. Self-hosted components report back to the same billing engine, so nothing needs manual reconciliation between the two.

Balance enforcement runs from one central ledger, not per site, so usage from anywhere is checked against the same live balance. The system is built for high availability across regions, so that central check doesn’t become a single point of failure.

Yes. PortaAIM supports bring-your-own-key, so your provider contracts and credentials stay with you, not routed through a third party. Useful for teams with data protection or sovereignty requirements around provider credential ownership.

Pricing, resellers, and multi-tenant

Yes. Rate plans can combine a committed baseline with usage-based charges on top, the same way telecom tariffs have worked for decades.

Yes. Metered usage and flat capacity charges run through the same rating engine, so they don’t end up as two separate systems to reconcile.

Tokens-per-minute and requests-per-minute limits from a provider are pooled and allocated across your tenants, so one customer can’t silently use up capacity that another one needs.

Yes. Reseller settlement runs in the same system regardless of whether the underlying capacity is yours or sourced from another provider.

Yes. The end-customer portal can be fully white-labelled, so your customers see your brand, not PortaOne’s.

More posts about AI​

Join those
'in-the-know'

Never miss an update, software release, webinar, best practice or anything else.

Search PortaOne

Search

Hot topics