All news

Why Pay-Per-Token Pricing Doesn't Scale for Continuous Cybersecurity

· by Alias Robotics

When the security workload decides how much AI reasoning is required, a metered unit turns workload variability into cost variability. Token price is an input metric; the decision-relevant one is cost per validated security outcome.

Pay-per-token doesn't scale for continuous cybersecurity: the same security alert investigated under token-metered inference versus continuous subscription access, showing the same workload under two different economic models

Cybersecurity AI is becoming more agentic, more continuous and more operational. Yet many foundation-model APIs still price inference as a metered input. That creates a question security leaders should examine before comparing model price lists: what happens when the security workload itself determines how much AI work is required?

Security workload sets the consumption curve

A chatbot conversation usually has a visible human rhythm. A user asks, the model answers, the interaction pauses. Many security workflows do not behave that way.

A SOC investigation can begin with one alert and expand into correlation across identities, endpoints and cloud services. Vulnerability validation can require code inspection, tool execution, retesting and evidence collection. DFIR may pull large artefacts into context and revisit earlier hypotheses as telemetry changes.

Agentic systems add another layer. An agent can choose tools, process outputs, retry failures and hand work to other agents. Parallel workflows can improve coverage and speed, but may also increase model calls and context. Alias Robotics' work on cybersecurity agents and multi-scaffold cybersecurity architectures reflects this shift from single-turn assistance towards systems that act, observe and iterate.

None of this means that every cybersecurity task necessarily consumes more tokens than an ordinary AI task. Architecture matters. Context management matters. Tool design matters. Caching matters. The important point is more precise: when additional investigative effort requires additional billable inference, security workload variability becomes AI cost variability.

That matters because cybersecurity demand is not smooth. A normal operating period and a live incident should not be expected to generate the same workload. Alias Robotics' regional machine-scale security assessment is one illustration of how automation can expand the number of assets and findings handled in parallel. Greater operational scale changes the economic question as much as the technical one.

Four constraints that buyers should never confuse

AI pricing conversations often collapse several different constraints into the word “usage”. They are not the same.

  • Per-token billing means cost changes with the number of billable input and output tokens consumed.
  • Rate limits constrain throughput, commonly through measures such as requests per minute or tokens per minute. They can exist whether billing is metered or subscription-based.
  • Usage quotas or subscription allowances define an included quantity of consumption that can renew, be exhausted or trigger another commercial tier.
  • Concurrency and infrastructure limits govern how many workflows can execute simultaneously, or how much compute capacity is available at once.

An API can be pay-per-token without stopping when consumption rises; the bill simply grows as usage grows, subject to the provider's other controls. Equally, an “unlimited” subscription can remove per-token billing while retaining operational rate limits.

For procurement and FinOps, “How much does one million tokens cost?” is only one question. “What happens when an incident multiplies investigative activity?” is the operational one.

Tokens can be cheap and the economic question still matters

The argument against relying on token price as the primary buying metric is not that tokens are universally expensive. Some are remarkably cheap.

As reviewed on 11 September 2026, OpenAI's GPT-5.6 Luna was listed at $0.20 per million input tokens and $1.20 per million output tokens for standard text processing, with higher multipliers for requests above 272,000 input tokens. Anthropic's Claude Sonnet 5 was $2 input and $10 output per million tokens. Z.ai's GLM-5.3 was $1.40 input and $4.40 output per million tokens, while Mistral Small 4 was $0.15 input and $0.60 output per million tokens under standard API pricing.

These are not total-cost comparisons. Caching, tools, processing modes, support and architecture can all change realised cost. They simply show that the argument survives low unit prices: a variable unit multiplied by a workload-dependent quantity still produces variable cost.

Agent architecture can change the bill more than model price

A strong Cybersecurity AI cost model therefore cannot stop at the model rate card.

Endor Labs tested the same model against the same 34 AppSec prompts across 12 large open-source projects using two agent configurations. One had access to precomputed deterministic security evidence; the other had to derive facts from repositories and public web research. Endor reported 6.6 million tokens versus 79.5 million, alongside 1,380 versus 6,165 tool calls.

This vendor-run controlled benchmark is not a universal estimate. It demonstrates a mechanism: harness design, tooling and reconnaissance can materially change inference consumption even when the model and task are held constant.

RunReveal offers a production example without giving us a universal number. Across more than 50,000 investigations and activities, a random cross-section using recent Claude models showed roughly $1 to $3 per Sonnet investigation. RunReveal explicitly says this is not a scientific comparison and that other teams will have different prompts, tools, harnesses and models.

That caveat is the point. There is no single “cost of an AI investigation”. There is a cost function shaped by the model, architecture, evidence available, investigation path and commercial unit being metered.

Chart comparing variable metered AI cost against fixed subscription access as security workload grows from normal operations to complex investigation and incident spike, with rate limits, quotas and concurrency shown as separate operational constraints
As security workload increases, metered AI costs can become more variable. Fixed subscription access changes that relationship, while operational limits remain a separate consideration.

When the incident gets harder, what should happen to the budget?

Pay-per-token does not inherently stop an investigation. Nor is subscription pricing automatically superior.

The concern is incentive alignment.

If each additional loop, validation step or parallel agent creates marginal cost, teams eventually need governance around that cost. Sensible organisations will set budgets, alerts, approval thresholds and spending controls. Those controls are financially rational. But in security they can create an awkward incentive: the investigation can become more expensive precisely when uncertainty, complexity or incident severity requires more work.

The response is not to ignore cost. Wasteful inference is still waste. But investigative depth should be governed primarily by risk, evidence, authorisation and expected security value, not fear of another increment on the inference bill.

A security team should optimise how an investigation runs. It should be more cautious about optimising away investigation itself.

The market is already experimenting with different units of value

Token billing is not the only consumption model emerging in security.

Torq publishes an AI credit model that assigns fixed credit amounts to activities such as an autonomous case investigation or an AI-agent execution. Dropzone AI describes pricing around investigation capacity, publishing up to 4,000 full investigations per AI analyst per year.

These models have their own quotas and trade-offs and are not directly comparable with API token pricing. What they show is more interesting: vendors are experimenting with commercial units closer to operational work than raw inference volume.

That should encourage buyers to ask a better question:

What behaviour does the pricing unit incentivise?

Unlimited.cybersecurity is an economic claim, not an infinite-capacity claim

CSI PRO takes a different approach for included alias2-mini usage. The current plan describes Unlimited* tokens with alias2-mini under a fixed subscription, with the asterisk tied to “good usage” under the licence terms.

That removes per-token billing for the included alias2-mini usage. It does not remove operational limits.

The current Product License Terms & Conditions publish maximum rate limits of 500,000 tokens per minute and 60 requests per minute. The CSI documentation explains that PRO keys are deliberately rate-limited to protect shared inference infrastructure; saturated keys can be throttled or temporarily rejected, with automatic backoff handling the cooldown.

So the distinction is explicit:

Unlimited token access is not unlimited throughput.

Buyers should not infer unlimited concurrency either. Billing, throughput, quotas and concurrency remain separate dimensions.

This is the useful meaning of Unlimited.cybersecurity: remove the marginal token meter from included model usage so that the amount of analysis is not priced one token at a time, while remaining transparent about the operational limits that still govern the service.

Economics is only one axis of Cybersecurity AI

Pricing should also be separated from capability and access policy.

General-purpose AI providers increasingly support authorised cybersecurity work through specialised programmes. OpenAI now operates Daybreak Blue and Daybreak Red, with GPT-5.6 Cyber available through Daybreak Red for approved advanced security workflows. Anthropic says Claude Sonnet 5 participates in its Cyber Verification Program.

The market distinction is therefore not “general-purpose providers prohibit offensive security”. That would be inaccurate.

Alias models are positioned as cybersecurity-specialised models, and Alias' licence explicitly includes authorised penetration testing, security-task automation and defensive use while prohibiting unauthorised attacks and malicious activity. Buyers should evaluate model capability, access controls, authorisation requirements and economics as separate procurement dimensions.

What security buyers should measure beyond token price

Token price still helps estimate marginal inference cost. It is simply insufficient on its own.

A CISO, CTO or procurement team evaluating Cybersecurity AI should ask what the commercial model meters; whether deeper investigations increase billable consumption; what rate limits, quotas and concurrency constraints apply; how costs behave during incident spikes; whether tool calls and large outputs add billable context; and whether the architecture can reduce repeated reasoning through better evidence and orchestration.

Then bring the analysis back to the outcome.

For a SOC, that might be a validated alert disposition with sufficient evidence for human review. For AppSec, a confirmed exploitable vulnerability and a verified remediation. For authorised offensive security, a validated exposure within scope. For DFIR, a defensible finding with traceable evidence and analyst oversight.

This is not an industry-standard accounting metric. It is a better executive frame:

Token price is an input metric. For security buyers, the more decision-relevant metric is cost per validated security outcome, considered alongside time, quality and human review.

Continuous Cybersecurity AI changes the unit economics because the workload can decide how much reasoning is required. The commercial model should make that reality easier to govern, not harder.

Explore CSI PRO and see how predictable access to specialised Cybersecurity AI changes the economics of continuous security.


Sources

All news