We recently published a post on “tokenomics” and “denial of wallet” attacks — a new type of DDoS attack aimed at overspending AI tokens. But before security specialists begin protecting their colleagues’ AI processes from this new threat, they should remember that the work of the information security department itself can become a target for such an attack. Since security teams are now actively testing and implementing various types of AI-based automation, the threat of external interference aimed at paralyzing the defense system directly affects their own processes and technologies.
Elastic has estimated that, depending on the “agent-based SOC” architecture, triaging a single alert from a host can cost $0.69 for an ensemble of highly specialized agents or $3.42 for a general-purpose agent with 14 different skills.
However, these averages mask a huge range of variation. At any given moment, the basic processing of a simple alert (say, a thousand tokens) can turn into a cascading investigation involving processing millions of log entries, API calls, and other related data — which amounts to millions of tokens in a matter of minutes.
Elastic itself acknowledges that it uses the Claude Sonnet 4.6 model in its SOC. This immediately raises the question: what happens when the budget is exhausted or the model’s security filters are triggered? In the worst-case scenario, the following occurs: the AI SOC launches an “orchestra of agents” to investigate suspicious events that resembles lateral movement in the infrastructure. The agent calls the Anthropic API, but the department’s monthly budget has been exhausted, and the API returns errors. The investigation grinds to a halt, and people find out about it half a day later. Moreover, if the constantly changing “security filters” of a major AI provider flag the SOC’s requests as malicious, the same effect is possible even without the budget being exhausted — as demonstrated by Hugging Face’s experience investigating the breach of their infrastructure by OpenAI.
Of course, an attacker has a vested interest in such an outcome. They can directly influence the defenders’ costs by injecting malicious prompts into the data available to them (DNS TXT records, HTTP headers, code comments, filenames, and so on), counting on the fact that these prompts will be processed by information security systems. For example, in the GhostJacking attack case, WAF logs serve as the vector for prompt injection, and in academic study, another method was proposed: researchers managed to attack the LLM security guardrails themselves, causing their requests to increase latency by 148 and the number of tokens consumed by 63. In other words, the system theoretically continues to function correctly, but the latency in processing alerts and the associated costs increase significantly.
Dangerous silent failure
In agent-based scenarios, budget exhaustion doesn’t always manifest as an explicit error with a “request limit exceeded” message and an immediate alert to the responsible parties. Under certain configurations, “guerrilla” scenarios can arise: a request to the model times out, the system automatically switches to a cheaper model, and as a result the quality of its outputs drops, and subagents and periodic tasks stop running. Meanwhile, the security team may continue to believe for hours on end that everything is working as before.
How to control AI costs in cybersecurity
The task of managing expenses, of course, isn’t simply a matter of “preventing overspending”. The industry has already fallen into this trap when implementing SIEM: when companies had to pay based on the volume of telemetry collected, they began to selectively — and not always effectively — limit log collection. As a result, they ended up with detection blind spots. The same situation could arise with tokens — but faster and on a larger scale. A poorly designed cost-cutting policy could compromise the depth of investigations and the quality of responses. And in the absence of such a policy, the decision to cut something will be made not by the architect or the CISO during the system design phase, but by the analyst on duty at 3am. Or it may be an automatic limiter set by the system vendor.
A few simple principles can help avoid predictable cost overruns, and better manage situations that are truly unpredictable.
Don’t feed the AI model anything that can be verified with a standard query. Everything that’s repetitive and known in advance should be handled the old-fashioned way — with rules and data queries. These can be developed using AI, but they must operate deterministically. The model should only be used where an assessment of an ambiguous situation is truly needed.
Use strict limits and aggressively alert users when they are exceeded. Limits should be combined: a limit per task, a daily limit, and so on. Alerts about limit exceedances must immediately reach both on-duty analysts and those responsible for the system as a whole.
Decide in advance what happens when a limit is reached. Whether to stop and save money — while losing some control — or to continue and pay, is a decision that must be made before an incident occurs.
Limit agents’ permissions and the set of tools available to them. The fewer actions the system has access to, the more effectively it works on a narrow task, the fewer opportunities it has to inflate costs, and the less likely it is to be “exploited” by an outsider.
Strictly scan external, untrusted data. Incoming requests, inquiries, messages, and comments, as well as various technical fields capable of containing arbitrary text (DNS records, HTTP headers, file names) can not only lead to prompt injection, but also deliberately inflate the workload. Limit their size, and monitor the load they generate.
To avoid service outages from external providers and maintain full control over your organization’s data, evaluate the possibility of using local models deployed on-premises at least for information security needs. While these models may not be state-of-the-art, their narrow specialization and fine-tuning with company data will deliver decent performance without unexpected failures.
In a more complex architecture, the use of LLMs can be divided into several levels. The first line of defense (routing, noise filtering, data enrichment) is handled by compact on-premise models. This significantly reduces budget volatility. The second level consists of a large local model that handles the substantive work: linking events, testing hypotheses, and piecing together the picture of an incident. It’s also not billed per request, but its throughput is limited; therefore, instead of budget overruns, overloads and event queues may occur. Only isolated, truly difficult cases are forwarded to advanced cloud-based models — preferably with the explicit approval of a human analyst. In this scheme, variable costs remain, but they’re limited to a small number of cases, and each one represents a conscious choice rather than a side effect of a script running overnight.
Essentially, all of this replicates the classic multi-tiered SOC model — where the first line filters, and the third line investigates — only the roles are played by models rather than people.
AI
Tips