{"id":56523,"date":"2026-10-07T15:13:03","date_gmt":"2026-10-07T19:13:03","guid":{"rendered":"https:\/\/www.kaspersky.com\/blog\/?p=56523"},"modified":"2026-10-07T15:13:03","modified_gmt":"2026-10-07T19:13:03","slug":"tokenomics-ai-and-cybersecurity","status":"publish","type":"post","link":"https:\/\/www.kaspersky.com\/blog\/tokenomics-ai-and-cybersecurity\/56523\/","title":{"rendered":"The direct impact of &#8220;tokenomics&#8221; on cybersecurity processes"},"content":{"rendered":"<p>We recently published a <a href=\"https:\/\/www.kaspersky.com\/blog\/tokenomics-ai-cost-ddos\/56455\/\" target=\"_blank\" rel=\"noopener nofollow\">post on \u201ctokenomics\u201d<\/a> and \u201cdenial of wallet\u201d attacks \u2014 a new type of DDoS attack aimed at overspending AI tokens. But before security specialists begin protecting their colleagues\u2019 AI processes from this new threat, they should remember that the work of the information security department itself can become a target for such an attack. Since security teams are now actively testing and implementing various types of AI-based automation, the threat of external interference aimed at paralyzing the defense system directly affects their own processes and technologies.<\/p>\n<p>Elastic has <a href=\"https:\/\/www.elastic.co\/security-labs\/blog\/agentic-soc-token-budget-architecture\" target=\"_blank\" rel=\"noopener nofollow\">estimated<\/a> that, depending on the \u201cagent-based SOC\u201d architecture, triaging a single alert from a host can cost $0.69 for an ensemble of highly specialized agents or $3.42 for a general-purpose agent with 14 different skills.<\/p>\n<p>However, these averages mask a huge range of variation. At any given moment, the basic processing of a simple alert (say, a thousand tokens) can turn into a cascading investigation involving processing millions of log entries, API calls, and other related data \u2014 which amounts to <a href=\"https:\/\/www.securityweek.com\/the-ai-token-costs-that-can-break-cybersecurity\/\" target=\"_blank\" rel=\"noopener nofollow\">millions of tokens in a matter of minutes<\/a>.<\/p>\n<p>Elastic itself acknowledges that it uses the Claude Sonnet 4.6 model in its SOC. This immediately raises the question: what happens when the budget is exhausted or the model\u2019s security filters are triggered? In the worst-case scenario, the following occurs: the AI SOC launches an \u201corchestra of agents\u201d to investigate suspicious events that resembles lateral movement in the infrastructure. The agent calls the Anthropic API, but the department\u2019s monthly budget has been exhausted, and the API returns errors. The investigation grinds to a halt, and people find out about it half a day later. Moreover, if the constantly changing \u201csecurity filters\u201d of a major AI provider flag the SOC\u2019s requests as malicious, the same effect is possible even without the budget being exhausted \u2014 as demonstrated by <a href=\"https:\/\/fortune.com\/2026\/07\/20\/hugging-face-turns-to-chinese-open-source-ai-to-fend-off-autonomous-ai-cyber-attack-after-american-ai-guardrails-stymie-defense\/\" target=\"_blank\" rel=\"noopener nofollow\">Hugging Face\u2019s experience<\/a> investigating the breach of their infrastructure by OpenAI.<\/p>\n<p>Of course, an attacker has a vested interest in such an outcome. They can directly influence the defenders\u2019 costs by injecting malicious prompts into the data available to them (DNS TXT records, HTTP headers, code comments, filenames, and so on), counting on the fact that these prompts will be processed by information security systems. For example, in the <a href=\"https:\/\/thehackernews.com\/2026\/08\/threatsday-ghostjacking-ai-attacks.html\" target=\"_blank\" rel=\"noopener nofollow\">GhostJacking<\/a> attack case, WAF logs serve as the vector for prompt injection, and in academic study, another method was proposed: researchers managed to attack the <a href=\"https:\/\/arxiv.org\/abs\/2606.14517\" target=\"_blank\" rel=\"noopener nofollow\">LLM security guardrails <\/a><a href=\"https:\/\/arxiv.org\/abs\/2606.14517\" target=\"_blank\" rel=\"noopener nofollow\">themselves<\/a>, causing their requests to increase latency by 148 and the number of tokens consumed by 63. In other words, the system theoretically continues to function correctly, but the latency in processing alerts and the associated costs increase significantly.<\/p>\n<h2>Dangerous silent failure<\/h2>\n<p>In agent-based scenarios, budget exhaustion doesn\u2019t always manifest as an explicit error with a \u201crequest limit exceeded\u201d message and an immediate alert to the responsible parties. Under certain configurations, \u201cguerrilla\u201d scenarios can arise: a request to the model times out, the system automatically switches to a cheaper model, and as a result the quality of its outputs drops, and subagents and periodic tasks stop running. Meanwhile, the security team may continue to believe for hours on end that everything is working as before.<\/p>\n<h2>How to control AI costs in cybersecurity<\/h2>\n<p>The task of managing expenses, of course, isn\u2019t simply a matter of \u201cpreventing overspending\u201d. The industry has already fallen into this trap when implementing SIEM: when companies had to pay based on the volume of telemetry collected, they began to selectively \u2014 and not always effectively \u2014 limit log collection. As a result, they ended up with detection blind spots. The same situation could arise with tokens \u2014 but faster and on a larger scale. A poorly designed cost-cutting policy could compromise the depth of investigations and the quality of responses. And in the absence of such a policy, the decision to cut something will be made not by the architect or the CISO during the system design phase, but by the analyst on duty at 3am. Or it may be an automatic limiter set by the system vendor.<\/p>\n<p>A few simple principles can help avoid predictable cost overruns, and better manage situations that are truly unpredictable.<\/p>\n<p><strong>Don\u2019t feed the AI model anything that can be verified with a standard query. <\/strong>Everything that\u2019s repetitive and known in advance should be handled the old-fashioned way \u2014 with rules and data queries. These can be developed using AI, but they must operate deterministically. The model should only be used where an assessment of an ambiguous situation is truly needed.<\/p>\n<p><strong>Use strict limits and aggressively alert users when they are exceeded. <\/strong>Limits should be combined: a limit per task, a daily limit, and so on. Alerts about limit exceedances must immediately reach both on-duty analysts and those responsible for the system as a whole.<\/p>\n<p><strong>Decide in advance what happens when a limit is reached. <\/strong>Whether to stop and save money \u2014 while losing some control \u2014 or to continue and pay, is a decision that must be made before an incident occurs.<\/p>\n<p><strong>Limit agents\u2019 permissions and the set of tools available to them. <\/strong>The fewer actions the system has access to, the more effectively it works on a narrow task, the fewer opportunities it has to inflate costs, and the less likely it is to be \u201cexploited\u201d by an outsider.<\/p>\n<p><strong>Strictly scan external, untrusted data. <\/strong>Incoming requests, inquiries, messages, and comments, as well as various technical fields capable of containing arbitrary text (DNS records, HTTP headers, file names) can not only lead to prompt injection, but also deliberately inflate the workload. Limit their size, and monitor the load they generate.<\/p>\n<p>To avoid service outages from external providers and maintain full control over your organization\u2019s data, evaluate the possibility <strong>of using local models deployed on-premises<\/strong> at least for information security needs. While these models may not be state-of-the-art, their narrow specialization and fine-tuning with company data will deliver decent performance without unexpected failures.<\/p>\n<p>In a more complex architecture, <strong>the use of LLMs can be divided into several levels<\/strong>. The first line of defense (routing, noise filtering, data enrichment) is handled by compact on-premise models. This significantly reduces budget volatility. The second level consists of a large local model that handles the substantive work: linking events, testing hypotheses, and piecing together the picture of an incident. It\u2019s also not billed per request, but its throughput is limited; therefore, instead of budget overruns, overloads and event queues may occur. Only isolated, truly difficult cases are forwarded to advanced cloud-based models \u2014 preferably with the explicit approval of a human analyst. In this scheme, variable costs remain, but they\u2019re limited to a small number of cases, and each one represents a conscious choice rather than a side effect of a script running overnight.<\/p>\n<p>Essentially, all of this replicates the classic multi-tiered SOC model \u2014 where the first line filters, and the third line investigates \u2014 only the roles are played by models rather than people.<\/p>\n<input type=\"hidden\" class=\"category_for_banner\" value=\"mdr\">\n","protected":false},"excerpt":{"rendered":"<p>We examine the impact of AI dependency on an organization\u2019s cybersecurity. <\/p>\n","protected":false},"author":2722,"featured_media":56524,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1999,3051],"tags":[1140],"class_list":["post-56523","post","type-post","status-publish","format-standard","has-post-thumbnail","category-business","category-enterprise","tag-ai"],"hreflang":[{"hreflang":"x-default","url":"https:\/\/www.kaspersky.com\/blog\/tokenomics-ai-and-cybersecurity\/56523\/"},{"hreflang":"ru","url":"https:\/\/www.kaspersky.ru\/blog\/tokenomics-ai-and-cybersecurity\/42810\/"}],"acf":[],"banners":"","maintag":{"url":"https:\/\/www.kaspersky.com\/blog\/tag\/ai\/","name":"AI"},"_links":{"self":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56523","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/users\/2722"}],"replies":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/comments?post=56523"}],"version-history":[{"count":1,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56523\/revisions"}],"predecessor-version":[{"id":56525,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56523\/revisions\/56525"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/media\/56524"}],"wp:attachment":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/media?parent=56523"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/categories?post=56523"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/tags?post=56523"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}