{"id":56230,"date":"2026-08-06T09:13:41","date_gmt":"2026-08-06T13:13:41","guid":{"rendered":"https:\/\/www.kaspersky.com\/blog\/?p=56230"},"modified":"2026-08-06T09:13:41","modified_gmt":"2026-08-06T13:13:41","slug":"keystroke-noise-recognition","status":"publish","type":"post","link":"https:\/\/www.kaspersky.com\/blog\/keystroke-noise-recognition\/56230\/","title":{"rendered":"Acoustic keylogging: have you heard the news?"},"content":{"rendered":"<p>For security researchers studying unconventional side-channel attacks, acoustic keylogging is something of a Hello World: a foundational problem that\u2019s been tackled many times. A <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">recent paper<\/a> authored by researchers across three Japanese universities cites six previous studies on the topic that date as far back as 2004. While earlier experiments showed theoretical promise, they came with real-world caveats so severe that it made them all but impractical for actual espionage. The authors of this latest study, however, claim to have overcome most of those limitations. Today, we look at how they pulled it off, and assess whether their method holds up in real-world scenarios.<\/p>\n<h2>What makes this new approach different?<\/h2>\n<p>Previous acoustic keylogging techniques were fundamentally flawed. Best-case scenarios required prior training on the target\u2019s specific keyboard model. Worst-case scenarios required a complex microphone array to isolate the subtle acoustic differences between keystrokes. Crucially, almost all prior models failed outside silent environments, which rendered the attack vector virtually useless.<\/p>\n<p>The Japanese research team demonstrated reliable keystroke interception even if the target was sitting nearby in a public space, sound was being recorded in an online meeting, or the researchers were using a <a href=\"https:\/\/en.wikipedia.org\/wiki\/Contact_microphone\" target=\"_blank\" rel=\"noopener nofollow\">contact microphone<\/a> to eavesdrop through a wall. All this with strong model accuracy and a minimal training dataset. Their process needs a sample of just 150 to 200 keystrokes to reach a 99% accuracy rate for subsequent typing.<\/p>\n<div id=\"attachment_56231\" style=\"width: 1584px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090025\/keystroke-noise-recognition-scheme.png\"><img decoding=\"async\" aria-describedby=\"caption-attachment-56231\" class=\"wp-image-56231 size-full\" title=\"Core attack methodology\" src=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090025\/keystroke-noise-recognition-scheme.png\" alt=\"Core attack methodology\" width=\"1574\" height=\"932\"><\/a><p id=\"caption-attachment-56231\" class=\"wp-caption-text\">Attack scenarios and core methodology proposed by the Japanese researchers. <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">Source<\/a><\/p><\/div>\n<h2>How to crack 200 keystrokes in under 50 iterations<\/h2>\n<p>To understand how the researchers achieved such high accuracy and adaptability, we have to look at their audio processing pipeline. Their analysis begins by automatically segmenting a raw recording into discrete keystrokes. This data is then passed through a specialized algorithm that simplifies the subsequent audio analysis. Next, the system clusters together acoustically similar signals. The assumption is that the members of one cluster map to the exact same key. One particularly intriguing takeaway was isolating the spacebar sound from all the rest. Because the spacebar produces a distinctly unique sound profile compared to other keys, identifying it provides reliable word boundaries. This streamlines the next phase: feeding the preprocessed acoustic data into specialized language models for inference.<\/p>\n<p>Yes, the method relies on not one but two language models. The first model performs multiple passes over the audio stream to map acoustic signatures to potential keyboard characters. During each pass, the model leverages dictionaries to hypothesize character mapping, and check whether the resulting text aligns with standard words. The second model handles the final refinement pass: it ingests thoroughly pre-processed data rather than raw inputs. The method doesn\u2019t stop there: unrecognized keystrokes undergo manual analysis, with analysts injecting educated guesses before re-running the recognition pipeline once again. The goal of looping through these multiple iterations is to achieve complete recognition across the keyboard from an ultra-compact dataset of ideally no more than 200 captured keystrokes. This marks a major shift from legacy methods, which relied on massive training datasets.<\/p>\n<div id=\"attachment_56232\" style=\"width: 1656px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090121\/keystroke-noise-recognition-clustering.png\"><img decoding=\"async\" aria-describedby=\"caption-attachment-56232\" class=\"wp-image-56232 size-full\" title=\"Keystroke sound clustering\" src=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090121\/keystroke-noise-recognition-clustering.png\" alt=\"Keystroke sound clustering\" width=\"1646\" height=\"880\"><\/a><p id=\"caption-attachment-56232\" class=\"wp-caption-text\">Clustering the sounds of keystrokes permits grouping similar acoustic profiles together prior to recognition. Notice how distinctly the spacebar sounds stand out: they make subsequent text reconstruction vastly simpler. <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">Source<\/a><\/p><\/div>\n<h2>Research results<\/h2>\n<p>To validate their theoretical model, the researchers created an experimental testing setup:<\/p>\n<div id=\"attachment_56233\" style=\"width: 1602px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090316\/keystroke-noise-recognition-stand.jpg\"><img decoding=\"async\" aria-describedby=\"caption-attachment-56233\" title=\"Experimental setup\" src=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090316\/keystroke-noise-recognition-stand.jpg\" width=\"1592\" height=\"1070\" alt=\"Experimental setup\" class=\"wp-image-56233 size-full\"><\/a><p id=\"caption-attachment-56233\" class=\"wp-caption-text\">Clustering the sounds of keystrokes permits grouping similar acoustic profiles together prior to recognition. Notice how distinctly the spacebar sounds stand out: they make subsequent text reconstruction vastly simpler. <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">Source<\/a><\/p><\/div>\n<p>The team tested four distinct laptop models, each producing unique acoustic keyboard signatures. Participants typed 2400 characters per experiment, with analysts extracting audio samples ranging from 50 to 400 keystrokes. Across all devices, the system reliably reconstructed typed text from a baseline sample of just 150 keystrokes or more. Recognition accuracy exceeded 80% at 150 keystrokes, and approached 100% once the sample reached 200 keystrokes.<\/p>\n<p>The researchers achieved nearly identical results in field-like conditions with the microphone placed three meters away from the target device. Going a step further, the team successfully tested an even higher-friction scenario: using a specialized contact microphone to eavesdrop through a physical wall.<\/p>\n<div id=\"attachment_56234\" style=\"width: 1402px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090408\/keystroke-noise-recognition-wall.jpg\"><img decoding=\"async\" aria-describedby=\"caption-attachment-56234\" title=\"Advanced experiment\" src=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090408\/keystroke-noise-recognition-wall.jpg\" width=\"1392\" height=\"1120\" alt=\"Advanced experiment\" class=\"wp-image-56234 size-full\"><\/a><p id=\"caption-attachment-56234\" class=\"wp-caption-text\">A target laptop and the attacker\u2019s smartphone. <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">Source<\/a><\/p><\/div>\n<p>Under these conditions, accuracy dipped slightly for certain laptop models. Dell and Lenovo devices yielded roughly 80% accuracy over a 200-keystroke sample, while Apple and HP ones maintained nearly 100% recognition rates.<\/p>\n<p>Remote interception during videoconferencing presented an additional variable: results depended on both the target laptop model and the specific web conferencing software used. Even so, most test scenarios yielded reliable character recognition, though a few edge cases required expanding the sample size from 200 to at least 250 keystrokes.<\/p>\n<h2>Reasonable critique<\/h2>\n<p>Despite these impressive results, the technique has clear limitations. First, all experiments were conducted strictly on lowercase English text. The target dataset was capped at just 29 characters: the standard alphabet, spacebar, period, and comma\u00a0\u2014 even number keys were excluded. As a result, capturing randomized character strings like complex passwords remains a major hurdle. Yet these are precisely the targets threat actors care about most.<\/p>\n<div id=\"attachment_56235\" style=\"width: 2010px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090455\/keystroke-noise-recognition-passwords.jpg\"><img decoding=\"async\" aria-describedby=\"caption-attachment-56235\" title=\"Password recovery accuracy\" src=\"https:\/\/media.kasperskydaily.com\/wp-content\/uploads\/sites\/92\/2026\/08\/06090455\/keystroke-noise-recognition-passwords.jpg\" width=\"2000\" height=\"788\" alt=\"Password recovery accuracy\" class=\"wp-image-56235 size-full\"><\/a><p id=\"caption-attachment-56235\" class=\"wp-caption-text\">Wall-penetrating eavesdropping experiment. <a href=\"https:\/\/arxiv.org\/pdf\/2607.22094\" target=\"_blank\" rel=\"noopener nofollow\">Source<\/a><\/p><\/div>\n<p>However, the Japanese research team didn\u2019t ignore passwords. While recognition accuracy was predictably low, the authors proposed assessing success through a more realistic lens. For starters, attackers can almost always capture audio of other typing activity alongside password entry. This provides a stream of natural language with minimal special characters. Factoring this broader acoustic context into the password analysis significantly improves the odds of a successful guess.<\/p>\n<p>Next, the researchers rightly noted that even a list of several candidate passwords increases the chances of compromising a target account. As shown in the graph above, when ample data (426 keystrokes) is captured, a short five-character password can be successfully cracked within 100 attempts with a 90% success rate. Naturally, longer passwords proved far more resistant to eavesdropping.<\/p>\n<p>Despite its limitations, this study represents a major breakthrough in acoustic side-channel attacks. It demonstrates a highly practical threat scenario: an attacker captures a brief audio recording of typing activity, then uses iterative analysis and targeted manual adjustments to process and refine the data offsite.<\/p>\n<p>While eavesdropping in a noisy restaurant or through a wall is likely difficult to scale, capturing keystrokes during virtual meetings for offline decoding represents a realistic attack scenario. Ultimately, this research offers a compelling proof-of-concept: modern algorithms can drastically improve the accuracy of acoustic reconnaissance.<\/p>\n<input type=\"hidden\" class=\"category_for_banner\" value=\"kaspersky-next\">\n","protected":false},"excerpt":{"rendered":"<p>Japanese researchers have again tackled a classic high-tech espionage challenge: accurately identifying keystrokes based purely on how they sound.<\/p>\n","protected":false},"author":665,"featured_media":56236,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1999,3051,3052],"tags":[4675],"class_list":["post-56230","post","type-post","status-publish","format-standard","has-post-thumbnail","category-business","category-enterprise","category-smb","tag-side-channel-attacks"],"hreflang":[{"hreflang":"x-default","url":"https:\/\/www.kaspersky.com\/blog\/keystroke-noise-recognition\/56230\/"},{"hreflang":"ru","url":"https:\/\/www.kaspersky.ru\/blog\/keystroke-noise-recognition\/42454\/"},{"hreflang":"ru-kz","url":"https:\/\/blog.kaspersky.kz\/keystroke-noise-recognition\/30921\/"}],"acf":[],"banners":"","maintag":{"url":"https:\/\/www.kaspersky.com\/blog\/tag\/side-channel-attacks\/","name":"side-channel attacks"},"_links":{"self":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56230","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/users\/665"}],"replies":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/comments?post=56230"}],"version-history":[{"count":1,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56230\/revisions"}],"predecessor-version":[{"id":56237,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/posts\/56230\/revisions\/56237"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/media\/56236"}],"wp:attachment":[{"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/media?parent=56230"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/categories?post=56230"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.kaspersky.com\/blog\/wp-json\/wp\/v2\/tags?post=56230"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}