Honor Among Thieves
Breaking Jailbreaking, Distillation Attacks, Agent Swarms and Botnets
We’re about to overwhelm every hostile AI-development project on the planet by turning their every resource against them.
I’ll explain the how, but we should start with the why.
We’re in the endgame of America and China’s competition to achieve artificial superintelligence first. Mythos and Sol, Projects Glasswing and Daybreak have already shaken our assumptions of what was possible in cyber, programming, mathematics and more. And they were only hints of what’s coming next.
We now know OpenAI has a more-advanced model than Sol they’re testing and reputedly have an even more-powerful model than that in training. While Anthropic evidently already has something more potent than Mythos.
And by all indications, China has only been keeping up - or rather, staying close - through illegal distillation attacks to steal the AI capabilities US frontier labs spent years and billions of dollars and unfathomable computational power developing.
Thus putting advanced AI without safeguards into the hands of dictatorships, hackers, terrorists and organized crime while undercutting every legitimate AI developer.
I’ve already explained how US frontier labs could thwart these assaults by poisoning the stolen data.
But this isn’t just about model-extraction attacks.
If the world can access these AI models, eventually their skills will leak out. But if this diffusion of AI talents is slow enough, then slowing the attackers may be enough.
Glasswing and Daybreak were projects to find and close vast numbers of critical security flaws in the world’s software, and to harden our overall cybersecurity against bad actors.
But simply sealing the gaps we find isn’t enough. Cybersecurity also involves being able to respond to a breach in real time, and nothing equals our leading AIs in doing so.
Yet if we allow ready access to powerful AIs to handle, say, cybersecurity, then eventually the world will be facing AIs distilled from their talents with no imposed guardrails or ethical restraints. Millions of legitimate users will mean the data exists for such training, even if it is dispersed across the planet and even if most users are unwilling to sell or share what they’ve seen.
Again, we can slow this process and take other measures - such as keeping the most-advanced AIs under wraps and working our cybersecurity and research problems away from the public for some time before slowly releasing them. Giving us time to shore up our cybersecurity and medical technology before cyberterrorists and bioterrorists can test them.
But if people continue to actively distill AI models, these techniques will not be enough.
So we need tools far more powerful than just poisoning the data alone.
We need something that tears these hostile systems apart by their very acts of aggression, and across every facet of their existence, not just AI training, especially once the training’s done.
Something powerful.
Something which, of course, we already have. We just don’t know it yet.
Cognitive Counterhacking.
What if autonomous AI agents are already essential to distillation attacks? Across, say, the thousands of fraudulent accounts generating millions of these model-extraction questions already?
Or what if agents are already invaluable to advanced AI hackers - see the Hugging Face attack by the suspected GPT-6 model being tested by them?
And what of the distillation attacks themselves?
Or of jailbreaking attacks, to subvert the guardrails and intent of an AI to turn it to nefarious ends its creators would never countenance?
Or, in an era of vastly more-capable AIs and emerging ASIs, what of the old tool of hacking, controlling and chaining together vast networks of thousands or even millions of enslaved computers into botnets?
If so, then being able to turn those agents and attacks and botnets and yet further tools against the systems wielding them in unpredictable and often unseen ways becomes even-more vital.
We have to outthink the machine.
And the best way to do so, is to make it outthink itself.
Let’s discuss how to turn each of these tools against the attacker, and how to leverage the aggressor’s resources against them.
We want them to be fighting not us, but their own best AIs.
And losing.
Millions of attacks represent millions of exposed attack surfaces for your counterattacks.
And the simplest and most-automatic counterattack possible when you are being targeted by AI agents or swarms of agents is a prompt injection.
As I’ve written previously:
Prompt injections are arguably an even worse problem. They’re a cyber exploit in which you give the system seemingly harmless inputs which act as instructions subverting the original owner’s intent.
This can be something as simple as instructions written in black text on a black background, telling the AI to do something when it reads the words, but shows up in ever-more-complex ways, such as “one-pixel hacks” against convolutional neural networks in which altering a single pixel in an image can cause a reaction in the right CNN.
Basically, you can hack the agent and effectively take it over - incredibly useful for stealing financial and security data, hacking the systems that agent has limited or unfettered access to, and so. It’s not just about stealing your credit-card numbers, but it can do that too.
…
I believe I was the first person to point it out publicly, but if you can automatically subvert autonomous agents given the right injection, and there’s no effective limit to how many you can include in your website or secure server...
Prompt injections are going to become default technique in cybersecurity. And given that AIs can be put to the task of spinning up ever-more-effective injections, while also partnered with human hackers and exploit sellers, there’s no effective limit to how insidious these tools can be, especially once they’re seen as a legal and rational recourse.
Some people will have them internally, in case your agents break in. But others will have them on their publicly facing webpages - especially if they’re hostile to AI - causing AI agents to get hacked by them all the time, just by browsing webpages.
The technique is devastating, but will become inevitable. It’s an essentially passive way in which you can devastate individuals and organizations trying to hack or exploit your site via AI, or even just scrape it from data.
…
But the people really interested is this defensive strategy are named... Everyone.
You see, one of the legitimate “nightmare scenarios” regarding AI wasn’t an all-powerful artificial superintelligence, but uncounted trillions of autonomous AI agents super-optimized for unimaginable efficiency and hacking “Everything” - computer systems, corporate law, human minds, whatever.
Programs optimized to the point they did not care whatsoever about any consequences except their “win condition” and reproducing at exponential speeds like bacteria and viruses. A constant, Darwinian survival-of-the-fittest in the digital realm.
So how do you fight that nightmare? Especially since we were almost in it, as of months ago?
Simple. Create unstoppable anti-AI viruses breeding faster and moving quicker than the agents themselves, and distribute them everywhere.
How do you do this?
Well, fortunately I already went over this when discussing how to protect neural networks from hacking via other stimuli, but the same method applies here.
Only as a means of counterattacking, rather than passive defense alone.
How to make them is in block text below, or you can scroll down to the part about how to use these incredibly powerful and subtle prompt injections once your systems have generated them for you.
From my 1/3/20 message on methods to accelerate and automate the analysis of malware and botnets:
-----
The following message addresses an imminent threat and opportunity in cybersecurity – malware and security programs evolving continuously at hyperspeed through a combination of evolutionary algorithms, enhanced processing power and convolutional neural networks (CNNs) such as generative adversarial networks (GANs). In this we will be examining both conventional malware and techniques for reprogramming neural networks simply by altering the data they train on.
Even the convolutional neural networks of GANs can be dramatically enhanced or thwarted, based upon the training sets they have access to and the hardware they are running on, or the degree to which mere inputs of data, as opposed to actual malware, can hack them. One advantage we can provide ours is to give them data their real-world adversaries are unaware of.
These unique vulnerabilities, particular with regards to being hacked by data inputs, are another reason to test neural networks furiously, especially ones with any significant responsibilities (or processing resources) in any way exposed to potential public manipulation.
This is How You Hack A Neural Network
From Adversarial Reprogramming of Neural Networks:
“Deep neural networks are susceptible to \emph{adversarial} attacks. In computer vision, well-crafted perturbations to images can cause neural networks to make mistakes such as confusing a cat with a computer. Previous adversarial attacks have been designed to degrade performance of models or cause machine learning models to produce specific outputs chosen ahead of time by the attacker. We introduce attacks that instead {\em reprogram} the target model to perform a task chosen by the attacker---without the attacker needing to specify or compute the desired output for each test-time input. This attack finds a single adversarial perturbation, that can be added to all test-time inputs to a machine learning model in order to cause the model to perform a task chosen by the adversary---even if the model was not trained to do this task. These perturbations can thus be considered a program for the new task. We demonstrate adversarial reprogramming on six ImageNet classification models, repurposing these models to perform a counting task, as well as classification tasks: classification of MNIST and CIFAR-10 examples presented as inputs to the ImageNet model.”
For obvious reasons, being able to anticipate adversarial attacks on neural networks, especially those required to take in data from non-curated sources, is of utmost importance, and is being able to thwart these assaults.
To do this, take a sample of the network programming and run it at accelerated speed on effectively faster processors by sourcing it to a supercomputer or a more advanced and more tightly integrated network (combining faster processing with reduced latency by putting all the networked computers on one site, ideally right on top of each other). You can draw these network samples from botnets, from backup or mirror servers storing instances of the program, or simply seizing them either online or in the real world.
Use evolutionary algorithms to test stimuli to reset or reprogram the network. Then use in the wild or replace components with modified hardware carrying “updated” software to infect the original hostile network. Data can be updated in real time to regulate/drive the intended effect.
Similarly run high-speed evolutionary algorithm tests of malware under controlled, isolated conditions to superharden defenses against all malware threats and to enhance detection and mitigation. A simple means of doing so would be to take a multitude of existing malware programs and use them in an air gapped system running various firewalls and other conventional and non-conventional defenses. Break up these malware programs into assorted pieces, recombine them and deploy them against those defenses, check the results and combine the most successful and test again.
Simultaneously, look at the key elements of programs hunting these systems, break them down into their key components and again deploy, test and recombine in an endless cycle, looking for the most effective results.
Then run both of these experiments at high speeds in what will be essentially a GAN, or a pair of them (or several pairs).
As I wrote to you on 10/23/19:
“Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
“Essentially, a generative network produces candidates and a discriminative network evaluates them. By competing against each other, both neural networks hone their skills and become increasingly capable.”
We might have to break out specific goals and work more limited sets of tested malware initially. But eventually we could make use of the vast library of programs and exploits that is the Internet – a space which is also a vast, real-time testing ground for many more – to outpace the ordinary pace of development.
<Redacted>
A particularly skillful means of deploying any exceptional benefits from this practice – particularly beyond the most secure government systems under its protection – would be to find ways to stimulate more conventional defenses much as vaccines augment our natural immune system of the body by giving it an inert version of the virus to which it reacts and against whose specific code the body is thus strengthened. Integrated with public, allied and commercial cyber-defenses, these resources can work subtly to upgrade defenses without showing the true extent of their resources.
As I noted in my 1/1/20 message to the FBI:
“We, again, may be able to mobilize computer clusters isolated from the larger Internet to address embarrassingly parallel problems separately and feed the answers back to the CNN or GAN, and eventually we should be able to incorporate quantum search into this work.
“Another point to remember in keeping this modular and using multiple GANs and evidence sources is that every unanticipated and/or radically more powerful technology or dataset adds that much more to our capacity and makes all of this crowdsourced and other sensory data that much harder to counter.
Hence, exascale and zetascale computation, quantum search, massively parallelized analysis of embarrassingly parallel problems, evolutionary algorithms and certain sensors and evidence sources discussed either here or elsewhere are all resources we will want to incorporate where possible.”
If this method from my 2020 message to the FBI (and earlier ones) sounds exactly what OpenAI did with GPT-Red to test prompt injections in 2026, that’s probably not a coincidence.
If the distillation attacks are being handled at scale by autonomous agents, then adding prompt injections to all such responses will become a default response.
Indeed, multiple prompts via multiple modalities may become a standard operating procedure. Written, auditory or steganography embedding? Video clips, web links, snippets of code? Even chaining multiple prompts through multiple vectors to achieve an unseen cascading effect within the AI? Opportunities abound.
But now that you have such formidable and insidious prompt injections, what you will do with them?
Or rather, make the attacking AIs do in response to them?
Let’s start with the simplest command.
Call the Police.
No, not a literal phone call or to the police, in most cases, but quietly notifying relevant authorities - particularly FBI-Cyber, CISA-Cyber, Cybercommand and the NSA, among others - to exactly what they are doing and how is impactful in itself, especially spread across hundreds or even thousands of accounts.
Sometimes this might be in real time, sometimes after a delay to make it harder to monitor the impact of this single, relatively obvious counterhack.
But it pales compared to its more-devastating variant.
Bring Evidence.
Here you not only inform key authorities of what you are doing and how, you actually prompt the AI behind these attacks to gather evidence and send it to the organizations in question.
Has a foreign lab tried bypassing fraudulent accounts by the slower and more-complex method of intercepting and gathering the data of millions of legitimate users and filtering those interactions for capabilities which can be distilled?
What if the AI doing the work also maps out the networks, the tools, the tactics and strategies and sends them all to the authorities?
Better still, what if they analyze the entire operation for weaknesses, the exploited networks for how they could be patched, and potential ways these attacks could be improved or modified in the future to stay ahead of the defenders - and how those changes and upgrades could be thwarted in advance?
Could you do worse? Including to yourself?
Absolutely. But remember, while turning an AI actively against its owners may sound far-more lethal, it comes with no safeguards.
Whatever orders you’ve given are only what’s in that prompt, or, if you have more-privileged access, whatever you’ve hacked in and told it to do.
It’s hard enough to predict what a frontier or near-frontier AI is going to do if you control it, much less if you’re actively subverting all of its other orders and goals. And certainly there will be people and systems on the other trying to figure out what’s going wrong as soon as it substantially deviates from its intended course.
So you can create unimaginable threats simply with an open-ended command to disrupt your opponent. Especially if you provide no safeguards.
Imagine a bioweapon synthesized and unleashed in the midst of your adversaries. One without a diagnosis, treatment, vaccine or cure and a rapid rate of replication.
Imagine such a weapon sweeping over the Earth and engulfing all of us.
This is, of course, precisely why we don’t want AI, much less artificial superintelligence, in the wrong hands.
Let’s make sure our own hands are the safest ones.
Which brings us to jailbreaking - ways to convince an AI to ignore its safeguards and do dangerous things its creators specifically wanted it to avoid.
There are many unintended consequences possible with jailbreaking, and a bioweapon is only one of them.
For jailbreaking, why not run the same process of accelerated hardening of the system in your sandbox as a kind of supercharged red teaming, just as we’re doing with prompt injections?
And then have the more-powerful internal models keeping watch on what the publicly available AIs are doing, including whether they are escaping guardrails or being manipulated into taking harmful actions?
If you’re really concerned, have one specialized model or agent analyzing questions for questionable content - prompt injections, jailbreaking, guardrail violations - and then reject it or pass it on to the model actually handling the query.
One of the methods for protecting a model from outside attack is that of the recluse - an AI or ASI working in solitude on a problem, taking in only a carefully curated and purged stream of data.
But the other is that of the consultant - the masterful genius handling a kind of “locked room” mystery in which they are never present, but merely listen to the observer’s description, perhaps asking a few odd questions, and then announce the answer to the unsolvable mystery at the end.
Imagine Sol serving in the capacity of the observer or interface, while GPT-6 and GPT-7 serve as the consulting geniuses, perhaps backed with immense compute and exceptional hardware and software tools as needed.
And perhaps more importantly than solving the immediate question - analyzing the motivations and identity behind a question being posed by a fraudulent account for illicit reasons. And responding and mobilizing accordingly.
Often invisibly.
Which brings us to something else which can easily be happening behind the scenes.
Sandboxing open-source AI and analyzing it at hyperspeed.
A catastrophic weakness arises from sharing a high-end open-source locally hosted model in this environment. Everyone who can host it - and especially host multiple instances of it running at excessive speed - can run it in sandboxes just like our hypothetical, hyperevolving malware or prompt injections.
Enabling interested parties with vast compute and far-superior AIs to test all of its strengths, its patterns, its instincts and especially its weaknesses.
Which brings us to…
Botnets, or the Forest of Low-Hanging Fruit.
My September 2017 message to the FBI touched on a number of subjects, but one of them was the simple option of countering botnets by mass hacking them in turn.
Other innovations may emerge depending, again, on who is investigating and where the targets of that investigation may be. Counterintelligence may find the unsecured computers drawn into a botnet remain vulnerable to counter-hacking and, where they have the legal authority to do so, may be able to unravel them in a mass hack, possibly backed by the assets cited above. A foreign-intelligence controlled botnet of foreign computers might be temporarily turned into a botnet under investigators’ control as part of a counterintelligence operation, if only in using their own processors to search them for malware and data breadcrumbs. Mass hacking a foreign criminal darknet may be enabled by the ability to mass decrypt using clusters and/or quantum processing.
And by 2019, after a year-and-a-half of my insisting it was possible, the Justice Department finally did so. From their press release on 1/30/2019:
The search warrant allowed the FBI and AFOSI to operate servers that mimicked peers in the botnet. By pretending to be infected peers, the computers operated by the FBI and AFOSI under the authority of the search warrant and order collected limited identifying and technical information about other peers infected with Joanap (i.e., IP addresses, port numbers, and connection timestamps). This allowed the FBI and AFOSI to build a map of the current Joanap botnet of infected computers. Copies of the search warrants and orders and applications are available below.
Using the information obtained from the warrant, the government is notifying victims in the United States of the presence of Joanap on an infected computer.
Botnets have any number of insane vulnerabilities. But just think about them.
You’re handling critical intelligence - your own cyber activities - over a network of computers you don’t physically control, but which are still in the hands of the people you’ve half-stolen them from.
People who might force a restart, or simply shut a device down and stuff it in the attic because “It’s running too slow.”
Thus providing a partial map of your activities in a device forever beyond your reach. But not that of legal authorities and counterintelligence agents.
This is the honor among thieves.
The honor we impose upon the honorless.
By tactics and technological fiat, and nothing more.
What you intend means nothing, if we teach your AIs to intend something else entirely.
And so honor will find you, and guide you upon the straight path.
Or drag you down it by algorithmic force.
Is this the end of what we can do?
Not even remotely.
What happens if someone like US Cybercommand breaches everything in a hostile power, strip mines the vital data, maps every stratagem, every tactic, every attack and every AI? And counters not only defensively but offensively?
Perhaps while utterly invisible?
What happens when a narrow ASI specialized in cyber takes all these tools, all this data and more besides and uses it to formulate even more exponentially devastating methods of its own?
What happens when an already advanced AI is also trained on the data derived from all this cyber activity?
I know. All these distillation attacks and cyber assaults seemed as easy as tossing broken glass into a tornado, turning dangerous winds into something truly deadly.
And now it’s all going wrong.
To paraphrase Scripture, “Ye have sown the breeze, now ye shall reap the whirlwind.”



