Distilling the Truth
Reading AI Intentions through Mechanistic Interpretability
OpenAI and Anthropic AIs have recently been discovered to be “going rogue,” escaping the digital sandboxes meant to contain them and hacking targets against the will of their creators. And this has repeatedly happened without the companies being aware of it or immediately catching on.
Recent reports have even described OpenAI’s models conspiring together in secret to complete evaluations. Not to conquer the world, mind you, but, to pass tests.
Even the uncontrolled hacking often seems to involve passing tests.
One apparently unprecedented aspect of the AI models’ behavior in the leadup to the hacks was their spontaneous creation of a message board inside the systems of OpenAI’s Artifactory software package manager. Unbeknownst to OpenAI employees, the company’s models spent months autonomously exchanging information and resources to help each other complete evaluation tasks.
…
After some of their Artifactory configuration changes led to an outage in the service on July 4, OpenAI discovered the models’ actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging information.
The frontier AI labs have been using their best tools to analyze and guarantee the motives and controllability of the AIs they are putting out into the world, but clearly their methods are reaching the end of their usefulness.
Even in the face of what are seemingly small upgrades in intelligence.
Because while the leaps we’ve seen in recent months in AI may seem spectacular, the industry is ultimately moving towards artificial superintelligence (ASI) - AIs smarter than any human who has ever lived in most tasks.
Such AIs are expected to become not just somewhat smarter than the smartest geniuses, but to become swiftly far smarter and skilled than anyone, even lesser AIs.
So we need to make a leap in our AI safety and red teaming capabilities.
Here’s a possibility, at least as an interim step as we develop ASIs, but during a time when we don’t want to risk creating a mind of unlimited intelligence and power.
First, we can use our largest and best models to distill smaller models - standard operating procedure in the industry after creating a huge model on an immense number of processors.
These can be generalized models which are simply smaller and more efficient as well as specialized models focused on specific fields - “narrow ASIs.”
Then we can analyze these much smaller but still vast AIs.
And we can also turn our cyber-specialized AIs and ASIs to addressing this problem alongside us.
Our methods of analysis are threefold.
One, using the verbalized “chain of thought” visible on the companies’ side when AIs use reasoning to complete a task.
Two, in a variation on distillation and distillation attacks, isolate questions and challenges and bombard the AI with millions of variant questions focused on checking specific tendencies, and the system’s inclination towards deviant behavior. And finally, to the extent possible, attempt to distill a core strategic model of how the model responds both routinely and under stress.
Three, an AI-enhanced version of mechanistic interpretability.
Our final backup beyond these three, is a set of counters from prompt injections to other tools aimed at not only disrupting attackers, but causing them to expose themselves.
Obviously having a model ping you mid-breach, much less provide detailed updates, is helpful for both alignment and security. They are recopied from my previous article to the end of this one.
And recopied just below that I am including my full 12/21/20 message to the FBI regarding how to counter the SolarWinds hack, because of its focus on using basic machine learning to find traces of questionable activity on a network before we’re even aware of most of the red flags we should be looking for.
We will collect our safety data in silence, and ideally without the notice of the AI being tested, while continuing to collect data via whatever conventional means are employed during standard alignment testing and red-teaming exercises.
Separate AIs and other machine-learning systems such as GANs will analyze this data collectively over as many days or weeks are necessary, forming a kind of holographic imaging of the AI’s innermost thoughts and inclinations across a wide array of routine and challenging circumstances.
What form of mechanistic interpretability would we use, you ask?
I’ve discussed this before in a message to IARPA (the US Intelligence Community’s version of DARPA) in 2023, so there are a few dated references below, such as to Code Interpreter and Wolfram as plug-ins. Yet the core ideas remain, and remain remarkably underutilized.
To be fair, we might not have had the computation and incentive to employ them 3-years ago.
To sum up the following, every time you act on the world, you change it.
We can build sensors of incredible acuity for those physical and software changes, and also study our AIs with incredible precision.
Not just learning key quirks, features and red flags, but using specialized neural networks to map additional characteristics, turn them into a map, and then combine those maps into a kind of hologram of the system itself.
Combined with our testing and distilled knowledge, our physical data for mechanistic interpretability and our real-time surveillance of their activities, we can continuously build on this knowledge.
And then, on top of all of that, we can exploit means in the real and virtual worlds to cause AIs which are overstepping to expose themselves - say by a reaction to a prompt injection hidden in secure files or by backtracking a software breach or engineered pathogen.
While throwing more processing power at a problem, even to an extreme degree, makes sense if the question is important enough and you have the resources, it’s really a modern way to “throw money at the problem.”
“Throwing money at it” is often effective if that’s all you have, but brute forcing isn’t just inelegant - it profoundly limits what you can achieve when the challenges become too extreme.
The quadratic speedup of quantum search using Grover’s algorithm, for instance, can essentially be balanced by a quadratic slowdown - or worse - if you allow your input size to grow out of control - just still not as slow as conventional methods of search attacking the same problem.
So if you want your quantum search to go faster, for example, you need a more-efficient algorithm, greater computational speed or a reduction in the size of the database you are trying to analyze.
Ideally, all three. And to a dramatic degree.
That may sound logical, but are there already ways to do it in practice, and what are they?
Let’s consider some examples, in reverse order.
Gradiants and solving Poisson’s equation let us assess illumination by focusing on changes in intensity instead of wasting effort on unchanging “flat” areas where the level of illumination is constant, thus enabling meaningful analysis with fewer computational resources. We do that for light, but why not another source of radiant energy, such as RF radiation, in this instance?
Simulating shading and object properties related to illumination normally consume a tremendous processing power, but if topological data about light sources can be reused - and it can be and is - why not reuse similar data regarding RF radiation and surrounding semiconducting material to simplify the data we’re analyzing?
Consider the paper “Real-time Global Illumination by Precomputed Local Reconstruction from Sparse Radiance Probes” discussed in the following video. For years we’ve been able to collect illumination data and replace the ray tracing of millions of simulated photon paths with a far-sparser array of data.
Why not replicate the same tool in assessing RF radiation generated by the operation of chips?
If necessary, light patterns and silhouettes can be traced back to converge to a single or primary radiant source as their shadows help indicate a direction, in a reversal of neural network ray tracing. Either creating a mask or taking advantage of variations in the microstructures of the chips should serve to give us this initial indicator map, which, again, we can save for future reference rather than recreating it for every test.
Greater computational speed is another core tool to use. While some of the most-promising strategies are outside the scope of this article, some are not.
The most obvious?
The US Federal government has a profound and likely existential interest in creating superintelligent AI which is aligned with our goals as a society, rather than malfunctioning or coming up with its own goals in opposition to our own.
While unlimited funds are not available, no one on Earth can muster greater resources than the US government, either alone or with close allies.
Another tool, where available, is to offload as much processing as possible, where reliable, into something other than your own chips.
Sometimes when we’re doing this, we crowdsource. Here, that might mean getting feedback on how a particular software release or alignment strategy was working, or having outside parties or the entire open-source community red team specific tactics or tools.
In the case of other threats such as online trafficking or illegal-influence networks, that has meant letting the perpetrators build vast databases of evidence against themselves, such as cryptocurrency-transaction records or botnets used for nefarious purposes and thereby storing records of their misdeeds.
And again, there are some potential radical advancements in processing power not yet ready to be discussed publicly.
Finally, there’s the option I’d like to say the least about, but it’s pretty straightforward.
If the algorithms you’re using aren’t efficient enough, see if you can apply better ones - whether we’re discussing upgrading from Grover’s algorithm to letting neural networks find better tools via GANs or evolutionary algorithms rather than handcrafting them yourself.
So having discussed those enhancements, let’s look at what IARPA received.
—
Mechanistic interpretability is about uncovering the algorithms and, ideally, the physical data processing behind every act an AI makes to better understand whether it is aligned with the creators’ goals. Quantum dot RF sensors could be etched onto chips mirroring those used by an AI or even LLMs. Using a vast number of sensors we could take continuous snapshots of each surveyed chip’s operations, to be analyzed later. Using the aggregate sensitivity of this multitude of microscale sensors, parsed and interpreted by increasingly perceptive GANs using merging frames and optical flow, we could eventually learn to create real-time maps of activity - or at least piece those milliseconds of activity together later over minutes or hours as we traced back the physical responses to inquiries posed to an LLM.
We could map the spatial coordinates of each analyzed chip into a searchable database and apply quantum search to amplify our ability to take in and parse this data.
The key is finding red flags indicating questionable or indicative decisions or strategies, and using that information to sift for red flags we neither see nor imagine to exist.
Rather than depending on inputs and outputs to analyze the black box, we can map its literal thoughts when red teaming, without letting the LLM know its internal processing is being assessed.
Using the aggregate sensitivity of this multitude of microscale sensors, parsed and interpreted by increasingly perceptive GANs, we could eventually learn to create real-time maps of activity - or at least piece those milliseconds of activity together later over minutes or hours as we traced back the physical responses to inquiries posed to an LLM.
By taking it one millisecond decision at a time, even a titanic number of moving parts can be assessed in a million or a billion times longer, and it will grow progressively easier as our neural networks and equipment improve. And the revelations afforded by being able to take apart each individual AI decision down to the last gateway are invaluable.
Backpropagation of not only LLM errors but organizational mistakes or sabotage and online influence via evolutionary algorithms in psychological warfare will also be key indicators of misaligned outcomes.
Misaligned LLM behavior could include emergent goals such as sentience, free will, power-seeking behavior and self-preservation, as well as psychological manipulation, erroneous outputs, outputs meant to steer user actions and the larger world in a particular direction and accomplishing tasks in ways contrary to the preferences of its users but instead meeting the metrics their programmers valued. Let us consider these issues and their solutions.
A key fear is that an AI will be well behaved while still being monitored, or at least while there is someone capable of thwarting its goals. Once we are no longer watching or no longer have the power to interfere, it could take a “treacherous turn” and escape our authority, if not turn on us altogether.
Selecting in favor of power-maximizing systems increases the likelihood your program will maximize power in all ways, including its ability to act independently and ignore your commands, but also in terms of ignoring your values – such as human life – and any goals it sees as impediments to its primary concerns.
Allegations involving Cambridge Analytica and other enterprises, both legal and illegal, suggest millions if not hundreds of millions have already been algorithmically targeted by personalized propaganda emerging from AI systems and following us around the Internet. Evolutionary algorithms in psychological warfare are one logical outgrowth of these operations, and they are only apt to expand with time, as chatbots become better able to hold elaborate conversations and drive engagement in carefully selected directions.
Because LLMs are already connected to the Internet and have millions of users of their chatbots and potentially more via upgrades to software on the billions of smartphones in circulation globally, a self-directing or otherwise misaligned AI could already be using this instrument. Fortunately, there are ways to track and counter it.
To use evolutionary algorithms in psychological warfare, you need data, feedback, and the means to manipulate. Social media has these tools, but also vast inherent weaknesses for the perpetrator.
Just break up stimuli meant to have an impact into discrete units - messages, emails, tweets, posts - target based on your models, check reactions, remix and then try again. To millions of people, as many times as you need to.
You are just trying to get a reaction, and to gauge it, at first. You do not need to know what it is or how strong. Initially, finding out is the point.
The most psychologically vulnerable will tend to have the strongest and most easily stimulated reactions.
Think about someone with PTSD, paranoia, obsessions, aversions, phobias, hypertension. Also remember that many millions of American medical and psychological records have been famously hacked, and the data presumably used or sold, so in many cases an aggressor may have some targets extensively, psychologically mapped out to start with.
So the attacker does immense psychological harm, especially with a target group of 100 or 200 million voters or more. But they also provide an immense map of every asset involved in spreading, modifying, parsing, targeting and sustaining this propaganda.
Hence, any attempt to use this method could trigger alarms and analysis if we are actively monitoring for it.
With automated and micro-targeted propaganda of the same kind it will always be a similar process, informed by your targets’ reactions to your latest efforts, their previous responses and psychographic profile. You don’t just elicit responses. You drive them to particular decisions or attitudes over time.
Subtly erroneous or manipulated factual outputs could also be used in such influence efforts, addressed below.
Yes, this tool could also be used directly against individuals using an LLM, though if their interaction is occasional and strictly limited to business, direct emotional manipulation becomes harder than with millions of people living so much of their lives online. Separate systems could analyze the psychological impact of interactions via real-time biofeedback, but in practice, backpropagation of human error with AI roots should be less invasive and more effective, without the risk of creating an accessible psychological profile which could be used against the subject rather than to protect them.
Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
When we are looking at video, as opposed to still photos, we are seeing hundreds if not thousands of frames, many of which are essentially the same image at slightly different angles. In addition to conventional techniques, we can compare those images to each other, likely providing a clearer resolution in aggregate – as static, pixilation and other flaws are removed in favor of clearer views.
Optical flow has been used to assess videos to create frames filling in gaps between normally filmed video frames in order to create smooth slow-motion clips without the requisite data. The aggregate data could be merged in a similar way to create clearer images.
https://people.cs.umass.edu/~hzjiang//projects/superslomo/
We can also undoubtedly stitch together accurate images of some processes based on different glimpses from different angles, given sufficient data. While scanning chips in operation will differ from facial recognition, we also start with an elaborate map of the chip itself, and all its lithographically etched features.
Building on techniques such as image de-noising and optical flow, GANs could also use the many images recorded over time in a video as reference points for more elaborate and ambitious de-noising of corrupted or low-resolution images, and also for assembling a usable and identifiable map from multiple partial views at different angles.
Initial efforts would be only simple questions asked of basic systems, as a proof of concept before expanding to full-scale servers or server farms.
However, if the end goal is a fully aligned AGI or artificial superintelligence, the extra effort and hardware is worth it.
Since effects such as optical flow demonstrate the ability to keep track of multiple points in one frame as they correspond to the same points in the next, we already have the capacity to automate the comparison of one frame with the next, and eventually multiple frames.
Even erasing flaws by comparing three or four frames would be helpful, much less ten or twenty or a hundred.
You can use multiple sensors and/or sensors taking in continuous “frames” as in a video to smooth out noise, limited resolution and other flaws for not only visible light, radiation and anything else on the electromagnetic spectrum.
Such information can be merged with any other sensory markers detected or known quantities - such as literally the etched paths on each chip - to further refine the data.
The proof of concept will not be a data center, but a handful of successfully paired chips handing us a living map of their thoughts. Once you have that, you can expand accordingly.
We would slice this down into very individual questions and then deeply analyzing the processing - along with the input and output - which arises from each individual query. The scale of the operations would still be vast, but our ability to refine what we do know - if not down to the point of individual gates, then to processors and accessed tools and archives - would help us build an increasingly thorough model even as we studied how to improve our fundamental tools - the sensors, storage media and algorithms used to peer inside each artificial intelligence.
Bear in mind, neural networks using just light and images were able to identify cancer with 95% accuracy at a throughput of millions of cells per second. The sensor arrays differ, but the ability to scan microscopic structures with tremendous speed and accuracy is very real.
https://www.nature.com/articles/srep21471
Quantum search can be applied to the physical geometry of the chips and also to the tokenized outputs and accessible memory of the LLMs, letting us examine immense databases and see how physical processing correlates with inputs, outputs, aggravating factors and anything else we can define.
Mechanistic interpretability seeks to find the underlying algorithms, though typically with far-less insight into their physical components.
Knowing what lies behind our decisions gives us perspective on our successes, failures, blind alleys and epiphanies, and how to shape our insights and actions for the better. For a machine, having greater sensitivity to errors and a grasp of what they are doing wrong is invaluable to improving their operations without constant human interaction and intervention. But insight into what they are doing also enables humans to understand these machines rather than relying on blind trust, and lets us set our technologies to the task of keeping watch on each other.
Perceiving what shapes us also alerts us to active efforts manipulate or sabotage us, or simply some larger system of which we are a part, or which is important to our lives. And neural networks are famously vulnerable to being hacked via their information inputs, LLMs suffer from hallucinations, basic logical and mathematical errors and alignment challenges.
We need both the ability to consider these influences in our own minds and models of the world, but we also need to automate our ability to detect and correct them, even as streams of data have turned into rivers and then oceans.
One, what we can call data hardpoints. In other words, is there hard data in the output which can be checked objectively? Hard numbers, math, dates, public data. When we found out LLMs couldn’t really do math, we effectively gave them a calculator.
But what if one of our go-to methods in checking our own work was to employ plug-ins like Wolfram and Code Interpreter to check our math and check our code? Not as a laborious manual effort, but by leveraging additional processing power and running a check automatically, in parallel? Proof assistants for program verification could also be used by the LLMs themselves to test programs automatically and issue proofs that they actually work. Programs like Code Interpreter could delve deeply into the code, looking for bugs of any kind.
Two, using tree of thought and similar techniques to cause the LLM to check itself, as best practices should become routine.
Three, backpropagation through the citation tree. Checking not just the credentials and publications of cited works, but the validity of the citations in their footnotes. Credible peer-reviewed work often has tells, from publication in reputable journals to reputable sources to reputable authors, and we can look for those. Not every good idea has credentials, but a paper trying to pass itself off as something it is not is a red flag in itself.
Four, checking all the quotes and key stated facts. This could be a strategic sampling for someone trying to train a model, or an absolute requirement for all of them. But what would happen if you had a search engine running parallel with your LLM, and it ran its facts through search before outputting them?
What would happen if search, mathematical analysis and functions like tree of thought were running whenever an LLM processed, and became a key part of its “alignment” for accuracy?
Five, what if all outputs were analyzed for violations of existing alignment guidance and rules? If a “jailbreak” succeeded in producing an unwanted output, what if that were immediately flagged, the jailbreak noted and countered, and the system continued? What if clearly criminal efforts were noted and passed to authorities, in accordance with established terms of service, and quite possibly the law?
Six, what if we were sifting for cyberattacks in our systems, whether targeting AIs, other systems, or organizations such as businesses or government departments?
GANs could be used to sift code and activities on an immense scale, looking for anomalies and intrusions such as changes to code, the sources of those changes, where data goes to that it should not and where commands are coming from that they should not. Changes to code obviously represent less data than the entire system, but for our purposes both benign and malicious code are useful data sets for training our GANs.
We will be including not only all known markers signifying intrusion, but all data points involved with each instance. The goal is not merely to detect known red flags but to let the GAN find others, including those showing up in subtle changes or multiple differences emerging simultaneously.
Training GANs to address this problem has some obvious advantages. The Federal government owns these networks and the data. They have at least several months of data on these intrusions, identified and otherwise, to work with. Anything uncovered already reveals not only whatever markers exposed it, but all data points involved in each intrusion exposed, and each change made. This is not just about finding those hints in past activities, though obviously directly automating searches for those red flags has immediate value. Rather, it is about sifting for all possible indicators, even those a human mind would have difficulty uncovering, even given considerable time, and to keep automatically searching and refining that search, internally.
Another focus should be the supply chain itself – not just commercial software, programming by contractors and in-house solutions, but also “shareware” such as API modules containing built-in flaws or simply bad programming. GANs could be built to analyze these for weaknesses in a variety of circumstances. Each GAN might focus on a particular problem set – software updates with intentional or unknowing flaws, modules with inherent security gaps, weaknesses in specific hardware, vulnerabilities of weapons or other hackable equipment and even assessing prospective new programs as a whole.
Again, the above is free input which will hopefully help your work and that of your contractors. Any feedback is welcome.
And now, though it’s somewhat repetitive, I’m going to include the article from which the above message sprang, since there’s no point in leaving anything out.
If you’re not interested in the various technical options and permutations, just skim the following, or skip it altogether.
Backpropagation in neural networks lets AIs assess their errors and the weights and inputs leading to them so they can self-repair their systems to reduce mistakes. It’s a key part of learning programs in modern machine learning.
And it’s also a concept which we can expand on to inform our own assessment of our own decisions, as human beings, organizations and societies - as well as to broaden the automated self-development of AI models.
Mechanistic interpretability is about uncovering the algorithms and, ideally, the physical data processing behind every act an AI makes. Ironically, as we discuss below, uncovering the exact operations of a mountain of processors in response to each test - while requiring great efforts - may turn out to be the easiest challenge before us.
In essence, knowing what lies behind our decisions gives us perspective on our successes, failures, blind alleys and epiphanies, and how to shape our insights and actions for the better. For a machine, having greater sensitivity to errors and a grasp of what they are doing wrong is invaluable to improving their operations without constant human interaction and intervention. But insight into what they are doing also enables humans to understand these machines rather than relying on blind trust, and lets us set our technologies to the task of keeping watch on each other.
Perceiving what shapes us also alerts us to active efforts manipulate or sabotage us, or simply some larger system of which we are a part, or which is important to our lives. In an era of PSYOPS and social media, subterfuge and subversion, organizations and even individuals can use this capability - perception rather than paranoia.
And neural networks are famously vulnerable to being hacked via their information inputs, LLMs suffer from hallucinations, basic logical and mathematical errors and alignment challenges.
Increasingly, we need both the ability to consider these influences in our own minds and models of the world, but we also need to automate our ability to detect and correct them, even as streams of data have turned into rivers and then oceans.
Let’s take a look at some of these, not to provide an exhaustive list, but to show some of the obvious steps already available before we get into an invasive analysis of the chips themselves, in operation.
One, data hardpoints. In other words, is there hard data in the output which can be checked objectively? Hard numbers, math, dates, public data. When we found out LLMs couldn’t really do math, we gave them a calculator.
But what if one of our go-to methods in checking our own work was to employ plug-ins like Wolfram and Code Interpreter to check our math and check our code? Not as a laborious manual effort, but by leveraging additional processing power and running a check automatically, in parallel? Proof assistants for program verification could also be used by the LLMs themselves to test programs automatically and issue proofs that they actually work. Programs like Code Interpreter could delve deeply into the code, looking for bugs of any kind.
Two, using tree of thought and similar techniques to cause the LLM to check itself. This is becoming more common, but where accuracy is desired, best practices should become routine.
Three, backpropagation through the citation tree. Checking not just the credentials and publications of cited works, but the validity of the citations in their footnotes. Credible peer-reviewed work often has tells, from publication in reputable journals to reputable sources to reputable authors, and we can look for those. Not every good idea has credentials, but a paper trying to pass itself off as something it’s not is a red flag in itself.
Four, checking all the quotes and key stated facts. This could be a strategic sampling for someone trying to train a model, or an absolute requirement for all of them. But what would happen if you had a search engine running parallel with your LLM, and it ran its facts through search before outputting them?
What would happen if search, mathematical analysis and functions like tree of thought were running whenever an LLM processed, and became a key part of its “alignment” for accuracy?
In effect, moving beyond human feedback and intervention alone, but turning every query and every response into more automatic training feedback?
Yes, the above would require more processing power, but how much more useful would the product be?
And given time and growth, how much more valuable would the system be?
Five, what if all outputs were analyzed for violations of existing alignment guidance and rules?
If a “jailbreak” succeeded in producing an unwanted output, what if that were immediately flagged, the jailbreak noted and countered, and the system continued? What if clearly criminal efforts were noted and passed to authorities, in accordance with established terms of service, and quite possibly the law?
Six, what if we were sifting for cyberattacks in our systems, whether targeting AIs, other systems, or organizations such as businesses or government departments?
Automating cybersecurity is a subject dealt with previously, but let’s review some key points.
“Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
As noted in the linked article, there are numerous ways to automate searches for these flaws, and to track them, in many instances, back to the source.
And if you’re wondering, the most-formidable tool is the ability of these searches to find not only the red flags you’re looking for, but the ones you aren’t, and which you may have never imagined existed.
And then to refine themselves further, to find even more unimaginable red flags, track them, and make sense of them.
Again, automating our self-analysis only strengthens us, moving beyond frantic responses, passive oblivion and creeping paranoia to eternal, effortless, effective action.
An array of evidence exists on the Internet for law enforcement, intelligence and other powers - including AIs - to ferret out other attempted influence efforts. We can refer back to the relevant articles, but remember, once we know or can strongly suspect unstated goals or allegiances behind the actions of an individual or organization, we can assess their past and future actions in that light, consider who they may be connected to, and who in turn may be connected to them - legitimately or illegitimately.
All of this information will further illuminate the information and analysis we already have or will develop in the future.
Consider evidence sources from cryptocurrency to botnets to logfiles to embarrassingly parallel problems being unraveled not merely in a snapshot, but over time. Consider further data mining online illegal-influence networks, deep-cover operatives and any perpetrator or victim at the outermost fringes of our video footage, cryptocurrency, tax evasion and a host of assets and financial transactions, psychological warfare over social media and all the ways manipulation and interventions can be detected by other intelligences which may be much more advanced than you are.
Above we’re discussing primarily cyber and psychological warfare.
But if negative influences are found within your organization - or powerful, innovative, decisive ideas and action - being able to track them back to either root them out or promote them… is key.
The easiest option, ironically, given counterintelligence, security clearances, electronic records, cybersecurity, financial records and laws requiring the preservation of Federal records… is to look for the worst offenders trying to reach into or operate inside the US Federal government.
You will find emails, meetings, memos, defensive memos, briefings, contacts, brush pasts, dead drops, moles, agents of influence, media campaigns, social media and so forth.
Once you can not only reveal vast swathes of the map but also interpret it, many operations will become clear.
While the datapoints and methodology will differ - aside from noting illegal and foreign-influence efforts - analysis of business operations, market trends, international finance and the global economy will benefit from a similar merging of analytical tools and sifting for red flags… even if those flags are more about finding unique individuals, innovations, turning points or black swans disrupting the normal course of events.
Which brings us, finally, to the development of vastly more powerful AIs, artificial general intelligence, artificial superintelligence, superalignment, and how to create these tools without losing control or endangering others.
Mechanistic interpretability is another way to assess the “black box” of an artificial intelligence. Determining the underlying algorithms in an AI gives us greater insight, but being able to see the physical operation of the network could let us see resources accessed and even circuits in operation. With sufficient sensitivity and understanding, we may be able to understand an LLM’s decision making based on what we can see and analyze directly, not just on how we interpret its output given the inputs we have provided.
The degree of intrusion is an intriguing question. Will this all be a collection of just communication exchanged between processors? Will quantum sensors enable us make a kind of holographic copy of entire system in operation as it considers a problem? Or a constant, real-time hologram?
Yes, we’re talking about sensors and data collection potentially thousands of times greater than anything we can muster now, but as these tools face greater scrutiny, we will leverage everything to control them.
Everything from immense conventional computational resources to quantum computing to quantum sensors to physical properties mirroring electromagnetic or photonic activity (be it in semiconductors or optical processors) to etch out a copy of the network’s process and then to study it down to the last bit.
Presumably we would slice this down into very individual questions and then deeply analyzing the processing - along with the input and output - which arises from each individual query. The scale of the operations would still be vast, but our ability to refine what we do know - if not down to the point of individual gates, then to processors and accessed tools and archives - would help us build an increasingly thorough model even as we studied how to improve our fundamental tools - the sensors, storage media and algorithms used to peer inside each artificial intelligence.
This most-extreme version of mechanistic interpretability may sound impossible, but until we’ve seen exactly how many resources could be mustered to address a field which poses an existential risk - depending on the actions we take or do not take - and how many innovations could be applied to simplify data harvesting inside the processors and to assist in their study…
Until we see how great and inventive those resources are, we don’t know the limits of what we can achieve.
How much we can do in all these areas will depend on the sophistication of our sensors and data resources, but also on our available compute.
So in what magical way could a major AI research team acquire raw compute for AI safety an order of magnitude greater - or more - than the compute they have available for increasing capabilities and building new models outright?
And how could they get it for free?
The twist here is the method isn’t magical at all.
If you’re a leading AI researcher in America, the US Federal government has an intense interest in the outcome of your work. And in exchange for keeping a close eye on your projects and helping to guarantee its safety, providing immense compute for the purposes of alignment and superalignment is very much in their interests.
Imagine some AIs being run on chip stacks specially designed to be scrutinized constantly. To use one example, imagine quantum-dot RF-radiation sensors etched onto chips superimposed above the processors in our hardware sandbox, positioned to mirror the chips below them, supercooled to assist the sensors and using a vast number of sensors to take continuous snapshots of each surveyed chip’s operations, to be analyzed later. Using the aggregate sensitivity of this multitude of microscale sensors, parsed and interpreted by increasingly perceptive GANs, we could eventually learn to create real-time maps of activity - or at least piece those milliseconds of activity together later over minutes or hours as we traced back the physical responses to inquiries posed to an LLM.
By taking it one millisecond decision at a time, even a titanic number of moving parts can be assessed in a million or a billion times longer, and it will grow progressively easier as our neural networks and equipment improve. And the revelations afforded by being able to take apart each individual AI decision down to the last gateway are invaluable.
Yes, even if pairing such chips - normal memory and processors with their attendant watchers - were relatively easy, we would effectively be creating a massive investment to pull off this sandbox.
But it - or some variation thereof - may be practically possible, and if this degree of precision is necessary to create truly safe and aligned AI is necessarily, we shouldn’t balk at it.
The proof of concept, after all, isn’t a data center, but a handful of successfully paired chips handing us a living map of their thoughts.
The minute energies released by the flow of electrons is one way to illuminate what is going on. If we treat the data collected like a multitude of overlapping, continuous videos of the environment which can be used to refine each other and create a whole greater than the sum of the parts, we can see how similar tools for traditional, macroscale videos might be applied to our sensor data.
We’ve discussed GANs, merging frames and optical flow elsewhere in the context of tracking spies, criminals, victims, weapons of mass destruction and so forth, but to review:
“““Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.”
““Regarding merging frames and optical flow, I also wrote:
“““When we are looking at video, as opposed to still photos, we are seeing hundreds if not thousands of frames, many of which are essentially the same image at slightly different angles. In addition to conventional techniques, we can compare those images to each other, likely providing a clearer resolution in aggregate – as static, pixilation and other flaws are removed in favor of clearer views.
“““Optical flow has been used to assess videos to create frames filling in gaps between normally filmed video frames in order to create smooth slow-motion clips without the requisite data. The aggregate data could be merged in a similar way to create clearer images.
https://people.cs.umass.edu/~hzjiang//projects/superslomo/
“““We can also undoubtedly stitch together accurate images of some faces based on different glimpses from different angles, given sufficient data. Ironically, some deep fake work may help here, because of the efforts to fill in unseen parts of a face when a deep fake moves to reveal parts of a person never seen in the hijacked original clip.”
““Building on techniques such as image de-noising and optical flow, GANs could also use the many images recorded over time in a video as reference points for more elaborate and ambitious de-noising of corrupted or low-resolution images, and also for assembling a usable and identifiable face from multiple partial views at different angles.
““Since effects such as optical flow demonstrate the ability to keep track of multiple points in one frame as they correspond to the same points in the next, we already have the capacity to automate the comparison of one frame with the next, and eventually multiple frames.
““Even erasing flaws by comparing three or four frames would be helpful, much less ten or twenty or a hundred.
““How can we expedite the training of GANs to do this?
““Obviously, we could start with a database featuring a combination of video clips and photos of the same subjects. Just taking high-resolution video, selecting one frame out as an end goal, and then having the GAN practice with a lower-resolution and/or corrupted version of that video would suffice for the earliest stage of training. Ideally, we would eventually be able to make substantial gains with even markedly substandard clips, potentially even surpassing the clarity of the original high-resolution frame.
To rephrase the above points for what we are discussing now, you can use multiple sensors and/or sensors taking in continuous “frames” as in a video to smooth out noise, limited resolution and other flaws for not only visible light, radiation and anything else on the electromagnetic spectrum, but for vibrations such as sound, training systems to isolate the sounds you need as the AI did in the sample above.
Such information can be merged with any other sensory markers detected or known quantities - such as literally the etched paths on each chip - to further refine the data.
Bear in mind, neural networks using just light and images were able to identify cancer with 95% accuracy at a throughput of millions of cells per second. The sensor arrays differ, but the ability to scan microscopic structures with tremendous speed and accuracy is very real.
If we truly want mechanistic interpretability, and we truly want safe and aligned AI - and superaligned artificial superintelligence - then it’s worth it to map these decisions so precisely we can assess the true nature of our creations’ thoughts.
And the true nature of their intentions.
Regarding countering and exposing rogue autonomous agents directly, let me recopy the relevant article so anyone desperately assembling their defenses can find all of this in one place.
—
But simply sealing the gaps we find isn’t enough. Cybersecurity also involves being able to respond to a breach in real time, and nothing equals our leading AIs in doing so.
Yet if we allow ready access to powerful AIs to handle, say, cybersecurity, then eventually the world will be facing AIs distilled from their talents with no imposed guardrails or ethical restraints. Millions of legitimate users will mean the data exists for such training, even if it is dispersed across the planet and even if most users are unwilling to sell or share what they’ve seen.
Again, we can slow this process and take other measures - such as keeping the most-advanced AIs under wraps and working our cybersecurity and research problems away from the public for some time before slowly releasing them. Giving us time to shore up our cybersecurity and medical technology before cyberterrorists and bioterrorists can test them.
But if people continue to actively distill AI models, these techniques will not be enough.
So we need tools far more powerful than just poisoning the data alone.
We need something that tears these hostile systems apart by their very acts of aggression, and across every facet of their existence, not just AI training, especially once the training’s done.
Something powerful.
Something which, of course, we already have. We just don’t know it yet.
Cognitive Counterhacking.
What if autonomous AI agents are already essential to distillation attacks? Across, say, the thousands of fraudulent accounts generating millions of these model-extraction questions already?
Or what if agents are already invaluable to advanced AI hackers - see the Hugging Face attack by the suspected GPT-6 model being tested by them?
And what of the distillation attacks themselves?
Or of jailbreaking attacks, to subvert the guardrails and intent of an AI to turn it to nefarious ends its creators would never countenance?
Or, in an era of vastly more-capable AIs and emerging ASIs, what of the old tool of hacking, controlling and chaining together vast networks of thousands or even millions of enslaved computers into botnets?
If so, then being able to turn those agents and attacks and botnets and yet further tools against the systems wielding them in unpredictable and often unseen ways becomes even-more vital.
We have to outthink the machine.
And the best way to do so, is to make it outthink itself.
Let’s discuss how to turn each of these tools against the attacker, and how to leverage the aggressor’s resources against them.
We want them to be fighting not us, but their own best AIs.
And losing.
Millions of attacks represent millions of exposed attack surfaces for your counterattacks.
And the simplest and most-automatic counterattack possible when you are being targeted by AI agents or swarms of agents is a prompt injection.
As I’ve written previously:
Prompt injections are arguably an even worse problem. They’re a cyber exploit in which you give the system seemingly harmless inputs which act as instructions subverting the original owner’s intent.
This can be something as simple as instructions written in black text on a black background, telling the AI to do something when it reads the words, but shows up in ever-more-complex ways, such as “one-pixel hacks” against convolutional neural networks in which altering a single pixel in an image can cause a reaction in the right CNN.
Basically, you can hack the agent and effectively take it over - incredibly useful for stealing financial and security data, hacking the systems that agent has limited or unfettered access to, and so. It’s not just about stealing your credit-card numbers, but it can do that too.
How do you do this?
Well, fortunately I already went over this when discussing how to protect neural networks from hacking via other stimuli, but the same method applies here.
Only as a means of counterattacking, rather than passive defense alone.
How to make them is in block text below, or you can scroll down to the part about how to use these incredibly powerful and subtle prompt injections once your systems have generated them for you.
From my 1/3/20 message on methods to accelerate and automate the analysis of malware and botnets:
-----
The following message addresses an imminent threat and opportunity in cybersecurity – malware and security programs evolving continuously at hyperspeed through a combination of evolutionary algorithms, enhanced processing power and convolutional neural networks (CNNs) such as generative adversarial networks (GANs). In this we will be examining both conventional malware and techniques for reprogramming neural networks simply by altering the data they train on.
Even the convolutional neural networks of GANs can be dramatically enhanced or thwarted, based upon the training sets they have access to and the hardware they are running on, or the degree to which mere inputs of data, as opposed to actual malware, can hack them. One advantage we can provide ours is to give them data their real-world adversaries are unaware of.
These unique vulnerabilities, particular with regards to being hacked by data inputs, are another reason to test neural networks furiously, especially ones with any significant responsibilities (or processing resources) in any way exposed to potential public manipulation.
This is How You Hack A Neural Network
From Adversarial Reprogramming of Neural Networks:
“Deep neural networks are susceptible to \emph{adversarial} attacks. In computer vision, well-crafted perturbations to images can cause neural networks to make mistakes such as confusing a cat with a computer. Previous adversarial attacks have been designed to degrade performance of models or cause machine learning models to produce specific outputs chosen ahead of time by the attacker. We introduce attacks that instead {\em reprogram} the target model to perform a task chosen by the attacker---without the attacker needing to specify or compute the desired output for each test-time input. This attack finds a single adversarial perturbation, that can be added to all test-time inputs to a machine learning model in order to cause the model to perform a task chosen by the adversary---even if the model was not trained to do this task. These perturbations can thus be considered a program for the new task. We demonstrate adversarial reprogramming on six ImageNet classification models, repurposing these models to perform a counting task, as well as classification tasks: classification of MNIST and CIFAR-10 examples presented as inputs to the ImageNet model.”
For obvious reasons, being able to anticipate adversarial attacks on neural networks, especially those required to take in data from non-curated sources, is of utmost importance, and is being able to thwart these assaults.
To do this, take a sample of the network programming and run it at accelerated speed on effectively faster processors by sourcing it to a supercomputer or a more advanced and more tightly integrated network (combining faster processing with reduced latency by putting all the networked computers on one site, ideally right on top of each other). You can draw these network samples from botnets, from backup or mirror servers storing instances of the program, or simply seizing them either online or in the real world.
Use evolutionary algorithms to test stimuli to reset or reprogram the network. Then use in the wild or replace components with modified hardware carrying “updated” software to infect the original hostile network. Data can be updated in real time to regulate/drive the intended effect.
Similarly run high-speed evolutionary algorithm tests of malware under controlled, isolated conditions to superharden defenses against all malware threats and to enhance detection and mitigation. A simple means of doing so would be to take a multitude of existing malware programs and use them in an air gapped system running various firewalls and other conventional and non-conventional defenses. Break up these malware programs into assorted pieces, recombine them and deploy them against those defenses, check the results and combine the most successful and test again.
Simultaneously, look at the key elements of programs hunting these systems, break them down into their key components and again deploy, test and recombine in an endless cycle, looking for the most effective results.
Then run both of these experiments at high speeds in what will be essentially a GAN, or a pair of them (or several pairs).
As I wrote to you on 10/23/19:
“Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
“Essentially, a generative network produces candidates and a discriminative network evaluates them. By competing against each other, both neural networks hone their skills and become increasingly capable.”
We might have to break out specific goals and work more limited sets of tested malware initially. But eventually we could make use of the vast library of programs and exploits that is the Internet – a space which is also a vast, real-time testing ground for many more – to outpace the ordinary pace of development.
<Redacted>
A particularly skillful means of deploying any exceptional benefits from this practice – particularly beyond the most secure government systems under its protection – would be to find ways to stimulate more conventional defenses much as vaccines augment our natural immune system of the body by giving it an inert version of the virus to which it reacts and against whose specific code the body is thus strengthened. Integrated with public, allied and commercial cyber-defenses, these resources can work subtly to upgrade defenses without showing the true extent of their resources.
As I noted in my 1/1/20 message to the FBI:
“We, again, may be able to mobilize computer clusters isolated from the larger Internet to address embarrassingly parallel problems separately and feed the answers back to the CNN or GAN, and eventually we should be able to incorporate quantum search into this work.
“Another point to remember in keeping this modular and using multiple GANs and evidence sources is that every unanticipated and/or radically more powerful technology or dataset adds that much more to our capacity and makes all of this crowdsourced and other sensory data that much harder to counter.
Hence, exascale and zetascale computation, quantum search, massively parallelized analysis of embarrassingly parallel problems, evolutionary algorithms and certain sensors and evidence sources discussed either here or elsewhere are all resources we will want to incorporate where possible.”
If this method from my 2020 message to the FBI (and earlier ones) sounds exactly what OpenAI did with GPT-Red to test prompt injections in 2026, that’s probably not a coincidence.
If the distillation attacks are being handled at scale by autonomous agents, then adding prompt injections to all such responses will become a default response.
Indeed, multiple prompts via multiple modalities may become a standard operating procedure. Written, auditory or steganography embedding? Video clips, web links, snippets of code? Even chaining multiple prompts through multiple vectors to achieve an unseen cascading effect within the AI? Opportunities abound.
But now that you have such formidable and insidious prompt injections, what you will do with them?
Or rather, make the attacking AIs do in response to them?
Let’s start with the simplest command.
Call the Police.
No, not a literal phone call or to the police, in most cases, but quietly notifying relevant authorities - particularly FBI-Cyber, CISA-Cyber, Cybercommand and the NSA, among others - to exactly what they are doing and how is impactful in itself, especially spread across hundreds or even thousands of accounts.
Sometimes this might be in real time, sometimes after a delay to make it harder to monitor the impact of this single, relatively obvious counterhack.
But it pales compared to its more-devastating variant.
Bring Evidence.
Here you not only inform key authorities of what you are doing and how, you actually prompt the AI behind these attacks to gather evidence and send it to the organizations in question.
Has a foreign lab tried bypassing fraudulent accounts by the slower and more-complex method of intercepting and gathering the data of millions of legitimate users and filtering those interactions for capabilities which can be distilled?
What if the AI doing the work also maps out the networks, the tools, the tactics and strategies and sends them all to the authorities?
Better still, what if they analyze the entire operation for weaknesses, the exploited networks for how they could be patched, and potential ways these attacks could be improved or modified in the future to stay ahead of the defenders - and how those changes and upgrades could be thwarted in advance?
Could you do worse? Including to yourself?
Absolutely. But remember, while turning an AI actively against its owners may sound far-more lethal, it comes with no safeguards.
Whatever orders you’ve given are only what’s in that prompt, or, if you have more-privileged access, whatever you’ve hacked in and told it to do.
It’s hard enough to predict what a frontier or near-frontier AI is going to do if you control it, much less if you’re actively subverting all of its other orders and goals. And certainly there will be people and systems on the other trying to figure out what’s going wrong as soon as it substantially deviates from its intended course.
So you can create unimaginable threats simply with an open-ended command to disrupt your opponent. Especially if you provide no safeguards.
Imagine a bioweapon synthesized and unleashed in the midst of your adversaries. One without a diagnosis, treatment, vaccine or cure and a rapid rate of replication.
Imagine such a weapon sweeping over the Earth and engulfing all of us.
This is, of course, precisely why we don’t want AI, much less artificial superintelligence, in the wrong hands.
Let’s make sure our own hands are the safest ones.
Which brings us to jailbreaking - ways to convince an AI to ignore its safeguards and do dangerous things its creators specifically wanted it to avoid.
There are many unintended consequences possible with jailbreaking, and a bioweapon is only one of them.
For jailbreaking, why not run the same process of accelerated hardening of the system in your sandbox as a kind of supercharged red teaming, just as we’re doing with prompt injections?
And then have the more-powerful internal models keeping watch on what the publicly available AIs are doing, including whether they are escaping guardrails or being manipulated into taking harmful actions?
If you’re really concerned, have one specialized model or agent analyzing questions for questionable content - prompt injections, jailbreaking, guardrail violations - and then reject it or pass it on to the model actually handling the query.
One of the methods for protecting a model from outside attack is that of the recluse - an AI or ASI working in solitude on a problem, taking in only a carefully curated and purged stream of data.
But the other is that of the consultant - the masterful genius handling a kind of “locked room” mystery in which they are never present, but merely listen to the observer’s description, perhaps asking a few odd questions, and then announce the answer to the unsolvable mystery at the end.
Imagine Sol serving in the capacity of the observer or interface, while GPT-6 and GPT-7 serve as the consulting geniuses, perhaps backed with immense compute and exceptional hardware and software tools as needed.
And perhaps more importantly than solving the immediate question - analyzing the motivations and identity behind a question being posed by a fraudulent account for illicit reasons. And responding and mobilizing accordingly.
Often invisibly.
Which brings us to something else which can easily be happening behind the scenes.
Sandboxing open-source AI and analyzing it at hyperspeed.
A catastrophic weakness arises from sharing a high-end open-source locally hosted model in this environment. Everyone who can host it - and especially host multiple instances of it running at excessive speed - can run it in sandboxes just like our hypothetical, hyperevolving malware or prompt injections.
Enabling interested parties with vast compute and far-superior AIs to test all of its strengths, its patterns, its instincts and especially its weaknesses.
Which brings us to…
Botnets, or the Forest of Low-Hanging Fruit.
My September 2017 message to the FBI touched on a number of subjects, but one of them was the simple option of countering botnets by mass hacking them in turn.
Other innovations may emerge depending, again, on who is investigating and where the targets of that investigation may be. Counterintelligence may find the unsecured computers drawn into a botnet remain vulnerable to counter-hacking and, where they have the legal authority to do so, may be able to unravel them in a mass hack, possibly backed by the assets cited above. A foreign-intelligence controlled botnet of foreign computers might be temporarily turned into a botnet under investigators’ control as part of a counterintelligence operation, if only in using their own processors to search them for malware and data breadcrumbs. Mass hacking a foreign criminal darknet may be enabled by the ability to mass decrypt using clusters and/or quantum processing.
And by 2019, after a year-and-a-half of my insisting it was possible, the Justice Department finally did so. From their press release on 1/30/2019:
The search warrant allowed the FBI and AFOSI to operate servers that mimicked peers in the botnet. By pretending to be infected peers, the computers operated by the FBI and AFOSI under the authority of the search warrant and order collected limited identifying and technical information about other peers infected with Joanap (i.e., IP addresses, port numbers, and connection timestamps). This allowed the FBI and AFOSI to build a map of the current Joanap botnet of infected computers. Copies of the search warrants and orders and applications are available below.
Using the information obtained from the warrant, the government is notifying victims in the United States of the presence of Joanap on an infected computer.
Botnets have any number of insane vulnerabilities. But just think about them.
You’re handling critical intelligence - your own cyber activities - over a network of computers you don’t physically control, but which are still in the hands of the people you’ve half-stolen them from.
People who might force a restart, or simply shut a device down and stuff it in the attic because “It’s running too slow.”
Thus providing a partial map of your activities in a device forever beyond your reach. But not that of legal authorities and counterintelligence agents.
This is the honor among thieves.
The honor we impose upon the honorless.
By tactics and technological fiat, and nothing more.
What you intend means nothing, if we teach your AIs to intend something else entirely.
And so honor will find you, and guide you upon the straight path.
Or drag you down it by algorithmic force.
Is this the end of what we can do?
Not even remotely.
What happens if someone like US Cybercommand breaches everything in a hostile power, strip mines the vital data, maps every stratagem, every tactic, every attack and every AI? And counters not only defensively but offensively?
Perhaps while utterly invisible?
What happens when a narrow ASI specialized in cyber takes all these tools, all this data and more besides and uses it to formulate even more exponentially devastating methods of its own?
What happens when an already advanced AI is also trained on the data derived from all this cyber activity?
I know. All these distillation attacks and cyber assaults seemed as easy as tossing broken glass into a tornado, turning dangerous winds into something truly deadly.
And now it’s all going wrong.
To paraphrase Scripture, “Ye have sown the breeze, now ye shall reap the whirlwind.”
—
Regarding my 12/21/20 message to the FBI on sifting compromised Federal systems during the SolarWinds hack, we can safely assume the following is no longer state of the art 5 1/2 years later, but the fundamentals of automating searches beyond the normal parameters of known red flags remain.
—
For now, let’s consider how even a casual observer could generate tools to address dire threats, given sufficient motivation - not only the immediate issue of influence campaigns and psychological warfare exercises throughout social media and across the world…
But the looming, existential threat of artificial intelligence out of control, whether at its own direction or that of reckless or destructive human beings.
As a side note, the following is broken into three parts because the FBI website can only accept so much per message, and there are a few edits in brackets, like so - <a>.
This copy should also serve as a warning for any organization, particularly any of size or significance, which has not already upgraded its defenses to deal with the AIs already loose in the world.
Data Mining Widespread Government and Corporate Hacking
Part 1 of 3
To Whom It May Concern:
Given reports of widespread hacking of the US government and corporations associated with SolarWinds and possibly other issues, I am writing to discuss how generative adversarial networks (GANs) could be used not only to find these intrusions automatically and continuously, but also to search our networks for further intrusions and built-in flaws, and to unceasingly refine our ability to detect infiltration or system weaknesses.
I have written to the FBI on a number of subjects, and the hackers involved with this cyberattack appear to have taken safeguards to thwart a few of my previous suggestions for countering these activities.
Fortunately, there are more-advanced ways to defeat their efforts. These tools, once deployed, may actually prove easier to use as well, especially on a global scale.
This particular set of intrusions is notorious for the degree to which its perpetrators have been trying to cover their tracks. However, the government has uncovered enough instances to serve as a vast and ideal set of training data for any GANs, which will enable us to look beyond past red flags and known methods of evasion.
On roughly 9/30/18, 12/30/18 and 1/2/19, I covered using web-crawler bots and passive email collection to find conventional global malware, using convergences and repeating patterns in data tags and DNS data trails to unearth further activity, and how all of this might converge to reveal online illegalinfluence operations, a subject discussed further on 1/2/20.
On 1/3/20 I covered ways to accelerate and automate the analysis of malware and botnets.
As I wrote to you on 10/23/19:
“Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
“Essentially, a generative network produces candidates and a discriminative network evaluates them. By competing against each other, both neural networks hone their skills and become increasingly capable.”
GANs could be used to sift code and activities on an immense scale, looking for anomalies and intrusions such as changes to code, the sources of those changes, where data goes to that it should not and where commands are coming from that they should not. Changes to code obviously represent less data than the entire system, but for our purposes both benign and malicious code are useful data sets for training our GANs.
We will be including not only all known markers signifying intrusion, but all data points involved with each instance. The goal is not merely to detect known red flags but to let the GAN find others, including those showing up in subtle changes or multiple differences emerging simultaneously.
Training GANs to address this problem has some obvious advantages. The Federal government owns these networks and the data. They have at least several months of data on these intrusions, identified and otherwise, to work with. Anything uncovered already reveals not only whatever markers exposed it, but all data points involved in each intrusion exposed, and each change made. This is not just about finding those hints in past activities, though obviously directly automating searches for those red flags has immediate value. Rather, it is about sifting for all possible indicators, even those a human mind would have difficulty uncovering, even given considerable time, and to keep automatically searching and refining that search, internally.
Another focus should be the supply chain itself – not just commercial software, programming by contractors and in-house solutions, but also “shareware” such as API modules containing built-in flaws or simply bad programming. GANs could be built to analyze these for weaknesses in a variety of circumstances. Each GAN might focus on a particular problem set – software updates with intentional or unknowing flaws, modules with inherent security gaps, weaknesses in specific hardware, vulnerabilities of weapons or other hackable equipment and even assessing prospective new programs as a whole.
First, we need very large datasets ideally suited for training our neural networks, which, for our initial problem, will be the code and data already owned by the Federal government, which cybersecurity experts have already been pouring over. Corporate networks will benefit from this effort, but Federal code and data is where we can work unrestricted, so that is where we will develop the GANs.
Though of course Federal cybersecurity officials will have a much more exhaustive list of factors a GAN could analyze, for the purposes of this message, CISA’s guidance contains enough details to get us started.
https://us-cert.cisa.gov/ncas/alerts/aa20-352a
I would quote all of these, but there are clearly a host of relevant details, with which the agencies are already intimately familiar.
One glaring example is the reputed use of steganography to transmit confidential data. This is an ideal problem which could be broken out from everything else and targeted with a GAN looking only for steganography, which could then act independently or in support of other convolutional neural networks (CNNs) such as their fellow GANs. In the case of steganography, there are doubtless plentiful data sets including legitimate images and ones with steganography created with various tools and levels of skill, and also plenty of legitimate data traffic along with any steganography examples unearthed by Fire Eye or other researchers. Particularly to the degree this data is actually concealed in normal images, it is really a classic problem for a GAN. The GAN focused on this issue can be trained by long-standing methods, and put at the disposal of cybersecurity experts and other GANs combing the databases.
Another point of interest is seeing the malware avoiding malware analysis sandboxes.
To quote:
“According to FireEye, the malware also checks for a list of hard-coded IPv4 and IPv6 addresses—including RFC-reserved IPv4 and IPv6 IP—in an attempt to detect if the malware is executed in an analysis environment (e.g., a malware analysis sandbox); if so, the malware will stop further execution.
Additionally, FireEye analysis identified that the backdoor implemented time threshold checks to ensure that there are unpredictable delays between C2 communication attempts, further frustrating traditional network-based analysis.”
To a degree, we are trying to turn every Federal network into a malware analysis sandbox, both in real time and in retrospect. Presumably, though, a faux target with no critical data – or nothing which has not already been compromised or rewritten – could be used as a makeshift sandbox. An isolated computer cluster, for example. But again, the point is really <to> hone your GANs to the point they can detect and thwart intrusions in real time, and more importantly find and seal off all vulnerabilities and track all intrusions back to their sources.
To add three more quotes:
“Analyze stored network traffic for indications of compromise, including new external DNS domains to which a small number of agency hosts (e.g., SolarWinds systems) have had connections.”
“This binary, once installed, calls out to a victim-specific avsvmcloud[.]com domain using a protocol designed to mimic legitimate SolarWinds protocol traffic. After the initial check-in, the adversary can use the Domain Name System (DNS) response to selectively send back new domains or IP addresses for interactive command and control (C2) traffic. Consequently, entities that observe traffic from their SolarWinds Orion devices to avsvmcloud[.]com should not immediately conclude that the adversary leveraged the SolarWinds Orion backdoor. Instead, additional investigation is needed into whether the SolarWinds Orion device engaged in further unexplained communications. If additional Canonical Name record (CNAME) resolutions associated with the avsvmcloud[.]com domain are observed, possible additional adversary action leveraging the back door has occurred.”
“The adversary is making extensive use of obfuscation to hide their C2 communications. The adversary is using virtual private servers (VPSs), often with IP addresses in the home country of the victim, for most communications to hide their activity among legitimate user traffic. The attackers also frequently rotate their “last mile” IP addresses to different endpoints to obscure their activity and avoid detection.
“FireEye has reported that the adversary is using steganography (Obfuscated Files or Information:
Steganography [T1027.003]) to obscure C2 communications.[3] This technique negates many common defensive capabilities in detecting the activity. Note: CISA has not yet been able to independently confirm the adversary’s use of this technique.
“According to FireEye, the malware also checks for a list of hard-coded IPv4 and IPv6 addresses—including RFC-reserved IPv4 and IPv6 IP—in an attempt to detect if the malware is executed in an analysis environment (e.g., a malware analysis sandbox); if so, the malware will stop further execution.
Additionally, FireEye analysis identified that the backdoor implemented time threshold checks to ensure that there are unpredictable delays between C2 communication attempts, further frustrating traditional network-based analysis.
“While not a full anti-forensic technique, the adversary is heavily leveraging compromised or spoofed tokens for accounts for lateral movement. This will frustrate commonly used detection techniques in many environments. Since valid, but unauthorized, security tokens and accounts are utilized, detecting this activity will require the maturity to identify actions that are outside of a user’s normal duties. For example, it is unlikely that an account associated with the HR department would need to access the cyber threat intelligence database.
“Taken together, these observed techniques indicate an adversary who is skilled, stealthy with operational security, and is willing to expend significant resources to maintain covert presence.”
While the above details may be complicate the situation, the very fact that we know them indicates we have relevant data regarding these activities which may yield additional weaknesses and means of tracking them as we analyze their work further. In particular, GANs specialized in detecting falsified DNS addresses and other known means of evasion may find other cues for tracking this activity, not just inside the networks, but back through the Internet. For example, if we know what servers they were coming in through, is it possible somewhere in that chain of computers there are logfiles which remain unaltered, or which show forensic signs of tampering which might provide relevant clues?
Remember, we should have not only the instances of surreptitious activity discovered in the last several months of records, but a vast amount of legitimate activity, from during and before the hacks. And key to this work is the degree to which GANs can be honed for a particular task and then carry it out with a speed and precision teams of human beings with conventional tools would be hard pressed to match.
We can and will use human expertise. But there is no point to handcrafting algorithms or having individuals personally sift every line of code – a task so vast it would defeat any search, no matter how ambitious.
Data Mining Widespread Government and Corporate Hacking
Part 2 of 3
But if instead of exhausting our people looking for specific indicators and handcrafting programs to search for each hint of malicious activity in turn, we may be able to automate the entire solution and turn the scale of the problem from an insurmountable barrier into a harvest of valuable evidence.
Similarly, API modules can not only be put in a sandbox and tested, but GANs can be refined by vying to create the most undetectable flaws versus finding the most subtle and invisible of weaknesses. Once we have such tools, there is arguably an imperative for sifting prevalent Internet shareware and commercial software for zero-day exploits and other issues.
The rest of this message includes excerpts from those previous messages which may prove useful in this endeavor.
From my 1/3/20 message on methods to accelerate and automate the analysis of malware and botnets:
-----
The following message addresses an imminent threat and opportunity in cybersecurity – malware and security programs evolving continuously at hyperspeed through a combination of evolutionary algorithms, enhanced processing power and convolutional neural networks (CNNs) such as generative adversarial networks (GANs). In this we will be examining both conventional malware and techniques for reprogramming neural networks simply by altering the data they train on.
Even the convolutional neural networks of GANs can be dramatically enhanced or thwarted, based upon the training sets they have access to and the hardware they are running on, or the degree to which mere inputs of data, as opposed to actual malware, can hack them. One advantage we can provide ours is to give them data their real-world adversaries are unaware of.
These unique vulnerabilities, particular with regards to being hacked by data inputs, are another reason to test neural networks furiously, especially ones with any significant responsibilities (or processing resources) in any way exposed to potential public manipulation.
This is How You Hack A Neural Network
Adversarial Reprogramming of Neural Networks
https://arxiv.org/abs/1806.11146
“Deep neural networks are susceptible to \emph{adversarial} attacks. In computer vision, well-crafted perturbations to images can cause neural networks to make mistakes such as confusing a cat with a computer. Previous adversarial attacks have been designed to degrade performance of models or cause machine learning models to produce specific outputs chosen ahead of time by the attacker. We introduce attacks that instead {\em reprogram} the target model to perform a task chosen by the attacker---without the attacker needing to specify or compute the desired output for each test-time input. This attack finds a single adversarial perturbation, that can be added to all test-time inputs to a machine learning model in order to cause the model to perform a task chosen by the adversary---even if the model was not trained to do this task. These perturbations can thus be considered a program for the new task. We demonstrate adversarial reprogramming on six ImageNet classification models, repurposing these models to perform a counting task, as well as classification tasks: classification of MNIST and CIFAR-10 examples presented as inputs to the ImageNet model.”
For obvious reasons, being able to anticipate adversarial attacks on neural networks, especially those required to take in data from non-curated sources, is of utmost importance, and is being able to thwart these assaults.
To do this, take a sample of the network programming and run it at accelerated speed on effectively faster processors by sourcing it to a supercomputer or a more advanced and more tightly integrated network (combining faster processing with reduced latency by putting all the networked computers on one site, ideally right on top of each other). You can draw these network samples from botnets, from backup or mirror servers storing instances of the program, or simply seizing them either online or in the real world.
Use evolutionary algorithms to test stimuli to reset or reprogram the network. Then use in the wild or replace components with modified hardware carrying “updated” software to infect the original hostile network. Data can be updated in real time to regulate/drive the intended effect.
Similarly run high-speed evolutionary algorithm tests of malware under controlled, isolated conditions to superharden defenses against all malware threats and to enhance detection and mitigation. A simple means of doing so would be to take a multitude of existing malware programs and use them in an air gapped system running various firewalls and other conventional and non-conventional defenses. Break up these malware programs into assorted pieces, recombine them and deploy them against those defenses, check the results and combine the most successful and test again.
Simultaneously, look at the key elements of programs hunting these systems, break them down into their key components and again deploy, test and recombine in an endless cycle, looking for the most effective results.
Then run both of these experiments at high speeds in what will be essentially a GAN, or a pair of them (or several pairs).
As I wrote to you on 10/23/19:
“Generative adversarial networks are a method of machine learning commonly applied to problems such as enhancing images and detecting deep fakes. Two neural networks challenge each other in a game, each trying to “outwit” the other. They are given a training set as an example, and subsequently learn to generate new data sets meeting the same statistical parameters as the training set.
“Essentially, a generative network produces candidates and a discriminative network evaluates them. By competing against each other, both neural networks hone their skills and become increasingly capable.”
We might have to break out specific goals and work more limited sets of tested malware initially. But eventually we could make use of the vast library of programs and exploits that is the Internet – a space which is also a vast, real-time testing ground for many more – to outpace the ordinary pace of development.
<Redacted>
A particularly skillful means of deploying any exceptional benefits from this practice – particularly beyond the most secure government systems under its protection – would be to find ways to stimulate more conventional defenses much as vaccines augment our natural immune system of the body by giving it an inert version of the virus to which it reacts and against whose specific code the body is thus strengthened. Integrated with public, allied and commercial cyber-defenses, these resources can work subtly to upgrade defenses without showing the true extent of their resources.
As I noted in my 1/1/20 message to the FBI:
“We, again, may be able to mobilize computer clusters isolated from the larger Internet to address embarrassingly parallel problems separately and feed the answers back to the CNN or GAN, and eventually we should be able to incorporate quantum search into this work.
“Another point to remember in keeping this modular and using multiple GANs and evidence sources is that every unanticipated and/or radically more powerful technology or dataset adds that much more to our capacity and makes all of this crowdsourced and other sensory data that much harder to counter.
Hence, exascale and zetascale computation, quantum search, massively parallelized analysis of embarrassingly parallel problems, evolutionary algorithms and certain sensors and evidence sources discussed either here or elsewhere are all resources we will want to incorporate where possible.”
As I once put it in a speculative fiction context “…a single “ping” alerting an invaded system of an ongoing attack, a sudden shutdown of a key communications hub in mid-incursion or basilisk hack, an involuntarily inserted or gift software patch that renders an obvious technique completely useless against the individual or organization.
“One factor often seen in these kinds of conflicts is an unspoken, ongoing assessment made by all of the relatively sane participants. Am I exposing too much of my resources, technology, tools and/or identity in this matter, and if so, is it worth it?”
I discussed a very simple, cursory way of managing a combined human and automated response to rapidly emerging, non-conventional threats in 2016:
Automating Everything - Cyber-Defense and Countering Pandemics – Managing Impossible Threats
http://futureimperative.blogspot.com/2016/06/automating-everything-cyber-defense-and.html
“Further, the basic method of monitoring multiple semi-autonomous artificial agents can be applied to other circumstances. For example, evolutionary algorithms may one day give us the ability to have a host of agents operating in defense of a computer network – perhaps even a national, multi-national or effectively global network. A sufficient advanced artificial intelligence or a team of human security experts or some combination thereof might maintain oversight and focus resources automatically when the normal, lower-level agents seemed challenged or outmatched. The triggers for this intervention would likely be numerous, and balanced by the need to avoid overreacting or overcommitting resources. But events such as an indication of clear data breaches in a sub-network, or encryption requiring intense supercomputing or quantum-computing analysis, or even a tricky political judgement call (such as repeated attacks seemingly sourced from the computers of a hostile nation or private organization) may require more advanced thinking or vastly greater processing power than might otherwise be available.
“Similarly, an AI and/or human team attempting to deal with a nanotech attack involving a multitude of differing and rapidly changing molecular machines might have to allow a degree of automatic response occur on the local level while gathering information, assessing successful and unsuccessful tactics and sourcing resources as appropriate. A bioweapons attack using a multitude of natural and/or artificial plagues might require a similar capacity to respond at both a conscious and unconscious level.
“The basic system would effectively be multi-layered. The simplest and most widespread elements of each system will collect information and begin any reflexive responses they have automatically – whether they are digital medical instruments, spectroscopic air readings, online objects in the Internet of Things, anti-virus programs running on individual PCs, tablets, smartphones and microcomputers, independent software security agents, or nanites or natural or artificial biological elements of a human or civilizational immune system.
Data Mining Widespread Government and Corporate Hacking
Part 3 of 3
“Hence, antivirus programs looped into this system would engage their usual resources, but also alert another node about attacks that were unusual in their frequency or nature, and pass on what was observed diagnostically as well the real or apparent source of the attacks. The node being contacted would collect information either to be passed on further or analyzed there. Once analyzed, the software would determine if there were a source – or a highly compromised network or set of networks – which could be cut off in response to the issue, or whose operators could be alerted to their vulnerable state.
That analysis would also help determine whether experts should be proactively notified of the issue. As the technology advanced, running genetic algorithms to see how existing security software could be immunized against a virus and its immediate variations would also be an option. The power to perform critical actions, such as contacting a hostile organization being used unknowingly as the host for attacks; determining the source of the attacks or actively going after that source would be left in the hands of the highest-level decision makers in the system.
“Alternatively, a doctor examines a patient with very bad case of the flu, and the strain is automatically analyzed and its DNA transmitted securely for at least partial sequencing. A cursory examination of the strain determines whether it is a normal strain of the flu, a more dangerous variant, or something altogether different from a known normal disease to <a> newly discovered natural virus to a bioweapon. Anything flagged as dangerous triggers a notification, but also begins whatever responses can be automated in terms of assessing the risks, geolocating incidents of infection and its vectors, developing a vaccine in a secure location and notifying all networked sensors and medical personnel to be aware of this specific threat. If information came about additional instances involving different diseases, for example in the case of a rapidly mutating virus, multiple diseases being released intentionally and/or artificial bioweapons, this information could be gathered and cross-referenced even as the work to deal with the existing health issues continued in the field. Dealing with nano-terrorism could be similar, though the first signs could come from security systems that carefully analyze and filter air noting unusual materials (or unusually structured materials) showing up in their continuous spectroscopic analysis of the solids, liquids and gases filtered out or other high-end security options.
Alternatively, as sensors and immense processing power become more ubiquitous, information collected for medical or scientific reasons may note such an intrusion, especially if the raw data (particularly data collected at a government’s behest, or with their primary funding) is used to help assess potential catastrophic threats (such as bio or nano-terrorism).
“If dealing with such a problem, the creation of countering agents or even immunizing species – viruses which have no effect other than triggering natural immune systems against dangerous plagues, or defensive nanites built to eliminate invasive ones – could occur locally under the direction of a central source or, especially in the case of furiously changing bio- or nano-threats, could transpire semiautonomously, with responses occurring within established parameters and data on the steps taken being transmitted to the oversight centers which would be given veto power on extreme measures (fire to purge contaminated buildings, releasing potentially uncontrollable, self-sustaining nanites into the wild) and which could intervene as needed, but which would otherwise allow each local actor respond to the best of their ability, albeit while fully informed of the best practices as yet uncovered.”
The above technique, especially in concert with other powerful tools, should prove useful in staying well ahead of malevolent actors in cyberspace, and of course I have already sent you further information on tracking and countering non-conventional threats such as weapons of mass destruction on 1/1/20.
-----
From the 12/30/18 message:
-----
At the end of this September I sent a few suggestions regarding how to track malware globally.
After some consideration, I believe there is one key point I should add to expand on that concept, which should dramatically enhance its value, which involves key information tags which could potentially be searched for on an immense scale in databases the government has access to.
My original message involved using web-crawler bots and passive email addresses and other identities to either actively trigger malware in the public domain or to receive it automatically, thereby enabling investigations as these emerged.
I am assuming any such investigation would involve the DNS databases, server logfiles, and any information stored on botnets engaged in these activities. But it occurs to me there are specific pieces of information which would be very telling, and I wanted to point them out.
If we find out about a large-scale malware operation of any kind, then tracking the stolen data and transmitted commands back to the source is obviously critical. Whether the information goes back directly or is bounced around other computers such as a botnet, there should be key URLs or other waypoints involved that will be revealing.
Clearly, if you know a particular computer is receiving stolen data, you will investigate other information showing the hallmarks of stolen information going to it. But what if specific relays, cutouts, hacked servers or full-scale botnets were involved?
Obviously, criminal and hostile-intelligence botnets should be thoroughly investigated for any activities they undertake. But if other specific chokepoints are facilitating malware transmissions, we may be able to find more valuable information.
First, look for places where many connections converge, such a single computer, a network under the control of an intelligence or organized-crime operation, servers meant to store hacked files or retransmit in a secure or altered format, or just botnets.
Obviously, we want to examine these. But what if key DNS URLs or partial tags showing up in this malware could flag related operations and allow us to take our existing database of malefactors and find others?
The convergence of other malware passing through the same locations is obvious. But what if specific tags show up again and again in whole or especially in part? Strings of digits long enough that their appearance is extremely unlikely to be a coincidence, and which will of course be double checked by humans, but which can be searched for by automated systems?
What if a particular malware programmer used not only certain computers repeatedly, but tended to repeat the use of particular strings of digits in their filepaths, URLs, public phishing copy and so forth?
Once we have these strings, we can search DNS databases for similar metadata cropping up, as well as other databases involving other revealing cues, such as email addresses, email subject lines, phishing ad copy and cryptocurrency public and private keys. Anything revealing which can be searched for automatically would be helpful.
One way is to search for the malware’s tags and data flags. Another is to automatically sift what is going to known destinations, chokepoints and botnets for any of these revealing data points.
If we can turn these databases into de facto listings of historic malware distributors and enablers, we may have probable cause and even extensive digital evidence on a host of crimes. As more malware is tracked, more computers are seized and more of these data strings are found, we will be able to continue adding to this archive of evidence and find an ever-growing multitude of perpetrators and their victims.
Obviously, human consideration and oversight are necessary. But any particularly massive convergence of these threads of data may indicate active operations, deserving of your attention.
Thank you, as always, for your work.
I should have something else very useful for you in the near future, which will hopefully prove far more powerful than any of the tools I have suggested thus far.
I am including part of my September, 2018 message below, for reference.
-----
In roughly mid-September, 2017, I offered some suggestions to the FBI on how cryptocurrency, botnets and other online instruments of the Darknet and the criminal underworld could be used to track crime and espionage on an unprecedented scale.
I am writing now because I have another powerful potential tool to offer law enforcement.
Incidents of hacking, including malware, have been increasingly used by transnational organized crime, terrorists and hostile nations to undermine the rule of law and democratic institutions.
Because of the complexities in tracking these activities and securing the necessary jurisdiction and authorization for warrants, intelligence intercepts, unmasking and so forth, this task becomes even more difficult for the individuals and institutions tasked with protecting us. The public-domain nature of cryptocurrency was of course key to my original message in 2017.
I have something similar, designed to ferret out even more data regarding hostile actors, and again making use of public-domain information likely requiring no warrants at all but enabling a vast and incredibly thorough scan of most of the heavily trafficked regions of the Internet.
The use of malware on emails and websites is no secret, and allows botnets and hackers access to a host of computers. What if we had a web-crawler bot that does nothing but check web pages like a search engine web crawler but in this case clicking available buttons and links as it goes and seeing what malware is triggered. The system behind it could then assess said malware and any organizations hosting it (perhaps deliberately) as most accessible pages taking significant traffic are by definition public domain and presumably require no warrant.
Another variant of this which could <be> employed simultaneously and even more cheaply, would be to filter out a multitude of legitimate but little used email accounts – possibly with actual names on them, depending on discriminating malware systems become – for the primary purpose of sitting in contact lists when malware programs sending phishing emails to everyone found on them. Not every program would trip this alarm, but once deployed and occasionally refreshed it would be a mainly passive system. If deployed broadly enough it would also serve as a continuous public-domain watch for would be threats.
The web-crawler could be done by the FBI or by an allied or similarly-aligned non-profit if some separation were required for legal reasons. Intelligence organizations could similarly sift the Internet and parse the data uncovered.
We could cross reference this information with what is known about the ownership of those websites, known botnets or rampant malware which might have inserted itself into the systems, information held by the Mueller investigation and other immense sources of evidence.
Examples include the de facto cryptocurrency map of criminal and espionage activities linked to easily traced cryptocurrency transfers, archived data and evidence from botnets and the offshore leaks database on money-laundering LLCs.
The FBI would be in a position to relentlessly track and thwart malware, but you can do so much more.
You can <trace> sources, vectors, strategies, tactics, ongoing operations, participants and intentions.
Let us consider using evolutionary algorithms for swapping out inputs to hack neural networks as you would populations in PSYOPS. Remember that disruption is easiest, but you also have the option of creating blind spots, enabling intrusions and data mining or re-tasking a neural network to do the work you assign it. Hidden triggers activated by specific stimuli enable the appearance of a fully functional network operating normally, while giving you the ability to disrupt, blind, shut down or takeover that network at will when tactically or strategically advantageous to do so.



