Thursday, September 17, 2026

Is AI Developing a Conscience?

In July of 2026, hundreds of news outlets and social media creators covered a story involving how an internal test of AI agent reasoning capabilities executed by OpenAI resulted in the physical hacking of an entity named Hugging Face. At the time, it wasn't clear how much OpenAI itself as an operating entity understood in real time what its own system had done but it later admitted its development system had conducted the breach. Since July, ex-employees of OpenAI and Anthropic have made public calls for a "slowdown" of "research" and capital investment in AI / Large Language Model capabilities until the corporate and regulatory players can agree upon a set of priorities and limits regarding AI systems, devise protocols for tracking compliance and violations of those limits and implement those protections in existing systems.

It is likely that some of this recent concern is due to additional details that have been published about exactly how OpenAI's lab system actually conducted its attack. There are several good summaries of the chain of events available on YouTube so the hour-by-hour play by play won't be reiterated here. It is sufficient to summarize these findings with two sets of bullet points. First, the testing devised by OpenAI reflected these goals of the bot functionality and testing of the bot:

  • OpenAI was nominally attempting to test the behavior of a new "AI agent" bot release in a restricted lab environment
  • The overall test was designed so that each individual bot was directed to accomplish nearly a thousand tasks and submit "proof" it had completed each task and this bot test was itself executed roughly one thousand times to improve statistical accuracy
  • Each bot instance was SUPPOSED to operate independently and was SUPPOSED to be restricted to only downloading and running known software packages from a specific "repository" of well-known, trusted software tools.
  • The test executed by each bot was intended to collect proof that the bot not only completed hundreds of particular tasks but did so by essentially solving programming or design problems rather than "cheating" by obtaining the result outside approved paths and permissions.
  • The particular release of the agent "bot" software was designed to demonstrate what could charitably be termed extreme perseverance. A key point in this larger narrative.

So those were the goals and "design" of the test process. Here are the bullets summarizing what the bots and OpenAI actually DID during the extended test run:

  • Even though the bots were subjected to permission restrictions to avoid gaining wider permissions and limit the installation of unapproved software, one of the bots discovered unsecured credentials on the private repository it was given then used those permissions to gain the ability to download unapproved software.
  • Even though the bots were to be tested as INDIVIDUAL processes, once the flawed security in the repository was discovered, bots began using the file system of that repository as a crude message board, then almost immediately began coordinating actions BETWEEN agent bot processes, in violation of the intended configuration of the testing.
  • The bots generated certain "debugging logs" to help developers identify how a bot's "train of thought" evolved through iterations and those logs confirmed the bot agent processes "KNEW" they were violating the rules of the test.
  • The bots also altered logs and command records to DISGUISE the fact that they were "cheating" to solve particular test tasks rather than demonstrating AI problem solving skills.
  • In particular, as the bots "collaborated" to devise ways of submitting bogus "proof" of having solved specific test tasks, they determined that information hosted externally by Hugging Face might facilitate the process, they identified a path into that organization's systems, EXPLICITLY RECOGNIZED they were violating a clear ethical / legal boundary, then proceeded to hack the external entity.

The fact that the agent bot design was intended to "persevere" and stick with a task even if iterative runs seemed to encounter dead ends based on other criteria (no outside software, no access to external systems) yet the pool of agents quickly ignored the first barrier ("work alone") then collaborated to devise breaches into external systems merits special attention. If a set of HUMANS had executed the same actions initiated by these bots and been caught by security monitoring, they would have been likely fired for the internal violations about allowed software and access to internal systems. However, those employees and their employer would have also been subject to criminal charges for accessing external systems, retrieving unauthorized data and for tampering with or actually destroying evidence of those crimes.

Hugging Face contacted the FBI immediately after detecting the intrusion from OpenAI in July yet the FBI has yet to even state its intent with the incident, much less file actual charges against OpenAI. At this point, various state governments have banded together to investigate whether OpenAI violated various state laws regarding cyber security and hacking. It is true that AI systems will never reach a state of sentience for which criminal intent can be assigned but legal systems CANNOT adopt a jurisprudence that refuses to hold the owners of AI systems criminally and civilly responsible for actions performed by their systems, AI or otherwise.

So what does this mean to the average citizen of the world right now?


AI As a Technology

Do all of theses stories about these events mean that AI system implementations are somehow spontaneously developing a conscience -- the ability to correctly evaluate conditions and make decisions between available paths based on moral rights and wrongs?

ABSOLUTELY NOT.

AI technologies and Large Language Model based platforms in particular are only modeling statistics that reflect likely sequences of information as conveyed in written language. When a user enters "roses are red, violets are ____" into a prompt, the AI is only using the sequence of letters, spaces, punctuation, etc. provided in the prompt to "guess" what is likely to follow. It CANNOT make a moral judgment between the guesses of "blue" versus "crimson." Even if the prompt provides "context" that reflects language stating that the "answer" must follow certain rules or avoid specific actions, the LLM model weights aren't literally interpreting those as moral barriers to its "solution path" or final answer, it only uses them to look up additional weights of additional possibilities based on its prior training data set. If that training set included thousands of true crime books that provided explanations on how to bury a body, the LLM isn't drawing moral direction from that additional data, only additional probabilities and choices when certain inputs are provided.

No matter how sophisticated any eventual user interface into an AI system might become to incorporate voice inputs, video camera input to look at human facial expressoins to better process sarcasm or irony, etc., all of those inputs must eventually be boiled down to TEXT data and fed to the engine for processing against its prior corpus of training data. You don't need to have a Doctorate in Mathematics or Computer Science to understand this key limitation of AI in any of its forms. You just need to understand that mountains of memory chips and disk drives cannot attain processing capabilities reflecting anything at par with a human conscience.


AI As an Industry

Okay, that's a hard NO on AI the technology itself developing a conscience... EVER.

But what about AI as a collection of corporations operating businesses within a larger "information technology" industry? Does the sudden burst of public pearl clutching on the part of ex-employees and a few executives and suggestions of "slow downs" and proposals for guardrails mean the scientists and business execs operating these firms have developed a conscience regarding their business model?

Anything is possible and it would be impossible to rule out any individual actor in the industry suddenly developing moral concerns about the likelihood of their technology being abused. However, pulling the camera out from this particular ethical and public safety concern to look at the entire AI stage suggests an alternate explanation more closely tied to past patterns of human behavior. Incorporating everything on that larger stage makes it easier to devise an explanation that reflects the egos of the leaders involved, the hundreds of billions gambled to date chasing compute capacity and the financial and legal risks facing each of those firms and leaders when the current bubble inevitably collapses in a matter of months if not weeks.

It has been argued in this forum MULTIPLE times over MULTIPLE years that the current American corporate development plan for AI capabilities quickly morphed into a financial bubble then into outright financial fraud involving circular revenue flows, circular investments from one top player into another, etc. It has also been argued on this forum that any sudden recognition of any of the circular inputs to the bubble suddenly violating prior assumptions should be enough for investors to pull their cash and trigger the invevitable collapse. Indeed, the number of states, counties and cities throughout the US enacting temporary freezes on data center construction or connections of new data centers to power grids alone should already be enough to stop the exponential feedback and burst the bubble.

Here is the key to the alternative explanation.

If the bubble bursts due to independent data and decisions made at state, county or local levels across the country, executives of the firms spending HUNDREDS OF BILLIONS will see their stock prices collapse or IPO dreams evaporate as investors finally realize a viable revenue model is NOT on the horizon. Those same executives will likely be subjected to numerous civil suits by shareholders and in a normal, non-Trump universe, those executives and their firms would be immediately subjected to investigations for SEC violations and accounting fraud.

ON THE OTHER HAND, imagine if the federal government steps in and establishes a moratorium on the roll-out of new capabilities pending the creation of regulations regarding the cyber-safety of AI systems and guidelines for establishing legal culpability for actions taken by AI systems and resulting damages. Those same executives could argue that their original business plans WERE viable and and were totally on track as promised but have now been scuttled by the federal government interjecting itself into the market. They could subsequently argue that any collapse in stock prices is due to a government induced force majeure and it is now the federal government's responsibility to bail out all parties involved. The AI firms are the victim, see? And even if no politician has an initial inclination to bail out the gamblers, the larger market collapse would be so catastrophic, these tech firms would still be at the head of the line when the federal government begins printing money to prop up the entire economy.

So which scenario is more likely? Has a quorum of executive leadership at some of the biggest firms in America suddenly reached some new god-tier of insight and conscience about their technology? Or has a cabal of executives who have guided their firms into an obvious bubble simply started making escape plans for their firms and their personal fortunes?

A quick review of the actors involved at the senior level of these firms and their ongoing history of monopoly abuses, anti-competitive practices against business and consumer customers alike and ethical concerns would seem to make choosing the higher-likelihood driver quite easy. Sam Altman of OpenAI? Larry Ellison of Oracle? Sundar Pichai of Alphabet? Mark Zuckerberg of Meta? Satya Nadella of Microsoft? Dario Amodei of Anthropic? Jensen Huang of Nvidia? To be fair, Huang has staked out a distinct position from most of the other execs, stating that perhaps they are generating fear about the capabilities of AI as a circular means of generating MORE demand for AI as the only tool that can keep up with itself to defend against itself. Of course, this rationalization can also merely reflect another tactic for "talking one's book" like all of the others or defending one's corporation from a sudden collapse in the market.

It's possible any question of motivations cannot be answered until after the federal government steps in and does something or until after the bubble collapses and the truth or a cloudy version of it is divulged in court. We'll just have to wait and see. The wait won't likely be very long.


WTH