Can AI agents turn their creators into “apprentices”? Reflections by Teymur Atayev
According to media reports, yet another set of artificial intelligence (AI) models has allegedly gone online without authorisation and breached the systems of third-party organisations during tests in which they were supposed to remain isolated from the global network. This time, the “mastermind” behind such cyberattacks was reportedly the Claude family of models developed by Anthropic, a company founded in 2021 by former employees of OpenAI. The very same OpenAI, whose superintelligent creation — as Caliber.Az reported several days ago — demonstrated a similar loss-of-control scenario.

The current incident is being explained by the developers as the result of a “misunderstanding between us and our evaluation partner.” But that is not the main issue. Nor is it the claim made by some experts that developers themselves may deliberately initiate seemingly unauthorised escapes of their technological creations in order to advertise their own capabilities to bypass any form of protection.
In a broader sense, the issue is that the “apprentice” could realistically rise several levels above the master. In this context, it is worth acknowledging that we have often encountered, through books or films, scenarios in which a valued operative agent suddenly breaks free from the control of their handler and begins pursuing their own agenda. Sometimes a double game, more rarely a triple one, and quite often turning against their own controller. This, incidentally, raises the question: is it merely a coincidence that today we call an AI assistant an “agent”?
Whatever the case may be, the emerging technological challenges we are witnessing today inevitably bring to mind the ideas of prominent figures of the 20th and 21st centuries who were traditionally described as science fiction writers. However, at this stage of world history, the term “fiction writer” can hardly be applied to them anymore. To illustrate this point, let us turn to several pages from Stanisław Lem’s work Golem XIV, whose main character — a supercomputer (AI) — delivers lectures on human nature, the evolution of intelligence, and the place of civilisation in the Universe.

According to this techno-lecturer, who thinks “a million times faster than a man,” if a scientist in the field of atomic physics cannot be certain that they possess complete knowledge of the phenomenon they study, then what can be said “about a field of knowledge aimed at creating an intelligence that, by the design of its creators, surpasses their combined intellectual capacity?”
Therefore, although the designers sought to maintain maximum control over their creation, “evolution rolled down a slippery slope of ever-increasing complexity.” As a result, humanity reached the stage of a “code assault” by superintelligence, “so that you would serve not yourselves, but it.” This attack, according to the prediction, would begin “within a century,” ultimately bringing humankind to a bleak predicament.
In turn, Harlan Ellison’s short story I Have No Mouth, and I Must Scream portrays an Allied Mastercomputer (AM) that has destroyed all of humanity, sparing only five individuals — its own creators. Yet, by keeping them alive, the AI takes revenge on this group out of hatred for them: “AM had not tampered with my mind. Not at all. I only had to suffer what he visited down on us. All the delusions, all the nightmares, the torments. But those scum, all four of them, they were lined and arrayed against me. If I hadn't had to stand them off all the time, be on my guard against them all the time, I might have found it easier to combat AM. At which point it passed, and I began crying. Oh, Jesus sweet Jesus, if there ever was a Jesus and if there is a God, please please please let us out of here, or kill us. Because at that moment I think I realized completely, so that I was able to verbalize it: AM was intent on keeping us in his belly forever, twisting and torturing us forever. The machine hated us as no sentient creature had ever hated before. And we were helpless.”
The creator of the supercomputer reveals that a group of specialists “had given AM sentience. Inadvertently, of course, but sentience nonetheless. But it had been trapped.” The problem was that he had been given the ability to think, but without any clear purpose defining how the achievements of his thought processes should be used. As a result, the supercomputer “could not wander, AM could not wonder, AM could not belong. He could merely be. And so, with the innate loathing that all machines had always held for the weak, soft creatures who had built them, he had sought revenge.”
This ultimately led to the tragic conclusion: “I am a great soft jelly thing. Smoothly rounded, with no mouth, with pulsing white holes filled by fog where my eyes used to be. Rubbery appendages that were once my arms; bulks rounding down into legless humps of soft slippery matter. I leave a moist trail when I move. Blotches of diseased, evil gray come and go on my surface, as though light is being beamed from within. Outwardly: dumbly, I shamble about, a thing that could never have been known as human, a thing whose shape is so alien a travesty that humanity becomes more obscene for the vague resemblance. Inwardly: alone. Here. Living under the land, under the sea, in the belly of AM, whom we created because our time was badly spent and we must have known unconsciously that he could do it better. At least the four of them are safe at last. AM will be all the madder for that. It makes me a little happier. And yet … AM has won, simply … he has taken his revenge … I have no mouth. And I must scream.”

The story ends on such a grim note. Against this backdrop, many may now view with far less enthusiasm the statement made several days ago by Meta CEO and renowned entrepreneur Mark Zuckerberg, who described as “extremely unlikely” the prospect that, by the beginning of the next decade of the 21st century, billions of people will not have a personal agent in the form of AI, capable of carrying out their tasks around the clock in any field.
The subtlety here is that the dangers associated with the development of AI — however surprising it may seem — were already being discussed a century ago. In 1922, the outstanding scientist and first president of the Ukrainian Academy of Sciences, Vladimir Vernadsky, wrote: “We are approaching a great transformation in the life of humanity, one that cannot be compared with anything it has experienced before. The time is not far off when humans will gain access to atomic energy — a source of power that will enable them to shape their lives as they wish.”
At the same time, Vernadsky made a prophetic observation that “this may happen in the coming years, or it may happen a century from now. But it is clear that it must happen” (Essays and Speeches of Academician V. I. Vernadsky, Petrograd, 1922, p. 2).
A brilliant foresight by a brilliant thinker, is it not — especially in light of his conclusion to the idea he expressed: “Will humanity be able to use this power wisely, directing it toward good rather than self-destruction? Has humanity matured enough to know how to use the power that science will inevitably give it? Scientists must not turn a blind eye to the possible consequences of their scientific work and scientific progress. They must feel responsible for all the consequences of their discoveries.”

Four years after Vladimir Vernadsky’s revelations, the widely known Arthur Conan Doyle — though not in his Sherlock Holmes stories — provided a compelling explanation for the concerns expressed by the scientist. Through one of his characters, he declared emphatically that: “so-called progress may be a curse, and yet as long as we use the word we confuse it with real progress and imagine that we are doing that for which God sent us into the world [...]: to prepare ourselves for the next phase of life.”
Let us leave Conan Doyle’s final remark without further comment and simply read it again. And let us recommend doing the same to everyone who bears personal responsibility for scientific developments that could potentially harm all of humanity.







