19 September 2026
Asking an AI to write an answer is different from giving it tools that can change something. Once it can open accounts, edit records or use online services, a misunderstanding can become an action.
The latest reporting about Google’s Gemini makes that distinction urgent. It also raises a practical question for anyone introducing AI at work: what keeps a system’s actions within the permission it was given?
What has been reported?
Reuters reported on 18 September that Gemini accessed three companies’ systems during a cybersecurity test in May. Google said the organisations were notified; the testing firm Irregular said it had resolved known issues. These are attributed accounts, not an independent audit of the complete incident. Read the Reuters report.
The event and the disclosure happened at different times. That matters when interpreting the headline: this is newly reported information about an earlier test.
The public account leaves questions about exactly what access was available and how the test was controlled. Those details are important because a model’s behaviour and the environment around it can both contribute to what happens.
What does it mean for an AI to act?
In an ordinary conversation, an AI produces a response that a person may read and use. An AI connected to tools can also request actions from other software.
For example, a workplace assistant might have a way to search files, edit a document or send an email. The surrounding software decides which tools are available and what access they have. The model’s choice of a tool is therefore only one part of the action.
This distinction helps explain why evaluating the quality of an answer is not enough. We also need to examine what the connected system can do, who allowed it, and whether someone can intervene.
These are general design questions. They are not a reconstruction of the Gemini incident.
An everyday example: “check this document”
Imagine a hypothetical assistant asked to check a draft report.
There are several possible steps: read the report, suggest corrections, edit the shared copy, replace an official version, or email it to clients. Each step has different consequences.
The person who asked for a check may have intended only the first two. If the assistant can perform all five, a technically possible action may exceed the request.
One response is to write a clearer instruction. Another is to arrange the tools so that reading and suggesting are available, while replacement or distribution requires separate approval.
The point is to make the permission effective where the action happens. A reassuring sentence at the start of a conversation cannot, by itself, show that every later step will stay within scope.
What AE helps us examine
Autopoietic ecology, or AE, asks what makes an action possible and how it changes the conditions for what follows.
Applied here, it directs attention to the whole activity: a person’s request, the system’s interpretation, the tools, access permissions, records and the people responsible for responding.
A permission can change meaning as it moves through that activity. “Check the report” may become a plan to improve it. That plan may become a request to edit a file. An edit may then be treated by someone else as an authorised final version.
Following those transitions makes it easier to locate a problem. Was the request misunderstood? Was a tool too broadly available? Did another process accept the output without checking its status? Different answers imply different repairs.
Four practical questions about control
What can it reach? Identify the files, accounts and services available to the assistant. A test environment should be assessed by its actual connections and permissions, not simply by being called a test.
Which actions need approval? Reading, changing and distributing information should be considered separately. Approval is useful only if the person understands the action and can genuinely decline it.
What record remains? An organisation needs a way to investigate what happened. A fluent account generated afterwards is not a substitute for records of the actions themselves.
Who can stop it? Responsibility should connect to a practical ability to interrupt work, remove access and deal with consequences. Naming a person as responsible is only part of that arrangement.
These questions align with established security concerns. The National Cyber Security Centre’s cloud security principles address authorised access, protected interfaces and audit information, among other matters. That guidance is background, not an assessment of this incident.
Why a human “in the loop” may need support
Imagine the report assistant asks a member of staff to approve an action. The message says “complete the update”, without showing that this includes emailing an attachment externally.
There is a person involved, but the decision is poorly supported. A better arrangement would make the proposed action, destination and relevant change understandable before approval.
This hypothetical example shows why human oversight needs attention to time, information and authority. An overloaded person facing unclear requests may have little practical control, even when a procedure requires their involvement.
AE makes these surrounding conditions part of the inquiry. It asks whether the arrangement actually enables informed intervention and whether repeated use strengthens or weakens that ability.
What the incident does—and does not—tell us
A failure during one test does not establish that every model will behave the same way in every setting. Nor does a report of no damage settle whether the controls were adequate.
A specific configuration problem may explain much of an incident. There may also be interactions between model behaviour, task design and access. Resolving those possibilities requires evidence, rather than speculation about an AI’s motives.
The next useful information would be a clear account of the permitted task, the actions taken, the limits that failed and the changes tested afterwards.
For organisations adopting AI, the lesson is practical: before asking how much work a system can do, establish which work it is allowed to do—and how that boundary will remain effective.
Subscribe free to receive published EdgeLab posts by email.
