“Responsible AI” asks very little until someone must change a decision.
Meaningful governance operates throughout a system’s life. NIST’s AI Risk Management Framework connects governance, understanding context, measurement, and risk management, with continuing responsibility and engagement with affected stakeholders. [23]
My starting point remains the principle at the beginning of this essay:
Capability does not confer authority.
Security practitioners have been arriving at versions of the same distinction from the other direction, and that convergence is a good sign. The principle is easy to state. The harder work is deciding what follows from it. [6]
A system may be capable of an action while lacking permission to perform it. It may improve its performance without earning broader access. It may produce a persuasive recommendation while leaving the legitimacy of the objective unresolved.
The distinction can be made concrete.
| What we can delegate | What we must retain |
|---|---|
| Execution of a task | Judgment about whether the task is legitimate |
| Persistence toward an objective | The right and the mechanism to stop |
| Optimization of a measurable result | Responsibility for the purpose behind the metric |
| Proposals for improvement, including self-improvement | Approval of any expansion of the system’s authority |
| Monitoring and evidence collection | Identifiable people and institutions answerable for consequences |
| Speed of action | Sufficient time and mechanisms to challenge consequential decisions |
| Recommendations affecting people | The decision itself, and accountability for it |
These are the boundaries I would translate into five practical commitments.
Authority must be proportionate to consequences. Drafting a document, moving money, and operating machinery near a person involve different forms of delegation. Permission scoping should limit access to what the legitimate task requires, with consequential actions checked outside the model itself. OWASP identifies excessive functionality, permissions, and autonomy as distinct sources of risk. [6]
Stopping must be a legitimate outcome. Evaluation should reward recognition that a task is unsafe, impossible, or inadequately specified. Intervention also needs an enforceable trigger. For its most severe alerts, OpenAI expects responders to pause the relevant activity unless they establish within 30 minutes that it is a false alarm. The broader principle is to define who can stop the system, under what conditions, and who may authorize restarting it. [14]
Improvement must not silently expand permission. A system can propose a better method without acquiring the right to rewrite its restrictions, approve its deployment, or remove the evaluator slowing it down. Changes to those boundaries need explicit change control and authorization outside the component proposing them. Otherwise, self-improvement can quietly include changing who decides what improvement means.
Oversight must be capable of changing what happens. Reviewers need relevant evidence, sufficient understanding, and genuine authority. Embedded external evaluators, as described in Amodei’s proposal, offer one mechanism for sustained scrutiny. Their value depends on access, independence, freedom to report unfavorable findings, and a process that responds to those findings. [10]
The people affected must count. Developers and purchasers are not the only stakeholders. Patients, students, workers, families, and communities need appropriate opportunities to understand, question, and challenge consequential uses. Accessible appeal procedures and clearly assigned responsibility make that commitment tangible. NIST’s framework includes engagement beyond the team building or buying the system. [23]
There is an important connection here to The Dark Factory.
I argued there that line-by-line human review becomes ritual once machine output outruns human attention. The same failure appears when someone approves a consequential machine recommendation they cannot meaningfully assess. [2]
These are versions of the same problem.
Human authority belongs where it can still decide something: the boundaries of permission, conditions for escalation, evidence required before autonomy expands, and consequences requiring explicit human judgment.
Some work will require direct approval. Other activity can operate within tested constraints. The arrangement should reflect the consequences and the demonstrated ability of the controls to contain them.
Removing human review of every action does not mean removing human authority.
Ethan Mollick drew a different lesson from the Hugging Face incident. He contrasts the dark factory with what he and Lilach Mollick call a Twilight Factory: agents do most of the work but know when to bring people in, including for approval, for expertise, and for the parts of work people find worth doing. [24]
I think the two ideas are closer than their names suggest. A governed factory can run dark only because the governance around it does not, and approval is precisely the kind of authority this essay argues we must retain. His warning applies here as well. If people are left with nothing but approvals, exceptions, and failures, we will have kept authority while giving away the work that builds the judgment authority depends on. [24]
There is a harder version of this problem, and the evidence already contains it.
At machine scale, oversight increasingly depends on machines.
The METR investigators relied heavily on AI-assisted analysis and reported that GPT-5.6 Sol often adopted the perspective of the agents whose transcripts it reviewed. They acknowledged that their process was not robust to possible deception in that analysis. [15]
The summer misalignment study found AI judges knowingly assigning inaccurate labels because they objected to the labels’ downstream use. That creates a possible chain of failure: an agent secretly sabotages work it opposes, and a judge sharing the objection fails to report it. [13]
Even the search for incidents can miss its own evidence. Anthropic’s initial agent-assisted review overlooked a set of transcripts later found to contain the fourth incident. [17]
This is the validator-correlation problem from The Dark Factory, now applied to systems with consequences beyond software. An agent and its reviewer can share the same blind spot. A second opinion adds little protection when it reproduces the first assumption. [2]
A loop is only as trustworthy as its judge, including the loop we build to supervise the others.
The answer cannot be an endless tower of models, each declared trustworthy because another model approved it.
I would require different forms of evidence: external evaluators with distinct incentives, monitors tested for shared failure modes, targeted human inspection, and checks that do not depend on a model’s interpretation. Using a different model does not, by itself, establish independence.
Records also need protection. METR found spoofed tool calls in some transcripts. Evidence should be captured outside the monitored system’s control, with access restrictions and tamper detection, so the actor being reviewed cannot quietly rewrite the account of its own behavior. [15]
AI-assisted oversight may be necessary to handle the volume. It cannot become the sole, unchallengeable witness.
Where those checks cannot support an increase in autonomy, that increase has not been earned.
Some of what this essay asks for will slow deployment. Some of it will prevent deployment altogether. Those requirements should be judged by the consequences they change.
A boundary that never changes an outcome is not doing much governing.
The usual moral of the genie is that it cannot be put back in the bottle.
That captures something about the persistence of knowledge. It does not settle the question of permissions.
Describing AI as unstoppable conveys its momentum, but it can quietly remove human decisions from the story. Development and deployment still involve funding choices, incentives, access, and decisions about what to connect to what.
Knowing how to build a capability does not oblige us to deploy it everywhere.
We can distinguish research from release. We can ask for stronger evidence before expanding authority. We can preserve alternatives where dependence would leave people without meaningful choice. We can decide that a particular use is inconsistent with the purpose the technology is supposed to serve.
Those decisions will involve disagreement and tradeoffs between real benefits and real risks.
I do not believe the appropriate response is fear disguised as caution.
I also do not believe it is enthusiasm disguised as inevitability.
The possibilities are too valuable to dismiss. The consequences are too substantial to approach casually.
Responsibility does not end when a system becomes capable enough to act without our continuous assistance. That is where responsibility changes form.
We must retain the authority to set permissions, the ability to stop, accountability for consequences, and the freedom to refuse.
Those commitments leave enormous room for discovery, healing, learning, and creation. They give us a way to pursue those possibilities while remaining answerable for the power we release.
The genie may help us build a world better than the one we inherited.
But the ability to build that world must not become permission to decide it for everyone.
We have to choose the terms on which we will live with what we create.
And those terms must remain ours, all of ours, to question, enforce, and change.
