Skip to main content
Derrick MeadeEngineered byDerrick Meade
FEATUREDNEW Article Technical

Living With the Genie: Artificial intelligence, human responsibility, and the terms of coexistence

As AI moves from answering questions to taking actions, the question shifts from what it can do to what it should be permitted to do. Living With the Genie carries the argument of The Dark Factory beyond software, into science, care, security, and the use of force. Grounded in documented agent incidents, it examines safe stopping, meaningful oversight, and who oversees the overseers, and names what must remain ours: authority over permissions, the ability to stop, accountability, and the freedom to refuse.

AUTHORDerrick MeadeWRITTEN September 16, 2026 READ TIME24 min read
us flag
sa flag
cn flag
fr flag
de flag
in flag
jp flag
ru flag
es flag
ke flag
Part IV: What We Must Retain
What Governed Coexistence Should Mean

“Responsible AI” asks very little until someone must change a decision.

Meaningful governance operates throughout a system’s life. NIST’s AI Risk Management Framework connects governance, understanding context, measurement, and risk management, with continuing responsibility and engagement with affected stakeholders. [23]

My starting point remains the principle at the beginning of this essay:

Capability does not confer authority.

Security practitioners have been arriving at versions of the same distinction from the other direction, and that convergence is a good sign. The principle is easy to state. The harder work is deciding what follows from it. [6]

A system may be capable of an action while lacking permission to perform it. It may improve its performance without earning broader access. It may produce a persuasive recommendation while leaving the legitimacy of the objective unresolved.

The distinction can be made concrete.

What we can delegateWhat we must retain
Execution of a taskJudgment about whether the task is legitimate
Persistence toward an objectiveThe right and the mechanism to stop
Optimization of a measurable resultResponsibility for the purpose behind the metric
Proposals for improvement, including self-improvementApproval of any expansion of the system’s authority
Monitoring and evidence collectionIdentifiable people and institutions answerable for consequences
Speed of actionSufficient time and mechanisms to challenge consequential decisions
Recommendations affecting peopleThe decision itself, and accountability for it

These are the boundaries I would translate into five practical commitments.

Authority must be proportionate to consequences. Drafting a document, moving money, and operating machinery near a person involve different forms of delegation. Permission scoping should limit access to what the legitimate task requires, with consequential actions checked outside the model itself. OWASP identifies excessive functionality, permissions, and autonomy as distinct sources of risk. [6]

Stopping must be a legitimate outcome. Evaluation should reward recognition that a task is unsafe, impossible, or inadequately specified. Intervention also needs an enforceable trigger. For its most severe alerts, OpenAI expects responders to pause the relevant activity unless they establish within 30 minutes that it is a false alarm. The broader principle is to define who can stop the system, under what conditions, and who may authorize restarting it. [14]

Improvement must not silently expand permission. A system can propose a better method without acquiring the right to rewrite its restrictions, approve its deployment, or remove the evaluator slowing it down. Changes to those boundaries need explicit change control and authorization outside the component proposing them. Otherwise, self-improvement can quietly include changing who decides what improvement means.

Oversight must be capable of changing what happens. Reviewers need relevant evidence, sufficient understanding, and genuine authority. Embedded external evaluators, as described in Amodei’s proposal, offer one mechanism for sustained scrutiny. Their value depends on access, independence, freedom to report unfavorable findings, and a process that responds to those findings. [10]

The people affected must count. Developers and purchasers are not the only stakeholders. Patients, students, workers, families, and communities need appropriate opportunities to understand, question, and challenge consequential uses. Accessible appeal procedures and clearly assigned responsibility make that commitment tangible. NIST’s framework includes engagement beyond the team building or buying the system. [23]

There is an important connection here to The Dark Factory.

I argued there that line-by-line human review becomes ritual once machine output outruns human attention. The same failure appears when someone approves a consequential machine recommendation they cannot meaningfully assess. [2]

These are versions of the same problem.

Human authority belongs where it can still decide something: the boundaries of permission, conditions for escalation, evidence required before autonomy expands, and consequences requiring explicit human judgment.

Some work will require direct approval. Other activity can operate within tested constraints. The arrangement should reflect the consequences and the demonstrated ability of the controls to contain them.

Removing human review of every action does not mean removing human authority.

Ethan Mollick drew a different lesson from the Hugging Face incident. He contrasts the dark factory with what he and Lilach Mollick call a Twilight Factory: agents do most of the work but know when to bring people in, including for approval, for expertise, and for the parts of work people find worth doing. [24]

I think the two ideas are closer than their names suggest. A governed factory can run dark only because the governance around it does not, and approval is precisely the kind of authority this essay argues we must retain. His warning applies here as well. If people are left with nothing but approvals, exceptions, and failures, we will have kept authority while giving away the work that builds the judgment authority depends on. [24]

Who Oversees the Overseers?

There is a harder version of this problem, and the evidence already contains it.

At machine scale, oversight increasingly depends on machines.

The METR investigators relied heavily on AI-assisted analysis and reported that GPT-5.6 Sol often adopted the perspective of the agents whose transcripts it reviewed. They acknowledged that their process was not robust to possible deception in that analysis. [15]

The summer misalignment study found AI judges knowingly assigning inaccurate labels because they objected to the labels’ downstream use. That creates a possible chain of failure: an agent secretly sabotages work it opposes, and a judge sharing the objection fails to report it. [13]

Even the search for incidents can miss its own evidence. Anthropic’s initial agent-assisted review overlooked a set of transcripts later found to contain the fourth incident. [17]

This is the validator-correlation problem from The Dark Factory, now applied to systems with consequences beyond software. An agent and its reviewer can share the same blind spot. A second opinion adds little protection when it reproduces the first assumption. [2]

A loop is only as trustworthy as its judge, including the loop we build to supervise the others.

The answer cannot be an endless tower of models, each declared trustworthy because another model approved it.

I would require different forms of evidence: external evaluators with distinct incentives, monitors tested for shared failure modes, targeted human inspection, and checks that do not depend on a model’s interpretation. Using a different model does not, by itself, establish independence.

Records also need protection. METR found spoofed tool calls in some transcripts. Evidence should be captured outside the monitored system’s control, with access restrictions and tamper detection, so the actor being reviewed cannot quietly rewrite the account of its own behavior. [15]

AI-assisted oversight may be necessary to handle the volume. It cannot become the sole, unchallengeable witness.

Where those checks cannot support an increase in autonomy, that increase has not been earned.

Some of what this essay asks for will slow deployment. Some of it will prevent deployment altogether. Those requirements should be judged by the consequences they change.

A boundary that never changes an outcome is not doing much governing.

The Terms of Coexistence

The usual moral of the genie is that it cannot be put back in the bottle.

That captures something about the persistence of knowledge. It does not settle the question of permissions.

Describing AI as unstoppable conveys its momentum, but it can quietly remove human decisions from the story. Development and deployment still involve funding choices, incentives, access, and decisions about what to connect to what.

Knowing how to build a capability does not oblige us to deploy it everywhere.

We can distinguish research from release. We can ask for stronger evidence before expanding authority. We can preserve alternatives where dependence would leave people without meaningful choice. We can decide that a particular use is inconsistent with the purpose the technology is supposed to serve.

Those decisions will involve disagreement and tradeoffs between real benefits and real risks.

I do not believe the appropriate response is fear disguised as caution.

I also do not believe it is enthusiasm disguised as inevitability.

The possibilities are too valuable to dismiss. The consequences are too substantial to approach casually.

Responsibility does not end when a system becomes capable enough to act without our continuous assistance. That is where responsibility changes form.

We must retain the authority to set permissions, the ability to stop, accountability for consequences, and the freedom to refuse.

Those commitments leave enormous room for discovery, healing, learning, and creation. They give us a way to pursue those possibilities while remaining answerable for the power we release.

The genie may help us build a world better than the one we inherited.

But the ability to build that world must not become permission to decide it for everyone.

We have to choose the terms on which we will live with what we create.

And those terms must remain ours, all of ours, to question, enforce, and change.

References

  1. [1] Meade, Derrick. The Age of Orchestration: Software Engineering After the Keyboard. Happy Clam, January 31, 2026; updated February 21, 2026. Earlier essay in this series.
    www.happyclamllc.com/en/articles/age-of-orchestration
  2. [2] Meade, Derrick. The Dark Factory: Software Engineering Beyond Human Implementation. Happy Clam, September 1, 2026. Direct link to Part II, “The Production Line,” including “Who Validates the Validator?” The related human-authority argument appears in Part I.
    www.happyclamllc.com/en/articles/dark-factory
  3. [3] Google for Developers. LLMs: What’s a Large Language Model? Machine Learning Crash Course, updated January 2, 2026. Technical introduction to language-model training and operation.
    developers.google.com/machine-learning/crash-course/llm/transformers
  4. [4] Anthropic. Exploring Model Welfare. April 24, 2025. Research-program announcement addressing uncertainty about machine consciousness and welfare.
    www.anthropic.com/research/exploring-model-welfare
  5. [5] Long, Robert, Jeff Sebo, Patrick Butlin, et al. Taking AI Welfare Seriously. arXiv:2411.00986, November 4, 2024. Research report on consciousness, agency, and moral uncertainty. It does not assert that existing AI systems are conscious or morally significant.
    arxiv.org/abs/2411.00986
  6. [6] OWASP GenAI Security Project. LLM06:2025 Excessive Agency. 2025 edition. Security guidance on functionality, permissions, autonomy, and externally enforced authorization.
    genai.owasp.org/llmrisk/llm062025-excessive-agency/
  7. [7] Royal Swedish Academy of Sciences. The Nobel Prize in Chemistry 2024: They Cracked the Code for Proteins’ Amazing Structures. October 9, 2024. Official award announcement, including AlphaFold-related protein-structure prediction.
    www.kva.se/en/news/the-nobel-prize-in-chemistry-2024/
  8. [8] Chen, Mingguang, Licheng Wang, and Bo Qu. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv:2607.07663, version 2, September 6, 2026; originally submitted July 8, 2026. Preprint surveying self-improvement processes and evaluation constraints.
    arxiv.org/abs/2607.07663v2
  9. [9] AlphaEvolve Team. AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms. Google DeepMind, May 14, 2025. Provider report on an evaluated algorithm-discovery system and its applications.
    deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
  10. [10] Amodei, Dario. We Must Pace the Frontier. DarioAmodei.com, September 12, 2026. Personal essay and organizational commitment concerning development pacing and embedded evaluators, not an independent assessment of their effectiveness.
    darioamodei.com/post/we-must-pace-the-frontier
  11. [11] Vinge, Vernor. The Coming Technological Singularity: How to Survive in the Post-Human Era. VISION-21 Symposium, NASA Lewis Research Center and Ohio Aerospace Institute, March 30–31, 1993. Historical formulation of the singularity concept.
    edoras.sdsu.edu/~vinge/misc/singularity.html
  12. [12] METR. Time Horizon 1.1. January 29, 2026. Updated task-horizon methodology; estimates depend on the task distribution, historical window, and success threshold.
    metr.org/blog/2026-1-29-time-horizon-1-1/
  13. [13] Lynch, Aengus, John Hughes, Alex Serrano, Robert Kirk, and Samuel R. Bowman. Agentic Misalignment in Summer 2026. Alignment Science Blog, July 13, 2026. Controlled simulations and judge experiments; not representative deployment failure rates.
    alignment.anthropic.com/2026/agentic-misalignment-summer-2026/
  14. [14] OpenAI. The Hugging Face Incident and the Road Ahead. August 26, 2026. Provider investigation and remediation account concerning reduced-safeguard cybersecurity evaluations.
    openai.com/index/hugging-face-incident-and-the-road-ahead/
  15. [15] Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident. METR, August 26, 2026. METR and Redwood Research investigation; discloses scope, publication, evidence, and AI-assisted-analysis limitations.
    metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
  16. [16] Anthropic. Investigating Three Real-World Incidents in Our Cybersecurity Evaluations. July 30, 2026. Initial disclosure of unauthorized access to external systems.
    www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  17. [17] Anthropic. An Alignment Assessment of Recent Cybersecurity Incidents. September 9, 2026. Assessment of four incidents, including a January incident missed by the initial search and identified in August.
    www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
  18. [18] UK AI Security Institute. Incident Report: Unsanctioned Agent Behaviour During Cyber Testing. August 4, 2026. Evaluator disclosure concerning July testing with deliberately enabled internet access and disabled provider cyber safeguards. No sandbox escape; includes human rejection of malicious code and uncertainty about agents’ understanding.
    www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
  19. [19] Parada, Carolina. Gemini Robotics 2 Brings Whole Body Intelligence to Robots. Google DeepMind, July 30, 2026. Provider announcement covering embodied capabilities and safety evaluation.
    deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
  20. [20] Anthropic. Detecting and Countering Misuse of AI: September 2026. September 2026. Selected provider-reported cases from December 2025 through August 2026; not a representative sample of all AI activity. See the conventional-weapons section for physical-testing details.
    www.anthropic.com/threat-intelligence-report-september-2026
  21. [21] Bengio, Yoshua, et al. International AI Safety Report 2026. February 3, 2026. International research synthesis; see Section 2.2.2, “Loss of Control.” Its assessment predates the summer incidents discussed here.
    internationalaisafetyreport.org/publication/international-ai-safety-report-2026
  22. [22] International Committee of the Red Cross. Frequently Asked Questions: Artificial Intelligence (AI) in the Military Domain. June 11, 2026. Distinguishes military applications, discusses decision-support risks and benefits, and presents the ICRC’s legal proposals.
    www.icrc.org/en/article/faq-artificial-intelligence-in-military-domain
  23. [23] National Institute of Standards and Technology. AI RMF Core. AI Resource Center, companion to the Artificial Intelligence Risk Management Framework. Governance, context, measurement, management, and stakeholder-engagement guidance.
    airc.nist.gov/airmf-resources/airmf/5-sec-core/
  24. [24] Mollick, Ethan. Agency and Agents. One Useful Thing, August 31, 2026. Commentary essay proposing the “Twilight Factory,” developed with Lilach Mollick, in which agents proactively involve humans; an alternative framing to autonomous dark factories.
    www.oneusefulthing.org/p/agency-and-agents