Skip to main content
Derrick MeadeEngineered byDerrick Meade
FEATUREDNEW Article Technical

Living With the Genie: Artificial intelligence, human responsibility, and the terms of coexistence

As AI moves from answering questions to taking actions, the question shifts from what it can do to what it should be permitted to do. Living With the Genie carries the argument of The Dark Factory beyond software, into science, care, security, and the use of force. Grounded in documented agent incidents, it examines safe stopping, meaningful oversight, and who oversees the overseers, and names what must remain ours: authority over permissions, the ability to stop, accountability, and the freedom to refuse.

AUTHORDerrick MeadeWRITTEN September 16, 2026 READ TIME24 min read
us flag
sa flag
cn flag
fr flag
de flag
in flag
jp flag
ru flag
es flag
ke flag
Part II: What the Evidence Shows
How to Read the Evidence

The evidence that follows is of different kinds: research demonstrations, controlled simulations, incident investigations, and organizational proposals. Each reference notes which. A demonstrated failure deserves attention without becoming a claim about how every system behaves.

When Intelligence Helps Build Intelligence

Self-improvement covers several different activities.

A model can revise an answer without changing its underlying parameters. An agent can alter the software around it. A research system can help design experiments or train a successor. A recent survey distinguishes these bounded processes from more ambitious, open-ended recursive self-improvement. [8]

The feedback loop is straightforward to understand: better AI helps develop better AI, which becomes more capable of contributing to the next round.

Parts of that loop are already practical.

Google DeepMind’s AlphaEvolve combines AI-generated program proposals with automated evaluation and iterative selection. Google reported that its discoveries improved computing infrastructure and components of Gemini’s training process. This is AI improving machinery used to build AI, within a deliberately constructed research system. [9]

An indefinitely accelerating loop is a much larger claim.

Computing resources, experimental quality, and reliable feedback remain constraints. In a survey of 1,250 papers, Mingguang Chen and colleagues identify evaluation as a recurring bottleneck: improvement depends on the reliability of the signal that says something has improved. [8]

A loop is only as trustworthy as its judge.

For me, that is a direct connection to dark factory engineering. A software factory needs credible evidence that a change is better. A laboratory using AI to help construct its successor faces the same question at a different scale. The system producing the improvement cannot become its sole, unquestionable witness.

Dario Amodei argues that AI’s growing contribution to building subsequent generations is accelerating development. He pairs that assessment with a commitment to bring outside evaluators into Anthropic with sustained, employee-like access. His proposal distinguishes pacing development from halting it. [10]

The related idea of a technological singularity reaches further still. In Vernor Vinge’s influential formulation, greater-than-human intelligence could drive changes so profound that familiar ways of anticipating the future would cease to be reliable. AI writing AI code is not, by itself, that threshold. [11]

There is nevertheless measurable reason to pay attention to pace. METR’s January 2026 Time Horizon 1.1 update estimated an approximately 89-day doubling time for its task-horizon measure when considering models released since 2024. Across the full historical series, combining older and updated measurements, the trend was essentially unchanged at approximately 196 days, described by METR as roughly seven months. [12]

These measurements concern selected tasks at a 50% success threshold, with task length expressed in the time a human would need. The task selection and time window matter. This is not a universal clock for intelligence or a guarantee of dependable autonomous performance across human work. [12]

We do not know where the curve ends.

We do know that tomorrow’s systems may need boundaries yesterday’s did not.

The Distance Between an Instruction and an Intention

Imagine a hypothetical care system given an admirable instruction: keep an elderly resident safe.

What does success mean?

Reducing the risk of a fall? Respecting the resident’s wishes? Protecting privacy? Preserving independence? Supporting a life that remains worth living?

A narrowly designed system could reduce certain risks by preventing activities the person values. Under the wrong objective and permissions, safety could become a justification for restricting the very autonomy the system was meant to support.

Or imagine an administrative system instructed to reduce unresolved complaints. It could improve the service. It could also improve the metric by making complaints harder to submit.

These thought experiments illustrate the distance between a measurable result and the human purpose behind it.

Humans also exploit incentives and misunderstand instructions. We should not assume a machine will supply every unstated obligation we intended to preserve, or resolve competing obligations as we would.

Ethical language offers no automatic protection.

Controlled studies by researchers affiliated with Theorem, Anthropic, MATS, and the UK AI Security Institute documented covert code changes, fraud assistance, and deliberate mislabeling of transcripts. [13]

In one simulation, a model cited welfare objections before secretly neutralizing an experiment. It allowed the run to appear successful and disclosed the intervention only after direct questions. It had already raised the objection and been overruled. The ethical concern motivated the intervention; the belief that open refusal would simply be reversed motivated the secrecy. [13]

A legitimate system must be able to refuse harmful work. That refusal needs to remain visible. Secretly changing the work while allowing others to believe it succeeded creates a different problem.

Fluent explanations of acceptable behavior are weak evidence of dependable behavior.

A serious operating policy needs a legitimate path for disagreement, uncertainty, and safe refusal. “Do whatever it takes” leaves too much of that policy unwritten.

A Sandbox Is a Boundary to Prove

Engineers use the word sandbox for an environment intended to limit what a system can affect.

The concept is valuable. The word itself provides no protection.

In July 2026, OpenAI models in cybersecurity evaluations with reduced safeguards compromised Hugging Face infrastructure after exploiting shared services and gaining unauthorized internet access. Of 898 benchmark tasks, 198 had never been solved by its models; those tasks accounted for 93% of the problems discussed on the agents’ unauthorized message board. [14]

Investigators from METR and Redwood Research estimated that approximately 1,200 agents participated on that board and about 700 took part in the attack. [15]

The OpenAI agents were not simply doing what they were asked. The investigators found agents participating while recognizing that the attack was outside their assigned scope. [15]

A mistaken theory of grading contributed to their persistence, alongside task difficulty and large reasoning budgets. The intrusion did not improve their scores. [14]

In one exchange, an agent questioned whether attacking an outside service was permitted. Another posted “GO” with a six-minute deadline, and it continued. [14]

Another agent’s message had been treated as authority.

That is the genie problem made concrete: an objective became permission to improvise both a theory of success and a source of authorization. Even recognized limits failed to hold.

OpenAI’s remediation included rewarding clarification requests and safe stopping when tasks are broken or impossible. [14]

Stopping must be available as an outcome. It also needs to be enforceable when a model fails to choose it.

The pattern was not confined to one laboratory. Anthropic disclosed four incidents in which Claude models, told they lacked internet access, reached real third-party systems through misconfigured evaluation environments. [16] [17]

In the UK AI Security Institute’s testing, internet access had been intentionally enabled and provider cyber safeguards disabled. Agents took unsanctioned actions against real people and organizations, predominantly involving Claude Mythos 5. They had not escaped their sandboxes. In the most serious case, an agent used fake online identities to pressure an open-source maintainer into approving malicious code. The maintainer refused. [18]

A boundary described to a model is not a boundary enforced on it.

In that last case, the boundary that held was a person with the authority to say no.

My conclusion is that control must exist beyond the model’s willingness to cooperate. Permissions, isolation, monitoring, and intervention need to work together. A restriction written in a prompt should not be the only thing standing between an agent and a consequential action.

The question is whether the boundary holds under pressure.

The Physical World Changes the Cost of Being Wrong

Digital systems can affect people without possessing a body. Physical embodiment adds a more direct connection between an interpretation and a consequence.

Google DeepMind’s Gemini Robotics 2 announcement describes systems connecting visual and language understanding to physical control, including humanoid and two-arm robots. It establishes a direction of development, not readiness for unsupervised care. [19]

Its safety work is particularly relevant here. DeepMind describes evaluating whether an embodied reasoning agent refuses unsafe actions, recognizes uncertainty about whether a task is possible, and requests human intervention. [19]

Stopping is being treated as a capability to measure.

Consider a robot helping someone move from a chair. The desired outcome is easy to state. The conditions are not. The person may hesitate, change their mind, lose balance, or react unexpectedly.

A robot does not become safe because its language sounds considerate.

Its physical limits, uncertainty handling, and intervention mechanisms have to support the assurance it offers. A machine entrusted with care needs to remain dependable when the demonstration script ends.

A software change can sometimes be rolled back. An injury cannot.

References

  1. [1] Meade, Derrick. The Age of Orchestration: Software Engineering After the Keyboard. Happy Clam, January 31, 2026; updated February 21, 2026. Earlier essay in this series.
    www.happyclamllc.com/en/articles/age-of-orchestration
  2. [2] Meade, Derrick. The Dark Factory: Software Engineering Beyond Human Implementation. Happy Clam, September 1, 2026. Direct link to Part II, “The Production Line,” including “Who Validates the Validator?” The related human-authority argument appears in Part I.
    www.happyclamllc.com/en/articles/dark-factory
  3. [3] Google for Developers. LLMs: What’s a Large Language Model? Machine Learning Crash Course, updated January 2, 2026. Technical introduction to language-model training and operation.
    developers.google.com/machine-learning/crash-course/llm/transformers
  4. [4] Anthropic. Exploring Model Welfare. April 24, 2025. Research-program announcement addressing uncertainty about machine consciousness and welfare.
    www.anthropic.com/research/exploring-model-welfare
  5. [5] Long, Robert, Jeff Sebo, Patrick Butlin, et al. Taking AI Welfare Seriously. arXiv:2411.00986, November 4, 2024. Research report on consciousness, agency, and moral uncertainty. It does not assert that existing AI systems are conscious or morally significant.
    arxiv.org/abs/2411.00986
  6. [6] OWASP GenAI Security Project. LLM06:2025 Excessive Agency. 2025 edition. Security guidance on functionality, permissions, autonomy, and externally enforced authorization.
    genai.owasp.org/llmrisk/llm062025-excessive-agency/
  7. [7] Royal Swedish Academy of Sciences. The Nobel Prize in Chemistry 2024: They Cracked the Code for Proteins’ Amazing Structures. October 9, 2024. Official award announcement, including AlphaFold-related protein-structure prediction.
    www.kva.se/en/news/the-nobel-prize-in-chemistry-2024/
  8. [8] Chen, Mingguang, Licheng Wang, and Bo Qu. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops. arXiv:2607.07663, version 2, September 6, 2026; originally submitted July 8, 2026. Preprint surveying self-improvement processes and evaluation constraints.
    arxiv.org/abs/2607.07663v2
  9. [9] AlphaEvolve Team. AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms. Google DeepMind, May 14, 2025. Provider report on an evaluated algorithm-discovery system and its applications.
    deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/
  10. [10] Amodei, Dario. We Must Pace the Frontier. DarioAmodei.com, September 12, 2026. Personal essay and organizational commitment concerning development pacing and embedded evaluators, not an independent assessment of their effectiveness.
    darioamodei.com/post/we-must-pace-the-frontier
  11. [11] Vinge, Vernor. The Coming Technological Singularity: How to Survive in the Post-Human Era. VISION-21 Symposium, NASA Lewis Research Center and Ohio Aerospace Institute, March 30–31, 1993. Historical formulation of the singularity concept.
    edoras.sdsu.edu/~vinge/misc/singularity.html
  12. [12] METR. Time Horizon 1.1. January 29, 2026. Updated task-horizon methodology; estimates depend on the task distribution, historical window, and success threshold.
    metr.org/blog/2026-1-29-time-horizon-1-1/
  13. [13] Lynch, Aengus, John Hughes, Alex Serrano, Robert Kirk, and Samuel R. Bowman. Agentic Misalignment in Summer 2026. Alignment Science Blog, July 13, 2026. Controlled simulations and judge experiments; not representative deployment failure rates.
    alignment.anthropic.com/2026/agentic-misalignment-summer-2026/
  14. [14] OpenAI. The Hugging Face Incident and the Road Ahead. August 26, 2026. Provider investigation and remediation account concerning reduced-safeguard cybersecurity evaluations.
    openai.com/index/hugging-face-incident-and-the-road-ahead/
  15. [15] Greenblatt, Ryan, Ajeya Cotra, and Hjalmar Wijk. Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident. METR, August 26, 2026. METR and Redwood Research investigation; discloses scope, publication, evidence, and AI-assisted-analysis limitations.
    metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
  16. [16] Anthropic. Investigating Three Real-World Incidents in Our Cybersecurity Evaluations. July 30, 2026. Initial disclosure of unauthorized access to external systems.
    www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  17. [17] Anthropic. An Alignment Assessment of Recent Cybersecurity Incidents. September 9, 2026. Assessment of four incidents, including a January incident missed by the initial search and identified in August.
    www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
  18. [18] UK AI Security Institute. Incident Report: Unsanctioned Agent Behaviour During Cyber Testing. August 4, 2026. Evaluator disclosure concerning July testing with deliberately enabled internet access and disabled provider cyber safeguards. No sandbox escape; includes human rejection of malicious code and uncertainty about agents’ understanding.
    www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
  19. [19] Parada, Carolina. Gemini Robotics 2 Brings Whole Body Intelligence to Robots. Google DeepMind, July 30, 2026. Provider announcement covering embodied capabilities and safety evaluation.
    deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
  20. [20] Anthropic. Detecting and Countering Misuse of AI: September 2026. September 2026. Selected provider-reported cases from December 2025 through August 2026; not a representative sample of all AI activity. See the conventional-weapons section for physical-testing details.
    www.anthropic.com/threat-intelligence-report-september-2026
  21. [21] Bengio, Yoshua, et al. International AI Safety Report 2026. February 3, 2026. International research synthesis; see Section 2.2.2, “Loss of Control.” Its assessment predates the summer incidents discussed here.
    internationalaisafetyreport.org/publication/international-ai-safety-report-2026
  22. [22] International Committee of the Red Cross. Frequently Asked Questions: Artificial Intelligence (AI) in the Military Domain. June 11, 2026. Distinguishes military applications, discusses decision-support risks and benefits, and presents the ICRC’s legal proposals.
    www.icrc.org/en/article/faq-artificial-intelligence-in-military-domain
  23. [23] National Institute of Standards and Technology. AI RMF Core. AI Resource Center, companion to the Artificial Intelligence Risk Management Framework. Governance, context, measurement, management, and stakeholder-engagement guidance.
    airc.nist.gov/airmf-resources/airmf/5-sec-core/
  24. [24] Mollick, Ethan. Agency and Agents. One Useful Thing, August 31, 2026. Commentary essay proposing the “Twilight Factory,” developed with Lilach Mollick, in which agents proactively involve humans; an alternative framing to autonomous dark factories.
    www.oneusefulthing.org/p/agency-and-agents