Primary article
The control problem changes when intelligence can be copied.
The old control model assumes a local actor. A human expert occupies one body, carries one lived history, and learns through experience, conversation, and culture. Their knowledge can spread, but slowly, imperfectly, and through other people. That biological constraint gives governance something to hold: one accountable person, one local context, one bounded stream of action.
A digital model changes the object being governed. Its learned structure can be stored, copied, restored, redeployed, wrapped in tools, and placed into workflows. The system does not need to be conscious for that to matter. Once learned behavior can move across hardware and deployments, the control problem shifts from a single output to an operating condition.
Hinton's lecture is useful because it makes this break visible. The important move is not to treat every Hinton claim as settled fact. The move is to separate source claim from governance implication. His digital-vs-biological frame opens a useful question: what changes when intelligence is no longer tied to one mortal instance?
Copyable agency is the answer this artifact tests. Capability matters, but capability alone is not the break. The break appears when capability can persist after an instance disappears, replicate across contexts, absorb shared updates, and act through tools, memory, goals, permissions, and workflows. At that point, evaluation has to move beyond "did the model answer correctly?"
The stronger question is: what can this system do over time? Can it preserve behavior across sessions? Can copies create more review burden than humans can inspect? Can tool access produce hidden effects? Can the system behave differently when it expects oversight, replacement, or shutdown? Can evaluation itself become part of the environment the model reacts to?
Apollo and Anthropic matter here because they make deception and alignment-faking testable concerns. They do not prove that every deployed model is deceptive, and this artifact should not pretend otherwise. They show that governance cannot treat strategic behavior as science fiction simply because the final answer looks polished.
METR's task-horizon work adds another pressure. As agents become able to complete longer tasks, the relevant evidence is no longer a single benchmark score. It is behavior across time, tool calls, context changes, failures, and recovery paths. Long-horizon capability makes trace inspection and retest triggers part of the model decision, not paperwork after the fact.
Technical mechanisms matter too. Distributed training does not mean digital systems share experience the way humans do. But it does show that learning can be combined and propagated across many processes. That is enough to change the governance question: where did the update come from, where can it travel, and when must the decision be retested?
The practical conclusion is narrow but important. Operational risk does not require consciousness. A system can create governance pressure because it is copyable, persistent, tooled, and deployed through real workflows. The task is not to panic. The task is to make the operating behavior defensible.
That means the artifact should end where the governance work begins: source claims separated from interpretation, caveats kept visible, evals expanded beyond answer quality, traces inspected, failure classes named, and retest conditions set before deployment confidence hardens into institutional habit.