Skip to content
all episodes
3 AImigo · episode 09

What Can We Trust an AI Agent With? Boundaries of Autonomy

1:52:54
Conversation

What we discussed on the recording

The hosts begin with the scope of delegation: an action, a task, a block of work, a role, and an end-to-end outcome. Aleksey distinguishes specific tasks, improvement directions, and areas of responsibility. Evgeny describes gradual automation of content releases at Flo Health; Alexander describes a marketing process constrained by approvals. Generation cannot accelerate the entire workflow while expert requirements and acceptance rules remain solely with people at its end.

New models provide a reason to revisit accumulated instructions. Aleksey experiments with fewer constraints and restores what failures show is necessary. Comparisons need representative tasks and quality criteria. Evgeny describes an iOS-to-Android game migration checked through emulators; Alexander distinguishes this experience from responsibility for a published product. In the podcast-cover example, the agent takes on repetitive work while the person supplies missing photographs. Permissions, costs, and consequences determine acceptable independence.

Discussing METR’s task horizon, Alexander separates human task duration from agent runtime. A public evaluation cannot replace checks of a specific workflow. The conversation about context leads to preserving the reasons behind architecture decisions and behavioral criteria. Aleksey warns against instructions that prematurely close off exploration, while Evgeny discusses regenerating an implementation from preserved knowledge as a test of whether that knowledge is sufficient.

The resulting assessment includes accepted tasks, interventions, verification costs, recovery time, and the scope of potential damage. Preparing policies, maintaining evaluations, and manually rescuing the process are work too. Evgeny proposes identifying recurring problems through session telemetry. The closing discussion considers managing people and agents: engineers need delegation skills, while the continuing need for managers and the prospect of flatter structures depend on organizational scale and design.

Agent autonomyTask delegationResult verificationHuman oversight