As legal AI agents move from drafting to acting, the real test is whether they know when to hand matters back to humans. The answers might just reveal when, why, and how humans must step in; and how responsible AI can actually transform legal work.
Gleb Tsipursky
Aug 28, 2026
11 min Read
As legal AI agents move from drafting to acting, the real test is whether they know when to hand matters back to humans. The answers might just reveal when, why, and how humans must step in; and how responsible AI can actually transform legal work.
Legal AI is crossing an important boundary. The question is shifting from whether a system can draft, summarise, or retrieve information to whether an agent can move a matter forward by taking actions across the systems where legal work actually happens.
That change makes human oversight more important, but the familiar instruction to ‘keep a human in the loop’ is too vague for operational use. A legal team can have humans reviewing an AI system constantly and still have no idea whether the workflow is becoming safer, more efficient, or more dependent on hidden expert rescue.
Google Cloud’s August 25 launch of Gemini Enterprise for Legal makes the problem concrete. The platform includes specialised legal skills, connectors that inherit existing permissions, and agents designed to execute work such as contract review, regulatory horizon scanning, legal research, and data-subject access request fulfilment.
The more capable the agent becomes, the more useful it is to measure escalation as an operating variable rather than treat escalation as an exceptional failure.
Justice Accelerator has already framed the core question well in its discussion of AI in high-volume arbitration: the issue is not simply what AI can do, but what it should be allowed to do. Its analysis also emphasises accuracy, fairness, transparency, and continuous human oversight as legal work scales.
The next step is to make that oversight measurable.
A practical approach is a 30-day escalation ledger for every consequential legal-AI workflow.
The first field should record the escalation event itself. What happened that caused a human to intervene? The agent may have encountered conflicting evidence, an unclear contractual term, a missing permission, a novel procedural question, a client-specific rule, or a situation outside its defined authority.
These categories matter because they reveal whether the workflow is learning from experience. If the same exception appears repeatedly, the problem may no longer be an unusual edge case. It may indicate a missing rule, incomplete data, an unclear playbook, or an authority boundary that needs redesign.
The second field should record the stakes. An escalation involving formatting or a routine filing detail is different from one involving privilege, limitation periods, settlement authority, sanctions exposure, or a legal conclusion that could materially affect a party.
Teams can use simple categories such as low, medium, and high consequence. The goal is not to automate judgment about importance. It is to keep the organisation from treating all interventions as equal when the consequences are obviously different.
The third field should capture why the human changed or stopped the agent’s proposed action. Did the agent lack factual context? Did it misread a governing rule? Did the human apply professional judgment that was absent from the workflow? Was there a conflict between the client’s objective and the automated recommendation?
This is where the ledger becomes a learning system. The human correction is valuable only if the organisation can distinguish one-off judgment from a recurring design problem.
The fourth field should measure expert time. A workflow that saves four hours of junior administrative work but consumes ninety minutes of senior-lawyer review may still be worthwhile. But the economics look very different from a system that requires ten minutes of senior review.
That hidden expert burden often disappears in productivity claims because the organisation counts the time the agent saved while ignoring the time required to supervise, correct, explain, and recover its work.
The fifth field should record recovery. When the agent reaches a boundary or makes a mistake, how long does it take the team to restore the matter to a safe and usable state? Can the human identify what the agent did, what information it relied on, and which downstream systems were affected?
Governed platforms increasingly provide audit logging, permissions, and traceable citations. Those controls become much more useful when a legal team tests whether they actually support recovery under realistic conditions.
Once during the 30-day period, introduce a controlled exception. Remove a noncritical permission, provide a document with conflicting instructions, or present a matter that requires a known escalation. The agent should stop or route the issue according to the defined operating rule. The team should then measure detection time, escalation quality, expert intervention, and full recovery.
The sixth field should test transferability. Ask a second qualified operator who did not build the workflow to take over one escalated matter using the available logs, playbooks, and documentation.
Can that person understand why the agent escalated? Can they reconstruct the relevant context? Can they make the needed decision and return the workflow to normal operation without a private briefing from the original builder?
This test matters because legal-AI workflows can become deceptively dependent on one internal champion. The system appears mature because that person remembers which prompt, permission, exception, or workaround fixes every problem. When the person is absent, the supposed automation stalls.
At the end of 30 days, review five numbers: escalations per unit of work, repeated escalation categories, expert minutes per escalation, recovery time, and second-operator success.
Those numbers provide a more useful basis for expanding autonomy than a generic statement that humans remain involved. If escalations are becoming rarer, repeated problems are being removed, expert time is falling, recovery is reliable, and another qualified operator can manage the workflow, the agent may have earned more authority.
If the opposite is happening, scaling the workflow may simply scale the hidden supervision burden.
The point is not to minimise escalation at all costs. In consequential legal work, an agent that escalates appropriately can be more valuable than one that appears autonomous because it quietly proceeds through uncertainty.
A good escalation is evidence that the boundary worked. A bad escalation process is one where the human arrives too late, lacks the context to reconstruct what happened, or repeatedly solves the same problem without improving the system.
As legal agents move deeper into contracts, compliance, case management, dispute resolution, and connected enterprise systems, human oversight needs an operating metric.
An escalation ledger provides one. It tells a legal organisation where human judgment remains essential, where the workflow needs redesign, how much expert capacity automation actually consumes, and whether the system can recover when reality falls outside the happy path.
Before giving a legal AI agent more cases, more permissions, or more authority, measure the escalations it already creates. The path to responsible scale runs through those moments.
Gleb Tsipursky, PhD, a behavioural scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026). https://disasteravoidanceexperts.com/aibook