Dream Hatch Labs
← All articles

AI implementation

Human Approval, Retries, and Recovery: Designing Reliable Agent Workflows

Build approval and recovery into agent workflows with explicit action proposals, version checks, idempotency, durable status, and exception handling.

Dream Hatch Labs

Human approval is useful when a person can inspect a concrete proposed action and understand its consequences. A vague “continue?” prompt does not establish which record, amount, recipient, or operation the person approved.

Combine approval with explicit action state and recovery behavior. Otherwise, a properly approved action can still be applied twice, executed against changed data, or reported as complete when its outcome is unknown.

Define the proposal before requesting approval

Consider an agent preparing a replacement shipment. The proposal should identify the verified order, replacement items, destination, relevant policy, and any material cost. Give it a stable proposal identifier and record the version of the underlying data.

The reviewer should be able to approve, reject, or request a correction. A changed address or quantity creates a materially different proposal; do not silently reuse the earlier approval.

Keep the approval decision in application state. It should not depend on the model remembering a sentence from the conversation.

Model the action lifecycle

A small illustrative lifecycle is:

proposed → awaiting approval → approved → executing → confirmed
                   ↓                         ↓
                rejected                outcome unknown
                                             ↓
                                     reconcile or escalate

Store timestamps, actors, proposal version, and operation identifier. A task awaiting reconciliation is different from a failed task and different from a confirmed action.

Before execution, verify that the approval is still applicable, the reviewer had authority, and the relevant record has not changed in a way that invalidates the proposal. The exact checks depend on the business system.

Handle a timeout without guessing

The shipping API accepts the replacement but the response is lost. The agent sees a timeout. Retrying with a new operation identifier may create a second shipment.

Where supported, reuse an idempotency key bound to the original action. Query the downstream system for status and record the confirmed result. If the downstream API lacks that capability, design a reconciliation process or route the uncertain case to an operator before retrying.

Exactly-once business behavior is an integration requirement to verify, not a promise created by an agent prompt or checkpoint.

Distinguish retries from compensation

A retry attempts to complete the same intended action. A compensating action attempts to address the consequences of an action that already occurred, such as cancelling a mistakenly created shipment.

Compensation may have its own cost, permissions, and failure modes. Do not assume every operation can be undone. Record which actions are reversible, which need a new approval, and which require manual intervention.

Use bounded retries with appropriate delays for transient failures. Repeating an invalid request or a permission denial indefinitely is not recovery. Return meaningful error categories so the workflow can choose the right response.

Give the operator enough evidence

An exception view should show the proposed action, approval, attempts, downstream references, and remaining uncertainty. It should support a verified resolution rather than forcing someone to read a long conversation and infer what happened.

Protect this evidence according to its contents. Operational visibility should not expose customer records to people who would otherwise lack access.

Test interruptions at each boundary

Interrupt the process before approval, after approval, during execution, and after the downstream system accepts the action. Resume the task and check for duplicate effects, stale approvals, and misleading completion messages.

Test a changed record, a revoked reviewer permission, a rejected proposal, and a correction that changes the action. Verify the final state in the business system, not only the agent's response.

Use the same cases in your agent evaluation suite. For the wider architecture, see why agents fail in production. Dream Hatch Labs can help implement the workflow and integrations around the actions your agent needs to perform.