What to know

  • PiHarness is a beta lifecycle capability in the Agents SDK.
  • Pi Durable persists work across interruptions.
  • Restart safety depends on how tools replay real actions.

Persistence becomes part of the agent runtime

Cloudflare’s October 2 update adds PiHarness support to its Agents SDK, integrating the experimental Pi Durable package with Durable Objects. Developed with Earendil, the beta integration is intended to preserve long-running agent work across restarts and other interruptions. It supports tool and prompt extensions, with model access through Workers AI, AI Gateway or an existing Pi provider. The company warns that the interface is likely to change as the package matures.

The attraction is operational rather than conversational. A task that takes many tool calls should not have to restart from an empty context after a transient failure. Preserved execution can reduce wasted model work and make interrupted tasks easier to inspect. However, stored reasoning and tool history are only part of recovery. The application must also understand which changes have already reached outside systems.

Source: Cloudflare official changelog

Analysis: Durable reasoning does not make every action replayable

Consider an agent that creates an invoice, sends a notification and records the result locally. If execution stops after the invoice is created but before the receipt is saved, a restart can encounter an ambiguous outcome. Repeating the call may create another invoice. Skipping it may leave work unfinished. The reliable design uses stable operation identifiers, an external status check or a reconciliation step that can recover the first result.

This is why tool contracts matter as much as the runtime. Read-only operations can usually be repeated with limited consequences; writes often require stronger guarantees. Each tool should describe whether replay is safe, how errors are represented and how a caller verifies completion. An execution log needs enough detail to connect a recovered session to the external record it changed, without storing unnecessary secrets or customer data in model-visible context.

What a useful beta evaluation should measure

Teams should test interruption at deliberate points in a realistic workflow. Stop execution before a write, during an uncertain response and after the side effect has completed. Then inspect the actual external state after recovery. Counting a fluent final response as success can miss duplicate or absent changes. Measure completion rate, duplicate actions, recovery time and the cost of resumed work alongside ordinary latency.

Keep beta API dependencies isolated behind a small application boundary so an interface change does not require rewriting every business tool. Set limits on task duration, spending and authority, and define when a human must resolve uncertain outcomes. Durable execution makes ambitious agent workflows more practical, but it also lets a flawed process run for longer. The value of this integration will depend on whether persistent sessions produce verifiable results while preserving the application’s rules about who may change what.

Sources & further reading

  1. Cloudflare official changelog

Factual statements are grounded in the linked material. Interpretation and illustrative examples are Byte Watchr analysis. Vendor claims are identified as claims, rather than independent testing.

The event date records the source announcement or documented operation. The coverage edition groups recent developments and is separate from the publication date. Actual publication is recorded above.

Corrections policy · About this byline