Wiki · Research
Public agent-framework synthesis
Shelf
Research

Public agent-framework synthesis

Six pinned extractions test the workflow-derived catalogue against independently built agent architectures. This is a comparison of mechanisms and boundaries, not a framework ranking. Each cell below links to an extraction whose Snapshot records its source repository, exact commit, licence, read date, measurements, and read boundary.

1. Six-framework comparison

Template sectionSWE-agentLangGraphOpenAI Agents SDKPydantic AIMicrosoft Agent FrameworkOpenHands
2. Core loopOne model → command → observation trajectory, with explicit format, cost, time, and retry exits (trace).A Pregel-style graph advances durable supersteps until no task remains or an interrupt halts it (trace).A small runner loops over model output, tools, guardrails, handoffs, and final output in both Python and TypeScript (trace).Prompt, model-request, and tool/output graph nodes retry invalid output and stop on a typed result or bound (trace).Executors exchange workflow messages over edges; a group-chat host selects participants and owns termination/output (trace).Canvas creates a conversation/worktree; Agent Server loops model steps, action events, tools, observations, approval waits, and terminal states (trace).
3. State and persistenceA .traj file is an append-after-step record, but incomplete runs restart rather than resume (state).Checkpoints, pending writes, channels, namespaces, and durability mode define replayable graph state (state).Run state can serialize interruption; session protocols persist conversation history, with stronger atomic history capabilities only in TypeScript (state).Core continuation is caller-supplied message history; optional durable-execution adapters move I/O into Temporal, DBOS, or Prefect (state).Workflow checkpoints capture executor state and pending requests through a host-selected store; emitted events are separate (state).Base state plus an event log restores a conversation; a restart marks an unmatched in-flight action as ambiguous error rather than rolling it back (state).
4. Tools / outside worldA compact agent-computer interface normalizes command schemas, state, truncation, empty output, and errors; the external runtime is another package (tools).Nodes and tasks may perform effects, but checkpoint replay does not make those effects exactly once (tools).Function tools, hosted tools, MCP, computer use, and guardrails share runner events; effect policy stays with the host/tool (tools).Typed inputs and dependencies validate the call boundary; ordinary callbacks still execute in the application process unless a durable adapter moves them (tools).AITool schema and invocation middleware support approval and telemetry, while deployment/auth/storage remain host concerns (tools).Local tools have host filesystem access; worktrees isolate Git state, while all-in-one Docker and optional DockerWorkspace expose different mount/network/process boundaries (tools).
5. CompositionNot applicable: the bounded default path is one agent; no graph/handoff/sub-agent primitive was found in the named read boundary (composition).Subgraphs and parallel graph tasks compose through channels, namespaces, and checkpointed supersteps (composition).Handoffs transfer continuation ownership; agent-as-tool runs a nested agent and returns to the caller (composition).Not applicable in the bounded core: another agent can be called as ordinary tool code, but no ownership-changing primitive was found (composition).Graph edges, executors, group-chat hosts, participant sessions, and explicit output mapping separate routing from disclosure (composition).Multiple backends/conversations compose operationally; optional delegated child conversations run concurrently but share the parent workspace path (composition).
6. Human in loopShellAgent supports terminal takeover and records human actions in the same trajectory, but the inspected loop has no pre-effect approval protocol (human boundary).An interrupt becomes checkpointed state with an addressable resume command; approver identity, authorization, and UI are application-owned (human boundary).Pending tool approvals live in serializable run state; the application approves/rejects and owns authority, UI, and persistence (human boundary).Deferred requests carry call identity and validated arguments into a later run; the application supplies approver, storage, and UI (human boundary).Typed request/response messages can be checkpointed, and one client binds approval to the original action; the host still supplies the human authority and ledger (human boundary).Canvas supplies live inspection, stop/resume, and risk buttons, but the server receives a boolean decision rather than the displayed action identity (human boundary).
7. BargainMinimal, inspectable coding loop; limited recovery and product infrastructure (bargain).Precise replay and durable graph semantics; replay-safe effects remain the author’s burden (bargain).Portable small primitives and escape hatches; operational durability and approval authority remain host work (bargain).Strong typing, injection, and deterministic seams; provider variance and durable-engine constraints remain visible (bargain).Rich workflow, checkpoint, approval, and telemetry machinery; enterprise properties remain opt-in and host-owned (bargain).A composed product supplies backend/workspace/history/Git/run-control surfaces; deployment shape still does not imply rollback or least privilege (bargain).

The pins are SWE-agent 3ea751c087f32b16e039a2233dd6eefecef325d5, LangGraph 644815f9e5bc52ad8f7a5227a456227e9c3e639b, OpenAI Agents Python 1a0c08868aec2a18eba964e5a07da4270a490c25, OpenAI Agents TypeScript d85dd2c144cd99bfdfa0111975cc759c00d56a77, Pydantic AI 9a602b3216b2cde46bfe29c1d32927eb36c501d6, Microsoft Agent Framework 12621e0a746517068300f7b9445225c3ee2406ea, Agent Canvas dc99e98615de4ace821692773b00a7f50d476e50, and OpenHands SDK 46ad3d43dc385b2e7975c0935f157153930ebb16; each resolves from its linked Snapshot above.

2. Convergences

A loop is a bounded state machine, even when its syntax is not a loop

All six turn probabilistic model output into named state transitions and terminal conditions. The linear subjects expose turn/cost/retry limits; graph subjects terminate on graph/manager state; the product adds pause, stuck, budget, and error states (SWE-agent loop, LangGraph loop, OpenAI loop, Pydantic loop, Microsoft loop, OpenHands loop). Opinion. I do not treat “until done” as an execution contract. A repository procedure invoking an agent should state the bounds and the state produced when each bound fires.

The tool boundary is agent-facing and effect-unsafe by default

Every subject converts a model decision into a structured action and converts execution back into model-visible state. Their strongest differences are schema richness and runtime location, not an exactly-once guarantee (SWE-agent ACI, LangGraph replay boundary, Pydantic tool boundary, OpenHands isolation boundary). Opinion. I treat tool schema, error shape, truncation, environment access, and replay behavior as one declared interface; validating arguments does not make the resulting effect reversible.

Continuity has at least three independent layers

Five extractions independently warn that conversation history, resumable machine state, and documentary intent are not interchangeable (SWE-agent warning, LangGraph warning, OpenAI warning, Pydantic warning, Microsoft warning). OpenHands adds a fourth useful distinction: a reconstructed event history can survive while the environment before an ambiguous tool effect cannot (OpenHands recovery). Opinion. I retain the existing session-handoff pattern only as narrative continuity; it must say explicitly that it is not a machine checkpoint, event store, or conversation-memory API.

A human pause is state; human authority is still external

LangGraph, OpenAI Agents SDK, Pydantic AI, Microsoft Agent Framework, and OpenHands can all stop on a pending decision and continue later. None of the inspected boundaries supplies the complete set of authenticated actor identity, authorization policy, durable retention, UI, and accountability ledger (LangGraph HITL, OpenAI HITL, Pydantic HITL, Microsoft HITL, OpenHands HITL). Microsoft’s binding client and OpenHands’ boolean response provide the positive and negative cases for one narrower invariant: the decision must bind to the exact surfaced request.

Composition needs ownership and disclosure semantics

An OpenAI handoff, an agent-as-tool call, a LangGraph subgraph, a Microsoft executor edge, and an OpenHands delegated child can look tool-shaped while differing in who continues, which state is shared, and what becomes public output (OpenAI composition, LangGraph composition, Microsoft composition, OpenHands composition). Opinion. I consider a routing declaration incomplete until it states continuation ownership, shared state, and the output allow-list.

3. Divergences

Opinion. I use the mechanisms column for extracted evidence and the reconciliation column for my synthesis judgement about what belongs in rungs.

ChoiceMechanisms observedReconciliation
Durability unittrajectory step · graph superstep/checkpoint · serializable run state/session history · caller history/external durable activity · workflow checkpoint · event log/base state (comparison §1)No universal default. Each mechanism must name what becomes durable together and what re-executes.
Tool contractcompact command ACI · arbitrary node code · multiple tool/provider protocols · typed dependency/schema boundary · middleware-wrapped AITool · terminal/editor/browser inside a workspace (tool rows)Shared principles are bounds, validation, effect declaration, and explicit errors; tool breadth is product/framework specific.
Compositionnone in two bounded cores · graph/subgraph · ownership-changing handoff or nested tool · executor graph/group-chat host · shared-workspace child conversations (composition row)Admit ownership/output practices, not a preferred graph or delegation topology.
Testing seamscripted models and loop traces appear in several repos; Pydantic AI makes deterministic substitution plus separate provider evidence the clearest contract (Pydantic tests)Add a testing pattern, while preserving the existing rule that a fake alone cannot prove the real boundary.
IsolationFive bounded reads leave ordinary tool execution to the host/runtime; OpenHands alone exposes local-host, deployment-container, and per-workspace-container choices (OpenHands tools)Execution-boundary declaration is commensurable with repo instructions. Container lifecycle and sandbox implementation are product architecture, not a rungs module pattern.
Run controlLibrary hosts receive events/interruption objects; OpenHands composes history, live stream, terminal/browser/Git views, stop/resume, and confirmation UI (OpenHands human boundary)The need to expose pending state is portable; a product UI/live-tail architecture is not commensurable with the workflow catalogue.

Opinion. I take the most important divergence to be category, not implementation: the first five subjects are agent libraries or focused runtimes, while OpenHands is a product that must package persistence, workspaces, credentials, ingress, and user control (product residue). Silently treating those product mechanisms as workflow-module defaults would merge the two corpora at the point where their responsibilities differ most.

4. What nobody solved

  1. Exactly-once outside-world effects or general rollback. LangGraph defines replay but still requires idempotent/recorded effects; Pydantic’s durable adapters inherit engine retry rules; OpenHands records an interrupted action as ambiguous rather than reversing it (LangGraph state, Pydantic state, OpenHands state).
  2. Semantic truth. Typed outputs, guardrails, structured tools, and retries can reject malformed values; none establishes that a well-formed model answer is true (Pydantic output gate, OpenAI guardrail bargain).
  3. Complete approval accountability. The strongest request binding still expects a host to authenticate the approver and retain decisions; the shipped UI counter-example sends only a boolean (Microsoft approval, OpenHands approval).
  4. A portable isolation and resource contract. OpenHands documents several useful boundaries, but local mode has host access and its inspected Docker workspace sets no CPU, memory, PID, or read-only limit; the five other reads do not supply a common sandbox contract (OpenHands boundary).
  5. One continuity artifact for intent, replay, events, and environment. The convergence above establishes that these are separate layers, not that combining them is desirable or possible (continuity convergence).
  6. Comparable real-run capacity or cost. OpenHands’ synthetic concurrency test deliberately does not establish real tool/model capacity, and the corpus did not run or benchmark subjects (OpenHands concurrency, corpus scope).

5. Catalogue reconciliation

SW, LG, OA, PA, MF, and OH below refer to the six pinned extractions in §1. “New (merged)” means the candidate is adjudicated as a clause of another admitted id rather than copied under two names.

Opinion. I make the outcomes and resulting changes as synthesis decisions; their evidence cells link to the pinned observations that make each decision reviewable.

Pattern idOutcomeEvidenceResulting catalogue change
narrowest-anchor-loopconfirmedSWAdd independent source SW; definition and rung stay.
prompt-writes-artifactconfirmedSWAdd SW and clarify that durable progress may be written before successful completion; rung stays 2.
session-handoffconfirmed (scope narrowed)SW, LG, OA, PA, MFRetain rung 1; state that this is narrative continuity, not checkpoint, event log, or conversation memory.
skill-neighboursconfirmed (strengthened)OAAdd OA; require ownership/return semantics as part of the neighbour boundary.
contract-test-baseconfirmed (strengthened)PAAdd PA; require separate real-boundary evidence because a deterministic fake alone proves only the loop.
structural-gatesnot commensurablePA analogyTyped model-output validation supports the same reasoning but is not evidence about repository link/id/path gates; no source or rung change.
scope-disciplinenot commensurableMF analogyNamespaced workflow state is analogous to work ownership, not evidence for backlog scope; no change.
worktree-lifecycleconfirmed (scope narrowed)OHAdd OH; state that worktree lifecycle coordinates Git state and is not a process, filesystem, credential, or network sandbox.
candidate: agent-facing-interfacenewSWAdmit agent-facing-interface under invocation, rung 2.
candidate: bounded-agent-loopnewSW, PA, OHAdmit bounded-agent-loop under invocation, rung 2.
candidate: durable-superstepnot commensurableLGKeep as architecture evidence only; rungs’ documentary procedures do not own a transactional checkpoint unit.
candidate: replay-safe-side-effectnewLG, PA, OH counter-exampleAdmit under workflows, rung 3: a resumable procedure names replay and effect handling.
candidate: interrupt-as-statenew (merged)LGFold stable pending identity/resume state into resumable-approval-state; do not create a duplicate id.
candidate: ownership-changing-handoffnewOA, MFAdmit under invocation, rung 2.
candidate: protocol-with-escape-hatchnewOA, PA portability warningAdmit under workflows, rung 2.
candidate: resumable-approval-statenewOA, PA, MFAdmit under workflows, rung 2, including the merged interrupt-state clause.
candidate: deterministic-model-substitutionnewPAAdmit under testing, rung 2, paired with separate real-boundary evidence.
candidate: typed-output-gatenewPAAdmit under gates, rung 1; distinguish structural validity from semantic truth.
candidate: approval-bound-to-requestnewMF positive case, OH counter-exampleAdmit under workflows, rung 2: immutable id/arguments, authorized decision, one-time consumption.
candidate: event-stream-not-audit-lognewMF, OHAdmit under session continuity, rung 1, as an accountability-boundary warning.
candidate: explicit-output-designationnewMFAdmit under workflows, rung 2: public output is an allow-list, not incidental graph connectivity.
candidate: isolation-boundary-declarationnewOHAdmit under instructions, rung 0: execution unit, crossings, and absent controls must be named.
candidate: event-log-plus-live-tailnot commensurableOHRetain as product-architecture evidence; no current rungs module owns a replayable UI transport.
candidate: run-control-surfacenot commensurableOHRetain as product residue; terminal/browser/Git UI is not a workflow-module default.
candidate: shared-workspace-subagentsdemotedOHReject as a catalogue default: parallel writers need explicit ownership or serialization first.

Every candidate named by the six extractions is adjudicated above. Opinion. I make no existing rung changes: the independent evidence changes definitions and source confidence, while every admitted practice fits the maturity threshold of its target module. Product-only candidates remain in this synthesis so the absence from the catalogue is a decision rather than an omission.

6. Consequences for rungs

The catalogue changes are documentation inputs to the shipped modules; this item does not edit modules/. The affected surfaces are instructions, workflows, skills, gates, session, and testing. WI-029 will decide which definitions require changes to manifests, templates, skills, gates, or module versions.

No ADR is admitted. The admission rule fails because this work item and the canonical catalogue already own the classification, and reversing a documentation-only reconciliation later is not materially more expensive. An ADR would duplicate the reconciliation table rather than constrain a separate future decision.

Opinion. I take the strongest finding to be non-commensurability. OpenHands shows that backend selection, sandbox implementation, event-tail transport, and run-control UI are real product work. The workflow catalogue should make execution boundaries visible, but it should not pretend that installing a repository module supplies those product capabilities.