Wiki · Research
SWE-agent
Shelf
Research

SWE-agent

1. Snapshot

This snapshot fixes the source boundary used by every claim and count below.

FieldValue
RepositorySWE-agent/SWE-agent
Pinned commit3ea751c087f32b16e039a2233dd6eefecef325d5
Date read2026-08-15
LicenceMIT — LICENSE
LanguagesPython package with shell/Python tool bundles — pyproject.toml, tools/
Measured scale409 tracked files; 100 tracked .py files; 12,413 lines across those Python files

Measured 2026-08-15 at the pinned commit, in PowerShell:

(git ls-files | Measure-Object -Line).Lines
(git ls-files -- '*.py' | Measure-Object -Line).Lines
git ls-files -- '*.py' | ForEach-Object { (Get-Content -LiteralPath $_ | Measure-Object -Line).Lines } | Measure-Object -Sum

The first two commands prove the tracked-file and Python-file counts. The third produced 12,413; it measures physical lines in every tracked Python file, including tests, not source-only or logical LOC.

2. The core loop

The default loop is in DefaultAgent.run: setup establishes the environment, tools, initial prompt history, and trajectory path; then run calls step until StepOutput.done and writes the trajectory after each step (setup).

One turn is the following concrete path:

  1. DefaultAgent.step sends the processed messages history to forward_with_handling.
  2. DefaultAgent.forward queries the model, parses its response into thought and action through ToolHandler, then calls handle_action.
  3. DefaultAgent.handle_action blocks forbidden commands, handles the explicit exit action, sends the guarded command to the environment, captures observation and state, and detects retry, forfeit, or submission signals.
  4. step appends the result to model history, updates exit/submission/model statistics, and appends a durable trajectory step. add_step_to_history gives empty and oversized observations explicit feedback shapes instead of passing raw output unchanged.

Errors remain part of the loop rather than bypassing it. forward_with_handling requeries format, blocklist, content-policy, and shell-syntax failures, while cost, context, execution-time, repeated timeout, environment, and API failures terminate through an attempted patch submission (forward_with_handling). Executable tests exercise step-by-step history and the cost/context/format/blocklist exit statuses (tests/test_agent.py).

3. State and persistence

The live agent keeps three distinct in-memory records: model-facing history, an append-only trajectory, and summary info (DefaultAgent.__init__). messages filters history to the named agent and applies configured history processors before a model call; processing does not replace the stored history (messages).

The .traj file is the durable boundary. get_trajectory_data serializes trajectory, full history, summary info, replay configuration, and environment name, and save_trajectory writes that JSON after every completed step (get_trajectory_data and save_trajectory, run).

That durability is not checkpoint recovery. Batch mode skips a trajectory with a final exit status, but deletes and reruns an empty, unreadable, early_exit, or status-less trajectory (RunBatch.should_skip). Replay constructs a new run that reissues recorded actions; it does not restore the old process and continue its model history from the crash point (run_replay.py).

Bounded absence check, 2026-08-15. rg -n 'resume|checkpoint|restore|replay' sweagent tests docs found replay, repository reset, and batch skip/delete paths, but no live-agent resume or checkpoint loader. This establishes absence only within those three tracked directories at the pinned commit.

4. Tools and the outside world

A tool bundle is a directory whose config.yaml is validated into named Command objects (Bundle). Each command has a signature, documentation, and typed arguments; it can be rendered as a function-calling schema (Command). ToolConfig.commands combines the built-in bash command with bundle commands and rejects duplicate names, while ToolConfig.tools exposes their schemas (ToolConfig).

At setup, ToolHandler uploads each bundle into the runtime, installs it, extends PATH, and checks that each command exists (ToolHandler.install). At turn time, parsing is delegated to the configured parser, commands pass a blocklist, multiline input is guarded, and execution crosses SWEEnv.communicate (ToolHandler, SWEEnv.communicate).

The isolation guarantee belongs below this repository: SWEEnv sends actions to a SWE-ReX deployment/runtime, and pyproject.toml declares swe-rex>=1.4.0 (swe_env.py, pyproject.toml). This extraction therefore establishes the boundary to the sandbox, not SWE-ReX’s containment properties.

Submission is also a tool protocol. The default configuration loads review_on_submit_m; its submit command first writes /root/model.patch and emits review feedback, and only a later stage prints the sentinel that handle_submission treats as terminal (config/default.yaml, tools/review_on_submit_m/bin/submit, handle_submission).

5. Composition

The default architecture constructs one named agent around one model and later attaches one environment during setup (DefaultAgent, setup). Bounded absence check, 2026-08-15. rg -n 'subagent|sub-agent|handoff|delegate' sweagent found no agent-to-agent delegation primitive in the live package at the pinned commit.

Composition exists one level above the turn as retry-and-select. RetryAgent creates a fresh DefaultAgent for an attempt, hard-resets the environment between attempts, records each attempt, and asks a score or chooser loop whether to retry and which result to retain (RetryAgent._setup_agent and _next_attempt, RetryAgent.run). The boundary is therefore sequential attempts plus reviewer state, not peers sharing control of one trajectory.

6. The human in the loop

DefaultAgent has no per-action approval step in the traced loop. The explicit intervention mode is the separate ShellAgent: Ctrl+C swaps the model for HumanModel, Ctrl+D restores the original model, and an AI-generated terminal result is handed to the human for final submission (ShellAgent). HumanModel reads actions from the terminal, including multiline commands, and returns them through the same model-response interface (HumanModel).

Because takeover still runs step, human actions and observations enter the same history and trajectory as model actions (ShellAgent.run, DefaultAgent.step). The durable record identifies the action, observation, and agent name, but the inspected types carry no separate approval decision or approver identity (types.py).

7. The abstraction bargain

Opinion. I think SWE-agent’s strongest abstraction is the agent-computer interface, not the while loop. The loop is small; most of the mechanism is spent making tools typed, installable and visible, shaping empty/truncated observations, rejecting unsafe interaction shapes, and turning submission into an explicit sentinel protocol. The premises are the tool and history paths traced in sections 2 and 4.

Opinion. I would treat .traj as an audit and replay artifact, not as durable execution. Saving after every step is valuable, but batch mode’s response to an incomplete artifact is deletion and a fresh run. Calling that “resume” would erase the most important persistence boundary found in section 3.

Opinion. I think retry-and-select is deliberately cheaper than multi-agent collaboration: each attempt owns a normal trajectory and the meta-layer shares only the problem, reset environment, budget, submissions, and reviewer result. That makes the cost legible, but it cannot express agents cooperating inside one turn.

8. What rungs takes

These verdicts are inputs to WI-017; they do not change the catalogue here.

Pattern idVerdictPinned evidenceReason
narrowest-anchor-looptakeconfig/default.yamlOpinion. I read the default “find/read → reproduce → edit → rerun” instruction as independent architecture-level support for the existing workflow pattern.
prompt-writes-artifacttakeDefaultAgent.runOpinion. I would strengthen the catalogue’s durable-output rationale with a system that writes the full trajectory after every step, not only at successful completion.
session-handofftake-as-warningget_trajectory_data, RunBatch.should_skipOpinion. I take the counter-example: a detailed durable log still is not a resumable handoff when incomplete work is discarded and restarted.
candidate: agent-facing-interfacetakeCommand.get_function_calling_tool, add_step_to_historyOpinion. I think tools, errors, empty output, truncation, and state should be designed as one agent-facing interface; the existing catalogue has no definition for that boundary.
candidate: bounded-agent-looptakeforward_with_handling, ToolConfigOpinion. I would extract the explicit cost, context, wall-time, consecutive-timeout, and format-retry limits as one termination-budget pattern rather than leave “until done” unbounded.