Wiki · Research
Follow-on public-agent research
Shelf
Research

Follow-on public-agent research

The follow-on corpus approved in WI-018 tests boundaries that the fixed six-framework corpus did not centre: durable agent-managed memory, repeatable evaluation and optimization, git-native/local product operation, and interoperability across process or ownership boundaries.

This is one corpus with three evidence tracks, not a ranking and not an expansion of frameworks/. Every extraction uses the same comparison spine and one track template. Synthesis compares like subjects within a track before it makes any cross-track claim.

Corpus question

Which practices survive product, evaluation, and interoperability boundaries; which claims are specific to one kind of evidence; and what does each result confirm, contradict, add to, or keep separate from the rungs pattern catalogue?

The names and questions below are selection hypotheses. They become findings only when their work item records pinned evidence.

Index

WorkTrackSelection questionState
Shared spine · product template · evaluation template · protocol template · WI-019MethodWhich fields are genuinely comparable, and which evidence rules vary by track?Ready for extraction
Letta Code · WI-020Durable/local productWhere do identity, memory writes, archival retrieval, and continual learning live and survive?In progress · pinned ec4e23a85d4aa2a449ed5c7fb0801a0be1bd68d0
Inspect AI · WI-021Evaluation/optimizationWhat makes an agent evaluation reproducible, isolated, inspectable, and aggregatable?Queued
Aider · WI-022Durable/local productHow do repository context, editing, validation, and Git history constrain one coding loop?Queued
goose · WI-023Durable/local productHow do local execution, extensions, MCP/ACP, and session isolation meet at the product boundary?Queued
Google ADK · WI-024Durable/local productHow do delegation, sessions, evaluation, and multi-language public contracts evolve together?Queued
MCP · WI-025Interoperability protocolWhich tool/context lifecycle, capabilities, errors, and trust responsibilities cross the client-server boundary?Queued
A2A · WI-026Interoperability protocolWhich discovery, task, artifact, streaming, and identity semantics cross independently operated agents?Queued
DSPy · WI-027Evaluation/optimizationHow do metrics, traces, examples, and optimizers turn an agent program into an improvement loop?Queued
Follow-on synthesis · WI-028All threeWhich results reconcile within a track and which are not commensurable across tracks?Queued

Evidence labels

Every material claim begins with or is unambiguously governed by one of these labels. A link alone does not identify what kind of claim the link supports.

LabelAdmitted evidenceWhat it can establishBoundary it cannot cross
NormativeVersioned specification text pinned to a full source commitWhat a conforming implementation is required, recommended, or permitted to doThat any implementation conforms, or that application policy is safe
ImplementedPinned source path and symbol, preferably paired with its executable testWhat the pinned code path doesThat an optional path is the default, or that hosted/current behaviour matches the pin
ExecutedNamed command against the pinned checkout, with date, inputs, environment, and resultWhat that bounded run demonstratedGeneral performance, portability, or absence outside the run boundary
MeasuredReproducible command, result, date, and exact path/ref scopeA count or property the command computesQuality, importance, or another adjacent interpretation
DocumentedDocumentation pinned with the same source snapshotWhat the project claims or instructsThat implementation or conformance was verified
OpinionExplicit synthesis of cited premisesA judgement useful to rungsA factual implementation or normative claim

For protocol work, Normative outranks a reference implementation when the question is what the protocol requires. For product and evaluation work, implementation or executable tests are needed for behavioural claims. Documentation remains useful evidence of a public contract, but must stay labelled as documentation.

Method

  1. Choose the track before reading. Copy the shared spine and the named track addendum into the extraction. Do not change tracks merely because an awkward result does not fit the hypothesis.
  2. Freeze every authority. Record a full commit SHA, read date, licence file, and read boundary before extracting claims. If a subject spans multiple repositories or specification and implementation sources, snapshot each separately.
  3. Trace one mechanism end to end. The item plan names the mechanism. Follow it through state, external effects, evidence artifacts, failure, and recovery rather than inventorying features.
  4. Try to disprove the selection question. Every extraction ends its analysis with the strongest counter-evidence found. An absence claim names the directories, symbols, tests, and search terms inspected.
  5. Separate authority from operation. A normative requirement, optional negotiated capability, reference implementation, and application policy are four different claims. So are a documented benchmark, a locally executed evaluation, and a general performance conclusion.
  6. Reconcile once. Subject items may propose existing or candidate pattern ids but do not edit pattern-catalog.md. WI-028 adjudicates them with all eight subjects in view.

Cross-track synthesis rule

Shared-spine fields may be compared across all eight subjects. Track-addendum fields are compared within their track first. A cross-track row must choose one of four outcomes:

  • commensurable — the claims share an authority type and boundary;
  • analogy only — the mechanism is useful vocabulary but not evidence for the other track;
  • contradiction — equivalent claims under equivalent boundaries disagree; or
  • not commensurable — authority, subject, or guarantee differs enough that one verdict would be misleading.

“Not commensurable” is a result, not a missing cell. The synthesis must state which boundary blocks comparison.

Template fit checks

These checks exercise the method without making source findings:

  • Product: Letta Code’s proposed memory question maps to the continuity-layer matrix, which prevents conversation history, recovery state, documentary intent, repository state, and agent-managed long-term memory from collapsing into “memory”.
  • Evaluation: Inspect AI’s proposed reproducibility question maps to the evaluation contract, which separates task definition, execution environment, evidence log, scoring, aggregation, and optimizer feedback.
  • Protocol: MCP’s proposed boundary question maps to the protocol authority table, which keeps normative requirements, optional capabilities, reference behaviour, and application policy separate.

No source was read to perform these fit checks; the examples come from the accepted work-item questions and test template coverage only.

Licence and quotation

Record the licence from a file at the pinned commit; if it cannot be established, write “not established”. Quote sparingly and cite the exact pinned artifact. The corpus extracts mechanisms and warnings, not source code or specification prose for reuse in rungs modules.