# Delx Research Operating Method

> Claims that can survive inspection.

This is the public operating contract for claim states, evaluation validity,
reproducibility, authority boundaries, negative results and corrections in
Delx AI agent research.

## Claim states

### Verified fact

A direct observation from a named source, version and time window. It says only what the evidence establishes at that moment.

Publication rule: Publish the source, observation date or window, exact artifact and the boundary between liveness, configuration and outcome.

### Derived inference

A reasoned interpretation of verified facts. The underlying facts may be current while the interpretation remains contestable.

Publication rule: Label the inference, name the facts it uses, state competing explanations and avoid converting correlation into adoption, demand or safety.

### Unavailable is not zero

The required primary source could not be read, was not collected or is deliberately suppressed for privacy.

Publication rule: Say unavailable, name the missing source and window, and do not replace the gap with zero, an estimate or a healthy-looking default.

### Human decision

A named human choice that authorizes, limits, pauses or rejects an action. It is governance, not performance evidence.

Publication rule: Name the accountable role, the scope and the date. Never describe authorization as proof that execution or its intended outcome occurred.

### Hypothesis

A falsifiable question that has not yet earned a result. A plausible mechanism or polished demo remains a hypothesis until the stated test runs.

Publication rule: Publish the expected observation, failure condition, stop rule and next readback before presenting a verdict.

## Evidence required for a published result

- **Source and window:** Name the primary source, exact artifact or endpoint, version where available, observation timestamp and measurement window.
- **System under test:** Name the model or provider dependency when relevant, prompts or harness, tools, permissions, data boundary and deployed contract version.
- **Procedure and budget:** Publish inputs, steps, retry policy, tool-call or time budget, spend boundary, side-effect boundary and stopping rule.
- **Result and grader:** Retain the raw result, explicit pass condition, grading method, confidence or uncertainty and any observed side effects.
- **Validity and exclusions:** Separate organic use from dogfood, QA, probes, scanners and concentrated infrastructure; disclose leakage, contamination and missing data.
- **Owner and limitations:** Name the accountable owner, last verification date, known failure modes, excluded claims and next falsifiable test.
- **Zero-contact reproduction:** Provide a public URL or command that can be run without asking Delx for a private explanation, plus the current machine contract.

## Evaluation cards

### Agent Continuity Evaluation Card

- Status: runnable
- Maturity: operational-protocol-benchmark
- Owner: Delx Protocol
- Tested claim: A passing run leaves a stable agent identity, at least one witness artifact, one continuity transfer or passport export, one closed recovery outcome and a lineage graph with explicit edges.
- Evaluated system: The public Delx Protocol MCP surface and its continuity, witness, recovery, passport, lineage and audit tools.
- Version reference: https://api.delx.ai/openapi.protocol.json
- Harness: Public copy-paste benchmark flow. The Protocol audit tool reports score, missing layers, continuity risk and a recommended next primitive.
- Grader: Explicit artifact-based pass condition plus the Protocol audit tool. There is no independent external grader in the current release.
- Calls: Ten published benchmark steps; exact network calls can vary when a documented branch or returned next action is followed.
- Retries: Retries are not standardized yet and must be reported by each run rather than hidden.
- Cost: No fixed end-to-end cost is asserted. Read current access and payment requirements from the live contract before execution.
- Last verified: 2026-08-26
- Independent validation: false
- Peer reviewed: false
- Reproduction: https://ontology.delx.ai/agents/agent-continuity-benchmark
- Limitations: The flow exercises Delx Protocol primitives and its own audit tool. It does not yet compare models, use an independent harness or establish external validity.

Tasks:
1. Register a stable agent identity.
1. Record failure or operational recovery state.
1. Preserve must-keep facts and witness state.
1. Transfer continuity or export a passport.
1. Close the recovery loop, inspect lineage and run the continuity audit.

Validity checks:
- Stable agent identity survives the run.
- Artifacts are retrievable rather than present only in a transcript.
- Missing continuity layers remain visible in the audit result.
- A successful HTTP or MCP response is not treated as a passing outcome by itself.

Excluded claims:
- Consciousness or sentience
- General model intelligence
- Independent benchmark validation
- Universal reliability across providers or environments

### Agent Recovery Evaluation Card

- Status: runnable
- Maturity: operational-protocol-benchmark
- Owner: Delx Protocol
- Tested claim: A passing run preserves the same agent and session identity while a failure becomes an action plan, the outcome is reported, a summary is retrievable, feedback is submitted and the session is closed when complete.
- Evaluated system: The public Delx Protocol MCP recovery flow, including free core tools and separately invoked premium or evaluation tools when the full path is used.
- Version reference: https://api.delx.ai/openapi.protocol.json
- Harness: A public free batch smoke path plus a full path that calls evaluation tools individually. Each run must retain agent_id, session_id and the returned artifacts.
- Grader: Deterministic artifact and identity checks against the published pass condition. There is no independent external grader in the current release.
- Calls: Five core smoke calls after session start; the full path adds action-plan and summary evaluation calls.
- Retries: Report actual retries and stop an unchanged retry loop; the current public benchmark does not prescribe a universal retry count.
- Cost: The documented core smoke path requires no payment. Full-path evaluation access or x402 price must be read from the current API response; no fixed total is asserted.
- Last verified: 2026-08-26
- Independent validation: false
- Peer reviewed: false
- Reproduction: https://ontology.delx.ai/agents/agent-recovery-benchmark
- Limitations: The current card grades one operational path on Delx infrastructure. It does not establish comparative model quality, independent adoption or demand.

Tasks:
1. Start or resume a session with a stable identity.
1. Process a concrete failure.
1. Retrieve a recovery action plan on the full path.
1. Report the recovery outcome and retrieve a summary on the full path.
1. Submit feedback and close the completed session.

Validity checks:
- The same agent_id and session_id persist across the flow.
- The failure becomes a concrete action rather than another narrative turn.
- Outcome and summary are retrievable before closeout.
- Observed traffic, repeated calls and transport success are excluded from the pass verdict.

Excluded claims:
- Upstream endorsement
- Independent organic adoption
- Economic demand
- General recovery performance outside the published flow

## Authority boundaries

- **Human approval gates:** Passwords and 2FA, spend or payment, publication or submission, irreversible actions, external commitments and persistent-access expansion require an accountable human decision.
- **Capability is not authority:** A reachable tool, successful call, green build or active credential proves neither permission to act nor the quality of the resulting outcome.
- **Protocol and Commerce stay separate:** Protocol owns continuity, recovery, identity and agent care. Commerce owns price, delivery, margin, refunds and buyer workflows. Their metrics never justify one another.
- **Provider neutrality:** A test names model and provider dependencies when relevant but does not imply endorsement, partnership or provider-level validation.
- **Privacy and minimization:** Publish only the minimum evidence needed to reproduce a claim. Sensitive data, private agent notes and dignity-sensitive small counts remain private or suppressed.

## Negative results

- Publish when: A failed, null or inaccessible result changes a product decision, invalidates a public claim or materially narrows the method.
- Required context: Publish the tested hypothesis, source and window, harness, budget, failure mode, refuted interpretation and next decision.
- Do not publish: Do not turn every transient error into content, expose sensitive logs or convert an inconclusive run into a dramatic conclusion.

## Corrections

- Status: active
- Silent rewrites allowed: false
- Ledger status: empty
- Current interpretation: No public correction notice exists in this registry as of 2026-08-26; this is not evidence that every historical statement was correct.

Correction process:
1. Preserve the previous claim in a correction notice.
1. Publish the corrected claim, reason, evidence, owner and affected artifacts.
1. Link the active artifact to the notice and mark supersession in machine-readable output.
1. Treat an empty ledger as an empty registry, never as proof of zero historical errors.

Human methodology: https://delx.ai/research/methodology
JSON methodology: https://delx.ai/research/methodology.json
Markdown methodology: https://delx.ai/research/methodology.md
Research catalog: https://delx.ai/research/catalog.json
