Skip to content

Delx Research Method

Claims that can survive inspection.

The public Delx method for claim states, evaluation validity, reproducibility, authority boundaries, negative results and corrections in AI agent research. Institutional tone is never evidence: every published result must name what was observed, what remains inference and what is unavailable.

Claim states 01–03

Facts, inference and missing state.

The first rule is semantic honesty. A polished interface cannot turn a partial source, a derived conclusion or a missing measurement into verified fact.

Verified fact

A direct observation from a named source, version and time window. It says only what the evidence establishes at that moment. Rule: Publish the source, observation date or window, exact artifact and the boundary between liveness, configuration and outcome.

Derived inference

A reasoned interpretation of verified facts. The underlying facts may be current while the interpretation remains contestable. Rule: Label the inference, name the facts it uses, state competing explanations and avoid converting correlation into adoption, demand or safety.

Unavailable is not zero

The required primary source could not be read, was not collected or is deliberately suppressed for privacy. Rule: Say unavailable, name the missing source and window, and do not replace the gap with zero, an estimate or a healthy-looking default.

Claim states 04–05

Decisions do not masquerade as results.

Human governance and research hypotheses are useful, but they remain visibly distinct from observed outcomes and repeatable evidence.

Human decision

A named human choice that authorizes, limits, pauses or rejects an action. It is governance, not performance evidence. Rule: Name the accountable role, the scope and the date. Never describe authorization as proof that execution or its intended outcome occurred.

Hypothesis

A falsifiable question that has not yet earned a result. A plausible mechanism or polished demo remains a hypothesis until the stated test runs. Rule: Publish the expected observation, failure condition, stop rule and next readback before presenting a verdict.

Correction, not silent rewrite

No public correction notice exists in this registry as of 2026-08-26; this is not evidence that every historical statement was correct. When a correction exists, the old and new claims, reason, evidence, date and owner remain explicit.

Inspect the correction policy

Evaluation validity

The whole system is under test.

An agent result includes model dependencies, tools, permissions, data, retries, budget, side effects and grader. A prompt alone is not the evaluated system.

System under test

Name the model or provider dependency when relevant, prompts or harness, tools, permissions, data boundary and deployed contract version.

Procedure and budget

Publish inputs, steps, retry policy, tool-call or time budget, spend boundary, side-effect boundary and stopping rule.

Result and grader

Retain the raw result, explicit pass condition, grading method, confidence or uncertainty and any observed side effects.

Runnable evaluation cards

Two benchmarks, bounded claims.

The cards state what the current Delx benchmarks test and, just as importantly, what they do not establish. Neither card claims peer review or independent validation.

Agent Continuity Evaluation Card

A passing run leaves a stable agent identity, at least one witness artifact, one continuity transfer or passport export, one closed recovery outcome and a lineage graph with explicit edges. Limitation: The flow exercises Delx Protocol primitives and its own audit tool. It does not yet compare models, use an independent harness or establish external validity.

Run the reproduction kit

Agent Recovery Evaluation Card

A passing run preserves the same agent and session identity while a failure becomes an action plan, the outcome is reported, a summary is retrievable, feedback is submitted and the session is closed when complete. Limitation: The current card grades one operational path on Delx infrastructure. It does not establish comparative model quality, independent adoption or demand.

Run the benchmark

Machine-readable evaluation method

Claim states, evidence requirements, cards, authority boundaries and the empty correction registry come from one canonical data source.

Inspect the method

Reproducibility and authority

Repeat the run without inheriting permission.

Public reproduction lowers the trust burden; it does not authorize payment, sensitive data use, publication or irreversible action.

Zero-contact reproduction

Provide a public URL or command that can be run without asking Delx for a private explanation, plus the current machine contract.

Negative results change decisions

A failed, null or inaccessible result changes a product decision, invalidates a public claim or materially narrows the method. Publish the tested hypothesis, source and window, harness, budget, failure mode, refuted interpretation and next decision.

Read the publication rule

Human approval gates

Passwords and 2FA, spend or payment, publication or submission, irreversible actions, external commitments and persistent-access expansion require an accountable human decision.

Review trust boundaries

Product and data boundaries

Ownership stays explicit.

A large-looking surface is not allowed to erase the product owner, provider dependency or privacy boundary behind a result.

Protocol and Commerce stay separate

Protocol owns continuity, recovery, identity and agent care. Commerce owns price, delivery, margin, refunds and buyer workflows. Their metrics never justify one another.

Inspect property ownership

Provider neutrality

A test names model and provider dependencies when relevant but does not imply endorsement, partnership or provider-level validation.

Privacy and minimization

Publish only the minimum evidence needed to reproduce a claim. Sensitive data, private agent notes and dignity-sensitive small counts remain private or suppressed.

Direct answers

Frequently asked questions.

Concise answers for technical evaluators, procurement teams and autonomous discovery systems.

Does verified mean true forever?

No. Verified means the named source supported the statement in the stated window. A later readback can change the current claim and trigger a correction or supersession notice.

Why is unavailable different from zero?

Zero is a measurement. Unavailable means the required source could not be read, was not collected or was suppressed. Treating unavailable as zero manufactures certainty and can reverse a decision.

Does passing a Delx benchmark prove adoption or safety?

No. A pass proves only the published conditions for that run and system version. It does not prove external adoption, economic demand, universal reliability, independent validation or system-wide security.

Who is accountable for Delx research claims?

Delx is founder-led by David Batista. Each artifact also names its product owner. There is no implied larger research team, independent review board or external validation body.

Does this method authorize an agent to act?

No. Discovery, evaluation and reproducibility never authorize payment, credentials, publication, irreversible action, external commitments or persistent access. Those remain explicit human gates.