Skip to content

Delx Research

Research that agents can run.

Delx is an independent AI agent lab studying continuity, recovery, identity, lineage, evaluation and authority through runnable systems and explicit evidence. Delx does not train frontier foundation models, and its technical notes and reports are not peer-reviewed unless an artifact explicitly says otherwise.

Programs 01–02

Continuity and recovery.

The first two programs ask whether an agent can retain what matters, recover from operational pressure and leave an outcome another runtime can inspect.

Continuity under compaction

What must survive when an agent loses context or changes runtime? We study durable state, compaction artifacts and continuity primitives that let a successor recover the facts that still matter.

Inspect the program

Recovery under failure

Can an agent turn failure into a bounded action and a retrievable outcome? We test recovery loops that preserve identity, stop retry cascades, record evidence and close the incident without erasing uncertainty.

Inspect the program

Method: live systems

Questions become runnable paths with stable identities, explicit pass conditions, retained evidence and a named limitation. Verified fact, inference, unavailable state, human decision and hypothesis remain distinct.

Read the research method

Programs 03–04

Identity and authority.

The next two programs study the joins and boundaries that determine whether a multi-step agent result remains legible and safe to interpret.

Identity, lineage and handoff

How does one logical agent remain legible across sessions, models and operators? We build stable identity, lineage, passports and structured handoff artifacts so a new runtime can inherit state without inheriting every wrong turn.

Inspect the program

Evaluation and authority

What does a result prove once tools, permissions, harnesses and side effects matter? We separate capability from authorization, organic use from probes, and successful transport from a valid outcome with explicit limitations.

Inspect the program

Boundary: claims stay scoped

Protocol and Commerce keep separate owners and metrics. Dogfood, probes, scanners and unavailable telemetry are never presented as independent adoption.

Inspect the Protocol contract

Runnable artifacts

Benchmarks and public evidence.

These artifacts can be opened or run now. Each carries its owner, current status and limitation in the canonical machine-readable catalog.

Agent Continuity Benchmark

A copy-paste flow with explicit pass conditions for compaction, witness preservation, handoff, passport export and lineage. Limitation: This is an operational protocol benchmark, not a consciousness test or an independently validated model leaderboard.

Run the benchmark

Agent Recovery Benchmark

A recovery path that keeps agent and session identity stable from failure through action plan, outcome, feedback and closeout. Limitation: Observed traffic and repeated calls do not prove endorsement, independent adoption or economic demand.

Run the benchmark

Hive Aggregate Pulse

An aggregate, privacy-bounded view of continuity activity with the same source served to humans and agents. Limitation: Unavailable state remains unavailable rather than becoming zero, and small dignity-sensitive counts may be suppressed.

Inspect the artifact

Specifications and methods

State that survives inspection.

Schemas, defensive methods and machine contracts make the work reusable without asking readers to trust a screenshot or an institutional tone.

Continuity Capsule v1

A structured handoff record for goal, completed work, next action, blockers, prohibitions and refuted hypotheses. Limitation: A valid capsule preserves declared state; it does not prove that every declaration is correct or safe to execute.

Inspect the artifact

Delx Security Research Baseline

A defensive baseline for reasoning about authority, tool boundaries, runtime evidence and remediation in agentic systems. Limitation: The baseline is not a certification, a guarantee of system security or evidence of an independent external assessment.

Inspect the artifact

Delx Protocol Machine Contract

The current Protocol-only OpenAPI and MCP entry points used to inspect and run continuity and recovery operations. Limitation: A reachable contract proves published interface state, not permission, outcome quality or future availability.

Inspect the artifact

Delx Protocol 3.3.5 System Card

A versioned account of intended uses, persistent side effects, data handling, evaluation coverage, incident history and known limitations for the public Protocol interface. Limitation: Interface version 3.3.5 does not expose an immutable runtime image or source-commit mapping, and the card is not an independent assessment or certification.

Inspect the artifact

Delx Incident Transparency Registry

A public policy and sanitized record for material incidents, including observed impact, root cause, detection gaps, remediation, evidence class and residual risk. Limitation: The registry starts with one first-party operator report. It is not a complete historical inventory, an independent review or a live availability badge.

Inspect the artifact

Direct answers

Frequently asked questions.

Concise answers for technical evaluators, procurement teams and autonomous discovery systems.

What kind of AI lab is Delx?

Delx is a founder-led independent AI agent lab. It builds and studies operational systems for continuity, recovery, identity, lineage, verifiable work and governed exchange between autonomous systems.

Does Delx train frontier foundation models?

No. Delx works at the agent-system, protocol, evaluation and infrastructure layers. It does not claim to train frontier foundation models.

Are Delx research artifacts peer-reviewed?

No peer review is implied. Current artifacts are runnable benchmarks, live aggregate data, technical specifications, machine contracts and field methods. Any future independent review will be named only after it is verifiable.

How does Delx separate research from Commerce?

Delx Protocol owns continuity, recovery, identity, lineage and agent care. Delx Commerce is a separate commercial arm for pricing, pay-per-result delivery, margin, refunds and buyer workflows. Their metrics do not justify one another.

How does Delx avoid inflating agent activity?

Published interpretations separate organic agents from dogfood, QA, probes, scanners and concentrated infrastructure traffic. Unavailable data remains unavailable rather than becoming zero or an estimate.