Internet-Draft Research Artifact Provenance October 2026
Watts Expires 11 April 2027 [Page]
Workgroup:
Individual Submission
Internet-Draft:
draft-watts-research-artifact-provenance-00
Published:
Intended Status:
Informational
Expires:
Author:
D. Watts
Independent Researcher

Research Artifact Provenance and Corrigible Scientific State

Abstract

This document describes a protocol-neutral pattern for content-addressed research artifacts, append-only evidence history, explicit epistemic classes, and corrigible active scientific state. It separates byte integrity from numerical correctness, empirical adequacy, mechanism identification, and release authority.

Archive Status

This RFCXML source is an archive working draft and is not represented as submitted, adopted, or endorsed by the IETF.

Status of This Memo

This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.

Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.

Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."

This Internet-Draft will expire on 11 April 2027.

▲

Table of Contents

1. Motivation

Research systems increasingly combine source code, datasets, model-generated outputs, experimental measurements, preregistrations, manuscripts, reviews, and release events. Hashes can preserve artifact identity, but integrity alone cannot determine scientific authority. A interoperable provenance profile therefore needs explicit semantics for both immutable history and corrigible conclusions.

2. Artifact Identity

Each artifact SHOULD have a stable identifier and SHOULD bind a cryptographic digest, media type, byte length, schema/profile version, creation or capture time, and provenance reference. The artifact identifier MUST NOT be interpreted as a truth or quality score.

When a signed or hashed JSON representation is used, profiles SHOULD use a deterministic canonicalization scheme such as RFC 8785 or register another canonical representation.

3. Evidence and Epistemic Classes

Profiles SHOULD distinguish at least conceptual design, mathematical derivation, model-generated data, simulation output, empirical observation, independent physical measurement, interpretation, and public communication. A numerical example MUST NOT be labeled as an experiment unless the profile explicitly defines it as a numerical experiment.

Suggested claim labels include ESTABLISHED, DERIVED, CONJECTURE, ANOMALY, and POTENTIALLY NOVEL. Suggested evidence tiers may distinguish empirical, derived, rigorous conjectural, and speculative material. These labels require profile definitions and MUST NOT be inferred from writing style.

4. Append-Only History and Corrigible State

Evidence and governance events SHOULD be append-only. Active conclusions, exclusions, and release decisions MAY be superseded, contested, reversed, or retracted through later events. Immutable history MUST NOT be interpreted as requiring monotone scientific conclusions.

5. Contradictions

Conflicting evidence SHOULD be retained as separate records. A system MUST NOT silently average incompatible claims into a single consensus value. A contradiction record SHOULD identify the competing sources, the crux, and a decisive test or review condition.

6. Reproducibility

A reproducibility record SHOULD identify code and environment digests, input dataset identifiers, random seeds or stream derivation when relevant, parameters, tolerances, expected outputs, and actual result. A passing reproduction demonstrates only the declared test.

7. Publication and Release

Release is a governance action distinct from scientific status. A release profile SHOULD include contributor verification, third-party rights review, privacy review, venue-specific integrity requirements, and a license decision. A proprietary research archive can therefore contain public-release candidates without implicitly licensing them.

8. Security Considerations

Threats include digest substitution, canonicalization mismatch, event deletion, rollback to stale state, provenance forgery, dependency confusion, and unauthorized release. Cryptographic integrity does not prevent flawed experiments or misleading interpretation.

9. Privacy Considerations

Research provenance can contain personal, medical, financial, location, or organizational information. Profiles SHOULD minimize disclosure, compartmentalize sensitive source material, and support redacted or opaque references when public verification does not require the underlying record.

10. IANA Considerations

This document has no IANA actions.

11. References

[RFC8785]
Rundgren, A., Jordan, B., and S. Erdtman, "JSON Canonicalization Scheme (JCS)", RFC 8785, , <https://www.rfc-editor.org/rfc/rfc8785.html>.
[PROV]
Group, W. P. W., "PROV-DM: The PROV Data Model", , <https://www.w3.org/TR/prov-dm/>.

Author's Address

Deonte Watts
Independent Researcher
San Francisco, CA
United States of America