<?xml version="1.0" encoding="UTF-8"?>
<rfc category="info" docName="draft-watts-research-artifact-provenance-00" ipr="trust200902" submissionType="IETF" version="3">
  <front>
    <title abbrev="Research Artifact Provenance">Research Artifact Provenance and Corrigible Scientific State</title>
    <seriesInfo name="Internet-Draft" value="draft-watts-research-artifact-provenance-00"/>
    <author fullname="Deonte Watts" initials="D." surname="Watts"><organization>Independent Researcher</organization><address><postal><city>San Francisco</city><region>CA</region><country>US</country></postal><email>deonte@goodshyt.fun</email><uri>https://orcid.org/0009-0005-8586-3650</uri></address></author>
    <date year="2026" month="October" day="8"/><area>Applications and Real-Time</area><workgroup>Individual Submission</workgroup>
    <abstract><t>This document describes a protocol-neutral pattern for content-addressed research artifacts, append-only evidence history, explicit epistemic classes, and corrigible active scientific state. It separates byte integrity from numerical correctness, empirical adequacy, mechanism identification, and release authority.</t></abstract>
    <note title="Archive Status"><t>This RFCXML source is an archive working draft and is not represented as submitted, adopted, or endorsed by the IETF.</t></note>
  </front>
  <middle>
    <section><name>Motivation</name><t>Research systems increasingly combine source code, datasets, model-generated outputs, experimental measurements, preregistrations, manuscripts, reviews, and release events. Hashes can preserve artifact identity, but integrity alone cannot determine scientific authority. A interoperable provenance profile therefore needs explicit semantics for both immutable history and corrigible conclusions.</t></section>
    <section><name>Artifact Identity</name>
      <t>Each artifact SHOULD have a stable identifier and SHOULD bind a cryptographic digest, media type, byte length, schema/profile version, creation or capture time, and provenance reference. The artifact identifier MUST NOT be interpreted as a truth or quality score.</t>
      <t>When a signed or hashed JSON representation is used, profiles SHOULD use a deterministic canonicalization scheme such as RFC 8785 or register another canonical representation.</t>
    </section>
    <section><name>Evidence and Epistemic Classes</name>
      <t>Profiles SHOULD distinguish at least conceptual design, mathematical derivation, model-generated data, simulation output, empirical observation, independent physical measurement, interpretation, and public communication. A numerical example MUST NOT be labeled as an experiment unless the profile explicitly defines it as a numerical experiment.</t>
      <t>Suggested claim labels include ESTABLISHED, DERIVED, CONJECTURE, ANOMALY, and POTENTIALLY NOVEL. Suggested evidence tiers may distinguish empirical, derived, rigorous conjectural, and speculative material. These labels require profile definitions and MUST NOT be inferred from writing style.</t>
    </section>
    <section><name>Append-Only History and Corrigible State</name>
      <t>Evidence and governance events SHOULD be append-only. Active conclusions, exclusions, and release decisions MAY be superseded, contested, reversed, or retracted through later events. Immutable history MUST NOT be interpreted as requiring monotone scientific conclusions.</t>
    </section>
    <section><name>Contradictions</name>
      <t>Conflicting evidence SHOULD be retained as separate records. A system MUST NOT silently average incompatible claims into a single consensus value. A contradiction record SHOULD identify the competing sources, the crux, and a decisive test or review condition.</t>
    </section>
    <section><name>Reproducibility</name>
      <t>A reproducibility record SHOULD identify code and environment digests, input dataset identifiers, random seeds or stream derivation when relevant, parameters, tolerances, expected outputs, and actual result. A passing reproduction demonstrates only the declared test.</t>
    </section>
    <section><name>Publication and Release</name>
      <t>Release is a governance action distinct from scientific status. A release profile SHOULD include contributor verification, third-party rights review, privacy review, venue-specific integrity requirements, and a license decision. A proprietary research archive can therefore contain public-release candidates without implicitly licensing them.</t>
    </section>
    <section><name>Security Considerations</name>
      <t>Threats include digest substitution, canonicalization mismatch, event deletion, rollback to stale state, provenance forgery, dependency confusion, and unauthorized release. Cryptographic integrity does not prevent flawed experiments or misleading interpretation.</t>
    </section>
    <section><name>Privacy Considerations</name>
      <t>Research provenance can contain personal, medical, financial, location, or organizational information. Profiles SHOULD minimize disclosure, compartmentalize sensitive source material, and support redacted or opaque references when public verification does not require the underlying record.</t>
    </section>
    <section><name>IANA Considerations</name><t>This document has no IANA actions.</t></section>
  </middle>
  <back><references><name>References</name><reference anchor="RFC8785" target="https://www.rfc-editor.org/rfc/rfc8785.html"><front><title>JSON Canonicalization Scheme (JCS)</title><author fullname="A. Rundgren"/><author fullname="B. Jordan"/><author fullname="S. Erdtman"/><date year="2020"/></front><seriesInfo name="RFC" value="8785"/></reference><reference anchor="PROV" target="https://www.w3.org/TR/prov-dm/"><front><title>PROV-DM: The PROV Data Model</title><author fullname="W3C Provenance Working Group"/><date year="2013"/></front></reference></references></back>
</rfc>
