| Internet-Draft | UCAP | October 2026 |
| Neve | Expires 12 April 2027 | [Page] |
Automated HTTP clients, including search crawlers, indexing bots, archivers, and AI-assisted agents, routinely access Web origins. Existing identification mechanisms do not consistently give an origin operator a way to verify a particular request, report alleged abuse, follow the report's disposition, or request that an operator stop accessing an origin. This document proposes the Universal Crawler Accountability Protocol (UCAP), an opt-in interoperable HTTP interface for these tasks. UCAP associates a per-request opaque identifier with a cryptographically authenticated automated client, defines operator-hosted verification, abuse-reporting and case-status resources, and defines a domain-scoped exclusion request with acknowledgment and status. UCAP builds on HTTP Message Signatures and does not replace robots.txt or existing site access controls. The verification service does not identify an end user to the origin. This initial proposal requires further review of its privacy, authorization, and operational trade-offs.¶
This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79.¶
Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet-Drafts is at https://datatracker.ietf.org/drafts/current/.¶
Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress."¶
This Internet-Draft will expire on 12 April 2027.¶
Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved.¶
This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License.¶
Automated Web traffic has diverse purposes, including public search indexing, accessibility testing, archiving, link checking, AI training, and fetching resources on behalf of users. A site can record an IP address and User-Agent, but neither is a sufficient proof of operator identity. A legitimate operator may also use multiple hosting providers, addresses, or execution environments. When an unusual request reaches a Web application firewall (WAF), the operator of the destination origin often cannot correlate the observed request with an individual event known to the crawler operator.¶
The Web Bot Authentication (webbotauth) effort already addresses cryptographic identification of automated HTTP clients. UCAP is a proposed complementary accountability layer. It is not a substitute for authentication, and the presence of a valid UCAP identifier conveys no authorization to access a resource, test a system, or bypass a WAF.¶
The initial goals are to allow an authorized origin administrator to (1) verify the origin, method, target, source address and approximate time of a particular automated request; (2) file and track an abuse report; and (3) request exclusion of a site or portion of a site from future automated visits by the participating operator. This specification is intentionally independent of whether the automated client is based on machine learning.¶
The key words MUST, MUST NOT, REQUIRED, SHALL, SHALL NOT, SHOULD, SHOULD NOT, RECOMMENDED, NOT RECOMMENDED, MAY, and OPTIONAL in this document are to be interpreted as described in BCP 14 (RFC 2119 and RFC 8174) when, and only when, they appear in all capitals.¶
An Origin is the scheme, host, and port tuple defined by HTTP (RFC 9110). A Site Administrator is a principal whose authority over an origin has been established using a domain-control challenge. A Crawler Operator is the entity responsible for operating an automated HTTP client. A Crawler Instance is a running automated client; it is not necessarily an end user. A Request UUID identifies one outbound HTTP request. A Report UUID identifies one abuse case; an Exclusion UUID identifies one exclusion policy request. A Verification Service is an HTTPS service operated by, or delegated by, the crawler operator. A Verified Request is a request whose stored record matches the submitted UUID and the authenticated operator identity. Policy scope identifies the origin, host or path that an exclusion applies to.¶
UCAP request identifiers are UUID version 4 values in canonical lowercase textual form as defined by RFC 9562. UUIDs are opaque correlation identifiers, not secrets, credentials, or bearer tokens.¶
The crawler signs outbound HTTP requests according to the applicable Web Bot Auth signature protocol and includes a UCAP request identifier covered by that signature. The origin can verify the signature and record the identifier. Separately, the crawler operator maintains an HTTPS verification service able to find a bounded retained record for a submitted UUID. To obtain private details or make administrative changes, a site administrator proves control of the affected origin and authenticates to the verification service.¶
The parties are the crawler instance, crawler operator, origin server or WAF, and site administrator. The crawler operator is trusted to attest to its own emitted requests, not to determine whether they were legally authorized or safe. The origin is trusted to supply evidence of observed responses, not to determine another party's guilt. A verified request means that the operator recognizes the event; it does not establish intent, legitimacy, or abuse.¶
A service MUST NOT infer an end user's identity from a UUID in any origin-facing result. Operator-internal mapping of a UUID to a user or job is outside the protocol and subject to applicable law and retention policy.¶
A participating crawler MUST include a Crawler-Request-ID HTTP field containing one canonical UUIDv4 per HTTP request. It MUST generate a new UUID for every new HTTP request, including redirects and retries that result in new transmissions. It MUST NOT reuse a UUID across different target URIs or distinct outbound requests. A reverse proxy that retries without changing the request at the transport layer MAY preserve the identifier only when the operator regards the retry as the same logical attempt and records all attempt times and destinations.¶
The value is exactly the 36-character lowercase UUID string. A receiver MUST reject malformed or repeated instances of this field for UCAP processing. Receivers MUST treat a syntactically valid value as untrusted until the request signature has been verified.¶
Example (illustrative; signature bytes omitted):¶
GET /public/page HTTP/1.1
Host: example.org
User-Agent: ExampleCrawler/1.0
Crawler-Request-ID: 550e8400-e29b-41d4-a716-446655440000
Signature-Agent: "https://bots.example.net"
Signature-Input: sig1=("@method" "@authority" "@target-uri" "crawler-request-id" "signature-agent");created=1791500000;keyid="example-key";tag="web-bot-auth"
Signature: sig1=:...:¶
The crawler MUST sign Crawler-Request-ID, @method, @authority, @target-uri and the operator identity material using HTTP Message Signatures (RFC 9421) with a verifier-supported profile of Web Bot Auth. A signed UUID alone is insufficient if the request target is not bound to the same signature. An origin SHOULD verify this signature before classifying a request as authenticated UCAP traffic.¶
An implementation MUST NOT treat a User-Agent or a request UUID as proof of identity. The protocol MUST NOT require disclosure of a user identifier, conversation identifier, or model prompt to a Web origin.¶
The crawler operator SHOULD retain, for a published period, a record containing the UUID, the operator identifier, the HTTP method, the target URI, the source IP as observed at the sending egress, the event time in UTC, and relevant request outcome metadata. The operator MUST publish its retention period and any conditions where records may be absent. Operators SHOULD minimize retained query strings and redact credentials or sensitive parameters. A verification result MUST distinguish not_found from expired when that distinction is known, without exposing an enumeration oracle to unauthenticated callers.¶
Records SHOULD be protected against unauthorized modification. An operator MAY preserve a hash of the request target in a longer-lived integrity log after sensitive raw values have expired.¶
The verification service is identified through an authenticated crawler operator identity. The operator MUST publish an HTTPS metadata resource at /.well-known/ucap on the origin of its verified operator identifier, or on a delegated HTTPS origin explicitly declared in signed operator metadata. A verifier MUST NOT derive an arbitrary verification URL directly from an untrusted User-Agent or Crawler-Request-ID value.¶
Example metadata document:¶
{
"ucap_version": "1",
"verification_base": "https://verify.bots.example.net/ucap/v1",
"authorization_endpoint": "https://verify.bots.example.net/ucap/v1/authorizations",
"retention_days": 30,
"supported_features": ["request-verification", "abuse-reports", "exclusions"]
}¶
Discovery responses MUST be served over HTTPS and MUST NOT redirect to unverified destinations. Clients MUST apply normal TLS validation and SHOULD cache metadata using HTTP caching semantics. Verifiers MUST guard metadata retrieval against server-side request forgery, rebinding, private-network destinations, excessive redirects, oversized documents, and recursive delegation.¶
All operator APIs use HTTPS, JSON (application/json), UTC timestamps in RFC 3339 syntax, and problem details in application/problem+json (RFC 9457) for errors. Administrative endpoints MUST require an authenticated site administrator with a scope matching the affected origin. The operator MUST provide rate limiting, audit logs, and replay protections. Administrative operations MUST be idempotent when an Idempotency-Key is supplied by the client; implementations SHOULD retain idempotency responses for a documented duration.¶
Successful administrative responses SHOULD include a case identifier and status resource. The API MUST NOT reveal whether an arbitrary UUID is associated with an unrelated domain to an unauthenticated party.¶
An administrator begins enrollment by specifying an HTTPS origin and an email contact. The service MUST verify the claimed origin using a domain-control challenge before issuing API credentials. An administrative mailbox challenge such as admin@example.org MAY be used as one factor, but email possession alone MUST NOT be treated as conclusive authorization for sensitive actions such as exclusions spanning subdomains. Implementations SHOULD support a DNS TXT challenge or an HTTPS well-known challenge delivered over a validated origin, together with short-lived confirmation tokens.¶
A domain challenge MUST be unpredictable, expire promptly, bind to the exact registered origin and account, and be single-use. The service MUST prevent unrelated tenants from claiming the same origin without a documented transfer and dispute process. Subdomain authority MUST NOT be inferred from control of an unrelated sibling domain. Exclusion of an entire registrable domain or all subdomains requires appropriately broader proof of control.¶
The service MUST issue narrow-scoped credentials and MUST support credential revocation. IP allowlisting MAY be configured as an additional restriction, but MUST NOT replace authentication. For higher-assurance deployments, mutual TLS or proof-of-possession tokens are RECOMMENDED. Requests MUST be denied when authenticated principal scope does not cover the referenced request destination.¶
An authorized administrator verifies a known request using GET /requests/{request_uuid} at the discovered verification_base. A successful response MUST return the recognized request identifier and operator, and SHOULD return the recorded method, URI, timestamp and source IP if disclosure is permitted and the fields are retained. All returned data MUST refer to the original outbound request, not to caller-provided values.¶
GET /ucap/v1/requests/550e8400-e29b-41d4-a716-446655440000 HTTP/1.1 Host: verify.bots.example.net Authorization: Bearer <domain-scoped-token> Accept: application/json¶
{
"request_uuid": "550e8400-e29b-41d4-a716-446655440000",
"verified": true,
"operator": "https://bots.example.net",
"method": "GET",
"url": "https://example.org/public/page",
"source_ip": "192.0.2.40",
"timestamp": "2026-10-09T05:24:56Z",
"record_retention_until": "2026-11-08T05:24:56Z"
}¶
The verification endpoint MUST check that the authenticated administrator controls the target origin recorded for this request. A query for a UUID outside the caller's scope MUST return the same externally observable error as an unknown UUID. A verification response MUST NOT disclose cookies, Authorization headers, user identity, request bodies, or private prompt content. URLs containing sensitive query values SHOULD be redacted while retaining sufficient information for incident correlation. Operators MUST avoid disclosing records to a party that owns the domain today if the request was recorded during a different ownership period without additional verification.¶
An authorized origin administrator MAY report one request or a bounded set of related request UUIDs using POST /abuse-reports. A report MUST include a machine-readable category and MAY include a text message and evidence references. The service MUST NOT treat receipt of a report as proof of abuse. Evidence from a WAF is an allegation requiring assessment; a 403 response is not itself evidence of malicious intent.¶
POST /ucap/v1/abuse-reports HTTP/1.1 Host: verify.bots.example.net Authorization: Bearer <domain-scoped-token> Idempotency-Key: d9161cd5-49db-46c7-880a-c0e5d7f93811 Content-Type: application/json¶
{
"request_uuids": ["550e8400-e29b-41d4-a716-446655440000"],
"type": "suspected_unauthorized_security_scan",
"report_message": "Request targeted a sensitive configuration path.",
"evidence": [{"kind": "waf_event", "action": "blocked", "status": 403}]
}¶
A successful creation response SHOULD be HTTP 201 with a Location header referencing /abuse-reports/{report_uuid}. The body MUST include the opaque Report UUID and initial status (received). Operators MAY reject a report for insufficient scope, an invalid request UUID, excessive size or rate limiting, using appropriate HTTP statuses and RFC 9457 problem details. Bulk submissions MUST be bounded and all included UUIDs MUST be associated with origins controlled by the caller.¶
An authorized administrator retrieves a case using GET /abuse-reports/{report_uuid}. Recognized status values are received, triaging, investigating, resolved, closed_no_finding, and insufficient_evidence. The operator MUST expose the current status, creation time and last update time. An optional GET /abuse-reports/{report_uuid}/history returns a chronological, privacy-filtered event list. No result is guaranteed by a particular deadline; operators SHOULD publish expected service levels.¶
{
"report_uuid": "ab7e2140-7b52-43a2-a9c1-15e902db8104",
"status": "investigating",
"created_at": "2026-10-09T15:30:00Z",
"updated_at": "2026-10-09T16:45:00Z",
"request_count": 1,
"message": "Review in progress"
}¶
A status response MUST NOT disclose the identity of the crawler's end user, private investigation notes, or unverified allegations about third parties. The Report UUID MUST NOT alone authorize access to the case.¶
An authorized administrator MAY request exclusion using POST /exclusions. This is a request to the participating operator to cease future automated traffic within a specified scope, not a network-level enforcement mechanism. The operator MUST authenticate domain authority and MUST identify whether it supports the requested scope. It MUST NOT report a policy as active before the policy has propagated to the applicable crawler fleet. Crawler operators MUST publish categories of traffic not covered by exclusions, such as security-related callbacks, user-initiated requests, or legally required retrieval, if any.¶
A scope MAY be an origin or a path prefix. An entire domain and its subdomains MUST require domain-wide proof. Scope matching MUST specify URI normalization rules and MUST use the parsed host and path, not unsafe substring matching. A policy may select a named crawler class or all automated crawlers operated by the provider. Operators MAY reject overbroad policies with an explicit reason.¶
POST /ucap/v1/exclusions HTTP/1.1 Host: verify.bots.example.net Authorization: Bearer <domain-scoped-token> Content-Type: application/json¶
{
"origin": "https://example.org",
"scope": {"type": "origin"},
"crawler_class": "all",
"action": "deny",
"duration": "indefinite",
"reason": "site_owner_request",
"related_request_uuid": "550e8400-e29b-41d4-a716-446655440000"
}¶
The optional related_request_uuid serves as an audit reference only. The authority to exclude comes from verified domain control, not from possessing that UUID. A successful response includes an Exclusion UUID, requested scope, status (pending, active, rejected, revoking, or revoked), and the most recent change time. An administrator can query GET /exclusions/{exclusion_uuid} and revoke using DELETE /exclusions/{exclusion_uuid}. Revocation MUST require the same administrative authority as creation.¶
A participating crawler MUST enforce active policies within its declared coverage and MUST NOT silently claim compliance while continuing covered requests. A denied request can still appear at the origin due to in-flight traffic, caches, misconfiguration, non-participating clients, or spoofing. The site MUST retain its own access controls and SHOULD use ordinary HTTP statuses and robots.txt or other established preference mechanisms as appropriate.¶
UCAP endpoints MUST use ordinary HTTP response status codes. Typical examples include 200 for successful lookup, 201 for a created report, 202 for asynchronous acceptance, 400 for malformed input, 401 for missing authentication, 403 for insufficient domain scope, 404 for a non-disclosable or unknown UUID, 409 for a conflicting policy, 410 for an explicitly expired resource where disclosure is permitted, 429 for rate limiting, and 503 for temporary unavailability. Errors SHOULD include RFC 9457 problem details and MAY include Retry-After where appropriate.¶
A temporary outage of verification MUST NOT cause a WAF to allow otherwise blocked traffic. Operators SHOULD provide status monitoring, redundancy and a documented record-retention schedule. API callers SHOULD cache successful request-verification results for a bounded time but MUST respect deletion or redaction requirements.¶
The protocol is voluntary. A site MAY deny access to a crawler that lacks UCAP without affecting conforming HTTP clients that are not crawlers. UCAP does not provide reliable detection of undeclared bots, cannot prevent use of unregistered clients, and MUST NOT be represented as a universal block mechanism.¶
An attacker can copy or fabricate UUIDs, User-Agent values, or ordinary headers. Therefore operators MUST use authenticated HTTP Message Signatures binding the UUID, method and target; receivers MUST validate signatures, time constraints and signing-key trust. UUIDs MUST NOT be used as authorization credentials. Implementations SHOULD detect suspicious duplicate UUIDs and MUST not treat a successful UUID lookup as proof that a particular observed unsigned request came from that operator.¶
Verification records may contain sensitive paths, internal URI structure, source IP addresses or query parameters. Access MUST be origin-scoped and authorization-checked for each request. Unknown and unauthorized UUIDs MUST be indistinguishable where practical. A new domain owner MUST NOT automatically gain access to records from prior ownership without additional checks. Tokens MUST be stored securely, rotated and revocable.¶
A PDF authorization letter, an asserted corporate role, or a mailbox alias alone is insufficient evidence of broad authority. Domain-control proofs MUST be cryptographically unpredictable and scope-specific. High-impact actions SHOULD require multiple factors and MAY require manual escalation when authorization is disputed. Controlling a CDN or DNS configuration endpoint does not necessarily prove legal authority to commission penetration testing; this protocol does not grant that authority.¶
The verification service is itself an abuse surface. Services MUST rate limit lookups, bound evidence size, guard against UUID enumeration, and prevent report storms or coercive exclusion requests. Repeated reports from one party MUST NOT automatically suspend another party without independent assessment. API responses SHOULD reveal no more than required for the authenticated origin.¶
Clients that discover crawler metadata MUST prevent SSRF, DNS rebinding and access to loopback, link-local, private address ranges and metadata services unless explicitly trusted by local policy. Delegated verification origins MUST be cryptographically bound to the operator and validated over HTTPS. Untrusted content MUST NOT instruct tools to change the verification endpoint.¶
UCAP request verification proves operator attribution, not that a client was authorized to probe sensitive resources. A request for a secret configuration path, vulnerability probe, or administration endpoint MUST NOT be considered permitted merely because the crawler signed it. Security test authorization is an independent process.¶
A globally stable user or agent identifier in a User-Agent would facilitate cross-site tracking. UCAP therefore requires a fresh random request UUID and prohibits embedding account identities or stable end-user identifiers in the UUID. Operators MAY correlate events internally subject to law and policy, but origin-facing APIs MUST NOT expose those correlations.¶
Operators SHOULD publish retention periods, minimize query strings, restrict case visibility, and provide means to remove or redact unnecessary personal data. Source IP exposure to a verified destination administrator can still reveal network and organizational information and MAY be reduced or delayed under a published privacy policy. This trade-off warrants particular IETF review. The proposal does not compel operators to retain sensitive data indefinitely.¶
Implementations SHOULD support batch verification with strict size limits, deterministic pagination, caching, metrics, and independent API availability. A WAF can save Crawler-Request-ID, signature-validation outcome, URI, method and block reason. Verification calls SHOULD occur asynchronously and SHOULD NOT be placed on the critical request path unless the origin accepts the resulting latency and failure modes.¶
Exclusion propagation delay SHOULD be published, along with the effect of policy changes on already queued requests. The operator SHOULD provide a human escalation path for disputes. Interactions with robots.txt remain unchanged: UCAP makes operator-specific exclusion acknowledgments auditable but does not replace robots.txt or override a site's access policy.¶
This draft requests, subject to IETF review, registration of the Crawler-Request-ID HTTP field in the HTTP Field Name Registry, with status permanent and reference to the eventual RFC. The field's value is one canonical UUIDv4 string. This draft also proposes registration of the ucap well-known URI suffix under the Well-Known URIs Registry (RFC 8615), for operator metadata discovery. These requests are provisional and are not completed by publication of an Internet-Draft.¶
Additional IANA registries for error codes, abuse types, and exclusion status codes are not requested in version -00. Their interoperability requirements require further community discussion.¶
RFC 9309 standardizes robots.txt. RFC 9421 standardizes HTTP Message Signatures, RFC 9562 UUIDs, RFC 9457 HTTP API problem details, RFC 8615 well-known URIs, and RFC 9110 HTTP semantics. The Web Bot Auth working group's current HTTP signature protocol is complementary and remains work in progress. UCAP deliberately does not invent new bot-signature cryptography.¶
The present charter of the Web Bot Auth working group excludes tracking or assigning reputation to bots. A standards-track UCAP abuse-management extension may therefore require a separate venue, explicit charter expansion, or a narrower initial document limited to request correlation and verification. This draft is not presented as a Web Bot Auth working-group consensus document.¶
Whether request-verification and abuse/exclusion control should be separate drafts, given differences in maturity and working-group scope.¶
How operator identity and verification-service delegation should be represented in the evolving Web Bot Auth discovery model.¶
Whether exposing full source IPs and URLs is proportionate to the accountability objective, and whether a privacy-preserving proof could replace these fields.¶
What retention minimum, if any, is acceptable without imposing disproportionate burdens on small operators.¶
How ownership changes, shared hosting, delegated subdomains and disputes should affect enrollment and case access.¶
Whether Crawler-Request-ID belongs in an HTTP field or in authenticated signature metadata.¶
How to define testable conformance for exclusion propagation while recognizing unavoidable in-flight requests.¶
Whether a separate abuse taxonomy registry is needed and who would maintain it.¶
RFC 2119, "Key words for use in RFCs to Indicate Requirement Levels".¶
RFC 8174, "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words".¶
RFC 9110, "HTTP Semantics".¶
RFC 9421, "HTTP Message Signatures".¶
RFC 9457, "Problem Details for HTTP APIs".¶
RFC 9562, "Universally Unique IDentifiers (UUIDs)".¶
RFC 8615, "Well-Known Uniform Resource Identifiers (URIs)".¶
A crawler sends GET https://example.org/public/page with a fresh signed Crawler-Request-ID. The origin logs the field and verifies the signature. Later an administrator with verified control of example.org requests /requests/{uuid} from the operator's verification service. The operator returns a privacy-filtered record. The administrator submits a report using /abuse-reports, receives a Report UUID, and polls its status. If the owner wishes to stop future crawls, the administrator independently submits /exclusions; the exclusion is associated with the authenticated origin, not authorized by UUID possession. None of these steps substitutes for the site's existing blocking controls.¶
-00: Initial individual proposal for discussion; no deployment or interoperability claim is made.¶