> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/4.-the-4-stage-verification-engine/4.1-structural-gateway.md).

# 4.1 Structural Gateway

Stage 1 converts an untrusted workbench submission into a structurally admissible, sanitized candidate. It establishes assignment identity, representation, and replay status before expensive evaluation. Admission is not certification of factual accuracy, human authorship, or entitlement to payment.

**Implementation status:** the current /annotate interface simulates the audit with timers. It does not call the production gateway specified here. These are engineering requirements for that integration. The registry already enforces uniqueness of supplied hashes, but cannot inspect remote payloads.

## Intake architecture

```
/annotate: decision + ratings + rationale + optional Golden Patch
                           |
                  authenticated request
                           v
       bounded volatile buffer -> strict UTF-8 / JSON parser
                           |
       assignment authorization + pinned schema validation
                           |
       field-aware PII / credential detection and redaction
                           |
       semantic checks + post-redaction schema validation
                           |
       deterministic encoding -> content SHA-256
                           |
       atomic reservation + sanitized record + durable outbox
                           v
                  Stage 1 receipt -> Stage 2
```

The browser is untrusted even when it runs the official UI. Disabled buttons and displayed wallet addresses are not server authorization. Resolve the campaign, task version, contributor, response identifiers, rubric version, and payout recipient from authenticated assignment state. Reject substitutions rather than allowing submitted values to redefine a payable assignment.

Bind each request to an assignment identifier and an operation identifier for retries. A campaign allowing ties needs an explicit tie outcome; it must not fabricate a chosen/rejected pair. Golden Patch eligibility and bonus amounts are campaign decisions, not arbitrary fields a contributor can override.

Set separate limits on request bytes, decompressed bytes, nesting depth, array cardinality, text length, processing time, and concurrent requests. Reject unsupported compression and malformed encoding. Authenticate before expensive parsing or inference. Configure proxies, tracing, crash reports, and middleware to exclude raw bodies: sanitization cannot undo an upstream log write.

Before sanitization, the service may process bounded transient memory but must not persist raw bodies in queues, temporary files, provider requests, or archives. Strict no-persistence operation also requires host policies for swap and crash dumps. Application code alone cannot guarantee that process memory never reaches disk.

## Deterministic schema enforcement

Pin the input and annotation schema digests in the campaign. A mutable schema URI is a locator, not a version. Permit only approved local references or controlled registry resolution; arbitrary remote schema fetching can expose internal services.

The [JSON Schema Draft 2020-12 specification](https://json-schema.org/draft/2020-12) defines the type and validation vocabulary used in [Custom Schemas & Rubrics](/3.-enterprise-and-frontier-ai-labs/3.2-custom-schemas.md). Apply an actual validator, followed by cross-field application checks. Successful JSON parsing is only the first step.

| Layer               | Enforcement                                                       | Failure handling                        |
| ------------------- | ----------------------------------------------------------------- | --------------------------------------- |
| Parsing             | Unique keys, valid UTF-8, finite numbers, bounded nesting         | Reject malformed input                  |
| Required attributes | Assignment, decision, ratings, rationale, schema version          | Report safe field-path errors           |
| Nested rubric       | Every required dimension is an integer from 1 through 5           | Reject missing or invalid ratings       |
| Text bounds         | Campaign-defined minimum and maximum lengths                      | Request correction                      |
| Unknown properties  | Reject extra properties in each controlled object                 | Reject unrecognized fields              |
| Patch mode          | Require correction text when correction mode is selected          | Reject contradictory state              |
| Identity            | Task and response versions match coordinator state                | Reject stale or unauthorized submission |
| Semantics           | Outcome, patch, and selected/rejected response IDs are consistent | Reject or escalate                      |

Reject duplicate JSON object keys rather than accepting whichever value a parser encounters last. Disable silent coercion, implicit defaults, and extra-property removal unless the campaign explicitly versions those transformations. A string containing a digit is not an integer rating.

Character bounds and transport byte bounds are different constraints. Multilingual strings and escaped characters can satisfy one while exceeding the other. Validate both. The authoritative prompt and candidate responses should already be sanitized before distribution; a returned judgment does not authorize replacing those enterprise inputs.

After redaction, repeat schema validation and semantic checks. If a transformation changes an executable answer or destroys evidence required by the rubric, quarantine it for correction. Never certify one representation and archive a different one under the old receipt.

## Content hashes and unambiguous encoding

The conceptual formula for a decisive pairwise comparison is:

```
payload_hash = SHA-256(prompt + response_chosen + response_rejected + patch)
```

Here addition must denote a defined encoding, not plain concatenation. Fields ("ab", "c") and ("a", "bc") otherwise yield identical bytes without breaking SHA-256. Length prefixes solve that boundary ambiguity.

The following complete reference profile uses UTF-8 and eight-byte big-endian byte lengths. It hashes already sanitized strings, preserves whitespace and Unicode exactly, and represents absence of a patch as an empty string. This proposed profile is not an already deployed gateway standard.

```python
import hashlib

DOMAIN = b"ARCV:pairwise-content:v1\x00"

def payload_hash(
    prompt: str,
    response_chosen: str,
    response_rejected: str,
    patch: str = "",
) -> str:
    digest = hashlib.sha256(DOMAIN)
    for value in (prompt, response_chosen, response_rejected, patch):
        if not isinstance(value, str):
            raise TypeError("Content fields must be strings")
        encoded = value.encode("utf-8", errors="strict")
        if len(encoded) >= 2**64:
            raise ValueError("Field exceeds encoding limit")
        digest.update(len(encoded).to_bytes(8, "big"))
        digest.update(encoded)
    return digest.hexdigest()

if __name__ == "__main__":
    assert payload_hash("ab", "c", "") != payload_hash("a", "bc", "")
    assert payload_hash("task", "better", "worse") == payload_hash(
        "task", "better", "worse", ""
    )
    assert payload_hash("task", "better", "worse") != payload_hash(
        "task", "worse", "better"
    )
    print(payload_hash("task", "better", "worse"))
```

Multi-turn tasks need a separate profile committing to the complete system, user, and tool context. Ties need canonical A/B order and an explicit tie label under a distinct domain. Do not reuse a decisive-pair encoding with an arbitrary winner.

| Identity              | Purpose                        | Commitment                                        |
| --------------------- | ------------------------------ | ------------------------------------------------- |
| Request identity      | Reconcile transport retries    | Authenticated assignment and operation ID         |
| Content fingerprint   | Screen exact answer duplicates | Declared sanitized content fields                 |
| Final artifact digest | Verify archived data           | Exact stored bytes including envelope/compression |

The four-field fingerprint excludes ratings and justification. Those fields must still be committed by the complete accepted record and its receipt. Changing a rationale invalidates prior evaluation evidence even if the content fingerprint stays unchanged.

For canonical JSON, use a declared conforming profile such as [RFC 8785 JCS](https://www.rfc-editor.org/rfc/rfc8785.html). Sorted keys alone are not full JCS, which also specifies primitive serialization. JCS and the length-prefixed encoding above are different formats and cannot share interchangeable digests.

## Recent-submission pool and replay control

The recent pool is an application admission index, not the Arc transaction mempool. Its low-latency lookup reduces duplicate evaluation work, but correctness requires a durable atomic reservation. Separate request identity from a campaign-scoped content policy: independently assigned reviewers may legitimately inspect identical responses.

```
BEGIN TRANSACTION
  read assignment + operation identity
  identical authenticated retry -> return original receipt/status
  same operation, conflicting body -> IDEMPOTENCY_CONFLICT
  enforce unique applicable content reservation
  insert sanitized candidate
  insert evaluation job into transactional outbox
COMMIT
```

A separate existence check followed by insertion permits races. Use a uniqueness constraint or atomic compare-and-set, and commit the candidate with its outbox entry. At-least-once job delivery must produce one logical admission and one payment authorization.

Transient leases can expire abandoned work; durable assignment/payment records cannot disappear with cache eviction. Reconcile ownership before releasing a reservation. Content-level duplicate detection is distinct from semantic similarity: inconsequential whitespace can change an exact hash, while the same short correct answer can legitimately recur.

The current contract's mapping is `usedHashes(bytes32)`, not `registeredHashes`. It rejects repeated supplied digests globally within a registry instance. A second mapping, `usedArweaveTxIds(bytes32)`, reserves Keccak-256 of the storage identifier. Both reservations remain after rejection. The gateway must reconcile these on-chain constraints with pending off-chain operations.

## Heuristic PII and credential sanitization

Inspect parsed fields, not raw JSON syntax. Use bounded patterns and parser-backed validation. Regexes are candidate detectors; they cannot establish a universal zero-PII guarantee.

| Category                    | Detection approach                                                   | Canonical replacement          |
| --------------------------- | -------------------------------------------------------------------- | ------------------------------ |
| Email addresses             | Address candidates, field context, internationalized-form handling   | \[REDACTED\_PII:EMAIL]         |
| Phone numbers               | Candidate patterns and regional number validation                    | \[REDACTED\_PII:PHONE]         |
| IPv4                        | Dotted-decimal candidates validated by an IP parser                  | \[REDACTED\_PII:IPV4]          |
| IPv6                        | Compressed/full candidates validated by an IP parser                 | \[REDACTED\_PII:IPV6]          |
| Social security identifiers | Jurisdiction-specific formats and exclusions                         | \[REDACTED\_PII:NATIONAL\_ID]  |
| Other national IDs          | Country/type rules and checksums where defined                       | \[REDACTED\_PII:NATIONAL\_ID]  |
| API secrets                 | Prefix families including sk- and ghp\_, secret field names, context | \[REDACTED\_PII:API\_KEY]      |
| Bearer tokens               | Authorization-header and embedded credential contexts                | \[REDACTED\_PII:BEARER\_TOKEN] |

There is no single valid national-ID pattern. Prefix matching misses unrecognized credentials, while phone-like values can be dates or task IDs. Networking examples can legitimately contain IP literals. Campaign rules must distinguish authorized task material from information requiring removal; uncertainty goes to quarantine rather than an invented clean result.

Use reviewed expressions without catastrophic backtracking. Merge overlapping spans deterministically, prioritize credential matches over generic strings, and replace full spans. The sanitizer should be idempotent:

```
sanitize(sanitize(x)) = sanitize(x)
```

Contributor-supplied redaction tags are not trusted scrubber receipts. Revalidate transformed fields and preserve only category counts, safe field paths, policy versions, and disposition. Never copy matched secrets into exceptions, metrics labels, permanent lineage, or a public reversible replacement map. Detected credentials require a confidential remediation path; removing a token from text does not revoke it.

## Receipts, errors, and release criteria

The Stage 1 receipt binds assignment identity, sanitized artifact digest, encoding version, schema digest, sanitizer version, operation identity, and disposition. Authenticate it as service evidence. A different artifact must not reuse the same receipt.

Structural errors stop progression; ambiguous redactions cause quarantine; infrastructure failures produce retryable holds. None automatically implies malicious behavior. Test duplicate-key parsing, concurrent replay, cache eviction, conflicting retries, nested missing ratings, decompression limits, sanitizer timeouts, overlapping matches, and redaction-induced invalid code before enabling the service.

The output is a sanitized candidate and a receipt, not a payout instruction. Continue with [Stage 2: Interaction Telemetry](/4.-the-4-stage-verification-engine/4.2-interaction-telemetry.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/4.-the-4-stage-verification-engine/4.1-structural-gateway.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
