> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/5.-smart-contracts-and-arc-network-settlement/5.2-hash-registry-and-replay-defense.md).

# 5.2 Hash Registry & Replay Defense

ARCV separates provenance registration from payment. Registration reserves identifiers and binds a supplied content commitment to a campaign and recipient. Settlement consumes the Pending state of that registered record exactly once. These mechanisms address different replay paths and must not be conflated with proof of content truth or global cross-chain uniqueness.

## The actual mapping and ABI

The implementation declares:

```solidity
mapping(bytes32 => bool) public usedHashes;
mapping(bytes32 => bool) public usedArweaveTxIds;
```

The generated read functions are usedHashes(bytes32) and usedArweaveTxIds(bytes32). There is no registeredHashes mapping or getter. Calling a proposed registeredHashes selector against this contract does not access an alias; clients must use the compiled ABI.

The digest reservation is set during submitDataset, before validation or payment. Consequently, usedHashes returning true means “reserved by a registered submission,” not “settled,” “approved,” or “available from Arweave.” A Pending or Rejected record retains the same reservation as a Paid record.

The storage-identifier mapping uses Keccak-256 over the UTF-8/ASCII identifier string. That is distinct from the SHA-256 commitment to payload bytes:

```
payload_digest = SHA-256(exact committed artifact bytes)
storage_key    = Keccak-256(bytes(arweave_tx_id))
```

The contract does not calculate payload\_digest. It accepts an authorized bundler's nonzero bytes32. Cryptographic assurance therefore requires a consumer to independently retrieve and hash the intended artifact.

## Registration algorithm and failure precedence

The submitDataset call checks authorization before entering its body. Body checks then execute in source order:

| Order | Condition                                                | Revert                           |
| ----- | -------------------------------------------------------- | -------------------------------- |
| 1     | Caller has BUNDLER\_ROLE                                 | AccessControlUnauthorizedAccount |
| 2     | Bounty sponsor is nonzero                                | UnknownBounty                    |
| 3     | Recipient is nonzero and not the registry                | InvalidAddress                   |
| 4     | Payload digest is nonzero; identifier passes shape check | InvalidPayload                   |
| 5     | usedHashes\[digest] is false                             | DuplicateHash                    |
| 6     | usedArweaveTxIds\[storage\_key] is false                 | DuplicateArweaveTxId             |
| 7     | pendingCount + paidCount is below targetCount            | CapacityExhausted                |

A request violating several conditions reports the first reached failure. Do not assume a duplicate error will be returned for an unknown bounty or unauthorized caller.

After checks succeed, the function reserves both identifiers, increments pendingCount and datasetCount, writes the Dataset tuple, and emits DatasetSubmitted. The transaction has no external call. If any later checked operation reverts, earlier writes are reverted as well.

The storage identifier must be exactly 43 bytes from A–Z, a–z, 0–9, underscore, and hyphen. This tests shape, not archival existence, finality, or strict canonical base64url decoding. In particular, it does not validate all unused trailing encoding bits of a 32-byte identifier. An ingestion service should validate the selected archive protocol's canonical representation and existence before invoking the contract.

Identifier strings are case-sensitive. Do not lowercase a transaction ID to “normalize” it. Different encodings or aliases can create different mapping keys even if an external service resolves them similarly; protocol-specific normalization is an off-chain responsibility.

## Reservation and payment idempotency

```
new payload hash + new storage ID
              |
              v
          Pending record
          /            \
       Paid           Rejected
        |                 |
  both reservations remain true
        |                 |
  no second payment   no later payment
```

The hash and storage-ID reservations never clear. Rejection releases funded capacity but preserves both identities. Corrected work therefore needs a genuinely revised artifact and a new valid archival identifier. Adding meaningless padding solely to evade rejection is not a legitimate correction.

Payment replay is blocked separately by status. validateAndDisburse requires Pending and moves the record to Paid before the native-value call. A second successful payment for the same dataset cannot occur under those transitions. If the recipient rejects value, the entire transaction reverts and the record remains Pending, so an operational retry is permitted.

| Attempt                                                 | Implemented result                       |
| ------------------------------------------------------- | ---------------------------------------- |
| Same digest in another campaign                         | DuplicateHash                            |
| Same digest for a different recipient                   | DuplicateHash                            |
| New digest with existing storage ID                     | DuplicateArweaveTxId                     |
| Previously rejected digest with new storage ID          | DuplicateHash                            |
| Pay a Paid or Rejected record                           | DatasetNotPending                        |
| Retry a transaction that reverted on recipient transfer | May succeed while record remains Pending |
| Equivalent content with new bytes and new storage ID    | Not detected by these mappings           |

The namespace is global across parallel bounties within one registry. Two transactions attempting the same digest are serialized by EVM execution: the first successful registration reserves it, and the later one reverts. This is not a race-prone off-chain check-then-insert process.

The stronger global restriction also limits legitimate reuse. A public benchmark artifact cannot be registered repeatedly for separate rewards under the identical digest. The system must define payable identities and consensus assignments coherently rather than discover this restriction after upload.

## Byte identity and canonicalization

Hash the exact representation declared by the archival profile. Whitespace, JSON property order, line endings, Unicode representation, compression headers, and encryption nonces can change byte identity without changing apparent semantics.

[Stage 1](/4.-the-4-stage-verification-engine/4.1-structural-gateway.md) defines a proposed content fingerprint for early duplicate screening. [Stage 4](/4.-the-4-stage-verification-engine/4.4-settlement-and-disbursement.md) distinguishes record hashes, Merkle roots, plaintext batch digests, and stored-artifact digests. The registry's dataSha256 should bind the declared retrievable artifact; it must not ambiguously alternate between those meanings.

The following complete standard-library command verifies a local artifact against an independently trusted SHA-256 value. It processes bounded chunks and exits unsuccessfully on mismatch. Save it as verify\_digest.py and supply the file path and the actual expected digest as command-line arguments.

```python
import argparse
import hashlib
import hmac
import re
from pathlib import Path

def digest_file(path: Path) -> str:
    digest = hashlib.sha256()
    with path.open("rb") as stream:
        for chunk in iter(lambda: stream.read(1024 * 1024), b""):
            digest.update(chunk)
    return digest.hexdigest()

def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("artifact", type=Path)
    parser.add_argument("expected_sha256")
    args = parser.parse_args()
    expected = args.expected_sha256.removeprefix("0x")
    if not re.fullmatch(r"[0-9a-fA-F]{64}", expected):
        parser.error("Expected a 32-byte SHA-256 digest")
    actual = digest_file(args.artifact)
    if not hmac.compare_digest(actual, expected.lower()):
        raise SystemExit("SHA-256 mismatch")
    print(actual)

if __name__ == "__main__":
    main()
```

Obtain the expected digest from a verified registry record or a manifest authenticated by that record. A checksum delivered beside an untrusted file by the same attacker does not establish provenance. A matching digest proves equality of bytes to the commitment, not factual accuracy, legal ownership, or quality.

## Arweave-to-Arc lineage

DatasetSubmitted binds datasetId, bountyId, dataSha256, arweaveTxId, and contributor in a public event. The Dataset storage tuple preserves the same identity with mutable status. DatasetPaid then binds the datasetId to the contributor and transferred amount.

```
Verified registry deployment
        |
DatasetSubmitted event + datasets(datasetId)
        |
        +--> SHA-256 commitment
        +--> archival identifier
        +--> campaign and contributor
        |
Retrieved artifact -- local hash equality
        |
Authenticated manifest -- shard digests / optional Merkle root
        |
DatasetPaid event + Paid state -- payment transaction receipt
```

A payment event alone does not include the payload hash or storage identifier. Indexers must join through datasetId, verify the emitting registry, and retain chain ID, block hash, transaction hash, and log index. Global dataset IDs are not unique across different deployments.

If the registered artifact is a manifest, first hash its exact bytes against dataSha256, then follow its authenticated shard commitments. If it is a direct dataset file, verify that file directly. Do not assume a manifest without checking its format/version.

The contract has no Merkle-root field or proof verifier. A root may be indirectly bound inside an authenticated manifest, but it is not independently checked by validateAndDisburse. A future shared-batch payout protocol would need leaf ownership, claim identities, proof rules, and per-claim spent state. One current dataset remains one reward, not a multi-contributor distribution.

## Cross-chain and fork boundaries

A mapping is local to one contract's state on one chain. Deploying the same source elsewhere creates an independent registry; identical digests can be registered there. There is no cross-chain messenger or globally shared spent set.

Chain-aware transaction signing helps distinguish transactions for different chain IDs, but it is not dataset replay protection. Independent chains can accept equivalent logical work through different signed transactions. A fork retaining the same chain ID and compatible state is not resolved by an application pretending the hash namespace is global.

For off-chain authorization and future signature designs, bind at least:

```
authorization domain =
  chain ID
  verifying registry address
  protocol/version identifier
  bounty ID
  assignment or claim identity
  contributor
  payload commitment
  nonce and validity boundary
```

This is a design requirement for a future signed authorization scheme, not an EIP-712 verifier currently exposed by the registry. Domains prevent accidental cross-context use when actually enforced; they do not establish which fork is economically canonical.

Indexers need an explicit network finality and reorganization policy. A log observed before finality can disappear or be replaced. Payment reconciliation should verify the receipt, canonical block identity, and current dataset state before marking final success. Do not permanently erase off-chain replay records merely because one RPC endpoint temporarily omits a transaction.

## Operator trust and admission requirements

Because an authorized bundler controls submitted commitments, it can reserve a hash before a legitimate registration or occupy campaign capacity with false metadata. Persistent reservations prevent duplicate reuse but provide no built-in appeal or reset. Secure assignment binding, archival verification, role custody, and off-chain admission queues are essential.

Do not treat digest uniqueness as proof of unique human work: two different envelopes can conceal duplicated answers. Conversely, do not expose hashes of low-entropy private fields as an anonymization mechanism; an observer can guess candidate values. The archive profile and privacy controls must operate before the commitment becomes public.

The implemented defense is precise: identifier uniqueness within one registry plus terminal payment status per dataset. Claims of universal cross-chain replay immunity, semantic deduplication, or verified Merkle claims would exceed the contract. The [test matrix](/5.-smart-contracts-and-arc-network-settlement/5.3-security-and-test-suite.md) identifies which of these transitions have executable coverage.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/5.-smart-contracts-and-arc-network-settlement/5.2-hash-registry-and-replay-defense.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
