> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/2.-contributor-and-annotator-guide/2.4-reputation-and-gold-traps.md).

# 2.4 Reputation & Gold-Traps

ARCV's intended integrity system combines reviewer calibration, independent task overlap, anomaly detection, and evidence-based adjudication. Its purpose is to improve the reliability of accepted data while distinguishing honest error, uncertain tasks, and deliberate manipulation. Speed, disagreement, or account age alone cannot establish malicious intent.

**Implementation status:** the current workbench displays illustrative reputation and a timed audit sequence. It does not enforce dwell-time thresholds, record a production interaction history, inject hidden benchmarks, maintain strikes, or disqualify wallets. The registry has no contributor staking, slashing, or reputation mapping. The mechanisms below define an intended operating model and its requirements, not an active enforcement service.

## Threat model and trust boundaries

A contributor can make an honest mistake, misunderstand a rubric, copy an answer, automate submissions, or coordinate with other accounts. An enterprise can also supply a defective task, an incorrect reference, or an inconsistent rubric. An integrity system must examine both sides of that boundary rather than assuming every disagreement proves contributor failure.

A wallet identifies a signing key, not a unique person. One operator can control many wallets, and a legitimate team can share network infrastructure. Wallet counts, IP addresses, and browser characteristics are therefore insufficient proofs of independence. A system claiming sybil resistance must state which additional identity, qualification, assignment, and correlation controls it actually enforces.

| Risk                         | Relevant evidence                                                            | Insufficient evidence by itself |
| ---------------------------- | ---------------------------------------------------------------------------- | ------------------------------- |
| Blind or instant submission  | Server-issued assignment timing plus poor task-specific evidence             | A single short interval         |
| Coordinated answer farming   | Repeated correlated outputs and assignment relationships                     | Shared network access           |
| Deliberate poisoning         | Substantiated manipulative content and a repeated or decisive attack pattern | One mistaken preference         |
| Unqualified participation    | Adequate domain-specific benchmark history                                   | A single difficult question     |
| Defective gold reference     | Reproducible counterexample or adjudicated reference error                   | Disagreement alone              |
| Automated superficial review | Combined cadence, quality, and task-evidence anomalies                       | Lack of mouse movement          |

Cryptographic provenance makes a submitted record attributable and tamper-evident under its verification assumptions. It does not prove that the record is true, independently authored, or created by a unique human. Reputation provides additional evidence; it does not replace content validation.

## Dwell-time telemetry and cadence analysis

### What elapsed time can establish

A production coordinator can record assignment issuance and submission receipt using a trusted service clock. Their difference establishes a server-observed interval, not the time a person spent reading. Loading, network delays, interruptions, and idle periods are included unless separately accounted for. Browser timestamps can add context but are untrusted and modifiable.

```
Assignment issued -> content delivered -> interaction milestones -> submitted
       |                                                          |
       +---------------- server-observed interval ----------------+

Client signals: focus changes, rating changes, revision counts
                         |
                         v
                  anomaly assessment
                         |
                         v
       content checks + benchmarks + independent adjudication
```

An immediate submission can be inconsistent with meaningful evaluation of a newly issued complex task. It may also indicate a timing bug, a replayed event, or preloaded content. Flag the anomaly and inspect its context. A bot can wait before submitting; a long interval does not prove human review.

No numeric dwell threshold is enforced by the current page. A future release must define thresholds by task class, loading exclusions, interrupted-session handling, and accessibility accommodations before treating them as submission gates. The modal's 1.5-second animation is not an interaction measurement.

### Reading velocity heuristics

A basic diagnostic can compare visible content volume with active viewing time:

```
nominal reading rate = visible word count / active viewing minutes
```

This ratio is a heuristic, not a calibrated measure of cognition. Code, mathematics, repeated boilerplate, and prose impose different reading demands. Expert familiarity can shorten evaluation, while assistive technology may produce different focus patterns. A universal speed cutoff would confuse these cases.

A robust review should ask whether the contributor's evidence shows task-specific understanding. Correct identification of a subtle ordering defect carries more information than elapsed time alone. Use velocity anomalies to prioritize review and combine them with quality evidence rather than automatically equating speed with abuse.

### Sliders, keyboard events, and revisions

The proposed interaction layer can record coarse events such as a rating change, editor revision count, focus transition, and form readiness. It should not capture raw keystrokes, clipboard contents, unrelated browsing, or credentials. A minimal telemetry record identifies the assignment, event type, relative timing, and policy version without copying secret-bearing input.

Interaction events are not proof of humanity: scripts can synthesize them. Conversely, a keyboard-only or assistive-technology user may produce few pointer events. Requiring arbitrary mouse motion or needless slider changes can reward ritual rather than evaluation quality.

The current form checks selection, minimum Expert rationale length, patch length when applicable, active request state, and AI-review acknowledgment. It does not require a slider change or inspect keystroke cadence before unlocking submission. Ratings begin at 3 and can remain there. Documentation of proposed telemetry must not be read as a description of those existing gates.

## Silent benchmark injection

Gold traps are hidden benchmark assignments with independently established expected outcomes. Their placement is concealed to prevent selective effort on known tests. The policy, data use, and possible consequences should nevertheless be disclosed; concealed placement does not justify undisclosed sanctions.

The requested **10–15% rate is a proposed standard range**. The enterprise form currently offers 5%, 10%, and 20% selections with a 10% default and states that insertion is not implemented. A production campaign must reconcile its configured rate with its published policy before enforcement. The current queue does not contain a measured operational benchmark frequency.

### Sampling mechanics

There are two distinct ways to interpret a percentage:

| Assignment method           | Meaning for 100 delivered tasks at 10%                | Operational implication                                            |
| --------------------------- | ----------------------------------------------------- | ------------------------------------------------------------------ |
| Fixed randomized quota      | Exactly 10 benchmark tasks, with randomized positions | Predictable quota; short-window patterns must not reveal placement |
| Independent random sampling | Expected count 10, but actual count varies            | Requires minimum scored evidence before consequential decisions    |

For independent draws, if `K` is the benchmark count, `N` the number of delivered tasks, and `p` the insertion probability:

```
E[K] = N * p
Var[K] = N * p * (1 - p)
P(K = 0) = (1 - p)^N

N = 10 and p = 0.10:
E[K] = 1
P(K = 0) = 0.9^10 = approximately 34.87%
```

A short queue can therefore contain no benchmarks without a system malfunction. Campaigns must define whether skipped, expired, replaced, and invalidated tasks belong in the frequency denominator. Report delivered frequency separately from completed scored frequency; they answer different questions.

The assignment service should determine hidden placement server-side. Shipping a benchmark flag or reference answer to the browser would undermine concealment. A browser-to-provider Copilot must receive only authorized task context, not the hidden reference or scoring metadata.

### Reference construction and maintenance

A benchmark needs a versioned prompt, rubric, expected result or evaluation procedure, accepted alternatives, tolerance, domain, and approval record. Arithmetic and deterministic code can support reproducible reference checks. Open-ended preferences generally require expert adjudication and uncertainty handling; describing all such references as mathematically proven is inaccurate.

Reference validation must be independent of the contributor answer being scored. For code, a reference may include boundary tests and an explanation of the expected behavior. For grounded question answering, preserve the approved source context and the claim it supports. Avoid turning stylistic preferences into undisclosed factual ground truth.

Retire leaked, obsolete, ambiguous, or disputed items. If a reference is wrong, invalidate affected scores and recompute dependent eligibility decisions. Immutability of an audit record should preserve the correction history, not prevent correction of an erroneous judgment.

## Benchmark accuracy and consensus deviation

### Accuracy with an explicit denominator

For binary accepted benchmark outcomes:

```
benchmark accuracy = correct valid benchmark judgments / valid scored benchmarks
```

Nine correct judgments out of ten yield 90%. One correct judgment out of one yields 100% with much less evidence. Zero scored benchmarks means insufficient evidence, not perfect accuracy. A reputation display should include the observation window, domain, number of scored items, and benchmark version alongside the percentage.

Partial-credit tasks require an explicit weighting rule. Do not mix a fractional rubric score into a binary accuracy denominator without documenting the transformation. Exclude invalidated references consistently from both numerator and denominator and preserve why they were removed.

A simple uncertainty illustration uses the standard error approximation for independent binary observations:

```
p_hat = correct / n
standard error = sqrt(p_hat * (1 - p_hat) / n)

90 / 100: p_hat = 0.90, standard error = 0.03
9 / 10:   p_hat = 0.90, standard error = approximately 0.0949
```

This approximation is not an enforcement rule and is unreliable near boundaries or with small samples. Correlated tasks further weaken independence assumptions. Production policy needs an appropriate interval method and sufficient evidence rather than treating the same headline percentage as equally reliable at every sample size.

### Agreement is different from truth

For overlapping pairwise tasks, define whether agreement means agreement with another contributor, the majority, or an adjudicated label. Ties and abstentions need explicit treatment. A simple disagreement fraction can be written as mismatches divided by comparable independently adjudicated items, but the comparator and inclusion rules must be fixed first.

If three contributors choose A and two choose B, the 60% majority is not proof that A is correct. The minority may have noticed a real defect. Conversely, identical answers from coordinated accounts may create strong-looking agreement without independent evidence.

For 1–5 criterion scores, an optional diagnostic is absolute distance from an adjudicated reference divided by 4. A score of 2 against reference 5 gives `3 / 4 = 0.75` normalized deviation. This measures disagreement under a chosen numeric convention; it does not prove poisoning. Different dimensions should retain separate evidence rather than disappearing into one unexplained reputation number.

The displayed **98.4% Swarm Agreement** is static demonstration content. It is neither calculated from submitted tasks nor accompanied by an operational sample window. The application must not use it as a verified credential without a corresponding scoring service.

## Qualification tiers and compensation

A production qualification system should distinguish a new contributor with little evidence, a contributor calibrated for a task class, and a specialist with demonstrated domain competence. These are design categories, not currently enforced tier assignments. A high score in code evaluation should not automatically authorize clinical or legal evaluation.

The enterprise interface describes expert gating using domain credentials and consensus above 95%; that selection does not establish an implemented credential verifier or assignment filter. Likewise, the workbench's **Tier-2 Specialist** badge is illustrative.

Per-task compensation must be fixed or disclosed for the accepted assignment before work begins. The present difference between Rapid `0.25 USDC` and Expert `0.50 USDC` is a mode-based demonstration, not a reputation multiplier. The additional Golden Patch `0.50 USDC` is a correction bonus. No active tier-to-payout multiplier schedule exists in the inspected implementation.

A future specialist premium must specify eligibility evidence, applicable domains, base rate, premium calculation, funding, and the effect of a status change on already assigned work. It must not retroactively reduce an earned amount because a later task changed a reputation score.

## Strike escalation and wallet disqualification

The following is a proposed adjudication framework. Numeric strike counts, minimum sample sizes, expiry periods, and permanent-ban criteria are not implemented. The enterprise UI's below-85% benchmark pause language is a proposed rule without an operational scoring or enforcement service.

| Stage                      | Evidence required                                                                               | Intended response                                              |
| -------------------------- | ----------------------------------------------------------------------------------------------- | -------------------------------------------------------------- |
| Corrective feedback        | A substantiated isolated error                                                                  | Explain the failed criterion and offer recalibration           |
| Temporary assignment pause | Repeated validated failures with sufficient relevant evidence, or credible suspected compromise | Stop new assignments in the affected scope while reviewing     |
| Restricted qualification   | Persistent domain-specific deficiencies or adjudicated violations                               | Limit task access or require requalification                   |
| Campaign disqualification  | Established manipulation or repeated serious noncompliance                                      | Revoke campaign eligibility with a recorded decision           |
| Permanent wallet exclusion | Adjudicated severe or persistent abuse under an explicit published policy                       | Maintain an exclusion record with a correction and appeal path |

Permanent exclusion must not be inferred from a single failed benchmark or unusual cadence. A precautionary pause is not a final malicious-intent finding. Distinguish erroneous work from intentional poisoning and preserve the evidence supporting that distinction.

A decision record should identify the wallet, affected assignment references, applicable policy version, evidence, adjudicator, scope, duration, and review outcome. Contributor notices can explain the violated criterion without exposing active hidden answers. If an appeal succeeds or a benchmark is invalidated, update reputation and eligibility consistently rather than merely adding an unread note.

A wallet-level denylist excludes an address, not necessarily the controlling person. Preventing re-entry through fresh wallets requires additional controls; the registry alone does not provide them. Document this limitation instead of presenting wallet exclusion as complete sybil resistance.

## Rejection, slashing, and payment finality

Rejecting a pending contribution, denying a future assignment, and confiscating collateral are different operations. The current registry supports rejection of pending datasets. It has no contributor stake to slash, no automated wallet blacklist, and no reputation-based asset seizure.

The bundler role controls registration, so a future coordinator can withhold ineligible assignments or registrations off-chain. That is an operator-enforced boundary, not an on-chain proof of contributor eligibility. There is no implemented reputation-triggered clawback of an already completed payout.

An earned payment, disputed pending reward, and unearned future opportunity must remain distinct in accounting and notices. A UI strike cannot silently reverse a settled native transfer. Any future staking or penalty mechanism requires its own specified contract behavior, authorization, and audit trail.

## Contributor operational practice

Apply the same care to every assignment, whether or not it might be a benchmark. Preserve independent judgment, inspect AI suggestions, and do not fabricate evidence of tests or source checks. Record uncertainty when it changes the decision and use skip or flag for defective assignments.

For a disputed score, retain the task reference and explain the precise criterion or reference defect without publishing confidential payloads. The current flag control is local and does not create a production support case. An operational integrity release must supply a real reporting and adjudication path before presenting these controls as enforceable contributor policy.

Return to [Workbench Overview](/2.-contributor-and-annotator-guide/2.1-workbench-overview.md) for current interface behavior or [Interaction Telemetry](/4.-the-4-stage-verification-engine/4.2-interaction-telemetry.md) for the verification-engine architecture.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/2.-contributor-and-annotator-guide/2.4-reputation-and-gold-traps.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
