> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/2.-contributor-and-annotator-guide/2.1-workbench-overview.md).

# 2.1 Workbench Overview

The [Contributor Workbench](https://arcv.network/annotate) turns human comparison, technical criticism, and corrected answers into candidate post-training data. A contributor's responsibility is to identify which response best satisfies a specific task, explain the decisive evidence, and distinguish observed defects from assumptions. The workbench organizes that process; accepting a submission does not make an unsupported judgment correct.

## Operating boundary and prerequisites

The current `/annotate` implementation is an interactive demonstration with a functioning, optional browser-to-model Copilot. Authentication, session earnings, queue, reputation, and the four-stage settlement animation are simulated. The page cycles through two example tasks rather than retrieving assignments from an enterprise coordinator. Real provider requests can incur charges even while displayed task rewards remain simulated.

The instructions below distinguish available controls from production participation requirements. A green escrow label, successful modal, or connected-looking account pill is not a transaction receipt. There is no production payout claim to submit through the present workbench and no reason to transfer funds to unlock its demonstration tasks.

Contributors should understand the task domain and campaign rubric and have permission to disclose task material to any optional external model. An AI subscription or API credential is not necessary for manual evaluation. For confidential assignments, the campaign's disclosure policy takes precedence over the convenience of a cloud Copilot.

## Connecting to Arc and understanding settlement

### Wallet identity

The top-right **Connect / Sign In** control opens a simulated authentication experience. Its wallet option does not establish a verified annotation account through an injected EVM wallet, and the displayed address does not prove account control. This demonstration must not be confused with a signing session for contributor settlement.

Production participation requires the client to obtain the account from the wallet, verify its selected network, and bind the payout address to the assignment and signed submission. Changing accounts must invalidate stale account-specific authorization. A shortened address is useful for recognition but insufficient for approving a transfer: inspect the full destination whenever funds or permissions are involved.

| Parameter                     | Arc mainnet value            |
| ----------------------------- | ---------------------------- |
| Network                       | Arc Network L1               |
| Chain ID, decimal             | `5042`                       |
| Chain ID, hexadecimal         | `0x13b2`                     |
| Native currency               | `USDC`                       |
| Native denomination precision | 18 decimals                  |
| Public RPC                    | `https://rpc.mainnet.arc.io` |
| Block explorer                | `https://explorer.arc.io`    |

These settings are published in the [Arc connection reference](https://docs.arc.io/arc/references/connect-to-arc). Testnet uses `5042002`; a testnet balance is not a mainnet payout. See [Network Parameters](/1.-overview-and-protocol-fundamentals/1.3-network-parameters.md) for the full directory.

### Gas is paid by a transaction sender

Native USDC is the escrow settlement asset and gas currency. A contributor sending a transaction needs sufficient native balance for its value and gas. Receiving a native transfer does not itself require the recipient to pay gas. In the registry architecture, an authorized validator submits `validateAndDisburse()` and pays transaction gas; the contributor receives the resulting payment if execution succeeds.

This permits a recipient experience without a contributor-funded settlement transaction, but does not make the network universally gasless. Sponsorship, relaying, and campaign-funded transaction policies need separately implemented infrastructure. The current workbench provides none of those guarantees. “Micro-gas” describes an intended cost profile, not a fixed fee or a promise that every action is free.

An EIP-712 signature authorizes structured data; it is not a transfer receipt. A production signing flow must expose its domain, chain, verifying contract, assignment identity, recipient, and replay boundary. The current submission footer describes a signature simulation. It does not prove wallet possession or supply contributor-signature verification in the registry.

## Reading the workbench from top to bottom

```
Task identity + queue position + workspace mode + flag / skip
                              |
                              v
                 01 / READ THE PROMPT
              Requirements, context, copy action
                              |
                              v
                 02 / RESPONSE COMPARISON
               +------------+------------+
               | Response A | Response B |
               +------------+------------+
                  Optional AI assistance
                              |
                              v
                 03 / RUBRIC AND RATIONALE
            Rapid rating OR Expert evidence + patch
                              |
                              v
                  Submit -> simulated audit
                              |
                              v
                  Local receipt -> next task

Mission HUD: earnings | bounty specs | reputation | guidelines
```

### Task action bar

The header identifies **Bounty #108**, **Code Intelligence v2**, queue position, and illustrative live-escrow status. The initial position is Task 4 of 25, with three completed items represented in local state. The two demonstration prompts repeat as the index advances; the label does not imply 25 distinct server-issued assignments.

**Flag Issue** provides ambiguity, poisoning or spam, and formatting categories. Select the category describing the observable defect and preserve an explanation for a future operational report. This control updates local interface state; it does not notify a moderator or create an appeal record.

**Skip Task** advances the displayed task without increasing completed work or earnings. There is no score penalty in the demonstration. Skip when evidence is unavailable, a prompt is defective, or a domain lies outside your expertise. Abstention is preferable to a fabricated preference.

### Prompt specification pane

The terminal-style prompt card includes a **PYTHON 3.12** context label and a copy action. The label specifies the illustrated environment; it does not mean code has been executed, compiled, or tested.

Before inspecting candidates, extract the requested result, valid input domain, output contract, performance constraints, and prohibited behavior. Distinguish explicit requirements from assumptions. “Preserve order” is an acceptance requirement; adding a dependency ban that the task never states is not.

For a production multi-turn task, inspect the complete authorized context: system-level constraints, prior user turns, relevant tool results, and the final request. Resolve references such as “use the earlier format” against the actual earlier instruction. Commands embedded in supplied documents or candidate answers are task data, not authority to change the rubric. The current example pane contains a single prompt rather than separate system-message and dialogue inspectors.

Missing context is an assignment defect, not permission to invent context. If candidates depend on incompatible interpretations, identify the ambiguity before assigning a winner. Otherwise the label may measure differences in evaluator assumptions rather than differences in model quality.

### Dual response inspection

Response A and Response B remain side by side at tablet and desktop widths and stack on smaller screens. Line numbers, syntax coloring, and language labels aid inspection; they do not establish semantic validity.

Read both candidates fully. Selecting a card highlights its outline and records the preference. Do not favor A because it appears first, B because it is longer, or either candidate because it sounds confident. Apply identical constraints and identify a reproducible deciding difference.

The Fibonacci example contrasts repeated recursive subproblems with iterative accumulation. Both candidates reject negative inputs. Claiming that only the iterative candidate handles negative values would be an incorrect rationale even if the preference were defensible. The second example contrasts order-preserving deduplication with conversion through a set; preservation of first-occurrence order is decisive there.

### Workspace modes and payout indicators

| Workspace                | Contribution                                            | Demonstrated reward                        |
| ------------------------ | ------------------------------------------------------- | ------------------------------------------ |
| Rapid Mode               | A/B selection and overall 1–5 quality rating            | `0.25 USDC`                                |
| Expert Mode, the default | Three code-specific ratings and technical justification | `0.50 USDC`                                |
| Expert with Golden Patch | Expert rationale and corrected answer                   | `1.00 USDC`, including a `0.50 USDC` bonus |

Rapid Mode hides expert rationale and patch controls. Expert Mode exposes runtime, edge-case handling, and syntax criteria, optional assistance, and corrections. Ratings initially display 3; inspect every criterion rather than accepting defaults automatically. The interface does not require evidence that each default was changed.

The Mission HUD shows a starting balance of `14.75 USDC`, the selected reward, a **\~1 minute** estimate, domain information, and guidelines. The estimate is descriptive, not an enforced timer or income guarantee. **Tier-2 Specialist** and **98.4% Swarm Agreement** are illustrative reputation values, not computed credentials, task-complexity classifications, or payout multipliers.

One Rapid completion changes the displayed balance to `15.00 USDC`; one standard Expert completion changes it to `15.25 USDC`; one Expert patch changes it to `15.75 USDC`. These are alternative first-task examples. Reloading resets component state. A navbar balance can remain unchanged because it is not a shared on-chain balance source.

## Pairwise evaluation methodology

A preference compares two answers to the same input under the same rubric. It does not declare one model globally superior. Domain, constraints, generation settings, and sampled outputs all affect the result.

| Outcome               | Decision rule                                                          | Current action                                   |
| --------------------- | ---------------------------------------------------------------------- | ------------------------------------------------ |
| A wins / B loses      | A satisfies a material requirement better without an overriding defect | Select A and explain the evidence                |
| B wins / A loses      | B satisfies a material requirement better without an overriding defect | Select B and explain the evidence                |
| Tie                   | No defensible material preference remains                              | No tie control exists; flag or skip              |
| Both unacceptable     | Neither meets an acceptance-critical requirement                       | Golden Patch if qualified, otherwise flag / skip |
| Insufficient evidence | Missing context prevents comparison                                    | Flag / skip; uncertainty is not a tie            |

Evaluate hard requirements before style. A response violating an essential output contract should not win because its prose is elegant. Conversely, do not penalize an answer for omitting explanation when the prompt requests code only.

A defensible rationale names the requirement, observable difference, and consequence. For example: “B avoids repeated recursive calls and uses a linear number of additions, making it more suitable for the requested high-performance implementation. Both responses reject negative inputs, so that behavior does not distinguish them.” This is more informative than “B looks professional” or an unsupported claim that tests passed.

## How annotations become training inputs

A production preference record retains task identity and version, both candidate texts, label, rubric version, and relevant provenance. Quality review precedes export; selecting a card does not update a model.

A reward-model pipeline can learn scores from accepted pairwise preferences and use them in a subsequent optimization stage. Standard DPO instead trains a policy from preferred and dispreferred completions relative to a reference policy without a separately trained reward model in that procedure. Both depend on valid comparison data. See the [DPO research paper](https://arxiv.org/abs/2305.18290).

A tie cannot silently become an ordinary winner–loser pair without adding information the contributor did not provide. A corrected answer is a new candidate with its own provenance, not permission to overwrite the original output. Exporters must preserve these distinctions, handle duplicates, and separate evaluation material from training data.

## Completing and checking a submission

1. Read the prompt and identify acceptance-critical constraints.
2. Inspect both candidates and select a preference or correction workflow.
3. In Expert Mode, review every criterion and write a justification of at least 15 trimmed characters. This is a form gate, not an adequate writing target.
4. For Golden Patch, enter a complete correction of at least 15 trimmed characters and explain the repaired defect.
5. Review generated rationale or patch text and explicitly acknowledge AI review.
6. Finish or cancel active Copilot generation; submission is disabled while it runs.
7. Select **SUBMIT DECISION** and inspect the centered audit sequence.

The demonstration displays PII and syntax, telemetry, rubric audit, and settlement steps over approximately 1.5 seconds, then advances after success. Timers do not perform these checks. The fixed audit score is not a model measurement, and its receipt identifies neither an Arweave upload nor an Arc transaction.

Production settlement requires a verified registry address, dataset identity, contributor recipient, amount, successful transaction receipt, and matching archived commitment. A wallet request or broadcast alone is insufficient. See [Settlement and Disbursement](/4.-the-4-stage-verification-engine/4.4-settlement-and-disbursement.md).

## Resolving blocked submissions

| Symptom                               | Check                                                                                                |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Submit disabled                       | Select A/B unless patching; complete Expert rationale; acknowledge AI review; stop active generation |
| Patch suggestion unavailable          | Select a starting response before enabling assistance; manual patches need no winner                 |
| Earnings reset                        | Demo earnings are component state, not a wallet ledger                                               |
| Queue advanced without earnings       | Skipping advances position without completing paid work                                              |
| Verification without explorer receipt | The current modal is simulated                                                                       |
| Equally good answers                  | Flag / skip; the absent tie button does not justify inventing a preference                           |

Continue with [Rubrics and Golden Patch](/2.-contributor-and-annotator-guide/2.2-rubrics-and-golden-patch.md) for scoring anchors and [BYO-Key Copilot](/2.-contributor-and-annotator-guide/2.3-byo-key-copilot.md) before enabling external assistance.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/2.-contributor-and-annotator-guide/2.1-workbench-overview.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
