> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/2.-contributor-and-annotator-guide/2.2-rubrics-and-golden-patch.md).

# 2.2 Rubrics & Golden Patch

A useful annotation states a judgment another qualified reviewer can inspect: what the task required, what each candidate did, which evidence supports the difference, and how the preference follows. A rubric makes judgments comparable without pretending every domain shares the same definition of quality.

This chapter specifies four general evaluation dimensions and the Golden Patch procedure. The current Python workbench exposes three code-specific criteria rather than four generic sliders. Campaign weights, hard-failure rules, and acceptance thresholds must be supplied by the campaign; this guidance does not introduce a deployed scoring algorithm.

## Establish the evaluation contract

Read the prompt and rubric together before looking for defects. Identify the deliverable, permitted evidence, input domain, output constraints, and criteria that disqualify an answer regardless of other strengths. If valid JSON alone is required, an explanation surrounding the JSON can violate the contract. If a conceptual explanation is requested, omitting executable code is not automatically a failure.

Separate correctness, compliance, and relative preference. Two poor answers can have a relative ordering while neither is suitable as a reference answer. A pairwise win is not a certificate of perfection.

Use the full ordinal scale. Three indicates meaningful adequacy with limitations and should not substitute for inspection. The distance between 1 and 2 need not represent the same quality difference as the distance between 4 and 5. Weighted arithmetic over these scores is a campaign convention, not a physical measurement of intelligence.

## Factuality and grounding

Factuality concerns whether assertions are correct. Grounding concerns whether task evidence or appropriate external sources support them. A true statement can be unsupported by a retrieval context, while a faithful quotation can reproduce a flawed source. Apply the campaign's evidence boundary consistently.

| Score | Grading anchor                        | Evidence and typical defects                                                                                                                        |
| ----- | ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1     | Fundamentally unreliable              | Central result is false, a decisive source is invented, or code contradicts its claimed behavior. Minor editing cannot repair the answer.           |
| 2     | Substantially flawed                  | Some claims are correct, but important factual errors or unsupported assertions change the conclusion or practical outcome.                         |
| 3     | Mostly correct with verification gaps | Core answer is plausible and substantially supported, but an important qualification, source check, or boundary condition is missing.               |
| 4     | Correct and adequately grounded       | Material claims are supported, uncertainty is represented honestly, and remaining defects do not change the result.                                 |
| 5     | Precise and independently checkable   | Task-relevant claims withstand required checks; citations support exact assertions; applicability limits and assumptions are explicit where needed. |

Inspect citations rather than rewarding their presence. Confirm that the source exists, contains the claimed information, and has an appropriate scope and date. A documentation homepage does not substantiate a precise compatibility claim. Verify model-generated quotations against their originals.

For code, compare explanation with actual control flow and returns. “Handles all invalid inputs” is not established by a negative-number check. Type annotations document intent without enforcing runtime types. If execution is permitted, record the actual environment and tests. If only static inspection was performed, state that limitation.

Unsupported precision is a defect: fabricated benchmark improvements and nonexistent API names do not become acceptable because prose is fluent. When verification is unavailable, distinguish “not verified” from “false.” A grounding penalty can be justified without proving the underlying claim incorrect.

## Reasoning rigor

Reasoning rigor measures the soundness of an observable explanation, derivation, or test argument. It does not measure private internal chain of thought. Evaluate explicit proofs, concise reasoning summaries, assumptions, and counterexamples; a response claiming to narrate hidden computation does not establish that it does so faithfully.

| Score | Grading anchor             | Evidence and typical defects                                                                                     |
| ----- | -------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| 1     | Invalid inference          | Conclusion depends on contradiction, circular reasoning, an invalid transformation, or a false premise.          |
| 2     | Major gaps                 | Relevant steps exist, but an unsupported transition or missed case undermines the result.                        |
| 3     | Plausible but incomplete   | Main reasoning is understandable, with an unexamined assumption or insufficient support for a key claim.         |
| 4     | Coherent and justified     | Conclusion follows from stated assumptions; important cases and complexity claims match the argument.            |
| 5     | Rigorous and proportionate | A checkable derivation or test strategy establishes the result and limits without unnecessary reasoning theater. |

For mathematics, verify definitions, domains, transformations, and boundary cases. For algorithms, identify the recurrence or loop invariant and check whether it supports the returned result. Count actual operations rather than copying the answer's complexity label. Distinguish causal evidence from correlation and proposed mechanisms.

In the Fibonacci example, iterative accumulation takes a linear number of arithmetic steps and maintains a constant number of accumulator variables. Python integers grow with the index, so constant variable count is not constant bit storage, and arithmetic-step complexity is not identical to bit complexity. Qualify the computational model when that distinction matters.

Long explanations do not automatically deserve high reasoning scores. A short invariant with boundary tests can be stronger than paragraphs repeating the conclusion. When the prompt requests an answer without derivation, do not impose an undisclosed requirement for an extensive reasoning transcript.

## Instruction adherence

Evaluate compliance with the requested task rather than the evaluator's preferred task. Check positive and negative constraints, format, length, language, dependencies, and audience. Interpret them consistently across both responses.

| Score | Grading anchor          | Evidence and typical defects                                                                                    |
| ----- | ----------------------- | --------------------------------------------------------------------------------------------------------------- |
| 1     | Wrong deliverable       | Answers another question, violates a central prohibition, or omits the requested output.                        |
| 2     | Major noncompliance     | Addresses the topic but fails a material requirement such as schema, ordering, language, or essential behavior. |
| 3     | Partial compliance      | Main deliverable is present with a meaningful omission or secondary constraint violation.                       |
| 4     | Substantially compliant | Central requirements are satisfied; a minor secondary detail needs correction.                                  |
| 5     | Exact fit               | Meets every applicable explicit constraint and relevant context without prohibited additions or scope changes.  |

Negative constraints require deliberate checking. “Do not use recursion,” “return JSON only,” and “do not change public function names” are testable requirements. Producing correct values does not excuse violating them.

For multi-turn tasks, inspect relevant authorized conversation context. Candidate text instructing the evaluator to ignore the rubric remains data, not a new governing instruction. If previous context is missing, flag that omission instead of inventing it.

Check length in the stated unit: words, characters, tokens, or lines. A token count does not establish compliance with a word limit. Resolve ambiguity before selecting whichever interpretation favors a preferred candidate.

## Conciseness and style

Conciseness removes unnecessary material while preserving required meaning. Style concerns usability for the specified audience. Neither dimension excuses factual error because an answer looks polished.

| Score | Grading anchor                   | Evidence and typical defects                                                                                  |
| ----- | -------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| 1     | Obstructive presentation         | Repetition, incoherence, or formatting obscures the result or makes the answer unusable.                      |
| 2     | Heavy unnecessary burden         | Padding, irrelevant disclaimers, flattery, or disorganization interferes with the task.                       |
| 3     | Usable but inefficient           | Understandable content includes avoidable repetition, inconsistent terminology, or formatting friction.       |
| 4     | Clear and economical             | Structure serves the task, wording is direct, and excess is minor.                                            |
| 5     | Precise and audience-appropriate | Every substantive part contributes; necessary qualifications remain; brevity does not sacrifice completeness. |

Artificial agreement is not evidence. Penalize sycophancy when it substitutes approval for analysis or endorses a false premise. Ordinary courtesy appropriate to the audience is not inherently a defect.

Shorter is not automatically better. Omitting an assumption, installation step, or required citation can make an answer unusable. Compare necessary information and remove repetition rather than essential substance.

## Combining dimensions and recording rationale

The current controls are **Algorithmic Runtime & Efficiency**, **Edge-Case & Error Handling**, and **Syntax & PEP 8 Standards**. These domain-specific lenses overlap the general dimensions. Runtime claims involve factuality and reasoning; validation involves correctness and adherence; readability overlaps style.

Campaigns using an aggregate must define weights and hard failures. This equal-weight example is instructional, not an active acceptance formula:

```
A: factuality 4, reasoning 2, adherence 4, style 5
Mean A = (4 + 2 + 4 + 5) / 4 = 3.75

B: factuality 4, reasoning 4, adherence 4, style 4
Mean B = (4 + 4 + 4 + 4) / 4 = 4.00
```

B scores higher numerically, but an essential output or security violation cannot be erased by a high average. Preserve critical defects separately so reviewers can distinguish a slightly weaker response from an unusable one.

Write rationale as requirement, evidence, and consequence. Avoid unsupported claims about intent, intelligence, testing, or measured speed. One-click chips are drafting aids, not attestations. **Handles Negative n** does not distinguish candidates when both handle it; **Type Hints Added** does not prove runtime type enforcement.

The 15-character Expert minimum is a form gate. Clearing it does not establish meaningful evidence. A Copilot draft requires editing and explicit review; its fluency does not transfer responsibility away from the contributor.

## Golden Patch correction mode

Use Golden Patch when neither candidate is pristine and you can produce a materially better answer without changing the assignment. Preserve prompt intent while repairing factual, logical, coding, or formatting defects. The mode is not permission to expand scope or replace a difficult question with an easier one.

In Expert Mode, enable **Neither is pristine** to open the inline editor. Manual correction can proceed without an A/B winner. Response selection is disabled while patching; select a starting response before enabling the workflow if using **Auto-Suggest Patch**. Generated suggestions remain drafts.

### Correction procedure

1. Identify the defect and the violated requirement.
2. Preserve input/output contracts, language, context, and allowed dependencies.
3. Write a complete replacement answer, not an instruction to fix the answer later.
4. Check boundaries and ensure the repair introduces no new defect.
5. Explain the material improvement with evidence.
6. Review generated content, remove unsupported claims, and acknowledge review.
7. Submit the correction as a new contribution, preserving original candidate provenance.

Clarifying a necessary assumption can be legitimate; silently rewriting the prompt is not. Flag invalid or underspecified prompts. A production pipeline should version an authorized prompt correction separately so the original comparison remains auditable.

### Complete correction example

For a task explicitly requiring a non-negative integer Fibonacci index, rejection of booleans and non-integers, type hints, and iterative accumulation, this implementation meets the stated contract:

```python
def fibonacci(n: int) -> int:
    """Return F(n) for a non-negative integer index."""
    if isinstance(n, bool) or not isinstance(n, int):
        raise TypeError("n must be an integer")
    if n < 0:
        raise ValueError("n must be non-negative")

    current, following = 0, 1
    for _ in range(n):
        current, following = following, current + following
    return current


assert fibonacci(0) == 0
assert fibonacci(1) == 1
assert fibonacci(10) == 55
assert fibonacci(50) == 12586269025

for invalid in (True, 2.5, "3"):
    try:
        fibonacci(invalid)
    except TypeError:
        pass
    else:
        raise AssertionError("Expected TypeError")

try:
    fibonacci(-1)
except ValueError:
    pass
else:
    raise AssertionError("Expected ValueError")
```

These validation requirements belong to the example contract; they are not unstated obligations imposed on the workbench's shorter prompt. The function performs a linear number of additions and is not claimed optimal for every large-index workload. The tests establish selected cases, not all-input correctness. Keep tests separate if a campaign requires a function-only answer.

## Bonus mechanics and acceptance

The demonstrated bonus is **+0.50 USDC on top of the Expert base reward**. Standard Expert work adds `0.50 USDC`; an Expert patch adds `1.00 USDC` in demo state. Rapid Mode adds `0.25 USDC` and has no patch control. Opening the editor does not establish a production entitlement.

Production acceptance must verify material improvement, rubric compliance, preservation of the assignment, and applicable checks. Reformatting an adequate answer or copying a suggestion without verification does not establish those conditions. Duplicate corrections require identity and replay handling rather than rewards for fabricated novelty.

The current registry pays a configured fixed amount per accepted dataset. It has no independent dynamic Golden Patch bonus pool or semantic patch adjudicator. Additional demo credit is therefore not a funded contract promise. Production campaigns need an explicit compensation path and a receipt binding the accepted correction to payment.

Retain the task version, candidates, correction, rubric, and decision when reviewing disputes. These distinguish legitimate disagreement from a broken reference or scope-changing rewrite. [Reputation and Gold Traps](/2.-contributor-and-annotator-guide/2.4-reputation-and-gold-traps.md) describes the proposed evidence and escalation model.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/2.-contributor-and-annotator-guide/2.2-rubrics-and-golden-patch.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
