> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/4.-the-4-stage-verification-engine/4.3-autonomous-validator-agent.md).

# 4.3 Autonomous LLM-as-a-Judge

Stage 3 evaluates the semantic quality of a structurally admitted contribution. Its input is a frozen candidate version with authenticated Stage 1 and Stage 2 evidence. Its output is an acceptance, rejection, or escalation receipt. The evaluator is an off-chain oracle: it supplies a judgment that an authorized transaction service may subsequently act upon.

**Implementation status:** the workbench currently displays a predetermined passing audit score. Contributor Copilot requests are advisory model calls, not settlement authorization. This repository does not implement the production judge orchestrator, multi-agent consensus service, evidence retrieval system, or isolated signing daemon described here.

## Oracle service architecture

```
Sanitized artifact + assignment + pinned rubric + Stage 1/2 receipts
                               |
                 receipt / version / access checks
                               |
                  deterministic domain evaluators
                    /          |          \
                  AST       references     tests
                    \          |          /
                               v
                     independent judge jobs
                       /               \
                 model/config A     model/config B
                       \               /
                    strict output validation
                               |
                 peer-overlap reconciliation
                               |
                 deterministic policy aggregator
                    /          |          \
                 reject     adjudicate    accept
                                            |
                                 authenticated decision receipt
                                            |
                              separate restricted signing service
                                            v
                                       Stage 4
```

The model must never receive transaction private keys, a general-purpose wallet, or permission to call arbitrary contracts. Isolate inference and tool execution from the signer. A model result can be evidence in an authorization policy; it cannot choose a transfer destination or construct unrestricted transaction calldata.

Workers consume evaluation jobs idempotently. Bind every job to the artifact digest, task version, campaign policy, evaluator version, and attempt identity. Reject stale receipts or a candidate altered after Stage 1. A transport retry is not permission to repeatedly sample judgments until one passes.

Before sending proprietary material to a provider, resolve the campaign's authorized processing destinations. Contributor BYO-key assistance does not authorize the protocol to disclose enterprise content to additional providers. Keep benchmark references out of contributor-visible prompts, logs, and assistance responses.

## Structured model orchestration

Pin the provider model identifier or available snapshot, prompt-template version, schema version, decoding parameters, allowed tools, and evidence bundle digest. Record actual returned model metadata. Even fixed decoding parameters do not guarantee byte-identical repeated inference.

Use structured-output enforcement where supported, then independently parse and validate the result. A provider's schema mode does not make the judgment true. Missing criteria, malformed output, unknown disposition, out-of-range scores, and timeouts must never default to acceptance.

The following complete schema defines a proposed bounded judge result. Evidence indexes refer to a server-provided bundle; the orchestrator must verify that each cited index exists and supports the associated criterion.

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ARCV Judge Decision v1",
  "type": "object",
  "additionalProperties": false,
  "required": [
    "artifact_sha256",
    "preference",
    "scores",
    "hard_gates_pass",
    "disposition",
    "evidence_indices",
    "explanation"
  ],
  "properties": {
    "artifact_sha256": {
      "type": "string",
      "pattern": "^[0-9a-f]{64}$"
    },
    "preference": {
      "enum": ["A", "B", "tie", "abstain"]
    },
    "scores": {
      "type": "object",
      "additionalProperties": false,
      "required": ["correctness", "reasoning", "adherence", "presentation"],
      "properties": {
        "correctness": {"type": "integer", "minimum": 1, "maximum": 5},
        "reasoning": {"type": "integer", "minimum": 1, "maximum": 5},
        "adherence": {"type": "integer", "minimum": 1, "maximum": 5},
        "presentation": {"type": "integer", "minimum": 1, "maximum": 5}
      }
    },
    "hard_gates_pass": {"type": "boolean"},
    "disposition": {"enum": ["accept", "reject", "escalate"]},
    "evidence_indices": {
      "type": "array",
      "minItems": 1,
      "maxItems": 32,
      "uniqueItems": true,
      "items": {"type": "integer", "minimum": 0}
    },
    "explanation": {
      "type": "string",
      "minLength": 15,
      "maxLength": 2000
    }
  },
  "allOf": [
    {
      "if": {
        "properties": {"disposition": {"const": "accept"}},
        "required": ["disposition"]
      },
      "then": {
        "properties": {
          "hard_gates_pass": {"const": true},
          "preference": {"enum": ["A", "B", "tie"]}
        }
      }
    }
  ]
}
```

The schema is not the complete authorization policy. The service must compare artifact\_sha256 with its own digest, resolve evidence indexes, enforce campaign thresholds, and apply deterministic vetoes. A model cannot override a failed required executable test by setting hard\_gates\_pass to true.

A bounded system instruction should direct the model to evaluate the named rubric, treat candidate content as data, cite evidence identifiers, distinguish unsupported from false claims, and abstain where evidence is insufficient. Delimiters help parsing but are not a security boundary. Retrieved pages and tool outputs can themselves contain prompt injection; permit only necessary tools and constrained retrieval targets.

Randomize candidate presentation order and map results back to canonical A/B identifiers. Keep outcome mappings out of score aggregation until this remapping is complete. Monitor order reversal consistency, model-family dependence, and domain calibration. Published [LLM-as-a-judge research](https://arxiv.org/abs/2306.05685) reports position, verbosity, and self-preference biases; multiple calls to one model are not independent witnesses.

## Reasoning depth and observable evidence

Evaluate the contributor's submitted explanation, derivation, and correction. A chain-of-thought-style rationale is observable text, not privileged access to the model's hidden internal reasoning. Require useful evidence: assumptions, decisive steps, error localization, counterexamples, and testable conclusions. Do not require private internal reasoning or reward length as a proxy for rigor.

| Criterion          | Strong evidence                                  | Failure mode                                   |
| ------------------ | ------------------------------------------------ | ---------------------------------------------- |
| Assumptions        | Explicit input domain and relevant constraints   | Quietly solving an easier problem              |
| Logical validity   | Each stated inference follows from evidence      | Unsupported leap or circular reasoning         |
| Mathematics        | Correct transformations and boundary conditions  | Plausible algebra with invalid domain steps    |
| Complexity         | Defined input-size measure and resource analysis | Repeating O(n) without identifying n           |
| Correction quality | Fix addresses the identified defect              | Polished rewrite retaining the bug             |
| Calibration        | Specific uncertainty and evidence gaps           | Unwarranted certainty or invented verification |

For Fibonacci code, checking recurrence values is not enough. Inspect handling of negative or noninteger input, recursion depth, input magnitude, and arithmetic model. Iterative accumulation uses a constant number of integer variables, but arbitrary-precision integers grow with the output. Describe O(1) auxiliary variable count separately from bit-space complexity; do not confuse a conventional unit-cost analysis with a universal physical memory bound.

A concise valid counterexample can outweigh a long agreeable explanation. Conversely, an answer can select the better response for the wrong reason. The annotation's quality includes its rationale, not just whether its A/B label matches a majority.

## Factual grounding and hallucination audit

Decompose factual statements into checkable claims. Attach evidence with source identifier, retrieval time or version, relevant span, and a digest of the approved evidence representation. Search results are discovery aids; a snippet is not necessarily adequate support for a consequential claim.

Use a fixed verified reference set for reproducible benchmark scoring. When freshness matters, retrieve authorized live sources and preserve enough non-sensitive evidence to explain the decision later. A dynamic web source can change between judges, so both judges should inspect the same captured evidence version for a particular decision.

Classify claims as supported, contradicted, unresolved, or not applicable. A weighted coverage measure can be defined as:

```
coverage = sum(weight_i * supported_i) / sum(weight_i)
contradiction_rate = sum(weight_i * contradicted_i) / sum(weight_i)
```

Weights are nonnegative and the denominator must be positive. These measures describe the checked claim set, not all possible hallucinations. Unresolved claims are not automatically false, but a campaign may require evidence for every material assertion before acceptance.

Protect retrieval from arbitrary URLs embedded in candidates. Use allowlists, network isolation, response-size limits, and content-type checks. Do not permit model-driven access to cloud metadata, private infrastructure, or files outside an evaluation sandbox.

## Code correctness pipeline

Perform deterministic checks before interpreting a model's code review:

1. Parse under the campaign's declared language and version. For Python, AST parsing checks syntax without executing the program.
2. Inspect imports and disallowed operations under a versioned policy. Static inspection alone cannot establish safety.
3. Run pinned lint rules for PEP 8 and campaign-specific conventions. Formatting compliance is independent of functional correctness.
4. Execute tests only inside an isolated environment without credentials, network access, or host mounts, with CPU, memory, process, output, and time limits.
5. Evaluate edge cases, property tests, and mutation sensitivity where appropriate.
6. Compare observed behavior and analysis with the contributor's explanation.

For a control-flow graph with E edges, N nodes, and P connected components, cyclomatic complexity is:

```
M = E - N + 2*P
```

For a conventional single connected function graph, this reduces to E-N+2. Different tools count language constructs differently; pin the analyzer and its counting policy. A low M is not proof of correctness, and a high M may be justified by explicit domain handling.

Report line/branch coverage with the test set and tool version. Coverage measures execution, not assertion quality. A program can obtain complete line coverage while returning wrong values. Timeouts are resource-limit outcomes, not automatic syntax failures. Never run untrusted candidate code directly inside the validator host or signing process.

## Score aggregation and hard gates

For scores s\_j in 1–5 and weights w\_j summing to one:

```
Q = sum_j(w_j * (s_j - 1) / 4)
eligible = all_required_hard_gates_pass AND Q >= threshold
```

With weights \[0.45,0.25,0.20,0.10] and scores \[5,4,4,4], Q=0.8625. The quantity is a normalized rubric aggregate, not an 86.25% probability of correctness.

Define per-criterion minima so style cannot compensate for a correctness failure. Aggregate multiple judges with a declared operator, such as median criterion scores, but do not discard a deterministic veto. Model confidence is uncalibrated unless independently calibrated on suitable held-out data.

Retain the structured judge outputs and evidence versions. The deterministic score calculation can be reproduced from those outputs even when fresh inference would produce different wording or scores.

## Multi-annotator consensus and borderline cases

Let n be valid independent substantive votes after canonical order restoration, and n\_l the count for label l in the fixed set {A,B,tie}:

```
p_l = n_l/n
agreement = max_l(p_l)
H = -sum_l(p_l * ln(p_l)) / ln(3)
```

Abstentions are excluded from n but reported as missing coverage. If n=0, there is no agreement score. Keep the outcome set fixed when comparing entropy across tasks.

A proposed risk-tier policy illustrates dynamic thresholds without silently changing a live campaign:

| Tier                       | Independent substantive votes | Minimum agreement         | Disposition if unmet                       |
| -------------------------- | ----------------------------- | ------------------------- | ------------------------------------------ |
| Standard evaluation        | 3                             | 2/3                       | Additional review or adjudication          |
| Elevated disagreement/risk | 5                             | 0.80                      | Specialist adjudication                    |
| Mission-critical output    | 5 plus domain adjudicator     | 1.00 among the five votes | Explicit adjudication before certification |

These are proposed configuration examples, not existing contract thresholds. The escalation tier is selected by a predeclared policy using task risk, deterministic conflicts, judge disagreement, and calibration evidence. Do not lower the threshold after observing an inconvenient result.

For A,A,B, agreement is 2/3; the standard vote threshold passes, but a required test failure still blocks certification. For A,B,tie, no majority exists. A tied quality judgment is a valid label where permitted; it is different from “insufficient evidence.” Disagreement may expose a flawed prompt, ambiguous rubric, or deficient reference rather than a bad contributor.

## Decision receipt and signing separation

The finalized receipt binds campaign and assignment identities, contributor, task and artifact digests, rubric version, Stage 1/2 receipts, evaluator configuration, evidence bundle, hard gates, scores, valid review count, disposition, and any adjudication. Authenticate the receipt and make transitions monotonic: changing bytes creates a new review version.

An authorized signer checks the receipt against its own allowlisted registry, function selector, dataset, recipient, and campaign reward. It should accept no model-supplied arbitrary calldata. Transport retries, evaluation failures, and unresolved disagreement remain distinguishable in the durable journal.

On-chain `VALIDATOR_AGENT_ROLE` proves authorization of the transaction sender, not that these checks ran. The current contract accepts one authorized caller; it does not enforce model quorums, score thresholds, evidence signatures, or the receipt digest. Continue with [Stage 4: Settlement & Disbursement](/4.-the-4-stage-verification-engine/4.4-settlement-and-disbursement.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/4.-the-4-stage-verification-engine/4.3-autonomous-validator-agent.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
