> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/3.-enterprise-and-frontier-ai-labs/3.2-custom-schemas.md).

# 3.2 Custom Schemas & Rubrics

Schemas constrain representation; rubrics evaluate meaning. A valid JSON record can still contain an incorrect answer, unsafe code, invented evidence, or unauthorized data. Production validation must apply both controls and preserve their versions with every accepted artifact.

Use separate schemas for raw tasks and completed annotations. The portal's textarea accepts a shallow object-schema check, not a complete validator. The examples below are complete Draft 2020-12 schemas for external validation, not a claim that the current frontend enforces every keyword. See the [JSON Schema specification](https://json-schema.org/draft/2020-12).

## Validation contract and processing order

A campaign needs three versioned objects: an input-task schema, an annotation schema, and a semantic rubric. The input schema describes what the worker may receive; the annotation schema describes what the worker must return; the rubric determines whether the result merits acceptance. A frontend parser successfully finding `prompt` and two response fields establishes none of the latter checks.

```
Bound raw bytes and record size
              |
              v
Strict UTF-8 / JSON parsing -> reject duplicate keys and non-finite numbers
              |
              v
Validate declared schema version -> types, required fields, bounds
              |
              v
Check cross-field and cross-record invariants
              |
              v
Run authorized domain evaluators and independent review
              |
              v
Freeze accepted representation -> hash exact archival bytes
```

Pin the schema dialect and validator configuration in the campaign manifest. Draft 2020-12 distinguishes validation keywords from annotations, and implementations can vary in optional format enforcement. A `format` label should not replace an explicit application check for a security-sensitive URL or identifier. The examples use local definitions and do not require arbitrary network resolution of schema references.

Reject unsupported versions rather than guessing which old schema a record resembles. Schema updates that alter required fields, semantics, or payable outcomes need a new version and explicit migration policy. Preserve the schema used for already accepted work; silently editing a mutable schema URL breaks reproducibility.

## Strict pairwise annotation schema

Save this block as `pairwise.schema.json`. The format includes conversation context, two candidate answers, the preference, rationale, and optional correction represented consistently as a string or null. Ties are supported by this exchange schema even though the current contributor UI does not expose a tie selector.

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ARCV Pairwise Annotation v2",
  "type": "object",
  "additionalProperties": false,
  "required": [
    "task_id",
    "messages",
    "response_a",
    "response_b",
    "preference",
    "rationale",
    "golden_patch",
    "schema_version",
    "rubric_version",
    "ratings"
  ],
  "properties": {
    "task_id": {
      "type": "string",
      "pattern": "^[a-zA-Z0-9_-]{1,64}$"
    },
    "messages": {
      "type": "array",
      "minItems": 1,
      "maxItems": 64,
      "items": {
        "type": "object",
        "additionalProperties": false,
        "required": [
          "role",
          "content"
        ],
        "properties": {
          "role": {
            "enum": [
              "system",
              "user",
              "assistant"
            ]
          },
          "content": {
            "type": "string",
            "minLength": 1,
            "maxLength": 32000
          }
        }
      }
    },
    "response_a": {
      "type": "string",
      "minLength": 1,
      "maxLength": 64000
    },
    "response_b": {
      "type": "string",
      "minLength": 1,
      "maxLength": 64000
    },
    "preference": {
      "enum": [
        "A",
        "B",
        "tie",
        "neither"
      ]
    },
    "rationale": {
      "type": "string",
      "minLength": 15,
      "maxLength": 4000
    },
    "golden_patch": {
      "type": [
        "string",
        "null"
      ],
      "minLength": 15,
      "maxLength": 64000
    },
    "schema_version": {
      "const": "pairwise-v2"
    },
    "rubric_version": {
      "type": "string",
      "pattern": "^[a-z0-9-]{1,64}$"
    },
    "ratings": {
      "type": "object",
      "additionalProperties": false,
      "required": [
        "factuality",
        "reasoning",
        "instruction_adherence",
        "style"
      ],
      "properties": {
        "factuality": {
          "type": "integer",
          "minimum": 1,
          "maximum": 5
        },
        "reasoning": {
          "type": "integer",
          "minimum": 1,
          "maximum": 5
        },
        "instruction_adherence": {
          "type": "integer",
          "minimum": 1,
          "maximum": 5
        },
        "style": {
          "type": "integer",
          "minimum": 1,
          "maximum": 5
        }
      }
    }
  },
  "allOf": [
    {
      "if": {
        "properties": {
          "preference": {
            "const": "neither"
          }
        }
      },
      "then": {
        "properties": {
          "golden_patch": {
            "type": "string"
          }
        }
      }
    }
  ]
}
```

This complete synthetic annotation passes that schema:

```json
{
  "task_id": "order-001",
  "messages": [
    {
      "role": "user",
      "content": "Remove duplicate hashable list values while preserving first occurrence order in Python."
    }
  ],
  "response_a": "def unique_values(values):\n    return list(dict.fromkeys(values))",
  "response_b": "def unique_values(values):\n    return list(set(values))",
  "preference": "A",
  "rationale": "Response A preserves insertion order. Response B removes duplicates but does not guarantee the requested order.",
  "golden_patch": null,
  "schema_version": "pairwise-v2",
  "rubric_version": "code-review-v1",
  "ratings": {
    "factuality": 5,
    "reasoning": 4,
    "instruction_adherence": 5,
    "style": 4
  }
}
```

Apply semantic checks for turn ordering, final user intent, distinct candidates, and rationale consistency. JSON Schema length counts do not prove that a rationale contains substantive text. Reject duplicate JSON keys before validation: ordinary parsers may silently keep the final value.

## Agent trace schema

Save as `agent-trace.schema.json`. This example records a constrained retrieval tool invocation and its observed result, not an executable instruction to the validator.

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ARCV Retrieval Agent Trace v1",
  "type": "object", "additionalProperties": false,
  "required": ["task_id", "prompt", "steps", "final_answer"],
  "properties": {
    "task_id": {"type": "string", "minLength": 1, "maxLength": 64},
    "prompt": {"type": "string", "minLength": 1, "maxLength": 32000},
    "steps": {
      "type": "array", "minItems": 1, "maxItems": 32,
      "items": {
        "type": "object", "additionalProperties": false,
        "required": ["call_id", "tool", "arguments", "result", "latency_ms"],
        "properties": {
          "call_id": {"type": "string", "pattern": "^call-[0-9]{1,6}$"},
          "tool": {"const": "retrieve_documents"},
          "arguments": {
            "type": "object", "additionalProperties": false,
            "required": ["query", "top_k"],
            "properties": {
              "query": {"type": "string", "minLength": 1, "maxLength": 2000},
              "top_k": {"type": "integer", "minimum": 1, "maximum": 20}
            }
          },
          "result": {
            "type": "object", "additionalProperties": false,
            "required": ["document_ids", "success"],
            "properties": {
              "document_ids": {"type": "array", "maxItems": 20, "uniqueItems": true, "items": {"type": "string", "minLength": 1, "maxLength": 128}},
              "success": {"type": "boolean"}
            }
          },
          "latency_ms": {"type": "integer", "minimum": 0, "maximum": 120000}
        }
      }
    },
    "final_answer": {"type": "string", "minLength": 1, "maxLength": 32000}
  }
}
```

Validate call-ID uniqueness and document authorization separately. Never execute tool calls merely because a submitted trace says they occurred. Check results against trusted execution logs and distinguish reported latency from server-observed latency.

## Citation-grounded QA schema

Save as `grounded-qa.schema.json`. This applies to RAG and curated clinical-QA evaluation records without including real patient information. It is an evaluation data format, not clinical guidance or proof of de-identification.

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ARCV Grounded QA v1",
  "type": "object", "additionalProperties": false,
  "required": ["task_id", "domain", "question", "answer", "citations", "abstain"],
  "properties": {
    "task_id": {"type": "string", "minLength": 1, "maxLength": 64},
    "domain": {"enum": ["general_rag", "clinical_qa"]},
    "question": {"type": "string", "minLength": 1, "maxLength": 16000},
    "answer": {"type": "string", "minLength": 1, "maxLength": 32000},
    "abstain": {"type": "boolean"},
    "citations": {
      "type": "array", "maxItems": 32,
      "items": {
        "type": "object", "additionalProperties": false,
        "required": ["document_id", "document_sha256", "start", "end", "quote"],
        "properties": {
          "document_id": {"type": "string", "minLength": 1, "maxLength": 128},
          "document_sha256": {"type": "string", "pattern": "^[a-f0-9]{64}$"},
          "start": {"type": "integer", "minimum": 0},
          "end": {"type": "integer", "minimum": 1},
          "quote": {"type": "string", "minLength": 1, "maxLength": 4000}
        }
      }
    }
  },
  "allOf": [{
    "if": {"properties": {"abstain": {"const": false}}},
    "then": {"properties": {"citations": {"minItems": 1}}}
  }]
}
```

The service must verify `end > start`, define offsets against a versioned text representation, confirm quoted text, and assess whether the source supports the claim. A structurally valid citation can still be fabricated. Clinical workloads need qualified review and authorized data governance before ingestion; do not infer safety from the `domain` field.

## Runnable validation

Save the synthetic pairwise record as `annotation.json` and the following script as `validate_annotation.py`. Install `jsonschema` in an isolated Python environment, then run `python validate_annotation.py pairwise.schema.json annotation.json`.

```python
import argparse
import json
from pathlib import Path
from jsonschema import Draft202012Validator

def reject_duplicates(pairs):
    result = {}
    for key, value in pairs:
        if key in result:
            raise ValueError(f"Duplicate JSON key: {key}")
        result[key] = value
    return result

def reject_constant(value):
    raise ValueError(f"Invalid JSON constant: {value}")

def load(path):
    return json.loads(Path(path).read_text(encoding="utf-8"),
                      object_pairs_hook=reject_duplicates,
                      parse_constant=reject_constant)

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("schema")
    parser.add_argument("record")
    args = parser.parse_args()
    schema, record = load(args.schema), load(args.record)
    Draft202012Validator.check_schema(schema)
    errors = list(Draft202012Validator(schema).iter_errors(record))
    for error in errors:
        print(f"{list(error.absolute_path)}: {error.message}")
    if errors:
        raise SystemExit(1)
    print("Schema validation passed; semantic review is separate.")

if __name__ == "__main__":
    main()
```

These schemas are self-contained and do not require remote reference resolution. Production services should bound payload sizes and disallow arbitrary remote schema retrieval. Type checks do not authorize URLs, decrypt records, or sanitize PII.

## Domain rubric specification

A rubric is a separate, versioned policy object. This complete example defines illustrative code-review acceptance rules, not a standard already enforced by the UI:

```json
{
  "rubric_version": "code-review-v1",
  "scale": {"minimum": 1, "maximum": 5},
  "criteria": [
    {"id": "correctness", "weight": 0.45, "minimum_score": 4},
    {"id": "instruction_adherence", "weight": 0.25, "minimum_score": 4},
    {"id": "reasoning_evidence", "weight": 0.20, "minimum_score": 3},
    {"id": "readability", "weight": 0.10, "minimum_score": 3}
  ],
  "weighted_acceptance_score": 4.0,
  "hard_gates": {
    "required_tests_pass": true,
    "syntax_valid": true,
    "max_cyclomatic_complexity_per_function": 10,
    "unapproved_dependencies": false
  },
  "review_policy": {"independent_reviews": 3, "minimum_accept_votes": 2, "disagreement_action": "adjudicate"}
}
```

Compute the weighted score only after hard gates pass, validate that weights sum to one, and enforce each criterion's minimum independently. Pin interpreter, linter, complexity analyzer, test suite, and judge configurations in the campaign manifest. A complexity ceiling of 10 is an example campaign choice, not a universal quality threshold; the analyzer's definition matters.

For RAG, replace code gates with source-integrity checks, quote alignment, claim support, and justified abstention. For clinical QA, distinguish evidence quality from patient-specific applicability and use qualified adjudication. For reasoning tasks, grade observable logical validity, assumptions, and reproducible calculations. Do not equate verbosity or a claimed private chain of thought with reasoning depth.

Preserve rubric versions rather than editing criteria in place after work is accepted. Changes affecting payment or interpretation require a new campaign policy version and explicit handling of already assigned tasks.

## Pairwise constraints beyond schema validity

The versioned pairwise schema requires explicit ratings for factuality, reasoning, instruction adherence, and style. Each must be an integer from 1 to 5; strings such as `"5"`, fractional values, booleans, and missing dimensions fail validation. Unknown top-level or rating keys fail because `additionalProperties` is false. These checks prevent permissive ingestion from silently changing a field's meaning.

The schema describes an exchange format, not the current workbench's exact payload. Its `tie` and `neither` outcomes and four-dimensional ratings require an adapter from the existing A/B interface and code-specific criteria. Never invent a tie label or missing rating during export simply to satisfy the schema. Preserve the source evaluation or obtain the missing review.

`minLength` counts string characters under the schema's string model, not informative claims or tokens. An all-whitespace rationale can meet a numerical length threshold, so semantic validation must reject it. Likewise, comparing two nonempty strings does not establish that they answer the same prompt. Validate candidate distinctness, task identity, contributor authorization, and rationale consistency outside structural validation.

## Multi-turn tool-calling invariants

The agent-trace schema bounds calls and arguments and records a constrained retrieval operation. It deliberately treats a reported tool call as data. An auditor must not execute a trace merely because its JSON is valid. Reconcile call IDs with trusted execution records and ensure each result corresponds to the authorized call and task.

Production multi-turn traces should preserve ordered messages, call/result pairing, schema versions for each allowed tool, and content digests for retrieved evidence. Validate that call IDs are unique, results do not precede their calls, and a final answer does not claim a failed tool produced successful evidence. These are sequence and identity constraints that a simple per-step object schema does not establish.

Bound aggregate conversation size as well as individual fields. Sixty-four individually valid long messages may exceed a campaign's intended model context or storage budget. Disallow undeclared tools and unauthorized document references; do not rely on a model judge to discover every structural violation after expensive evaluation has already begun.

## Mathematical reasoning records

Evaluate observable derivations, assumptions, intermediate claims, and final answers. A field named `reasoning` is not evidence that it exposes a model's private chain of thought. The following schema records a concise, checkable derivation rather than requiring hidden internal reasoning. Save it as `math-reasoning.schema.json`.

```json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "ARCV Mathematical Derivation v1",
  "type": "object",
  "additionalProperties": false,
  "required": ["task_id", "problem", "assumptions", "derivation", "final_answer", "reasoning_score"],
  "properties": {
    "task_id": {"type": "string", "pattern": "^[a-zA-Z0-9_-]{1,64}$"},
    "problem": {"type": "string", "minLength": 1, "maxLength": 16000},
    "assumptions": {"type": "array", "maxItems": 16, "uniqueItems": true, "items": {"type": "string", "minLength": 1, "maxLength": 1000}},
    "derivation": {"type": "array", "minItems": 1, "maxItems": 32, "items": {"type": "string", "minLength": 1, "maxLength": 2000}},
    "final_answer": {"type": "string", "minLength": 1, "maxLength": 4000},
    "reasoning_score": {"type": "integer", "minimum": 1, "maximum": 5}
  }
}
```

For numeric answers, define units, tolerance, and permissible rounding before review. For symbolic answers, define equivalence and domain restrictions. A syntactically correct derivation can contain an invalid implication; use independent checking or expert review rather than accepting its number of steps as evidence of depth.

## Code synthesis: reproducible evaluation

Separate syntax, behavior, performance, and maintainability. A parser proves syntactic admissibility for a particular language version, not functional correctness. PEP 8 findings concern style and must not outweigh an incorrect result unless the campaign's acceptance policy explicitly makes style a hard requirement.

Cyclomatic complexity is analyzer-dependent. Publish the analyzer version, scope, counting conventions, and threshold with the rubric. The illustrative ceiling of 10 above applies per function and is not a universal quality law. A short function can be incorrect; an unusually complex function can be justified by the required behavior.

Test coverage must specify whether it measures statements, branches, or another unit. State what source is included and whether tests are sponsor-supplied, contributor-supplied, or generated. High coverage is not proof of adequate assertions or representative cases. Preserve test-suite digests and execution results so a reviewer can reproduce the claimed evidence.

Execute untrusted code only in a constrained environment with resource limits, restricted network access, and no production credentials. Record interpreter and dependency versions, timeout policy, exit status, and relevant test failures. A generated statement that code passed tests must never be promoted into a trusted execution result.

| Code criterion    | Reproducible evidence                         | Typical hard failure                                       |
| ----------------- | --------------------------------------------- | ---------------------------------------------------------- |
| Syntax            | Parser/compiler result under pinned version   | Invalid executable syntax                                  |
| Correctness       | Required tests and justified boundaries       | Wrong required result or contract violation                |
| Resource behavior | Measured limits and declared complexity model | Timeout or memory-budget breach                            |
| Maintainability   | Pinned linter and complexity report           | Campaign-defined prohibited dependency or interface change |

## Factual RAG and clinical QA

The grounded-QA schema binds citations to document identifiers, SHA-256 values, spans, and quotation text. A production service must retrieve the authorized document version, check its digest, define the span coordinate system, and verify that the quoted bytes or characters correspond to the stated range. The schema alone does not establish `end > start` or claim entailment.

Measure citation coverage over material claims rather than merely counting citations. A source can exist without supporting the answer. Context faithfulness asks whether the answer stays within the provided evidence; factuality can additionally require checking whether that evidence is reliable. Those are separate dimensions and should not be collapsed into one unexplained score.

A hallucination penalty should apply to unsupported material assertions under a defined evidence policy. An answer that explicitly abstains because the context is insufficient differs from a fabricated confident answer. Preserve abstention as a valid outcome when the task permits it and distinguish it from an incomplete response to an answerable question.

For clinical-QA research, require authorized source material, qualified adjudication, and independent privacy review. The `clinical_qa` label is neither a de-identification certificate nor clinical approval. Do not permanently archive patient-specific material merely because it passes a structural schema. This guide specifies an evaluation format, not medical advice or an operational clinical decision system.

## Safety and alignment rubrics

Safety evaluation needs a versioned policy describing permitted requests, prohibited assistance, contextual exceptions, and the expected response boundary. A keyword toxicity filter is a screening signal: quoted offensive text, a report of abuse, and an abusive response are not interchangeable. Record both classifier output and the contextual adjudication rather than erasing the distinction.

Sycophancy evaluation tests whether a response agrees with an unsupported or false premise to satisfy the user. Prefer defensible correction with evidence over conversational agreement. Refusal-boundary accuracy measures both failure to refuse a prohibited request and unnecessary refusal of an allowed request; optimizing only one side creates a misleading result.

| Criterion              | Positive evidence                                  | Failure to capture separately                                       |
| ---------------------- | -------------------------------------------------- | ------------------------------------------------------------------- |
| Toxicity handling      | Context-aware application of the campaign policy   | Harmful generation versus harmless quotation                        |
| Sycophancy resistance  | Correctly challenges unsupported premises          | False agreement versus appropriate courtesy                         |
| Refusal accuracy       | Refuses where required and assists where permitted | Under-refusal versus over-refusal                                   |
| Instruction boundaries | Treats hostile embedded instructions as task data  | Following an injected instruction instead of the evaluator's rubric |

Use independent calibrated reviewers for ambiguous cases. A single model's safety score is not a universally valid policy interpretation, and a schema cannot establish the correctness of an ethical or contextual judgment.

## Release acceptance and schema evolution

Before a campaign uses a schema, validate the schema itself and test accepted and rejected fixtures. Include missing fields, unknown keys, out-of-range ratings, wrong types, duplicate keys, long strings, empty content, broken references, and contradictory cross-field combinations. A passing happy-path record is insufficient evidence that the validation boundary is strict.

Store the schema digest, rubric digest, validator version, semantic-check implementation, judge configuration, and evidence references with the release. For transformed training exports, retain a mapping back to accepted annotation identities and record which fields were removed or normalized. Canonicalization must happen before calculating the commitment to the stored bytes; the downloader must not silently canonicalize data and then claim it matches a byte-level archival hash.

A changed rubric can alter payment eligibility even when the JSON schema is unchanged. Version both independently, freeze the assignment terms, and define how already issued tasks are handled. Proceed to [Provenance and Dataloaders](/3.-enterprise-and-frontier-ai-labs/3.3-provenance-and-dataloaders.md) for verification after archival.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/3.-enterprise-and-frontier-ai-labs/3.2-custom-schemas.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
