> For the complete documentation index, see [llms.txt](https://docs.arcv.network/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.arcv.network/4.-the-4-stage-verification-engine/4.2-interaction-telemetry.md).

# 4.2 Interaction Telemetry

Stage 2 evaluates interaction plausibility and contributor calibration. It combines assignment timing, privacy-minimized event features, hidden benchmark performance, and independent overlap judgments. These signals support risk-based admission; they cannot prove that a wallet represents one human or that a browser is free of automation.

**Implementation status:** the current workbench has no operational telemetry collector, dwell-time admission gate, gold-trap scheduler, or strike ledger. Its audit messages advance at 375 ms intervals and its reputation metrics are illustrative. This chapter specifies the proposed service and explicitly labels example policy parameters.

## Evidence architecture

```
Assignment coordinator -----> server issuance ledger
          |                             |
          | real task or hidden         | authenticated receive timestamp
          | benchmark assignment        |
          v                             v
Contributor browser -> bounded event batches -> timing / cadence features
                                                |
Protected benchmark store -----------------> reference scorer
                                                |
Independent review assignments ------------> deviation analysis
                                                |
                                     versioned policy evaluator
                                      /          |          \
                                  eligible     review      hold
                                      |          |          |
                                      v          +----------+
                                    Stage 3      reputation ledger
```

The coordinator must record when an assignment becomes available and when it receives the submission. Reference labels remain server-side. Do not expose benchmark identities through response metadata, HTML, JavaScript bundles, client logs, or Copilot context.

Each event batch binds an assignment, session, sequence number, and schema version. Bound event count and size; reject impossible sequencing and cross-assignment reuse. TLS and request authentication protect transport and attribution, but client-origin event data remain forgeable. A script can generate realistic input events or wait out a timer.

A deployment should distinguish optional interaction evidence from indispensable protocol state. Contributors using keyboard navigation, assistive technology, speech input, or a precomposed correction may have little pointer movement. A missing mouse trace must not be equated with an absent human.

## Dwell-time and decision velocity

Let the coordinator issue the assignment at server time t\_issue and receive the submission at t\_receive:

```
T_server = t_receive - t_issue
T_active = measure(union of valid foreground intervals)
v_read   = 60 * W / T_active
v_decide = J / T_active
```

Times are measured in seconds, W is the declared readable word count, and J is the count of substantive decisions. Velocity is undefined when active time is zero or invalid; do not substitute an artificial denominator and report a meaningful reading speed.

Calculate foreground intervals using a client monotonic clock, clip them to valid assignment-relative bounds, and merge overlaps before measuring their union. Absolute client and server clocks are not interchangeable. Server elapsed time includes network and idle delay; client active time is only a reported estimate. Neither measures comprehension directly.

The proposed minimum gate rejects a newly issued sub-second completion from immediate certification. An illustrative gate uses a server minimum of one second as an absolute floor and a calibrated task-class duration above that floor. The gate is not implemented by the frontend. Previously issued assignments, resumes, retries, and task caching must retain their original timing identity so transport retries do not look like new instant completions.

An excessively short valid duration produces a held submission with a reason such as INSUFFICIENT\_INTERACTION, not an automatic allegation of poisoning. A malformed timestamp or missing event batch produces an instrumentation review. Waiting longer can bypass a duration floor, so passing the gate is necessary evidence under a policy, not sufficient proof of human work.

Reading velocity must be calibrated per modality. Code, tables, equations, languages, and familiar repeated instructions have different inspection costs. A universal words-per-minute cutoff would punish expertise while remaining easy for bots to evade. Estimate thresholds from held-out legitimate sessions, including accessible workflows, and retain the model and calibration window used.

## Cadence, focus, and movement features

| Feature             | Calculation or observation                         | Interpretation limit                                    |
| ------------------- | -------------------------------------------------- | ------------------------------------------------------- |
| Action cadence      | Inter-event intervals and their distribution       | A script can add jitter; UI timers can cause regularity |
| Foreground dwell    | Union of relevant focused/visible intervals        | Foreground presence does not establish reading          |
| Rating activity     | Initial choice, changes, and final value           | Requiring unnecessary changes rewards fake activity     |
| Revision metadata   | Insert/delete counts, paste events, revision count | Counts do not prove independent reasoning               |
| Pointer movement    | Direction bins, pauses, normalized entropy         | Keyboard users may have no samples                      |
| Session consistency | Sequence continuity and assignment binding         | Signed transport does not authenticate a physical human |

For m nonempty-duration pointer direction bins with probabilities p\_k, use a fixed declared bin count M greater than one:

```
H_mouse = -sum(p_k * ln(p_k)) / ln(M)
0 * ln(0) := 0
CV_time = standard_deviation(delta_t) / mean(delta_t)
```

Entropy lies in \[0,1] for a normalized histogram. No pointer samples means unavailable, not zero entropy. Cadence coefficient of variation requires a positive mean and enough intervals. Low entropy or low CV can describe both automation and efficient repetitive human work. High entropy can be synthesized. Do not advertise these features as headless-browser detection guarantees.

Record aggregates rather than full keystroke contents or screen recordings. Deleted text can contain secrets that never reach the final submission. Focus duration and revision counts can serve an integrity purpose without retaining that content. Define a short, campaign-approved telemetry retention period; confidential traces must not enter the permanent training archive. Public lineage can bind a sanitized decision receipt without revealing raw behavioral data.

## Blind benchmark injection

The requested 10–15% range is a proposed benchmark allocation policy. The current enterprise selector offers 5%, 10%, and 20%, with a 10% default. Those UI values do not execute injection. Resolve campaign policy into one explicit probability or quota before launching a real scheduler.

For independent sampling, use server-side cryptographically secure randomness:

```
p in [0.10, 0.15]
u <- uniform random value in [0,1)
benchmark_assignment := (u < p)
K ~ Binomial(N, p)
E[K] = N*p
Var(K) = N*p*(1-p)
P(K = 0) = (1-p)^N
```

With N=100 and p=0.10, the expected count is 10 and variance is 9. With N=20 at that rate, the probability of no benchmark is approximately 12.16%. An average injection rate is therefore not guaranteed coverage for each short session.

A quota scheduler instead places round(p\*N) benchmark slots into a server-side random permutation. Define rounding, skip handling, resumptions, and the denominator. Count issued assignments separately from completed/scored ones. Avoid deterministic retry behavior that reveals a benchmark or enables unlimited skipping until a contributor reaches an ordinary task.

Match benchmark modality, difficulty, formatting, and response latency to ordinary tasks. Rotate leaked items and enforce exposure budgets. Use assignment-level blinding rather than merely hiding a badge. Do not pay production data bounties twice for synthetic calibration work; benchmark compensation, if offered, needs a separately funded campaign policy.

## Ground-truth construction and veracity

A mathematically grounded benchmark must have a reviewable reference: an independently checked derivation, an exhaustive result over a finite domain, or a formally established property under stated assumptions. Executable tests provide useful code evidence, but passing finite tests is not a proof for all inputs.

Open-ended preference labels are adjudicated references, not mathematical truths. Store acceptable alternative outcomes, criterion ranges, domain constraints, and an ambiguity disposition. Remove invalid or disputed benchmarks from scoring, and recompute affected reputation after correcting the reference.

For binary accepted-label alignment:

```
correct_i = 1 if contributor_outcome_i belongs to accepted_labels_i else 0
accuracy  = sum(correct_i) / n
```

Only valid scored benchmark assignments enter n. When n=0, return insufficient evidence. Report the reference version, domain, window, and sample count with accuracy; 100% from one trial is not the same evidence as 100% from hundreds.

A Wilson interval illustrates small-sample uncertainty. With observed proportion a=c/n and z=1.96:

```
d = 1 + z^2/n
center = (a + z^2/(2*n)) / d
radius = z * sqrt(a*(1-a)/n + z^2/(4*n^2)) / d
interval = [center-radius, center+radius]
```

For 9 correct out of 10, the interval is approximately \[0.596, 0.982]. Correlated or repeatedly exposed items violate the simple independent-binomial interpretation. The interval must not be presented as certainty about honesty.

## Peer deviation and independent agreement

For contributor i, compare criterion score s\_ij with a leave-one-out peer median g\_ij, using nonnegative weights summing to one:

```
D_i = sum_j(w_j * abs(s_ij - g_ij) / 4)
0 <= D_i <= 1
A_i = matching valid peer labels / number of valid peer labels
```

Exclude the contributor's own vote from the peer reference. If too few independent peers remain, report insufficient evidence. For an adjudicated gold reference, replace the peer median with the reference criterion score and identify that different evidence source.

D\_i measures rubric distance, not maliciousness. A\_i measures label agreement, not truth. A coordinated majority can be wrong; a correct minority may detect an edge case that others missed. Cluster-aware assignment and external reference checks reduce dependence, but a wallet address does not establish a unique person. Shared networks alone are also insufficient grounds to label accounts as sybils.

For categorical outcomes, count A, B, tie, and abstention according to the campaign policy; do not average category codes. Abstentions should reduce coverage rather than automatically count as endorsements of the majority.

## Reputation updates and escalation

The following is an explicit proposed policy model, not deployed wallet enforcement. For an adjudicated benchmark event, maintain bounded trust R in \[0,1]:

```
R_next = (1-alpha)*R_current + alpha*y
alpha = 0.10
y = 1 for a confirmed accepted benchmark result, otherwise 0
```

One confirmed failure from R=0.95 produces R=0.855. Another produces 0.7695. This sensitivity is why trust updates require valid references and an evidence window; a raw EWMA alone must not trigger permanent exclusion.

An example campaign escalation policy makes the operational consequences explicit:

| Evidence                                                       | Proposed consequence                                             |
| -------------------------------------------------------------- | ---------------------------------------------------------------- |
| One ordinary benchmark mismatch                                | Explain criterion failure; update calibrated score               |
| At least 20 valid recent benchmarks and accuracy below 85%     | Freeze new assignments pending requalification                   |
| Successful independent requalification                         | Restore eligible assignments under a documented recovery rule    |
| Repeated adjudicated integrity violation after requalification | Restrict the affected campaign/domain                            |
| Confirmed intentional poisoning or coordinated account abuse   | Adjudicate wallet disqualification with evidence and appeal path |

This example's 20-item minimum and recovery procedure are specification choices, not existing product guarantees. Disqualification is an assignment-access decision. Permanent wallet exclusion should require confirmed severe abuse, not merely a low score, missing mouse movement, or a disputed preference.

Maintain an append-only decision ledger with policy version, safe evidence references, scope, review date, and appeal outcome. Correct invalid benchmark penalties retroactively. Operational timeouts and telemetry defects are not malicious strikes.

The current contract has no stake, reputation, slashing, or wallet-blacklist mechanism. Off-chain exclusion cannot claw back a paid reward. On-chain slashing would require separately specified stake custody and contract logic. Stage 2 eligibility advances the immutable candidate version to [Stage 3: Autonomous Validator Agent](/4.-the-4-stage-verification-engine/4.3-autonomous-validator-agent.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.arcv.network/4.-the-4-stage-verification-engine/4.2-interaction-telemetry.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
