# Jev vs Laya — frozen before Laya exam inference

Use all 996 existing adapted items (58 discovery, 938 follow-up) and the original Jev 1.13.0 responses. Preserve question text, state, option order, and exact-match grading; do not regenerate distractors or use answers in inference. Report discovery, follow-up, counting/probability, and original task categories separately. This compares an archived Jev run with a new Laya run, not simultaneous production latency.

Primary Laya checkpoint: convaiinnovations/laya English, pinned revision in model-revision.json, SDK 0.3.4. This is the documented English checkpoint, selected before observing scores. No fine-tuning, answer feedback, external mathematical tools, or best-of-model selection.

Run two explicitly separate conditions:
1. `default`: identical public state/questions through unmodified Laya SDK defaults. Audit internal instruction, option and state truncation per judgment. This measures drop-in behavior, not equal effective context if truncated.
2. `full_input`: same public state/questions and SDK serialization template, with deterministic no-truncation token assembly. No text rewritten or reordered; no correct answers read. Evaluate one judgment at a time to bound memory. Maximum encoder limit is verified before execution; over-limit judgments are recorded as unsupported rather than clipped. This is a disclosed adapter outside stock SDK preprocessing, not a claim about default Laya performance.

Both conditions and all failures will be reported, without choosing whichever scores higher as the sole result. Local inference latency excludes model download/load and is not comparable to Jev hosted HTTP latency; report hardware/device, warmup separately, per-judgment times and total elapsed. Report full distributions and native confidence; native confidence definitions can differ between systems. No calibration fit on exam answers.

Source adaptation limitations (public training exposure, solution-aware distractor authoring, supplied-proof/answer-length cues, grouped items, no partial credit) remain unchanged. Offline scoring alone reads private keys. Preserve original evidence files untouched.
