# SemIf exam comparison — fixed before inference

Run all 996 frozen adapted items (1,000 independent API judgments) from the original discovery and thirteen-exam follow-up. Retain the original Jev 1.13.0 and Laya comparisons; no new distractors, rearranged choices, hints, solutions, or answer feedback.

SemIf source: ca3ba65f142967030ecb453346e94d6f476a69df (TheoLeeCJ/SemIf).
Model: Qwen/Qwen3.5-4B, revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Backend: official MLX direct scoring, original checkpoint precision, no quantization, thinking disabled by the official prompt. Native option-logit readout; no generated tokens or computational tools. Max prompt tokens 4096, enforced without truncation. Unsupported/failed requests count as incorrect and are retained. No score-driven retry, selection or changes. The original option ordering is preserved.

Adapter: one SemIf row per original independent judgment. Preserve the original `state` object exactly and map `instructions` to `question`. Choice criteria become the same ordered option IDs/descriptions. Noul judgments map to two options, false then true, using generic false/true descriptions; take the true probability with the original >=0.5 rule. Original select-all grading remains exact-set. SemIf's system prompt, JSON field names, chat template and answer-token labeling are its official native wrapper; this changes the interface serialization, not the mathematical question or alternatives. Record hashes and raw outputs.

Inference reads only frozen public request files. Offline scoring uses the existing answer evidence. Single first pass. Prefix reuse is not enabled, avoiding execution-shape probability changes. Model load/download and one unrelated warmup are outside timed inference. Report local timing separately from hosted Jev batch HTTP latency. Scores are conditional option probabilities, not calibrated correctness confidence.

The original adaptation limitations remain: public exposure may overlap training; alternatives were written after reading solutions; supplied proofs and answer-length cues can help recognition; no student-equivalent partial credit. Report all 996 items, the original 58, and the identical preexisting 380-item counting/probability subset.
