Two Scorers That Flip on Inputs They Should Ignore
An answer-order swing and a dialect-driven flag look unrelated, yet one cheap experiment most teams skip would size both before they reach a production threshold.
Why order leaks into the score
Archer Hume reconstructed Jev without weights or documentation. He used latency profiling, token accounting, option-order probes, reference-card placement and tokenizer fingerprinting, and none of 192 public tokenizers matched. His inferred design is a causal transformer that processes options listwise, so each option's representation conditions on the options placed before it. Under that design, order dependence is structural, not sampling noise. Any judge that reads candidates in sequence carries the same exposure unless it was trained on permuted orders. That includes multiple-choice scorers, pairwise LLM-as-judge setups, and classifiers prompted with a label list.
The item count behind the order example was not reported. Treat it as proof the effect exists, not as an estimate of how often it flips your decisions.
The detector case is the same bug with a different nuisance variable
For Pangram, the variable that should not matter is authorial style. Orélien says his prose uses Haitian and Caribbean repetition and rhythm. Critics quoted by Morning Brew say detectors flag too broadly and train mainly on European text. The vendor figure arrives with no unit, corpus, language mix or interval. By the rule of three (zero errors in n trials bounds the true rate near 3/n at 95% confidence), backing 1 in 24,000 takes about 72,000 human-written negatives with zero false positives. That count applies to each slice separately.
The stacked evidence is weaker than it looks. Several passages flagged, and academics using similar tools found high AI percentages. But if style drives the flag, every passage carries the trigger. The chance that all passages flag for a human author is then roughly the chance this style trips the detector, not the per-passage rate raised to the k-th power. Detectors trained on overlapping corpora also share error modes. Base rates make it worse. With an assumed 1% prior and 95% true-positive rate, the vendor's rate gives about 99.6% positive predictive value (the share of flags that are correct). A 2% false-positive rate on this slice gives about 32%.
| Dimension | Jev option order | Pangram dialect |
|---|---|---|
| Nuisance variable | Position of answer options | Author's dialect and rhythm |
| Reported evidence | One example, n unknown | Aggregate vendor rate, no interval |
| Error structure | Built into listwise processing | Correlated across one author's passages |
| Test that settles it | Flip rate under permutation | Per-slice false-positive rate with intervals |
The move: publish a flip rate beside accuracy
Sizing is cheap. Estimating a flip rate near 5% within ±2 points at 95% confidence needs about 460 items (1.96² × 0.05 × 0.95 / 0.02²). Concentrate them near your operating threshold, because flips elsewhere do not change decisions. Score each item under all cyclic shifts or 8–16 random orders. Then report the flip rate at threshold, the per-item score range, and the mean absolute change under reversal.
The open clone jaredpalmer/kev makes the ablation affordable. It pairs a Qwen base with a rank-16 LoRA adapter and a small pointer head, and a fine-tune costs about $1 on an H100. TypeSafe's SDK talks to it unchanged. Kev-27B scores 0.848 against Jev's 0.857 on unseen sources, with no reported interval. At roughly 1,000 items the 95% interval would be about ±0.022, so the gap proves neither parity nor loss. Kev-9B trails on MMLU, 0.74 against 0.90. Kev was trained on at most 384 state tokens, while its server accepts 65,536. It also ships unauthenticated, so never bind it to 0.0.0.0 without KEV_API_KEY set.
For AI-text filters in your data pipeline, what's at stake is coverage. A detector that reacts to repetition and rhythm will disproportionately remove non-standard dialect and ESL text. That quietly narrows the linguistic range of your training set.
A scorer's accuracy tells you nothing about its invariances; if you have never shuffled the options or sliced by dialect, you do not know your flip rate.
What to do
Permutation-test every production LLM judge or multiple-choice classifier this sprint: sample ~460 items near the operating threshold, score each under cyclic shifts or 8–16 random option orders, and report flip rate at threshold next to accuracy.
Audit per-slice false-positive rates, with Wilson 95% intervals, for any AI-text detector filtering your training corpus or annotator output before the next data refresh, using a human-verified hold-out of non-standard and ESL text sliced by language, dialect and register.
Run a permutation-augmentation ablation on a self-hosted Kev this quarter, fine-tuning with and without shuffled option orders and bucketing results by state-token length, with KEV_API_KEY enabled.