Science & Analytics

The Scientist

The Signal

Xiaomi keeps GRPO learning after all 16 rollouts pass by scoring LLM-written checklists.

The binary pass/fail reward gets multiplied by two checklist scores, and a listwise grader takes over where passes are rare. The recipe cost $2.6M of compute across 7,000+ environments and puts MiMo-V2.6-Pro level with Grok 4.7. Parity is the research result, and $0.13 a task against $2.73 is the production one you would budget around.

In Play

  1. Open-weight recipes beat open weights

    Xiaomi's MIT-licensed MiMo-V2.6-Pro ties Grok 4.7 at 46 on the Artificial Analysis Intelligence Index for $0.13 per task against $2.73, a 21x gap, per The Batch. MiMo-V2.6-Flash tops CyberBench at 75.36% for $0.05, roughly one-eightieth of Claude Fable 5.1's cost. Xiaomi also published the full RL recipe: $2.6M of compute and 7,000+ environments. The differentiator is now reward design and serving architecture, not the weights.

  2. Open-weight safety is a removable default

    Anthropic's Frontier Red Team and NIST's CAISI independently rated Zhipu's open-weight GLM-5.3 as near-frontier at exploit development, per Matt Johansen. CAISI called it roughly four months behind the US frontier. Its guardrails degrade from 64% bypass with a cover story to 100% once refusals are ablated from the weights, and ablated builds appeared within days. If you fine-tune or deploy open weights, the model card's refusals are a soft default, not a control.

  3. Your eval harness is the only arbiter

    Outflank found that Codex and Claude Code rewrite MCP tool definitions, annotations and search client-side, so the schema your model conditions on can differ from what your server sends. Gemini 4 Argon, GPT-6.1 Sol and Claude Sonnet 5.5 all list at $2/$10 per million tokens, so price no longer separates models; only a paired eval on your own tasks does. Treat the client as an experimental factor and log the post-client schema before you trust any eval delta.

  4. A macro regime shift breaks macro features

    The US 10-year Treasury yield hit 5.34%, its highest since 2002. The UK 30-year passed 6% for the first time since 1998, and oil ran from about $70 to about $100, per Morning Brew. Those readings likely sit outside most production training windows. Gradient-boosted trees cannot extrapolate past their largest training split, so a 5.34% input scores as the training maximum. The model silently clips the new regime, with no error signal until labels land months later.

Deep Dives

  1. Xiaomi published the reward fix that keeps GRPO learning

    Two mechanisms sit under the cost-capability collapse: a reward that multiplies pass/fail by LLM-written checklists, and a KV cache small enough that agent cost now turns on cache hits, not output.

    GRPO (group-relative policy optimization) normalizes rewards within each group of attempts. When all 16 rollouts for a prompt pass, a binary pass/fail reward scores every attempt identically. Each attempt then carries zero advantage and contributes no gradient, though the compute…

    3 action items

    ●
  2. Open-weight guardrails hit 100% bypass — plan controls as if they're gone

    Two independent evaluations, one from a lab with incentive to inflate the threat and one from a government body with none, agree an open model now builds exploits near the frontier, and its refusals vanish once you touch the weights.

    What converts this from vendor marketing into signal is the bias test it survives . Anthropic's Frontier Red Team has a commercial incentive to make freely downloadable weights look dangerous; its writeup ends with a pitch to put Claude in…

    3 action items

    ●
  3. The frontier priced to parity, so your eval harness is the only arbiter

    With three frontier models at identical list prices and a parser claiming 'up to 20%' fewer errors, the deciding evidence is a paired test on your own data — and the MCP client is quietly rewriting the schema you think you measured.

    The least-discussed detail in the Outflank write-up is the one with the largest effect on experimental validity. Outflank found that Codex and Claude Code rewrite MCP tool definitions, annotations and search mechanisms client-side . The schema the model conditions on…

    3 action items

    ●

The edition continues

Take the signal into the room.

Sign up or log in to read all 3 deep dives in full, plus the final take.

Read the full edition

Continue with LinkedIn