- Sources: primary, discussion
- Summary: CTGT distilled from DeepSeek V4 Flash as teacher into GPT-OSS-120B as student and reports a matched censorship gap of +45.45 for the teacher across 76 core-political prompt pairs, +32.02 across 152 pairs pooled, scored by four judges from four labs, against +2.58 for the Flash-taught student. The 20B model is released as open weights on Hugging Face, while the 120B that carries every headline score is reachable only through a playground, alongside the matched-pair evaluation, the prompts, and the judge rubric. CTGT states its own limits: safety training and refusal behaviour were not tested, one run did not reproduce at an identical seed, 84.87 percent re-running to 82.35 percent, and the self-distilled arm is described as gap-consistent and truncation-sensitive rather than fully reproduced.
- Why it matters: CTGT is an interested party, the same post benchmarks its own 120B at 83.61 percent on FinanceReasoning above Kimi K3 and Inkling and ends in a partnership solicitation, so the released 20B weights and the published rubric are the parts a team can check independently and the 120B numbers are not among them.
send feedback on this story