Qwen3.6-27B ExpThink

GRPO post-training of Qwen3.6-27B with a correctness-gated reward, aimed at shorter reasoning without losing accuracy.

Built with: LLM Post-Training, GRPO

Weights: ryankim17920/qwen3p6-27b-expthink-step60

  • Cut thinking tokens 52% in-domain / 34% OOD
  • Truncation 6.0% → 2.6%, while OOD accuracy rose 84% → 86%
  • Trained with PyTorch FSDP on 8×H100 via verl, with SGLang tensor-parallel rollouts

July – August 2026