Signal2026-07-23
Papers With Code

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Part of

Advanced Inverse Reward And Prompt Engineering