Ctrl + K
Log In
Train LLMs to reason by only rewarding correct answers — never punishing wrong ones — and it works as well as GRPO | BedrockNews