Imagine you're following GPS directions through a city, but the GPS is slightly wrong about which streets are one-way. If you blindly follow the turn-by-turn gradients, you'll drive into walls. But if instead you ask the GPS to just rank every intersection by how close it gets you to the restaurant, you can pick routes yourself without trusting its street-level turn signals. FERPO does exactly this to reinforcement learning critics: it uses the critic's value estimates to construct a target distribution over actions, then fits the policy to that target — never once differentiating the critic with respect to actions. The committed claim: you can do on-policy maximum-entropy RL in continuous control without taking action gradients through the critic, and match or beat methods that do. The paper derives an optimal target action distribution from a policy-improvement objective regularized by entropy and KL divergence, then fits the actor to this target by minimizing a forward-KL objective estimated via self-normalized importance sampling (SNIS). The KL regularization keeps the target distribution close to the rollout policy, which in turn keeps importance weights well-behaved — a practical necessity for stable training. The ladder context matters here. The primary comparison point is REPPO (Relative Entropy Pathwise Policy Optimization), which does differentiate through the critic. FERPO reports competitive performance on MuJoCo Playground and ManiSkill benchmarks, with sample-efficiency gains and faster actor updates than REPPO. The paper also compares against standard baselines in these environments, though the abstract doesn't name specific reward numbers — the claim is "competitive performance" rather than definitive SOTA on a single headline metric. Architecturally, FERPO belongs to the on-policy, entropy-regularized policy optimization family — think SAC's maximum-entropy objective but without the pathwise gradient trick. It leans on importance sampling rather than reparameterization, which means it trades the bias of potentially unreliable critic gradients for the variance of importance weights. The KL constraint is the load-bearing structural choice: it simultaneously promotes exploration (forward-KL covers multiple modes instead of collapsing onto one) and keeps the variance of the SNIS estimator tractable. The integrity story is reasonable but not airtight. The benchmarks (MuJoCo Playground and ManiSkill) are community-standard continuous control environments, and ablations are reported. Code is released on GitHub, which is a strong signal. However, these are same-team simulation results — no independent replication exists yet, and the environments, while standard, don't include the hardest manipulation or locomotion benchmarks in the field. The computational speedup claim (faster actor updates than REPPO) is a welcome practical metric, but wall-clock comparisons depend heavily on implementation details. The forward-KL vs reverse-KL distinction is the conceptual payload here, and it has legs beyond this paper. Reverse-KL objectives (standard in most policy gradient methods) are mode-seeking — they can collapse onto a single high-value mode and ignore others, which hurts exploration. Forward-KL is mean-seeking — it tries to cover all modes, which promotes exploration but can spread probability mass across suboptimal regions. FERPO's design bet is that the exploration benefit outweighs the covering cost, especially when KL regularization keeps things tight. The milestone question is straightforward: continuous control is a crowded field, and the real test is whether gradient-free-critic methods can scale to high-dimensional manipulation tasks with contact-rich dynamics — the kind of tasks where critic gradients are most unreliable and where the payoff of avoiding them is highest. The paper demonstrates viability on standard benchmarks; the next step is demonstrating clear wins on problems where pathwise methods visibly struggle.