Imagine you're training a new barista. You don't hand them a textbook on fluid dynamics — you make an espresso while they watch, then say 'your turn.' The catch is that watching you make an espresso simultaneously shows them hand trajectories, cup semantics, where to grip things, spatial layout of the machine, and the goal of 'hot coffee in cup.' A visual demonstration is radically overloaded with information, and nobody has clearly said which signal the robot is supposed to extract. That's the core problem this paper attacks. The committed claim: robotic in-context learning (ICL) has been operating without a proper problem definition. SimpleICL provides one — explicitly specifying what a robot should learn from a visual prompt — and then demonstrates that a minimalist architecture (visual prompt encoder + a low-cost data pipeline) achieves strong performance in both simulation and real-world manipulation without massive pre-training or specialized data infrastructure. This is a democratization claim, not a SOTA claim: the argument is that the field has been over-engineering the solution because it never properly defined the problem. The architecture is deliberately lightweight. A visual prompt encoder processes the demonstration video. No foundation-model-scale pre-training. No bespoke data infrastructure. The paper designs a low-cost data collection protocol that sidesteps the usual bottleneck of expensive teleoperation or large-scale demonstration datasets. The algorithmic family is supervised imitation learning with a vision-conditioned policy, leaning on the visual prompt as the task specification rather than language or reward shaping. What makes this paper interesting is the diagnostic work, not just the system. The authors run extensive experiments probing four discrimination properties of robot ICL: action discrimination (can the robot distinguish different trajectories?), semantic discrimination (does it understand object identity?), composition discrimination (can it parse multi-step tasks?), and affordance discrimination (does it infer how to interact with novel objects?). These experiments turn ICL from a vague capability claim into a testable, decomposable research program. That's the real contribution. The integrity picture is mixed in a characteristic way for robotics papers. Evaluation spans both simulation (likely ManiSkill or similar, though the abstract doesn't name the specific benchmark) and real-world experiments — which is better than simulation-only. But the baselines aren't named in the abstract, and without knowing which prior ICL methods they compare against (Voltron? RT-2? Octo?), the ladder is hard to assess. The promise to fully open-source data and training pipeline is significant; if delivered, it transforms this from a paper into a public benchmark. The milestone question for robot ICL broadly is whether visual prompting can generalize across task families and embodiments without per-task fine-tuning. SimpleICL demonstrates within-distribution discrimination, but the next concrete test is cross-embodiment transfer: can the same visual prompt encoder drive a different robot arm on a different manipulation task? That's the experiment that separates a nice framework from a paradigm. The obvious experiment not run: scaling to longer-horizon, multi-stage tasks in unstructured real environments (kitchen cleanup, not tabletop pick-and-place). The honest read is (a) — this is a resource constraint. Real-world long-horizon evaluation is expensive and slow. The framework's simplicity makes it a strong candidate for this test, and the authors are almost certainly planning it. The question is whether the minimalist architecture has enough representational capacity when task complexity scales.