Imagine you're trying to reverse-engineer a secret recipe by ordering dishes from a restaurant, but you can only order a few items, the menu only says 'chicken' or 'fish' (no ingredient lists), and you don't even know if they're using a wok or an oven. That's the threat model Dagger operates under: a black-box GNN API that returns only hard class labels, a tight query budget, and zero knowledge of the victim's architecture. The claim is that you can still build a surrogate model that closely mimics the victim — and do it better than any existing attack. The core mechanism is a two-phase decoupling strategy. Phase 1 separates information propagation from label supervision. Instead of needing the full graph structure (which a real API wouldn't hand over), Dagger pre-trains a surrogate encoder using decoupled message passing over sparse local subgraphs, handling isolated nodes that would otherwise be garbage inputs. It then uses manifold-level node mixup — interpolating between node representations to synthesize continuous training signals from the discrete hard labels the API returns. This is the key engineering insight: you can't get gradients from hard labels, but you can manufacture a smooth supervision landscape by mixing representations. Phase 2 freezes the encoder and fine-tunes only the classification head. The problem here is class imbalance: with a tiny query budget, some classes may barely appear or not appear at all. Dagger uses class-balanced sampling paired with logit adjustment to correct for this skew without burning additional queries on the victim API. The two-phase split means each phase can solve its own problem cleanly — structure learning doesn't fight with class balancing. The results are tested across four benchmark graphs (Cora, CiteSeer, Amazon-Photo, Coauthor-CS) and four victim backbones (GCN, GAT, GraphSAGE, APPNP). The headline numbers: up to 18.16% higher fidelity than the strongest baseline, using 12.23× fewer queries. Fidelity here means agreement between the surrogate's predictions and the victim's — the metric that matters for a functional clone. The backbone-agnostic results are important because real-world attackers don't know what architecture the target is running. The threat model deserves credit for its constraints. Most prior GNN stealing work assumes soft-label outputs (full probability vectors), large query budgets, and sometimes knowledge of the victim's architecture family. Dagger drops all of these. The hard-label-only setting is genuinely harder — you lose the rich gradient signal that soft labels provide. The authors identify four specific challenges (sparse structures, insufficient supervision, class imbalance, backbone mismatch) and address each with a named mechanism. This is clean problem decomposition. Integrity-wise, this is a same-team simulation study on public benchmarks. The baselines compared include recent GNN stealing methods, though the paper is under review and independent replication hasn't happened. The benchmark graphs are standard community datasets, not cherry-picked. The 12.23× query reduction is measured against the strongest baseline, not the weakest — a good sign. However, all evaluation is on citation and co-purchase graphs; no evaluation on heterogeneous or large-scale industrial graphs. The successor question is obvious: does Dagger work on graphs with millions of nodes and heterogeneous edge types — the kind deployed in production recommender systems and fraud detection? The authors didn't run this, most likely because (a) the benchmark graphs are standard for the subfield and (b) scaling to production-size graphs would require significant compute and possibly proprietary data. The defense side is also unaddressed — how well do existing GNN watermarking or fingerprinting defenses detect Dagger surrogates?