Imagine you run a restaurant kitchen where every dish gets inspected by a food critic before it leaves. The critic's opinion changes depending on their mood, but the health code doesn't. Now imagine the health inspector shows up and asks: 'Was every dish safe on Tuesday?' If the critic's notes are tangled up with the health code, you can't answer — you'd have to re-taste everything. This paper separates the tasting notes from the health code, so you can re-evaluate any dish against any standard, at any time, without re-tasting. The committed claim: point-in-time audits by LLM-based security tools (like RepoAudit) cannot provide continuous assurance for merge decisions because both the evidence (non-deterministic LLM outputs) and the interpretation (policy, jurisdiction, risk appetite) shift independently. The proposed fix is the Policy-Evidence-Execution Separation Pattern, implemented as the Trustworthy AI Posture (TAIP) Assurance Engine. This is an architectural contribution — a separation-of-concerns pattern for CI security gates — not a new detection method. The architecture is straightforward: a versioned 'Posture Tree' binds retained LLM audit evidence to policy profiles, so when policy changes (new compliance rules, new risk threshold) or model context changes (switching from gpt-4o-mini to gpt-4.1), assurance posture recomputes over already-collected evidence rather than re-running the auditor. This is the key insight — decouple the expensive, non-deterministic inference step from the cheap, deterministic policy evaluation step. The pattern is model-agnostic by design. The evaluation uses 80 unmodified RepoAudit executions across two OpenAI model configurations on a fixed Python Null Pointer Dereference benchmark. The headline numbers are latency-focused: 1.1 ms maximum policy-to-posture latency for a single context, 1.62 s for full recomputation across 1,000 independent decision gateways on a single host, against a declared 5 s budget. These are assurance-layer numbers only — they explicitly exclude the cost of actually running RepoAudit or calling OpenAI. The ladder question is tricky because there isn't a named prior system doing exactly this job. The closest comparison is ad-hoc CI gate scripts that re-run auditors on policy change, which is what this paper argues against. The paper doesn't benchmark detection quality at all — it benchmarks the assurance recomputation layer. This means the contribution lives or dies on whether the separation pattern is adopted, not on whether it beats an existing system on a shared metric. Integrity is mixed. The benchmark is narrow — one tool (RepoAudit), one bug class (Null Pointer Dereference), one language (Python), two models. The 80 executions provide evidence variance data, but the paper doesn't test adversarial policy configurations, doesn't test with non-OpenAI models, and doesn't test at enterprise scale beyond the single-host 1,000-context measurement. The latency numbers are clean but the generalizability envelope is small. Code availability and independent replication status are not stated in the abstract. The real value here is the pattern, not the numbers. If you're building or evaluating LLM-powered security gates in CI pipelines, the separation of policy from evidence from execution is the takeaway. The numbers prove the pattern is computationally feasible, not that it's battle-tested. The next credible milestone is adoption by a major CI platform or an enterprise deployment with multiple auditor tools, diverse bug classes, and adversarial policy churn.