Federico Viticci has been running local AI models for over a year, and his M5 Ultra Mac Studio review reads less like a hardware review and more like a field report from someone who built an entire research pipeline on local inference. The central claim is specific: the M5 Ultra's combination of 80-core GPU, 1.2 TB/s memory bandwidth, and Apple's unified memory architecture makes local models viable for sustained agentic loops — the kind where a model processes a system prompt, user context, tool descriptions, reasons through multi-turn exchanges, and dispatches subagents without choking on growing context windows. The numbers back this up convincingly. Token generation on the M5 Ultra runs ~70% faster than the M3 Ultra on average. Prompt processing — the bottleneck that killed local models in agentic workflows — is up 150%, a 2.5× improvement. Qwen3.8-Flash-Next clears 100 tokens/second on short prompts and sustains 60-85 tok/s even at 64K-256K context. These aren't synthetic benchmarks; Viticci ran these models inside Open Minis, Hermes Agent, and Codex with real tool calls and real multi-turn sessions. The proof-of-concept is his iOS and iPadOS 27 review research stack: 310 documents processed by a team of local agents running DeepSeek V4 Flash and olmOCR, operating 24/7 for 99 days, handling transcription, feature extraction, cross-referencing, and Notion API orchestration. Total cloud API cost: $0. That's the argument in one number. The agents did the research scaffolding; Viticci wrote the review himself. It's a clean separation of concerns that makes the local-AI case on economic grounds, not ideology. The RTX 5090 comparison is handled honestly. Viticci acknowledges the 5090 has higher memory bandwidth and retains an edge in raw throughput. But he frames the Mac Studio's advantage in terms the target audience actually cares about: thermal performance, noise, physical size, OS quality, and the unified memory architecture that lets you load larger models without VRAM partitioning headaches. This isn't a benchmarking article pretending to be objective — it's an experienced user explaining why the tradeoffs land where they do for his workflow. The review is notably transparent about its limitations. Viticci states upfront he's not an AI developer — he doesn't train or fine-tune models. He explicitly says this setup is bleeding edge and not something he'd recommend to anyone who just wants to pay $20/month for Claude. That kind of audience honesty is rare in tech reviews, and it makes the enthusiasm feel earned rather than performed. What's missing is the cost-benefit math stated plainly. The M5 Ultra Mac Studio with 256 GB starts at a price point that could fund years of frontier API access. Viticci acknowledges this but doesn't run the numbers side by side. For the 99-day, 310-document workload, the local approach clearly won on cost. But for someone running lighter workloads, the breakeven calculation would look very different. The review would be stronger with that analysis included. The piece functions best as a workflow document — a detailed account of how one power user integrated local AI into a professional writing operation and what hardware made it viable. It's less useful as a buying guide and more useful as a proof-of-existence: this workflow is now possible, here's what it looks like, here's what it costs in hardware, and here's the performance floor you need to make it work.