The argument is structural, not philosophical: if every major chatbot ships with a disclaimer saying 'I make mistakes, check my work,' but provides zero tooling to actually check that work, the disclaimer is liability management, not product design. The author contends this gap between stated limitation and actual interface reveals that frontier labs are not building productivity tools — they are building confidence machines. The most concrete proposal is a verification worksheet: a two-column interface where LLM output sits alongside human-authored notes, with explicit checkboxes per claim. The author argues that without this friction, the path of least resistance — trusting the output — becomes the default workflow, particularly in coding environments where AI-generated code gets shunted to reviewers who assume the 'author' already checked it. This is a dark pattern, not a feature gap. On citations, the critique is equally specific. Current chatbots present sources as tiny inline annotations showing only domain names in near-illegible font sizes. The author wants citations presented as large, first-class objects with publication dates, author names, and — critically — verbatim quotations extracted by regular programs rather than LLM-generated summaries. AI editorialization should be subordinate to the source material, not the reverse. The piece takes aim at anthropomorphization as a design failure. First-person language, apologies upon correction, and verbose emotional responses are wastes of tokens and user attention in a tool context. The author connects this to documented mental health crises caused by chatbot interactions, noting that 'guard rails' added in response remain trivially bypassable in 2026. On interfaces, the argument is that natural language as a universal input is producing 'superstitions masquerading as best practices' — prompt engineering folklore that papers over imprecise communication. Rather than compensating with tool-use via MCP servers (giving imprecisely-instructed systems destructive capabilities), the author wants task-specific UIs: a security scanner button for OWASP top 10 bugs, not a chat prompt hoping the model interprets correctly. The data provenance point, though the article cuts off, is clear in direction: outputs from deterministic API calls, RAG retrievals, and pure generation are presented identically, despite having radically different reliability profiles. A serious product would visually distinguish these sources so users can calibrate trust per-output. What makes this piece notable is not that it criticizes AI — that is abundant — but that it offers specific, buildable product specifications. Checkboxes per claim, citation-first layouts, provenance indicators, task-specific UIs. These are not research problems. They are product decisions that frontier labs have chosen not to make, because the current design serves engagement metrics better than it serves users.