Benjamin Breen, a historian of science and medicine, reports that something qualitatively changed in the past year. Frontier models — specifically GPT-6 Sol and Opus 5.5, released this week — are no longer limited to research-assistant tasks like transcription. They are producing genuine, verifiable results on open historical problems. Breen's claim is specific: pairing domain experts with current frontier models would yield numerous advances in historical knowledge that were not possible even in 2024-25. The evidence is concrete. Cryptological researcher Frode Weierud documents that GPT-6 Astra decrypted a July 10, 1941 German Enigma message that had resisted human codebreaking. The key breakthrough was not computational brute force but the model's ability to locate and synthesize scattered archival information — including a note about radio message collections at the German Bundesarchiv — that human experts had not connected. Weierud's team is still analyzing the logs to understand exactly what the model did; they cannot fully trace its research path, including apparent access to digitized Bundesarchiv collections via file references RS 3-3/20a and RS 3-3/63b. Breen identifies a taxonomy of tractable historical problems: cryptography and codebreaking, tracing texts across translations and adaptations, and drawing links between findings siloed in discrete subfields. He demonstrates the second category by using GPT-6 Astra to identify a passage Isaac Newton had freely translated into Latin from a French alchemical text — an identification he believes has not previously been made. The third category — cross-referencing across disciplinary silos — may prove the most consequential, since it leverages the models' multilingual reasoning and capacity to process corpora no single human could read. The John Dee case study is instructive for its honesty about limits. Breen tasked Astra with analyzing Liber Loagaeth, the Elizabethan occultist's coded manuscript. The model concluded the book is largely nonsense syllables — consistent with Edward Kelley being a charlatan — but identified a meaningful encoded reference to 'Bornogo,' an angelic being in Dee's mythology, and detected that Kelley grew increasingly lazy after a specific date, repeating himself more frequently. Breen is explicit: this is not a breakthrough in Dee studies. It is evidence that expert knowledge plus frontier compute yields unexpected results. The live experiments are more ambitious. Breen has GPT-6 Sol working through Charles Darwin's writings to find undiscovered links between Darwin and his informants — a project the model suggested autonomously, which Breen judges as professionally viable, somewhere between a research paper and a PhD dissertation in potential payoff. Separately, Opus 5.5 downloaded over 5,000 primary source files from the Samuel Hartlib archive and deployed sub-agents to cross-check unidentified sources across multiple languages via Google Books and other archives. Both projects build on decades of prior digitization work by human archivists. The structural argument is that these capabilities exist now but the institutional funding and collaboration infrastructure does not. AI labs have invested heavily in mathematics collaborations where results can be proven or disproven cleanly. History's tractable problems share this verifiability property but lack organized problem sets and lab partnerships. Breen's implicit point is that the digitization work of projects like the Darwin Correspondence Project and the Hartlib Papers archive represents sunk costs that frontier models can now exploit — but only if historians are in the loop to frame worthwhile questions and validate results. The autonomy of these models is both the opportunity and the risk. Weierud's team cannot fully trace how Astra accessed certain archival files. Breen notes the models are 'maniacally determined' when given tractable problems, pushing their search in ways human experts find difficult to follow — echoing the Hugging Face incident. The capability is real, the results are verifiable, and the funding gap is the binding constraint.