You know how GPS works in your car? It doesn't compute your position from scratch every second. It gets a rough fix from satellites, then refines it with dead reckoning between fixes. Coarse global method for the big picture, cheap local method for the fine detail. That's exactly what the 8087 does with tangent — CORDIC gets you 16 bits of accuracy by rotating a vector through precomputed special angles, then a rational polynomial cleans up the last 48 bits on the tiny residual angle. Two algorithms, one pipeline, and the handoff point is the design decision that matters. The committed claim here is not that CORDIC works or that Padé approximants work — both were well-established by 1980. The claim is that Intel's engineers found the optimal splice point between two algorithmic families, implemented it in 1648 micro-instructions on dedicated silicon, and delivered 64-bit accuracy tangent in 90 microseconds versus 13,000 microseconds in software on the 8086. That's a 144× speedup. The article, by Ken Shirriff, recovers this algorithm by physically decapping the chip and reading the microcode ROM under a microscope. The CORDIC half is a pseudo-division loop: the input angle is decomposed into a sum of special angles αₙ = arctan(2⁻ⁿ), each requiring only shifts and adds — no multiplier needed. Sixteen iterations yield 16 bits of accuracy and a tiny residual angle on the order of 2⁻¹⁶ radians. The decision bits (which angles were subtracted) are stored in a 16-bit shift register for later use in pseudo-multiplication, where the actual vector rotations are applied. The second half uses the Padé approximant 3x/(3−x²) to compute the tangent of that residual. Because the residual is so small, the error term — proportional to x⁴ — is less than 2⁻⁶⁴, which meets the 8087's 64-bit accuracy target. A Taylor series of comparable order would be less accurate for the same computation cost. And because FPTAN returns numerator and denominator separately, the division implicit in the rational approximation is free — the caller gets two values and can divide whenever convenient. The hardware architecture is a dedicated floating-point datapath operating on 80-bit values: an adder that doubles as the engine for multiplication, division, and square root via looping; a 64-bit barrel shifter; a constant ROM holding CORDIC angle tables; an exponent ROM; and eight stack registers plus temporaries. The microcode orchestrates the three-phase pipeline — pseudo-division, rational approximation, pseudo-multiplication — entirely in firmware, not in hardwired logic. This is the architectural family of microcode-controlled arithmetic coprocessors, a lineage running from the 8087 through the 80287 and 80387 and ending when FPUs were absorbed into the main CPU die. The integrity of the analysis is unusually strong for reverse-engineering work. Shirriff's method is physical inspection of the actual die, cross-referenced with the known IEEE 754 behavior of the instruction. The constants in the ROM are directly readable. The algorithm is verified by tracing the microcode execution path and confirming it reproduces correct tangent values to 64-bit precision. This is not simulation — it is forensic reconstruction from primary hardware evidence. The broader lesson is about hybrid algorithm design under hardware constraints. The 8087's engineers didn't choose between CORDIC and polynomial approximation; they recognized that each method's cost-accuracy curve has a different shape and spliced them at the crossover point. CORDIC is cheap per bit for the first ~16 bits but linear thereafter. The Padé approximant is expensive to set up but converges explosively on small inputs. The 2⁻¹⁶ handoff point is where the marginal cost curves cross. This splice-point thinking shows up everywhere in modern computing — from mixed-precision neural network inference to hybrid classical-quantum algorithms — making this 45-year-old design decision surprisingly instructive.