Imagine you're a waiter on a unicycle, carrying a tray with all the plates stacked on one side. You have two completely separate problems: not falling over right now, and knowing where Table 7 actually is. Your inner ear handles the first; your eyes handle the second. If you tried to use your eyes for both — balancing AND navigating — you'd crash before you left the kitchen. This paper builds exactly that split for a two-wheeled inverted pendulum (TWIP) robot with a side-mounted arm. The committed claim: a dual-loop hierarchical controller — LQR for fast onboard stabilization, ArUco-marker vision for slow global navigation — enables a TWIP robot to carry asymmetric payloads across uneven terrain without toppling, while maintaining accurate waypoint tracking that pure odometry cannot provide. This is not a novel control algorithm per se; it's a novel integration architecture that solves a real coupling problem other TWIP platforms quietly avoid by never mounting an arm. The architecture is clean and legible. The inner loop runs a Linear Quadratic Regulator (LQR) on pitch and yaw using IMU and encoder data at high frequency, rejecting disturbances like speed bumps and the lateral roll torque introduced by lifting a payload on one side. The outer loop uses an overhead camera tracking ArUco fiducial markers to provide ground-truth position, generating waypoints that the inner loop executes. The key insight is that TWIP odometry is fundamentally unreliable — the constant balancing corrections corrupt wheel encoder data — so external vision replaces it entirely for navigation. This is a classical control + computer vision fusion, not a learning-based approach. On the ladder, the paper demonstrates physical hardware results across four scenarios: flat navigation, speed bump rejection, dynamic seesaw traversal, and payload manipulation. They report successful completion of all tasks without toppling. However, there is no head-to-head comparison against a named alternative controller (pure PID, MPC, or a learned policy) on the same hardware. The paper cites prior TWIP work but does not benchmark against it quantitatively. This is the biggest gap: we know the system works, but we don't know how much better it works than the next best approach on the same task. Integrity is honest but limited. These are real physical experiments on custom-built hardware — not simulation-only — which is a genuine strength. The paper includes 9 figures and 2 tables documenting performance. But there's no statistical reporting (no repeated trials with variance), no ablation removing the vision loop or the LQR loop to prove each component is necessary, and no code or hardware design release mentioned. The validation is real-world but anecdotal: we see that it worked, not how robustly or repeatably. The milestone framing is implicit. The current system demonstrates feasibility on a single custom platform in a controlled indoor environment with overhead markers. The next concrete step is outdoor or GPS-denied environments where the overhead camera assumption breaks. Beyond that, the real unlock is dynamic payload estimation — knowing the arm's payload mass and offset in real time rather than pre-tuning the LQR gains. That's the bridge from lab demo to deployable mobile manipulation. The obvious experiment not run: an ablation study. Remove the ArUco loop and run pure odometry navigation; remove the LQR and run pure PID balance. This would quantify the contribution of each half of the hierarchy. The honest read is (a) — this is a 6-page conference paper from what appears to be a student team, and the hardware build itself consumed most of the effort. The integration is the contribution; the rigorous comparison is likely the journal version.