SCM (Screen Memories) is a desktop search tool for your personal media library that does something genuinely useful: it lets you describe a memory in plain language and find the photo or exact video frame that matches. It runs entirely on your Mac — no cloud uploads, no accounts, no telemetry — using CLIP, SigLIP, Whisper, and Tesseract locally. The ambition is clear: make every image and every second of video in your folders semantically searchable, offline. The five search modes are where the engineering lives. Files mode ranks whole photos and videos by vision-model cosine similarity with filename and phrase boosts. Scenes mode segments videos at shot boundaries and embeds midpoint frames, so you land on a timecode, not just a file. OCR mode does literal text matching against Tesseract extractions. Dialogue mode searches Whisper transcripts with tiered exactness — exact phrase, exact words within a window, or words anywhere in the video. An opt-in LLM mode (Qwen3 1.7B or Llama 3.2 3B via llama.cpp) lets you ask questions over the extracted evidence with cited answers. The video pipeline is the most technically ambitious piece. ffmpeg detects shot boundaries, then segments are sampled at densities you choose — from Eco (60s per point, 4-32 segments) up to Ultra Pro (2.5s per point, up to 2,048 segments per video). Each segment embeds its midpoint frame and caches a poster. Shot plans are fingerprinted by path, size, mtime, and config, so re-imports skip detection. This is real engineering for a solo project. The library management is thoughtful in ways that suggest the developer actually uses this daily. Content-hash deduplication survives renames. Watched folders auto-import with filesystem watchers plus re-sync on launch. Model switches trigger background re-embedding without blocking search. Named embedding snapshots let you roll back. The screenshot classifier is rename-proof across 20+ languages. The email tab reassembles OCR-fractured addresses including comma-for-dot noise and bracketed obfuscation. Privacy architecture is genuine, not marketing. The renderer runs in a sandboxed app:// bundle with contextIsolation, OS sandbox, and CSP pinned to self. Weights download once from Hugging Face and Tesseract CDN, SHA-256 verified, and everything after that is offline. The LLM sidecar binds to loopback only. No network calls after initial model download is a real commitment, not a checkbox. The friction is real too. This is an Electron app requiring Bun as package manager, with model downloads starting at ~435MB for default CLIP and climbing to ~850MB for the high-detail SigLIP model. Whisper adds another 150-300MB. Video indexing at higher density presets will eat CPU time and disk. The four vision model options (CLIP ViT-L/14@336, SigLIP-2-B/16, SigLIP-2-L/16@256, SigLIP-B/16@384) offer a genuine speed-quality tradeoff — SigLIP-2-B runs at 50-100ms per image versus 480ms+ for the detail models — but switching re-embeds your entire library. SCM sits in a small but growing category of local-first AI tools that treat your data as yours. It is not a product with a business model; it is an open-source tool with real depth, built by someone who clearly wanted this for themselves and built it properly.