SCM, a local macOS app, searches every photo and video frame by description

SCM is an Electron app for macOS, shared as a Show HN post, that searches every photo and every frame of video in a folder. Its README describes it as local-first: no accounts, no cloud, no uploads, with inference running on the user's Mac. Model weights download once, and everything after that is offline.
Search works in separate modes. In the main mode you describe a memory in plain language and a local vision model matches it against image embeddings. Typing first triggers an instant filename-keyword pass, then the vision model takes over. Results are scored by cosine similarity, with gated phrase and filename boosts, an honesty floor calibrated per model and a near-duplicate diversity filter. Every tile carries a "why it matched" badge (Visual match, Filename match and so on) and a hover tooltip with the per-component score breakdown. Chinese, Japanese and Korean queries are searched as overlapping bigrams, so "Taipei Station" also matches "Taipei" and "Station".
Video is handled by scene. ffmpeg scans each video for shot boundaries and builds a segment plan at a density the user picks in Settings, with each preset showing its measured time and disk cost. Each segment embeds its midpoint frame and keeps a poster image. A hit therefore lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps to that moment. A noise gate returns "no scene match" rather than filling the grid with junk, and each video contributes at most 3 scenes. Whole videos embed three frames (at 20, 50 and 80 percent) and average them; GIFs average middle frames. Shot plans are cached per file, so re-imports skip detection.
Two modes do not use the vision model at all. OCR mode matches the fraction of query tokens literally visible in each image's text, ignores the filename, and works even while the AI engine is warming up or offline; matched words are boxed in amber. Tesseract runs in its own worker. English is always on and 35 more languages can be toggled on (default: Simplified and Traditional Chinese, Japanese, Korean). Each language pack is about 2.4 to 5MB, and the default set is about 17MB. Dialogue mode is exact literal retrieval over Whisper transcripts, with no embeddings and no thresholds. It returns three tiers: Exact line (a contiguous phrase in one utterance), Exact words (all words in one utterance or within a window of 8 seconds or less) and Words spoken (all words in the same video). Opening a result seeks to the line. The default Whisper model is tiny.en (about 150MB); base.en (about 300MB) is the alternative, and switching re-transcribes every video.
There are four switchable vision models run through ONNX Runtime, chosen per library. The default CLIP model downloads about 435MB of weights on first use. Switching models re-embeds the whole library: the switch lands instantly, the rest fills in the background, and search falls back to filename keywords until it finishes. The README does not name the other three models.
An optional chat feature, Ask, is opt-in: nothing downloads or runs until it is enabled in Settings. It uses a llama.cpp sidecar bound to loopback that answers from evidence the app already extracted (dialogue lines, OCR text and filename keyword hits), with clickable numbered citations and streamed output with a live tokens-per-second readout. Prefixes such as /screenshots, /videos and /email narrow the corpus, and empty evidence stops the run before the model starts.
Around the search sit library tools. Users can save any query as a tab, up to 20, and the Screenshots and Email tabs can be toggled. The Email view finds photos whose OCR text contains an address, including noisy forms like "gmail,com" or "allen [at] gmail [dot] com". Screenshot classification is rename-proof and uses four signals in priority order: a manual override, a filename vocabulary (30+ localized OS screenshot names in 20+ languages), a PNG or JPEG metadata probe, and a source-folder hint. Import works by drag and drop, a keyboard shortcut or watched folders; problem files retry up to 3 times. Every file is SHA-256 hashed before copy, so renames do not create duplicates. Named embedding versions snapshot the whole searchable state, capped at 10, with restore.
On privacy, the README says media is copied into an app-managed library under ~/Library/Application Support/scm, with no telemetry, no accounts and no uploads. Only main-process workers download anything, once per thing: vision weights from Hugging Face, OCR language packs from the Tesseract CDN, Whisper weights, and, only on opt-in, the llama.cpp sidecar and GGUF chat models, sha256-verified at download time.
Install is via a Homebrew tap (allenv0/scm) for Apple Silicon on macOS 12 or later; the cask clears the macOS quarantine flag automatically on every install and upgrade. Building from source uses Bun and electron-builder, and the build produces SCM-0.2.4.dmg and SCM-0.2.4.zip. Without an Apple Developer identity in the keychain the DMG is unsigned, and macOS may require right-click then Open on first launch. The repo includes unit tests, Electron smoke tests and benchmark scripts that write JSON reports.
Key facts
- SCM is a macOS Electron app that searches photos and video scenes by natural-language description, OCR text and Whisper-transcribed dialogue, with inference running on the user's Mac.
- Default CLIP weights are about 435MB on first use, after which the app is offline; four vision models are switchable via ONNX Runtime, and a switch re-embeds the whole library in the background.
- Dialogue search is exact literal matching over Whisper transcripts (default tiny.en, about 150MB) in three tiers, and OCR mode needs no vision model.
- Video scenes are segmented with ffmpeg and each video contributes at most 3 scenes to results; opening a hit jumps to the timecode.
- Install is via a Homebrew tap on Apple Silicon with macOS 12 or later; the opt-in Ask chat uses a local llama.cpp sidecar.
Why it matters
Finding a specific photo or moment in a large personal library usually means remembering a filename or folder. SCM tries to replace that with description, visible text and spoken words, and does it without sending media anywhere. The README makes the privacy case central: no accounts, no cloud, no uploads, no telemetry. It also goes down to the scene level in video, so a result points at a timecode rather than just a file. The building blocks are CLIP-style image embeddings, Tesseract OCR and Whisper, so the novelty is in how they are combined and packaged into one desktop app.
Who it affects
People with large local collections of photos, screenshots and videos on a Mac, especially Apple Silicon users, who want search without a cloud service. The OCR language packs (35 optional, with Chinese, Japanese and Korean on by default) and CJK query handling make it relevant to non-English libraries. Developers can also read it as a reference for a fully local multi-modal indexing stack built with Electron, ONNX Runtime and Bun.
How to use it
The easiest route is Homebrew on Apple Silicon with macOS 12 or later: run brew tap allenv0/scm, brew trust allenv0/scm, then brew install --cask allenv0/scm/scm. Upgrades use brew upgrade --cask allenv0/scm/scm. Users who prefer least privilege can trust only the cask. To build from source, run bun install, then bun run dev, or bun run dist for installers (bun run dist:unsigned skips code-sign discovery). Import files with the import shortcut, drag and drop, or watched folders. Pick a video search density and OCR languages in Settings; the Ask chat stays off until enabled there. The source text states no price or license.
How solid is it
The source is the project README, so every capability is the author's own description and the text gives no benchmark numbers for search speed, accuracy or indexing time. Benchmarks are mentioned only as scripts that write JSON reports. The repo does describe a test battery: unit tests for ranking, dialogue exact-match, CJK tokens, the Whisper model ladder and MIME sniffing, plus Electron smoke tests across a dozen-plus scenarios. The build instructions refer to version 0.2.4, which suggests an early-stage project. The source does not name the author, state a license, or give release dates or user counts.
Risks and caveats
Switching vision models re-embeds the whole library, and until it finishes search falls back to filename keywords. Switching the Whisper model re-transcribes every video. First use downloads about 435MB of CLIP weights, plus Whisper and OCR packs, so a fully offline start is not possible. Media is copied into the app-managed library, which means extra disk use, and video density presets carry time and disk costs. The DMG builds unsigned without an Apple Developer identity, and macOS may require right-click then Open on first launch. Windows and Linux support is not mentioned; menu-bar and tray features are macOS-only. The opt-in Ask feature downloads a llama.cpp sidecar and GGUF chat models from GitHub and Hugging Face when enabled.
“Weights download once; everything after that is offline.”
— SCM project README