DigUp Brings On-Device Multimodal File Search to Apple Silicon Macs

A free, open-source Mac app uses Google DeepMind's new multimodal embedding model to search inside photos, PDFs, audio, video and code locally.

·
·
·
DigUp Brings On-Device Multimodal File Search to Apple Silicon MacsPRO
  • DigUp is a free Mac menu bar app for semantic search across files, built on EmbeddingGemma 2.
  • Indexes photos, PDFs, documents, audio, video and optionally code into one shared 768-dim vector space.
  • Runs fully on-device via llama.cpp on Metal, with a one-time 865 MB model download.
  • Cross-modal queries work: text can find video moments, images, or PDF pages interchangeably.
  • Handles 100+ languages for semantic search, with OCR limited to ~25 languages via Apple Vision.
  • Requires Apple Silicon and macOS 14+; MIT-licensed and uses EmbeddingGemma 2 under Apache 2.0.

DigUp brings multimodal file search to Apple Silicon Macs

DigUp is a free, MIT-licensed menu bar app that searches files by meaning across text, images, audio, video, PDFs, and source code. Spotlight already searches filenames, metadata, and indexed document text; DigUp adds cross-modal retrieval, allowing a phrase such as “a zebra in a field” to find a video frame or “termination clause” to locate the relevant page in a lease. Search and indexing run on-device after the initial model download.

DigUp logo showing text, image, audio, and video connected through a shared search system

One coordinate system, four media types

EmbeddingGemma 2 converts text, images, and audio into points within the same 768-dimensional vector space. Google describes the 740-million-parameter release as its first natively multimodal embedding model, designed for tasks such as retrieving video with a voice memo or matching an image to a written description.

Each indexed item becomes a list of 768 numbers that represents its meaning. Items with similar content land near one another, so a query such as “a dog on the beach” can rank a photograph, a PDF page, and a sampled video frame together without first generating captions. DigUp decomposes video into visual frames and audio before embedding those components.

DigUp indexes folders selected by the user, stores the resulting vectors in one SQLite file, and maintains a keyword index for literal matches. The project reports millisecond-scale query embedding, followed by comparison against stored vectors for pictures, document passages, PDF pages, video frames, and overlapping audio windows.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads