Lipflow Lets Mac Users Silently Mouth Words Into Any App

A new open-source Mac app reads your lips through the webcam so you can silently dictate text at your cursor, no mic required.

·
·
·
Lipflow Lets Mac Users Silently Mouth Words Into Any AppPRO
  • Lipflow is a local Mac app that lip-reads your silent mouthing via webcam and types text at your cursor
  • Built on Auto-AVSR trained on LRS3, with MediaPipe landmarks and MLX for Apple Silicon
  • Fine-tunes the lip reader to your face and language model to your phrasing during an 8-minute setup
  • LLM cleanup via Claude, local Qwen3-0.6B through MLX, Ollama, or offline rules disambiguates look-alike phonemes
  • Whisper mode fuses lips plus soft audio, dropping word error rate from 31.9% to 6.9% on test clips
  • MIT-licensed code, but LRS3 weights are non-commercial research use only

Lipflow turns silent mouthing into Mac dictation

Lipflow is an open-source macOS proof of concept that converts silently mouthed words into text at the active cursor. The app captures the user’s lips through a webcam, runs visual speech recognition on Apple Silicon, and inserts the result into the current app. It provides a low-noise dictation option for shared rooms and other places where speaking aloud is impractical.

Recognition and personalization run locally. Text cleanup can also remain offline, or users can send candidate transcriptions and recent dictations to Anthropic’s Claude. The code is MIT licensed, while the downloaded LRS3-trained model weights are restricted to non-commercial research use.

Right Option controls capture: hold the key while mouthing words, then release it to transcribe and paste. Double-tapping the key starts a hands-free session lasting up to 60 seconds. No audio is required in the default mode.

From webcam frames to cursor text

During capture, MediaPipe FaceLandmarker tracks facial landmarks on every frame, removing the need for a second detection pass after recording. Lipflow aligns each face with the model’s mean training face using the eyes, nose base, and mouth as anchors. It then extracts a 96 × 96 grayscale mouth crop and resamples the video to the model’s expected 25 frames per second.

Stage Implementation Purpose
Visual front end 3D-convolutional ResNet Extracts spatial and motion features from consecutive mouth frames.
Sequence encoder Conformer Combines local convolution with attention across the utterance.
Text decoder Transformer with CTC training Maps the visual sequence to text while learning alignment between frames and tokens.
Candidate search Subword language model and beam search Ranks several possible transcriptions for the cleanup stage.

Auto-AVSR, the underlying recognizer, reports a 19.1% word error rate on LRS3, a research corpus of clearly filmed speech. Word error rate counts substitutions, deletions, and insertions relative to a reference transcript, with lower values indicating fewer errors. Silent mouthing usually produces smaller lip movements than the spoken footage used by that benchmark.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads