Google DeepMind's TapNet Unifies Seven Video Tracking Models in One Repo

Google DeepMind's TAPNet repository bundles every generation of their point tracker, from TAP-Net to TAPNext++, in one open-source hub.

·
·
·
Google DeepMind's TapNet Unifies Seven Video Tracking Models in One RepoPRO
  • Google DeepMind's tapnet repo trends on GitHub, bundling the full TAP model family.
  • Includes TAP-Net, TAPIR, BootsTAPIR, TAPNext, TAPNext++, RoboTAP, and TRAJAN with Apache 2.0 weights.
  • TAPNext++ reaches 67.0% AJ on DAVIS at 512x512 with 40x longer stable tracking.
  • Checkpoints available in both JAX and PyTorch via Hugging Face.
  • Causal TAPIR runs around 17 fps on a 2018 mobile GPU with the included live webcam demo.
  • Ships TAP-Vid, RoboTAP, and TAPVid-3D benchmarks plus eleven Colab notebooks for offline, online, and clustering workflows.

Google DeepMind consolidates its point-tracking stack in TapNet

Google DeepMind’s TapNet repository has climbed GitHub’s trending list while serving as the official home for the lab’s Tracking Any Point models, benchmarks, checkpoints, and demos. For video and robotics developers, the repository replaces a search across separate papers and implementations with one maintained codebase.

With roughly 2,200 GitHub stars, the project includes JAX and PyTorch code, pretrained checkpoints, eleven Colab notebooks, TAP-Vid evaluation tools, and a live webcam demo. Support varies by model, so deployment choices still depend on the required framework, latency, resolution, and tracking horizon.

A pixel becomes a trajectory

Given a video and one or more query points placed on any frame, a Tracking Any Point model estimates the corresponding 2D location on every other frame. It also predicts whether each point remains visible, allowing a track to disappear behind an object and reappear later.

A TAP system combines pixel-scale localization, deformable-surface tracking, long-range association, and occlusion handling. A query can sit on fabric, masonry, skin, or a robot gripper, with the resulting trajectory describing how that precise surface location moves through the clip.

  • Input: a video plus query points defined by frame and image coordinates.
  • Output: per-frame coordinates and visibility estimates for every query.
  • Modes: offline tracking can inspect the full video, while causal tracking processes frames as they arrive.
  • Applications: robot imitation, video editing, dynamic 3D reconstruction, motion analysis, and generative-video evaluation.

Seven systems, distinct jobs

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads