GitHub's ModernBERT Classifier Catches Leaked Passwords Every Two Seconds

GitHub's new ModernBERT-based classifier evaluates candidate secrets in under 2ms, potentially doubling the number of leaked credentials blocked before entering repository history.

·
·
·
GitHub's ModernBERT Classifier Catches Leaked Passwords Every Two Seconds
  • GitHub launches a ModernBERT classifier for push protection that evaluates candidate secrets in under 2ms
  • Model uses surrounding code context to catch unstructured secrets like custom database passwords, not just known token patterns
  • Could more than double the number of secrets blocked before entering repository history, up from current 30% rate
  • Built with Microsoft Applied Sciences; more precise and cheaper than LLM-based detection pipelines
  • Private preview now; rolling out this month to Secret Protection customers on Enterprise Cloud and Teams, consumes AI credits
  • Also shipping in Enterprise Server 3.23 and the Copilot CLI /security-review command for individual developers

GitHub says a credential appears in public code roughly every two seconds, while the annual volume of exposed credentials has doubled in each of the past three years. To catch more leaks before they enter repository history, GitHub and Microsoft Applied Sciences fine-tuned a classifier that evaluates suspicious strings alongside their surrounding code. It targets password-like secrets with no standard prefix or format and processes a batch of candidates in less than two milliseconds.

GitHub is integrating the classifier into push protection, its existing system for blocking recognizable credentials before a push reaches the remote repository. The company estimates that adding contextual detection could more than double the share of secrets stopped at this boundary.

When secrets have no signature

GitHub says one in three pull requests now involves an AI agent, up from fewer than one in ten a year ago. That growth has increased the absolute number of credentials committed by mistake. Between Q2 2024 and Q2 2026, screened pushes increased 2.84-fold and pushes containing credentials increased 2.59-fold. The leak rate per push remained broadly stable, but the volume of exposed credentials grew with overall activity.

Credential cleanup often trails exposure by weeks. GitHub reports a mean manual revocation time of about 40 days, with roughly one in five secrets taking more than 90 days to revoke. Push protection stops about 30% of newly detected secrets before they enter repository history. The remaining 70% are discovered after exposure, when teams may need to rotate credentials, review access logs, and assess possible misuse.

Pattern scanners frequently miss unstructured secrets such as internal database passwords, custom service tokens, and credentials embedded in configuration files. These values lack recognizable prefixes such as sk- or ghp_. Detecting them requires examining variable names, file types, connection strings, and nearby code.

Four constraints at the push boundary

GitHub designed the system around four linked constraints because every check runs in the critical path of a push:

  • Precision: False positives interrupt work and make future warnings easier to dismiss.
  • Latency: Added inference time delays each affected push.
  • Throughput: The service must evaluate pushes across GitHub’s global workload.
  • Cost: Token-generating models require too much computation for this volume.

Optimizing any one measure can weaken another. A larger model may improve classification while increasing latency and infrastructure costs. A more aggressive threshold may catch additional secrets while blocking harmless placeholders. GitHub therefore needed a compact model that could make a single contextual judgment quickly and consistently.

A small model for a global hot path

GitHub chose a fine-tuned ModernBERT classifier, an encoder-only transformer that reads the candidate and its context, then returns a classification score. Its architecture avoids token-by-token generation, reducing the inference work required for each decision. GitHub says the resulting model is more precise than the LLM-based pipelines it evaluated and can process candidate batches in less than two milliseconds.

In GitHub’s published demonstration, the classifier blocked password-like values inside a database connection string, a Kubernetes Secret manifest, and a Dockerfile. It allowed the placeholder changeme. Those examples show how nearby syntax can distinguish a likely credential from a sample value that would look similar to a regular expression.

The rollout spans cloud, server, and Copilot

GitHub has placed the push protection integration in private preview and plans to extend it later this month to organizations using GitHub Secret Protection on Enterprise Cloud and GitHub Teams. Push-time classification will consume AI credits. Other deployments have different access and billing terms:

Surface Availability Access or cost
Push protection Private preview, with broader availability planned later this month Requires GitHub Secret Protection and consumes AI credits
Post-push scanning Organizations with AI secret detection will receive the new model automatically Alerts remain included at no additional cost
GitHub Enterprise Server Public preview in version 3.23 Supports AI-detected alerts in air-gapped environments
Copilot CLI and Copilot App Being added to the /security-review command Available to individuals without an organization-level plan

Fewer secrets reach git log

Repositories that use custom tokens, internal service passwords, or configuration files containing both real and placeholder values gain broader protection before a push writes those credentials to remote history. GitHub says push protection blocked a secret at least once per second during the past month. If the contextual classifier performs as projected in production, that rate could rise substantially.

Earlier detection reduces the number of incidents that require credential rotation, repository cleanup, log review, and follow-up alerts. As AI agents increase code and push volume, GitHub’s approach moves more screening into the push path with a model designed for low latency and platform-scale inference.

Trending
  • No trending articles

Comments

avatar

Next Reads