Gender-Classifier-Mini Hits 97.2% Accuracy and 778K Monthly Downloads

A compact SigLIP2 fine-tune hits 97.2% accuracy on binary gender classification, pulling 778K downloads with a 93M-parameter footprint.

·
·
·
Gender-Classifier-Mini Hits 97.2% Accuracy and 778K Monthly DownloadsPRO
  • Gender-Classifier-Mini hits 97.2% accuracy on binary gender classification with just 92.9M parameters.
  • Fine-tuned from google/siglip2-base-patch16-224 using the SiglipForImageClassification head.
  • Trained on the myvision/gender-classification dataset, roughly balanced across two classes.
  • Apache-2.0 licensed, 778K downloads, runs via standard transformers image-classification pipeline.
  • Precision and recall are nearly symmetric across classes, atypical for binary face models.
  • No demographic fairness breakdown published, limiting safe use to aggregate or research settings.

Gender-Classifier-Mini Draws 778,000 Monthly Downloads

Hugging Face’s model page reports more than 778,000 downloads in the platform’s rolling monthly counter and over 3,000 likes for Gender-Classifier-Mini. The ungated checkpoint fine-tunes Google’s SigLIP 2 vision encoder to assign an image one of two labels, female or male. Its model card reports 97.2% accuracy on a roughly balanced, 5,000-image evaluation split.

Hugging Face counts qualifying file requests, which can include repeated downloads from the same user or automated job. The figure signals substantial usage, though it does not represent 778,000 distinct users or deployments.

The 92.9-million-parameter model is compact enough for local inference on many laptop GPUs and CPUs, with latency determined by hardware and precision. Its weights occupy roughly 372 MB at FP32 before runtime overhead. The Apache 2.0 license permits commercial use, modification, and redistribution subject to its notice requirements, but it does not resolve rights associated with training data or laws governing sensitive-attribute inference.

Inside the 224-pixel encoder

Google’s SigLIP family learns shared image and text representations from paired web data. CLIP trains with a batchwise softmax objective in which matching pairs compete against other examples in the batch. SigLIP scores each image-text pair with an independent sigmoid loss. SigLIP 2 adds captioning, masked prediction, global-local learning, localization features, and support for varied resolutions and aspect ratios.

Gender-Classifier-Mini starts from the base checkpoint, which accepts 224-by-224-pixel inputs and divides them into 16-by-16-pixel patches. That produces a 14-by-14 patch grid before the encoder pools the visual representation. The fine-tuned

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads