SigLIP 2

Scores images against text, for zero-shot triage and text search across image collections.

Field
Cross-domain
Method
Foundation model
Data it takes
imagery
License
Apache-2.0

SigLIP 2 embeds images and text into one space, so an image can be scored against a written description with no training. That supports searching a photograph archive by description, and sorting a large collection into rough groups before anyone opens it.

This is the general-purpose counterpart to RemoteCLIP and DOFA-CLIP, which were trained on overhead imagery. For aerial and satellite scenes those two describe the content better.

Scores are relative within a comparison. They rank candidates and do not give a calibrated probability that a description is correct.

Upstream project → · Paper →

Catalog entry last checked 2026-09-08. All models →

Run SigLIP 2 on your data

The lab runs this on its own compute, in a container, with the inputs and parameters recorded alongside the result. Initial scoping conversations are free.

Contact the Lab