<< All versions
Skill v1.0.0
currentAutomated scan100/100mkurman/zorai/open-clip
──Details
PublishedJuly 29, 2026 at 12:08 PM
Content Hashsha256:2261a1eed91625ad...
Git SHAd0acbfaf3d62
──Files
Files (1 file, 1.4 KB)
SKILL.md1.4 KBactive
SKILL.md · 40 lines · 1.4 KB
version: "1.0.0" name: open-clip description: "OpenCLIP — open-source implementation of CLIP trained on LAION-5B/OpenCLIP datasets. Multi-head attention pooling, SigLIP loss variants, and wide model zoo (ViT, ConvNeXt, EVA). Community-driven." tags: [open-clip, multimodal, image-text, laion, zero-shot, embeddings, zorai]
Overview
OpenCLIP is an open-source reimplementation of CLIP trained on LAION-5B, LAION-400M, and DataComp. Provides larger and better architectures than the original: ViT-H/14, ConvNeXt, EVA-02, SigLIP. Full model transparency with flexible training customizations.
Installation
bash
uv pip install open-clip-torch
Encoding Images and Text
python
import open_clipimport torchfrom PIL import Imagemodel, _, preprocess = open_clip.create_model_and_transforms("ViT-H-14", pretrained="laion2b_s32b_b79k")tokenizer = open_clip.get_tokenizer("ViT-H-14")image = preprocess(Image.open("photo.jpg")).unsqueeze(0)text = tokenizer(["a dog", "a cat", "a car"])with torch.no_grad():image_features = model.encode_image(image)text_features = model.encode_text(text)logits = (image_features @ text_features.T).softmax(dim=-1)print(f"Predicted: class {logits.argmax().item()} with {logits.max():.2%}")