Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and π video, up to 5x faster than OpenAI CLIP and LLaVA πΌοΈ & ποΈ
-
Updated
Oct 30, 2025 - Python
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and π video, up to 5x faster than OpenAI CLIP and LLaVA πΌοΈ & ποΈ
Reproducible scaling laws for contrastive language-image learning (https://arxiv.org/abs/2212.07143)
Ultralytics fork of Apple MobileCLIP for fast image-text inference, training, evaluation, and an iOS demo.
Using Segment-Anything and CLIP to generate pixel-aligned semantic features.
Model merging, task-vector rebasin, and fine-tuning for vision and LLM models.
[Official] [IROS 2024] A goal-oriented planning to lift VLN performance for Closed-Loop Navigation: Simple, Yet Effective
Clipora is a powerful toolkit for fine-tuning OpenCLIP models using Low Rank Adapters (LoRA).
A performant cross-platform image viewer and editor built with Rust and Slint, with an extensible plugin system.
Text-to-image search with OpenCLIP, Docker, Flask, Faiss, etc. and a basic front-end.
A simple open-sourced SigLIP model finetuned on Genshin Impact's image-text pairs.
Mori_Cloud is a web platform that allows users to store, manage, and share memorable moments through images and text. Powered by artificial intelligence for smart image search, Mori_Cloud delivers a personalized, secure, and modern user experience.
use SAM and OpenCLIP to perform zero-shot object detection using COCO 2017 val split.
CLIP based Zero Shot Instance Segmentation
Multimodal RAG chatbot with voice input for fashion product search π€ποΈπβ‘
A doctor-assistive AI system that interprets medical knowledge and patient images simultaneously. It utilizes a Dual-Encoder architecture to cross-reference textbook theory with visual pathology, generating clinically grounded diagnoses.
Group images by provided labels using OpenAI/CLIP
VALORA AI is a Multimodal Pricing Prediction Model that uses textual and visual data to make precise predictions on product prices
Masked Multi-Component Gated Decomposition Architecture
Local-first semantic image explorer powered by OpenCLIP embeddings and Qdrant vector search, enabling natural-language retrieval across screenshots and photos.
An AI-powered computer vision system that automatically selects the best wedding photos from thousands of images.
Add a description, image, and links to the openclip topic page so that developers can more easily learn about it.
To associate your repository with the openclip topic, visit your repo's landing page and select "manage topics."