clip

39 projetos partilham este topic do GitHub

clip — X-AnyLabeling ★9.8kclipChinese-CLIP — ★6kVLMEvalKit — ★4.3kmmpretrain — ★3.8kzero_nlp — ★3.8kVLM_survey — ★3.1kclip-retrieval — ★2.8kcambrian — ★2kawesome-openai-vision-api-experiments — ★1.7kVideo-ChatGPT — ★1.5kawesome-vlm-architectures — ★1.3kUForm — ★1.2kvlms-zero-to-hero — ★1.2kStable-Diffusion-NCNN — ★1.1knatural-language-image-search — ★1kCLIP4Clip — ★1knatural-language-youtube-search — ★935Transformer-MM-Explainability — ★911aphantasia — ★790PaddleMIX — ★724Vision-Language-Models-Overview — ★665awesome-foundation-and-multimodal-models — ★637keras_cv_attention_models — ★627tidy — ★580clip.cpp — ★563cliport — ★546PicQuery — ★506Transformers-for-NLP-and-Computer-Vision-3rd-Edition — ★498diffusion-explainer — ★482OCRAutoScore — ★481CLIP_Surgery — ★479EVE — ★374Instruct2Act — ★374ViralCutter — ★351GenSim — ★350ViP-LLaVA — ★338FiT3D — ★329Disco_Diffusion_Local — ★314LodeDB — ★84Chinese-CLIP★ 6kVLMEvalKit★ 4.3kmmpretrain★ 3.8kzero_nlp★ 3.8kVLM_survey★ 3.1kclip-retrieval★ 2.8kcambrian★ 2kawesome-openai-vision-ap…★ 1.7kVideo-ChatGPT★ 1.5kawesome-vlm-architecture…★ 1.3kUForm★ 1.2kvlms-zero-to-hero★ 1.2kStable-Diffusion-NCNN★ 1.1knatural-language-image-s…★ 1kCLIP4Clip★ 1knatural-language-youtube…★ 935Transformer-MM-Explainab…★ 911aphantasia★ 790PaddleMIX★ 724Vision-Language-Models-O…★ 665awesome-foundation-and-m…★ 637keras_cv_attention_model…★ 627tidy★ 580clip.cpp★ 563cliport★ 546PicQuery★ 506Transformers-for-NLP-and…★ 498diffusion-explainer★ 482OCRAutoScore★ 481CLIP_Surgery★ 479EVE★ 374Instruct2Act★ 374ViralCutter★ 351GenSim★ 350ViP-LLaVA★ 338FiT3D★ 329Disco_Diffusion_Local★ 314LodeDB★ 84

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
X-AnyLabeling
Effortless data labeling with AI support from Segment Anything and other awesome models.
★ 9.8k
Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6k
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3k
mmpretrain
OpenMMLab Pre-training Toolbox and Benchmark
★ 3.8k
zero_nlp
中文nlp解决方案(大模型、数据、模型、训练、推理)
★ 3.8k
VLM_survey
Collection of AWESOME vision-language models for vision tasks
★ 3.1k
clip-retrieval
Easily compute clip embeddings and build a clip retrieval system with them
★ 2.8k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
awesome-openai-vision-api-experiments
Must-have resource for anyone who wants to experiment with and build on the OpenAI vision API 🔥
★ 1.7k
Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation…
★ 1.5k
awesome-vlm-architectures
Famous Vision Language Models and Their Architectures
★ 1.3k
UForm
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and…
★ 1.2k
vlms-zero-to-hero
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge…
★ 1.2k
Stable-Diffusion-NCNN
Stable Diffusion in NCNN with c++, supported txt2img and img2img
★ 1.1k
natural-language-image-search
Search photos on Unsplash using natural language
★ 1k
CLIP4Clip
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1k
natural-language-youtube-search
Search inside YouTube videos using natural language
★ 935
Transformer-MM-Explainability
[ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting…
★ 911
aphantasia
CLIP + FFT/DWT/RGB = text to image/video
★ 790
PaddleMIX
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end…
★ 724
Vision-Language-Models-Overview
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository.…
★ 665
awesome-foundation-and-multimodal-models
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples…
★ 637
keras_cv_attention_models
Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fas…
★ 627
tidy
Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art…
★ 580
clip.cpp
CLIP inference in plain C/C++ with no extra dependencies
★ 563
cliport
CLIPort: What and Where Pathways for Robotic Manipulation
★ 546
PicQuery
🔍 Search local images with natural language on Android, powered by OpenAI's CLIP model. / 在 Android…
★ 506
Transformers-for-NLP-and-Computer-Vision-3rd-Edition
Transformers 3rd Edition
★ 498
diffusion-explainer
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
★ 482
OCRAutoScore
OCR自动化阅卷项目
★ 481
CLIP_Surgery
[Pattern Recognition 25] CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
★ 479
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 374
Instruct2Act
Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model
★ 374
ViralCutter
Free tool to create viral videos from YouTube, generating clips optimized for TikTok and Instagram with…
★ 351
GenSim
Generating Robotic Simulation Tasks via Large Language Models
★ 350
ViP-LLaVA
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 338
FiT3D
[ECCV 2024] Improving 2D Feature Representations by 3D-Aware Fine-Tuning
★ 329
Disco_Diffusion_Local
Getting the latest versions of Disco Diffusion to work locally, instead of colab. Including how I run this on…
★ 314
LodeDB
World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and…
★ 84
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.