multimodal

70 proyectos comparten este topic de GitHub

multimodal — Janus ★17.8kmultimodalrerun — ★11.1kmlx-audio — ★7.6kPixelRAG — ★7kai-notes — ★6.2kVLM-R1 — ★6kalign-anything — ★4.7kFengshenbang-LM — ★4.1kMOSS-TTS — ★3.9ktorchscale — ★3.1kmaestro — ★2.7kOFA — ★2.6kstability-sdk — ★2.4kmPLUG-DocOwl — ★2.4kShow-o — ★2kQwen-VL-Series-Finetune — ★1.9kthepipe — ★1.5kOvis — ★1.5kAwesome-Multimodal-Research — ★1.4kLLaMA-Mesh — ★1.2kalan-sdk-cordova — ★1.1kAria — ★1.1kXrayGLM — ★1.1kMOVA — ★1.1kVisCPM — ★1.1kONE-PEACE — ★1.1kohmycaptcha — ★813lmms-engine — ★805papermage — ★800Multimodal-AND-Large-Language-Models — ★760Pluralistic-Inpainting — ★690MMRec — ★683Awesome-Reasoning-Foundation-Models — ★656SEED — ★642Seg-Zero — ★635tokenize-anything — ★601verl-omni — ★599blended-diffusion — ★589AI-Employe — ★585EVF-SAM — ★505Awesome-Multimodal-Modeling — ★500rerun★ 11.1kmlx-audio★ 7.6kPixelRAG★ 7kai-notes★ 6.2kVLM-R1★ 6kalign-anything★ 4.7kFengshenbang-LM★ 4.1kMOSS-TTS★ 3.9ktorchscale★ 3.1kmaestro★ 2.7kOFA★ 2.6kstability-sdk★ 2.4kmPLUG-DocOwl★ 2.4kShow-o★ 2kQwen-VL-Series-Finetune★ 1.9kthepipe★ 1.5kOvis★ 1.5kAwesome-Multimodal-Resea…★ 1.4kLLaMA-Mesh★ 1.2kalan-sdk-cordova★ 1.1kAria★ 1.1kXrayGLM★ 1.1kMOVA★ 1.1kVisCPM★ 1.1kONE-PEACE★ 1.1kohmycaptcha★ 813lmms-engine★ 805papermage★ 800Multimodal-AND-Large-Lan…★ 760Pluralistic-Inpainting★ 690MMRec★ 683Awesome-Reasoning-Founda…★ 656SEED★ 642Seg-Zero★ 635tokenize-anything★ 601verl-omni★ 599blended-diffusion★ 589AI-Employe★ 585EVF-SAM★ 505Awesome-Multimodal-Model…★ 500

Las líneas conectan a los miembros que están mediblemente relacionados entre sí. El tamaño de los puntos refleja las estrellas.

🧬 Miembros
Janus
Janus-Series: Unified Multimodal Understanding and Generation Models
★ 17.8k
rerun
Visualize, query, and stream to train on multimodal robotics data.
★ 11.1k
mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX…
★ 7.6k
PixelRAG
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
★ 7k
ai-notes
notes for software engineers getting up to speed on new AI developments. Serves as datastore for…
★ 6.2k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
align-anything
Align Anything: Training All-modality Model with Feedback
★ 4.7k
Fengshenbang-LM
★ 4.1k
MOSS-TTS
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS…
★ 3.9k
torchscale
Foundation Architecture for (M)LLMs
★ 3.1k
maestro
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
★ 2.7k
OFA
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a…
★ 2.6k
stability-sdk
SDK for interacting with stability.ai APIs (e.g. stable diffusion inference)
★ 2.4k
mPLUG-DocOwl
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4k
Show-o
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding…
★ 2k
Qwen-VL-Series-Finetune
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9k
thepipe
Get clean data from tricky documents, powered by vision-language models ⚡
★ 1.5k
Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and…
★ 1.5k
Awesome-Multimodal-Research
A curated list of Multimodal Related Research.
★ 1.4k
LLaMA-Mesh
Unifying 3D Mesh Generation with Language Models
★ 1.2k
alan-sdk-cordova
The Self-Coding System for Your App — Alan AI SDK for Cordova
★ 1.1k
Aria
Codebase for Aria - an Open Multimodal Native MoE
★ 1.1k
XrayGLM
🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model…
★ 1.1k
MOVA
MOVA: Towards Scalable and Synchronized Video–Audio Generation
★ 1.1k
VisCPM
[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) |…
★ 1.1k
ONE-PEACE
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One…
★ 1.1k
ohmycaptcha
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible…
★ 813
lmms-engine
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
★ 805
papermage
library supporting NLP and CV research on scientific papers
★ 800
Multimodal-AND-Large-Language-Models
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv…
★ 760
Pluralistic-Inpainting
[CVPR 2019]: Pluralistic Image Completion
★ 690
MMRec
A Toolbox for MultiModal Recommendation. Integrating 10+ Models...
★ 683
Awesome-Reasoning-Foundation-Models
✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models
★ 656
SEED
Official implementation of SEED-LLaMA (ICLR 2024).
★ 642
Seg-Zero
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
★ 635
tokenize-anything
[ECCV 2024] Tokenize Anything via Prompting
★ 601
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 599
blended-diffusion
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
★ 589
AI-Employe
Create browser automation as if you were teaching a human using GPT-4 Vision.
★ 585
EVF-SAM
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
★ 505
Awesome-Multimodal-Modeling
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 500
visualwebarena
VisualWebArena is a benchmark for multimodal agents.
★ 482
Spatial-MLLM
[NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based…
★ 477
OMML
Multi-Modal learning toolkit based on PaddlePaddle and PyTorch, supporting multiple applications such as…
★ 477
MultimodalRecSys
A curated list of awesome resources about multimodal recommender systems.
★ 470
ChatTS
[VLDB' 25] ChatTS: LLM for Time Series Understanding and Reasoning
★ 464
GLM-skills
Official skills for the GLM family of models.
★ 452
tsflex
Flexible time series feature extraction & processing
★ 443
DALLE-mtf
Open-AI's DALL-E for large scale training in mesh-tensorflow.
★ 431
Med-PaLM
Towards Generalist Biomedical AI
★ 431
VirConv
Virtual Sparse Convolution for Multimodal 3D Object Detection
★ 384
llark
Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner,…
★ 382
LLaVA-Interactive-Demo
LLaVA-Interactive-Demo
★ 380
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models,…
★ 379
gazelle
Joint speech-language model - respond directly to audio!
★ 374
Multimodal-Sentiment-Analysis
多模态情感分析——基于BERT+ResNet的多种融合方法
★ 369
VisionReasoner
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
★ 348
Awesome-Multimodal-Papers
A curated list of awesome Multimodal studies.
★ 341
cc2dataset
Easily convert common crawl to a dataset of caption and document. Image/text Audio/text Video/text, ...
★ 321
HPT
HPT - Open Multimodal LLMs from HyperGAI
★ 313
Book-of-MLM
《多模态大模型:新一代人工智能技术范式》配套教学资源
★ 310
Thinking-with-Visual-Primitives
Archived snapshot of Thinking-with-Visual-Primitives
★ 272
DeepEar
DeepEar | 顺风耳 An open-source framework for Deep Research and Financial Signal Tracking. …
★ 261
OpenSearch-VL
🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through…
★ 254
DiffThinker
[ICML 2026] Official repo for "DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models"
★ 185
Multimodal-Recommendation-Library
A Continuously Updated Library for Advanced Models for Multimodal Recommendation
★ 181
ICLR2026-Guide-CN
不想啃 5000+ 全文?我已经替你和 LLM 啃完了 — ICLR 2026 全景中文导读
★ 152
GEMS
GEMS: Agent-Native Multimodal Generation with Memory and Skills
★ 139
FlashVID
[ICLR 2026 Oral] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal…
★ 111
VisionTrim
[ICLR 2026] Official code repository for "⚡️VisionTrim: Unified Vision Token Compression for…
★ 54
🔗 Familias relacionadas

Medido a partir de los temas de GitHub compartidos por ambos proyectos, ponderado por cuán raros son cada uno de los temas.