data

48 projects share this GitHub topic

data — llama_index ★51kdatapandas-ai — ★23.7kairbyte — ★21.7kmage-ai — ★8.8kmachine-learning-roadmap — ★7.9kdata-juicer — ★6.7kDataFlow — ★6.7kmachine-learning-mindmap — ★6.3kEdit-Banana — ★5.4ksuperduper — ★5.3kllm-datasets — ★4.7kdatasets — ★4.6kdocetl — ★3.9kLazyLLM — ★3.9kAnyCrawl — ★3.4kspiceai — ★3kstats — ★3kweld — ★3kdeepnote — ★3kdatasets — ★2.8kBook6_First-Course-in-Data-Science — ★2.7kmito — ★2.6ksketch — ★2.3kawesome-streamlit — ★2.3kTigerBot — ★2.3kPython — ★2.2kquant-mind — ★2.1kfree-ai-resources — ★1.9kdiffgram — ★1.9kCurator — ★1.7klotus — ★1.7kdurable-streams — ★1.6ksynthetic-data-kit — ★1.6kcovid19_scenarios — ★1.4klhotse — ★1.1kmcap — ★1kdata-prep-kit — ★949vectordb — ★875VAD — ★869NeumAI — ★864firecrawl-app-examples — ★782pandas-ai★ 23.7kairbyte★ 21.7kmage-ai★ 8.8kmachine-learning-roadmap★ 7.9kdata-juicer★ 6.7kDataFlow★ 6.7kmachine-learning-mindmap★ 6.3kEdit-Banana★ 5.4ksuperduper★ 5.3kllm-datasets★ 4.7kdatasets★ 4.6kdocetl★ 3.9kLazyLLM★ 3.9kAnyCrawl★ 3.4kspiceai★ 3kstats★ 3kweld★ 3kdeepnote★ 3kdatasets★ 2.8kBook6_First-Course-in-Da…★ 2.7kmito★ 2.6ksketch★ 2.3kawesome-streamlit★ 2.3kTigerBot★ 2.3kPython★ 2.2kquant-mind★ 2.1kfree-ai-resources★ 1.9kdiffgram★ 1.9kCurator★ 1.7klotus★ 1.7kdurable-streams★ 1.6ksynthetic-data-kit★ 1.6kcovid19_scenarios★ 1.4klhotse★ 1.1kmcap★ 1kdata-prep-kit★ 949vectordb★ 875VAD★ 869NeumAI★ 864firecrawl-app-examples★ 782

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
llama_index
LlamaIndex is the leading document agent and OCR platform
★ 51k
pandas-ai
Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational…
★ 23.7k
airbyte
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses,…
★ 21.7k
mage-ai
🧙 Build, run, and manage data pipelines for integrating and transforming data.
★ 8.8k
machine-learning-roadmap
A roadmap connecting many of the most important concepts in machine learning, how to learn them and what…
★ 7.9k
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6.7k
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
★ 6.7k
machine-learning-mindmap
A mindmap summarising Machine Learning concepts, from Data Analysis to Deep Learning.
★ 6.3k
Edit-Banana
Edit Banana: A framework for converting statistical formats into editable.
★ 5.4k
superduper
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5.3k
llm-datasets
Curated list of datasets and tools for post-training.
★ 4.7k
datasets
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...
★ 4.6k
docetl
A system for agentic LLM-powered data processing and ETL
★ 3.9k
LazyLLM
Easiest and laziest way for building multi-agent LLMs applications.
★ 3.9k
AnyCrawl
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured…
★ 3.4k
spiceai
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query,…
★ 3k
stats
A well tested and comprehensive Golang statistics library package with no dependencies.
★ 3k
weld
High-performance runtime for data analytics applications
★ 3k
deepnote
Deepnote is a drop-in replacement for Jupyter with an AI-first design, sleek UI, new blocks, and native data…
★ 3k
datasets
🎁 7,400,000+ Unsplash images made available for research and machine learning
★ 2.8k
Book6_First-Course-in-Data-Science
★ 2.7k
mito
Jupyter extensions that help you write code faster: Context aware AI Chat, Autocomplete, and Spreadsheet
★ 2.6k
sketch
AI code-writing assistant that understands data content
★ 2.3k
awesome-streamlit
The purpose of this project is to share knowledge on how awesome Streamlit is and can be
★ 2.3k
TigerBot
TigerBot: A multi-language multi-task LLM
★ 2.3k
Python
This repository helps you learn Python and Machine Learning from scratch.
★ 2.2k
quant-mind
QuantMind is an intelligent knowledge extraction and retrieval framework for quantitative finance.
★ 2.1k
free-ai-resources
🚀 FREE AI Resources - 🎓 Courses, 👷 Jobs, 📝 Blogs, 🔬 AI Research, and many more - for everyone!
★ 1.9k
diffgram
The AI Datastore for Schemas, BLOBs, and Predictions. Use with your apps or integrate built-in Human…
★ 1.9k
Curator
Scalable data pre processing and curation toolkit for LLMs
★ 1.7k
lotus
Optimized Agentic and LLM Bulk Processing Over Your Data
★ 1.7k
durable-streams
The data primitive for the agent loop.
★ 1.6k
synthetic-data-kit
Tool for generating high quality Synthetic datasets
★ 1.6k
covid19_scenarios
Models of COVID-19 outbreak trajectories and hospital demand
★ 1.4k
lhotse
Tools for handling multimodal data in machine learning projects.
★ 1.1k
mcap
MCAP is a modular, performant, and serialization-agnostic container file format, useful for pub/sub and…
★ 1k
data-prep-kit
Open source project for data preparation for GenAI applications
★ 949
vectordb
Epsilla is a high performance Vector Database Management System
★ 875
VAD
Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our…
★ 869
NeumAI
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large…
★ 864
firecrawl-app-examples
🔥 This repository contains complete application examples, including websites and other projects, developed…
★ 782
swiftide
Fast, streaming indexing, query, and agentic LLM applications in Rust
★ 721
manuscript-core
Manuscript is a revolutionary blockchain data streaming framework. With Manuscript, you can seamlessly…
★ 690
bagofwords
Chat with your data - with memory, rules, and observability built in. Deploy in 2 minutes
★ 446
Data-Science-Hacks
Data Science Hacks consists of tips, tricks to help you become a better data scientist. Data science hacks…
★ 432
xpert
Xpert AI is an AI agents and data analysis platform for enterprises to make business decisions.
★ 421
lionagi
An intelligence orchestra
★ 400
marimo-pair
Drop agents inside running marimo notebook sessions
★ 364
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.