GitHub repo leaderboard by stars, growth rate and activity.
Code for Machine Learning for Trading, 3rd edition — from data sourcing to live execution.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Open Source Data Security Platform for Developers to Monitor and Detect PII, Anonymize Production Data and Sync it across environments.
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
Synthetic data curation for post-training and structured data extraction
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | Code for Machine Learning for Trading, 3rd edition — from data sourcing to live execution. | Jupyter Notebook | 20,862 | last pushed 3 hours ago | |
| 2 | Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more. | Python | 5,057 | last pushed 7 hours ago | |
| 3 | Open Source Data Security Platform for Developers to Monitor and Detect PII, Anonymize Production Data and Sync it across environments. | Go | 4,141 | last pushed 1 year ago | |
| 4 | Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers. | Python | 3,390 | last pushed 3 days ago | |
| 5 | 🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data. | Python | 2,209 | last pushed 5 hours ago | |
| 6 | Synthetic data curation for post-training and structured data extraction | Python | 1,727 | last pushed 1 week ago |
All · 11,658