GitHub repo leaderboard by stars, growth rate and activity.
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
| # | Repo | Language | Stars | 30-day trend | Last updated |
|---|---|---|---|---|---|
| 1 | [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond. | Python | 25,018 | last pushed 2 years ago | |
| 2 | Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model | Python | 3,634 | last pushed 1 year ago | |
| 3 | 🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability. | Python | 3,438 | last pushed 3 years ago | |
| 4 | InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions | Python | 2,927 | last pushed 1 year ago |
All · 11,658