| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 870 |
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
|
LLM Models, Inference & Training | Python | 545d | 36543 | - | +652 | +749 |
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 81 |
LLM inference in C/C++
|
LLM Models, Inference & Training | C++ | 1298d | 129830 | - | +87 | +741 |
| 339 |
🧠 Train a 64M-parameter LLM from scratch in just 2h!
|
LLM Models, Inference & Training | Python | 794d | 62862 | - | +67 | +840 |
| 50 |
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
|
LLM Models, Inference & Training | Python | 2891d | 166790 | - | +42 | +313 |
| 647 |
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
|
LLM Models, Inference & Training | Go | 285d | 43024 | - | +74 | +748 |
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 152 |
A high-throughput and memory-efficient inference and serving engine for LLMs
|
LLM Models, Inference & Training | Python | 1327d | 92912 | - | +71 | +555 |
| 520 |
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
|
LLM Models, Inference & Training | Go | 1053d | 49040 | - | +42 | +432 |
| 223 |
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
|
LLM Models, Inference & Training | Python | 1034d | 76990 | - | +98 | +450 |
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 43 |
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
|
LLM Models, Inference & Training | Go | 1190d | 181894 | - | +58 | +501 |
| 241 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
|
LLM Models, Inference & Training | Python | 1219d | 75169 | - | +37 | +211 |
| 376 |
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
|
LLM Models, Inference & Training | Python | 1160d | 59825 | - | +67 | +483 |
| 735 |
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
|
LLM Models, Inference & Training | Python | 726d | 39913 | - | +18 | +110 |
| 747 |
Kronos: A Foundation Model for the Language of Financial Markets
|
LLM Models, Inference & Training | Python | 455d | 39627 | - | +99 | +308 |
| 835 |
Hundreds of models & providers. One command to find what runs on your hardware.
|
LLM Models, Inference & Training | Rust | 225d | 37287 | - | +63 | +345 |
| 926 |
AirLLM 70B inference with single 4GB GPU
|
LLM Models, Inference & Training | Jupyter Notebook | 1204d | 35179 | - | +44 | +532 |
| 986 |
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
|
LLM Models, Inference & Training | Python | 882d | 33953 | - | +61 | - |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 29 |
An Open Source Machine Learning Framework for Everyone
|
LLM Models, Inference & Training | C++ | 3979d | 200603 | - | +26 | +373 |
| 51 |
Stable Diffusion web UI
|
LLM Models, Inference & Training | Python | 1498d | 165147 | - | 0 | +102 |
| 121 |
Tensors and Dynamic neural networks in Python with strong GPU acceleration
|
LLM Models, Inference & Training | Python | 3699d | 103485 | - | +51 | +330 |
| 391 |
The best ChatGPT that $100 can buy.
|
LLM Models, Inference & Training | Python | 350d | 58318 | - | +22 | +115 |
| 738 |
A generative speech model for daily dialogue.
|
LLM Models, Inference & Training | Python | 854d | 39880 | - | +4 | +19 |
| 793 |
DSPy: The framework for programming—not prompting—language models
|
LLM Models, Inference & Training | Python | 1358d | 38410 | - | +21 | +226 |
| 989 |
LLM Frontend for Power Users.
|
LLM Models, Inference & Training | JavaScript | 1327d | 33900 | - | +41 | +257 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 109 |
Robust Speech Recognition via Large-Scale Weak Supervision
|
LLM Models, Inference & Training | Python | 1473d | 109714 | - | +39 | +273 |
| 258 |
A latent text-to-image diffusion model
|
LLM Models, Inference & Training | Jupyter Notebook | 1510d | 73489 | - | +4 | +34 |
| 323 |
Deep Learning for humans
|
LLM Models, Inference & Training | Python | 4203d | 64343 | - | -2 | +17 |
| 441 |
Port of OpenAI's Whisper model in C/C++
|
LLM Models, Inference & Training | C++ | 1464d | 53999 | - | +22 | +157 |
| 453 |
Focus on prompting and generating
|
LLM Models, Inference & Training | Python | 1146d | 53224 | - | +19 | +88 |
| 548 |
Run frontier AI locally.
|
LLM Models, Inference & Training | Python | 826d | 47673 | - | +10 | +79 |
| 631 |
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
|
LLM Models, Inference & Training | Python | 3625d | 43940 | - | +1 | +56 |
| 641 |
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
|
LLM Models, Inference & Training | 1254d | 43342 | - | +51 | +354 | |
| 783 |
Easily train a good VC model with voice data <= 10 mins!
|
LLM Models, Inference & Training | Python | 1281d | 38615 | - | +21 | +191 |
| 804 |
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
|
LLM Models, Inference & Training | Python | 378d | 38119 | - | +90 | +265 |
| 815 |
Instant voice cloning by MIT and MyShell. Audio foundation model.
|
LLM Models, Inference & Training | Python | 1034d | 37717 | - | +17 | +111 |
| 879 |
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
|
LLM Models, Inference & Training | Python | 2895d | 36362 | - | +5 | +42 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 1029 |
Open Frontier Intelligence
|
LLM Models, Inference & Training | 63d | 8876 | - | +5 | +40 | |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 1240 |
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
|
LLM Models, Inference & Training | Python | 11d | 1941 | - | +69 | - |
| 1278 |
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
|
LLM Models, Inference & Training | Python | 35d | 1689 | - | +4 | +45 |
| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 870 |
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
|
LLM Models, Inference & Training | Python | 545d | 36543 | - | +652 | +749 |
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 747 |
Kronos: A Foundation Model for the Language of Financial Markets
|
LLM Models, Inference & Training | Python | 455d | 39627 | - | +99 | +308 |
| 223 |
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
|
LLM Models, Inference & Training | Python | 1034d | 76990 | - | +98 | +450 |
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 804 |
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
|
LLM Models, Inference & Training | Python | 378d | 38119 | - | +90 | +265 |
| 81 |
LLM inference in C/C++
|
LLM Models, Inference & Training | C++ | 1298d | 129830 | - | +87 | +741 |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 647 |
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
|
LLM Models, Inference & Training | Go | 285d | 43024 | - | +74 | +748 |
| 152 |
A high-throughput and memory-efficient inference and serving engine for LLMs
|
LLM Models, Inference & Training | Python | 1327d | 92912 | - | +71 | +555 |
| 1240 |
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
|
LLM Models, Inference & Training | Python | 11d | 1941 | - | +69 | - |
| 339 |
🧠 Train a 64M-parameter LLM from scratch in just 2h!
|
LLM Models, Inference & Training | Python | 794d | 62862 | - | +67 | +840 |
| 376 |
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
|
LLM Models, Inference & Training | Python | 1160d | 59825 | - | +67 | +483 |
| 835 |
Hundreds of models & providers. One command to find what runs on your hardware.
|
LLM Models, Inference & Training | Rust | 225d | 37287 | - | +63 | +345 |
| 986 |
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
|
LLM Models, Inference & Training | Python | 882d | 33953 | - | +61 | - |
| 43 |
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
|
LLM Models, Inference & Training | Go | 1190d | 181894 | - | +58 | +501 |
| 869 |
SGLang is a high-performance serving framework for large language models and multimodal models.
|
LLM Models, Inference & Training | Python | 995d | 36558 | - | +52 | +287 |
| 121 |
Tensors and Dynamic neural networks in Python with strong GPU acceleration
|
LLM Models, Inference & Training | Python | 3699d | 103485 | - | +51 | +330 |
| 641 |
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
|
LLM Models, Inference & Training | 1254d | 43342 | - | +51 | +354 | |
| 1239 |
|
LLM Models, Inference & Training | Python | 25d | 1942 | - | +45 | - |
| 926 |
AirLLM 70B inference with single 4GB GPU
|
LLM Models, Inference & Training | Jupyter Notebook | 1204d | 35179 | - | +44 | +532 |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 50 |
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
|
LLM Models, Inference & Training | Python | 2891d | 166790 | - | +42 | +313 |
| 520 |
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
|
LLM Models, Inference & Training | Go | 1053d | 49040 | - | +42 | +432 |
| 989 |
LLM Frontend for Power Users.
|
LLM Models, Inference & Training | JavaScript | 1327d | 33900 | - | +41 | +257 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 109 |
Robust Speech Recognition via Large-Scale Weak Supervision
|
LLM Models, Inference & Training | Python | 1473d | 109714 | - | +39 | +273 |
| 241 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
|
LLM Models, Inference & Training | Python | 1219d | 75169 | - | +37 | +211 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 1291 |
Qwen's most powerful open-source image generation model
|
LLM Models, Inference & Training | Python | 15d | 1609 | - | +27 | - |
| 29 |
An Open Source Machine Learning Framework for Everyone
|
LLM Models, Inference & Training | C++ | 3979d | 200603 | - | +26 | +373 |
| 391 |
The best ChatGPT that $100 can buy.
|
LLM Models, Inference & Training | Python | 350d | 58318 | - | +22 | +115 |
| 441 |
Port of OpenAI's Whisper model in C/C++
|
LLM Models, Inference & Training | C++ | 1464d | 53999 | - | +22 | +157 |
| 783 |
Easily train a good VC model with voice data <= 10 mins!
|
LLM Models, Inference & Training | Python | 1281d | 38615 | - | +21 | +191 |
| 793 |
DSPy: The framework for programming—not prompting—language models
|
LLM Models, Inference & Training | Python | 1358d | 38410 | - | +21 | +226 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 453 |
Focus on prompting and generating
|
LLM Models, Inference & Training | Python | 1146d | 53224 | - | +19 | +88 |
| 735 |
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
|
LLM Models, Inference & Training | Python | 726d | 39913 | - | +18 | +110 |
| 328 |
The simplest, fastest repository for training/finetuning medium-sized GPTs.
|
LLM Models, Inference & Training | Python | 1371d | 63437 | - | +17 | +144 |
| 815 |
Instant voice cloning by MIT and MyShell. Audio foundation model.
|
LLM Models, Inference & Training | Python | 1034d | 37717 | - | +17 | +111 |
| 345 |
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
|
LLM Models, Inference & Training | Python | 988d | 62239 | - | +16 | +223 |
| 611 |
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
|
LLM Models, Inference & Training | Rust | 1139d | 44698 | - | +16 | +113 |
| 887 |
A modular graph-based Retrieval-Augmented Generation (RAG) system
|
LLM Models, Inference & Training | Python | 915d | 36149 | - | +16 | +89 |
| 1269 |
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
|
LLM Models, Inference & Training | Python | 44d | 1755 | - | +16 | +144 |
| 848 |
Cross-platform, customizable ML solutions for live and streaming media.
|
LLM Models, Inference & Training | C++ | 2664d | 37114 | - | +14 | +86 |
| 859 |
Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
|
LLM Models, Inference & Training | Python | 1898d | 36939 | - | +13 | +72 |
| 509 |
Learn how to develop, deploy and iterate on production-grade ML applications.
|
LLM Models, Inference & Training | Jupyter Notebook | 2885d | 49643 | - | +11 | +70 |
| 514 |
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
|
LLM Models, Inference & Training | Go | 1290d | 49315 | - | +11 | +108 |
| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 339 |
🧠 Train a 64M-parameter LLM from scratch in just 2h!
|
LLM Models, Inference & Training | Python | 794d | 62862 | - | +67 | +840 |
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 870 |
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
|
LLM Models, Inference & Training | Python | 545d | 36543 | - | +652 | +749 |
| 647 |
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
|
LLM Models, Inference & Training | Go | 285d | 43024 | - | +74 | +748 |
| 81 |
LLM inference in C/C++
|
LLM Models, Inference & Training | C++ | 1298d | 129830 | - | +87 | +741 |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 152 |
A high-throughput and memory-efficient inference and serving engine for LLMs
|
LLM Models, Inference & Training | Python | 1327d | 92912 | - | +71 | +555 |
| 926 |
AirLLM 70B inference with single 4GB GPU
|
LLM Models, Inference & Training | Jupyter Notebook | 1204d | 35179 | - | +44 | +532 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 43 |
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
|
LLM Models, Inference & Training | Go | 1190d | 181894 | - | +58 | +501 |
| 376 |
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
|
LLM Models, Inference & Training | Python | 1160d | 59825 | - | +67 | +483 |
| 223 |
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
|
LLM Models, Inference & Training | Python | 1034d | 76990 | - | +98 | +450 |
| 520 |
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
|
LLM Models, Inference & Training | Go | 1053d | 49040 | - | +42 | +432 |
| 29 |
An Open Source Machine Learning Framework for Everyone
|
LLM Models, Inference & Training | C++ | 3979d | 200603 | - | +26 | +373 |
| 641 |
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
|
LLM Models, Inference & Training | 1254d | 43342 | - | +51 | +354 | |
| 835 |
Hundreds of models & providers. One command to find what runs on your hardware.
|
LLM Models, Inference & Training | Rust | 225d | 37287 | - | +63 | +345 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 121 |
Tensors and Dynamic neural networks in Python with strong GPU acceleration
|
LLM Models, Inference & Training | Python | 3699d | 103485 | - | +51 | +330 |
| 50 |
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
|
LLM Models, Inference & Training | Python | 2891d | 166790 | - | +42 | +313 |
| 747 |
Kronos: A Foundation Model for the Language of Financial Markets
|
LLM Models, Inference & Training | Python | 455d | 39627 | - | +99 | +308 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 869 |
SGLang is a high-performance serving framework for large language models and multimodal models.
|
LLM Models, Inference & Training | Python | 995d | 36558 | - | +52 | +287 |
| 109 |
Robust Speech Recognition via Large-Scale Weak Supervision
|
LLM Models, Inference & Training | Python | 1473d | 109714 | - | +39 | +273 |
| 769 |
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
|
LLM Models, Inference & Training | Python | 447d | 38910 | - | +7 | +271 |
| 804 |
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
|
LLM Models, Inference & Training | Python | 378d | 38119 | - | +90 | +265 |
| 989 |
LLM Frontend for Power Users.
|
LLM Models, Inference & Training | JavaScript | 1327d | 33900 | - | +41 | +257 |
| 793 |
DSPy: The framework for programming—not prompting—language models
|
LLM Models, Inference & Training | Python | 1358d | 38410 | - | +21 | +226 |
| 345 |
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
|
LLM Models, Inference & Training | Python | 988d | 62239 | - | +16 | +223 |
| 241 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
|
LLM Models, Inference & Training | Python | 1219d | 75169 | - | +37 | +211 |
| 783 |
Easily train a good VC model with voice data <= 10 mins!
|
LLM Models, Inference & Training | Python | 1281d | 38615 | - | +21 | +191 |
| 441 |
Port of OpenAI's Whisper model in C/C++
|
LLM Models, Inference & Training | C++ | 1464d | 53999 | - | +22 | +157 |
| 328 |
The simplest, fastest repository for training/finetuning medium-sized GPTs.
|
LLM Models, Inference & Training | Python | 1371d | 63437 | - | +17 | +144 |
| 1269 |
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
|
LLM Models, Inference & Training | Python | 44d | 1755 | - | +16 | +144 |
| 391 |
The best ChatGPT that $100 can buy.
|
LLM Models, Inference & Training | Python | 350d | 58318 | - | +22 | +115 |
| 611 |
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
|
LLM Models, Inference & Training | Rust | 1139d | 44698 | - | +16 | +113 |
| 815 |
Instant voice cloning by MIT and MyShell. Audio foundation model.
|
LLM Models, Inference & Training | Python | 1034d | 37717 | - | +17 | +111 |
| 735 |
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
|
LLM Models, Inference & Training | Python | 726d | 39913 | - | +18 | +110 |
| 514 |
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
|
LLM Models, Inference & Training | Go | 1290d | 49315 | - | +11 | +108 |
| 51 |
Stable Diffusion web UI
|
LLM Models, Inference & Training | Python | 1498d | 165147 | - | 0 | +102 |
| 481 |
We write your reusable computer vision tools. 💜
|
LLM Models, Inference & Training | Python | 1400d | 51071 | - | +4 | +90 |
| 887 |
A modular graph-based Retrieval-Augmented Generation (RAG) system
|
LLM Models, Inference & Training | Python | 915d | 36149 | - | +16 | +89 |
| 453 |
Focus on prompting and generating
|
LLM Models, Inference & Training | Python | 1146d | 53224 | - | +19 | +88 |
| 848 |
Cross-platform, customizable ML solutions for live and streaming media.
|
LLM Models, Inference & Training | C++ | 2664d | 37114 | - | +14 | +86 |
| 293 |
scikit-learn: machine learning in Python
|
LLM Models, Inference & Training | Python | 5886d | 67416 | - | +10 | +85 |
| 434 |
Open-Source Frontier Voice AI
|
LLM Models, Inference & Training | Python | 399d | 54530 | - | +6 | +81 |
| 548 |
Run frontier AI locally.
|
LLM Models, Inference & Training | Python | 826d | 47673 | - | +10 | +79 |
| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 1240 |
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
|
LLM Models, Inference & Training | Python | 11d | 1941 | - | +69 | - |
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 1239 |
|
LLM Models, Inference & Training | Python | 25d | 1942 | - | +45 | - |
| 870 |
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
|
LLM Models, Inference & Training | Python | 545d | 36543 | - | +652 | +749 |
| 1291 |
Qwen's most powerful open-source image generation model
|
LLM Models, Inference & Training | Python | 15d | 1609 | - | +27 | - |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 1269 |
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
|
LLM Models, Inference & Training | Python | 44d | 1755 | - | +16 | +144 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 1253 |
Everything I know about running LLMs locally
|
LLM Models, Inference & Training | Shell | 87d | 1853 | - | +5 | +12 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 747 |
Kronos: A Foundation Model for the Language of Financial Markets
|
LLM Models, Inference & Training | Python | 455d | 39627 | - | +99 | +308 |
| 1278 |
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
|
LLM Models, Inference & Training | Python | 35d | 1689 | - | +4 | +45 |
| 804 |
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
|
LLM Models, Inference & Training | Python | 378d | 38119 | - | +90 | +265 |
| 986 |
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
|
LLM Models, Inference & Training | Python | 882d | 33953 | - | +61 | - |
| 647 |
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
|
LLM Models, Inference & Training | Go | 285d | 43024 | - | +74 | +748 |
| 835 |
Hundreds of models & providers. One command to find what runs on your hardware.
|
LLM Models, Inference & Training | Rust | 225d | 37287 | - | +63 | +345 |
| 869 |
SGLang is a high-performance serving framework for large language models and multimodal models.
|
LLM Models, Inference & Training | Python | 995d | 36558 | - | +52 | +287 |
| 223 |
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
|
LLM Models, Inference & Training | Python | 1034d | 76990 | - | +98 | +450 |
| 926 |
AirLLM 70B inference with single 4GB GPU
|
LLM Models, Inference & Training | Jupyter Notebook | 1204d | 35179 | - | +44 | +532 |
| 1172 |
Run the full 2.78-trillion-parameter Kimi K3 model, DeepSeek V4.1 Flash or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
|
LLM Models, Inference & Training | C | 62d | 2457 | - | +3 | +21 |
| 989 |
LLM Frontend for Power Users.
|
LLM Models, Inference & Training | JavaScript | 1327d | 33900 | - | +41 | +257 |
| 641 |
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
|
LLM Models, Inference & Training | 1254d | 43342 | - | +51 | +354 | |
| 376 |
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
|
LLM Models, Inference & Training | Python | 1160d | 59825 | - | +67 | +483 |
| 339 |
🧠 Train a 64M-parameter LLM from scratch in just 2h!
|
LLM Models, Inference & Training | Python | 794d | 62862 | - | +67 | +840 |
| 1045 |
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
|
LLM Models, Inference & Training | Swift | 73d | 6841 | - | +6 | +46 |
| 520 |
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
|
LLM Models, Inference & Training | Go | 1053d | 49040 | - | +42 | +432 |
| 152 |
A high-throughput and memory-efficient inference and serving engine for LLMs
|
LLM Models, Inference & Training | Python | 1327d | 92912 | - | +71 | +555 |
| 81 |
LLM inference in C/C++
|
LLM Models, Inference & Training | C++ | 1298d | 129830 | - | +87 | +741 |
| 1029 |
Open Frontier Intelligence
|
LLM Models, Inference & Training | 63d | 8876 | - | +5 | +40 | |
| 793 |
DSPy: The framework for programming—not prompting—language models
|
LLM Models, Inference & Training | Python | 1358d | 38410 | - | +21 | +226 |
| 783 |
Easily train a good VC model with voice data <= 10 mins!
|
LLM Models, Inference & Training | Python | 1281d | 38615 | - | +21 | +191 |
| 121 |
Tensors and Dynamic neural networks in Python with strong GPU acceleration
|
LLM Models, Inference & Training | Python | 3699d | 103485 | - | +51 | +330 |
| 241 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
|
LLM Models, Inference & Training | Python | 1219d | 75169 | - | +37 | +211 |
| 735 |
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
|
LLM Models, Inference & Training | Python | 726d | 39913 | - | +18 | +110 |
| 815 |
Instant voice cloning by MIT and MyShell. Audio foundation model.
|
LLM Models, Inference & Training | Python | 1034d | 37717 | - | +17 | +111 |
| 887 |
A modular graph-based Retrieval-Augmented Generation (RAG) system
|
LLM Models, Inference & Training | Python | 915d | 36149 | - | +16 | +89 |
| 441 |
Port of OpenAI's Whisper model in C/C++
|
LLM Models, Inference & Training | C++ | 1464d | 53999 | - | +22 | +157 |
| 391 |
The best ChatGPT that $100 can buy.
|
LLM Models, Inference & Training | Python | 350d | 58318 | - | +22 | +115 |
| 848 |
Cross-platform, customizable ML solutions for live and streaming media.
|
LLM Models, Inference & Training | C++ | 2664d | 37114 | - | +14 | +86 |
| 611 |
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
|
LLM Models, Inference & Training | Rust | 1139d | 44698 | - | +16 | +113 |
| 453 |
Focus on prompting and generating
|
LLM Models, Inference & Training | Python | 1146d | 53224 | - | +19 | +88 |
| 109 |
Robust Speech Recognition via Large-Scale Weak Supervision
|
LLM Models, Inference & Training | Python | 1473d | 109714 | - | +39 | +273 |
| 859 |
Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
|
LLM Models, Inference & Training | Python | 1898d | 36939 | - | +13 | +72 |
| 43 |
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
|
LLM Models, Inference & Training | Go | 1190d | 181894 | - | +58 | +501 |
| 328 |
The simplest, fastest repository for training/finetuning medium-sized GPTs.
|
LLM Models, Inference & Training | Python | 1371d | 63437 | - | +17 | +144 |
| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 1269 |
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
|
LLM Models, Inference & Training | Python | 44d | 1755 | - | +16 | +144 |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 1278 |
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
|
LLM Models, Inference & Training | Python | 35d | 1689 | - | +4 | +45 |
| 870 |
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
|
LLM Models, Inference & Training | Python | 545d | 36543 | - | +652 | +749 |
| 647 |
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
|
LLM Models, Inference & Training | Go | 285d | 43024 | - | +74 | +748 |
| 926 |
AirLLM 70B inference with single 4GB GPU
|
LLM Models, Inference & Training | Jupyter Notebook | 1204d | 35179 | - | +44 | +532 |
| 339 |
🧠 Train a 64M-parameter LLM from scratch in just 2h!
|
LLM Models, Inference & Training | Python | 794d | 62862 | - | +67 | +840 |
| 835 |
Hundreds of models & providers. One command to find what runs on your hardware.
|
LLM Models, Inference & Training | Rust | 225d | 37287 | - | +63 | +345 |
| 520 |
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management.
|
LLM Models, Inference & Training | Go | 1053d | 49040 | - | +42 | +432 |
| 1172 |
Run the full 2.78-trillion-parameter Kimi K3 model, DeepSeek V4.1 Flash or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
|
LLM Models, Inference & Training | C | 62d | 2457 | - | +3 | +21 |
| 641 |
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
|
LLM Models, Inference & Training | 1254d | 43342 | - | +51 | +354 | |
| 376 |
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
|
LLM Models, Inference & Training | Python | 1160d | 59825 | - | +67 | +483 |
| 869 |
SGLang is a high-performance serving framework for large language models and multimodal models.
|
LLM Models, Inference & Training | Python | 995d | 36558 | - | +52 | +287 |
| 747 |
Kronos: A Foundation Model for the Language of Financial Markets
|
LLM Models, Inference & Training | Python | 455d | 39627 | - | +99 | +308 |
| 989 |
LLM Frontend for Power Users.
|
LLM Models, Inference & Training | JavaScript | 1327d | 33900 | - | +41 | +257 |
| 769 |
A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.
|
LLM Models, Inference & Training | Python | 447d | 38910 | - | +7 | +271 |
| 804 |
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
|
LLM Models, Inference & Training | Python | 378d | 38119 | - | +90 | +265 |
| 1045 |
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
|
LLM Models, Inference & Training | Swift | 73d | 6841 | - | +6 | +46 |
| 1253 |
Everything I know about running LLMs locally
|
LLM Models, Inference & Training | Shell | 87d | 1853 | - | +5 | +12 |
| 152 |
A high-throughput and memory-efficient inference and serving engine for LLMs
|
LLM Models, Inference & Training | Python | 1327d | 92912 | - | +71 | +555 |
| 793 |
DSPy: The framework for programming—not prompting—language models
|
LLM Models, Inference & Training | Python | 1358d | 38410 | - | +21 | +226 |
| 223 |
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
|
LLM Models, Inference & Training | Python | 1034d | 76990 | - | +98 | +450 |
| 81 |
LLM inference in C/C++
|
LLM Models, Inference & Training | C++ | 1298d | 129830 | - | +87 | +741 |
| 783 |
Easily train a good VC model with voice data <= 10 mins!
|
LLM Models, Inference & Training | Python | 1281d | 38615 | - | +21 | +191 |
| 1029 |
Open Frontier Intelligence
|
LLM Models, Inference & Training | 63d | 8876 | - | +5 | +40 | |
| 345 |
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
|
LLM Models, Inference & Training | Python | 988d | 62239 | - | +16 | +223 |
| 121 |
Tensors and Dynamic neural networks in Python with strong GPU acceleration
|
LLM Models, Inference & Training | Python | 3699d | 103485 | - | +51 | +330 |
| 815 |
Instant voice cloning by MIT and MyShell. Audio foundation model.
|
LLM Models, Inference & Training | Python | 1034d | 37717 | - | +17 | +111 |
| 441 |
Port of OpenAI's Whisper model in C/C++
|
LLM Models, Inference & Training | C++ | 1464d | 53999 | - | +22 | +157 |
| 241 |
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
|
LLM Models, Inference & Training | Python | 1219d | 75169 | - | +37 | +211 |
| 735 |
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
|
LLM Models, Inference & Training | Python | 726d | 39913 | - | +18 | +110 |
| 43 |
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
|
LLM Models, Inference & Training | Go | 1190d | 181894 | - | +58 | +501 |
| 611 |
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
|
LLM Models, Inference & Training | Rust | 1139d | 44698 | - | +16 | +113 |
| 109 |
Robust Speech Recognition via Large-Scale Weak Supervision
|
LLM Models, Inference & Training | Python | 1473d | 109714 | - | +39 | +273 |
| 887 |
A modular graph-based Retrieval-Augmented Generation (RAG) system
|
LLM Models, Inference & Training | Python | 915d | 36149 | - | +16 | +89 |
| 848 |
Cross-platform, customizable ML solutions for live and streaming media.
|
LLM Models, Inference & Training | C++ | 2664d | 37114 | - | +14 | +86 |
| 328 |
The simplest, fastest repository for training/finetuning medium-sized GPTs.
|
LLM Models, Inference & Training | Python | 1371d | 63437 | - | +17 | +144 |
| 514 |
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
|
LLM Models, Inference & Training | Go | 1290d | 49315 | - | +11 | +108 |
| 391 |
The best ChatGPT that $100 can buy.
|
LLM Models, Inference & Training | Python | 350d | 58318 | - | +22 | +115 |
| 859 |
Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
|
LLM Models, Inference & Training | Python | 1898d | 36939 | - | +13 | +72 |
| 949 |
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
|
LLM Models, Inference & Training | Python | 1582d | 34628 | - | +6 | +65 |
| 50 |
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
|
LLM Models, Inference & Training | Python | 2891d | 166790 | - | +42 | +313 |
| # | Repo | Category | Language | Age | Stars | Now | 24h | 7d |
|---|---|---|---|---|---|---|---|---|
| 798 |
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
|
LLM Models, Inference & Training | C | 89d | 38193 | - | +170 | +1373 |
| 1002 |
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
|
LLM Models, Inference & Training | Python | 11d | 27941 | - | +875 | +18337 |
| 1014 |
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
|
LLM Models, Inference & Training | Python | 70d | 13955 | - | +36 | +505 |
| 1018 |
Generates original ARC-AGI-1-style tasks distribution-matched to the public eval set.
|
LLM Models, Inference & Training | Python | 55d | 11113 | - | -6 | -57 |
| 1026 |
|
LLM Models, Inference & Training | Python | 60d | 9350 | - | +41 | +305 |
| 1029 |
Open Frontier Intelligence
|
LLM Models, Inference & Training | 63d | 8876 | - | +5 | +40 | |
| 1030 |
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
|
LLM Models, Inference & Training | C | 58d | 8781 | - | +44 | +647 |
| 1035 |
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
|
LLM Models, Inference & Training | Python | 11d | 7713 | - | +208 | +5416 |
| 1045 |
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
|
LLM Models, Inference & Training | Swift | 73d | 6841 | - | +6 | +46 |
| 1048 |
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
|
LLM Models, Inference & Training | Python | 9d | 6593 | - | +78 | +3018 |
| 1087 |
OpenAI Codex 桌面端/CLI 的可视化管理工具,具有Provider/API 切换、会话同步、提示词注入、Skills/MCP 管理、TOML 配置可视化的跨平台工具。
|
LLM Models, Inference & Training | Rust | 87d | 3988 | - | +21 | +337 |
| 1165 |
DeepSeek Harness 破甲:让所有模型都能破甲,不同模型可换不同提示词;默认提示词面向国模「小码酱」。Jailbreak for every model — swap prompts per model. 求 Star 收藏 ⭐
|
LLM Models, Inference & Training | JavaScript | 40d | 2591 | - | +95 | +825 |
| 1172 |
Run the full 2.78-trillion-parameter Kimi K3 model, DeepSeek V4.1 Flash or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
|
LLM Models, Inference & Training | C | 62d | 2457 | - | +3 | +21 |
| 1239 |
|
LLM Models, Inference & Training | Python | 25d | 1942 | - | +45 | - |
| 1240 |
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
|
LLM Models, Inference & Training | Python | 11d | 1941 | - | +69 | - |
| 1253 |
Everything I know about running LLMs locally
|
LLM Models, Inference & Training | Shell | 87d | 1853 | - | +5 | +12 |
| 1269 |
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
|
LLM Models, Inference & Training | Python | 44d | 1755 | - | +16 | +144 |
| 1278 |
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
|
LLM Models, Inference & Training | Python | 35d | 1689 | - | +4 | +45 |
| 1291 |
Qwen's most powerful open-source image generation model
|
LLM Models, Inference & Training | Python | 15d | 1609 | - | +27 | - |