模型部署与推理
把模型跑起来并跑得更快、更省显存的工具
数据截至 8/1 18:48(构建快照,正在获取最新)
按仍在维护排序:先筛掉近 30 天没有提交的项目,再按热度排。 高 star 但早已停更的项目不会出现在这里。
- 1
A high-throughput and memory-efficient inference and serving engine for LLMs
- 2
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
- 3
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
- 4
Making large AI models cheaper, faster and more accessible
- 5
Cross-platform, customizable ML solutions for live and streaming media.
- 6
SGLang is a high-performance serving framework for large language models and multimodal models.
- 7
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
- 8
ncnn is a high-performance neural network inference framework optimized for the mobile platform
- 9
Machine Learning Engineering Open Book
- 10
🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
- 11
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
- 12
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
- 13
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
- 14
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
- 15
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
- 16
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
- 17
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
- 18
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
- 19
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
- 20
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
- 21
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
- 22
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
- 23
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
- 24
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
- 25
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
- 26
Fast inference engine for Transformer models
- 27
Pre-trained Deep Learning models and demos (high quality and extremely fast)
- 28
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
- 29
Fast ML inference & training for ONNX models in Rust
另有 10 个项目因近期无提交或缺少数据未列入。
模型部署与推理要落到业务里,还差什么?
开源项目给的是能力,不是方案。数据怎么接、权限怎么管、上线后谁维护,这些才是落地的真正成本。