模型部署与推理

把模型跑起来并跑得更快、更省显存的工具

数据截至 8/1 18:48(构建快照,正在获取最新)

仍在维护排序:先筛掉近 30 天没有提交的项目,再按热度排。 高 star 但早已停更的项目不会出现在这里。

  1. 1
    vllmGitHub模型部署与推理Apache-2.0

    A high-throughput and memory-efficient inference and serving engine for LLMs

    87.8kstar今天有提交Python
  2. 2
    rayGitHub模型部署与推理Apache-2.0

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    43.4kstar今天有提交Python
  3. 3
    DeepSpeedGitHub模型部署与推理Apache-2.0

    DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

    42.8kstar今天有提交Python
  4. 4
    ColossalAIGitHub模型部署与推理Apache-2.0

    Making large AI models cheaper, faster and more accessible

    41.4kstar18 天前提交Python
  5. 5
    mediapipeGitHub模型部署与推理Apache-2.0

    Cross-platform, customizable ML solutions for live and streaming media.

    36.4kstar1 天前提交C++
  6. 6
    sglangGitHub模型部署与推理Apache-2.0

    SGLang is a high-performance serving framework for large language models and multimodal models.

    31.0kstar今天有提交Python
  7. 7
    llm-actionGitHub模型部署与推理Apache-2.0

    本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

    24.8kstar12 天前提交HTML
  8. 8
    ncnnGitHub模型部署与推理

    ncnn is a high-performance neural network inference framework optimized for the mobile platform

    23.6kstar1 天前提交C++
  9. 9
    ml-engineeringGitHub模型部署与推理CC-BY-SA-4.0

    Machine Learning Engineering Open Book

    18.5kstar今天有提交Python
  10. 10
    ts-patternGitHub模型部署与推理MIT

    🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.

    15.1kstar4 天前提交TypeScript
  11. 11
    TensorRTGitHub模型部署与推理Apache-2.0

    NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

    13.2kstar24 天前提交C++
  12. 12
    OpenLLMGitHub模型部署与推理Apache-2.0

    Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

    12.4kstar4 天前提交Python
  13. 13
    amazon-sagemaker-examplesGitHub模型部署与推理Apache-2.0

    Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.

    11.0kstar1 天前提交Jupyter Notebook
  14. 14
    LMCacheGitHub模型部署与推理Apache-2.0

    LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

    11.0kstar今天有提交Python
  15. 15
    serverGitHub模型部署与推理BSD-3-Clause

    The Triton Inference Server provides an optimized cloud and edge inferencing solution.

    10.9kstar今天有提交Python
  16. 16
    yolov3GitHub模型部署与推理AGPL-3.0

    PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.

    10.6kstar今天有提交Python
  17. 17
    skypilotGitHub模型部署与推理Apache-2.0

    The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

    10.4kstar今天有提交Python
  18. 18
    inferenceGitHub模型部署与推理Apache-2.0

    Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

    9.5kstar今天有提交Python
  19. 19
    oumiGitHub模型部署与推理Apache-2.0

    Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

    9.4kstar今天有提交Python
  20. 20
    BentoMLGitHub模型部署与推理Apache-2.0

    The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

    8.7kstar12 天前提交Python
  21. 21
    planoGitHub模型部署与推理Apache-2.0

    Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.

    6.9kstar1 天前提交Rust
  22. 22
    MooncakeGitHub模型部署与推理Apache-2.0

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    6.1kstar今天有提交C++
  23. 23
    whichllmGitHub模型部署与推理MIT

    Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

    6.1kstar3 天前提交Python
  24. 24
    kserveGitHub模型部署与推理Apache-2.0

    Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

    5.8kstar今天有提交Go
  25. 25
    gpustackGitHub模型部署与推理Apache-2.0

    A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

    5.4kstar今天有提交Python
  26. 26
    CTranslate2GitHub模型部署与推理MIT

    Fast inference engine for Transformer models

    4.6kstar28 天前提交C++
  27. 27
    open_model_zooGitHub模型部署与推理Apache-2.0

    Pre-trained Deep Learning models and demos (high quality and extremely fast)

    4.4kstar22 天前提交Python
  28. 28
    csghubGitHub模型部署与推理Apache-2.0

    CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️

    4.2kstar2 天前提交Vue
  29. 29
    ortGitHub模型部署与推理Apache-2.0

    Fast ML inference & training for ONNX models in Rust

    2.4kstar4 天前提交Rust

另有 10 个项目因近期无提交或缺少数据未列入。

模型部署与推理要落到业务里,还差什么?

开源项目给的是能力,不是方案。数据怎么接、权限怎么管、上线后谁维护,这些才是落地的真正成本。