Là một công cụ AI, InternLM/lmdeploy đã đạt 8.0k sao trên GitHub, ngôn ngữ Python. LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Tóm tắt dựng từ metadata GitHub của chính dự án — chưa có bài review TopGit. Trang sẽ tự động cập nhật khi bài review đầy đủ được xuất bản.
VÌ SAO CHƯA CÓ REVIEW
TopGit viết bài đầy đủ cho repo có nhiều sao nhất và được yêu cầu nhiều nhất. Trang này là snapshot trong thời gian chờ — xem README gốc ở tab READ ME.
[2026/04] PyPI has expanded the storage quota for LMDeploy and wheel uploads have resumed. v0.12.3 is now available on PyPI, so you can install it directly via pip install lmdeploy.
[2026/02] Support Qwen3.5
[2026/02] Support vllm-project/llm-compressor 4bit symmetric/asymmetric quantization. Refer here for detailed guide
2025
[2025/09] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the performmance of vLLM on H800 for openai gpt-oss models!
[2025/06] Comprehensive inference optimization for FP8 MoE Models
[2025/06] DeepSeek PD Disaggregation deployment is now supported through integration with DLSlime and Mooncake. Huge thanks to both teams!
[2025/04] Enhance DeepSeek inference performance by integration deepseek-ai techniques: FlashMLA, DeepGemm, DeepEP, MicroBatch and eplb
[2025/01] Support DeepSeek V3 and R1
2024
[2024/11] Support Mono-InternVL with PyTorch engine
[2024/10] PyTorchEngine supports graph mode on ascend platform, doubling the inference speed
[2024/09] LMDeploy PyTorchEngine adds support for Huawei Ascend. See supported models here
[2024/09] LMDeploy PyTorchEngine achieves 1.3x faster on Llama3-8B inference by introducing CUDA graph
[2024/08] LMDeploy is integrated into modelscope/swift as the default accelerator for VLMs inference
[2024/07] Support Llama3.1 8B, 70B and its TOOLS CALLING
[2024/07] Support InternVL2 full-series models, InternLM-XComposer2.5 and function call of InternLM2.5
[2024/06] PyTorch engine support DeepSeek-V2 and several VLMs, such as CogVLM2, Mini-InternVL, LlaVA-Next
[2024/05] Balance vision model when deploying VLMs with multiple GPUs
[2024/05] Support 4-bits weight-only quantization and inference on VLMs, such as InternVL v1.5, LLaVa, InternLMXComposer2
[2024/04] Support Llama3 and more VLMs, such as InternVL v1.1, v1.2, MiniGemini, InternLMXComposer2.
[2024/04] TurboMind adds online int8/int4 KV cache quantization and inference for all supported devices. Refer here for detailed guide
[2024/04] TurboMind latest upgrade boosts GQA, rocketing the internlm2-20b model inference to 16+ RPS, about 1.8x faster than vLLM.
[2024/04] Support Qwen1.5-MOE and dbrx.
[2024/03] Support DeepSeek-VL offline inference pipeline and serving.
[2024/03] Support VLM offline inference pipeline and serving.
[2024/02] Support Qwen 1.5, Gemma, Mistral, Mixtral, Deepseek-MOE and so on.
[2024/01] OpenAOE seamless integration with LMDeploy Serving Service.
[2024/01] Support for multi-model, multi-machine, multi-card inference services. For usage instructions, please refer to here
[2024/01] Support PyTorch inference engine, developed entirely in Python, helping to lower the barriers for developers and enable rapid experimentation with new features and technologies.
2023
[2023/12] Turbomind supports multimodal input.
[2023/11] Turbomind supports loading hf model directly. Click here for details.
[2023/11] TurboMind major upgrades, including: Paged Attention, faster attention kernels without sequence length limitation, 2x faster KV8 kernels, Split-K decoding (Flash Decoding), and W4A16 inference for sm_75
[2023/09] TurboMind supports Qwen-14B
[2023/09] TurboMind supports InternLM-20B
[2023/09] TurboMind supports all features of Code Llama: code completion, infilling, chat / instruct, and python specialist. Click here for deployment guide
[2023/08] TurboMind supports 4-bit inference, 2.4x faster than FP16, the fastest open-source implementation. Check this guide for detailed info
[2023/08] LMDeploy has launched on the HuggingFace Hub, providing ready-to-use 4-bit models.
[2023/08] LMDeploy supports 4-bit quantization using the AWQ algorithm.
[2023/07] TurboMind supports Llama-2 70B with GQA.
[2023/07] TurboMind supports Llama-2 7B/13B.
[2023/07] TurboMind supports tensor-parallel inference of InternLM.
Introduction
LMDeploy is a toolkit for compressing, deploying, and serving LLM, developed by the MMRazor and MMDeploy teams. It has the following core features:
Efficient Inference: LMDeploy delivers up to 1.8x higher request throughput than vLLM, by introducing key features like persistent batch(a.k.a. continuous batching), blocked KV cache, dynamic split&fuse, tensor parallelism, high-performance CUDA kernels and so on.
Effective Quantization: LMDeploy supports weight-only and k/v quantization, and the 4-bit inference performance is 2.4x higher than FP16. The quantization quality has been confirmed via OpenCompass evaluation.
Effortless Distribution Server: Leveraging the request distribution service, LMDeploy facilitates an easy and efficient deployment of multi-model services across multiple machines and cards.
Excellent Compatibility: LMDeploy supports KV Cache Quant, AWQ and Automatic Prefix Caching to be used simultaneously.
Performance
Supported Models
LLMs
VLMs
Llama (7B - 65B)
Llama2 (7B - 70B)
Llama3 (8B, 70B)
Llama3.1 (8B, 70B)
Llama3.2 (1B, 3B)
InternLM2 (7B - 20B)
InternLM3 (8B)
InternLM2.5 (7B)
Qwen1.5 (0.5B - 110B)
Qwen1.5 - MoE (0.5B - 72B)
Qwen2 (0.5B - 72B)
Qwen2-MoE (57BA14B)
Qwen2.5 (0.5B - 32B)
Qwen3, Qwen3-MoE
Qwen3-Next(80B)
Code Llama (7B - 34B)
ChatGLM2 (6B)
GLM-4 (9B)
GLM-4-0414 (9B, 32B)
CodeGeeX4 (9B)
YI (6B-34B)
Mistral (7B)
DeepSeek-MoE (16B)
DeepSeek-V2 (16B, 236B)
DeepSeek-V2.5 (236B)
DeepSeek-V3 (685B)
DeepSeek-V3.2 (685B)
Mixtral (8x7B, 8x22B)
Gemma (2B - 7B)
Phi-3-mini (3.8B)
Phi-3.5-mini (3.8B)
Phi-3.5-MoE (16x3.8B)
Phi-4-mini (3.8B)
MiniCPM3 (4B)
SDAR (1.7B-30B)
gpt-oss (20B, 120B)
GLM-4.7-Flash (30B)
GLM-5 (754B)
LLaVA(1.5,1.6) (7B-34B)
Qwen2-VL (2B, 7B, 72B)
Qwen2.5-VL (3B, 7B, 72B)
Qwen3-VL (2B - 235B)
Qwen3.5 (0.8B - 397B)
Qwen3-Omni (30B-A3B)
DeepSeek-VL (7B)
DeepSeek-VL2 (3B, 16B, 27B)
InternVL-Chat (v1.1-v1.5)
InternVL2 (1B-76B)
InternVL2.5(MPO) (1B-78B)
InternVL3 (1B-78B)
InternVL3.5 (1B-241BA28B)
Intern-S1 (241B)
Intern-S1-mini (8.3B)
Intern-S1-Pro (1TB)
Intern-S2-Preview (35B-A3B)
ChemVLM (8B-26B)
CogVLM-Chat (17B)
CogVLM2-Chat (19B)
MiniCPM-Llama3-V-2_5
MiniCPM-V-2_6
Phi-3-vision (4.2B)
Phi-3.5-vision (4.2B)
GLM-4V (9B)
GLM-4.1V-Thinking (9B)
Molmo (7B-D,72B)
Gemma3 (1B - 27B)
Llama4 (Scout, Maverick)
LMDeploy has developed two inference engines - TurboMind and PyTorch, each with a different focus. The former strives for ultimate optimization of inference performance, while the latter, developed purely in Python, aims to decrease the barriers for developers.
They differ in the types of supported models and the inference data type. Please refer to this table for each engine's capability and choose the proper one that best fits your actual needs.
Quick Start
Installation
It is recommended installing lmdeploy using pip in a conda environment (python 3.10 - 3.13):
Starting from v0.13.0, the default prebuilt wheels published on PyPI are built against CUDA 12.8, so pip install lmdeploy is sufficient for typical setups including GeForce RTX 50 series.
Offline Batch Inference
import lmdeploy
with lmdeploy.pipeline("internlm/internlm3-8b-instruct") as pipe:
response = pipe(["Hi, pls intro yourself", "Shanghai is"])
print(response)
[!NOTE]
By default, LMDeploy downloads model from HuggingFace. If you would like to use models from ModelScope, please install ModelScope by pip install modelscope and set the environment variable:
export LMDEPLOY_USE_MODELSCOPE=True
If you would like to use models from openMind Hub, please install openMind Hub by pip install openmind_hub and set the environment variable:
export LMDEPLOY_USE_OPENMIND_HUB=True
For more information about inference pipeline, please refer to here.
Tutorials
Please review getting_started section for the basic usage of LMDeploy.
For detailed user guides and advanced guides, please refer to our tutorials:
User Guide
LLM Inference pipeline
VLM Inference pipeline
LLM Serving
VLM Serving
Quantization
Advance Guide
Inference Engine - TurboMind
Inference Engine - PyTorch
Customize chat templates
Add a new model
gemm tuning
Long context inference
Multi-model inference service
Third-party projects
Deploying LLMs offline on the NVIDIA Jetson platform by LMDeploy: LMDeploy-Jetson
Example project for deploying LLMs using LMDeploy and BentoML: BentoLMDeploy
Contributing
We appreciate all contributions to LMDeploy. Please refer to CONTRIBUTING.md for the contributing guideline.
Acknowledgement
FasterTransformer
llm-awq
vLLM
DeepSpeed-MII
Citation
@misc{2023lmdeploy,
title={LMDeploy: A Toolkit for Compressing, Deploying, and Serving LLM},
author={LMDeploy Contributors},
howpublished = {\url{https://github.com/InternLM/lmdeploy}},
year={2023}
}
@article{zhang2025efficient,
title={Efficient Mixed-Precision Large Language Model Inference with TurboMind},
author={Zhang, Li and Jiang, Youhe and He, Guoliang and Chen, Xin and Lv, Han and Yao, Qian and Fu, Fangcheng and Chen, Kai},
journal={arXiv preprint arXiv:2508.15601},
year={2025}
}
License
This project is released under the Apache 2.0 license.
InternLM/lmdeploy có 8.0k sao GitHub — tải lại trang để xem số mới nhất, hoặc xem trực tiếp github.com/InternLM/lmdeploy. TopGit phản chiếu số sao của GitHub nhưng không cam kết đến từng phút.
InternLM/lmdeploy có những chủ đề gì?
GitHub topics của InternLM/lmdeploy: "codellama", "cuda-kernels", "deepspeed", "fastertransformer", "internlm", "llama", "llama2", "llama3", "llm", "llm-inference", "turbomind". TopGit xếp repo vào nhóm AI Tools.
InternLM/lmdeploy còn đang phát triển không?
Commit gần nhất trên InternLM/lmdeploy là 18 ngày trước (theo timestamp GitHub). Repo có 721 fork — một chỉ báo về mức độ quan tâm của cộng đồng.
InternLM/lmdeploy là gì?
InternLM/lmdeploy (InternLM/lmdeploy) là dự án Python trên GitHub. Theo mô tả gốc: LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
InternLM/lmdeploy so với các dự án AI Tools khác thế nào?
InternLM/lmdeploy được TopGit xếp vào nhóm AI Tools, với 8.0k sao GitHub và viết bằng Python. Xem trang chủ đề AI Tools trên TopGit để so sánh với các dự án tương tự theo số sao và mức độ hoạt động.
InternLM/lmdeploy viết bằng ngôn ngữ gì?
InternLM/lmdeploy chủ yếu viết bằng Python. Trường "language" của GitHub dựa trên phần lớn byte ở nhánh mặc định.
Vì sao InternLM/lmdeploy được xếp vào nhóm AI Tools?
TopGit xếp InternLM/lmdeploy vào nhóm AI Tools dựa trên GitHub topics và mô tả của repo (gắn thẻ: "codellama", "cuda-kernels", "deepspeed"). Việc phân loại dựa trên metadata thật của repo, không phải đoán theo cảm tính biên tập.
Đọc đầy đủ README ở tab phía trên.
Chưa chắc lmdeploy có hợp với bạn?
Để ChatGPT, Claude hoặc Perplexity tìm hiểu giúp — bấm bên dưới và xem AI nói gì về lmdeploy.