LMDeploy

LMDeploy

LMDeploy is an efficient toolkit for compressing, deploying, and serving LLMs.

FreeLinuxAPICLIOpen sourceKoreanLLM-basedMultimodal
Visit websitegithub.com
Compare with book-to-skillExplore LMDeploy alternatives

Overview

LMDeploy is an efficient toolkit for compressing, deploying, and serving LLMs. It features the TurboMind engine, supporting AWQ quantization, continuous batching, and optimized KV caching to maximize throughput and minimize latency for production-grade AI services.

Key features

  • TurboMind Inference Engine
  • AWQ 4-bit Quantization
  • Continuous Batching
  • KV Cache Management
  • Multi-GPU Distributed Inference
  • OpenAI-compatible API
  • PyTorch Engine Support

Pricing

FreeStarting price: Free (open source)
View pricing page

Verified on:

Use cases

  • Building high-performance LLM API servers
  • Model compression and quantization
  • Deploying real-time chatbot services

Who it is for

AI EngineersMLOps SpecialistsBackend Developers

Integrations

DockerKubernetesGradioFastAPI

Tags

MLOps

How we verified this

Company, pricing, and feature details come from the primary sources below and our latest verification pass. When sources disagree, the official source and the most recent check win.

Last verified 08/30/2026Verified sources: 1

Alternatives

Tools you can use instead