Skip to main content

Engine

The Engine class is the main entry point to the SGLang inference engine. It provides a Python API for text generation and embedding tasks.

Architecture

The engine consists of three components:
  1. TokenizerManager: Tokenizes requests and sends them to the scheduler
  2. Scheduler (subprocess): Receives requests, schedules batches, forwards them, and sends output tokens to the detokenizer
  3. DetokenizerManager (subprocess): Detokenizes output tokens and sends results back to the tokenizer manager
  • The HTTP server, Engine, and TokenizerManager all run in the main process
  • Inter-process communication is done through IPC via the ZMQ library

Initialization

Constructor Parameters

str
required
Path to the model on Hugging Face or local filesystem.
Optional[str]
default:"None"
Path to the tokenizer. Defaults to model_path if not specified.
int
default:"1"
Tensor parallelism size. Number of GPUs to use for model parallelism.
bool
default:"False"
Whether to trust remote code when loading the model.
Optional[int]
default:"None"
Maximum context length. Auto-detected from model config if not specified.
Optional[float]
default:"None"
Fraction of GPU memory to use for static allocation (model weights + KV cache).
str
default:"error"
Logging level. Options: “debug”, “info”, “warning”, “error”.
For the complete list of parameters, see ServerArgs.

Methods

generate

Generate text completions synchronously.
Optional[Union[List[str], str]]
Input prompt(s). Can be a single string or list of strings for batching.
Optional[Union[List[Dict], Dict]]
Sampling parameters. See SamplingParams for details.
Optional[Union[List[List[int]], List[int]]]
Token IDs for text. Use either prompt or input_ids, not both.
Optional[MultimodalDataInputFormat]
Image input(s) for multimodal models. Can be:
  • Single image (file path, URL, or base64 string)
  • List of images (one per request)
  • List of lists of images (multiple images per request)
Optional[Union[List[bool], bool]]
default:"False"
Whether to return log probabilities.
bool
default:"False"
Whether to stream the response token by token.
Optional[int]
default:"None"
Data parallel rank to route the request to when using data parallelism.
Returns: Union[Dict, Iterator[Dict]] - Response dictionary or iterator if streaming Response Format:

async_generate

Generate text completions asynchronously.
Parameters are identical to generate(). Returns Union[Dict, AsyncIterator[Dict]] for streaming.

encode

Generate embeddings for input text.
Union[str, List[str], List[Dict], List[List[Dict]]]
required
Text or messages to encode.
Optional[MultimodalDataInputFormat]
Image data for multimodal embedding models.
Optional[int]
Output embedding dimensions (if model supports dimensionality reduction).
Returns: Dict - Embeddings dictionary

async_encode

Asynchronous version of encode().

score

Score the probability of label tokens appearing after (query + item) pairs.
Optional[Union[str, List[int]]]
required
Query text or pre-tokenized token IDs.
Optional[Union[str, List[str], List[List[int]]]]
required
Item text(s) or pre-tokenized token IDs to score.
Optional[List[int]]
List of token IDs to compute probabilities for.
bool
default:"False"
Whether to normalize probabilities using softmax.
bool
default:"False"
If True, prepend items to query. Otherwise append items to query.
Returns: ScoreResult with scores and prompt_tokens fields.

Session Management

open_session

Open a session for multi-turn conversation with shared context.
int
required
Maximum string length capacity for the session.
Optional[str]
Optional session ID. A UUID is generated if not provided.
bool
default:"False"
Use low-overhead path for realtime streaming (append-only mode).
Optional[float]
Auto-close session after this many seconds of inactivity.
Returns: str - The session ID

close_session

Close a session and release its resources.

Weight Management

update_weights_from_disk

Update model weights from disk without restarting the engine.

load_lora_adapter

Load a LoRA adapter without restarting the engine.

unload_lora_adapter

Unload a LoRA adapter.

Profiling and Monitoring

get_server_info

Get server configuration and runtime information.

start_profile / stop_profile

Start and stop performance profiling.

flush_cache

Flush the KV cache.

Shutdown

shutdown

Shutdown the engine and all subprocesses.
You can also use the engine as a context manager:

Usage Examples

Basic Text Generation

Streaming Generation

Batch Generation

Multimodal Generation

See Also