Surogate Engine

Train and serve on one engine

Native C++/CUDA engines for NVIDIA GPUs. Pretrain, fine-tune and run reinforcement learning, then serve the result over the HTTP APIs your clients already speak.

136,200
training tok/s · 4× RTX 5090 · Qwen3-0.6B FP4 LoRA
2.53×
training throughput vs. Unsloth · 1× H100 · BF16 on both
802
serving tok/s, one user · 1× RTX 5090 · Qwen3.5-0.8B GGUF
7.0×
serving throughput vs. llama.cpp · 8× RTX 5090 · 16 users

Open source under Apache 2.0. Install on Linux x86_64 with Python 3.12, CUDA 13 and a supported NVIDIA GPU:

Installlinux · x86_64
curl -LsSf https://github.com/invergent-ai/surogate/releases/latest/download/install.sh | bash
source .venv/bin/activate

Speed you can measure

Fast for one user. Fast under load.

At 100 users, Qwen3.5-4B returns its first token in 40 ms at the median, against 230 ms for vLLM. Training Qwen3-0.6B with FP4 LoRA scales from 36,400 tok/s on one RTX 5090 to 136,200 on four.

Serving · decode tokens/s · RTX 5090
Training · tokens/s

Two engines, one workflow

Train it, then serve what you trained.

Training engine

From raw text to specialized models, with Python configuration and native C++/CUDA execution.

  • Pretraining & full fine-tuningTrain from scratch, continue pretraining, or update the full model with SFT.
  • LoRA & QLoRAAdapters on BF16 bases or FP8, NVFP4 and BnB/NF4 quantization, including stacked LoRA.
  • Native precision recipesBF16, hybrid FP8 and Blackwell NVFP4.
  • GRPO, DPO & distillationReinforcement learning with reward environments, preference training, and teacher-to-student distillation.
  • Multi-GPU & multi-nodeThreaded data parallelism, ZeRO sharding, and Ray across nodes.
  • Memory controlCPU offload for weights, gradients, optimizer state and activations, to train models larger than your cards.

Serving engine

A native C++/CUDA HTTP server built for quick responses, concurrent workloads, and efficient model placement.

  • Familiar APIsChat Completions, Completions, Responses and Messages endpoints, so existing clients connect unchanged.
  • Concurrent servingContinuous batching, chunked prefill, CUDA graphs, and up to 128 active sequences per model.
  • Speculative decodingMTP, and DFlash with a compatible drafter on a single GPU.
  • Native GGUF & Hugging Face checkpointsK-quants, Q8_0 and IQ formats, or safetensors in BF16, FP8 and NVFP4.
  • Runtime LoRALoad and unload adapters without restarting; pick one per request.
  • Large models on the hardware you haveMulti-GPU layer pipelines, CPU weight offload, and GPU expert caching for MoE models.

Run it

One command to serve, one config to train.

From the quickstart. Training needs an Ada, Hopper or Blackwell GPU; serving targets the RTX 40 and 50 series, L4/L40 and H100/H200.

Serve a modelchat completions
surogate serve Qwen/Qwen3.5-0.8B \
  --served-model-name surogate \
  --port 8080 --max-model-len 4096 \
  --max-num-seqs 16 --kv-capacity auto

curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "surogate", "stream": true,
       "messages": [{"role": "user", "content": "Hello"}]}'
Train, then serve itlora · sft
# train.yaml
model: Qwen/Qwen3-0.6B
output_dir: ./output
recipe: bf16        # fp8-hybrid for FP8; nvfp4 for Blackwell FP4
lora: true
lora_rank: 16
datasets:
  - path: mlabonne/FineTome-100k
    type: auto

surogate sft train.yaml
surogate merge --base-model Qwen/Qwen3-0.6B \
  --checkpoint-dir output --output merged
surogate serve merged --served-model-name my-finetune