Training engine
From raw text to specialized models, with Python configuration and native C++/CUDA execution.
- Pretraining & full fine-tuningTrain from scratch, continue pretraining, or update the full model with SFT.
- LoRA & QLoRAAdapters on BF16 bases or FP8, NVFP4 and BnB/NF4 quantization, including stacked LoRA.
- Native precision recipesBF16, hybrid FP8 and Blackwell NVFP4.
- GRPO, DPO & distillationReinforcement learning with reward environments, preference training, and teacher-to-student distillation.
- Multi-GPU & multi-nodeThreaded data parallelism, ZeRO sharding, and Ray across nodes.
- Memory controlCPU offload for weights, gradients, optimizer state and activations, to train models larger than your cards.