refactor ray integration, support save ckpt

2026-03-07 20:26:00 +08:00 · 2025-01-07 08:54:41 +00:00
parent 1e8e7be0a5
commit d8cac6f546
18 changed files with 215 additions and 161 deletions
--- a/examples/README_zh.md
+++ b/examples/README_zh.md
@@ -95,6 +95,12 @@ FORCE_TORCHRUN=1 NNODES=2 NODE_RANK=1 MASTER_ADDR=192.168.0.1 MASTER_PORT=29500
 FORCE_TORCHRUN=1 llamafactory-cli train examples/train_lora/llama3_lora_sft_ds3.yaml
 ```

+#### 使用 Ray 在 4 张 GPU 上微调
+
+```bash
+USE_RAY=1 llamafactory-cli train examples/train_full/llama3_lora_sft_ray.yaml
+```
+
 ### QLoRA 微调

 #### 基于 4/8 比特 Bitsandbytes/HQQ/EETQ 量化进行指令监督微调（推荐）