Yaowei Zheng
|
40d3691e9e
|
[misc] fix moe models (#9230)
|
2025-10-05 02:41:02 +08:00 |
|
h7878778h
|
09dedf144f
|
[npu] Redirect SDPA to torch_npu.npu_fusion_attention (opt-in, ZeRO-3 safe, no impact off NPU) (#8972)
|
2025-09-30 18:11:31 +08:00 |
|
Yaowei Zheng
|
6ffebe5ff7
|
[data] fix qwen omni plugin (#9204)
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-09-28 01:02:29 +08:00 |
|
xvxuopop
|
0761a4448f
|
[model] add qwen3-vl/qwen3-omni (#9196)
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-09-27 01:21:47 +08:00 |
|
Yaowei Zheng
|
80fe3a172d
|
[model] add dots ocr (#9176)
|
2025-09-21 23:34:19 +08:00 |
|
Yaowei Zheng
|
260b5625c3
|
[assets] update wechat (#9129)
|
2025-09-14 03:05:08 +08:00 |
|
Kingsley
|
610a3f1094
|
[data] Fix qwen_2vl with valuehead (#9078)
|
2025-09-14 02:22:20 +08:00 |
|
Yaowei Zheng
|
db223e3975
|
[misc] update readme (#9071)
|
2025-09-03 17:22:54 +08:00 |
|
Kingsley
|
185f0556d4
|
[model] support Internvl3_5 (#9028)
|
2025-08-28 17:12:00 +08:00 |
|
Kingsley
|
9c433f6b41
|
[model] fix kimivl (#9018)
|
2025-08-25 16:32:23 +08:00 |
|
Haian Huang(深度眸)
|
1664657d80
|
[model] Support Intern-S1-mini (#8976)
|
2025-08-20 23:52:51 +08:00 |
|
Kingsley
|
022a326ca4
|
[misc] update glm4v ligerkernel (#8978)
|
2025-08-20 23:39:56 +08:00 |
|
Yaowei Zheng
|
2c31279316
|
[assets] update wechat (#8962)
|
2025-08-19 02:55:09 +08:00 |
|
Zeju Qiu
|
003a2acb1a
|
[feature] adding orthogononal finetuning (OFT) to llama factory (#8623)
Co-authored-by: Zeju <zqiu@g003.internal.cluster.is.localnet>
Co-authored-by: Zeju <zqiu@login2.is.localnet>
Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>
|
2025-08-18 18:22:47 +08:00 |
|
Kingsley
|
893edb26d0
|
[model] support GLM4.5V (#8876)
|
2025-08-11 21:45:14 +08:00 |
|
Yaowei Zheng
|
b523543994
|
[data] fix template (#8827)
|
2025-08-06 06:58:09 +08:00 |
|
Yaowei Zheng
|
4dfad24902
|
[model] add gpt oss (#8826)
|
2025-08-06 05:56:46 +08:00 |
|
davidlightmysterion
|
c709c0378d
|
[train] fix adjusting logits size after adding special tokens (#8823)
|
2025-08-05 20:35:07 +08:00 |
|
Kingsley
|
52882d01c3
|
[model] support keye-vl-8b (#8776)
|
2025-07-29 21:24:08 +08:00 |
|
Kingsley
|
d6767f355a
|
[model] add glm4moe (#8689)
|
2025-07-25 19:53:45 +08:00 |
|
Yaowei Zheng
|
4b0ec83928
|
[deps] bump transformers to 4.49.0 (#8564)
|
2025-07-07 20:31:50 +08:00 |
|
Vivek Iyer
|
e0dfdb7dbb
|
Revert "[model] add lora dropout to unsloth" - requested feature already exists (#8554)
Co-authored-by: viyer <vivek_iyer2@apple.com>
|
2025-07-05 11:25:31 +08:00 |
|
Vivek Iyer
|
0686206020
|
[model] add lora dropout to unsloth (#8548)
Co-authored-by: viyer <vivek_iyer2@apple.com>
|
2025-07-04 14:56:36 +08:00 |
|
Kingsley
|
e9f70daabe
|
[model] add gemma3n (#8509)
|
2025-07-01 22:37:24 +08:00 |
|
Kingsley
|
d17a672251
|
[model] add GLM-4.1V (#8462)
|
2025-06-30 01:09:41 +08:00 |
|
Yaowei Zheng
|
2c26ce6ac4
|
Merge commit from fork
|
2025-06-26 13:55:42 +08:00 |
|
Yaowei Zheng
|
f276b9a963
|
[model] do not force load processor (#8457)
|
2025-06-25 19:43:00 +08:00 |
|
Yaowei Zheng
|
9af7915f7b
|
[model] add kimi vl 2506 (#8432)
|
2025-06-23 17:56:48 +08:00 |
|
Vivek Iyer
|
7b252b2368
|
[model] unsloth resume from checkpoint bug (#8423)
Co-authored-by: viyer <vivek_iyer2@apple.com>
|
2025-06-23 16:43:54 +08:00 |
|
Yaowei Zheng
|
ca75f1edf3
|
[model] fix vlm utils (#8388)
|
2025-06-17 01:08:49 +08:00 |
|
Yaowei Zheng
|
3a3bae1cfe
|
[data] fix qwen2vl pos ids (#8387)
|
2025-06-17 00:48:54 +08:00 |
|
Yaowei Zheng
|
31874e4f62
|
[version] release v0.9.3 (#8386)
|
2025-06-16 19:21:32 +08:00 |
|
Kingsley
|
31bca4d172
|
[model] support Mistral3.1 small 2503 (#8335)
|
2025-06-09 10:37:42 +08:00 |
|
Yaowei Zheng
|
9acab4949d
|
[model] fix model generate (#8327)
|
2025-06-07 08:47:50 +08:00 |
|
Vivek Iyer
|
32b4574094
|
[model] pushing FFT with unsloth (#8325)
Co-authored-by: viyer <vivek_iyer2@apple.com>
|
2025-06-07 08:20:58 +08:00 |
|
hoshi-hiyouga
|
9ae17cd173
|
[deps] update to transformers 4.52 (#8125)
|
2025-05-21 05:16:18 +08:00 |
|
hoshi-hiyouga
|
45030ff803
|
[model] switch to gptqmodel (#8108)
|
2025-05-19 22:25:40 +08:00 |
|
piamo
|
bc7f00f2c7
|
[model] update rope kwargs for yarn (#8101)
|
2025-05-19 20:07:54 +08:00 |
|
hoshi-hiyouga
|
994ab6424a
|
[misc] update liger kernel patch (#7966)
|
2025-05-06 20:32:16 +02:00 |
|
hoshi-hiyouga
|
3ae5da2a04
|
[model] fix dsv3 leaf node (#7879)
|
2025-04-28 18:11:09 +08:00 |
|
zhaop-l
|
df27d7e48a
|
[trainer] make projector trainable in freeze training (#7872)
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-28 13:19:37 +08:00 |
|
Kingsley
|
1157f4e246
|
fix attn patch for kimivl (#7867)
|
2025-04-27 23:12:28 +08:00 |
|
hoshi-hiyouga
|
2233b739fa
|
[model] fix vit gradient checkpointing (#7830)
|
2025-04-23 22:48:48 +08:00 |
|
hoshi-hiyouga
|
c1a7f2ebb2
|
[model] fix moe zero3 (#7826)
|
2025-04-23 15:30:49 +08:00 |
|
Kingsley
|
fa0eb91f1f
|
[data] fix internvl plugin (#7817)
|
2025-04-23 00:58:22 +08:00 |
|
Kingsley
|
2a564c25d1
|
[model] add arch check for InternVL (#7803)
|
2025-04-22 16:38:05 +08:00 |
|
hoshi-hiyouga
|
0e4ce039ee
|
[data] improve mmplugin (#7795)
|
2025-04-22 01:25:33 +08:00 |
|
hoshi-hiyouga
|
b07628dea5
|
[example] add bash usage (#7794)
|
2025-04-22 00:25:51 +08:00 |
|
flashJd
|
0ac641326b
|
[misc] fix new tokens adding (#7253)
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-21 23:19:02 +08:00 |
|
ddddng
|
c5ba9106ec
|
[model] fix gemma3 export (#7786)
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-21 23:07:11 +08:00 |
|