LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2026-06-18 05:08:54 +08:00

Author	SHA1	Message	Date
LittleYanlin	2d9bd2aa14	[fix] qwen3.5 projector path (#10242 ) Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>	2026-03-04 01:31:09 +08:00
Kingsley	816480012f	[fix] register visual part for Qwen3.5 (#10227 )	2026-02-28 16:39:24 +08:00
Yaowei Zheng	122cd46084	[model] update constants (#10220 )	2026-02-26 21:13:56 +08:00
Shanay Mehta	aab9b400bb	[model] Add DeepSpeed Z3 leaf module for Qwen3-Next (#10194 ) Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-24 19:54:37 +08:00
Shanay Mehta	184304b5b4	[model] add liger kernel support for Qwen3-Next (#10176 ) Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-10 21:47:48 +08:00
Xue Yadong	d3ebd5678d	[model] support GLM-OCR SFT (#10183 )	2026-02-10 21:41:01 +08:00
Shanay Mehta	ea644d04ec	[model] support GLM-4.7-Flash SFT (#10173 ) Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-09 10:40:44 +08:00
Hertz	b53d7037c2	[model] support youtu-vl model (#10152 )	2026-02-02 21:42:43 +08:00
浮梦	bf04ca6af8	[deps] adapt to transformers v5 (#10147 ) Co-authored-by: frozenleaves <frozen@Mac.local> Co-authored-by: hiyouga <hiyouga@buaa.edu.cn>	2026-02-02 12:07:19 +08:00
Jewon Lee	9640f79ae5	[fix] add visual.pos_embed to Qwen3-VL visual model keys (#10139 )	2026-01-27 16:33:01 +08:00
Yaowei Zheng	8abb8fb533	[v1] use async streamer (#9741 )	2026-01-09 16:12:07 +08:00
Vo Van Phuc	958fb523a2	[model] support LiquidAI's LFM2.5-VL vision-language model (#9729 )	2026-01-07 17:20:29 +08:00
Yaowei Zheng	d22de0d4bf	[v1] add renderer ut (#9722 )	2026-01-07 02:06:07 +08:00
Xunpeng Xiao	68119e5522	[misc] Add a PyTorch version warning for Conv3D. (#9715 )	2026-01-05 13:26:29 +08:00
Yaowei Zheng	8600530002	[misc] lint (#9710 )	2026-01-04 13:47:56 +08:00
Xunpeng Xiao	0087bc253b	[misc] Compatible with an empty architectures field in config.json (#9709 )	2026-01-04 12:11:35 +08:00
浮梦	16735b9e35	[v1] Refactor kernel plugin (#9669 ) Co-authored-by: frozenleaves <frozen@Mac.local>	2025-12-31 18:26:48 +08:00
Copilot	eceec8ab69	[deps] goodbye python 3.9 (#9677 ) Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: hiyouga <16256802+hiyouga@users.noreply.github.com> Co-authored-by: hiyouga <hiyouga@buaa.edu.cn>	2025-12-27 02:50:44 +08:00
Yaowei Zheng	55590f5ece	[misc] fix ci with uv (#9676 )	2025-12-27 01:39:13 +08:00
Xunpeng Xiao	3c17f2722c	[model] Update ernie_vl to adapt new version (#9665 )	2025-12-26 19:57:49 +08:00
Yaowei Zheng	a754604c11	[misc] fix accelerator (#9661 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-12-25 02:11:04 +08:00
Xunpeng Xiao	6a2eafbae3	[feat] Models trained and inferred with Mxfp4 are dequantized by default (#9652 ) Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>	2025-12-24 00:26:40 +08:00
Yaowei Zheng	84485406b7	[ci] disable pip cache for ci (#9654 )	2025-12-23 18:37:40 +08:00
thulyubh22	7901b2f32e	[model] efficient tuning for gpt-oss (#9354 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-12-23 16:28:38 +08:00
Hertz	4923f52a28	[model] support MiMo-V2-Flash model (#9637 )	2025-12-21 14:38:18 +08:00
浮梦	5204cd2bca	[misc] add version check for moe (#9633 )	2025-12-19 14:57:37 +08:00
Xunpeng Xiao	8c74dca76a	[feat] Models trained and inferred with FP8 are dequantized by default (#9627 )	2025-12-18 22:54:35 +08:00
tangefly	4fd94141a4	[model] Add Ministral3 (#9582 ) Co-authored-by: kingsley <kingsleydodonow@gmail.com>	2025-12-10 15:57:24 +08:00
DoubleWheat	cff4483392	[config] Fix RoPE scaling patch for resuming from a scaled model (#9588 )	2025-12-09 20:37:37 +08:00
Yaowei Zheng	5d56817e2b	[misc] lint (#9593 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-12-09 18:00:35 +08:00
xvxuopop	109162dc56	[fix] fix the issue when using fsdp2 with gradient checkpointing. (#9541 ) Co-authored-by: jin-yongxu <jinyongxu@h-partners.com>	2025-12-06 16:04:51 +08:00
Kingsley	22be45c78c	[misc] fix omni thinker load (#9552 ) Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>	2025-11-30 09:36:36 +08:00
浮梦	2b6f16f261	[model] temporarily support npu fused options on v0, powered by v1 kernels (#9520 ) Co-authored-by: frozenleaves <frozen@Mac.local>	2025-11-27 02:08:36 +08:00
Edge-Seven	9779b1f361	[misc] fix typos in some files (#9505 ) Co-authored-by: khanhkhanhlele <namkhanh20xx@gmail.com>	2025-11-18 20:36:01 +08:00
浮梦	d4e120423d	[data] fix qwen3omni moe model (#9501 ) Co-authored-by: frozenleaves <frozen@Mac.local>	2025-11-18 13:43:22 +08:00
Pory	10a446e373	[model] ktransformers qwen3 support (#9485 ) Co-authored-by: unknown <xiongchenhui@hisense.ad>	2025-11-13 20:09:44 +08:00
Yaowei Zheng	eaf963f67f	[model] update kt code (#9406 )	2025-11-05 15:27:22 +08:00
魅影	14abb75126	[model] enable using FA in npu (#9397 ) Co-authored-by: frozenleaves <frozen@Mac.local>	2025-11-04 19:32:30 +08:00
한송민	5a9939050e	[model] add deepstack_merger_list to Qwen3-VL vision_model_keys (#9399 )	2025-11-04 19:27:34 +08:00
Peilin Li	934b3084ee	[train] KTransformers SFT as backend engine for LLaMA-Factory (#9400 ) Co-authored-by: jimmy128 <jimmy128@noreply.gitcode.com> Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>	2025-11-04 15:54:12 +08:00
魅影	767b344fb4	[model] remove npu sdpa patch (#9368 ) Co-authored-by: frozenleaves <frozen@Mac.local>	2025-10-30 16:26:35 +08:00
Yaowei Zheng	d9d67ba62d	[misc] fix import error (#9299 )	2025-10-17 17:46:27 +08:00
Yaowei Zheng	a442fa90ad	[misc] fix import error (#9296 )	2025-10-17 10:54:30 +08:00
Ximing Xing	c867e28093	[model] adds semantic initialization support for special tokens (#9267 ) Co-authored-by: ximingxing <ximingxing@tencent.com>	2025-10-14 17:00:48 +08:00
Jiayi Mao	48974783da	[model]: add ernie4_5_moe support for DeepSpeed Zero3 training (#9262 )	2025-10-13 13:13:31 +08:00
Yaowei Zheng	40d3691e9e	[misc] fix moe models (#9230 )	2025-10-05 02:41:02 +08:00
h7878778h	09dedf144f	[npu] Redirect SDPA to torch_npu.npu_fusion_attention (opt-in, ZeRO-3 safe, no impact off NPU) (#8972 )	2025-09-30 18:11:31 +08:00
Yaowei Zheng	6ffebe5ff7	[data] fix qwen omni plugin (#9204 ) Co-authored-by: kingsley <kingsleydodonow@gmail.com>	2025-09-28 01:02:29 +08:00
xvxuopop	0761a4448f	[model] add qwen3-vl/qwen3-omni (#9196 ) Co-authored-by: kingsley <kingsleydodonow@gmail.com>	2025-09-27 01:21:47 +08:00
Yaowei Zheng	80fe3a172d	[model] add dots ocr (#9176 )	2025-09-21 23:34:19 +08:00

1 2 3 4 5 ...

288 Commits