LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2026-03-02 01:36:02 +08:00

Author	SHA1	Message	Date
hoshi-hiyouga	73198a6645	[misc] fix uv (#7913 )	2025-04-30 07:45:03 +08:00
hoshi-hiyouga	d4ee44bdef	[data] add eval_on_each_dataset arg (#7912 )	2025-04-30 06:56:43 +08:00
hoshi-hiyouga	6d2cde43e7	[data] replace eos token for base models (#7911 )	2025-04-30 06:52:28 +08:00
hoshi-hiyouga	11295cdea0	[data] improve mm plugin (#7910 )	2025-04-30 06:34:28 +08:00
hoshi-hiyouga	98f23c6584	[model] add qwen3 (#7885 )	2025-04-29 09:34:05 +08:00
Kingsley	db9559456c	[data] fix qwen2.5 omni template (#7883 )	2025-04-29 00:58:23 +08:00
hoshi-hiyouga	3ae5da2a04	[model] fix dsv3 leaf node (#7879 )	2025-04-28 18:11:09 +08:00
hoshi-hiyouga	d173cb50f5	[data] fix qwen2 omni plugin (#7875 )	2025-04-28 14:22:41 +08:00
zhaop-l	df27d7e48a	[trainer] make projector trainable in freeze training (#7872 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-28 13:19:37 +08:00
hoshi-hiyouga	bb5b83352b	[data] fix minicpmo vllm infer (#7870 )	2025-04-28 01:59:53 +08:00
Kingsley	1157f4e246	fix attn patch for kimivl (#7867 )	2025-04-27 23:12:28 +08:00
Eric Tang	ef03832cd4	[ray] add storage filesystem to ray config (#7854 )	2025-04-27 22:12:40 +08:00
hoshi-hiyouga	2233b739fa	[model] fix vit gradient checkpointing (#7830 )	2025-04-23 22:48:48 +08:00
hoshi-hiyouga	091d2539e8	Merge commit from fork	2025-04-23 16:38:27 +08:00
hoshi-hiyouga	c1a7f2ebb2	[model] fix moe zero3 (#7826 )	2025-04-23 15:30:49 +08:00
Kingsley	fa0eb91f1f	[data] fix internvl plugin (#7817 )	2025-04-23 00:58:22 +08:00
hoshi-hiyouga	49f9ed0232	[assets] update model readme (#7804 )	2025-04-22 16:43:56 +08:00
Kingsley	2a564c25d1	[model] add arch check for InternVL (#7803 )	2025-04-22 16:38:05 +08:00
Kingsley	7500e761d3	[misc] update internvl constants (#7801 )	2025-04-22 15:53:08 +08:00
hoshi-hiyouga	fddcd43c88	[trainer] support early stop (#7797 )	2025-04-22 01:59:33 +08:00
hoshi-hiyouga	0e4ce039ee	[data] improve mmplugin (#7795 )	2025-04-22 01:25:33 +08:00
hoshi-hiyouga	b07628dea5	[example] add bash usage (#7794 )	2025-04-22 00:25:51 +08:00
Juanxi Tian	12ada72ed4	[trainer] Add Muon Optimizer (#7749 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:38:37 +08:00
hoshi-hiyouga	416853dd25	[parser] support omegaconf (#7793 )	2025-04-21 23:30:30 +08:00
Changrui Chen	bd7bc31c79	[data] Fix wrong position ids with packed attention masks (#7754 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:19:36 +08:00
flashJd	0ac641326b	[misc] fix new tokens adding (#7253 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:19:02 +08:00
ddddng	c5ba9106ec	[model] fix gemma3 export (#7786 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:07:11 +08:00
Sachin Beldona	3b2d3794a5	[misc] fix bug in constant (#7765 ) Co-authored-by: Sachin Beldona <sbeldona@cs.cmu.edu>	2025-04-21 23:06:31 +08:00
hoshi-hiyouga	b605c20768	[assets] update wechat (#7792 )	2025-04-21 21:29:42 +08:00
hoshi-hiyouga	39169986ef	[trainer] fix pt loss (#7748 ) * fix pt loss * robust * fix * test	2025-04-17 03:15:35 +08:00
hoshi-hiyouga	86ebb219d6	[breaking] bump transformers to 4.45.0 & improve ci (#7746 ) * update ci * fix * fix * fix * fix * fix	2025-04-17 02:36:48 +08:00
hoshi-hiyouga	d222f63cb7	[infer] set env for vllm ascend (#7745 )	2025-04-17 01:08:55 +08:00
Kingsley	2e518f255f	[model] support intern-VL 2.5-3 series (#7258 ) * add internvl and rebase * fix for internvl2&3 * remove lines * fix video_inputs & lint * nit * add constants * remove lines * fix * fix error * pass ci * pass ci * skip internvl & nit	2025-04-17 00:31:30 +08:00
ENg-122	8f88a4e6a4	[misc] improve entrypoint (#7345 ) * 纯粹优化下入口代码，因为看到if else太多了 * Update cli.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-16 21:48:23 +08:00
leo-pony	b9263ff5ac	[infer] support vllm-ascend (#7739 )	2025-04-16 20:06:47 +08:00
hoshi-hiyouga	ee2ab093a7	[api] fix chat messages (#7732 )	2025-04-15 16:39:08 +08:00
hoshi-hiyouga	3df021d4d7	[deps] upgrade vllm (#7728 )	2025-04-15 14:57:40 +08:00
Joe Schoonover	e252abf051	[docker] patch docker-rocm (#7725 ) * Update Dockerfile * Fix typo * Fix syntax for /bin/sh conditional * Add build args to docker-compose * Change shell to /bin/bash This is required for "==" syntax in conditional string comparison	2025-04-15 13:36:39 +08:00
hoshi-hiyouga	1134baeedd	[assets] update model readme (#7724 )	2025-04-15 00:41:09 +08:00
Kingsley	2101399c94	[model] Support Kimi_VL thinking/instruct (#7719 ) * add kimi_vl * patch config * check version * Update mm_plugin.py * Update mm_plugin.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-15 00:21:58 +08:00
hoshi-hiyouga	3f91a95250	[misc] fix env vars (#7715 )	2025-04-14 16:04:04 +08:00
hoshi-hiyouga	7c61b35106	[misc] upgrade cli (#7714 )	2025-04-14 15:41:22 +08:00
hoshi-hiyouga	f518bfba5b	[deps] upgrade transformers (#7704 )	2025-04-13 18:11:34 +08:00
Yuxuan Zhang	8162f94db5	[model] add GLM-4-0414 (#7695 ) * Update README_zh.md * update	2025-04-13 17:10:45 +08:00
hoshi-hiyouga	1f0c52b73c	[deps] fix uv conflicts (#7686 ) * fix #7678 * Update setup.py * Update tests.yml * Update publish.yml * Update Makefile	2025-04-11 18:02:24 +08:00
Eric Tang	a8caf09c7f	[data] support for specifying a dataset in cloud storage (#7567 ) * add support for loading datasets from s3/gcs * add comments to readme * run linter and address comments * add option to pass in kwargs to ray init (i.e. runtime env) * address comment * revert mixed up changes	2025-04-10 11:31:35 +08:00
Eric Tang	bb8d79bae2	[ray] allow for specifying ray.init kwargs (i.e. runtime_env) (#7647 ) * ray init kwargs * Update trainer_utils.py * fix ray args --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-10 11:31:05 +08:00
Dain Kim	1c436c9f25	[bugfix] enable_gemma_liger_kernel (#7660 ) - The `enable_liger_kernel` function for the Gemma model series was not executed due to the existing `if` statement in the code. - Changed the line to an `elif` statement so that the `apply_liger_kernel` function is executed properly. resolved: #7628	2025-04-10 11:27:30 +08:00
jilongW	1b0934bccb	[misc] fix cuda warn on intel GPU (#7655 )	2025-04-09 21:37:54 +08:00
hoshi-hiyouga	4eec541857	[data] add coig-p dataset (#7657 )	2025-04-09 21:18:25 +08:00

1 2 3 4 5 ...

2624 Commits