LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2026-07-07 09:35:27 +08:00

Author	SHA1	Message	Date
Kingsley	fa0eb91f1f	[data] fix internvl plugin (#7817 )	2025-04-23 00:58:22 +08:00
hoshi-hiyouga	49f9ed0232	[assets] update model readme (#7804 )	2025-04-22 16:43:56 +08:00
Kingsley	2a564c25d1	[model] add arch check for InternVL (#7803 )	2025-04-22 16:38:05 +08:00
Kingsley	7500e761d3	[misc] update internvl constants (#7801 )	2025-04-22 15:53:08 +08:00
hoshi-hiyouga	fddcd43c88	[trainer] support early stop (#7797 )	2025-04-22 01:59:33 +08:00
hoshi-hiyouga	0e4ce039ee	[data] improve mmplugin (#7795 )	2025-04-22 01:25:33 +08:00
hoshi-hiyouga	b07628dea5	[example] add bash usage (#7794 )	2025-04-22 00:25:51 +08:00
Juanxi Tian	12ada72ed4	[trainer] Add Muon Optimizer (#7749 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:38:37 +08:00
hoshi-hiyouga	416853dd25	[parser] support omegaconf (#7793 )	2025-04-21 23:30:30 +08:00
Changrui Chen	bd7bc31c79	[data] Fix wrong position ids with packed attention masks (#7754 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:19:36 +08:00
flashJd	0ac641326b	[misc] fix new tokens adding (#7253 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:19:02 +08:00
ddddng	c5ba9106ec	[model] fix gemma3 export (#7786 ) Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-21 23:07:11 +08:00
Sachin Beldona	3b2d3794a5	[misc] fix bug in constant (#7765 ) Co-authored-by: Sachin Beldona <sbeldona@cs.cmu.edu>	2025-04-21 23:06:31 +08:00
hoshi-hiyouga	b605c20768	[assets] update wechat (#7792 )	2025-04-21 21:29:42 +08:00
hoshi-hiyouga	39169986ef	[trainer] fix pt loss (#7748 ) * fix pt loss * robust * fix * test	2025-04-17 03:15:35 +08:00
hoshi-hiyouga	86ebb219d6	[breaking] bump transformers to 4.45.0 & improve ci (#7746 ) * update ci * fix * fix * fix * fix * fix	2025-04-17 02:36:48 +08:00
hoshi-hiyouga	d222f63cb7	[infer] set env for vllm ascend (#7745 )	2025-04-17 01:08:55 +08:00
Kingsley	2e518f255f	[model] support intern-VL 2.5-3 series (#7258 ) * add internvl and rebase * fix for internvl2&3 * remove lines * fix video_inputs & lint * nit * add constants * remove lines * fix * fix error * pass ci * pass ci * skip internvl & nit	2025-04-17 00:31:30 +08:00
ENg-122	8f88a4e6a4	[misc] improve entrypoint (#7345 ) * 纯粹优化下入口代码，因为看到if else太多了 * Update cli.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-16 21:48:23 +08:00
leo-pony	b9263ff5ac	[infer] support vllm-ascend (#7739 )	2025-04-16 20:06:47 +08:00
hoshi-hiyouga	ee2ab093a7	[api] fix chat messages (#7732 )	2025-04-15 16:39:08 +08:00
hoshi-hiyouga	3df021d4d7	[deps] upgrade vllm (#7728 )	2025-04-15 14:57:40 +08:00
Joe Schoonover	e252abf051	[docker] patch docker-rocm (#7725 ) * Update Dockerfile * Fix typo * Fix syntax for /bin/sh conditional * Add build args to docker-compose * Change shell to /bin/bash This is required for "==" syntax in conditional string comparison	2025-04-15 13:36:39 +08:00
hoshi-hiyouga	1134baeedd	[assets] update model readme (#7724 )	2025-04-15 00:41:09 +08:00
Kingsley	2101399c94	[model] Support Kimi_VL thinking/instruct (#7719 ) * add kimi_vl * patch config * check version * Update mm_plugin.py * Update mm_plugin.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-15 00:21:58 +08:00
hoshi-hiyouga	3f91a95250	[misc] fix env vars (#7715 )	2025-04-14 16:04:04 +08:00
hoshi-hiyouga	7c61b35106	[misc] upgrade cli (#7714 )	2025-04-14 15:41:22 +08:00
hoshi-hiyouga	f518bfba5b	[deps] upgrade transformers (#7704 )	2025-04-13 18:11:34 +08:00
Yuxuan Zhang	8162f94db5	[model] add GLM-4-0414 (#7695 ) * Update README_zh.md * update	2025-04-13 17:10:45 +08:00
hoshi-hiyouga	1f0c52b73c	[deps] fix uv conflicts (#7686 ) * fix #7678 * Update setup.py * Update tests.yml * Update publish.yml * Update Makefile	2025-04-11 18:02:24 +08:00
Eric Tang	a8caf09c7f	[data] support for specifying a dataset in cloud storage (#7567 ) * add support for loading datasets from s3/gcs * add comments to readme * run linter and address comments * add option to pass in kwargs to ray init (i.e. runtime env) * address comment * revert mixed up changes	2025-04-10 11:31:35 +08:00
Eric Tang	bb8d79bae2	[ray] allow for specifying ray.init kwargs (i.e. runtime_env) (#7647 ) * ray init kwargs * Update trainer_utils.py * fix ray args --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-10 11:31:05 +08:00
Dain Kim	1c436c9f25	[bugfix] enable_gemma_liger_kernel (#7660 ) - The `enable_liger_kernel` function for the Gemma model series was not executed due to the existing `if` statement in the code. - Changed the line to an `elif` statement so that the `apply_liger_kernel` function is executed properly. resolved: #7628	2025-04-10 11:27:30 +08:00
jilongW	1b0934bccb	[misc] fix cuda warn on intel GPU (#7655 )	2025-04-09 21:37:54 +08:00
hoshi-hiyouga	4eec541857	[data] add coig-p dataset (#7657 )	2025-04-09 21:18:25 +08:00
hoshi-hiyouga	89a4f9ec7f	[assets] update readme (#7654 )	2025-04-09 18:27:38 +08:00
hoshi-hiyouga	1abd71b551	[assets] update readme (#7644 )	2025-04-09 01:06:06 +08:00
Kingsley	349c56c51c	[data] Fix bugs of `use_audio_in_video` in Qwen2.5 Omni (#7638 ) * cache _mm_inputs * nit * support for use_audio_in_video * remove cache * fix data * Update mllm_video_audio_demo.json	2025-04-08 18:40:10 +08:00
Shawn Tao	acb09fa3a3	[trainer] fix key error (#7635 )	2025-04-08 18:39:50 +08:00
Adarsh Shirawalmath	f75b91077b	[sglang] support transformers 4.51.0 (#7639 )	2025-04-08 18:39:23 +08:00
hoshi-hiyouga	c3c0efbaa0	[misc] fix packing and eval plot (#7623 )	2025-04-07 18:20:57 +08:00
hoshi-hiyouga	5115dc8c7f	[assets] update readme (#7612 )	2025-04-06 13:58:49 +08:00
hoshi-hiyouga	831e7f1cfd	[model] add llama4 (#7611 )	2025-04-06 13:42:31 +08:00
Kingsley	d4cfa9507e	[data] fix qwen2.5 omni plugin (#7578 ) * specific entry * Update mm_plugin.py * fix fps cal --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-02 23:58:39 +08:00
Kingsley	d32c6c014d	[data] fix qwen2.5 omni plugin (#7573 ) * align key with qwen2vl * nit && change scripts	2025-04-02 21:28:52 +08:00
gechengze	7b9deb9410	[trainer] fix batch processing in PPO trainer (#7576 )	2025-04-02 21:17:48 +08:00
hoshi-hiyouga	5e22597ff1	[infer] vllm video/audio inference (#7566 )	2025-04-02 02:27:04 +08:00
hoshi-hiyouga	2bfcad2394	[model] fix kv cache (#7564 )	2025-04-01 23:07:46 +08:00
Yu Shi Jie	a13b1bb49a	[model] fix use_cache patching for gemma3 multimodal (#7500 )	2025-04-01 16:06:48 +08:00
Ritesh Goru	d10467d178	[data] specify position_ids in PackedSupervisedDatasetProcessor for neat_packing (#7318 ) * use position_ids for neat_packing with fa2 * revert fa2 changes	2025-04-01 16:03:13 +08:00

1 2 3 4 5 ...

2609 Commits