LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2025-10-14 15:52:49 +08:00

Author	SHA1	Message	Date
hoshi-hiyouga	f518bfba5b	[deps] upgrade transformers (#7704 )	2025-04-13 18:11:34 +08:00
Yuxuan Zhang	8162f94db5	[model] add GLM-4-0414 (#7695 ) * Update README_zh.md * update	2025-04-13 17:10:45 +08:00
Eric Tang	a8caf09c7f	[data] support for specifying a dataset in cloud storage (#7567 ) * add support for loading datasets from s3/gcs * add comments to readme * run linter and address comments * add option to pass in kwargs to ray init (i.e. runtime env) * address comment * revert mixed up changes	2025-04-10 11:31:35 +08:00
hoshi-hiyouga	4eec541857	[data] add coig-p dataset (#7657 )	2025-04-09 21:18:25 +08:00
hoshi-hiyouga	1abd71b551	[assets] update readme (#7644 )	2025-04-09 01:06:06 +08:00
hoshi-hiyouga	c3c0efbaa0	[misc] fix packing and eval plot (#7623 )	2025-04-07 18:20:57 +08:00
hoshi-hiyouga	5115dc8c7f	[assets] update readme (#7612 )	2025-04-06 13:58:49 +08:00
hoshi-hiyouga	2bfcad2394	[model] fix kv cache (#7564 )	2025-04-01 23:07:46 +08:00
Billy Cao	00409ff28a	[data] shard the dataset to allow multiprocessing when streaming is enabled (#7530 ) * Shard the dataset when streaming to allow multiprocessing * Allow user to not set dataset_shards to ensure backward compatibility	2025-04-01 15:36:23 +08:00
Kingsley	7eed496336	[model] add Qwen2.5-Omni model (#7537 ) * preserve image_sizes * preserve image_sizes * init plugin * support audio-text2text lora * nit * support image/video-text2text, audio-text2text * remove args * remove lines * add docs && nit * remove some comments * fix && add merge part script * add license	2025-03-31 20:39:35 +08:00
hoshi-hiyouga	0583d06676	[model] add qwen2vl 32b & upgrade peft (#7469 ) * add qwen2vl 32b * fix ci * upgrade peft to 0.15 * fix ci * fix ci	2025-03-25 12:15:58 +08:00
hoshi-hiyouga	ca42c0c406	[assets] fix gemma3 readme (#7449 )	2025-03-24 10:31:25 +08:00
hoshi-hiyouga	d5915a7dd7	[assets] update videos (#7340 ) * Update README.md * Update README_zh.md	2025-03-17 15:48:02 +08:00
Hertz	ec1154662b	[model] support hunyuan 7b (#7317 ) * [Model]supported tencent-hunyuan model * [Model]supported tencent-hunyuan model(fix) * [Model]supported tencent-hunyuan model(fix)	2025-03-15 20:55:24 +08:00
Qiaolin Yu	a44a53ebec	[inference] support sglang backend (#7278 ) * Mimic SGLang offline Engine * Add more tests and args * Pass all current tests * Clean Code * fix sample_params * clean code * Fix Stream Chat * change sglang from engine mode to server mode * fix * Fix Review Issues * Use SGLang Built-In Utilities * Fix test SGLang * Some Doc Issue * fix sglang engine * add readme --------- Co-authored-by: Jin Pan <jpan236@wisc.edu> Co-authored-by: hiyouga <hiyouga@buaa.edu.cn>	2025-03-15 04:37:58 +08:00
hoshi-hiyouga	0be0d7796a	[assets] update video (#7287 )	2025-03-13 18:45:47 +08:00
hoshi-hiyouga	4b9d8da5a4	[model] support gemma3 (#7273 )	2025-03-13 01:35:23 +08:00
hoshi-hiyouga	522a3e8493	[infer] fix vllm args (#7235 ) Former-commit-id: 999be5b4512890b8cf4f45874a77e35cf35626f5	2025-03-11 01:15:35 +08:00
hoshi-hiyouga	71a1c1321a	[config] update args (#7231 ) Former-commit-id: f71a901840811bf560df671ec63a146ff99140c6	2025-03-10 23:04:43 +08:00
hoshi-hiyouga	9adc0a2c3f	[assets] update readme (#7209 ) Former-commit-id: d1631b38dad9ba3d41aebbb00e3500eb79b9e8e9	2025-03-07 17:27:49 +08:00
hoshi-hiyouga	d2f845d70d	[deps] upgrade vllm (#7183 ) Former-commit-id: 37678a3d64668c3b4a4bfefc054e3b9b40427c1a	2025-03-06 15:25:08 +08:00
hoshi-hiyouga	e62dae37fe	[assets] update wechat (#7106 ) Former-commit-id: 0ea430060994631e9fdb18fbbca0dd565a04fd66	2025-02-28 12:01:04 +08:00
leo-pony	b9f84900ee	[npu] update cann base image and torch 2.4 (#7061 ) * Update base npu container image version:The Python version required for Hugging Face Transformers is >= python3.10 * Fix the bug: arg type of INSTALL_DEEPSPEED shoud been string now. * Update Ascend CANN, CANN-Kernel and corresponding torch and torch-npu version * Upgrade torch-npu needs packages' version: torch==2.1.0 and torch-npu==2.4.0.post2 Former-commit-id: d6dafada58412b0c801e576ef4d8d96203f792af	2025-02-25 23:32:01 +08:00
hoshi-hiyouga	5f65558088	[misc] fix project toml (#7067 ) Former-commit-id: 28a668ff4e0beebfe5387362f5518c1d9343666f	2025-02-25 23:22:48 +08:00
hoshi-hiyouga	ee46011b34	[assets] update readme (#7051 ) Former-commit-id: c89a39bfc6a3f0aaa376cd1b221320f466aba617	2025-02-24 20:45:06 +08:00
hoshi-hiyouga	d55f420206	[assets] update wechat (#7019 ) Former-commit-id: 3d102fe7e0bfc23db7d75f90ebaf53216c54cc85	2025-02-20 20:32:33 +08:00
hoshi-hiyouga	331f53381f	[data] add r1 distill dataset (#6983 ) Former-commit-id: 1da5ee4edaa3896593b9cae488f0ac5917c3243e	2025-02-18 17:25:09 +08:00
hoshi-hiyouga	1d675a287d	[version] support transformers 449 (#6982 ) * support transformers 449 * fix mm plugin Former-commit-id: e9118a9df0839d24f6ddff5a0b55ef101a1d3d22	2025-02-18 17:05:40 +08:00
hoshi-hiyouga	8b8fdb3a85	[misc] update readme (#6918 ) Former-commit-id: f5823479bd51c39db668b68056be749af09894d1	2025-02-13 01:01:41 +08:00
hoshi-hiyouga	290057069e	[misc] update readme (#6917 ) Former-commit-id: 6bbed1d8c4189fb7bea40230e278c40bb5336fbd	2025-02-13 00:58:10 +08:00
hoshi-hiyouga	d58fcd094e	[misc] update readme (#6903 ) Former-commit-id: 830d028939149d54bc91b6bda110dfa5de949483	2025-02-11 22:51:26 +08:00
hoshi-hiyouga	88eafd865b	[misc] support export ollama modelfile (#6899 ) * support export ollama modelfile * update config * add system and num ctx Former-commit-id: 8c2af7466f4015f300b51841db11bcd2505ebf20	2025-02-11 19:52:25 +08:00
Zhangchi Feng	2047eab723	[da'ta] fix minicpmv plugin (#6890 ) * fix template name * tiny fix * support minicpm-o-2.6 * support inference of minicpmv * update readme * support dpo of minicpmv * update init audio * update init audio * [model]fix image process in minicpmo * fix no mm inputs Former-commit-id: cdd19ccd8cec460606b4545e886e932c1c5c5fe1	2025-02-11 13:30:44 +08:00
hoshi-hiyouga	94726bdc8d	[dataset] add openthought (#6866 ) Former-commit-id: 20c748a4f108c0087f0d85377a4aa99126a0beb0	2025-02-09 00:53:01 +08:00
Zhangchi Feng	8f401e37f8	[model] support audio (#6701 ) * support qwen2_audio * improve code * lint * fix * fix * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 5eacb5629e4d7733cd992a63747a1335f2c6a929	2025-02-05 04:59:09 +08:00
neavo	34746d6151	[readme] update flash attention installation instruction on win platform (#6788 ) * Update README_zh.md * Update README.md Former-commit-id: e48d1327fb39cc95f8fbfc746494f67a79471893	2025-02-01 12:43:29 +08:00
hoshi-hiyouga	a28261a866	[model] add mistral small models (#6786 ) Former-commit-id: e5e95c39bc4199fa89c67e34f9adaaa987058744	2025-02-01 04:31:38 +08:00
hoshi-hiyouga	800de98dc8	[model] add qwen2.5 vl models (#6779 ) Former-commit-id: ed46fb4f6194c30060b908092464dded12e5787c	2025-01-31 03:00:29 +08:00
hoshi-hiyouga	222423bcef	[breaking] support transformers 4.48 (#6628 ) Former-commit-id: f154ab175c513a4d7bb866bf2cffc34b77b50508	2025-01-31 01:36:33 +08:00
hoshi-hiyouga	e71737351f	[webui] improve webui & reasoning mode (#6778 ) Former-commit-id: 3f17fc0d7163372e0446f1a38792ff761e99b739	2025-01-31 00:09:21 +08:00
qvlehao	4f298894da	[model] add deepseek-R1 & show think process (#6767 ) Former-commit-id: 4dccb724af51208a001c96fefbdbf226be09e50c	2025-01-29 12:16:26 +08:00
hoshi-hiyouga	e4046bdd1f	[assets] update wechat (#6692 ) Former-commit-id: 70dba5fab6f4c9225758cafb646113d8e80ac084	2025-01-18 12:35:03 +08:00
hoshi-hiyouga	ef994600db	update readme (#6648 ) Former-commit-id: b47467276ab3174c50329b3c8b76823bc0a2249c	2025-01-15 11:06:19 +08:00
hoshi-hiyouga	7638f1070e	[optim] clean apollo (#6645 ) * clean apollo code * update readme Former-commit-id: 38b8ec4a99189483124b54df9d6bc6b0d318855a	2025-01-15 01:42:50 +08:00
Zhangchi Feng	66184762e8	update readme of MiniCPM-o (#6642 ) * fix template name * tiny fix * support minicpm-o-2.6 * support inference of minicpmv * update readme Former-commit-id: 68604050ae2c98aeef5e9a6b4d2c11a4eb609bfa	2025-01-14 21:22:35 +08:00
hoshi-hiyouga	41a9e231cb	lint (#6641 ) Former-commit-id: 79731ae13ecd17eb8646fb53162c81dddfef3b00	2025-01-14 18:40:07 +08:00
Haian Huang(深度眸)	1bb06e06df	Support InternLM3 Dense 8B Model (#6640 ) * support internlm3 * update * update * update * add hint Former-commit-id: 24ab7ae0944c5f373e9cac60f0332e704824a057	2025-01-14 18:07:27 +08:00
Zhangchi Feng	ae32c148d1	Support new features of MiniCPM-V (#6626 ) * fix template name * tiny fix * support minicpm-o-2.6 Former-commit-id: 53034a61c7654358f46916cbc370910fb2aeff3b	2025-01-14 00:26:19 +08:00
hoshi-hiyouga	2a05941b14	[inference] fix stop token for object detection (#6624 ) * fix stop token * update minicpm data pipeline * fix npu qlora examples Former-commit-id: 844919fadaa8a61dfae47020971ea80730b2346f	2025-01-13 21:34:20 +08:00
codingma	11c38b9173	add nf4 qlora support on Ascend NPU (#6601 ) * add nf4 qlora support on Ascend NPU * add transformers version check * add python>=3.10 requirement description for npu * tiny fix --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 7912d1acac5f10dab22145fe729a90c57aad8d85	2025-01-13 19:43:36 +08:00

1 2 3 4 5 ...

502 Commits