LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2026-07-31 21:26:09 +08:00

Author	SHA1	Message	Date
Kingsley	d4cfa9507e	[data] fix qwen2.5 omni plugin (#7578 ) * specific entry * Update mm_plugin.py * fix fps cal --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>	2025-04-02 23:58:39 +08:00
Kingsley	d32c6c014d	[data] fix qwen2.5 omni plugin (#7573 ) * align key with qwen2vl * nit && change scripts	2025-04-02 21:28:52 +08:00
hoshi-hiyouga	5e22597ff1	[infer] vllm video/audio inference (#7566 )	2025-04-02 02:27:04 +08:00
hoshi-hiyouga	2bfcad2394	[model] fix kv cache (#7564 )	2025-04-01 23:07:46 +08:00
Ritesh Goru	d10467d178	[data] specify position_ids in PackedSupervisedDatasetProcessor for neat_packing (#7318 ) * use position_ids for neat_packing with fa2 * revert fa2 changes	2025-04-01 16:03:13 +08:00
Billy Cao	00409ff28a	[data] shard the dataset to allow multiprocessing when streaming is enabled (#7530 ) * Shard the dataset when streaming to allow multiprocessing * Allow user to not set dataset_shards to ensure backward compatibility	2025-04-01 15:36:23 +08:00
Hao	d70b3b4bc5	[trainer] new kto mismatch pair creation strategy (#7509 )	2025-04-01 15:21:53 +08:00
hoshi-hiyouga	e76eba051d	[data] fix qwen2.5 omni collator (#7553 )	2025-04-01 00:15:12 +08:00
Kingsley	7eed496336	[model] add Qwen2.5-Omni model (#7537 ) * preserve image_sizes * preserve image_sizes * init plugin * support audio-text2text lora * nit * support image/video-text2text, audio-text2text * remove args * remove lines * add docs && nit * remove some comments * fix && add merge part script * add license	2025-03-31 20:39:35 +08:00
Kingsley	8da1d2fa71	[data] fix pixtral plugin (#7505 ) * preserve `image_sizes` * add comments	2025-03-27 17:06:40 +08:00
hoshi-hiyouga	7203365b80	[trainer] fix vlm loss for transformers 4.49 (#7448 )	2025-03-24 10:24:05 +08:00
hoshi-hiyouga	05b19d6952	[deps] upgrade transformers to 4.50.0 (#7437 ) * upgrade transformers * fix hf cache * fix dpo trainer	2025-03-23 17:44:27 +08:00
hoshi-hiyouga	63752fccf7	[assets] update wechat (#7361 )	2025-03-18 21:31:09 +08:00
hoshi-hiyouga	128b5b12b3	[data] fix template (#7349 )	2025-03-17 23:45:20 +08:00
Hertz	ec1154662b	[model] support hunyuan 7b (#7317 ) * [Model]supported tencent-hunyuan model * [Model]supported tencent-hunyuan model(fix) * [Model]supported tencent-hunyuan model(fix)	2025-03-15 20:55:24 +08:00
hoshi-hiyouga	93e6184cbe	[data] gemma3 plugin pan and scan (#7294 ) * gemma3 pan and scan * add test case * fix test	2025-03-13 23:29:23 +08:00
Ritesh Goru	480369a9f2	[data] efficient 4d_attention_mask creation in neat_packing (#7272 )	2025-03-13 03:31:12 +08:00
hoshi-hiyouga	650a9a9057	[misc] update format (#7277 )	2025-03-13 02:53:08 +08:00
hoshi-hiyouga	4b9d8da5a4	[model] support gemma3 (#7273 )	2025-03-13 01:35:23 +08:00
hoshi-hiyouga	264538cb26	[misc] upgrade format to py39 (#7256 )	2025-03-12 00:08:41 +08:00
hiyouga	478e8194d9	remove exit in preprocess Former-commit-id: f369b6ef41ffd9586ba568b88c5ff32a1af4bace	2025-03-11 15:08:25 +08:00
hoshi-hiyouga	16419b2834	[data] fix loader (#7207 ) * fix dataloader * add test case * fix type * fix ci * fix ci * fix ci * disable overwrite cache in ci Former-commit-id: e84af0e140b1aafd1a6d6fe185a8e41c8fc5f831	2025-03-07 17:20:46 +08:00
ZhangChuanhui	151ef48b40	[data] fix function formatter (#7201 ) Co-authored-by: zhangchuanhui <zhangchal@digitalchina.com> Former-commit-id: 3efb32b986170d2839e526640f85ba230715879a	2025-03-07 15:17:23 +08:00
hoshi-hiyouga	bb8aba5abf	[data] fix mm template (#7181 ) Former-commit-id: 648616d473c81d393592806307e3e25b159cb278	2025-03-06 15:18:32 +08:00
hoshi-hiyouga	7b985f55db	[trainer] update config (#7174 ) Former-commit-id: 9f535d0e3c4ee3cd0f1b65218c2eee5d03f43c6f	2025-03-05 23:32:54 +08:00
sirui.li	fd0357a26d	[data] fix qwen2audio plugin (#7166 ) * Update pairwise.py [data]Repair multimodal model dpo training * Update pairwise.py [data]repair multimodal model dpo training using deepcopy * Update pairwise.py * Update mm_plugin.py Former-commit-id: 86763dfdb8e9e5668c1ddd7e924e4be76bf78368	2025-03-05 18:03:36 +08:00
hoshi-hiyouga	31f9daa362	[data] use bicubic resampler (#7143 ) Former-commit-id: c708f19ab0ab57526134952afddaa90aae8decbf	2025-03-04 00:17:06 +08:00
rabbit	19a6916d80	[data] bailing template (#7117 ) * add bailing template * add bailing template * add bailing template --------- Co-authored-by: chengshiwen.csw@antgroup.com <chengshiwen.csw@antgroup.com> Former-commit-id: 4a36f5e0abb5a63f4b3b81560bb1ad0e6832d379	2025-03-03 15:33:22 +08:00
hoshi-hiyouga	065f7fb5da	[data] fix mllama (#7053 ) * fix mllama * fix test Former-commit-id: f5af20a63f3d59a6a68d323a7c6f68e551edb3a3	2025-02-24 22:05:38 +08:00
hoshi-hiyouga	c1d5073bd3	[model] add models (#7054 ) * add qwen25vl awq models * add moonlight Former-commit-id: ae3be2970fea8a35907202a313ab767381c44916	2025-02-24 22:05:13 +08:00
Zhangchi Feng	fcf75633a0	[data] fix MiniCPMV plugin (#6998 ) * fix template * fix bug in messages processing Former-commit-id: f98b828f53968fb9c72bff9e45510ad5586c4fab	2025-02-19 19:36:04 +08:00
hoshi-hiyouga	1d675a287d	[version] support transformers 449 (#6982 ) * support transformers 449 * fix mm plugin Former-commit-id: e9118a9df0839d24f6ddff5a0b55ef101a1d3d22	2025-02-18 17:05:40 +08:00
hoshi-hiyouga	be33ef67fb	[misc] fix script (#6977 ) Former-commit-id: 775efa1d8cbdb1b7d122be2a986d47f85214e0a1	2025-02-18 17:00:46 +08:00
hoshi-hiyouga	f5cd17881e	[data] update vlm args (#6976 ) Former-commit-id: c28e710636a0286d4b8a1d494529b25168a8f3ab	2025-02-18 02:12:51 +08:00
hoshi-hiyouga	c09b648934	[data] add min resolution option (#6975 ) Former-commit-id: 76bd9a98a2fb00f1a1d881e6e1364c02fd36d327	2025-02-18 01:40:46 +08:00
hoshi-hiyouga	f2fd9d1b25	[data] fix predict dataset (#6972 ) Former-commit-id: f9a82e527877b1ed47cabb3d34f4d155705f4048	2025-02-17 20:29:40 +08:00
Zhangchi Feng	167342af8a	[data] fix minicpmo template (#6946 ) Former-commit-id: 09e4438b58d5c1a5fdde37ff781c3d79461c4743	2025-02-15 00:37:41 +08:00
hoshi-hiyouga	a893505924	[misc] fix lora regex (#6944 ) * fix lora regex * fix Former-commit-id: 1d0ecbaee1b72f1e03154ddd4fcc8b7876e01f89	2025-02-14 21:38:43 +08:00
SrWYG	a4c4b8496f	[data] evaluate on each dataset (#5522 ) * [Update] loader.py , evaluate will run separate evaluations on each dataset. `If you pass a dictionary with names of datasets as keys and datasets as values, evaluate will run separate evaluations on each dataset. This can be useful to monitor how training affects other datasets or simply to get a more fine-grained evaluation` seq2seqtrainner support eval_dataset as Dict. * fix format * fix * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: cf00f78650a442c85678ce805e030d2b96cbecd7	2025-02-13 02:19:03 +08:00
Noah	38c9641777	[data] improve error handling (#6128 ) * sync from upstream * update * update * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 1569e6096fec07da5583f1a3435b0d23ae09b5ba	2025-02-13 01:39:41 +08:00
hoshi-hiyouga	46203856fc	[breaking change] refactor data pipeline (#6901 ) * refactor data * rename file Former-commit-id: 7a1a4ce6451cb782573d0bd9dd27a5e443e3a18b	2025-02-13 00:39:20 +08:00
marko1616	0c0cdc26bc	[trainer] fix llama3.2 vision kto train (#6904 ) Former-commit-id: 1563e89adc8988fc6e4250634a3f1e385979b0e5	2025-02-12 19:09:14 +08:00
hoshi-hiyouga	2581cc844b	[data] feat: auto template (#6905 ) * support auto template * add unittest Former-commit-id: 0c6c9150db6414a5a05527ea486dce6633dff4b3	2025-02-12 00:22:53 +08:00
hoshi-hiyouga	86063e27ea	[data] fix ollama template (#6902 ) * fix ollama template * add meta info * use half precision Former-commit-id: 1304bbea69d8c8ca57140017515dee7ae2ee6536	2025-02-11 22:43:09 +08:00
hoshi-hiyouga	88eafd865b	[misc] support export ollama modelfile (#6899 ) * support export ollama modelfile * update config * add system and num ctx Former-commit-id: 8c2af7466f4015f300b51841db11bcd2505ebf20	2025-02-11 19:52:25 +08:00
hoshi-hiyouga	3f7bd98bfa	[data] refactor template (#6896 ) Former-commit-id: f78d5a3eca947ed965ca2f6c87d60441b1a59867	2025-02-11 17:59:25 +08:00
hoshi-hiyouga	808ff89a2d	[data] refactor mm plugin (#6895 ) * refactor plugin * lint Former-commit-id: 1c8dcc3adca4a2e78f514f8bb70573dd1ca08746	2025-02-11 16:34:49 +08:00
HJ	6d7f1299bd	[data] fix qwen_2_5_vl video processing (#6868 ) * fix qwen_2_5_vl video processing * Update mm_plugin.py * Update mm_plugin.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 35f326dabdc8e84036296d2e3de1c84c67b8def8	2025-02-11 16:14:50 +08:00
Zhangchi Feng	2047eab723	[da'ta] fix minicpmv plugin (#6890 ) * fix template name * tiny fix * support minicpm-o-2.6 * support inference of minicpmv * update readme * support dpo of minicpmv * update init audio * update init audio * [model]fix image process in minicpmo * fix no mm inputs Former-commit-id: cdd19ccd8cec460606b4545e886e932c1c5c5fe1	2025-02-11 13:30:44 +08:00
HJ	e11b40c344	[data] fix: sharegpt converter (#6879 ) * fix-sharegpt-format * fix --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: ae8f8151ff750839998b50446f127061f240d41a	2025-02-10 21:59:12 +08:00

1 2 3 4 5 ...

339 Commits