LLaMA-Factory

mirror of https://github.com/hiyouga/LLaMA-Factory.git synced 2026-03-09 21:25:59 +08:00

Author	SHA1	Message	Date
Zhangchi Feng	69fcc8e0c0	[data] fix minicpmo template (#6946 ) Former-commit-id: 09e4438b58d5c1a5fdde37ff781c3d79461c4743	2025-02-15 00:37:41 +08:00
Eric Tang	413aa5944a	[ray] specify ray storage path (#6920 ) Former-commit-id: 4be6b66b1eaa79955e936ce2b747a8837ecd1e49	2025-02-14 21:55:41 +08:00
hoshi-hiyouga	1cda37892e	[misc] fix lora regex (#6944 ) * fix lora regex * fix Former-commit-id: 1d0ecbaee1b72f1e03154ddd4fcc8b7876e01f89	2025-02-14 21:38:43 +08:00
hoshi-hiyouga	6ebe81e04d	[misc] fix grad ckpt (#6931 ) Former-commit-id: deae1fc9a0bea5c8b8be1564cf9c81c9c02a0b3a	2025-02-13 23:27:51 +08:00
hoshi-hiyouga	a9b4e229af	[model] add liger kernel to qwen2_5 vl (#6930 ) * add liger kernel to qwen2_5 vl * fix patch * fix patch Former-commit-id: 828776d155986166498dfc907194f64436571106	2025-02-13 23:05:54 +08:00
Billy Cao	680648098b	[trainer] fix gen_kwarg to eval during training (#5451 ) * Correctly pass gen_kwarg to eval during model runs * fix * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 845d16122496311e08263610a6a922f82604de7b	2025-02-13 02:35:06 +08:00
SrWYG	d9ea4baf00	[data] evaluate on each dataset (#5522 ) * [Update] loader.py , evaluate will run separate evaluations on each dataset. `If you pass a dictionary with names of datasets as keys and datasets as values, evaluate will run separate evaluations on each dataset. This can be useful to monitor how training affects other datasets or simply to get a more fine-grained evaluation` seq2seqtrainner support eval_dataset as Dict. * fix format * fix * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: cf00f78650a442c85678ce805e030d2b96cbecd7	2025-02-13 02:19:03 +08:00
Noah	f1c2ae9d47	[data] improve error handling (#6128 ) * sync from upstream * update * update * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 1569e6096fec07da5583f1a3435b0d23ae09b5ba	2025-02-13 01:39:41 +08:00
hoshi-hiyouga	a7448ef8f3	[misc] update readme (#6918 ) Former-commit-id: f5823479bd51c39db668b68056be749af09894d1	2025-02-13 01:01:41 +08:00
hoshi-hiyouga	98b20233ae	[misc] update readme (#6917 ) Former-commit-id: 6bbed1d8c4189fb7bea40230e278c40bb5336fbd	2025-02-13 00:58:10 +08:00
hoshi-hiyouga	efef4eaefc	[breaking change] refactor data pipeline (#6901 ) * refactor data * rename file Former-commit-id: 7a1a4ce6451cb782573d0bd9dd27a5e443e3a18b	2025-02-13 00:39:20 +08:00
Eric Tang	19ba216571	[misc] support for launching LLaMA-Factory with `uv run` (#6907 ) * yay * uv with ray temporary commit * remove ray specific code for now * cleanup Former-commit-id: 1a9cab6de49e300bf9c747eefbb11d693592b477	2025-02-13 00:38:44 +08:00
Eric Tang	24ad208345	[example] fix path to ray example (#6906 ) Former-commit-id: e9bee3ef045d85051da04e6ad581a23a9e1a9551	2025-02-13 00:29:32 +08:00
hoshi-hiyouga	8cbfa350fd	[misc] fix grad ckpt func (#6916 ) Former-commit-id: 35e069a52b3d7cfd9b0107574b09265eb2290f0b	2025-02-13 00:17:18 +08:00
marko1616	a23e16ae89	[trainer] fix llama3.2 vision kto train (#6904 ) Former-commit-id: 1563e89adc8988fc6e4250634a3f1e385979b0e5	2025-02-12 19:09:14 +08:00
hoshi-hiyouga	0a0d7671e0	[data] feat: auto template (#6905 ) * support auto template * add unittest Former-commit-id: 0c6c9150db6414a5a05527ea486dce6633dff4b3	2025-02-12 00:22:53 +08:00
hoshi-hiyouga	c14899b0a8	[misc] update readme (#6903 ) Former-commit-id: 830d028939149d54bc91b6bda110dfa5de949483	2025-02-11 22:51:26 +08:00
hoshi-hiyouga	c5649d7149	[data] fix ollama template (#6902 ) * fix ollama template * add meta info * use half precision Former-commit-id: 1304bbea69d8c8ca57140017515dee7ae2ee6536	2025-02-11 22:43:09 +08:00
hoshi-hiyouga	ca5cd8276c	[misc] support export ollama modelfile (#6899 ) * support export ollama modelfile * update config * add system and num ctx Former-commit-id: 8c2af7466f4015f300b51841db11bcd2505ebf20	2025-02-11 19:52:25 +08:00
hoshi-hiyouga	5954d4cfbc	[data] refactor template (#6896 ) Former-commit-id: f78d5a3eca947ed965ca2f6c87d60441b1a59867	2025-02-11 17:59:25 +08:00
codingma	185319802f	support ollama modelfile export (#4686 ) Former-commit-id: 15cca102a7fc0d08b5d049cf264acc6fa576b104	2025-02-11 17:52:24 +08:00
hoshi-hiyouga	64069e75d7	[data] refactor mm plugin (#6895 ) * refactor plugin * lint Former-commit-id: 1c8dcc3adca4a2e78f514f8bb70573dd1ca08746	2025-02-11 16:34:49 +08:00
HJ	02264873fd	[data] fix qwen_2_5_vl video processing (#6868 ) * fix qwen_2_5_vl video processing * Update mm_plugin.py * Update mm_plugin.py --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 35f326dabdc8e84036296d2e3de1c84c67b8def8	2025-02-11 16:14:50 +08:00
hoshi-hiyouga	fc43b83490	[assets] update wechat (#6892 ) Former-commit-id: 0b268cc903a583ae78cb7e63d2bdc4602d7220fc	2025-02-11 13:56:26 +08:00
Zhangchi Feng	5cf98e2084	[da'ta] fix minicpmv plugin (#6890 ) * fix template name * tiny fix * support minicpm-o-2.6 * support inference of minicpmv * update readme * support dpo of minicpmv * update init audio * update init audio * [model]fix image process in minicpmo * fix no mm inputs Former-commit-id: cdd19ccd8cec460606b4545e886e932c1c5c5fe1	2025-02-11 13:30:44 +08:00
HJ	96b6865ba5	[data] fix: sharegpt converter (#6879 ) * fix-sharegpt-format * fix --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: ae8f8151ff750839998b50446f127061f240d41a	2025-02-10 21:59:12 +08:00
hoshi-hiyouga	4d30538c03	[data] fix mllama collator (#6874 ) Former-commit-id: c694fa3d66651c6ce547fa72c8260c46a406126b	2025-02-09 22:42:25 +08:00
hoshi-hiyouga	78779441f0	[test] align test cases (#6865 ) * align test cases * fix function formatter Former-commit-id: a68f5e22d0391c80a9a826dc83967255be572032	2025-02-09 01:03:49 +08:00
hoshi-hiyouga	9204641049	[dataset] add openthought (#6866 ) Former-commit-id: 20c748a4f108c0087f0d85377a4aa99126a0beb0	2025-02-09 00:53:01 +08:00
hoshi-hiyouga	c322512037	[deps] upgrade vllm (#6857 ) Former-commit-id: 4bd50f65a3d62528768561019fda2723d045c7fd	2025-02-08 15:02:28 +08:00
hoshi-hiyouga	f5df7f8453	fix qwen2vl plugin (#6855 ) Former-commit-id: fd13b7138ab3f4da0a429a327b9d076bcb70b944	2025-02-08 10:59:10 +08:00
hoshi-hiyouga	38c52f20f7	[misc] allow extra args (#6831 ) Former-commit-id: 0fd3a5295cb4e08a4e57e860e82103364c28fba8	2025-02-06 12:38:08 +08:00
Zhangchi Feng	46a1786595	[model] support audio (#6701 ) * support qwen2_audio * improve code * lint * fix * fix * fix --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 5eacb5629e4d7733cd992a63747a1335f2c6a929	2025-02-05 04:59:09 +08:00
Yueqi Song	14d6188852	[data] allow thought in function call (#6797 ) * Update template.py * Update template.py * use formatter * fix regex --------- Co-authored-by: hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 3a31af6e920683ec074da93b1719e29f5d4cffd6	2025-02-05 02:26:23 +08:00
hoshi-hiyouga	40b6e9045d	[misc] update license year & fix llama pro (#6814 ) * fix llamapro script * change year Former-commit-id: d9ae594178796994d400a5f207d6499712816f89	2025-02-05 01:53:33 +08:00
Yueqi Song	34420988c5	[data] fix qwen tool template (#6796 ) * Update tool_utils.py * fix unittest --------- Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn> Former-commit-id: 02bb78a792112f5151b3a96ddde2528823855288	2025-02-05 00:02:00 +08:00
Zhangchi Feng	5f9e4d01bd	[data] fix minicpmv plugin (#6801 ) * fix template name * tiny fix * support minicpm-o-2.6 * support inference of minicpmv * update readme * support dpo of minicpmv * update init audio * update init audio * [model]fix image process in minicpmo Former-commit-id: 8f704c8b6228ef50f828014f85dce67fda868660	2025-02-04 21:20:15 +08:00
neavo	48ddec0591	[readme] update flash attention installation instruction on win platform (#6788 ) * Update README_zh.md * Update README.md Former-commit-id: e48d1327fb39cc95f8fbfc746494f67a79471893	2025-02-01 12:43:29 +08:00
hoshi-hiyouga	eb24fb5d4d	[misc] update workflows (#6787 ) Former-commit-id: 15add6b250149e2aeabdc62d7dca69fc06054e01	2025-02-01 04:54:42 +08:00
hoshi-hiyouga	e335c548c1	[model] add mistral small models (#6786 ) Former-commit-id: e5e95c39bc4199fa89c67e34f9adaaa987058744	2025-02-01 04:31:38 +08:00
hoshi-hiyouga	1132aaa53c	[model] add qwen2.5 vl models (#6779 ) Former-commit-id: ed46fb4f6194c30060b908092464dded12e5787c	2025-01-31 03:00:29 +08:00
hoshi-hiyouga	46068b3324	[breaking] support transformers 4.48 (#6628 ) Former-commit-id: f154ab175c513a4d7bb866bf2cffc34b77b50508	2025-01-31 01:36:33 +08:00
hoshi-hiyouga	5abd0168fa	[webui] improve webui & reasoning mode (#6778 ) Former-commit-id: 3f17fc0d7163372e0446f1a38792ff761e99b739	2025-01-31 00:09:21 +08:00
qvlehao	88546138e1	[model] add deepseek-R1 & show think process (#6767 ) Former-commit-id: 4dccb724af51208a001c96fefbdbf226be09e50c	2025-01-29 12:16:26 +08:00
yinpu	5062b099f7	fix: avoid redundant normalization in DPO's SFT loss calculation (#6722 ) Former-commit-id: 971a8ccbdacf130763d40c7ef82a711b2fc1292f	2025-01-21 13:38:02 +08:00
engchina	f34390b596	[webui] support ja (#6698 ) * add support for japanese language * add support for japanese language --------- Co-authored-by: engchina <atjapan2015@gmail.com> Former-commit-id: 88692e403f9b5085dd0c7c2b2c68656c5da50dd4	2025-01-20 19:46:38 +08:00
hoshi-hiyouga	87db2a849a	[model] support yarn (#6693 ) Former-commit-id: 8c412abc44a4c61b683465e36c6288580d980250	2025-01-18 13:56:09 +08:00
hoshi-hiyouga	be3525910d	[assets] update wechat (#6692 ) Former-commit-id: 70dba5fab6f4c9225758cafb646113d8e80ac084	2025-01-18 12:35:03 +08:00
hoshi-hiyouga	51fe95c0be	[misc] update mm plugin (#6691 ) Former-commit-id: 00303338d6927b1fda58b23340a31a8fa009f706	2025-01-17 23:04:26 +08:00
hoshi-hiyouga	01fa59bb6b	disable valset by default (#6690 ) Former-commit-id: a1a94f364e33d1d73852f74eda4fa581e6b16533	2025-01-17 21:09:30 +08:00

1 2 3 4 5 ...

2612 Commits