43 Commits

Author SHA1 Message Date
hzhaoy
e1751f6398 fix #4579
Former-commit-id: 677c86594e4ea904fde0a557852daf54636b06ae
2024-06-27 13:49:57 +08:00
hiyouga
28e613efd0 fix #4458
Former-commit-id: 8d6cd69ac43afd4bd7c14bd02b0061455827ac9e
2024-06-26 19:52:35 +08:00
hiyouga
ad0304e147 fix #4379
Former-commit-id: cc016461e63a570142b56d50a5d11e55a96ab8db
2024-06-25 02:31:44 +08:00
hiyouga
a225b5a70c tiny fix about badam
Former-commit-id: 095fab58d3692607c9e78747b4218ae1abcf5aaf
2024-06-25 01:54:53 +08:00
hoshi-hiyouga
fe6ef6400c Merge pull request #4352 from Ledzy/main
[Enhancement] Support ZeRO-3 when using BAdam

Former-commit-id: d0f953bf5bdbfd49acc82ff055bd54889241761a
2024-06-25 01:49:13 +08:00
hiyouga
7be502c5c5 update readme
Former-commit-id: e507e60638b2e8c66f24805b3b28f6b9f98f5924
2024-06-24 18:22:12 +08:00
hiyouga
7735456561 fix templates
Former-commit-id: 4cff6a4ad55b24bf57db6be5cf817180c1ea5626
2024-06-19 17:44:05 +08:00
Jonery
c779899f7b Cleaner integration.
Former-commit-id: 5c2ff1b749a265dd3c979189ec491d8ac911a6f6
2024-06-19 12:29:40 +08:00
Jonery
3a5eacb4cf Support distributed BAdam.
Former-commit-id: 0f72aac8c9227e33ad20d2b1641b1c9faae16a5f
2024-06-18 12:27:47 +08:00
Jonery
5d59f6562a Merge remote-tracking branch 'upstream/main'
Former-commit-id: ea1f3ba5e030504e07053484f50f4cbdb37808bc
2024-06-17 18:44:51 +08:00
Jonery
756566342d adapt for badam with ds zero3
Former-commit-id: 33b437277846d4f0b64c13a0bc892ef4f345a21e
2024-06-17 18:18:10 +08:00
hiyouga
ce4a27a5f7 fix tol
Former-commit-id: 46093b5786611d99adf1fd3d42926a728fc629f8
2024-06-16 01:38:44 +08:00
hiyouga
f25b8626bf support pissa
Former-commit-id: 8c1046d78ac6c8f9429b73617e35e1eccb35138f
2024-06-16 01:08:12 +08:00
hiyouga
c0c6b8075a tiny fix
Former-commit-id: 38b6b0f52edeb8ba45aa03b415b3c0c1b0e0c1e4
2024-06-16 01:06:41 +08:00
hiyouga
2946153cea add license
Former-commit-id: d87108daa68bd40174b262be1ca65fe6e1b7ab56
2024-06-15 17:54:33 +08:00
hiyouga
ab66ae8cd2 fix #4295
Former-commit-id: 78589cf90c6e12e612f269b1c771f19f3dad83d2
2024-06-15 04:34:55 +08:00
hiyouga
a3f4925c2c add test cases
Former-commit-id: b27269bd2b52fb9d43cde8a8b7f293099b0127a2
2024-06-15 04:05:54 +08:00
hiyouga
8fccaf20c5 fix #4221
Former-commit-id: 6baafd4eb3147ad9f7d2952b8eb27c5486940f36
2024-06-13 02:48:21 +08:00
hiyouga
81ed4d8abf fix #4209
DeepSpeed ZeRO3 has inflight param error when calling model.eval()


Former-commit-id: cf9f2d6c42b5a37038c9eededbb767eae6a3f67d
2024-06-13 02:25:50 +08:00
hiyouga
833aa324c2 clean code
Former-commit-id: 2ed8270112755971e3f2dfd2f29c5939b077330a
2024-06-13 01:58:16 +08:00
hiyouga
5834651c4a fix #4198
Former-commit-id: 89f2bd8c8c035181927bd530a7ffc733407d674c
2024-06-11 15:38:38 +08:00
hiyouga
ca9468ff04 tiny fix
Former-commit-id: f8d8690bf4c2981f3151b4ccf07daeb4f3cd38a9
2024-06-07 05:19:21 +08:00
hiyouga
4f3c89a6eb fix ppo trainer save zero3 model
accelerator.get_state_dict(ds_model) should be called at all ranks


Former-commit-id: 4489d73ac75c8dbc002fc16c854148994d432c3a
2024-06-07 05:14:19 +08:00
hiyouga
f76d427332 fix ppo in trl 0.8.6
Former-commit-id: 2702d7e952523b584d67c8901888b492d4a79b14
2024-06-07 04:48:29 +08:00
hiyouga
d3196318be fix #4120
Former-commit-id: f9e818d79cf686cb34789327add7ed1f749966c6
2024-06-07 04:18:05 +08:00
hiyouga
8da149ba40 rename files
Former-commit-id: 74f96efef9bcd63f65d0190c901ff9be54ccd350
2024-06-07 00:09:06 +08:00
hiyouga
368695483d fix ppo+zero3 #3108
Former-commit-id: 76c61905b20f69fac5c7a6c4ea9450bf33d3b1f2
2024-06-06 23:30:07 +08:00
hiyouga
e0aadd4b34 fix ppo dataset bug #4012
Former-commit-id: 149610c636bbb974e546d13fa302884ea65a6d38
2024-06-06 19:03:20 +08:00
hiyouga
e898d8bbc4 update trainers
Former-commit-id: fad2591e314093335ef1c301d0a70f0cbe935728
2024-06-06 18:45:49 +08:00
hiyouga
a16786d8ba fix #4090
Former-commit-id: 67fe822324a9f830175e44f89acdd9d759b38852
2024-06-06 00:50:32 +08:00
hiyouga
6f7b6ae0c3 remove gc warnings in DPO&KTO
Former-commit-id: f9a206509ec8cd3abfad8bd924c9387317a4ead8
2024-06-03 22:53:54 +08:00
hoshi-hiyouga
5d96cf146e Update trainer.py
Former-commit-id: 24499f40dc1d9db448a3328d2a75c60eec27feb9
2024-06-03 22:08:38 +08:00
enji.zhou
e58aca0602 fix KTO Trainer Sampler
Former-commit-id: 34a2c5087a174a807e5a11cae3748bcaaaf13550
2024-06-03 21:32:38 +08:00
Uminosachi
0de4e1e9e2 Set scheduler_specific_kwargs to get_scheduler
Former-commit-id: 14e97dc1192f6cf94ab99eb3a9b8c64029040384
2024-05-31 13:45:39 +09:00
hiyouga
468d0e7ed1 10x generate in ppo w/ zero3
https://github.com/huggingface/trl/pull/1483

Former-commit-id: 65cd8bdbdbe1b19250ecd813aeb72c8e00ef2f9c
2024-05-29 00:23:23 +08:00
hiyouga
bfac965f9c update dpo, kto trainer
Former-commit-id: 7c8e01bb74bb2d2da5dba5059a9c262e4730b802
2024-05-29 00:14:29 +08:00
hiyouga
14f6cc2b7c clean kto trainer
Former-commit-id: 900e1ea622a2ffa45c5e2a359471962563fabca7
2024-05-28 21:43:26 +08:00
hiyouga
4807c11db8 support SimPO #3900
Former-commit-id: cb63b32986c43f97994211ec34dc5928fc3bb9d7
2024-05-26 23:46:33 +08:00
hiyouga
3e729798df refactor data preprocessing, fix mllm rlhf
Former-commit-id: 3a023bca2a502810a436cfba7708df164754ea62
2024-05-24 04:08:25 +08:00
hiyouga
519d2511ae improve data process logger
Former-commit-id: a851056229f37391023627180b5712ed64ae3528
2024-05-18 22:02:42 +08:00
hiyouga
13d7b48efe improve KTO impl., replace datasets
Former-commit-id: c450ee87a35ff9235f9b695b0de2e042b2971178
2024-05-18 03:44:56 +08:00
enji.zhou
03956053b8 add kto
Former-commit-id: db1d5a4f51faae61fe18666057353747b01f5b8d
2024-05-17 13:09:17 +08:00
hiyouga
cae823ddf0 rename package
Former-commit-id: 308edbc4260d45907b4a9d3a45ec21d83e48aacb
2024-05-16 18:39:08 +08:00