hoshi-hiyouga
|
7754024e9b
|
Update trainer.py
Former-commit-id: 8565d4b43db905374c328ae57c71fc226980d14f
|
2024-06-03 22:08:38 +08:00 |
|
enji.zhou
|
b4913569a8
|
fix KTO Trainer Sampler
Former-commit-id: 39eb1bfa272011554322e9bb2534f83b68282a70
|
2024-06-03 21:32:38 +08:00 |
|
hiyouga
|
9b551309de
|
update dpo, kto trainer
Former-commit-id: 4a6cc3c7046f8b27d05ea53ef216bab6fa7ebfaf
|
2024-05-29 00:14:29 +08:00 |
|
hiyouga
|
9fed4a2ef4
|
clean kto trainer
Former-commit-id: 76402bd78cbd3a99a544f0ac019468b569b0e1d1
|
2024-05-28 21:43:26 +08:00 |
|
hiyouga
|
b0d9966663
|
support SimPO #3900
Former-commit-id: 6b954ce60155cf8334150b795cfc4bb63ca74c8b
|
2024-05-26 23:46:33 +08:00 |
|
hiyouga
|
bf59383783
|
refactor data preprocessing, fix mllm rlhf
Former-commit-id: 53ff2dd24f9121ea30c95063bb72e49a9b31e980
|
2024-05-24 04:08:25 +08:00 |
|
hiyouga
|
2bff90719b
|
improve KTO impl., replace datasets
Former-commit-id: e56a57ddcf061de6e4acc8679f7dbf0b68364986
|
2024-05-18 03:44:56 +08:00 |
|
enji.zhou
|
66b5634ebf
|
add kto
Former-commit-id: ec51986cf70b0bdd79b8141e45916670fb97a08e
|
2024-05-17 13:09:17 +08:00 |
|