hiyouga
|
8c841d6f76
|
tiny fix
Former-commit-id: 0ee159654ac6339c162745b004e2152ba6fe3c81
|
2023-08-18 13:07:35 +08:00 |
|
hiyouga
|
68c30094e1
|
support ppo score norm (trl 0.5.1.dev required)
Former-commit-id: 2b25db6d260ec1532281a592e873579346c7d21c
|
2023-08-18 12:02:42 +08:00 |
|
hiyouga
|
a8dd39ad08
|
fix PPO trainer #551 , update readme
Former-commit-id: faead74849470cebae9e37cde5fab2a71b32aa43
|
2023-08-18 11:43:10 +08:00 |
|
hiyouga
|
c0e887c85a
|
update training resuming
Former-commit-id: 2ec75c31f609e65116ac3b621eeb7d8ccbf69135
|
2023-08-18 01:41:17 +08:00 |
|
hoshi-hiyouga
|
c856724d3f
|
Merge branch 'main' into main
Former-commit-id: 870d2c7bf74d0da5a927bef4b8b01d15cc66a3e9
|
2023-08-18 01:37:23 +08:00 |
|
hiyouga
|
ff5ebb6a86
|
support bf16 ppo #551
Former-commit-id: 092088967de7409a2d51847cfc7afc83a8887320
|
2023-08-18 00:40:32 +08:00 |
|
hiyouga
|
a089d2665a
|
fix ChatGLM2 ppo #527 #528
Former-commit-id: 60d6ad64d7c9f6445b0df8de0153c3a311974198
|
2023-08-18 00:34:59 +08:00 |
|
hiyouga
|
bb7028f7e2
|
fix generation bug #532
Former-commit-id: c071121e67374e5f09798db57cfc8668617a36ae
|
2023-08-17 22:21:34 +08:00 |
|
hiyouga
|
4319b5172c
|
fix streaming in pt stage #548 #549
Former-commit-id: 050e992bee2a9293cc7399b578de807b5bf9bddc
|
2023-08-17 17:59:26 +08:00 |
|
hiyouga
|
22dec02b5f
|
fix generation
Former-commit-id: 66a0300d312ef91c24fcf80667fa3b0bb8e1a342
|
2023-08-16 22:39:54 +08:00 |
|
hiyouga
|
1449d3182c
|
fix system prompt
Former-commit-id: 411e775aa939bdd154a3f1e92921ede90d989f18
|
2023-08-16 01:35:52 +08:00 |
|
hiyouga
|
0684a5adc8
|
fix ChatGLM RLHF
Former-commit-id: 4e43e887e432ceb7e9287b4e309b63af3c3ba1bf
|
2023-08-15 11:19:20 +08:00 |
|
hiyouga
|
4d6534a665
|
alert pad_token source
Former-commit-id: f26a84e0d927d2554890daf431a93652e18f4235
|
2023-08-15 00:07:56 +08:00 |
|
hiyouga
|
1ce4a6109f
|
fix ChatGLM lm_head #494
Former-commit-id: bf0048abdaeb2b9592d38ac991704ad014370b47
|
2023-08-14 14:14:48 +08:00 |
|
hiyouga
|
fff747ee5a
|
web UI integrating RLHF
Former-commit-id: 137fd146b90f89a1164b56e6d507b30b1f5c2437
|
2023-08-14 10:48:47 +08:00 |
|
hiyouga
|
0366115d29
|
fix #480
Former-commit-id: ec15ca8fffacba2c34e1849c5ce90ca9989d66a2
|
2023-08-14 00:23:56 +08:00 |
|
hiyouga
|
03537dda06
|
tiny fix
Former-commit-id: 50a34c043de6d9e1410291e1d8c1ea9d53754e9e
|
2023-08-12 22:02:43 +08:00 |
|
hiyouga
|
4bbbee4370
|
fix rope scaling
Former-commit-id: 2e0dd36700ec5e8294581c1db4b9431f755fc5f8
|
2023-08-12 22:00:01 +08:00 |
|
hiyouga
|
56d6448c4a
|
update readme
Former-commit-id: 6fa381400c21fa249cebcdff8c3afd72f8de20b3
|
2023-08-12 21:00:11 +08:00 |
|
hiyouga
|
a15a602809
|
support rope scaling, fix #475 #476 #478
Former-commit-id: 337d5f68b72230e545e7a94ca789187c7a2b7187
|
2023-08-12 20:46:27 +08:00 |
|
hiyouga
|
9644286027
|
fix unusual output of 8bit models #278 #391
Former-commit-id: 337ce5272b81f5561162beb08814b0e5abf23703
|
2023-08-12 00:25:29 +08:00 |
|
hiyouga
|
98aa629843
|
Release v0.1.6
Former-commit-id: 43c8b3c3c8bfb2e32d17fb3e8b194938e37d54bd
|
2023-08-11 23:25:57 +08:00 |
|
hiyouga
|
7ada4f5f6f
|
support DPO training (2305.18290)
Former-commit-id: 6d98de148e4af63a7028dfaeb6cf86eb56a4488f
|
2023-08-11 03:02:53 +08:00 |
|
hiyouga
|
6eb4120464
|
support val set in streaming mode
Former-commit-id: faed15b58ed00b1e09bb091e7eee48f5ef7c508b
|
2023-08-09 23:00:26 +08:00 |
|
niuba
|
186d7f2d67
|
add last_checkpoint support
Former-commit-id: 9f1977e4de00b14a9d1b555c25bcaf12998d5046
|
2023-08-09 16:39:27 +08:00 |
|
hiyouga
|
3a821e5f0c
|
fix sft trainer
Former-commit-id: 08cc888b1569572d0cd20bcf3f07e20072a0311a
|
2023-08-09 16:35:03 +08:00 |
|
hiyouga
|
b6316b460f
|
fix tokenizer #417
Former-commit-id: 01aa678311bfd213a4b410a4e0ff09f48a0d40a1
|
2023-08-08 23:59:41 +08:00 |
|
hiyouga
|
5ada0ee44c
|
fix bug
Former-commit-id: c13ce66021b21e015871b84489eeafa127a424a4
|
2023-08-08 17:55:55 +08:00 |
|
hiyouga
|
ab66cc8e37
|
fix chatml template #408
Former-commit-id: 21e0cc3f44c35ae689b00b274391492f413725ac
|
2023-08-08 17:44:39 +08:00 |
|
hiyouga
|
1ede9595c1
|
fix #376
Former-commit-id: a5b01257ba3323bcb2dd0217fb89a387e39ddbec
|
2023-08-07 13:58:59 +08:00 |
|
hiyouga
|
1d0ccf63f2
|
update trainer
Former-commit-id: 0d39b53a5164e34d22fe0a492eaa0d7ac63102fe
|
2023-08-07 13:34:35 +08:00 |
|
hiyouga
|
577b3cd0de
|
fix qwen eos token
Former-commit-id: 770830c67886f5872b39b9608949ec62d4616b27
|
2023-08-06 13:31:17 +08:00 |
|
hiyouga
|
2d153e35ef
|
fix mtloader
Former-commit-id: ca48c2c02c3cfa9afb99971b50daeda9cf14e7cb
|
2023-08-03 19:29:02 +08:00 |
|
hiyouga
|
7358098e32
|
fix qwen inference
Former-commit-id: 823f0de0ca0a92b6f48a90e5ffe57a48dc018f1d
|
2023-08-03 16:31:55 +08:00 |
|
hiyouga
|
bf15c4a03c
|
support Qwen-7B, fix InternLM-7B inference
Former-commit-id: 25d2ca29ecb70cbfd5206333c667042a0c4d2e5a
|
2023-08-03 15:53:32 +08:00 |
|
hiyouga
|
4c37e03ca6
|
fix webui
Former-commit-id: e87630ef77977b2879f1199b9a421acbbbb32a51
|
2023-08-03 12:43:12 +08:00 |
|
hiyouga
|
18f73169fd
|
modify code structure
Former-commit-id: 6369f9b1751e6f9bb709ba76a85f69cbe0823e5d
|
2023-08-02 23:17:36 +08:00 |
|
hiyouga
|
f5696a08c1
|
fix PPO trainer
Former-commit-id: 21982a7d4dd9b7c3a1145b481f02b9990e32dc00
|
2023-08-02 19:10:23 +08:00 |
|
hiyouga
|
2fdef55143
|
update ppo trainer
Former-commit-id: c27136a83e167465d3f825e40f10c7b9fcfbf97a
|
2023-08-02 18:46:41 +08:00 |
|
hiyouga
|
e6ba6f6d61
|
fix memory leak of PPO trainer
Former-commit-id: 38410894a5ebf0b043b55a6bd5cca3cd0a44b27d
|
2023-08-02 17:41:34 +08:00 |
|
hiyouga
|
ab3f685330
|
fix RM save model
Former-commit-id: 8104cc2425431eb1cddccf3909855296116f922b
|
2023-08-01 11:56:17 +08:00 |
|
hiyouga
|
5faad2a64c
|
release v0.1.4
Former-commit-id: 81f84aaf2e120e39edb28ef42893939fc9a184e2
|
2023-08-01 10:08:47 +08:00 |
|
hiyouga
|
44c31d6064
|
fix inference
Former-commit-id: 55dc2bdd3eaa552c655e584fc3cbbf017c7bc3e7
|
2023-08-01 00:06:48 +08:00 |
|
hiyouga
|
1c2422df31
|
fix arg check
Former-commit-id: 2c5c73de9ebc88e2d04e80754781c94a571133a0
|
2023-07-31 23:48:57 +08:00 |
|
hiyouga
|
5cab5c1b36
|
update readme
Former-commit-id: d99cda254e5025ff3f968d256197ab031bfabef1
|
2023-07-31 23:42:32 +08:00 |
|
hiyouga
|
63123a9098
|
support streaming data, fix #284 #274 #268
Former-commit-id: 819cc1353599e5fa45658bc56dd0dbe4b258b197
|
2023-07-31 23:33:00 +08:00 |
|
hiyouga
|
913284c29f
|
fix #242
Former-commit-id: 80a346e29beb49e8935b786e2af1059fdc4954b2
|
2023-07-25 17:04:02 +08:00 |
|
hiyouga
|
36d04e248e
|
release v0.1.3
Former-commit-id: 62c68bcbf591516e8f90b47810bea6f710fd23f6
|
2023-07-21 16:48:34 +08:00 |
|
hiyouga
|
3855557f8c
|
fix save function
Former-commit-id: 1d6beb0c8490a7531ffdf7a2819410597b200d12
|
2023-07-21 14:09:07 +08:00 |
|
hiyouga
|
1f4a65afea
|
update web UI, support rm predict #210
Former-commit-id: 92cc6b655dc91b94d5bf9d8618c3b57d5cf94333
|
2023-07-21 13:27:27 +08:00 |
|