Yaowei Zheng
|
0894b4f37e
|
[misc] lint (#9636)
|
2025-12-20 16:19:39 +08:00 |
|
mrhaoxx
|
964569751f
|
[kt] refactor ktransformers integration (#9632)
|
2025-12-18 21:26:04 +08:00 |
|
tangefly
|
4fd94141a4
|
[model] Add Ministral3 (#9582)
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-12-10 15:57:24 +08:00 |
|
Yaowei Zheng
|
5d56817e2b
|
[misc] lint (#9593)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-09 18:00:35 +08:00 |
|
Hertz
|
c1f5f8fff6
|
[model] support GLM4.6v (#9586)
|
2025-12-09 11:06:42 +08:00 |
|
tangefly
|
739954910a
|
[deps] Update for Transformers v5 (#9569)
|
2025-12-08 01:13:32 +08:00 |
|
Ben Feuer
|
1c44b60e3e
|
[feat] fp8 training (#8960)
Co-authored-by: Benjamin Feuer <penfever@gmail.com>
Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>
|
2025-10-01 14:32:53 +08:00 |
|
Yaowei Zheng
|
52488ac974
|
[deps] upgrade transformers to 4.56.1 (#9128)
|
2025-09-14 02:26:39 +08:00 |
|
XLXW
|
1ada15981a
|
[feature] add support for dft loss (#8917)
|
2025-08-15 23:29:57 +08:00 |
|
hoshi-hiyouga
|
9ae17cd173
|
[deps] update to transformers 4.52 (#8125)
|
2025-05-21 05:16:18 +08:00 |
|
hoshi-hiyouga
|
416853dd25
|
[parser] support omegaconf (#7793)
|
2025-04-21 23:30:30 +08:00 |
|
hoshi-hiyouga
|
39169986ef
|
[trainer] fix pt loss (#7748)
* fix pt loss
* robust
* fix
* test
|
2025-04-17 03:15:35 +08:00 |
|
hoshi-hiyouga
|
86ebb219d6
|
[breaking] bump transformers to 4.45.0 & improve ci (#7746)
* update ci
* fix
* fix
* fix
* fix
* fix
|
2025-04-17 02:36:48 +08:00 |
|
Shawn Tao
|
acb09fa3a3
|
[trainer] fix key error (#7635)
|
2025-04-08 18:39:50 +08:00 |
|
hoshi-hiyouga
|
c3c0efbaa0
|
[misc] fix packing and eval plot (#7623)
|
2025-04-07 18:20:57 +08:00 |
|
hoshi-hiyouga
|
7203365b80
|
[trainer] fix vlm loss for transformers 4.49 (#7448)
|
2025-03-24 10:24:05 +08:00 |
|
hoshi-hiyouga
|
650a9a9057
|
[misc] update format (#7277)
|
2025-03-13 02:53:08 +08:00 |
|
hoshi-hiyouga
|
264538cb26
|
[misc] upgrade format to py39 (#7256)
|
2025-03-12 00:08:41 +08:00 |
|
Billy Cao
|
58e9ca8aa0
|
[trainer] fix gen_kwarg to eval during training (#5451)
* Correctly pass gen_kwarg to eval during model runs
* fix
* fix
---------
Co-authored-by: hiyouga <hiyouga@buaa.edu.cn>
Former-commit-id: 845d16122496311e08263610a6a922f82604de7b
|
2025-02-13 02:35:06 +08:00 |
|
hoshi-hiyouga
|
c2022431aa
|
[misc] update license year & fix llama pro (#6814)
* fix llamapro script
* change year
Former-commit-id: d9ae594178796994d400a5f207d6499712816f89
|
2025-02-05 01:53:33 +08:00 |
|
hoshi-hiyouga
|
222423bcef
|
[breaking] support transformers 4.48 (#6628)
Former-commit-id: f154ab175c513a4d7bb866bf2cffc34b77b50508
|
2025-01-31 01:36:33 +08:00 |
|
hoshi-hiyouga
|
2a05941b14
|
[inference] fix stop token for object detection (#6624)
* fix stop token
* update minicpm data pipeline
* fix npu qlora examples
Former-commit-id: 844919fadaa8a61dfae47020971ea80730b2346f
|
2025-01-13 21:34:20 +08:00 |
|
hiyouga
|
647c51a772
|
imporve log
Former-commit-id: a6abf375975ffea3d51e1b944c9855b5f62ffac8
|
2025-01-08 09:56:10 +00:00 |
|
hiyouga
|
2aaf3697d7
|
fix #6499
Former-commit-id: dffc607220ff6dac15cf501ac9a3cdbe80c25211
|
2025-01-02 11:28:54 +00:00 |
|
hiyouga
|
88b1874c04
|
fix #6448
Former-commit-id: 04f78e85af5af14b4c195936623e426a6a128af2
|
2024-12-27 16:54:39 +00:00 |
|
hiyouga
|
a897d46049
|
support report custom args
Former-commit-id: d41254c40a1c5cacf9377096adb27efa9bdb79ea
|
2024-12-21 21:42:45 +00:00 |
|
hoshi-hiyouga
|
0a869c4ed4
|
Merge pull request #6401 from Zeyi-Lin/hiyouga/swanlab
feat: add swanlab for experiment tracking and visualization.
Former-commit-id: e65fe507f7643bf40b0fc462805c7b7f8ef6b738
|
2024-12-21 14:09:33 +08:00 |
|
hiyouga
|
0385c60177
|
fix #6391
Former-commit-id: 067ba6e6cb4d8a1d95bba0a108f73008416a2865
|
2024-12-19 12:16:38 +00:00 |
|
hiyouga
|
01eeae50b5
|
support disable shuffling
Former-commit-id: 9d8c35fd6b838ede0bd6827c6c6121f2cba2b11b
|
2024-12-19 08:53:21 +00:00 |
|
hiyouga
|
7eeeffdb8a
|
add swanlab
Former-commit-id: c85a77c8a8824a56a67d56b97b4877fcd6edeb3d
|
2024-12-19 07:12:31 +00:00 |
|
hiyouga
|
19ebc0e7a2
|
support control eos, fix #6345
Former-commit-id: cb0f8399356bf372f3b7963f2565c3d504be0923
|
2024-12-17 10:42:05 +00:00 |
|
hiyouga
|
aacd9642f5
|
fix #6348
Former-commit-id: 83e552320909f4775377889f1512994b7e638a7e
|
2024-12-17 10:06:46 +00:00 |
|
hiyouga
|
fb22651faf
|
fix mrope
Former-commit-id: 55bee1d333549ca19858b3f5c1b7b86926e5fb09
|
2024-12-12 15:08:17 +00:00 |
|
hiyouga
|
c1768cfb14
|
support batch infer in vllm
Former-commit-id: 3ef5ed3b9a44eed2f7e3ff221dfc343d0a97c0b5
|
2024-12-04 13:50:00 +00:00 |
|
Ting
|
87b1f851f1
|
code refactor
Former-commit-id: ee3f85aa9677d0aeecb3bc396530d2cd7c50dce5
|
2024-11-19 20:33:18 +08:00 |
|
Ting
|
a20c2b6ecf
|
update
Former-commit-id: a3e8ca53e654136242197a2da872cc0e5cf67880
|
2024-11-19 19:10:07 +08:00 |
|
Ting
|
fee94e1c54
|
support efficient tokens calculation on sft/dpo
Former-commit-id: b157d5cccdeb42412b8b440d25d5bdfa8a50be68
|
2024-11-19 17:15:47 +08:00 |
|
hiyouga
|
2bb3255e74
|
fix dpo metrics
Former-commit-id: 57029280da825a39fbf5a05097921b861f126669
|
2024-11-02 20:59:01 +08:00 |
|
hiyouga
|
093eda2ad6
|
support rank0 logger
Former-commit-id: 84528eabe560091bfd866b6a0ca864085af7529b
|
2024-11-02 18:31:04 +08:00 |
|
hiyouga
|
8185eb1890
|
fix incorrect loss value for vlms
Former-commit-id: 0aa29a71ce958343a2086090d647eb63b8f5f5be
|
2024-10-30 08:56:46 +00:00 |
|
hiyouga
|
7f71276ad8
|
add docstrings, refactor logger
Former-commit-id: c34e489d71f8f539028543ccf8ee92cecedd6276
|
2024-09-08 00:56:56 +08:00 |
|
hiyouga
|
af178cbcd1
|
update get template
Former-commit-id: 21ea0d0786f91c0bce79630963e66b815a6792a0
|
2024-09-04 22:36:20 +08:00 |
|
hiyouga
|
7056087e92
|
lazy image load
Former-commit-id: cdd733b575411e003bc5ffd6560dd8eff8aa09cf
|
2024-09-04 02:27:08 +08:00 |
|
hiyouga
|
2e1396cd6b
|
lint
Former-commit-id: d821d933e6cb982d648a69f85f6ad01d0560ed70
|
2024-09-03 00:46:25 +08:00 |
|
hiyouga
|
b5e9df5df8
|
fix #5324
Former-commit-id: f7aa06c9c0b18c28419ea5792410915d3f322cbf
|
2024-09-02 23:56:21 +08:00 |
|
hoshi-hiyouga
|
7367c6ec21
|
fix trainer predict
Former-commit-id: 2790790cd26c6743105555a60523b89f367ebce3
|
2024-09-02 10:15:29 +08:00 |
|
hoshi-hiyouga
|
6579ec8c4c
|
remove .cpu()
Former-commit-id: 35c57cc9dcba305d40282a9757ddc23968c210ac
|
2024-09-02 10:10:53 +08:00 |
|
hiyouga
|
2f6fc27c8b
|
remove visual_inputs, fix qlora
Former-commit-id: be30c01c4f1482520ece770bd54c6a4837c26f0a
|
2024-08-31 00:24:51 +08:00 |
|
hiyouga
|
d789b667d7
|
optimize predict vram
Former-commit-id: a577e44eee351b3ed8011a33ae01cd713354ff97
|
2024-08-30 23:08:45 +08:00 |
|
moontidef
|
33a90b9026
|
fix: rename optimzer to optimizer
Former-commit-id: 186dc1fde822e6a603ac273538741ea3853f243e
|
2024-08-07 10:05:01 +08:00 |
|