Yaowei Zheng
|
80fe3a172d
|
[model] add dots ocr (#9176)
|
2025-09-21 23:34:19 +08:00 |
|
Kingsley
|
1c675522fd
|
[data] Fix glm4v mmplugin when not expand (#9029)
|
2025-08-27 11:45:24 +08:00 |
|
Kingsley
|
1639e4b587
|
[data] fix gemma3n mmplugin (#8627)
|
2025-07-14 13:32:57 +08:00 |
|
Kingsley
|
542fa97a72
|
[data] support glm4.1v video training (#8571)
|
2025-07-08 16:29:44 +08:00 |
|
Yaowei Zheng
|
4b0ec83928
|
[deps] bump transformers to 4.49.0 (#8564)
|
2025-07-07 20:31:50 +08:00 |
|
Kingsley
|
e9f70daabe
|
[model] add gemma3n (#8509)
|
2025-07-01 22:37:24 +08:00 |
|
Kingsley
|
d17a672251
|
[model] add GLM-4.1V (#8462)
|
2025-06-30 01:09:41 +08:00 |
|
Yaowei Zheng
|
48897e5b16
|
[data] fix audio reader (#8448)
|
2025-06-24 20:53:20 +08:00 |
|
Kingsley
|
31bca4d172
|
[model] support Mistral3.1 small 2503 (#8335)
|
2025-06-09 10:37:42 +08:00 |
|
Kingsley
|
c224d17cb2
|
[data] support nested images input for videos (#8264)
|
2025-06-03 20:26:29 +08:00 |
|
Kingsley
|
f08b748199
|
[data] fix internvl plugin when using PIL images (#8129)
|
2025-05-22 01:32:59 +08:00 |
|
hoshi-hiyouga
|
9ae17cd173
|
[deps] update to transformers 4.52 (#8125)
|
2025-05-21 05:16:18 +08:00 |
|
hoshi-hiyouga
|
beae231af6
|
[doc] add no build isolation (#8103)
|
2025-05-19 19:25:13 +08:00 |
|
Kingsley
|
52b23f9e56
|
[data] add forward compatibility for video_utils in Transformers 4.52.0 (#8077)
|
2025-05-16 17:41:04 +08:00 |
|
hoshi-hiyouga
|
68fc068cab
|
[data] fix kimi vl template (#8015)
|
2025-05-11 20:45:19 +08:00 |
|
hoshi-hiyouga
|
c566e39b7d
|
[data] fix base plugin (#7924)
|
2025-04-30 16:28:05 +08:00 |
|
hoshi-hiyouga
|
11295cdea0
|
[data] improve mm plugin (#7910)
|
2025-04-30 06:34:28 +08:00 |
|
Kingsley
|
db9559456c
|
[data] fix qwen2.5 omni template (#7883)
|
2025-04-29 00:58:23 +08:00 |
|
hoshi-hiyouga
|
3ae5da2a04
|
[model] fix dsv3 leaf node (#7879)
|
2025-04-28 18:11:09 +08:00 |
|
hoshi-hiyouga
|
d173cb50f5
|
[data] fix qwen2 omni plugin (#7875)
|
2025-04-28 14:22:41 +08:00 |
|
hoshi-hiyouga
|
bb5b83352b
|
[data] fix minicpmo vllm infer (#7870)
|
2025-04-28 01:59:53 +08:00 |
|
Kingsley
|
fa0eb91f1f
|
[data] fix internvl plugin (#7817)
|
2025-04-23 00:58:22 +08:00 |
|
Kingsley
|
7500e761d3
|
[misc] update internvl constants (#7801)
|
2025-04-22 15:53:08 +08:00 |
|
hoshi-hiyouga
|
0e4ce039ee
|
[data] improve mmplugin (#7795)
|
2025-04-22 01:25:33 +08:00 |
|
hoshi-hiyouga
|
39169986ef
|
[trainer] fix pt loss (#7748)
* fix pt loss
* robust
* fix
* test
|
2025-04-17 03:15:35 +08:00 |
|
hoshi-hiyouga
|
86ebb219d6
|
[breaking] bump transformers to 4.45.0 & improve ci (#7746)
* update ci
* fix
* fix
* fix
* fix
* fix
|
2025-04-17 02:36:48 +08:00 |
|
Kingsley
|
2e518f255f
|
[model] support intern-VL 2.5-3 series (#7258)
* add internvl and rebase
* fix for internvl2&3
* remove lines
* fix video_inputs & lint
* nit
* add constants
* remove lines
* fix
* fix error
* pass ci
* pass ci
* skip internvl & nit
|
2025-04-17 00:31:30 +08:00 |
|
Kingsley
|
2101399c94
|
[model] Support Kimi_VL thinking/instruct (#7719)
* add kimi_vl
* patch config
* check version
* Update mm_plugin.py
* Update mm_plugin.py
---------
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-15 00:21:58 +08:00 |
|
Kingsley
|
349c56c51c
|
[data] Fix bugs of use_audio_in_video in Qwen2.5 Omni (#7638)
* cache _mm_inputs
* nit
* support for use_audio_in_video
* remove cache
* fix data
* Update mllm_video_audio_demo.json
|
2025-04-08 18:40:10 +08:00 |
|
hoshi-hiyouga
|
c3c0efbaa0
|
[misc] fix packing and eval plot (#7623)
|
2025-04-07 18:20:57 +08:00 |
|
hoshi-hiyouga
|
831e7f1cfd
|
[model] add llama4 (#7611)
|
2025-04-06 13:42:31 +08:00 |
|
Kingsley
|
d4cfa9507e
|
[data] fix qwen2.5 omni plugin (#7578)
* specific entry
* Update mm_plugin.py
* fix fps cal
---------
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-02 23:58:39 +08:00 |
|
Kingsley
|
d32c6c014d
|
[data] fix qwen2.5 omni plugin (#7573)
* align key with qwen2vl
* nit && change scripts
|
2025-04-02 21:28:52 +08:00 |
|
hoshi-hiyouga
|
5e22597ff1
|
[infer] vllm video/audio inference (#7566)
|
2025-04-02 02:27:04 +08:00 |
|
hoshi-hiyouga
|
2bfcad2394
|
[model] fix kv cache (#7564)
|
2025-04-01 23:07:46 +08:00 |
|
Kingsley
|
7eed496336
|
[model] add Qwen2.5-Omni model (#7537)
* preserve image_sizes
* preserve image_sizes
* init plugin
* support audio-text2text lora
* nit
* support image/video-text2text, audio-text2text
* remove args
* remove lines
* add docs && nit
* remove some comments
* fix && add merge part script
* add license
|
2025-03-31 20:39:35 +08:00 |
|
Kingsley
|
8da1d2fa71
|
[data] fix pixtral plugin (#7505)
* preserve `image_sizes`
* add comments
|
2025-03-27 17:06:40 +08:00 |
|
hoshi-hiyouga
|
7203365b80
|
[trainer] fix vlm loss for transformers 4.49 (#7448)
|
2025-03-24 10:24:05 +08:00 |
|
hoshi-hiyouga
|
93e6184cbe
|
[data] gemma3 plugin pan and scan (#7294)
* gemma3 pan and scan
* add test case
* fix test
|
2025-03-13 23:29:23 +08:00 |
|
hoshi-hiyouga
|
650a9a9057
|
[misc] update format (#7277)
|
2025-03-13 02:53:08 +08:00 |
|
hoshi-hiyouga
|
4b9d8da5a4
|
[model] support gemma3 (#7273)
|
2025-03-13 01:35:23 +08:00 |
|
hoshi-hiyouga
|
264538cb26
|
[misc] upgrade format to py39 (#7256)
|
2025-03-12 00:08:41 +08:00 |
|
hoshi-hiyouga
|
bb8aba5abf
|
[data] fix mm template (#7181)
Former-commit-id: 648616d473c81d393592806307e3e25b159cb278
|
2025-03-06 15:18:32 +08:00 |
|
hoshi-hiyouga
|
7b985f55db
|
[trainer] update config (#7174)
Former-commit-id: 9f535d0e3c4ee3cd0f1b65218c2eee5d03f43c6f
|
2025-03-05 23:32:54 +08:00 |
|
sirui.li
|
fd0357a26d
|
[data] fix qwen2audio plugin (#7166)
* Update pairwise.py
[data]Repair multimodal model dpo training
* Update pairwise.py
[data]repair multimodal model dpo training using deepcopy
* Update pairwise.py
* Update mm_plugin.py
Former-commit-id: 86763dfdb8e9e5668c1ddd7e924e4be76bf78368
|
2025-03-05 18:03:36 +08:00 |
|
hoshi-hiyouga
|
31f9daa362
|
[data] use bicubic resampler (#7143)
Former-commit-id: c708f19ab0ab57526134952afddaa90aae8decbf
|
2025-03-04 00:17:06 +08:00 |
|
hoshi-hiyouga
|
065f7fb5da
|
[data] fix mllama (#7053)
* fix mllama
* fix test
Former-commit-id: f5af20a63f3d59a6a68d323a7c6f68e551edb3a3
|
2025-02-24 22:05:38 +08:00 |
|
Zhangchi Feng
|
fcf75633a0
|
[data] fix MiniCPMV plugin (#6998)
* fix template
* fix bug in messages processing
Former-commit-id: f98b828f53968fb9c72bff9e45510ad5586c4fab
|
2025-02-19 19:36:04 +08:00 |
|
hoshi-hiyouga
|
1d675a287d
|
[version] support transformers 449 (#6982)
* support transformers 449
* fix mm plugin
Former-commit-id: e9118a9df0839d24f6ddff5a0b55ef101a1d3d22
|
2025-02-18 17:05:40 +08:00 |
|
hoshi-hiyouga
|
be33ef67fb
|
[misc] fix script (#6977)
Former-commit-id: 775efa1d8cbdb1b7d122be2a986d47f85214e0a1
|
2025-02-18 17:00:46 +08:00 |
|