Kingsley
|
fef86fa7fe
|
[data] fix qwen3omni audio length calculation (#9467)
|
2025-11-12 10:37:15 +08:00 |
|
Yaowei Zheng
|
3ae15da9c0
|
[misc] lint code (#9395)
|
2025-11-03 22:08:59 +08:00 |
|
魅影
|
215580c77d
|
[data] fix mm pluigin for qwen omni video training (#9388)
Co-authored-by: frozenleaves <frozen@Mac.local>
|
2025-11-03 11:44:27 +08:00 |
|
Xiaosu Zhu
|
129e918106
|
[data] Fix Qwen3VL plugin (#9297)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Yaowei Zheng <hiyouga@buaa.edu.cn>
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-10-26 16:07:04 +08:00 |
|
Yaowei Zheng
|
6ffebe5ff7
|
[data] fix qwen omni plugin (#9204)
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-09-28 01:02:29 +08:00 |
|
xvxuopop
|
0761a4448f
|
[model] add qwen3-vl/qwen3-omni (#9196)
Co-authored-by: kingsley <kingsleydodonow@gmail.com>
|
2025-09-27 01:21:47 +08:00 |
|
Yaowei Zheng
|
80fe3a172d
|
[model] add dots ocr (#9176)
|
2025-09-21 23:34:19 +08:00 |
|
Kingsley
|
1c675522fd
|
[data] Fix glm4v mmplugin when not expand (#9029)
|
2025-08-27 11:45:24 +08:00 |
|
Kingsley
|
1639e4b587
|
[data] fix gemma3n mmplugin (#8627)
|
2025-07-14 13:32:57 +08:00 |
|
Kingsley
|
542fa97a72
|
[data] support glm4.1v video training (#8571)
|
2025-07-08 16:29:44 +08:00 |
|
Yaowei Zheng
|
4b0ec83928
|
[deps] bump transformers to 4.49.0 (#8564)
|
2025-07-07 20:31:50 +08:00 |
|
Kingsley
|
e9f70daabe
|
[model] add gemma3n (#8509)
|
2025-07-01 22:37:24 +08:00 |
|
Kingsley
|
d17a672251
|
[model] add GLM-4.1V (#8462)
|
2025-06-30 01:09:41 +08:00 |
|
Yaowei Zheng
|
48897e5b16
|
[data] fix audio reader (#8448)
|
2025-06-24 20:53:20 +08:00 |
|
Kingsley
|
31bca4d172
|
[model] support Mistral3.1 small 2503 (#8335)
|
2025-06-09 10:37:42 +08:00 |
|
Kingsley
|
c224d17cb2
|
[data] support nested images input for videos (#8264)
|
2025-06-03 20:26:29 +08:00 |
|
Kingsley
|
f08b748199
|
[data] fix internvl plugin when using PIL images (#8129)
|
2025-05-22 01:32:59 +08:00 |
|
hoshi-hiyouga
|
9ae17cd173
|
[deps] update to transformers 4.52 (#8125)
|
2025-05-21 05:16:18 +08:00 |
|
hoshi-hiyouga
|
beae231af6
|
[doc] add no build isolation (#8103)
|
2025-05-19 19:25:13 +08:00 |
|
Kingsley
|
52b23f9e56
|
[data] add forward compatibility for video_utils in Transformers 4.52.0 (#8077)
|
2025-05-16 17:41:04 +08:00 |
|
hoshi-hiyouga
|
68fc068cab
|
[data] fix kimi vl template (#8015)
|
2025-05-11 20:45:19 +08:00 |
|
hoshi-hiyouga
|
c566e39b7d
|
[data] fix base plugin (#7924)
|
2025-04-30 16:28:05 +08:00 |
|
hoshi-hiyouga
|
11295cdea0
|
[data] improve mm plugin (#7910)
|
2025-04-30 06:34:28 +08:00 |
|
Kingsley
|
db9559456c
|
[data] fix qwen2.5 omni template (#7883)
|
2025-04-29 00:58:23 +08:00 |
|
hoshi-hiyouga
|
3ae5da2a04
|
[model] fix dsv3 leaf node (#7879)
|
2025-04-28 18:11:09 +08:00 |
|
hoshi-hiyouga
|
d173cb50f5
|
[data] fix qwen2 omni plugin (#7875)
|
2025-04-28 14:22:41 +08:00 |
|
hoshi-hiyouga
|
bb5b83352b
|
[data] fix minicpmo vllm infer (#7870)
|
2025-04-28 01:59:53 +08:00 |
|
Kingsley
|
fa0eb91f1f
|
[data] fix internvl plugin (#7817)
|
2025-04-23 00:58:22 +08:00 |
|
Kingsley
|
7500e761d3
|
[misc] update internvl constants (#7801)
|
2025-04-22 15:53:08 +08:00 |
|
hoshi-hiyouga
|
0e4ce039ee
|
[data] improve mmplugin (#7795)
|
2025-04-22 01:25:33 +08:00 |
|
hoshi-hiyouga
|
39169986ef
|
[trainer] fix pt loss (#7748)
* fix pt loss
* robust
* fix
* test
|
2025-04-17 03:15:35 +08:00 |
|
hoshi-hiyouga
|
86ebb219d6
|
[breaking] bump transformers to 4.45.0 & improve ci (#7746)
* update ci
* fix
* fix
* fix
* fix
* fix
|
2025-04-17 02:36:48 +08:00 |
|
Kingsley
|
2e518f255f
|
[model] support intern-VL 2.5-3 series (#7258)
* add internvl and rebase
* fix for internvl2&3
* remove lines
* fix video_inputs & lint
* nit
* add constants
* remove lines
* fix
* fix error
* pass ci
* pass ci
* skip internvl & nit
|
2025-04-17 00:31:30 +08:00 |
|
Kingsley
|
2101399c94
|
[model] Support Kimi_VL thinking/instruct (#7719)
* add kimi_vl
* patch config
* check version
* Update mm_plugin.py
* Update mm_plugin.py
---------
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-15 00:21:58 +08:00 |
|
Kingsley
|
349c56c51c
|
[data] Fix bugs of use_audio_in_video in Qwen2.5 Omni (#7638)
* cache _mm_inputs
* nit
* support for use_audio_in_video
* remove cache
* fix data
* Update mllm_video_audio_demo.json
|
2025-04-08 18:40:10 +08:00 |
|
hoshi-hiyouga
|
c3c0efbaa0
|
[misc] fix packing and eval plot (#7623)
|
2025-04-07 18:20:57 +08:00 |
|
hoshi-hiyouga
|
831e7f1cfd
|
[model] add llama4 (#7611)
|
2025-04-06 13:42:31 +08:00 |
|
Kingsley
|
d4cfa9507e
|
[data] fix qwen2.5 omni plugin (#7578)
* specific entry
* Update mm_plugin.py
* fix fps cal
---------
Co-authored-by: hoshi-hiyouga <hiyouga@buaa.edu.cn>
|
2025-04-02 23:58:39 +08:00 |
|
Kingsley
|
d32c6c014d
|
[data] fix qwen2.5 omni plugin (#7573)
* align key with qwen2vl
* nit && change scripts
|
2025-04-02 21:28:52 +08:00 |
|
hoshi-hiyouga
|
5e22597ff1
|
[infer] vllm video/audio inference (#7566)
|
2025-04-02 02:27:04 +08:00 |
|
hoshi-hiyouga
|
2bfcad2394
|
[model] fix kv cache (#7564)
|
2025-04-01 23:07:46 +08:00 |
|
Kingsley
|
7eed496336
|
[model] add Qwen2.5-Omni model (#7537)
* preserve image_sizes
* preserve image_sizes
* init plugin
* support audio-text2text lora
* nit
* support image/video-text2text, audio-text2text
* remove args
* remove lines
* add docs && nit
* remove some comments
* fix && add merge part script
* add license
|
2025-03-31 20:39:35 +08:00 |
|
Kingsley
|
8da1d2fa71
|
[data] fix pixtral plugin (#7505)
* preserve `image_sizes`
* add comments
|
2025-03-27 17:06:40 +08:00 |
|
hoshi-hiyouga
|
7203365b80
|
[trainer] fix vlm loss for transformers 4.49 (#7448)
|
2025-03-24 10:24:05 +08:00 |
|
hoshi-hiyouga
|
93e6184cbe
|
[data] gemma3 plugin pan and scan (#7294)
* gemma3 pan and scan
* add test case
* fix test
|
2025-03-13 23:29:23 +08:00 |
|
hoshi-hiyouga
|
650a9a9057
|
[misc] update format (#7277)
|
2025-03-13 02:53:08 +08:00 |
|
hoshi-hiyouga
|
4b9d8da5a4
|
[model] support gemma3 (#7273)
|
2025-03-13 01:35:23 +08:00 |
|
hoshi-hiyouga
|
264538cb26
|
[misc] upgrade format to py39 (#7256)
|
2025-03-12 00:08:41 +08:00 |
|
hoshi-hiyouga
|
bb8aba5abf
|
[data] fix mm template (#7181)
Former-commit-id: 648616d473c81d393592806307e3e25b159cb278
|
2025-03-06 15:18:32 +08:00 |
|
hoshi-hiyouga
|
7b985f55db
|
[trainer] update config (#7174)
Former-commit-id: 9f535d0e3c4ee3cd0f1b65218c2eee5d03f43c6f
|
2025-03-05 23:32:54 +08:00 |
|