| Yu, Jiazuo, Zhuge, Yunzhi, Zhang, Lu, Huang, Zichen, Zhou, Wei, Wang, Dong, Lu, Huchuan, He, You and Chen, Long 2026. One aligned LLM to serve them all: A transfer recipe for training VLMs without visual-language re-alignment. International Journal of Computer Vision 134 (7) , 331. 10.1007/s11263-026-02930-z |
Abstract
Vision Language Models (VLMs) with different vision encoders, trained under the two-stage paradigm of pre-training and fine-tuning, have shown varying strengths in different Visual Question Answering (VQA) tasks. However, it remains unclear whether an LLM that has already been aligned during vision-language pre-training can be reused without re-alignment when combined with a different vision encoder for a specific task. To investigate this, we systematically explore VLMs with five different vision encoders using various combinations of LLMs and pre-training strategies. Specifically, we first demonstrate that the aligned LLM with a general-purpose vision encoder can effectively enhance downstream VQA performance with task-specific encoders. Secondly, we investigate several alignment strategies between the aligned LLM and new task-specific encoders. These include (i) feature distillation from the projector layer using both general and task-specific encoders, (ii) a two-stage training strategy with varying the proportion of pre-training data. We find the aligned LLM has acquired transferable vision-language alignment capabilities, such that when combined with new encoders, it no longer requires additional alignment strategy. Evaluated across 13 task metrics after transfer learning with 5 different vision encoders, this new training recipe reduces pre-training time by 2 to 9 hours while achieving comparable or even superior performance.
| Item Type: | Article |
|---|---|
| Date Type: | Publication |
| Status: | Published |
| Schools: | Schools > Computer Science & Informatics |
| Publisher: | Springer |
| ISSN: | 0920-5691 |
| Date of Acceptance: | 15 August 2026 |
| Last Modified: | 06 Jul 2026 10:45 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/187915 |
Actions (repository staff only)
![]() |
Edit Item |




Dimensions
Dimensions