VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.
- https://huggingface.co/blog/vlms
- https://en.wikipedia.org/wiki/Multimodal_learning