Many of the models we have today seem to only perform OCR on the images you send and use the text retrieved for context when answering. However, Qwen-VL, and I guess Gemini now? Are different, they seem to "understanding" the image I send with my prompt. They manage to capture spatial relationships, objects, and semantics from the image, it's very impressive. I’ve been telling my friends about the Qwen3-VL model option in Qwen Chat for a while because I feel like it’s underrated.