▲ 2 pointsText or pixels? On the token efficiency of visual text inputs in multimodal LLMsarxiv.orgby hhs·9mo ago·0 comments·view on hn ↗