▲ 2 points
back
1 comments
While this excels at human understanding of natural language tasks, it seems generalizable across problem domains. Biomedical microscopy images for disease detection, for example. The "cross-modal" representation works! Just as effectively as it does for single modes such as image, text and sound. A very interesting finding ;)