Do you have lots of unlabelled data, and if so, do you do any self-supervised pre-training?
Have you ever considered releasing the backbone weights for the pre-trained models you have? No idea if this would be possible without giving up core IP, but I know I'm personally dying for an alternative to Imagenet (COCO for you?) trained on a big dataset.
Are the images in your set diverse enough that you'd expect the backbone to be a good general feature extractor?