The optimal case was always loading up the coprocessor memory and letting it crunch the data until it finishes. One nice thing was the low cognitive load at a time CUDA was not a great experience - it just pretended to be a normal computer, if a bit memory-starved. A classic “lopsided machine”, good for some things, not great for most others.
And I agree - it was getting a lot better by the time it was killed. I’d love to find a reasonably priced Phi machine, but it seems the latest batches, with proper virtualisation support, went straight into supercomputers.