Hamming w/xor+popcount is the only thing I can make numpy do faster than float32 dot products :)
int8s, float16s are all fairly slow. I suppose it’s because BLAS does float32/64 very fast.
int8s, float16s are all fairly slow. I suppose it’s because BLAS does float32/64 very fast.
Separately, I’m a huge fan of your writing about search! It’s been very helpful while I’ve been building https://scour.ing.