Making np.searchsorted up to 25× Faster in NumPy 2.5
We explore how to speed up binary search by batching independent searches with NumPy's vectorized operations. We then reformulate the algorithm so all searches progress together with only $O(1)$ additional memory, port it to C++, and achieve up to a 25× speedup over NumPy 2.4's implementation.