خرید بک لینک

Vote count: 0

I want to hash numpy arrays without copying the data into a bytearray first.

Specifically, I have two contiguous read-only numpy arrays A and B that are parallel, A.shape[0] == B.shape[0] and A[i] corresponds to B[i]. I want to make a constant time function that maps an arbitrary A[i] to its B[i]. The natural choice is to make a dictionary like:

lookup = dict(zip(A, B))

Unfortunately numpy arrays are not hashable. There are ways to hash numpy arrays, but ideally I'd like the equality to be preserved so I can use it in a dictionary without writing manual collision detection.

The referenced article does point out that I could do:

lookup = dict(zip((a.data.tobytes() for a in A), B))

And then look up B with lookup[ap.data.tobytes()], however the tobytes method retus a copy of the data pointed to by A.data, hence doubling the memory usage.

What I'd love to do is something like:

lookup = dict(zip((a.data for a in A), B))

This would ideally use a pointer to the underlying memory instead of the array object or a copy of its bytes, but since A.dtype == int, this fails with:

ValueError: memoryview: hashing is restricted to formats 'B', 'b' or 'c'

This is fine, we can cast it Aview = A.view(np.byte), now we have:

>>> Aview.flags
#   C_CONTIGUOUS : True
#   F_CONTIGUOUS : False
#   OWNDATA : False
#   WRITEABLE : False
#   ALIGNED : True
#   UPDATEIFCOPY : False

>>> Aview.data.format
# 'b'

However, when trying to hash this, it still errors with:

TypeError: unhashable type: 'numpy.ndarray'

A possible solution (inspired by this) would be to define:

class _wrapper(object)
    def __init__(self, array):
        self._array = array
        self._hash = hash(array.data.tobytes())

    def __hash__(self):
        retu self._hash

    def __eq__(self, other):
        retu self._hash == other._hash and np.all(self._array == other._array)

Then I could do:

lookup = dict(zip((_wrapper(a) for a in A), B))
lookup[_wrapper(ap)]

But this seems inelegant. Is there a way to tell python to just interpret the memoryview as a bytearray and hash it normally without having to make a copy to a bytestring or having numpy intercede and abort the hash?

Other things I've tried:

  1. The format of A does allow me to map each row into a distinct integer, but for very large A the space of possible arrays is larger than np.iinfo(int).max, and while I can use python's integer types, this is ~100x slower than just hashing the memory.

  2. I also tried doing something like:

    Aview = A.view(np.void(A.shape[1] * A.itemsize)).squeeze()

    However, even though A.flags.writeable == False, A[0].flags.writeable == True. When trying hash(A[0]) python raises TypeError: unhashable type: 'writeable void-scalar'. I'm unsure if it's possible to mark scalars as read-only, or otherwise hash a void scalar, even though most other scalars seem hashable.

asked 15 secs ago

برچسب: نویسنده: استخدام کار تاريخ: چهارشنبه 6 مرداد 1395 ساعت: 22:30

صفحه بندی