build tool is now blocked on file-based checksum and i am L O C K E D the fuck in
and more importantly, inodes are much nicer for several practical reasons!
in particular, inodes are how you can abstract away a tree-like object database from any particular hashing algorithm. this isn't just a security concern--it in fact makes the case in itself that so-called "universal" identifiers in content-addressed stores are fundamentally flawed
but possibly even more importantly, an inode can be allocated before a checksum is even possible to calculate. and a collection of inodes can be transferred from one connected repository to another completely losslessly
the distinction between a symlink and a directory hard link (the latter of which is currently illegal) is that a directory link is in fact a kind of capability--not only does the filesystem ensure it's always valid by efficient reference counting, it's also guaranteed to represent the same data at all times, no matter who dereferences it
think about that for a moment and then consider our original problem: hashing a directory tree to verify it represents the same data
furthermore, just how important is it that the precise directory hierarchy is the same when we compare tarball checksums? sure, it's tangentially relevant, if you can trigger variant behavior depending upon paths. but we're not using cargo here
i see the data stored in files as fundamentally distinct from its arrangement into some sort of graph structure. and i don't think it serves any end users, or maintainers, or packagers, or anyone else to enforce that in our measurement of filesystem integrity
oh this is so cooking. this is a ratatouille. anton ego is shook
i wasn't even trying to go for this. but yeah obviously it makes everything make so much more sense if the graph is distinct from the file resources
i was actually trying to move on to file hashing which is actually much less obvious to me and also something i think i need to solve sooner rather than later
what i was going to say:
i have a "simple" answer which considers a file as an unstructured contiguous bit string of known size. but one reason my directory checksum is very useful is the ability to maintain a hash tree independently of the input data that attests to the integrity of every recursive component.
this especially means you can merge hash trees (e.g. composing directories to form a chroot) and diff tree states (e.g. to capture output from a process execution) without reference to any object database
in other words you can create an algebra of filesystem states which is far more than merely "reproducible", and can actually describe what changed and whether that was correct
and all that is even more true with the flattened representation described above where each inode vertex has its directed graph relationships described completely independent of its identity
what i was trying to do was to shit on the "sponge construction" used in sha-3. it's sof ucking funny
https://en.wikipedia.org/wiki/Sponge_function
The sponge function "absorbs" (in the sponge metaphor) all blocks of a padded input string as follows:
Sis initialized to zero- for each
r-bit blockBofP(string)
Ris replaced withR XOR B(using bitwise XOR)Sis replaced byf(S)
what is f(S)? well,
fproduces a pseudorandom permutation of the2^bstates fromS.
in other words, f is the actual fucking hash function
like very literally f is producing a permutation (i.e. an encryption) that is predicated only upon the bits of S instead of a separate key. that's the actual definition of a hash function
that's why it's so deceptive to say "a hash function produces a smaller output". no it doesn't!!!!
; printf '' | wc -c
0
; printf '' | sha256sum
e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 -
that's an encryption! and that sucks if i actually want to know what my data is, as opposed to wanting my data to be secret. using the same construct for cryptographic signatures is completely fucked when applying it to data integrity
and it's also just completely not the way hashes are used literally ANYWHERE else. the reason bloom filters work is because the hash functions are crafted to achieve collisions without having to calculate pairwise bit-for-bit equality over each pair of database entries!
radix sort gets better than n log n runtime for this exact reason!
so i refuse to call it a "hash". it's a "checksum" at best