build tool is now blocked on file-based checksum and i am L O C K E D the fuck in
but note that this "parent"/"child" terminology is a relative distinction, and says nothing about the identity of each vertex.
in fact, in every filesystem, this is what an "inode" designates--a unique vertex identifier. but this isn't exposed in any useful way at all
i have crashed out about the inode deception many times
anyway, the problem statement that gives you a "root" (i.e. a specific inode to start your paths from--that's all the / directory is, in a virtualized/sandboxed scenario) makes this problem fully deterministic. this requires another small but important inference, which is that directory cycles must necessarily occur as the result of a depth-first traversal
(i said "breadth-first" earlier, which was actually wrong, but only because the distinction didn't matter at all there)
this is to say: directory traversals from a fixed root do not consider arbitrary "back edges" (not yet). and a cycle therefore only arises if you can keep going "downwards" and find a vertex you've seen before
in fact i was boasting earlier. i totally don't know how directory cycles should work yet
but instead of going downwards, what happens if we went upwards from the leaves?
importantly, leaves are definitionally files and not directories. so we're not repeating ourselves the same way.
and more importantly, we have no choice about how we pick the upward path from a leaf!
so we would obviously like to assume that we can only ever have one cycle at a time. but in fact it's very easy to describe a mutually recursive data structure by accident--this is why i believe it's important to support this!
and more importantly, inodes are much nicer for several practical reasons!
in particular, inodes are how you can abstract away a tree-like object database from any particular hashing algorithm. this isn't just a security concern--it in fact makes the case in itself that so-called "universal" identifiers in content-addressed stores are fundamentally flawed
but possibly even more importantly, an inode can be allocated before a checksum is even possible to calculate. and a collection of inodes can be transferred from one connected repository to another completely losslessly
the distinction between a symlink and a directory hard link (the latter of which is currently illegal) is that a directory link is in fact a kind of capability--not only does the filesystem ensure it's always valid by efficient reference counting, it's also guaranteed to represent the same data at all times, no matter who dereferences it
think about that for a moment and then consider our original problem: hashing a directory tree to verify it represents the same data
furthermore, just how important is it that the precise directory hierarchy is the same when we compare tarball checksums? sure, it's tangentially relevant, if you can trigger variant behavior depending upon paths. but we're not using cargo here
i see the data stored in files as fundamentally distinct from its arrangement into some sort of graph structure. and i don't think it serves any end users, or maintainers, or packagers, or anyone else to enforce that in our measurement of filesystem integrity
oh this is so cooking. this is a ratatouille. anton ego is shook
i wasn't even trying to go for this. but yeah obviously it makes everything make so much more sense if the graph is distinct from the file resources