|

File Identity Is Not Content Identity: How ReadMachine Uses BLAKE3

ReadMachine DevLog #001

A library file can have several kinds of identity at the same time.

The filesystem can tell us which resource we are looking at. The application can store a stable Book ID. The file contents can tell us whether two different files are probably the same document.

Those are different questions.

ReadMachine currently supports PDF documents, but this part of the architecture is not meant to depend on PDF as a format. The same problem will exist when the library grows to other document types.

That is why ReadMachine does not try to solve discovery, presentation, and duplicate detection with one identifier.

Today the flow has three separate stages:

  1. metadata-only filesystem reconciliation;
  2. bounded background preparation for library presentation;
  3. content fingerprinting for duplicate detection.

The important part is not BLAKE3 by itself. It is where each kind of identity belongs.

Stage 1: discover files without reading their contents

The first job is simple: find supported files and reconcile them with the library database.

The current scanner asks Foundation for filesystem metadata:

let keys: Set<URLResourceKey> = [
    .isRegularFileKey,
    .fileSizeKey,
    .fileResourceIdentifierKey,
]

It walks the library directory, ignores hidden files, keeps supported documents, and records the file size plus fileResourceIdentifier when the filesystem provides one.

The current implementation only accepts .pdf files here, but the identity logic itself is more general than that format check.

For each discovered file, ReadMachine creates a scan-time fingerprint:

let fingerprint = "file:\(entry.stableIdentifier ?? relative)"

If the filesystem gives us a resource identifier, we use it as the stronger scan-time signal. Otherwise, the relative path is the fallback.

This stage does not read the document bytes.

That distinction matters because reconciliation should stay focused on one task: which library records still match which filesystem entries?

If scanning also parsed documents, generated covers, calculated content hashes, and resolved duplicates, one operation would own too many unrelated responsibilities.

A filesystem resource ID is useful, but it is not content identity

fileResourceIdentifier is a good example of why one identifier cannot answer every question.

Apple describes it as an identifier for a filesystem resource. It can help ReadMachine recognize the same resource when the path changes.

But it is not a permanent document identity. Apple’s Core Foundation documentation also notes that filesystem resource identifiers are not persistent across system restarts.

A copied document makes the difference even clearer.

These two paths:

library/design-systems.pdf
library/archive/design-systems-copy.pdf

can point to two different filesystem resources even when the file contents are identical.

So the filesystem can help answer:

Which resource am I looking at?

It cannot answer:

Does this file contain the same document as another file?

ReadMachine keeps those questions separate.

Stage 2: prepare documents for the library UI

The scanner itself is metadata-only, but that does not mean ReadMachine waits until the Reader opens a document before reading any bytes.

That was true in an earlier version of the architecture. It is not true anymore.

After reconciliation completes, WorkspaceViewModel starts a background presentation-preparation pass for available books.

The current coordinator keeps at most two preparation tasks active at once:

for _ in 0..<2 {
    guard let book = iterator.next() else { break }
    group.addTask { await store.preparePresentation(for: book) }
}

For the current PDF implementation, preparePresentation reads the file and builds a PDFDocument:

guard !Task.isCancelled,
    let data = try? Data(contentsOf: url),
    let document = PDFDocument(data: data)
else { return nil }

It can then persist presentation data such as:

  • title;
  • author;
  • page count;
  • cover image;
  • password-protected state.

This can materialize a provider-backed file in the background.

That is intentional.

ReadMachine needs complete metadata and covers for Home and the sidebar, and keeping files permanently unmaterialized created its own product problems. The answer was not to make discovery heavier. The answer was to create a separate, bounded preparation stage after discovery.

This distinction is important:

reconciliation stays metadata-only, while presentation preparation is allowed to read document contents.

The active library generation also owns this work, so stale preparation can be cancelled or ignored when the user switches libraries.

Stage 3: content identity is still a separate problem

Reading a file for presentation does not automatically make presentation identity the same thing as duplicate identity.

ReadMachine currently calculates its content fingerprint after the Reader successfully loads the selected document.

In ReaderView, after the document is available, it asks the workspace model to run duplicate detection:

model.detectDuplicateContent(for: book, at: url)

The store then moves fingerprint work off the main actor:

fingerprint = try await Task.detached(priority: .utility) {
    try FastFingerprint.make(for: url, fileSize: Int64(size))
}.value

This trigger may evolve as ReadMachine adds more formats and more shared document-processing infrastructure.

The important architectural rule is more stable:

content identity is not part of filesystem reconciliation.

It is its own stage with its own cost, persistence, and failure model.

The fingerprint does not hash the whole file

FastFingerprint uses BLAKE3, but it does not feed the whole file into the hash function.

It reads:

  • the first 64 KiB;
  • the last 64 KiB when the file is larger than one window;
  • the full file size encoded as big-endian bytes.

The implementation is small:

private static let windowSize = 64 * 1024

static func make(for url: URL, fileSize: Int64) throws -> String {
    let handle = try FileHandle(forReadingFrom: url)
    defer { try? handle.close() }

    let first = try handle.read(upToCount: windowSize) ?? Data()

    let last: Data
    if fileSize > Int64(windowSize) {
        try handle.seek(toOffset: UInt64(fileSize - Int64(windowSize)))
        last = try handle.read(upToCount: windowSize) ?? Data()
    } else {
        last = Data()
    }

    let hasher = BLAKE3()
    hasher.update(data: first)
    hasher.update(data: last)

    var length = fileSize.bigEndian
    hasher.update(data: withUnsafeBytes(of: &length) { Data($0) })

    return hasher.finalizeData()
        .map { String(format: "%02x", $0) }
        .joined()
}

The amount of file data read for the fingerprint stays bounded even when the document itself is very large.

For files between 64 KiB and 128 KiB, the first and last windows can overlap. That is fine for this design. The fingerprint input is defined as “start + end + size”, not as a fixed number of unique bytes.

Why BLAKE3?

BLAKE3 is a cryptographic hash function designed for high performance. Its tree structure allows work to scale well across large inputs.

ReadMachine currently uses the blake3-swift package, which wraps the BLAKE3 implementation for Swift.

But there is an important limitation here:

BLAKE3 is stronger than the fingerprint strategy around it.

The hash function can only hash the bytes we give it.

If two same-sized files have identical first and last 64 KiB regions but different bytes in the middle, ReadMachine gives BLAKE3 the same input for both files.

The result will also be the same.

That is not a BLAKE3 collision. It is a consequence of our sampling strategy.

This is why I prefer the term content fingerprint instead of full-file content hash.

Why not hash every byte?

A full-file hash gives a stronger equality signal.

It also makes fingerprint work proportional to file size.

That does not automatically make it wrong. In fact, ReadMachine may choose a stronger second-stage verification later.

But today the fingerprint has a narrower job: find a likely duplicate candidate.

The fingerprint is persisted on the Book record:

ALTER TABLE books ADD COLUMN contentFingerprint TEXT

Then ReadMachine looks for another available Book with the same content fingerprint.

If it finds one, it stores a duplicate relationship:

CREATE TABLE duplicate_candidates (
    contentFingerprint TEXT PRIMARY KEY NOT NULL,
    firstBookID TEXT NOT NULL REFERENCES books(id),
    secondBookID TEXT NOT NULL REFERENCES books(id),
    detectedAt DOUBLE NOT NULL
)

That does not immediately merge the records.

It only creates a candidate that the user can review.

For this use, bounded fingerprinting is a practical first signal.

If the resolution flow becomes automatic in the future, I would want a stronger final verification step before any destructive action.

Detection and resolution are deliberately separate

Two files can look like duplicate content while their application state is different.

One Book may have the title I want to keep. Another may have better author metadata. The selected canonical Book also determines which Book ID and attached application state survive as the active record.

So duplicate detection stops at a candidate.

The current resolution UI lets the user choose:

  • which source file to keep;
  • which title to keep;
  • which author to keep.

Only after confirmation does ReadMachine move the rejected source file into the library’s Trash directory.

The database update happens separately from the filesystem move. Because those two systems cannot share one transaction, the code uses compensation: if the database update fails after the file move, ReadMachine attempts to move the file back.

That transaction problem deserves its own DevLog entry, but it is also another reason not to treat “same fingerprint” as “merge now.”

The current implementation is PDF-based, but the identity model is not

Today some parts of this flow are naturally tied to PDFKit:

  • current file discovery accepts .pdf;
  • presentation preparation uses PDFDocument;
  • Reader load uses PDFKit;
  • duplicate-resolution copy calls the file a PDF.

Those are implementation details of the current format support.

The identity model itself is more general:

filesystem discovery
        ↓
metadata-only reconciliation
        ↓
bounded presentation preparation
        ↓
document becomes usable by the app
        ↓
content fingerprint
        ↓
duplicate candidate
        ↓
explicit resolution

A future EPUB, comic archive, or another document type can have a different presentation parser and reader while still using the same separation:

  • filesystem identity for reconciliation;
  • format-specific parsing for presentation;
  • content fingerprinting for duplicate detection;
  • application identity for the Book record.

That is the part I want to keep stable.

What I would change later

There are a few places where the current design can become stronger.

Add a second verification stage

The bounded fingerprint is intentionally not proof of full byte equality.

If duplicate resolution becomes automatic, or if the product starts making more destructive decisions, a full-file BLAKE3 hash or byte-for-byte comparison before the final action would reduce false positives.

Make fingerprint scheduling less Reader-specific

The FastFingerprint implementation already works on raw bytes and a URL. It does not depend on PDFKit.

The current trigger lives after Reader load because that was the point where the file was known to be locally usable when this feature was introduced.

Now that ReadMachine has a separate presentation-preparation stage, future work can decide whether fingerprinting should stay Reader-driven, join document preparation, or move into another shared document-processing queue.

That should be a product and performance decision, not something hidden inside filesystem reconciliation.

Keep file identity provisional

fileResourceIdentifier is useful, but it should remain a reconciliation signal rather than a permanent application identity.

ReadMachine already has its own Book UUID for application state. That is the right place for stable app identity.

One file, several identities

The useful lesson from this implementation is simple:

a file does not have one identity that solves every layer of the application.

The filesystem has an identity for a resource.

The application has a Book identity for durable state.

The document contents can produce a fingerprint for duplicate detection.

Those identities can overlap in purpose, but they should not be confused.

In ReadMachine, keeping them separate makes the architecture easier to reason about:

  • discovery can stay cheap and deterministic;
  • presentation work can be bounded and cancellable;
  • duplicate detection can evolve independently;
  • format-specific code does not have to define application identity;
  • a content match does not automatically become a destructive action.

BLAKE3 is one tool inside that design.

The more important decision is knowing which question we are asking before we choose the identifier.

Sources