How to save an APFS HDD

One morning my external hard drive stopped mounting. It’s a large USB drive, encrypted APFS, plugged into a Mac that shares it over the network and takes Time Machine backups from my other Macs. Several terabytes of data on it, some of it nowhere else.

Disk Utility’s First Aid ran for a minute and ended with this:

$ diskutil verifyVolume disk5
Started file system verification on disk5s1
Verifying file system
...
error: btn: invalid key order: minkey is less than index 0 (should be equal)
error: btn: unable to repair minkey
   Object map is invalid.
** The volume ... was found to be corrupt and cannot be repaired.
ShellScript

A note on device names, since they change from block to block below. APFS stacks two devices on one partition: disk4 is the USB drive and disk4s2 its partition, the “physical store” where the bytes live, while disk5 is the APFS container macOS builds on top of it and disk5s1 the volume you mount. Tools that read or patch raw blocks talk to /dev/rdisk4s2, the r meaning the raw, unbuffered device, and Apple’s checks talk to the container or the volume, disk5 and rdisk5s1. The backup image inside the drive gets its own pair in the same way, disk6s2 and disk7.

The threads I could find with that error either end unanswered or with the same advice: recover what you can and reformat1. I didn’t have a backup of this hard drive.

Five days later the drive mounts, passes First Aid, has had well over a hundred gigabytes of test writes and a real Time Machine backup land on it without a complaint, and nothing was erased. I did this with Claude Code, which did the reading, wrote the tools and did most of the thinking, while I approved each step that touched the disk. This is the long version, including the part where our first repair quietly broke it again. In short, what we did:

  1. Worked out from the logs when the drive broke and whether the hardware was failing.
  2. Read the unmountable drive with a patched open-source driver, and salvaged the lost index entries from older copies still on the disk.
  3. Copied everything out with its metadata intact, and verified the copy with a second read.
  4. Repaired the index in place, 59 blocks, after rehearsing on test images and on an overlay of the real drive, with Apple’s checker as the judge.
  5. Watched it corrupt itself again on the first writable mount, traced that to the free queue, and built the ownership check that would have caught it.
  6. Repaired it a second time, this time into fresh blocks, and put well over a hundred gigabytes of test writes through it, each verified.
  7. Fixed the Time Machine image inside the drive twice: its object map and free-space bitmaps, then the same free-queue problem by moving 2,275 blocks.
  8. Ran the first real backups, cleaned up what an interrupted one left behind, and checked the drive after each.

What had actually happened

The first decision was whether to treat this as a dying drive or a damaged file system, because the two call for opposite moves: a dying drive wants its data copied off at once and nothing else, a damaged file system can be studied. SMART was no help, since macOS cannot read it over USB on Apple silicon without a third-party driver. What we could do was read, and look at the shape of the damage. None of that needs the volume mounted. A raw partition device can be read block by block even when macOS refuses to mount it, and every APFS page carries a checksum, so you can tell a readable page from a damaged one without any help from the file system:

$ sudo dd if=/dev/rdisk4s2 bs=4096 skip=14127044 count=1 2>/dev/null | xxd | head -n 4
ShellScript

Every block of the index came back on the first try. More telling was what the damaged blocks held. They were not unreadable or garbled. Each was a valid, checksummed, older version of the right page. That came from a small scanner Claude wrote on the first day, which walks the object map from its root on the raw device and compares every page with what its parent expects (more on it in the next section). Its bad-child lines report pages with a lower transaction number than the parent points at, with the checksum intact:

$ apfs-omapscan /dev/rdisk4s2 > omapscan.log
$ grep "BAD CHILD" omapscan.log | head -n 1
BAD CHILD: parent 0xd7d42c index 100 -> block 0xd82f1b (expected level 0): first key differs from parent's key [found level 0, xid 1202685, 110 keys, ...]
ShellScript

A failing platter or a bad cable gives you read errors and broken checksums. Writes that never arrived, with the previous content still sitting there, point at the path between the Mac and the platters: a USB reset or a power dip while the drive still held the writes in its cache. That fitted the earlier USB drop-out and the one read that came back full of FF bytes, both on the hub dongle.

The kernel’s own view backs this up. The drive’s error and retry counters sit in the I/O registry, and the unified log keeps every disk and USB complaint. Both can be read with the volume unmounted:

$ ioreg -c IOBlockStorageDriver -r -w 0 -l | grep -A1 '"BSD Name" = "disk4"' | grep -o '"Errors (Read)"=[0-9]*\|"Errors (Write)"=[0-9]*\|"Retries (Read)"=[0-9]*\|"Retries (Write)"=[0-9]*'
"Errors (Write)"=0
"Retries (Read)"=0
"Errors (Read)"=0
"Retries (Write)"=0
$ log show --last 24h --style compact --predicate 'process == "kernel" AND (eventMessage CONTAINS[c] "I/O error" OR eventMessage CONTAINS "not readable" OR eventMessage CONTAINS "USBMSC")'
Timestamp               Ty Process[PID:TID]
ShellScript

An empty log and zero counters do not prove a drive healthy, but a drive that is failing rarely manages both.

The second job was working out when it broke. The logs on both Macs disagreed with my memory, so we asked the system log directly:

$ log show --start "<that morning>" --predicate 'eventMessage CONTAINS "disk5" OR eventMessage CONTAINS "apfs"' --style compact
ShellScript
  • First Aid had passed cleanly a week earlier.
  • The first failure came in the small hours of the morning. A disk image stored on the drive could no longer save one of its pieces. The error code was 92, which is what APFS returns when it detects corruption.
  • I noticed nine hours later, when Finder couldn’t read files. I tried to eject, the eject failed, and the Mac kernel panicked. For a while we blamed the panic. The logs showed the damage was already there.

Two earlier incidents turned up as well. Two weeks before, the drive had vanished from USB entirely. A week before, the kernel had asked the drive for a block and got back one filled with FF bytes, although the block on disk was fine when checked three quarters of an hour later.

A drive that drops off the bus, returns garbage on a read, and then loses writes is not a file system bug. The drive was connected through a USB hub dongle that also carried the Mac’s Ethernet, and every incident happened while another Mac was backing up to it over the network. I can’t prove the dongle did it. The drive now sits on a direct cable, and every write since has been clean.

How APFS finds things, briefly

APFS is the file system Apple has used on everything since 2017, the layer that turns a disk full of blocks into files, folders and permissions. It replaced the thirty-year-old HFS+ and was built for flash storage: it never overwrites data in place, it checksums its own bookkeeping, it can take snapshots of a whole volume in an instant, and it encrypts natively. Time Machine on modern macOS relies on those snapshots, which is why the backup image in this story is APFS too. Four of its concepts come up again and again below, so here they are in one place.

Save points. Besides your files, a file system keeps its own records about them: which blocks each file occupies, the folder tree, the names, the dates and permissions, and the map of which blocks on the disk are free. That is its bookkeeping, and it is what breaks when writes are lost, even when every byte of your files is still there. APFS never overwrites its bookkeeping in place. Every change writes new blocks and then a new save point (a checkpoint) that refers to them, and each save point carries a transaction number. If the Mac loses power halfway, the previous save point is still complete, and that is where the file system resumes.

The object map. The pages of that bookkeeping are referred to by an object ID, not by their position on the disk, because every change writes a page to a new block and the position would change constantly. The object map is the index that turns an ID into a position, the number of the block on the disk: “object 0x1234, as of transaction N, lives at the position of block B”. Lose an entry of that map and the page is still on the disk, but nothing can find it.

The file index. One big sorted tree, a B-tree, that holds every record about every file on the volume: the folder entries that give a name to each file, the file records with dates, permissions and sizes, and the extent records that say which blocks on the disk hold the file’s data. The tree is made of pages, usually called nodes, each one a 4 KB block holding a batch of sorted records. The nodes at the bottom hold the records; the nodes above them hold only pointers to the nodes below, so a lookup walks from the top down in a few steps. The pointers are object IDs, not positions, so every step goes through the object map. A node that the map cannot locate takes all its records with it, and every file below it in the tree becomes unreachable even though its data is untouched.

The free queue. Because nothing is overwritten in place, every change leaves an old block behind: the previous version of a page, or the data of a deleted file. APFS cannot free such a block at once, since the last save point still refers to it and must stay usable until the new one is safely on disk. So the block goes into a queue, tagged with the transaction number of the change that retired it, and is released a few transactions later, once no save point can refer to it any more. The queue is a small tree of its own, and it is part of the bookkeeping, so it too can be out of step with reality after lost writes. Nothing checks it against the rest. Remember this one. It bit us.

A few smaller terms turn up in the tool output below. An extent is one run of consecutive blocks holding part of a file; a file’s extent records say which runs it occupies. The free-space bitmap is the map with one bit per block saying whether it is free. In the output, oid is an object ID and xid a transaction number, both printed in hexadecimal. First Aid in Disk Utility is a front end for the command-line checker fsck_apfs; the two give the same verdicts, and the command form can be run on an unmounted disk or an attached image. A container is the APFS structure on a partition that holds one or more volumes and owns the free-space accounting for all of them.

Because nothing is overwritten, old versions of all of these linger on the disk until the space is reused. That turned out to be the raw material for the repair, and also the trap.

Reading a drive macOS refuses to mount

Mounting is what turns a disk into a folder you can open in Finder. macOS would not mount this volume because its own driver hit the broken object map and gave up. The way around it is a different driver: apfs-fuse is an open-source, read-only APFS implementation that runs as an ordinary program rather than inside the kernel, using FUSE (Filesystem in Userspace), the standard way to plug such a program into the operating system as a file system. On macOS it attaches through fuse-t, which needs no kernel extension, and it handles encrypted volumes if you give it the password. With small patches it built on this Mac.

The first mount attempt died at once. To see why, Claude added a small command-line lister to the driver’s source tree, apfs-ls, which uses the driver’s library to open the volume and list one folder without mounting anything, so every error shows up on the terminal. It is not a macOS tool; nothing Apple ships can read a volume in this state:

$ security find-generic-password -a "$VOLUME_UUID" -s "$VOLUME_UUID" -w | xxd -r -p | apfs-ls /dev/rdisk4s2 1
[apfs-ls] listing inode 1 ...
ERROR: GetNode: omap entry oid faa5b xid 125a30 not found.
ShellScript

Inode 1 is the root folder, the top of the whole file tree. The error says the object map has no entry for object faa5b as of transaction 125a30, and that object is the index node holding the root folder’s records. Without it, nothing below the root can be reached. (The password is piped from the keychain so it never appears on a command line.)

So the next question was what state the object map itself was in. The object map is a tree too, with its own top node, not to be confused with the root folder of the file system that the lister could not reach. Claude wrote a scanner, apfs-omapscan, that starts at the map’s top node, follows every pointer down, and checks each page it lands on against what the parent expects: the right level, the right first key, a valid checksum. Then it does the opposite as well, reading through the region of the disk where the map lives for pages that are valid but no longer linked from anywhere:

$ apfs-omapscan /dev/rdisk4s2 > omapscan.log
BAD CHILD: parent 0xd7d42c index 100 -> block 0xd82f1b (expected level 0): first key differs from parent's key [found level 0, xid 1202685, 110 keys, ...]
...
orphan index node @ 0xd87a7f: level 1, xid 1082784, 143 keys, oid 0xf6c66..0xfa126
scan: 9969 valid omap nodes in range; in the damaged key range: 165 orphan index nodes, 540 orphan leaves, 53287 mappings written
ShellScript

A “bad child” line means the parent page points at a block and the block holds a different page than the parent expects. Over the whole map there were 38 of them. They weren’t garbage. Each was a valid, checksummed object map page, just an older version of the one the parent wanted, because the newer version had never reached the disk. That is what a lost write looks like: the block still holds whatever was there before.

The orphan lines are the other half. An orphan is a page that is still intact on the disk but no longer linked from any parent, which in APFS means an older version that was superseded and whose block has not been reused yet. The scanner found 540 such bottom-level pages (“leaves”) from the damaged part of the map, and together they held 53,287 entries of the form “object X, version Y, lives in block Z”. Those are the entries the 38 lost pages used to hold, one version older.

The driver got a fallback built on them: when the live object map cannot resolve an object, look it up in the salvaged entries, newest version first, and accept the block it names only if the block’s checksum is valid and the object ID and version written inside it match. With that fallback, the root folder listed and the volume mounted read-only.

Copying it out first

The first rule of recovery is to get the data off before touching the disk, so the next step was a complete copy to the Mac’s internal drive. Copying through the FUSE mount with Finder or cp works but drops metadata: fuse-t presents the volume through an NFS layer that cannot carry creation dates, extended attributes or the file flags macOS uses. Worse, it turned out that the driver returned zeros for a block it failed to read and reported success, so a copy made that way could contain silently blank files.

So the copy went through a direct tool, apfs-copy, which reads the volume with the same library, never goes through a mount, writes each file with its real dates, attributes and flags, keeps hard links as hard links, names every file it cannot read instead of papering over it, and can resume where it stopped:

$ apfs-copy /dev/rdisk4s2 /Backups ~/recovered/Backups --resume
[apfs-copy] 1224381 files skipped (already copied)
[apfs-copy] done: 1097157 hard links, 0 compressed files, 350472 xattrs set
ShellScript

That run shows the tool resuming a folder of 1.2 million files it had already copied; across all folders the copy came to over two million files, more than a million of them hard links, that is, extra names for a file that already exists, which Time Machine folders are full of and which a naive copy turns into duplicates.

A copy is only as good as the read that produced it, and that read had gone through the hub dongle suspected of the damage. So once the drive was on a direct cable, a sampling pass re-read ten percent of the files plus every large file straight from the drive and compared them byte for byte with the copies. Over two hundred thousand files, no mismatches.

At this point the data was safe, and the sensible move was to erase the drive and copy everything back. I asked whether it could be repaired instead.

A mistake worth describing

Five folders were still unreadable, because their index nodes were among the ones the salvaged entries did not cover: the driver knew which object ID each folder’s node had, and had no entry anywhere saying which block held it. So we scanned more widely. Every index node carries its own object ID, transaction number and checksum in its first 32 bytes, so a node can be recognised without any map. The scan decrypted every block in the first half terabyte of the drive, the region where the index lives, kept every block whose header and checksum made it a valid file index node, and wrote down the three facts for each: object ID, version, block. That gave a second lookup table, built from the blocks themselves rather than from the object map. When the driver’s lookup for one of the five folders failed in the map and in the salvaged entries, it searched this table for the folder’s object ID, took the newest version not later than the one wanted, and checked the block’s checksum once more. That recovered the five folders. It also broke the next tool, which went into an infinite loop.

The reason took half a day to find. The drive holds Time Machine disk images, and each of those images contains a complete APFS file system of its own, with its own file index. The wide scan cannot tell the two apart: a valid index node is a valid index node. And since every APFS volume numbers its objects from the same starting point, some nodes from inside the images carried the same object IDs as nodes of the outer drive. When the lookup code had several candidate blocks for one ID, it preferred the one with the newest transaction number, and sometimes that was a block belonging to a backup image.

For half a day this looked like the drive had “mixed versions” all through its index, as if pages from different points in time had been stitched together. It hadn’t. The fix was to accept a candidate only if it sits in the region where the drive’s own index lives and if its keys fit under the parent that points at it. After that, the audit of the whole file index came back with nothing out of place, apart from five blocks that had been lost for good.

The first repair

A lost write leaves the right block holding old content. So a repair does not need to move anything. It writes the correct content into the same block, and no pointer anywhere has to change. The correct content was reconstructable: for each lost object map page, the salvaged orphan one version older plus the scan of the index told us exactly which entries the page should hold.

The plan for the real drive came to:

What the patch fixesCount
Object map pages at the bottom level whose newest version was lost37
Pages that were in the right place but still held outdated entries19
Lost page one level up, all 140 of the pages it points at found intact1
Blocks to write in total59
Size of the patch242 KB

The outdated pages are the nasty kind. They pass every structural check, since they are valid pages in the right place; you only find them by comparing each entry with the blocks that exist.

Nothing was written to the drive until three things had been done.

Tests on fake damage. We built encrypted test volumes, saved an image of one at an earlier moment, kept writing to it, then copied the earlier content back into chosen blocks of the final image. That reproduces lost writes exactly, and you know the right answer. The repair had to get Apple’s fsck_apfs to say “appears to be OK” with every file’s checksum matching, and the undo had to restore the damaged image byte for byte, including after a repair interrupted halfway.

A rehearsal on the real data. macOS can attach a disk with a shadow file, so that reads come from the disk and writes go to a separate file. It refuses above roughly 8 TB. The replacement was a small FUSE layer, an overlay, that opens the drive read-only and keeps every write in memory. On that overlay the 59 blocks went in, and macOS mounted the volume by itself for the first time in two days.

Apple’s own tool as the judge. First Aid on the overlay now passed the object map and complained only about the five lost index nodes. Run in repair mode on the whole container, still on the overlay, it removed those five, fixed 188 free-space records, and reported the volume and container OK.

Then the real thing, in two stages. Stage one wrote the 59 blocks to the drive, with the originals saved first and each block read back. It took under a minute and was fully reversible. Stage two was Apple’s repair, which is not reversible:

$ sudo python3 omap_repair.py apply /dev/rdisk4s2 wd.patch wd-real.undo
$ sudo fsck_apfs -y /dev/rdisk5
...
** The volume /dev/rdisk5s1 ... appears to be OK.
** The container /dev/rdisk5 appears to be OK.
ShellScript

The drive mounted. We spent the next day on the Time Machine image inside it, which had its own, older damage (more on that below). Then, on the evening of the third day, I started a backup to the drive. It failed after four minutes.

It broke again, and this time it was our fault

The backup had died on a missing file inside one of the backup images, which was odd but explainable: that piece had been lost in the original incident. The worrying part came from running First Aid afterwards:

$ diskutil verifyVolume disk5
...
error: btn: invalid o_oid (0x...), expected (0x...)
error: btn: invalid o_xid ...
** The volume ... was found to be corrupt and cannot be repaired.
ShellScript

Each of those lines means the object map points at a block for some index node, and the block now holds a node with a different ID or version. 106 pages of the file index were affected. Not old content this time. Newer content, written in the first minutes after the repaired drive had been mounted read-write, hours before the backup. The drive, the cable and the kernel had reported nothing.

We had a dump of the index taken before the repair. Comparing the 106 damaged pages against it gave the answer in an hour, and it is the one thing I’d want anyone attempting this to know.

When the writes were lost, the newest copies of those 106 index pages had vanished, and the salvaged entries our repair put back pointed at the older surviving copies, so the repair “re-linked” those. But APFS had already put those older blocks in the free queue, because from its point of view they had been replaced by the newer copies. Neither our tool nor Apple’s fsck_apfs -y looks at the queue. The first time the volume was mounted writable, the queue did its job: it released the blocks, and macOS reused them for whatever it wrote next, over the top of the index pages we had just made live again.

What the comparison showedCount
Index pages damaged by the first writable mount106
Of those, pages our repair had re-linked to an older copy104
Of those, whose block was sitting in the free queue before the repair104
Blocks released by the queue and overwritten by macOS106

The two pages that were not re-linked by the repair had been double-released: their blocks had stale queue entries as well as live ones, so the queue freed them twice. So the repair had been right about the content and wrong about one piece of bookkeeping, and that was enough. The lesson is blunt: when you re-link an old copy of anything in APFS, you must also take its block out of the free queue, and you must rehearse a read-write mount with writes on the real data, not just on test images.

The second repair, with a gate this time

The older copies were gone, overwritten. But the pre-repair dump still held the exact versions of all 106 pages, so the second repair put them into fresh, free blocks that sat in no queue, and pointed the object map at the new positions. On an encrypted volume each block is encrypted with its own block number mixed into the key, so a page moved to a new block has to be decrypted and re-encrypted for its new address; the tool does that. Three more blocks were patched to mark the new positions as allocated in the free-space bitmap.

Before writing, we built the check that would have caught the first repair. Call it the ownership check. Its first argument names a dump of the index made by a companion tool, which is how the encrypted drive’s file records reach it without a mount. It reads the whole drive’s bookkeeping, every object map entry, every file index node, every record of which blocks each file occupies, the free-space bitmaps and both free queues, and from that works out who owns every block on the disk. Then it insists on three things:

$ uv run refcheck.py wd-refs-w8 /dev/rdisk4s2
(a) referenced but marked FREE in the bitmap: 0 ranges, 0 blocks
(b) referenced blocks that sit in a free queue: 0
(c) blocks with two different owners: 0 overlaps
RESULT: PASS   (19 s)
ShellScript

In words: nothing in use is marked free, nothing in use is waiting to be freed, and no block is claimed by two owners. Run against a reconstruction of the drive as the first repair had left it, the check fails on rule (b) with exactly 107 blocks: the 106 that got destroyed plus one more that had not been overwritten yet. Run against the second repair’s plan, it passes. Apple’s checker says OK for both states. That difference is why I’d never again take “appears to be OK” as the only word.

Then the writing was rehearsed on a local copy of the index with a read-write mount and gigabytes of real file operations, checked again, and finally applied to the drive: 154 blocks in all, with undo saved.

Proving it holds up under writes

After that we did the write tests we should have done the first time, all on the real drive. Each test was followed by the same gate: the ownership check, Apple’s check, and a byte-for-byte compare of every index page that existed before the test against its state after, which proves that writing new data did not disturb anything old:

$ ./realtest.sh gate w7 w8
[apfs-refs] finished: 0 structural issues in total
RESULT: PASS   (19 s)
The volume /dev/rdisk5s1 ... appears to be OK
The container /dev/disk4s2 appears to be OK
1. old index nodes still in use: 567780 unchanged; content changed: 0; at a different block: 0
GATE w8: PASS
ShellScript
TestResult
First writable mount, with a small workload of files created, renamed, overwritten and deletedpass
Remount, verify every file from the first test, more churn, deletepass
100 GiB written in large files, read back after a remountidentical
A throw-away sparse bundle created on the drive with a large file tree inside, re-attached and checked with First Aidpass
4,000 random pre-existing files compared with the copy made before the repairidentical

Well over a hundred gigabytes written in a night, zero errors, zero retries, every old index page unchanged.

The backup inside the drive

The file-server Mac’s own Time Machine image on the drive had stopped opening a week before all this. It is a sparse bundle: on the outside a folder of 8 MB files called bands, on the inside one big virtual disk with an APFS volume on it, holding 60 backups going back most of a year. Each backup is a snapshot of that inner volume. Every check of the image starts the same way, attaching it read-only so nothing can touch it:

$ hdiutil attach -readonly -nomount "/Volumes/<drive>/Time Machine/<name>.sparsebundle"
/dev/disk6          	GUID_partition_scheme
/dev/disk6s1        	EFI
/dev/disk6s2        	Apple_APFS
ShellScript

The inner volume had lost writes too. Its 14 newest save points were unusable, and 320 object map pages were gone. That is a harder version of the same problem, because of the snapshots: each snapshot must keep seeing the index exactly as it was, so the object map holds several versions of most objects side by side, 8.09 million entries for about 5.1 million objects. Rebuilding a lost page means working out which versions of which objects it held, which depends on which snapshots still need them.

The plan took three attempts. The second resolved every current object correctly and still failed its own check, because of how the map’s pages are divided. The code assumed each page covered a range of object IDs; in fact a page covers a range of object ID and version pairs, so the versions of one object can straddle two pages. With that fixed, almost every surviving page turned out to be exactly where it belonged, and the patch shrank from 938 blocks to 381.

Apple’s checker then found 247 blocks in use but marked free in the free-space bitmaps. That is harmless while the volume is read-only, and fatal on the first write, because the next backup would be handed those blocks as free space. Apple’s repair mode fixes it, but insists on walking all 60 snapshots first, which on this image takes days. A small tool did the narrow job instead: set the 247 bits in the bitmaps, lower the free counts to match, recompute the checksums. 81 blocks.

What the first fix wroteBlocks
Rebuilt object map pages381
Save points retired, so that checks start from the last good one14
Free-space bitmap corrections81
Total476

Then came the free-queue lesson from the main drive, and we ran the new ownership check against this image too. (The capture: argument tells the checker to keep a local copy of every block it reads, so that later runs and other tools can work from the copy instead of seeking across the HDD again.) It failed, of course, in the same way:

$ uv run refcheck.py - capture:/dev/rdisk6s2:mini.pack
(a) referenced but marked FREE in the bitmap: 0 ranges, 0 blocks
(b) referenced blocks that sit in a free queue: 2275 (1859 file-index nodes, 416 object-map nodes)
(c) blocks with two different owners: 3 overlaps
RESULT: FAIL
ShellScript

A writable mount would have wrecked it just like the drive. Here the fix was to move rather than to edit the queue: each of the 2,275 blocks was copied to a free block, and the object map entry or the parent page that pointed at the old block was pointed at the new copy. One constraint shaped the choice of new blocks. A sparse bundle only stores the bands that were ever written to, and the tool that patches band files cannot create new bands, so the new blocks had to fall inside bands that already exist. The patch came to 2,482 blocks. It was rehearsed three times on an overlay of the image, each time with two writable mounts in between so that macOS could finish an object map clean-up it had been in the middle of when the image broke, and each time the ownership check and Apple’s check passed afterwards.

On the fifth day it went into the real image, with the bundle detached and the affected bands copied aside first:

$ python3 bundle_patch.py apply "<name>.sparsebundle" 209735680 mini-reloc2-real.undo mini-reloc2.patch mini-reloc2-space.patch
blocks checked: 2482, differing from the rehearsal's originals: 0
wrote 2482 blocks (2482 distinct) into 159 band files; originals saved in mini-reloc2-real.undo; read-back mismatches: 0
$ sudo fsck_apfs -n -S /dev/rdisk7
...
** The volume /dev/rdisk7s1 ... appears to be OK.
** The container /dev/rdisk7 appears to be OK.
ShellScript

The -n tells Apple’s checker to report without repairing, and -S to skip its pass over the 60 snapshots, which alone takes days on this image. Then the whole battery again on the real image, this time with a few extra checks I had asked for: is every block the inner volume refers to really stored in a band on the drive, does every one of the 60 backups have a readable entry point, and does a random sample of 60,000 index pages hold what the map says. All passed.

One thing those extra checks turned up is worth knowing if you ever look inside a sparse bundle: about 1.7 million data blocks the inner file system refers to have nothing stored behind them. That is normal. Databases reserve space past their end and never write it, partially overwritten extents keep their old length, and a sparse bundle only stores blocks that were actually written. Reads of those return zeros, and nothing is missing.

The first real backup

That evening the drive got its first Time Machine backup since the repair. The last one before all this had failed within minutes. This one was a deep scan, because the local reference snapshot was lost with a reboot, so Time Machine compared every file against the last good backup and wrote only the changes. Its own log tells the two apart:

$ log show --info --last 3m --predicate 'process == "backupd" AND subsystem == "com.apple.TimeMachine"' | grep -E "Copied:|Projected"
	Backup Projected Stats: 13260179 items (p:1.79 TB)
	Copied: 82381 (p:44.42 GB) Propagated: 2239425 (p:194.74 GB)
ShellScript

“Propagated” is what it links to the previous backup without writing; “Copied” is what it writes. Most of the copied part turned out to be my own working folder for this repair, which is excluded now.

I stopped that run on purpose and removed its leftover folder by hand, which is harder than it sounds on an APFS backup volume. The folders belong to root, the Data folder inside and some folders below it carry the system no-unlink flag, and a few files keep the immutable flags they had on the Mac, so rm refuses even with sudo until the flags are cleared. Finder cannot do it at all, and tmutil delete does not know about interrupted folders:

$ V="/Volumes/<backup volume>"
$ sudo find "$V/<date-time>.interrupted" \( -flags +sunlnk -o -flags +schg -o -flags +uchg \) -print0 \
    | sudo xargs -0 chflags -h nosunlnk,noschg,nouchg
$ sudo /bin/rm -rfv "$V/<date-time>.interrupted"
ShellScript

The real backups are APFS snapshots of that volume, so a folder in the live volume can go without touching them. Three things to know before trying it. Only a few dozen entries per folder carry a flag, mostly the protected system and home folders near the top plus a handful of files that were immutable on the Mac, so find them rather than walking every entry with chflags -R; the walk is the slow part either way, but it needs doing only once. Some of the flagged entries are symbolic links, and chflags changes the link’s target unless you give it -h. And turn the automatic backups off first: one started by itself halfway through, renamed the folder under the running rm and opened a new one. The rm carried on by directory handle and finished in half an hour, which was far quicker than my estimate. It later turned out the rm is optional: once the flags are gone, Time Machine’s own cleanup deletes the folder at the end of its next completed backup.

After the removal the drive passed its gate again, the image passed the ownership check, and the drive passed another gate after that, all unattended overnight. The backup itself refused to start that night: Time Machine was asked to mount the image by itself a second after the drive had been remounted, and returned at once with “failed to mount destination”. With the image mounted first, the next morning’s run started normally.

That run taught me one more thing about interrupted backups. Time Machine resumes only its most recent interrupted attempt, and it carries that one folder forward from run to run by renaming it. The folder I had deleted was the head of a chain going back weeks, and it held the only copies of several trees that the last completed backup never had. So the morning run had to copy those trees again, about 230 GB, before it reached the parts it could link. Deleting the newest interrupted folder is not free; deleting older ones is.

The run was an incremental deep scan like the night before, and by the end its own counters settled the question of whether anything was being copied needlessly:

$ log show --info --last 3m --predicate 'process == "backupd" AND subsystem == "com.apple.TimeMachine"' | grep -E "Total Items"
	     2489921 Total Items Added (l: 232.63 GB p: 231.11 GB)
	     7291124 Total Items Propagated (recursive) (l: 535.52 GB p: 440.32 GB)
	     9781045 Total Items in Backup (l: 768.16 GB p: 671.43 GB)
ShellScript

Three files linked for every one written. It ended once with “device locked” after five hours, because a handful of files inside app containers are encrypted with keys macOS only holds while the screen is unlocked, and I had walked away; a resume with the Mac unlocked finished the job in twenty minutes, reusing everything the first pass had done. The drive took 236 GB of real backup writes over the day with zero errors and zero retries, and the gate afterwards found every pre-existing index page unchanged.

Two things Time Machine did at the end are worth knowing. It thinned the three backups from one day in September down to one, which is its normal rule and a reminder that it deletes snapshots on its own schedule. And it tried to delete every leftover interrupted folder on the volume and failed on all of them with “access denied”, because of the same no-unlink flag that had stopped my rm. It sets that flag itself and never clears it, which is how 28 such folders had accumulated since April. The fix is to clear the flag by hand once; the deletion itself then happens at the end of the next backup, which is Time Machine’s job and not mine.

Things that went wrong along the way

Besides the two big ones above:

  • A read-only audit sat in a loop for an hour while I was told it was two-thirds done. Nobody had checked when the last progress line was written.
  • A list of snapshots was passed to a tool as a single argument. It checked the first one and printed “done, 0 problems”.
  • A capture file written as a sparse file made APFS pre-allocate over a terabyte, filled the internal disk, and the Mac panicked twice in one day. Lesson: never rely on sparse files for scattered writes, and watch free space during long jobs.
  • Most time estimates were wrong, usually by a factor of two, in both directions.

None of these touched the drive, because nothing could. Every tool opened it read-only until the one step that was meant to write, and that step was always a separate script that needed my approval.

What I’d tell someone in the same spot

  • “Cannot be repaired” from First Aid means it won’t rebuild an index. It says nothing about whether your files are still there. Mine were.
  • Stop writing to the disk. Old copies of the index are your raw material, and they get reused.
  • Get the data off before attempting any repair, and verify the copy with a second read.
  • If you re-link old copies, reconcile the free queue. Apple’s checker will not do it for you, and it will tell you the volume is fine right up to the first write.
  • Rehearse a read-write mount with real writes on an overlay of the real data. Test images are not enough.
  • Keep external drives off shared hubs. Give a file server one network path.
  • Have a second copy. This whole exercise exists because I didn’t.

Tools

The tools are a patched apfs-fuse plus a set of small C++ and Python programs: an object map scanner, an offline directory lister, a metadata-preserving copier, the ownership checker, the overlay, the repair planners and the band-file patcher. They are on GitHub at xmwa/apfs-rescue , with a README that says what each one does and, in bold, that you use them at your own risk. They were written for one drive and one image, and some of them write to raw devices. Read the drive first, copy everything off, and rehearse on an overlay before any of the writers touch a real disk.

Where it ended

A week after First Aid gave up, the drive mounts, passes every check I can throw at it, has taken hundreds of gigabytes of verified writes, and holds a fresh Time Machine backup. Nothing on it was erased, and the copy made on the first day sits on another disk as the second copy I should have had all along. The drive goes back to the Mac it serves, on a direct cable this time, with a monitoring script to come.

What I would keep from the week is not the repair itself but the order of operations: read before you touch, copy before you repair, check with something that understands the free queue before you trust “appears to be OK”, and rehearse every write on something that cannot hurt the original. The tools are on GitHub for anyone in the same spot, with the warning that they were built for one drive. And the cause? The hub dongle is the suspect, the direct cable is the fix, and SMART from a machine that can read it is the one question still open.

  1. Apple Community: Could not mount “LaCie” (unable to repair minkey)iBoysoft: How to fix the invalid APFS object map error[↩]

Posted

in

,

by

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

← 🧭