JESVS

Multitasking Through Disk Time: The Hard Drive Fix That Unfroze BMAmiga

There’s a particular way a retro machine betrays its age. Not by being slow — slowness is period-correct — but by being slow wrongly. My bug report read: “HD access is very slow, and it slows down the whole machine while it’s accessing. This isn’t the case on a real Amiga (multitasking!).”

That parenthesis was the whole bug report, really. On a real A600, you open a drawer and the machine keeps living: the driver task issues a command, goes to sleep on the drive’s interrupt, and every other task keeps running while the platter spins. On BMAmiga, the machine visibly stopped. Not stuttered — stopped, for as long as the disk was busy. If you’re new here: BMAmiga is a from-scratch OCS Amiga — 68000 interpreter plus full chipset — booting bare-metal on a Raspberry Pi, no OS underneath. The hard drive is a Gayle-IDE emulation backed by a hardfile image on the SD card.

What follows is the round that fixed it: a diagnosis that turned out to be a philosophy problem, a fix copied from WinUAE’s homework, one spectacular livelock detour, and a benchmark session conducted by typing shell commands into a Pi across a LAN with nobody in front of the TV.

The bug was a timing model

Chasing it through the code took three passes, and they all converged on the same paragraph. When the emulated 68k writes the ATA command register — MOVE.B #$20, $A0001C, “read sector” — the emulator handled it like this:

command write
  └→ task file decode
       └→ start transfer
            └→ fill read-ahead window
                 └→ hdf_read()            ← blocking SD card read
                      └→ Circle EMMC driver (busy-polled, PIO, word by word)

All of it, synchronously, inside the guest’s own instruction. The command write instruction did not return to the 68k until the sectors were in RAM. That has two consequences, and they compound:

Zero emulated time elapsed. No scanlines, no colour clocks — the scheduler’s clock didn’t move while the card answered. From the Amiga’s point of view, the drive was infinitely fast: status reads showed a finished command instantly, DRQ and all.

Unbounded host time was stolen. On the Pi, each fetch is a real EMMC transaction — and Circle’s driver does a card-status handshake on every call, then copies the data word-by-word through a FIFO register in a busy-poll loop. Measured later on the bench: 560 µs for a single 512-byte fetch, up to 16 ms for a bulk one. Every one of those microseconds happened inside the frame’s walk — inside chipset_frame(), with the audio pump, the input drain, everything parked behind it.

And here’s the twist that makes it a whole-machine slowdown rather than a disk slowdown: the frame loop paces to 50 Hz with audio as the master clock. When a frame overruns because the walk was busy waiting on a memory card, the pacer doesn’t re-sync to wall time — it repays the deficit by running frames back-to-back. The stolen milliseconds thus propagate past the frame they stalled. The user sees the entire machine running at two-thirds speed for as long as the HD is busy. Both reported symptoms — “HD is slow” and “machine is slow” — were the same stolen milliseconds, viewed from two angles.

The deepest irony: the guest never slept. On real hardware the driver sleeps on INTRQ while the drive seeks; other tasks run. Here, completion was instantaneous in emulated time, so there was nothing to sleep through — the guest was always awake, and the host ate the entire latency instead. The emulator had faithfully implemented a drive faster than any disk ever made, and the machine suffered for it.

Preparation and publication

WinUAE, as usual, had already solved this years ago and left the answer lying in its source: when the guest writes a command, WinUAE sets BSY, does the file I/O on a worker thread, and publishes DRQ/INTRQ one or two scanlines later through a per-scanline hook. The emulated timeline never blocks; any task that isn’t polling BSY keeps executing. That’s the real machine’s multitasking, and it’s made of two separable ideas:

Deferred publication. A command now prepares at the command write and publishes two scanlines later from a new ide_line() hook, called by the scheduler walk beside the floppy’s per-line tick. BSY is up in between, so a status poller waits and an interrupt waiter sleeps — exactly as on hardware, because that two-scanline window is where real drive time lives.

Deferred I/O. Preparation shouldn’t touch the card either. The platform now registers an async backend: a read submits one bulk fetch of the whole command into the sector buffer, and a write submits one bulk store when the data phase ends. The only thing the walk does per scanline is poll a flag. Actually moving the bytes is plat_ide_pump()’s job, and it runs from the frame loop’s slack — once at the top of every frame with a ~1.5 ms budget, and once per iteration of the pace loop, which is otherwise dead time anyway. Host latency stops being stolen time and becomes emulated BSY scanlines the guest sleeps through. Small reads (≤ 8 sectors) still execute inline at submit: drawer walks are latency-critical, and one bounded transaction was the historical stall profile anyway.

If the pump ever starves — a dead card, or a scene that leaves no slack — a timeout fails the command with ERR|ABRT after ~200 ms instead of hanging the guest. A drive that reports errors is a drive the OS knows how to retry.

There was also a pure cache win hiding in the same code: the read-ahead window was reset on every command, so Workbench’s one-sector directory reads each paid a full card round trip. The window now lives on absolute image addresses and carries across commands — seventeen sequential one-sector commands cost three fetches, not seventeen. That one is pinned by a test, as is the fix’s little brother: the old window-validity check refetched a sector at every window boundary, a quiet 12% tax.

The detour: a one-line bug, three years deep

The first bench of the new code booted to a dead machine. Not slow — dead, parked inside the ROM’s scsi.device at a level-2 interrupt that never ended. The serial heartbeat told the story in one register: INTREQ had the PORTS bit pinned, the INTENA master bit was off, and the CPU was executing RTE → autovector → RTE → autovector, forever. An interrupt storm.

The root cause was beautiful. Since the beginning, our Gayle emulation had carried a “direct wire”: the IDE interrupt line forced 68k level 2 past the INTENA master bit, on the theory that Gayle’s IRQ goes straight to the CPU and Paula isn’t in that path. With synchronous commands this was harmless — the line only ever rose while the driver was actively running, and its next status read dropped it. But deferred publication raised INTRQ two scanlines after the command write — and the ROM’s driver had, by then, entered a section with interrupts disabled. Line up, master down, handler exits without servicing, autovector re-fires. Livelock, on demand.

The reference settled it in minutes: WinUAE delivers the Gayle IDE interrupt through INTREQ PORTS like any Paula interrupt — master-gated (safe_interrupt_set, for those following along at home). No wire. We deleted the bypass, the PORTS mirror with its follow-the-line semantics was already the complete and correct delivery, and the machine booted to Workbench off the hardfile. Every gate we have — the 1.3 boot that must stay byte-identical, the walk-equivalence harness across three disks and with a hardfile attached — passed clean.

The lesson I’m keeping: a divergence from the reference that survives for years isn’t correct, it’s unreachable. The first change that reaches it turns it into a wedge. Deferred timing was that change.

Numbers, or it didn’t happen

The whole round was built measurement-first: the hardfile backend now reports per-second count, total microseconds and worst single fetch on the 1 Hz heartbeat (hd 24/128432us/15919max ...), which meant every claim below was read off the live machine, not asserted.

Boot from HD. The heaviest second of the whole boot — RDB parse, FFS mount, Workbench loading — moves 24 fetches totalling 128 ms of SD time, worst single fetch 15.9 ms. Frames: 50/50. Emulated-frame cost: 5 ms of the 20 ms budget. That same second on the old build spent its 128 ms inside the walk — six and a half frames’ worth of machine time, gone, every second of the way up.

Interactive use. Opening drawers, listing directories: invisible. 50 fps, ≤ 6 ms. Sustained reads peaked at 243 ms/s of card traffic at 47–50 fps.

The copy test is where honesty is required. I typed a scripted suite into the bench’s AmigaShell — over the network, through the emulator’s web console keyboard, which deserves its own post — building a 418 KB file with join and copying it volume-to-volume. The heaviest seconds dipped to 32–41 fps. The heartbeat’s CPU-phase field showed 14–20 ms of interpreter time: the emulated 68k was genuinely busy, because on this architecture PIO is CPU work — the 68000 moves every disk word through the task file, and a same-volume copy is two full data streams plus filesystem bookkeeping through it. A real A600 feels a copy too. What was left of host-side theft was the SD card’s write path: this card programs a 4 KB chunk in ~60 ms (~66 KB/s effective — a Circle driver question, batched multi-block writes are the next lever), and the pump rides that as best it can.

For contrast: the old build ran all of that card time inline. The same copy was multi-hundred-millisecond freezes stacked on the same CPU load. Slow during a copy is a real A600. Frozen is a bug. We’re at real A600.

The shape of the fix

The commit that landed is modest — a few hundred lines across the core, the Pi platform layer and the docs — but the idea is worth keeping in your pocket for any emulator, or any soft-realtime system:

Never do real work inside emulated time. Emulated time must be spendable only on the guest’s own execution. Host I/O belongs between frames, and its latency should surface to the guest as state it can wait on — BSY, a flag, an interrupt — not as time stolen from the scheduler.

The Amiga knew this in 1985. Its driver sleeps, the disk interrupts, the machine lives. It only took an emulator to make me appreciate the design.

Next on the list: that EMMC write path, and teaching the pump a second core for the builds that have one to spare. The hardfile no longer stops the machine — now it just needs to stop being slow.

← volver a posts
↑