Chasing the Beam: How BMAmiga's Fading Tearing Bug Died
The bug report was almost polite: “the emulator struggles with fades — from going to white to dark, it shows some kind of tearing, something alike vsync issues.” On the bench TV, a fast Workbench fade didn’t just look wrong; a horizontal seam tore across the picture and swept from bottom to top, taking about eight seconds per round trip.
Eight seconds. That number is diagnostic gold, and I’ll come back to it. But first, the ground rules of the battlefield.
The machine
BMAmiga is an OCS Amiga
emulator written from scratch — 68000 interpreter plus full chipset, with a
colour-clock-scheduled bus and slot-exact DMA arbitration — booting bare-metal
on a Raspberry Pi. No Linux underneath; the firmware loads kernel8.img and
the next thing that runs is the emulator. That matters here, because “how does
a frame get to the TV” is a question I get to answer at the metal: the
framebuffer is the Pi’s, the pacing is mine, and every wart in between is mine
to fix.
The frame loop paints 752×574 (a full overscan PAL frame, one pixel per Amiga
hires unit) into the scanout buffer, then waits for the next 20 ms deadline.
Simple. And on static screens, flawless. The fade was the tell: a fade changes
COLOR00 once per vblank, which means every pixel on screen changes every
frame. There is no content that stress-tests a presentation path harder.
First rule of emulator debugging: repro it, then measure it
You can’t photograph a moving seam with any dignity, so I built the
instrument: demo/fade.s, a bootable ADF whose entire payload is a grey
staircase — COLOR00 stepping one level per vblank, borders and all, no
bitplanes, no copper, nothing that can lie to you. Any non-uniform row in a
captured frame is signal. Next to it, tools/fadecheck.py, a row-level census
that classifies frames as uniform, banded (with boundary position), or other,
and stats the staircase.
Host first — 90% of BMAmiga debugging happens on the host, where the same core runs natively and you can trace every custom-register write with a beam position stamped on it. Verdict: 698 uniform frames out of 700. The emulator applies vblank palette writes exactly where it should. The core was innocent. (The kick13 host boot tried to interfere with its own recoverable-alert drama, but that’s another story; kick31 boots clean.)
So the tear lived on the Pi side: in the space between “Denise painted row R” and “the TV lit row R”. Which is to say: in time.
The TV that lies about being 50 Hz
A seam that sweeps means painter and scanout run at different frequencies: each frame, the paint phase slips a little against the beam, and the scanout catches the framebuffer mid-update at a slowly rotating row. The sweep rate is the frequency difference made visible.
Time to measure the enemy. I added a boot-time probe: two back-to-back
WaitForVerticalSync mailbox calls — the wait’s own latency cancels
wait-to-wait, so the gap is the display’s true frame period.
VSYNC: period(us)=020046
The mode configured as “50 Hz” runs at 49.885 Hz. The firmware’s CVT-generated timing lands 0.23% off nominal — 46 µs per frame. Against a 50.000 Hz painter, that’s a phase drift of 46 µs/frame, and 574 rows of visible frame at ~62.7 µs/row gives you… a full-screen sweep every ~400 frames. Eight seconds. The bug report had contained the measurement all along.
The fixes that didn’t survive contact
What do you do with two clocks that disagree? You either sync one to the other, or you make the swap atomic. I tried the atomic one first, because it’s the textbook answer.
Page-flipping. Request the framebuffer with double virtual height, render
into the hidden half, SetVirtualOffset to swap. A latched swap can’t tear —
that’s the whole religion of page-flipping. Except: this Pi 4 firmware applies
SetVirtualOffset immediately. No vblank latch. Flip mid-scan and you get a
seam at a fixed row, every frame, forever. The bench verdict — “it looks even
worse” — was not a regression I could argue with. (It took me an embarrassingly
long afternoon of pixel forensics to prove the flip had even run, but more on
the build system shortly.)
Waiting for vblank before each swap, then. WaitForVerticalSync is
edge-based, so: wait, swap in blanking, done — except the mailbox round trip
costs ~10 ms and jitters. Gate a 2 ms blanking window with a ±4 ms fuse.
Frame rates of 25, then 37 fps, with the swap landing mid-scan anyway. The
mailbox is a fine instrument for slow questions and an unusable clock.
Fixing the mode. If the display ran exactly 50.000 Hz, painter and scanout
would share a crystal and everything would be phase-stable. I computed an
explicit hdmi_timings — 864×625 total pixels at exactly 27.000 MHz is
precisely 50 Hz, arithmetic, not negotiation. The probe confirmed
period(us)=019999. The TV hated every line of it: the picture rolled, black
bars flickered, the drive-LED strip surfaced a quarter of the way down the
screen like a drowned thing. Retired. (The code and the exact numbers stay in
config.txt as a comment, waiting for a television with better manners.)
I should confess the two ways the hunt poisoned itself, because they’re now landmines in the HANDBOOK and they’d poison yours too:
- The capture card was gaslighting me. The HDMI grabber’s scaler swaps source frames at a fixed internal line whenever content changes — a stationary “seam” at row ~175 in every census, flip or no flip, identical across kernels whose timing I had wildly changed. The control that cracked it: freeze the picture (open the menu) and the seam vanished. An instrument must pass a negative control before it’s allowed to have opinions.
- The build system sold me the same binary twice. Circle’s application
objects compile into the source tree (
src/circle/kernel.o), and neithermake cleantargets nor CFLAGS changes rebuild them. Several “flag A vs flag B” experiments that day were one binary being compared with itself. The ritual is nowfind src -name '*.o' -deletebefore any flag-switched build, and never trusting a build whose md5 didn’t move.
CHASE: don’t fight the beam, outrun it
Which leaves the option that doesn’t need the firmware’s permission.
The beam is a fixed, physical thing: it walks the frame top-down, once per 20.046 ms, whether I talk to it or not. The framebuffer is mine to write any time. So: start the paint inside the vertical blanking — and paint faster than the beam scans. The painter leads; the beam reads finished rows; no swap, no latch, no offset, nothing for the immediate-apply firmware to sabotage.
The numbers make it comfortable. The paint takes ~16 ms for all 574 rows (~28 µs/row); the beam spends ~18.5 ms on active scan (~32 µs/row). Start the paint ~1.5 ms after the vblank edge — a good 1.7 ms before active scan begins — and the lead doesn’t just hold, it grows by ~4 µs every row. The beam crosses the finish line ~4 ms after the painter is done admiring their work.
Two problems remained, both solved with modest machinery:
Where’s the edge? The mailbox can’t time it, but it can measure it: the
edge is predicted on the local microsecond timer and re-calibrated by one
WaitForVerticalSync every 8th frame (latency subtracted, drift absorbed).
The frame period itself is measured at boot — the real one, 20046 µs, not the
nominal one — and the prediction advances by that. This is genlock with the
display as master, built from two mailbox calls and arithmetic.
The cache was the final boss. The first CHASE build still showed small
horizontal lines in the upper two-thirds — faint, sometimes tinged with
colour, at random positions. Classic one-frame-stale reads: the framebuffer is
mapped cacheable (uncached paint costs 15-25 ms/frame; cached, ~0.1), and the
batched cache-flush after painting began around row ~440 — every row the beam
had already passed was read from DRAM still holding the previous frame’s
data. The fix is the pattern the render-worker core had used all along:
flush each row the moment it’s painted, from the row hook, one
CleanDataCacheRange per scanline. The beam now trails a continuously
freshening framebuffer by nearly two milliseconds at its closest.
The frame loop reads like the design now:
wait for predicted vblank edge (+1.5 ms, audio FIFO pumped all the way)
paint 574 rows top-down, flushing each row as it lands
re-anchor the edge prediction every 8th frame
Bench verdict, in order: picture correctly positioned, no bars. The sweep: gone — painter and beam now share a clock. The artefacts: gone. fps 50. The fade demo runs its grey staircase for hours, smooth as an oscilloscope trace.
Why this is my favourite kind of bug
Nothing here needed a faster CPU or a bigger buffer. It needed knowing where the beam is — the same knowledge every Amiga demo coder from 1989 had to develop, learned here again from first principles on hardware that doesn’t officially expose it. The emulator’s chip-level timing model meant the guest side was exonerated in an afternoon; the war was entirely in the last 20 milliseconds between my framebuffer and your retina.
That’s the pitch for BMAmiga, really. A from-scratch OCS Amiga — cycle-scheduled bus, slot-exact DMA, raw-MFM floppies, Gayle IDE — running bare-metal on a Raspberry Pi, where “vsync” is not a flag you set but a raster you learn to chase. When the machine tears, you get to find out why, all the way down to the pixelvalve.
The chase is committed
(02404e5), along with the
fade repro, the census tool, and two new landmines so the next ghost hunt
starts warmer. The beam is still out there, sweeping. We just paint faster
than it does.