JESVS

Chasing the Beam: How BMAmiga's Fading Tearing Bug Died

The bug report was almost polite: “the emulator struggles with fades — from going to white to dark, it shows some kind of tearing, something alike vsync issues.” On the bench TV, a fast Workbench fade didn’t just look wrong; a horizontal seam tore across the picture and swept from bottom to top, taking about eight seconds per round trip.

Eight seconds. That number is diagnostic gold, and I’ll come back to it. But first, the ground rules of the battlefield.

The machine

BMAmiga is an OCS Amiga emulator written from scratch — 68000 interpreter plus full chipset, with a colour-clock-scheduled bus and slot-exact DMA arbitration — booting bare-metal on a Raspberry Pi. No Linux underneath; the firmware loads kernel8.img and the next thing that runs is the emulator. That matters here, because “how does a frame get to the TV” is a question I get to answer at the metal: the framebuffer is the Pi’s, the pacing is mine, and every wart in between is mine to fix.

The frame loop paints 752×574 (a full overscan PAL frame, one pixel per Amiga hires unit) into the scanout buffer, then waits for the next 20 ms deadline. Simple. And on static screens, flawless. The fade was the tell: a fade changes COLOR00 once per vblank, which means every pixel on screen changes every frame. There is no content that stress-tests a presentation path harder.

First rule of emulator debugging: repro it, then measure it

You can’t photograph a moving seam with any dignity, so I built the instrument: demo/fade.s, a bootable ADF whose entire payload is a grey staircase — COLOR00 stepping one level per vblank, borders and all, no bitplanes, no copper, nothing that can lie to you. Any non-uniform row in a captured frame is signal. Next to it, tools/fadecheck.py, a row-level census that classifies frames as uniform, banded (with boundary position), or other, and stats the staircase.

Host first — 90% of BMAmiga debugging happens on the host, where the same core runs natively and you can trace every custom-register write with a beam position stamped on it. Verdict: 698 uniform frames out of 700. The emulator applies vblank palette writes exactly where it should. The core was innocent. (The kick13 host boot tried to interfere with its own recoverable-alert drama, but that’s another story; kick31 boots clean.)

So the tear lived on the Pi side: in the space between “Denise painted row R” and “the TV lit row R”. Which is to say: in time.

The TV that lies about being 50 Hz

A seam that sweeps means painter and scanout run at different frequencies: each frame, the paint phase slips a little against the beam, and the scanout catches the framebuffer mid-update at a slowly rotating row. The sweep rate is the frequency difference made visible.

Time to measure the enemy. I added a boot-time probe: two back-to-back WaitForVerticalSync mailbox calls — the wait’s own latency cancels wait-to-wait, so the gap is the display’s true frame period.

VSYNC: period(us)=020046

The mode configured as “50 Hz” runs at 49.885 Hz. The firmware’s CVT-generated timing lands 0.23% off nominal — 46 µs per frame. Against a 50.000 Hz painter, that’s a phase drift of 46 µs/frame, and 574 rows of visible frame at ~62.7 µs/row gives you… a full-screen sweep every ~400 frames. Eight seconds. The bug report had contained the measurement all along.

The fixes that didn’t survive contact

What do you do with two clocks that disagree? You either sync one to the other, or you make the swap atomic. I tried the atomic one first, because it’s the textbook answer.

Page-flipping. Request the framebuffer with double virtual height, render into the hidden half, SetVirtualOffset to swap. A latched swap can’t tear — that’s the whole religion of page-flipping. Except: this Pi 4 firmware applies SetVirtualOffset immediately. No vblank latch. Flip mid-scan and you get a seam at a fixed row, every frame, forever. The bench verdict — “it looks even worse” — was not a regression I could argue with. (It took me an embarrassingly long afternoon of pixel forensics to prove the flip had even run, but more on the build system shortly.)

Waiting for vblank before each swap, then. WaitForVerticalSync is edge-based, so: wait, swap in blanking, done — except the mailbox round trip costs ~10 ms and jitters. Gate a 2 ms blanking window with a ±4 ms fuse. Frame rates of 25, then 37 fps, with the swap landing mid-scan anyway. The mailbox is a fine instrument for slow questions and an unusable clock.

Fixing the mode. If the display ran exactly 50.000 Hz, painter and scanout would share a crystal and everything would be phase-stable. I computed an explicit hdmi_timings — 864×625 total pixels at exactly 27.000 MHz is precisely 50 Hz, arithmetic, not negotiation. The probe confirmed period(us)=019999. The TV hated every line of it: the picture rolled, black bars flickered, the drive-LED strip surfaced a quarter of the way down the screen like a drowned thing. Retired. (The code and the exact numbers stay in config.txt as a comment, waiting for a television with better manners.)

I should confess the two ways the hunt poisoned itself, because they’re now landmines in the HANDBOOK and they’d poison yours too:

  • The capture card was gaslighting me. The HDMI grabber’s scaler swaps source frames at a fixed internal line whenever content changes — a stationary “seam” at row ~175 in every census, flip or no flip, identical across kernels whose timing I had wildly changed. The control that cracked it: freeze the picture (open the menu) and the seam vanished. An instrument must pass a negative control before it’s allowed to have opinions.
  • The build system sold me the same binary twice. Circle’s application objects compile into the source tree (src/circle/kernel.o), and neither make clean targets nor CFLAGS changes rebuild them. Several “flag A vs flag B” experiments that day were one binary being compared with itself. The ritual is now find src -name '*.o' -delete before any flag-switched build, and never trusting a build whose md5 didn’t move.

CHASE: don’t fight the beam, outrun it

Which leaves the option that doesn’t need the firmware’s permission.

The beam is a fixed, physical thing: it walks the frame top-down, once per 20.046 ms, whether I talk to it or not. The framebuffer is mine to write any time. So: start the paint inside the vertical blanking — and paint faster than the beam scans. The painter leads; the beam reads finished rows; no swap, no latch, no offset, nothing for the immediate-apply firmware to sabotage.

The numbers make it comfortable. The paint takes ~16 ms for all 574 rows (~28 µs/row); the beam spends ~18.5 ms on active scan (~32 µs/row). Start the paint ~1.5 ms after the vblank edge — a good 1.7 ms before active scan begins — and the lead doesn’t just hold, it grows by ~4 µs every row. The beam crosses the finish line ~4 ms after the painter is done admiring their work.

Two problems remained, both solved with modest machinery:

Where’s the edge? The mailbox can’t time it, but it can measure it: the edge is predicted on the local microsecond timer and re-calibrated by one WaitForVerticalSync every 8th frame (latency subtracted, drift absorbed). The frame period itself is measured at boot — the real one, 20046 µs, not the nominal one — and the prediction advances by that. This is genlock with the display as master, built from two mailbox calls and arithmetic.

The cache was the final boss. The first CHASE build still showed small horizontal lines in the upper two-thirds — faint, sometimes tinged with colour, at random positions. Classic one-frame-stale reads: the framebuffer is mapped cacheable (uncached paint costs 15-25 ms/frame; cached, ~0.1), and the batched cache-flush after painting began around row ~440 — every row the beam had already passed was read from DRAM still holding the previous frame’s data. The fix is the pattern the render-worker core had used all along: flush each row the moment it’s painted, from the row hook, one CleanDataCacheRange per scanline. The beam now trails a continuously freshening framebuffer by nearly two milliseconds at its closest.

The frame loop reads like the design now:

wait for predicted vblank edge (+1.5 ms, audio FIFO pumped all the way)
paint 574 rows top-down, flushing each row as it lands
re-anchor the edge prediction every 8th frame

Bench verdict, in order: picture correctly positioned, no bars. The sweep: gone — painter and beam now share a clock. The artefacts: gone. fps 50. The fade demo runs its grey staircase for hours, smooth as an oscilloscope trace.

Why this is my favourite kind of bug

Nothing here needed a faster CPU or a bigger buffer. It needed knowing where the beam is — the same knowledge every Amiga demo coder from 1989 had to develop, learned here again from first principles on hardware that doesn’t officially expose it. The emulator’s chip-level timing model meant the guest side was exonerated in an afternoon; the war was entirely in the last 20 milliseconds between my framebuffer and your retina.

That’s the pitch for BMAmiga, really. A from-scratch OCS Amiga — cycle-scheduled bus, slot-exact DMA, raw-MFM floppies, Gayle IDE — running bare-metal on a Raspberry Pi, where “vsync” is not a flag you set but a raster you learn to chase. When the machine tears, you get to find out why, all the way down to the pixelvalve.

The chase is committed (02404e5), along with the fade repro, the census tool, and two new landmines so the next ghost hunt starts warmer. The beam is still out there, sweeping. We just paint faster than it does.

← volver a posts
↑