Skip to content

Compile whole VU1 programs and run sound IRX modules natively (DQ8 fork) - #254

Draft
Sinan-Karakaya wants to merge 73 commits into
ran-j:mainfrom
Sinan-Karakaya:main
Draft

Sinan-Karakaya wants to merge 73 commits into
ran-j:mainfrom
Sinan-Karakaya:main

Conversation

@Sinan-Karakaya

Copy link
Copy Markdown
Contributor

Draft, to show what my fork has been doing for Dragon Quest VIII and to see how you'd like it brought in. With it the game reaches its 30 FPS cap in the field on an M1 Pro, with sound.

The main pieces:

  • A VU pipeline model with compiled regions and loops, and on top of it whole VU1 programs compiled from microcode recorded while the game runs (NEON and SSE2 fast paths, interpreter fallback, differential tests against the interpreter).
  • The game's own sound modules (LIBSD, SDRDRV, the Standard Kit, PCMPLAY) running unmodified on a small emulated IOP with both SPU2 cores, in ps2xIOP/src/lle.
  • An ordered GS worker thread, movie audio kept in stream order in the MPEG stub, and scheduler and memory fixes found along the way.

It also carries #205 to #209, which can go in on their own.

It predates #244 and conflicts with it in 20 files. The IOP part overlaps with the new emulator, so rather than keep two I'd move the SPU2 and the kernel calls these drivers need onto yours. CI will be red as well: MSVC can't build the fork's runtime yet (ucontext.h, weak symbols) and the GCC job fails at an LTO link step. Happy to split all of this up in whatever order suits you.

Sinan-Karakaya and others added 30 commits August 17, 2026 19:10
On an ELF with no symbols and no DWARF, parse() carves functions from JAL
targets. Those carvings end at the next JAL target or, for the last one in a
region, at the end of the code section, so on a single-PROGBITS executable
they can run straight through interleaved rodata.

loadGhidraFunctionMap() appended its rows to the same vector and then purged
auto-named entries only where no map row shared the start address. Since both
the carvings ("sub_") and the names Ghidra exports by default ("FUN_") count
as auto-generated, a carving that shared a start with a map row survived the
purge and then won the "larger end" tie-break, so the imprecise bounds
replaced the ones the map had just supplied.

Collect the map rows into a local vector, drop every auto-named carving once
the map has parsed, and append the rows afterwards. Entries named from
symbols or DWARF are unaffected.

On a 3 MB Metrowerks-built PS2 executable with an 11,491-row map, 5,613
functions (48.8%) had been emitted with inflated bounds; the worst grew from
368 bytes to 0x51 KB and produced 22 MB of C++ decoding string data as
instructions. Output for that function is now 19 KB and total output drops
from 235 MB to 180 MB.
Not every toolchain points e_entry at an instruction. Metrowerks CodeWarrior
for PS2 emits a crt0 data table there -- scratchpad addresses and size words --
with the first real instruction some way past it, so no function covers the
entry address and neither entryName nor getFunctionName() resolves.

The emitter threw in that case, which aborted the run after every per-function
source had already been written but before register_functions.cpp,
ps2_recompiled_functions.h and ps2_recompiled_stubs.h were generated, leaving
an output directory that looks complete and is not.

Skip the entry registration with a warning instead. Synthesizing a name would
emit a table reference to a definition that was never generated and fail at
link time, and the start address for such a binary has to come from
configuration regardless.
BEQ and BNE already compare the full GPR, but BLEZ, BGTZ, BLTZ and BGEZ (and
their likely/and-link variants) were emitted against the low word only. The
R5900 compares the whole 64-bit register, so any value whose upper half is
significant takes the wrong branch.

Compilers reach these opcodes through the dsll32/dsra32 sign-extension idiom,
which leaves a canonical value and hides the bug; code that keeps a genuine
64-bit quantity in the register does not.
ps2xRecomp and ps2xAnalyzer reached ps2xRuntime through CMAKE_SOURCE_DIR,
which is the top-level source directory of whatever build is running. That
holds only when this repository is itself the top level; adding it to another
project with add_subdirectory() made both components look for
ps2xRuntime/cmake/ReleaseMode.cmake under the consuming project and fail at
configure time.

Use CMAKE_CURRENT_SOURCE_DIR-relative paths, as ps2xRecomp already does for
its ps2xRuntime include directory and ps2xRuntime does for its own cmake
include. Note that each component declares its own project(), so
PROJECT_SOURCE_DIR is not an alternative here.
When a computed jump cannot be resolved to a jump table, every instruction in
the function becomes an entry point, because the jump could land on any of
them. That fallback was also applied to JALR, which is not a jump but a call:
it transfers control to another function and returns to the instruction after
the delay slot. That return address is already queued as a resume target a few
lines above, so nothing else in the function needs to be reachable from
outside.

Indirect calls are ordinary code -- function pointers, virtual dispatch,
callbacks -- so the fallback fired constantly. On a 3 MB PS2 executable, 2,210
of the 2,422 unresolved sites were JALR, and 1,078 of the 1,282 affected
functions contained no unresolved jump at all.

Restrict the fallback to JR. Promoted entries drop from 189,876 to 1,688,
registered table entries from 156,783 to 75,386, the generated registration
file from 13 MB to 6 MB, and total output from 180 MB to 163 MB. Every
indirect call site in real code keeps its return-address resume entry (the
only sites that lose one are bogus functions carved out of rodata, where the
address is outside the function anyway).
A syscall can hand control back to the scheduler before the instruction
after it runs. SetSyscall lets the guest install its own handler for a
syscall number; dispatchSyscallOverride then suspends the calling
thread and queues that handler as a GuestInvocation. When the
invocation finishes, EeScheduler resumes the parent thread at the
address the generated code stored just before calling handleSyscall --
the instruction right after the syscall.

The analyzer never marked that address as an entry point. It queues
resume entries for JAL and JALR only, so no generated function could be
re-entered there, EeScheduler's hasFunction() check failed, and the
thread was made dormant instead of resumed. The thread simply stops;
because the scheduler then drains normally and run() returns, it looks
like a clean shutdown rather than a fault, which makes it awkward to
recognise.

This is reachable during early boot on a real title. Dragon Quest VIII
hits it in crt0: the Metrowerks startup code installs a handler for
syscall 0x83 and immediately issues it, and execution ends there,
roughly ten functions into the binary.

Note the offset is +4, not the +8 used for JAL and JALR -- syscall has
no delay slot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dispatchSyscallOverride returns bool, and its caller uses that to
decide whether the built-in handler still needs to run. The success
path -- the one that actually queues the guest's handler as an
invocation -- fell off the end of the function without returning.

That is undefined behaviour, and the practical failure mode is bad:
whatever happened to be in the return register decided whether the
built-in syscall ran in addition to the guest's override, so a game
that overrides a syscall could get the effect applied twice, or not at
all, depending on the build.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ps2_runtime.h includes <smmintrin.h> unconditionally, and the
recompiler emits SSE4.1-only intrinsics (_mm_blendv_ps and friends) for
the COP2 and FPU select idioms. SSE4.1 is therefore a hard requirement
of the codebase, not a tuning option.

Nothing sets it for GCC or Clang. MSVC does not need it -- its
intrinsics are not gated behind a target feature -- and the only place
any x86 feature is named is /arch:AVX2 inside EnableFastReleaseMode,
which is MSVC-only and applies to Release and RelWithDebInfo only. GCC
and Clang default to the plain x86-64 baseline, which is SSE2, so on a
stock Linux or macOS toolchain the affected translation units fail with
"always_inline function ... requires target feature 'sse4.1'".

Adds EnableX86SimdBaseline and applies it to ps2_runtime in every
configuration rather than in EnableFastReleaseMode, since a Debug build
needs it just as much. PUBLIC, so recompiled game code linking against
ps2_runtime inherits it -- that code is where most of the SSE4.1
intrinsics actually are. Guarded on the target processor so ARM builds
are unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
logDeci2Text stops after 256 lines and then silently drops everything
else. That is a sensible default against a game that spams kputs, but
it makes the runtime unusable as the candidate side of a boot-trace
diff: real captures run to many thousands of lines, and a silent
truncation partway through looks exactly like the guest stopping.

Adds PS2X_DECI2_LOG_LIMIT, with 0 meaning unlimited. Unset or
unparseable keeps the existing 256.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Guest code reads COP0 Status to decide whether interrupts are enabled,
and two different bits are involved:

  IE  (bit 0)   the architectural MIPS interrupt enable, set once by the
                kernel during boot and normally left set.
  EIE (bit 16)  the EE-specific enable that `ei` and `di` toggle.

We never execute the boot ROM, so nothing was setting either one, and
R5900Context started with Status at zero.

That is not cosmetic. libkernel's StartThread opens with
`mfc0 Status; xori 1; andi 1` and bails out with -1 when IE is clear --
its "you must call iStartThread from an interrupt handler" guard. With
Status at zero that guard fired every time, so every StartThread failed.
Dragon Quest VIII hits this during boot: it creates its CD streaming
thread, gets -1, prints "Can't start thread for streaming." and then
deadlocks with every thread blocked and none runnable. Nothing in the
runtime logs anything, because from its point of view the guest simply
asked a question and got an answer.

EIE matters for the matching reason: DIntr reports whether it was set so
the caller knows whether to pair it with an EIntr. Starting at zero makes
DIntr always answer "already disabled", so the re-enable never happens.

Two changes, both needed:

- R5900Context's constructor now sets Status to EIE | IE rather than 0.
  BEV is deliberately left clear -- that selects the boot exception
  vectors, which is the pre-handoff state, not this one.

- PS2Runtime's constructor no longer memsets m_cpuContext. R5900Context
  already zeroes itself before applying its reset values, so the memset
  only threw those values away. Threads created later were unaffected
  because EeScheduler::startThread assigns `R5900Context{}`, which is
  why this presented as "the main thread cannot start threads" rather
  than something more obviously global.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ran-j pointed out on ran-j#211 that invokeCurrent is
[[noreturn]] -- it throws EeDispatcherTransfer to unwind back to the
scheduler -- so control never reaches the end of the function and the
return I added was dead code. He closed the PR; this drops the change so
the branch matches upstream again.
decodeFunction stops at function.end, so when the last instruction inside
the range is a branch its delay slot falls outside and the emitter
substitutes a NOP (makeSyntheticDelaySlot). A real instruction disappears
with no diagnostic.

For a jr $ra epilogue that instruction is the addiu $sp, $sp, N, so every
call to the function leaks stack. In DQ8 that was __umoddi3, called from
_strtol_r while parsing a text asset: after a few calls _strtol_r read its
own saved $ra from the wrong offset, got 0, and walked the main thread off
the end of crt0. The symptom looked like a scheduler hang.

286 functions in DQ8's map end on a branch, 165 of them where the next
mapped function starts more than one word later.
dispatchIrq seeded the handler invocation with handler.sp, which is
whatever $sp happened to be when the guest called AddIntcHandler. By the
time the handler fires that thread has long since moved on and is using
those addresses for something else, so the handler builds its frame on top
of live data.

In DQ8 the VBLANK_END handler landed inside the main thread's stack and
zeroed a saved $ra, and the thread returned to 0. The dispatcher already
does the right thing for an invocation whose $sp is 0 -- it hands out an
invocation stack -- so passing a recorded $sp was bypassing it. Same fix
for the GS VSync callback and the Alarm event, which shared the pattern.

That alone would have moved the corruption rather than removed it:
reserveAsyncCallbackStack carves invocation stacks downwards from the top
of RAM, and the EE kernel puts a game's initial stack there too (DQ8's main
thread runs at 0x01F40000..0x02000000). reserveGuestStackFromAsyncPool now
pulls the pool below any guest stack that reaches into it, called wherever
a thread's stack is recorded.
Two related seams for games that do more than the runtime assumes.

Guest-owned heap. A game shipping its own allocator (newlib and MSL both
do) grows it through sbrk/EndOfHeap and hands out addresses the runtime
must not also be allocating from. setGuestHeapCeiling makes EndOfHeap
report the guest's ceiling rather than the runtime arena's limit, and
SetupHeap no longer calls configureGuestHeap when one is set -- the guest
is describing its own heap there, and honouring it dragged the runtime
arena back on top of the guest's first chunk. guestHeapHardLimit exposes
the top of the allocatable range, which guestHeapLimit cannot report before
the heap is configured, so a caller can split the range without hardcoding
the constant.

Found in DQ8: the runtime's 128-byte GIF packet from sceGsResetGraph landed
on the guest heap's top chunk header, dlmalloc read a zero size, and its
own guard (set_head(top, PREV_INUSE), "will force null return from malloc")
did exactly that.

Overlay dispatch. The generated function table is one dense array over one
address range, so it holds one function per guest address. Titles that
stream code into a fixed arena break that. FunctionRegionResolver lets a
host that translates each image separately say which table covers an
address right now; nothing about overlay formats or loading enters the
runtime.
The VIF interpreters decoded the bit and set VIFn_STAT.INT, but nothing
turned that into an interrupt, so a game that asks to be told when its
display list completes waits forever. The acknowledge side already worked:
a VIFn_FBRST write with STC clears the STAT bits.

PS2Memory latches the request (bit 0 VIF0, bit 1 VIF1) and
EeScheduler::processPendingEvents drains it into dispatchIrq, mirroring how
EE timer interrupts already travel from advanceEeTimers to the scheduler.

DQ8 sends one such VIFcode -- 0x80000000, a VIF NOP with the i bit -- then
blocks on a semaphore its INTC cause 5 handler signals. Without this the
game stops with the title artwork loaded and never draws. With it, it
renders.

Not emulated: hardware stalls VIF1 at the interrupt and resumes on the STC
write, while the interpreter runs on as it already did when it only set
STAT.INT. DQ8 puts the i bit on the last VIFcode of the packet so there is
nothing left to stall; a game that interrupts mid-packet would need it.
SMODE2.FFMD selects what the frame buffer holds, not how the CRT scans it,
and PresentFromLocalMemory had it inverted. FFMD=0 buffers a whole frame
whose two fields are alternate rows, so weaving to a progressive host frame
is just reading every row; FFMD=1 buffers half the display height and each
field reads all of it, which is where line-doubling belongs. Applying the
FFMD=1 rule to an FFMD=0 game threw away half the vertical detail.

Also snaps the two PCRTC read circuits together when they cover the same
buffer one row apart. That configuration is a deliberate vertical flicker
filter and merging it faithfully costs half a row of blur; PCSX2 snaps by
default. This is the one deliberate departure from hardware here, and it is
isolated -- deleting it costs about 1.6 dB.

Measured on six DQ8 title-screen GS dumps: identical adjacent rows
224/447 -> 0/447, PSNR 28.97-29.12 dB -> 29.53-29.67 dB. The win holds under
both resample directions. The FFMD=1 branch has no test coverage -- every
dump is FFMD=0 -- and is reasoned from the spec rather than exercised.
Two defects that both trace back to the frame rate, and one hazard beside
them.

Input. readState() sampled IsKeyDown() on the EE thread once per scePadRead,
which is a level read of raylib state refreshed once per presented frame. At
7 fps that is a 140 ms window and a press-and-release inside it was never
observed at all.

Edge detection alone does not fix this: PollInputEvents copies
current->previous and then calls glfwPollEvents, and GLFW's callback sets the
key to 1 on press and 0 on release, so a tap whose press and release land in
the same poll leaves both states at 0. Measured, 5 ms taps fired on 0 of 51
frames for IsKeyDown and IsKeyPressed alike. What survives is keyPressedQueue,
which the callback appends to regardless. ps2PadPollHost() therefore runs on
the render thread and publishes two atomics -- a held mask from IsKeyDown and
a latch drained from GetKeyPressed -- and readState() ORs them, consuming the
latch exactly once. Taps of 5/20/60 ms went from 0/15, 3/15, 6/15 seen to
15/15 in all three cases, one guest frame of button-down each.

Declared in a new ps2_pad_host.h rather than on PSPadBackend: ps2_pad.h is
reachable from ps2_runtime.h, so putting it there rebuilds the whole
recompiled corpus for an input change.

Presentation. The window is resizable and SetTextureFilter was never called,
so a non-integer nearest-neighbour scale duplicated some pixel columns and
dropped others -- at 700x500 only 510 of 512 source columns survived, and
1920x1080 gave runs of 2 and 3 pixels. One-pixel font stems came out uneven.
Upscales now snap to floor(fitScale) with point sampling, bilinear is used
only below 1.0 where integer scaling has nowhere to go, and the destination
origin is floored because odd centring offsets were an independent half-pixel
source. Verified over 15 window sizes: 8 of 12 upscales mangled before, 0
after, all 512 columns present with uniform run widths. The window now opens
at 2x when that fits in 90% of the monitor.

Also SetExitKey(KEY_NULL): Escape is bound to Circle, and raylib's default
exit key is Escape, so pressing Circle closed the game.
… pool

Pending invocations are pushed onto the running thread between any two guest
function dispatches, including partway through one already in flight, so they
nest rather than run in sequence. DQ8's movie streaming queues SIF RPC
completion callbacks faster than they retire and reached depth 13 on one
thread, exhausting the invocation stack pool and throwing.

Past four deep the rest simply stay on m_pendingInvocations and are picked up
as the depth drops. Real nesting is at most a couple deep -- an interrupt
during a callback -- so this bounds the pathological case without changing
legitimate behaviour, and an unbounded depth cannot be supported anyway: the
pool is the fixed gap between the runtime arena and the guest's own stack.

Measured on DQ8: 16 stacks consumed then failure, all but two of them thread
12 at depths 0 through 13, all GuestInvocationKind::RpcCallback. With the
bound the game runs past the movie open into its sound settings screen.
Three defects that between them left DQ8's logo movie as a frozen frame that
never terminated. All are reachable by any game that plays a PSS movie.

1. A runtime built without FFmpeg deadlocks. The stub decoder's feed() returns
   false and the failure path sets decoderFailed = false, which is right for
   the FFmpeg build where false means "no sequence header yet, resync". With no
   decoder that is not recoverable, so sceMpegGetPicture parks on a frame that
   can never arrive. decoderFailed was declared, cleared in two places and read
   in five as the give-up signal, but never once set: the path was dead code.
   Both decoder variants now carry kAvailable and feedElementaryStream raises
   the flag. Both demux entry points include it in the completeExternalWait
   predicate, without which a thread already parked -- the usual case, since
   the first GetPicture precedes the first demux -- is never woken.

2. Nothing could report end of stream to a game that streams with plain
   sceCdRead. The two available signals are a program-end code in the PSS and
   currentCdStreamEofSeen, and the latter is only ever raised from sceCdStPause
   /StRead/StStart/StStop. DQ8 contains no sceCdSt* function at all and its
   .MVI files carry no program-end code -- they end in 0xFF sector padding --
   so the movie decoded all 270 pictures and then hung forever. sawSequenceEnd
   already tracked the video sequence-end code 00 00 01 B7 but was read only by
   a debug log; it now drives the wait predicate and the guest's own end flag.

   The flag itself is inner[0x00], reached through mpeg[0x40]. The game's
   sceMpegIsEnd equivalent is three instructions -- return **(mpeg+0x40) -- and
   the stub wrote every neighbouring field but that one.

3. mpeg[0x08] gates the caller's VRAM upload rather than counting pictures.
   DQ8 uploads its movie buffers only while it reads zero, so writing a running
   counter there stopped the movie updating after the very first frame. It now
   reports whether this call produced a picture.

Also stop sceMpegReset carrying streamEnded forward unless the sceCdSt*
producer is actually driving. appendPssBytes refuses new data while that flag
is set until cdStreamGeneration advances, and only notifyMpegCdStreamStart
advances it -- so for a plain-sceCdRead game the first ended stream wedged
every later movie.
m_invocationStackTops keys a 16 KiB stack per (thread, depth) and is only ever
find()-ed and emplace()-d, never erased. reserveAsyncCallbackStack behind it is
a bump allocator that walks m_asyncCallbackStackTop downwards with no free
path. So every guest thread that ever takes one invocation holds its stacks
forever, and the pool -- 0x01F00000 to 0x01F40000, exactly 16 slots -- only
ever drains.

DQ8 starts a fresh thread per movie and deletes the old one, so it hit the wall
on the second: "EE invocation stack space exhausted" the moment TAKA_N.MVI
opened. Instrumenting the allocation showed 14 of 16 slots held by seven live
threads before the second movie was even a factor, thread 12 alone holding four
at depths 0 through 3.

releaseInvocationStacks returns a thread's stacks to a free list when its
record is erased, from both destruction paths, and invocationStackTop pops that
list before falling back to the pool. reset() returns them too: it clears
m_threads and restarts ids at kFirstThreadId, which would otherwise strand
every existing entry permanently.

Recycling at record-erase time is safe. deleteThread requires the thread to be
Dormant already, and exitCurrent throws EeDispatcherTransfer immediately
afterwards, so unwinding completes before another invocation can claim the
slot. terminateThread deliberately does not release -- it only makes the thread
dormant and the id can be restarted.

Measured on DQ8: delete thread 9 and delete thread 13 both recycle, thread 13
reuses the freed slots, and the allocation that threw on every prior attempt
succeeds. Exhaustion count zero across a full boot.
PresentationFrame assumed the software backend's fixed 640-wide staging
buffer: callers de-strided by that constant, so a backend returning tightly
packed rows -- or rows at any other width -- was misread as skewed.

Adds rowPitchBytes, and a GSPresentationMode distinguishing host pixels from a
backend that presented through its own swapchain and has no pixel payload to
copy. GS::copyLatchedHostPresentationFrame now uses the frame's own pitch and
refuses one narrower than its width rather than reading past the end.

Needed by a hardware backend, which presents at its own internal resolution:
a 4x render target hands back 2048x1792 rows, not 640-strided ones.
PresentationFrame already had a BackendNative mode; nothing could reach it,
because every frame went out through raylib. The runtime now takes an external
presenter: with one installed it opens no window, creates no framebuffer
texture, and does no drawing. It still latches each frame -- that is the call
that reaches the backend's Present() -- and then hands control to the presenter
to service its window.

Between frames it sleeps briefly instead of spinning. raylib's frame limiter
used to block in EndDrawing(); without a replacement the loop burns a core the
EE thread wants.

Audio moves out of the windowed branch. It has no window and an external
presenter does not replace it, so initialising it only in the raylib path left
the audio backend permanently not-ready.

ps2PadPublishHostState() lets a host that is not raylib feed the same latching
path, so press edges between two guest polls still register.
mpegDemuxBackpressured() returns true once the decoder is 8 pictures ahead, and
the demux stubs then returned 0 consumed and nothing else. DQ8's producer loop
reads that as "not now" and asks again immediately, so the movie thread span:
670,171 calls to sceMpegDemuxPssRing per 30 presented frames -- about 22,300
per frame -- against 0.5 sceMpegGetPicture calls per frame of actual work.

That was the movie's frame time. Measured with the graphics side fully
accounted for: the GS backend was 6% of wall, GIF packet decode ~0%, and the
896 tile transfers per frame cost almost nothing.

Both demux entry points now rotate the ready queue before returning. This is
deliberately a yield and not a park: the comment on mpegDemuxBackpressured
warns that parking the producer can leave a consumer asleep with nobody to wake
it, which is real -- Code Veronica wakes its video thread before every demux
call. Rotating the queue lets the consumer run and comes back, so the spin
becomes a scheduling point without changing who wakes whom.

DQ8's attract movie goes from 0.8 fps to 9.1, and demux calls per frame from
~22,300 to ~1,025.

Still 1,025 spins per frame, so this is a large improvement rather than a cure.
The remaining fix is a real producer/consumer handshake -- park the producer on
space-available and wake it from sceMpegGetPicture -- which needs care for
exactly the reason the yield was chosen here.

Also adds timers on sceMpegGetPicture, writeDecodedFrameToGuest and the demux
entry points, and on the GIF frontend's packet and native-image-upload paths.
None of this was visible before; the demux call count is what identified the
bug, and it was only findable because the graphics side had already been
measured and eliminated.
The previous commit called rotateReadyQueue from the demux stub and returned
normally. That clears the current thread and requests a reschedule, so the next
syscall reached bindMainContextForSyscall with no thread running and tripped
its assert -- the game aborted on every movie. transferIfRequested must follow,
which is what the RotateThreadReadyQueue syscall does, and the return value has
to be set before it because the transfer throws.

Verified: 0 aborts across a 700-second run that plays the attract movie
through, 582 pictures served.

Also tried and reverted: parking the producer on a space-available wait and
completing it from sceMpegGetPicture. That deadlocks -- 4 pictures served
instead of 658 -- which is exactly what the comment on mpegDemuxBackpressured
predicts. Both the attempt and its result are recorded in
docs/notes/movie-decode.md so it is not tried a third time.

Honest numbers: 0.8 fps before, 1.2-1.7 fps now, demux calls halved from 7.6M
to 3.4M over the run. That is an improvement, not a fix. Demux is now 6% of
wall and the backend 6%; the other ~88% is still unattributed, and the
principled next step -- buffering PSS bytes and decoding lazily so the producer
is never refused -- has not been attempted.
… event pump

DQ8's movie producer thread is a spin loop, and correctly so -- it feeds audio
and waits for its CD ring, ~5,150 iterations per movie frame, which is about
what a real 294 MHz EE would manage. Our cost per iteration was the problem.

Three measured costs, each removed:

publishSnapshot() built and sorted three vectors of kernel state on every one
of ~40 scheduler operations, for state only the debug panel and an
aggressive-log tick ever read. 45% of EE thread time during a movie. It now
builds only when snapshot() has asked; the idle path publishes unconditionally
so a quiet scheduler still converges for a poller.

The MPEG stub refused demux bytes once 8 pictures were decoded ahead -- 134,563
refusals out of 134,569 calls per 30 frames -- and the guest's loop
(remaining -= consumed; bgtz remaining) cannot terminate on a zero count.
Demux now queues elementary stream and sceMpegGetPicture pulls decode from it,
so lookahead is bounded by buffered bytes rather than decoded pictures: ~25 KB
per frame on the wire against ~917 KB decoded. Refusals go to zero and demux
calls per frame from 4,551 to 13. flushDecoderIfEnded() has to wait for the
queue to drain, or a program-end code drains the decoder while the whole movie
is still buffered and every later packet fails with AVERROR_EOF.

processPendingEvents() ran after every guest dispatch and took three mutexes,
read the clock, walked m_deadlines and heap-allocated a deque to find nothing.
24% of EE thread time, now 6%.

Also: transferIfRequested() skips the unwind when the thread it would
reschedule is the one already executing, and counters for dispatches, transfers
and refusals so the remaining cost stays visible. DQ8_EE_DISPATCH_HISTOGRAM
names the guest addresses the executor re-enters most.

Logo movie 1.4 -> 3.3 fps. Still not smooth: 30,908 context transfers per movie
frame at ~10 us each, where 60 fps needs ~540 ns. That is the exception-based
transfer itself, and fixing it means fibers rather than throw/unwind.
…itch

The EE scheduler transferred control by throwing EeDispatcherTransfer, which
destroyed the guest's whole C++ call chain and rebuilt it on the way back --
about 2.7 executor re-entries per transfer, ~10us each, where 60fps allows
~540ns. DQ8's movie thread needs ~31,000 transfers per movie frame, so the
movie ran at 1.4fps and _Unwind was 43% of EE thread time.

EeFiber gives each GuestThread its own stack. A yield now saves the
callee-saved registers and swaps the stack pointer: 22ns per round trip,
measured. Guest dispatches per movie frame fall 82,441 -> 353 and thrown
transfers per 30 frames 927,230 -> 2; the movie runs at 7.0fps.

transferIfRequested suspends, and so does blockCurrentResumable -- semaphore,
event flag and sleep waits, which only take a return value from makeReady() and
whose callers have nothing left to do. Waits carrying a completion still throw,
because the completion re-invokes its syscall on wake and would run twice
against a stack that is still there. blockCurrent stays [[noreturn]]: giving it
a return path makes the compiler emit ud2, and waitVSync/waitExternal take an
optional completion, so a flag on blockCurrent turns those into SIGILL.

x86-64 gets a hand-written switch, everything else falls back to ucontext
(correct, but a sigprocmask syscall: 461ns against 22ns). setjmp/longjmp across
stacks is faster still and deliberately unused -- _FORTIFY_SOURCE turns it into
an abort. The trampoline realigns the stack before calling, or the first
std::string inside a fiber faults on an aligned SSE spill. Stacks are mmapped
with a guard page so an overflow faults instead of corrupting its neighbour.

Also here: DQ8_PAD_SCRIPT replays input timed in guest vsync ticks so a script
lands in the same place whatever the frame rate; DQ8_GFX_SCREENSHOT_DIR writes
the presented frame as a PPM named after that same tick; DQ8_SKIP_MOVIES taps
START while a movie plays, which is the game's own skip -- reporting
end-of-stream from the MPEG stub instead leaves its streaming thread spinning
on a black screen.
…spatch

DQ8 reported no memory card in either slot because isValidMcPortSlot() demanded
slot == 0 and DQ8 passes slot=1 to every libmc call -- `addiu $a1, $zero, 0x1`
at each of its sceMcGetInfo/GetDir/Delete call sites. Everything downstream
already models a console with no multitap (one card per port, and the mc root
path ignores the slot entirely), so the validator was the odd one out. GetInfo
now answers type=2 free=8192 format=1, the no-card prompt is gone, and the game
goes on to look for its /BASLUS-21207dq8 save directory.

DQ8_MC_TRACE=1 logs the card calls. RUNTIME_LOG is compile-time and sits in a
header the whole recompiled corpus includes, so turning it on to watch one stub
would cost a full rebuild.

Two scheduler costs, found by sampling the EE thread while a movie runs:
selectReady() walked all 128 priority deques on every dispatch (~7% of EE
thread time) and now consults a one-bit-per-priority mask maintained by
refreshReadyMask(); advanceEeTimers() did up to four 64-bit divisions per
enabled timer on every guest safe point, and below one EE second the
whole-seconds term is zero and a sub-tick accumulation needs no division at
all. Movie 6.5 -> 7.2 fps for the two together.

WriteRunCT32/WriteRunZ32 write a run of pixels with the traits and lookup table
inlined, instead of a cross-TU call per pixel; measured no change, because
PixelStorageTraits is only 11% of the profile even as its top leaf. Kept as
strictly less work.

DQ8_SKIP_MOVIES now only lifts the 30fps presentation pacing. Its previous
behaviour -- tapping START -- does nothing, because DQ8's attract movies are
not button-skippable: 314 injected presses through one movie changed nothing.
ReadRowCT32/Z32/P8 mirror WriteRunCT32/Z32 for the read direction: texture
expansion walks whole rows, and the per-pixel Read entry points recompute
page, block row and column from (x, y) with three integer divisions per
texel -- 229k times for one 512x448 expand. Along a row only the in-page
column and the page column move, so the row walk is increments and the
swizzle table lookup is the only per-pixel work left. The address mask keeps
Read()'s wrap guard for regions past 4 MiB.
The movie thread yields ~4M times a second and each yield pumped events
twice. Three per-pump costs, each visible in its profile:

- The tail of processPendingEvents took m_eventMutex to ask whether any
  event was still queued. m_pendingEventCount now answers without the
  mutex; postEvent counts under the same mutex it pushes under, and the
  pump zeroes the count when it drains.
- takePendingVifInterrupts was an atomic exchange per pump; a plain load
  answers first because the pending set is almost always empty.
- advanceEeTimers walked all four timers per safe point even when the game
  never armed one; a cued-any flag maintained on timer writes gates it.

Measured on DQ8's boot-to-title: movie phases 38.8 -> 39.5-46.6 fps, the
settings screens and title stay at the 60 fps vsync cap. The remaining
per-yield cost is the fiber round trip and the syscall dispatch path.
Sinan-Karakaya and others added 30 commits September 14, 2026 10:19
Expose constant upper and lower register usage independently so native blocks fold readiness loops. Use guarded ARM64 fused vector product-sum arithmetic while retaining the scalar interpreter oracle and exact exceptional fallback. Emit deterministic external instantiations across private compilation units to shorten native rebuilds.

Verified 49 VU tests, 268 captured signatures and 788108 resumed raw-byte boundaries, plus 4112 perturbed field states. Paired heavy replays save 27-35 percent with readiness specialization beyond the SIMD stage. Seven integrated runtime suites and the generator coverage test pass. Guest cycles, callbacks and code invalidation are unchanged.
Sound drivers are the part of the IOP that high-level services reproduce
worst. Their RPC protocol is simple enough, but the sound itself comes
from how the driver programs the SPU2: envelopes, reverb, streaming
through AutoDMA, timers that pace the sequencer. Rewriting that for every
game would never end, so this runs the original modules instead.

NativeIop loads IRX files onto a small emulated IOP: an R3000A
interpreter, a high-level kernel with the calls sound drivers make
(threads, semaphores, event flags, mailboxes, alarms, hardware timers,
interrupts, the heap, SIF RPC and DMA, a little sysclib), and both SPU2
cores with their voices, reverb and 2 MiB of sound memory.

IOP time advances only with the SPU2's output, 768 cycles per 48 kHz
sample, and whoever pulls samples drives it. Work done to answer an RPC,
and interrupt, alarm and timer handlers, runs without moving the clock;
charging it to the clock made heavy RPC traffic run the drivers' timers
ahead of the audio, and music played fast.

Once a native module registers an RPC server, IopSubsystem hands requests
for that SID to it before any high-level service. A module that binds
back to an EE server can ask the host whether that server exists yet.
A runtime names the modules it wants run natively with
ps2_native_iop::setModules(). From then on:

- sceSifLoadModule of one of them also loads it on the emulated IOP. The
  high-level tracker still hands out the module id the game sees.
- IOP heap allocation and EE-to-IOP sceSifSetDma go to the emulated IOP's
  memory, which is where the modules look for what the game sends them.
- The first module to load opens a 48 kHz stereo stream whose callback
  runs the IOP. Volume 0 keeps it running silently, because games wait on
  their sound drivers whether anyone listens or not.

SDRDRV's callback thread binds to an RPC server on the EE. Answering that
bind at once left the thread spinning on a server that did not exist yet
and starved the rest of the driver, so the bind now waits until the EE has
really registered the server.

PS2X_AUDIO_DUMP=path keeps a raw copy of everything played, to compare
against a reference offline.
Two things made DQ8's movie audio sound chopped once its sound driver
actually ran.

The demux callbacks were queued on the EE scheduler and ran after
sceMpegDemuxPss() had returned. By then the caller had released the ring
space they point at, and a batch of queued invocations runs last to
first, so the audio came out shuffled. libmpeg calls them inside the
demux call, in stream order, and the stub now does too, with the return
value set first since the callbacks run before the caller resumes.

The demux also ran as far ahead as the picture queue allowed, far more
audio than DQ8 keeps: it holds a quarter second and drops what does not
fit. With an audio track, the demux now stays 150 ms of audio ahead of
what has played, measured in real time rather than in pictures, which
fall behind whenever decoding does.

With both, every chunk the game streams to the IOP during a movie matches
the movie's audio track byte for byte.
Each case pins something that was wrong at one point while getting DQ8's
drivers to play: AutoDMA blocks out of order across the DMA restarts a
driver makes for every half it refills, a stopped and restarted stream
not starting over, BVOL and AVOL swapped (a core's own input follows
BVOL), and LIBSD's halfword writes to BCR.
The block compiler covers regions of up to 16 pairs and some short
loops; everything between them runs on the interpreter a pair at a
time. In DQ8's field that per-pair work was most of what VU1 cost, and
VU1 was most of what the game thread cost.

This compiles programs whole, from the MSCAL entry to the E bit, out of
microcode recorded while the game runs (PS2_VU_PROGRAM_PROFILE). A
routine is the code reachable from one entry through branches; calls
and register jumps end it and their targets become routines of their
own. Routines are found at run time by PC and checked word for word
against VU memory, so microcode that changed falls back to the
interpreter.

Pairs still issue in order under the interpreter's stall rules. What
changes is where the work is done:

- Each pair's registers, latencies and hazards are decoded at compile
  time, and reads that can no longer stall, given the pairs before them
  in the block, are dropped from the readiness check.
- Results are written at issue and only their ready cycles are kept.
  MAC, status and clip results wait in a small queue that is folded when
  something reads them rather than every cycle.
- The clock, the flag queue and the kick state are locals, so they stay
  in registers. Whatever is still in flight when a routine returns goes
  back into the interpreter's pipelines, and a later pair or flush sees
  what it would have seen.
- On AArch64 the common FMAC cases use NEON, with range checks that send
  anything unusual to the interpreter's arithmetic. MAX, MINI, ITOF,
  FTOI, ABS and CLIP have small forms of their own everywhere: inlining
  the interpreter's whole upper switch into every pair made a routine
  take minutes to compile.

A routine only starts with at least 4096 cycles of budget. It hands back
at a block boundary when the budget would run out, or at a pair it
cannot compile.

On 768 programs captured in the field, the routines leave the same
state, VU memory and GIF packets as the interpreter at every budget the
replay checks, and run about 9 times faster on an M1 Pro (5 times on the
portable path). Compiling all of DQ8's recorded routines takes under ten
seconds. In the opening field this took the game from about 15 to about
22 completed frames per second, after which the renderer was the limit.
Game microcode cannot go into the repository, so vu_programs.py writes
synthetic programs in the layout a recording has. A few are written out
to reach particular paths (blocks split for length, an E bit on a
branch, a program handed back part-way) and the rest come from a seeded
generator mixing what the compiler accepts inside counted loops, forward
branches, calls and register jumps.

ps2_vu1_program_tests.cpp runs each of them compiled and interpreted,
whole, and then stopped by budgets the program outlasts and resumed in
uneven slices, comparing state, VU memory and GIF packets each time.
PS2_VU_PROGRAM_FIXTURES names the programs; without it the test has
nothing to run. PS2_VU_REQUIRE_PROGRAMS makes a missing routine a
failure.

Writing it turned up two bugs, fixed in the compiler itself: P could
take an older EFU result after three EFU operations in a row, and a
routine lookup could be answered from a cache entry left by another
interpreter whose code happened to sit at the same address.

test_compile_vu_programs.py covers how the generator cuts code into
routines and blocks.
The worker gave back a batch's queue space only once all of it had run.
A batch can hold most of a frame, so the producer stopped for the whole
batch and the worker then sat idle waiting for the next one. Giving back
space every 64 KiB keeps both of them busy.
processGIFPacket read the clock twice per packet for a total that only a
backend's stats report reads. A busy scene sends over ten thousand
packets a frame, so that cost a couple of percent of the game thread. A
backend that reports the numbers now sets g_gsFrontendTiming.
Both DQ8 pull requests move the submodule; this is the commit to use once both are in.
Running the callbacks inside the demux call puts them on the calling
guest thread's stack. The MPEG tests call the demux from the host, with
no guest call in progress, so there is no such stack and the scheduler
threw. Those callers get the callbacks queued, as before.
x86 hosts ran compiled routines on the plain C++ path, a little over
half as fast as NEON on the field captures. The SSE2 version takes the
same steps with the same range checks. SSE2 only has signed integer
compares, which is fine as nothing compared reaches 2^31, and movemask
gives the lanes in the opposite order to the MAC flags, hence the table.

Product sums have to round like the interpreter's acc + fs * ft, which
GCC and Clang fuse into one multiply-add when the target has one. So
the fast path fuses on AArch64 and with FMA enabled, and multiplies
then adds otherwise.

PS2X_VU_PROGRAM_SSE builds this path on AArch64 through sse2neon. Run
that way, all 768 field captures and the synthetic programs match the
interpreter, at 8 to 9 times its speed (NEON: 9, plain: 5). Reversing
the lane table or rounding product sums the other way makes both fail.
On x86-64 it has only been compiled here, with and without -mfma;
DQ8's Linux x86-64 CI job runs the synthetic programs.
A compiled routine ends at a register jump, and run() only looks the
target up, and records it if it is missing, once the routine before it
exists. A first recording is made with nothing compiled, so every program
runs on the interpreter and those targets never showed up. DQ8's main
field program starts at 0x960 and jumps on to 0x970-0x990, so a build
from one recording compiled the start and ran the rest on the
interpreter: 97% of VU1 work, about 6 FPS in the field. It took a second
recording, made with the first routines built, to find the targets.

While a profile is recorded, the interpreter now saves where JR and JALR
land, once per target and microcode generation since recordMissing hashes
the whole image. One pass along the field route compiles to the same 52
routines several passes produced before, and the field holds 30 FPS with
all of VU1 compiled.

setProfileDirectory lets the new test record into a directory of its own.
libmpeg's sceMpegCreate finishes by calling sceMpegReset, which clears
the end-of-stream word the game polls through sceMpegIsEnd. The stub
stopped short of that, so a decoder created in memory that an earlier
user had left non-zero could read as ended before its first picture.

Dragon Quest VIII creates the decoder for the movie after its first
fight in heap memory that gameplay has already used. Its decode thread
checks the end word before asking for a picture, finds it set, resets
and exits without decoding anything, and the player waits for the two
pictures it needs before starting. The screen stays black.

The new test fills the work buffer with 0xFF before sceMpegCreate and
fails without the reset.
Record VU jump targets so one recording pass is enough
Reset a new MPEG decoder the way sceMpegCreate does
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant