Compile whole VU1 programs and run sound IRX modules natively (DQ8 fork) - #254
Draft
Sinan-Karakaya wants to merge 73 commits into
Draft
Sinan-Karakaya wants to merge 73 commits into
Sinan-Karakaya wants to merge 73 commits into
Conversation
On an ELF with no symbols and no DWARF, parse() carves functions from JAL
targets. Those carvings end at the next JAL target or, for the last one in a
region, at the end of the code section, so on a single-PROGBITS executable
they can run straight through interleaved rodata.
loadGhidraFunctionMap() appended its rows to the same vector and then purged
auto-named entries only where no map row shared the start address. Since both
the carvings ("sub_") and the names Ghidra exports by default ("FUN_") count
as auto-generated, a carving that shared a start with a map row survived the
purge and then won the "larger end" tie-break, so the imprecise bounds
replaced the ones the map had just supplied.
Collect the map rows into a local vector, drop every auto-named carving once
the map has parsed, and append the rows afterwards. Entries named from
symbols or DWARF are unaffected.
On a 3 MB Metrowerks-built PS2 executable with an 11,491-row map, 5,613
functions (48.8%) had been emitted with inflated bounds; the worst grew from
368 bytes to 0x51 KB and produced 22 MB of C++ decoding string data as
instructions. Output for that function is now 19 KB and total output drops
from 235 MB to 180 MB.
Not every toolchain points e_entry at an instruction. Metrowerks CodeWarrior for PS2 emits a crt0 data table there -- scratchpad addresses and size words -- with the first real instruction some way past it, so no function covers the entry address and neither entryName nor getFunctionName() resolves. The emitter threw in that case, which aborted the run after every per-function source had already been written but before register_functions.cpp, ps2_recompiled_functions.h and ps2_recompiled_stubs.h were generated, leaving an output directory that looks complete and is not. Skip the entry registration with a warning instead. Synthesizing a name would emit a table reference to a definition that was never generated and fail at link time, and the start address for such a binary has to come from configuration regardless.
BEQ and BNE already compare the full GPR, but BLEZ, BGTZ, BLTZ and BGEZ (and their likely/and-link variants) were emitted against the low word only. The R5900 compares the whole 64-bit register, so any value whose upper half is significant takes the wrong branch. Compilers reach these opcodes through the dsll32/dsra32 sign-extension idiom, which leaves a canonical value and hides the bug; code that keeps a genuine 64-bit quantity in the register does not.
ps2xRecomp and ps2xAnalyzer reached ps2xRuntime through CMAKE_SOURCE_DIR, which is the top-level source directory of whatever build is running. That holds only when this repository is itself the top level; adding it to another project with add_subdirectory() made both components look for ps2xRuntime/cmake/ReleaseMode.cmake under the consuming project and fail at configure time. Use CMAKE_CURRENT_SOURCE_DIR-relative paths, as ps2xRecomp already does for its ps2xRuntime include directory and ps2xRuntime does for its own cmake include. Note that each component declares its own project(), so PROJECT_SOURCE_DIR is not an alternative here.
When a computed jump cannot be resolved to a jump table, every instruction in the function becomes an entry point, because the jump could land on any of them. That fallback was also applied to JALR, which is not a jump but a call: it transfers control to another function and returns to the instruction after the delay slot. That return address is already queued as a resume target a few lines above, so nothing else in the function needs to be reachable from outside. Indirect calls are ordinary code -- function pointers, virtual dispatch, callbacks -- so the fallback fired constantly. On a 3 MB PS2 executable, 2,210 of the 2,422 unresolved sites were JALR, and 1,078 of the 1,282 affected functions contained no unresolved jump at all. Restrict the fallback to JR. Promoted entries drop from 189,876 to 1,688, registered table entries from 156,783 to 75,386, the generated registration file from 13 MB to 6 MB, and total output from 180 MB to 163 MB. Every indirect call site in real code keeps its return-address resume entry (the only sites that lose one are bogus functions carved out of rodata, where the address is outside the function anyway).
A syscall can hand control back to the scheduler before the instruction after it runs. SetSyscall lets the guest install its own handler for a syscall number; dispatchSyscallOverride then suspends the calling thread and queues that handler as a GuestInvocation. When the invocation finishes, EeScheduler resumes the parent thread at the address the generated code stored just before calling handleSyscall -- the instruction right after the syscall. The analyzer never marked that address as an entry point. It queues resume entries for JAL and JALR only, so no generated function could be re-entered there, EeScheduler's hasFunction() check failed, and the thread was made dormant instead of resumed. The thread simply stops; because the scheduler then drains normally and run() returns, it looks like a clean shutdown rather than a fault, which makes it awkward to recognise. This is reachable during early boot on a real title. Dragon Quest VIII hits it in crt0: the Metrowerks startup code installs a handler for syscall 0x83 and immediately issues it, and execution ends there, roughly ten functions into the binary. Note the offset is +4, not the +8 used for JAL and JALR -- syscall has no delay slot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dispatchSyscallOverride returns bool, and its caller uses that to decide whether the built-in handler still needs to run. The success path -- the one that actually queues the guest's handler as an invocation -- fell off the end of the function without returning. That is undefined behaviour, and the practical failure mode is bad: whatever happened to be in the return register decided whether the built-in syscall ran in addition to the guest's override, so a game that overrides a syscall could get the effect applied twice, or not at all, depending on the build. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ps2_runtime.h includes <smmintrin.h> unconditionally, and the recompiler emits SSE4.1-only intrinsics (_mm_blendv_ps and friends) for the COP2 and FPU select idioms. SSE4.1 is therefore a hard requirement of the codebase, not a tuning option. Nothing sets it for GCC or Clang. MSVC does not need it -- its intrinsics are not gated behind a target feature -- and the only place any x86 feature is named is /arch:AVX2 inside EnableFastReleaseMode, which is MSVC-only and applies to Release and RelWithDebInfo only. GCC and Clang default to the plain x86-64 baseline, which is SSE2, so on a stock Linux or macOS toolchain the affected translation units fail with "always_inline function ... requires target feature 'sse4.1'". Adds EnableX86SimdBaseline and applies it to ps2_runtime in every configuration rather than in EnableFastReleaseMode, since a Debug build needs it just as much. PUBLIC, so recompiled game code linking against ps2_runtime inherits it -- that code is where most of the SSE4.1 intrinsics actually are. Guarded on the target processor so ARM builds are unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
logDeci2Text stops after 256 lines and then silently drops everything else. That is a sensible default against a game that spams kputs, but it makes the runtime unusable as the candidate side of a boot-trace diff: real captures run to many thousands of lines, and a silent truncation partway through looks exactly like the guest stopping. Adds PS2X_DECI2_LOG_LIMIT, with 0 meaning unlimited. Unset or unparseable keeps the existing 256. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Guest code reads COP0 Status to decide whether interrupts are enabled,
and two different bits are involved:
IE (bit 0) the architectural MIPS interrupt enable, set once by the
kernel during boot and normally left set.
EIE (bit 16) the EE-specific enable that `ei` and `di` toggle.
We never execute the boot ROM, so nothing was setting either one, and
R5900Context started with Status at zero.
That is not cosmetic. libkernel's StartThread opens with
`mfc0 Status; xori 1; andi 1` and bails out with -1 when IE is clear --
its "you must call iStartThread from an interrupt handler" guard. With
Status at zero that guard fired every time, so every StartThread failed.
Dragon Quest VIII hits this during boot: it creates its CD streaming
thread, gets -1, prints "Can't start thread for streaming." and then
deadlocks with every thread blocked and none runnable. Nothing in the
runtime logs anything, because from its point of view the guest simply
asked a question and got an answer.
EIE matters for the matching reason: DIntr reports whether it was set so
the caller knows whether to pair it with an EIntr. Starting at zero makes
DIntr always answer "already disabled", so the re-enable never happens.
Two changes, both needed:
- R5900Context's constructor now sets Status to EIE | IE rather than 0.
BEV is deliberately left clear -- that selects the boot exception
vectors, which is the pre-handoff state, not this one.
- PS2Runtime's constructor no longer memsets m_cpuContext. R5900Context
already zeroes itself before applying its reset values, so the memset
only threw those values away. Threads created later were unaffected
because EeScheduler::startThread assigns `R5900Context{}`, which is
why this presented as "the main thread cannot start threads" rather
than something more obviously global.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ran-j pointed out on ran-j#211 that invokeCurrent is [[noreturn]] -- it throws EeDispatcherTransfer to unwind back to the scheduler -- so control never reaches the end of the function and the return I added was dead code. He closed the PR; this drops the change so the branch matches upstream again.
decodeFunction stops at function.end, so when the last instruction inside the range is a branch its delay slot falls outside and the emitter substitutes a NOP (makeSyntheticDelaySlot). A real instruction disappears with no diagnostic. For a jr $ra epilogue that instruction is the addiu $sp, $sp, N, so every call to the function leaks stack. In DQ8 that was __umoddi3, called from _strtol_r while parsing a text asset: after a few calls _strtol_r read its own saved $ra from the wrong offset, got 0, and walked the main thread off the end of crt0. The symptom looked like a scheduler hang. 286 functions in DQ8's map end on a branch, 165 of them where the next mapped function starts more than one word later.
dispatchIrq seeded the handler invocation with handler.sp, which is whatever $sp happened to be when the guest called AddIntcHandler. By the time the handler fires that thread has long since moved on and is using those addresses for something else, so the handler builds its frame on top of live data. In DQ8 the VBLANK_END handler landed inside the main thread's stack and zeroed a saved $ra, and the thread returned to 0. The dispatcher already does the right thing for an invocation whose $sp is 0 -- it hands out an invocation stack -- so passing a recorded $sp was bypassing it. Same fix for the GS VSync callback and the Alarm event, which shared the pattern. That alone would have moved the corruption rather than removed it: reserveAsyncCallbackStack carves invocation stacks downwards from the top of RAM, and the EE kernel puts a game's initial stack there too (DQ8's main thread runs at 0x01F40000..0x02000000). reserveGuestStackFromAsyncPool now pulls the pool below any guest stack that reaches into it, called wherever a thread's stack is recorded.
Two related seams for games that do more than the runtime assumes. Guest-owned heap. A game shipping its own allocator (newlib and MSL both do) grows it through sbrk/EndOfHeap and hands out addresses the runtime must not also be allocating from. setGuestHeapCeiling makes EndOfHeap report the guest's ceiling rather than the runtime arena's limit, and SetupHeap no longer calls configureGuestHeap when one is set -- the guest is describing its own heap there, and honouring it dragged the runtime arena back on top of the guest's first chunk. guestHeapHardLimit exposes the top of the allocatable range, which guestHeapLimit cannot report before the heap is configured, so a caller can split the range without hardcoding the constant. Found in DQ8: the runtime's 128-byte GIF packet from sceGsResetGraph landed on the guest heap's top chunk header, dlmalloc read a zero size, and its own guard (set_head(top, PREV_INUSE), "will force null return from malloc") did exactly that. Overlay dispatch. The generated function table is one dense array over one address range, so it holds one function per guest address. Titles that stream code into a fixed arena break that. FunctionRegionResolver lets a host that translates each image separately say which table covers an address right now; nothing about overlay formats or loading enters the runtime.
The VIF interpreters decoded the bit and set VIFn_STAT.INT, but nothing turned that into an interrupt, so a game that asks to be told when its display list completes waits forever. The acknowledge side already worked: a VIFn_FBRST write with STC clears the STAT bits. PS2Memory latches the request (bit 0 VIF0, bit 1 VIF1) and EeScheduler::processPendingEvents drains it into dispatchIrq, mirroring how EE timer interrupts already travel from advanceEeTimers to the scheduler. DQ8 sends one such VIFcode -- 0x80000000, a VIF NOP with the i bit -- then blocks on a semaphore its INTC cause 5 handler signals. Without this the game stops with the title artwork loaded and never draws. With it, it renders. Not emulated: hardware stalls VIF1 at the interrupt and resumes on the STC write, while the interpreter runs on as it already did when it only set STAT.INT. DQ8 puts the i bit on the last VIFcode of the packet so there is nothing left to stall; a game that interrupts mid-packet would need it.
SMODE2.FFMD selects what the frame buffer holds, not how the CRT scans it, and PresentFromLocalMemory had it inverted. FFMD=0 buffers a whole frame whose two fields are alternate rows, so weaving to a progressive host frame is just reading every row; FFMD=1 buffers half the display height and each field reads all of it, which is where line-doubling belongs. Applying the FFMD=1 rule to an FFMD=0 game threw away half the vertical detail. Also snaps the two PCRTC read circuits together when they cover the same buffer one row apart. That configuration is a deliberate vertical flicker filter and merging it faithfully costs half a row of blur; PCSX2 snaps by default. This is the one deliberate departure from hardware here, and it is isolated -- deleting it costs about 1.6 dB. Measured on six DQ8 title-screen GS dumps: identical adjacent rows 224/447 -> 0/447, PSNR 28.97-29.12 dB -> 29.53-29.67 dB. The win holds under both resample directions. The FFMD=1 branch has no test coverage -- every dump is FFMD=0 -- and is reasoned from the spec rather than exercised.
Two defects that both trace back to the frame rate, and one hazard beside them. Input. readState() sampled IsKeyDown() on the EE thread once per scePadRead, which is a level read of raylib state refreshed once per presented frame. At 7 fps that is a 140 ms window and a press-and-release inside it was never observed at all. Edge detection alone does not fix this: PollInputEvents copies current->previous and then calls glfwPollEvents, and GLFW's callback sets the key to 1 on press and 0 on release, so a tap whose press and release land in the same poll leaves both states at 0. Measured, 5 ms taps fired on 0 of 51 frames for IsKeyDown and IsKeyPressed alike. What survives is keyPressedQueue, which the callback appends to regardless. ps2PadPollHost() therefore runs on the render thread and publishes two atomics -- a held mask from IsKeyDown and a latch drained from GetKeyPressed -- and readState() ORs them, consuming the latch exactly once. Taps of 5/20/60 ms went from 0/15, 3/15, 6/15 seen to 15/15 in all three cases, one guest frame of button-down each. Declared in a new ps2_pad_host.h rather than on PSPadBackend: ps2_pad.h is reachable from ps2_runtime.h, so putting it there rebuilds the whole recompiled corpus for an input change. Presentation. The window is resizable and SetTextureFilter was never called, so a non-integer nearest-neighbour scale duplicated some pixel columns and dropped others -- at 700x500 only 510 of 512 source columns survived, and 1920x1080 gave runs of 2 and 3 pixels. One-pixel font stems came out uneven. Upscales now snap to floor(fitScale) with point sampling, bilinear is used only below 1.0 where integer scaling has nowhere to go, and the destination origin is floored because odd centring offsets were an independent half-pixel source. Verified over 15 window sizes: 8 of 12 upscales mangled before, 0 after, all 512 columns present with uniform run widths. The window now opens at 2x when that fits in 90% of the monitor. Also SetExitKey(KEY_NULL): Escape is bound to Circle, and raylib's default exit key is Escape, so pressing Circle closed the game.
… pool Pending invocations are pushed onto the running thread between any two guest function dispatches, including partway through one already in flight, so they nest rather than run in sequence. DQ8's movie streaming queues SIF RPC completion callbacks faster than they retire and reached depth 13 on one thread, exhausting the invocation stack pool and throwing. Past four deep the rest simply stay on m_pendingInvocations and are picked up as the depth drops. Real nesting is at most a couple deep -- an interrupt during a callback -- so this bounds the pathological case without changing legitimate behaviour, and an unbounded depth cannot be supported anyway: the pool is the fixed gap between the runtime arena and the guest's own stack. Measured on DQ8: 16 stacks consumed then failure, all but two of them thread 12 at depths 0 through 13, all GuestInvocationKind::RpcCallback. With the bound the game runs past the movie open into its sound settings screen.
Three defects that between them left DQ8's logo movie as a frozen frame that never terminated. All are reachable by any game that plays a PSS movie. 1. A runtime built without FFmpeg deadlocks. The stub decoder's feed() returns false and the failure path sets decoderFailed = false, which is right for the FFmpeg build where false means "no sequence header yet, resync". With no decoder that is not recoverable, so sceMpegGetPicture parks on a frame that can never arrive. decoderFailed was declared, cleared in two places and read in five as the give-up signal, but never once set: the path was dead code. Both decoder variants now carry kAvailable and feedElementaryStream raises the flag. Both demux entry points include it in the completeExternalWait predicate, without which a thread already parked -- the usual case, since the first GetPicture precedes the first demux -- is never woken. 2. Nothing could report end of stream to a game that streams with plain sceCdRead. The two available signals are a program-end code in the PSS and currentCdStreamEofSeen, and the latter is only ever raised from sceCdStPause /StRead/StStart/StStop. DQ8 contains no sceCdSt* function at all and its .MVI files carry no program-end code -- they end in 0xFF sector padding -- so the movie decoded all 270 pictures and then hung forever. sawSequenceEnd already tracked the video sequence-end code 00 00 01 B7 but was read only by a debug log; it now drives the wait predicate and the guest's own end flag. The flag itself is inner[0x00], reached through mpeg[0x40]. The game's sceMpegIsEnd equivalent is three instructions -- return **(mpeg+0x40) -- and the stub wrote every neighbouring field but that one. 3. mpeg[0x08] gates the caller's VRAM upload rather than counting pictures. DQ8 uploads its movie buffers only while it reads zero, so writing a running counter there stopped the movie updating after the very first frame. It now reports whether this call produced a picture. Also stop sceMpegReset carrying streamEnded forward unless the sceCdSt* producer is actually driving. appendPssBytes refuses new data while that flag is set until cdStreamGeneration advances, and only notifyMpegCdStreamStart advances it -- so for a plain-sceCdRead game the first ended stream wedged every later movie.
m_invocationStackTops keys a 16 KiB stack per (thread, depth) and is only ever find()-ed and emplace()-d, never erased. reserveAsyncCallbackStack behind it is a bump allocator that walks m_asyncCallbackStackTop downwards with no free path. So every guest thread that ever takes one invocation holds its stacks forever, and the pool -- 0x01F00000 to 0x01F40000, exactly 16 slots -- only ever drains. DQ8 starts a fresh thread per movie and deletes the old one, so it hit the wall on the second: "EE invocation stack space exhausted" the moment TAKA_N.MVI opened. Instrumenting the allocation showed 14 of 16 slots held by seven live threads before the second movie was even a factor, thread 12 alone holding four at depths 0 through 3. releaseInvocationStacks returns a thread's stacks to a free list when its record is erased, from both destruction paths, and invocationStackTop pops that list before falling back to the pool. reset() returns them too: it clears m_threads and restarts ids at kFirstThreadId, which would otherwise strand every existing entry permanently. Recycling at record-erase time is safe. deleteThread requires the thread to be Dormant already, and exitCurrent throws EeDispatcherTransfer immediately afterwards, so unwinding completes before another invocation can claim the slot. terminateThread deliberately does not release -- it only makes the thread dormant and the id can be restarted. Measured on DQ8: delete thread 9 and delete thread 13 both recycle, thread 13 reuses the freed slots, and the allocation that threw on every prior attempt succeeds. Exhaustion count zero across a full boot.
PresentationFrame assumed the software backend's fixed 640-wide staging buffer: callers de-strided by that constant, so a backend returning tightly packed rows -- or rows at any other width -- was misread as skewed. Adds rowPitchBytes, and a GSPresentationMode distinguishing host pixels from a backend that presented through its own swapchain and has no pixel payload to copy. GS::copyLatchedHostPresentationFrame now uses the frame's own pitch and refuses one narrower than its width rather than reading past the end. Needed by a hardware backend, which presents at its own internal resolution: a 4x render target hands back 2048x1792 rows, not 640-strided ones.
PresentationFrame already had a BackendNative mode; nothing could reach it, because every frame went out through raylib. The runtime now takes an external presenter: with one installed it opens no window, creates no framebuffer texture, and does no drawing. It still latches each frame -- that is the call that reaches the backend's Present() -- and then hands control to the presenter to service its window. Between frames it sleeps briefly instead of spinning. raylib's frame limiter used to block in EndDrawing(); without a replacement the loop burns a core the EE thread wants. Audio moves out of the windowed branch. It has no window and an external presenter does not replace it, so initialising it only in the raylib path left the audio backend permanently not-ready. ps2PadPublishHostState() lets a host that is not raylib feed the same latching path, so press edges between two guest polls still register.
mpegDemuxBackpressured() returns true once the decoder is 8 pictures ahead, and the demux stubs then returned 0 consumed and nothing else. DQ8's producer loop reads that as "not now" and asks again immediately, so the movie thread span: 670,171 calls to sceMpegDemuxPssRing per 30 presented frames -- about 22,300 per frame -- against 0.5 sceMpegGetPicture calls per frame of actual work. That was the movie's frame time. Measured with the graphics side fully accounted for: the GS backend was 6% of wall, GIF packet decode ~0%, and the 896 tile transfers per frame cost almost nothing. Both demux entry points now rotate the ready queue before returning. This is deliberately a yield and not a park: the comment on mpegDemuxBackpressured warns that parking the producer can leave a consumer asleep with nobody to wake it, which is real -- Code Veronica wakes its video thread before every demux call. Rotating the queue lets the consumer run and comes back, so the spin becomes a scheduling point without changing who wakes whom. DQ8's attract movie goes from 0.8 fps to 9.1, and demux calls per frame from ~22,300 to ~1,025. Still 1,025 spins per frame, so this is a large improvement rather than a cure. The remaining fix is a real producer/consumer handshake -- park the producer on space-available and wake it from sceMpegGetPicture -- which needs care for exactly the reason the yield was chosen here. Also adds timers on sceMpegGetPicture, writeDecodedFrameToGuest and the demux entry points, and on the GIF frontend's packet and native-image-upload paths. None of this was visible before; the demux call count is what identified the bug, and it was only findable because the graphics side had already been measured and eliminated.
The previous commit called rotateReadyQueue from the demux stub and returned normally. That clears the current thread and requests a reschedule, so the next syscall reached bindMainContextForSyscall with no thread running and tripped its assert -- the game aborted on every movie. transferIfRequested must follow, which is what the RotateThreadReadyQueue syscall does, and the return value has to be set before it because the transfer throws. Verified: 0 aborts across a 700-second run that plays the attract movie through, 582 pictures served. Also tried and reverted: parking the producer on a space-available wait and completing it from sceMpegGetPicture. That deadlocks -- 4 pictures served instead of 658 -- which is exactly what the comment on mpegDemuxBackpressured predicts. Both the attempt and its result are recorded in docs/notes/movie-decode.md so it is not tried a third time. Honest numbers: 0.8 fps before, 1.2-1.7 fps now, demux calls halved from 7.6M to 3.4M over the run. That is an improvement, not a fix. Demux is now 6% of wall and the backend 6%; the other ~88% is still unattributed, and the principled next step -- buffering PSS bytes and decoding lazily so the producer is never refused -- has not been attempted.
… event pump DQ8's movie producer thread is a spin loop, and correctly so -- it feeds audio and waits for its CD ring, ~5,150 iterations per movie frame, which is about what a real 294 MHz EE would manage. Our cost per iteration was the problem. Three measured costs, each removed: publishSnapshot() built and sorted three vectors of kernel state on every one of ~40 scheduler operations, for state only the debug panel and an aggressive-log tick ever read. 45% of EE thread time during a movie. It now builds only when snapshot() has asked; the idle path publishes unconditionally so a quiet scheduler still converges for a poller. The MPEG stub refused demux bytes once 8 pictures were decoded ahead -- 134,563 refusals out of 134,569 calls per 30 frames -- and the guest's loop (remaining -= consumed; bgtz remaining) cannot terminate on a zero count. Demux now queues elementary stream and sceMpegGetPicture pulls decode from it, so lookahead is bounded by buffered bytes rather than decoded pictures: ~25 KB per frame on the wire against ~917 KB decoded. Refusals go to zero and demux calls per frame from 4,551 to 13. flushDecoderIfEnded() has to wait for the queue to drain, or a program-end code drains the decoder while the whole movie is still buffered and every later packet fails with AVERROR_EOF. processPendingEvents() ran after every guest dispatch and took three mutexes, read the clock, walked m_deadlines and heap-allocated a deque to find nothing. 24% of EE thread time, now 6%. Also: transferIfRequested() skips the unwind when the thread it would reschedule is the one already executing, and counters for dispatches, transfers and refusals so the remaining cost stays visible. DQ8_EE_DISPATCH_HISTOGRAM names the guest addresses the executor re-enters most. Logo movie 1.4 -> 3.3 fps. Still not smooth: 30,908 context transfers per movie frame at ~10 us each, where 60 fps needs ~540 ns. That is the exception-based transfer itself, and fixing it means fibers rather than throw/unwind.
…itch The EE scheduler transferred control by throwing EeDispatcherTransfer, which destroyed the guest's whole C++ call chain and rebuilt it on the way back -- about 2.7 executor re-entries per transfer, ~10us each, where 60fps allows ~540ns. DQ8's movie thread needs ~31,000 transfers per movie frame, so the movie ran at 1.4fps and _Unwind was 43% of EE thread time. EeFiber gives each GuestThread its own stack. A yield now saves the callee-saved registers and swaps the stack pointer: 22ns per round trip, measured. Guest dispatches per movie frame fall 82,441 -> 353 and thrown transfers per 30 frames 927,230 -> 2; the movie runs at 7.0fps. transferIfRequested suspends, and so does blockCurrentResumable -- semaphore, event flag and sleep waits, which only take a return value from makeReady() and whose callers have nothing left to do. Waits carrying a completion still throw, because the completion re-invokes its syscall on wake and would run twice against a stack that is still there. blockCurrent stays [[noreturn]]: giving it a return path makes the compiler emit ud2, and waitVSync/waitExternal take an optional completion, so a flag on blockCurrent turns those into SIGILL. x86-64 gets a hand-written switch, everything else falls back to ucontext (correct, but a sigprocmask syscall: 461ns against 22ns). setjmp/longjmp across stacks is faster still and deliberately unused -- _FORTIFY_SOURCE turns it into an abort. The trampoline realigns the stack before calling, or the first std::string inside a fiber faults on an aligned SSE spill. Stacks are mmapped with a guard page so an overflow faults instead of corrupting its neighbour. Also here: DQ8_PAD_SCRIPT replays input timed in guest vsync ticks so a script lands in the same place whatever the frame rate; DQ8_GFX_SCREENSHOT_DIR writes the presented frame as a PPM named after that same tick; DQ8_SKIP_MOVIES taps START while a movie plays, which is the game's own skip -- reporting end-of-stream from the MPEG stub instead leaves its streaming thread spinning on a black screen.
…spatch DQ8 reported no memory card in either slot because isValidMcPortSlot() demanded slot == 0 and DQ8 passes slot=1 to every libmc call -- `addiu $a1, $zero, 0x1` at each of its sceMcGetInfo/GetDir/Delete call sites. Everything downstream already models a console with no multitap (one card per port, and the mc root path ignores the slot entirely), so the validator was the odd one out. GetInfo now answers type=2 free=8192 format=1, the no-card prompt is gone, and the game goes on to look for its /BASLUS-21207dq8 save directory. DQ8_MC_TRACE=1 logs the card calls. RUNTIME_LOG is compile-time and sits in a header the whole recompiled corpus includes, so turning it on to watch one stub would cost a full rebuild. Two scheduler costs, found by sampling the EE thread while a movie runs: selectReady() walked all 128 priority deques on every dispatch (~7% of EE thread time) and now consults a one-bit-per-priority mask maintained by refreshReadyMask(); advanceEeTimers() did up to four 64-bit divisions per enabled timer on every guest safe point, and below one EE second the whole-seconds term is zero and a sub-tick accumulation needs no division at all. Movie 6.5 -> 7.2 fps for the two together. WriteRunCT32/WriteRunZ32 write a run of pixels with the traits and lookup table inlined, instead of a cross-TU call per pixel; measured no change, because PixelStorageTraits is only 11% of the profile even as its top leaf. Kept as strictly less work. DQ8_SKIP_MOVIES now only lifts the 30fps presentation pacing. Its previous behaviour -- tapping START -- does nothing, because DQ8's attract movies are not button-skippable: 314 injected presses through one movie changed nothing.
ReadRowCT32/Z32/P8 mirror WriteRunCT32/Z32 for the read direction: texture expansion walks whole rows, and the per-pixel Read entry points recompute page, block row and column from (x, y) with three integer divisions per texel -- 229k times for one 512x448 expand. Along a row only the in-page column and the page column move, so the row walk is increments and the swizzle table lookup is the only per-pixel work left. The address mask keeps Read()'s wrap guard for regions past 4 MiB.
The movie thread yields ~4M times a second and each yield pumped events twice. Three per-pump costs, each visible in its profile: - The tail of processPendingEvents took m_eventMutex to ask whether any event was still queued. m_pendingEventCount now answers without the mutex; postEvent counts under the same mutex it pushes under, and the pump zeroes the count when it drains. - takePendingVifInterrupts was an atomic exchange per pump; a plain load answers first because the pending set is almost always empty. - advanceEeTimers walked all four timers per safe point even when the game never armed one; a cued-any flag maintained on timer writes gates it. Measured on DQ8's boot-to-title: movie phases 38.8 -> 39.5-46.6 fps, the settings screens and title stay at the 60 fps vsync cap. The remaining per-yield cost is the fiber round trip and the syscall dispatch path.
This reverts commit 4b69ccf.
Expose constant upper and lower register usage independently so native blocks fold readiness loops. Use guarded ARM64 fused vector product-sum arithmetic while retaining the scalar interpreter oracle and exact exceptional fallback. Emit deterministic external instantiations across private compilation units to shorten native rebuilds. Verified 49 VU tests, 268 captured signatures and 788108 resumed raw-byte boundaries, plus 4112 perturbed field states. Paired heavy replays save 27-35 percent with readiness specialization beyond the SIMD stage. Seven integrated runtime suites and the generator coverage test pass. Guest cycles, callbacks and code invalidation are unchanged.
Sound drivers are the part of the IOP that high-level services reproduce worst. Their RPC protocol is simple enough, but the sound itself comes from how the driver programs the SPU2: envelopes, reverb, streaming through AutoDMA, timers that pace the sequencer. Rewriting that for every game would never end, so this runs the original modules instead. NativeIop loads IRX files onto a small emulated IOP: an R3000A interpreter, a high-level kernel with the calls sound drivers make (threads, semaphores, event flags, mailboxes, alarms, hardware timers, interrupts, the heap, SIF RPC and DMA, a little sysclib), and both SPU2 cores with their voices, reverb and 2 MiB of sound memory. IOP time advances only with the SPU2's output, 768 cycles per 48 kHz sample, and whoever pulls samples drives it. Work done to answer an RPC, and interrupt, alarm and timer handlers, runs without moving the clock; charging it to the clock made heavy RPC traffic run the drivers' timers ahead of the audio, and music played fast. Once a native module registers an RPC server, IopSubsystem hands requests for that SID to it before any high-level service. A module that binds back to an EE server can ask the host whether that server exists yet.
A runtime names the modules it wants run natively with ps2_native_iop::setModules(). From then on: - sceSifLoadModule of one of them also loads it on the emulated IOP. The high-level tracker still hands out the module id the game sees. - IOP heap allocation and EE-to-IOP sceSifSetDma go to the emulated IOP's memory, which is where the modules look for what the game sends them. - The first module to load opens a 48 kHz stereo stream whose callback runs the IOP. Volume 0 keeps it running silently, because games wait on their sound drivers whether anyone listens or not. SDRDRV's callback thread binds to an RPC server on the EE. Answering that bind at once left the thread spinning on a server that did not exist yet and starved the rest of the driver, so the bind now waits until the EE has really registered the server. PS2X_AUDIO_DUMP=path keeps a raw copy of everything played, to compare against a reference offline.
Two things made DQ8's movie audio sound chopped once its sound driver actually ran. The demux callbacks were queued on the EE scheduler and ran after sceMpegDemuxPss() had returned. By then the caller had released the ring space they point at, and a batch of queued invocations runs last to first, so the audio came out shuffled. libmpeg calls them inside the demux call, in stream order, and the stub now does too, with the return value set first since the callbacks run before the caller resumes. The demux also ran as far ahead as the picture queue allowed, far more audio than DQ8 keeps: it holds a quarter second and drops what does not fit. With an audio track, the demux now stays 150 ms of audio ahead of what has played, measured in real time rather than in pictures, which fall behind whenever decoding does. With both, every chunk the game streams to the IOP during a movie matches the movie's audio track byte for byte.
Each case pins something that was wrong at one point while getting DQ8's drivers to play: AutoDMA blocks out of order across the DMA restarts a driver makes for every half it refills, a stopped and restarted stream not starting over, BVOL and AVOL swapped (a core's own input follows BVOL), and LIBSD's halfword writes to BCR.
The block compiler covers regions of up to 16 pairs and some short loops; everything between them runs on the interpreter a pair at a time. In DQ8's field that per-pair work was most of what VU1 cost, and VU1 was most of what the game thread cost. This compiles programs whole, from the MSCAL entry to the E bit, out of microcode recorded while the game runs (PS2_VU_PROGRAM_PROFILE). A routine is the code reachable from one entry through branches; calls and register jumps end it and their targets become routines of their own. Routines are found at run time by PC and checked word for word against VU memory, so microcode that changed falls back to the interpreter. Pairs still issue in order under the interpreter's stall rules. What changes is where the work is done: - Each pair's registers, latencies and hazards are decoded at compile time, and reads that can no longer stall, given the pairs before them in the block, are dropped from the readiness check. - Results are written at issue and only their ready cycles are kept. MAC, status and clip results wait in a small queue that is folded when something reads them rather than every cycle. - The clock, the flag queue and the kick state are locals, so they stay in registers. Whatever is still in flight when a routine returns goes back into the interpreter's pipelines, and a later pair or flush sees what it would have seen. - On AArch64 the common FMAC cases use NEON, with range checks that send anything unusual to the interpreter's arithmetic. MAX, MINI, ITOF, FTOI, ABS and CLIP have small forms of their own everywhere: inlining the interpreter's whole upper switch into every pair made a routine take minutes to compile. A routine only starts with at least 4096 cycles of budget. It hands back at a block boundary when the budget would run out, or at a pair it cannot compile. On 768 programs captured in the field, the routines leave the same state, VU memory and GIF packets as the interpreter at every budget the replay checks, and run about 9 times faster on an M1 Pro (5 times on the portable path). Compiling all of DQ8's recorded routines takes under ten seconds. In the opening field this took the game from about 15 to about 22 completed frames per second, after which the renderer was the limit.
Game microcode cannot go into the repository, so vu_programs.py writes synthetic programs in the layout a recording has. A few are written out to reach particular paths (blocks split for length, an E bit on a branch, a program handed back part-way) and the rest come from a seeded generator mixing what the compiler accepts inside counted loops, forward branches, calls and register jumps. ps2_vu1_program_tests.cpp runs each of them compiled and interpreted, whole, and then stopped by budgets the program outlasts and resumed in uneven slices, comparing state, VU memory and GIF packets each time. PS2_VU_PROGRAM_FIXTURES names the programs; without it the test has nothing to run. PS2_VU_REQUIRE_PROGRAMS makes a missing routine a failure. Writing it turned up two bugs, fixed in the compiler itself: P could take an older EFU result after three EFU operations in a row, and a routine lookup could be answered from a cache entry left by another interpreter whose code happened to sit at the same address. test_compile_vu_programs.py covers how the generator cuts code into routines and blocks.
The worker gave back a batch's queue space only once all of it had run. A batch can hold most of a frame, so the producer stopped for the whole batch and the worker then sat idle waiting for the next one. Giving back space every 64 KiB keeps both of them busy.
processGIFPacket read the clock twice per packet for a total that only a backend's stats report reads. A busy scene sends over ten thousand packets a frame, so that cost a couple of percent of the game thread. A backend that reports the numbers now sets g_gsFrontendTiming.
Both DQ8 pull requests move the submodule; this is the commit to use once both are in.
Running the callbacks inside the demux call puts them on the calling guest thread's stack. The MPEG tests call the demux from the host, with no guest call in progress, so there is no such stack and the scheduler threw. Those callers get the callbacks queued, as before.
x86 hosts ran compiled routines on the plain C++ path, a little over half as fast as NEON on the field captures. The SSE2 version takes the same steps with the same range checks. SSE2 only has signed integer compares, which is fine as nothing compared reaches 2^31, and movemask gives the lanes in the opposite order to the MAC flags, hence the table. Product sums have to round like the interpreter's acc + fs * ft, which GCC and Clang fuse into one multiply-add when the target has one. So the fast path fuses on AArch64 and with FMA enabled, and multiplies then adds otherwise. PS2X_VU_PROGRAM_SSE builds this path on AArch64 through sse2neon. Run that way, all 768 field captures and the synthetic programs match the interpreter, at 8 to 9 times its speed (NEON: 9, plain: 5). Reversing the lane table or rounding product sums the other way makes both fail. On x86-64 it has only been compiled here, with and without -mfma; DQ8's Linux x86-64 CI job runs the synthetic programs.
A compiled routine ends at a register jump, and run() only looks the target up, and records it if it is missing, once the routine before it exists. A first recording is made with nothing compiled, so every program runs on the interpreter and those targets never showed up. DQ8's main field program starts at 0x960 and jumps on to 0x970-0x990, so a build from one recording compiled the start and ran the rest on the interpreter: 97% of VU1 work, about 6 FPS in the field. It took a second recording, made with the first routines built, to find the targets. While a profile is recorded, the interpreter now saves where JR and JALR land, once per target and microcode generation since recordMissing hashes the whole image. One pass along the field route compiles to the same 52 routines several passes produced before, and the field holds 30 FPS with all of VU1 compiled. setProfileDirectory lets the new test record into a directory of its own.
libmpeg's sceMpegCreate finishes by calling sceMpegReset, which clears the end-of-stream word the game polls through sceMpegIsEnd. The stub stopped short of that, so a decoder created in memory that an earlier user had left non-zero could read as ended before its first picture. Dragon Quest VIII creates the decoder for the movie after its first fight in heap memory that gameplay has already used. Its decode thread checks the end word before asking for a picture, finds it set, resets and exits without decoding anything, and the player waits for the two pictures it needs before starting. The screen stays black. The new test fills the work buffer with 0xFF before sceMpegCreate and fails without the reset.
Record VU jump targets so one recording pass is enough
Reset a new MPEG decoder the way sceMpegCreate does
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft, to show what my fork has been doing for Dragon Quest VIII and to see how you'd like it brought in. With it the game reaches its 30 FPS cap in the field on an M1 Pro, with sound.
The main pieces:
ps2xIOP/src/lle.It also carries #205 to #209, which can go in on their own.
It predates #244 and conflicts with it in 20 files. The IOP part overlaps with the new emulator, so rather than keep two I'd move the SPU2 and the kernel calls these drivers need onto yours. CI will be red as well: MSVC can't build the fork's runtime yet (
ucontext.h, weak symbols) and the GCC job fails at an LTO link step. Happy to split all of this up in whatever order suits you.