Playdate C optimization guide: cache, TCM and linker tricks
A post on the Playdate developer forum shares advanced optimization tricks the author found while working with @stonerl on a full-speed Gameboy emulator for the Playdate. The author says these are not the usual game-optimization tips, and are most useful for demanding code such as an emulator, a large simulation, a 3D renderer or a codec.
The first idea is a mental model: the Playdate's CPU is very fast but slow at reaching memory, especially memory that is not in the cache. This comes from research by @StiNKz, with results on Discord. Code that works mostly on registers and loads little data should run very fast, even with many loops and branch mispredictions.
Second, the instruction cache is small. The author notes that -O3 is not necessarily fastest and -Os might be much better. Rev A's instruction cache is only 4 kilobytes (apparently 16 kb for Rev B, the author adds with a question mark). If the core code, such as an emulator interpreter loop, a fragment renderer or a physics engine, does not fit, it will likely run worse than more compact code, even if the compact code needs many more CPU operations. The author's first improvement for Gameboy emulation was to squeeze the 20 kb emulator core into 2 kb by replacing a giant switch table with much less code and a few branches, and by using -Os. Behaviour was identical, but it ran much faster. Rare operations, such as almost-never-used opcodes, can live elsewhere. Core functions should sit at contiguous addresses, which can be done by tagging each with attribute((section(".text.
Third, a custom linker map gives finer control. Copy link_map.ld from the C_API/buildsupport folder, and add override LDSCRIPT=./link_map.ld to the makefile before including common.mk. The linker script can then place a custom section, all code from chosen object files, and align things (for example . = ALIGN(32);) before the remaining code. Fourth, the command nm Source/pdex.elf | sort > syms.txt lists every symbol with its address, so you can check how big the core code is and whether it fits in the icache.
Fifth is DTCM, a small region of tightly-coupled memory that is faster than normal memory. The stack sits in it, so a stack-allocated object is generally faster than one on the heap or a static variable. If you will operate on a struct for a while, it can pay to memcpy it onto the stack, work on it, and copy it back; @RPDev did this for PlayGB and got a big performance boost. To avoid the two copies, the author keeps data permanently at the low-address end of the stack, since the app does not control main. In the eventHandler on kEventInit, use __builtin_frame_address(0) to find the high end of the stack, then subtract less than 10 kb (0x2180 has been safe for the author) to reach the low end. The author calls this area the dtcm_mempool, which persists from update to update, and advises placing a canary at its start and end and checking both at least once per update to catch stack overflow. Other TCM memory, such as the framebuffer, can also hold data, but anything stored there appears on screen.
Sixth, code can go into TCM too (ITCM). The author says it is unclear whether this speeds anything up on Rev B and more research is needed, but on Rev A it is the secret sauce for efficient Gameboy emulation. The Playdate does not place code there automatically, so you copy it manually. Mark functions with a macro, #define _itcm attribute((section(".itcm"))) attribute((short_call)), and in the linker map expose __itcm_start and __itcm_end around *(.itcm). Then memcpy the whole block to the destination and call functions through the relocated address. The destination must be congruent to the original address mod 2, because the Cortex M7 uses the lowest bit of a function pointer to mark Thumb code, and for the same reason you cannot memcpy from a function address directly. Considerations: do not enable -fPIC, because the relocation table breaks relocated code; use short_call so ITCM functions calling each other use relative jumps; mark any function outside the ITCM region that ITCM code calls with attribute((longcall)); and flush the icache after copying code, using a function in the Playdate API. The author warns this will likely crash at first, so start with a single trivial function.
Seventh is the performance lottery: as code changes, some builds are, without an obvious pattern, faster or slower, sometimes significantly (like 50% faster). The author names two prominent causes. One is cache misalignment: the cache line is 32 bytes, so a function that straddles two lines is slower, and adding one function of an awkward size early on shifts every later function. The fix is to force 32-byte alignment now and then, with attribute((aligned(32))) or . = ALIGN(32); in the linker map, optionally plus a manual offset such as . += 0xC chosen by checking the nm output of the last fast build. The author puts . = ALIGN(32) before each source file and section, and also suggests -falign-loops=32. The other cause is branch prediction. The Cortex M7's details are not public, and by the author's own estimate, which they say could be completely wrong, there is a prediction table for 1024 unique instruction addresses. Two if statements a multiple of 1024 bytes apart would then be treated the same, causing mispredictions. The suggested remedy is to sprinkle a handful of . = ALIGN(1024); . += n; directives with n between 0 and 1024, ideally chosen from the address of the next symbol in the last good build. That wastes 4+ kilobytes, so it should not be overused. The author may be wrong about branch prediction being the cause, but says the ALIGN(1024) trick helped regardless, and that the performance lottery has almost disappeared since implementing these tricks.
The eighth section covers prefetching. The author has not observed any improvement from it, but may not be using it smartly enough. @FReDs72 describes a case where prefetching helped in a Discord thread, and the author suggests measuring frame time with the high-precision timer because gains are likely fine-grained. The available text ends partway through this section.
Key facts
- Rev A's instruction cache is only 4 KB (apparently 16 KB on Rev B, per the author). Core code that does not fit runs worse than more compact code, even if the compact code needs many more operations.
- The author shrank the Gameboy emulator core from 20 kb to 2 kb by replacing a giant switch table with a few branches and building with -Os; behaviour was identical but much faster.
- Tightly-coupled memory is faster: the stack lives in it, and a permanent area at the low end of the stack (about 0x2180 below the high end worked) can hold data, even code copied into ITCM, which the author calls the secret sauce for Rev A emulation.
- Build-to-build speed swings of around 50% (the performance lottery) come from cache-line misalignment (32-byte lines) and, by the author's unverified estimate, a 1024-address branch prediction table; alignment directives in a custom linker map almost removed them.
- The author has seen no gain from prefetching, though @FReDs72 reports a case where it helped.
Why it matters
The post documents low-level behaviour of the Playdate's Cortex M7 that the typical optimization advice does not cover. Its central finding is counterintuitive: on this hardware, smaller code can beat code that executes fewer operations, because a 4 KB instruction cache on Rev A punishes anything that does not fit. The author's example is concrete: a 20 kb emulator core cut to 2 kb ran the same but much faster. It also explains why speed seemed to change at random between builds, and offers ways to remove that noise.
Who it affects
Playdate developers writing C for performance-hungry programs. The author names emulators, large simulations (a factorio-like), 3D renderers and codecs as the most likely beneficiaries, and says the tips may be less useful to readers unfamiliar with the existing techniques. Developers of the Gameboy emulator and PlayGB are directly involved: @RPDev used the stack-copy trick in PlayGB.
How to use it
Start by measuring: run nm Source/pdex.elf | sort > syms.txt to see how big your core code is and whether it fits in the 4 KB cache. Try -Os instead of -O3, tag hot functions with a section attribute so they sit contiguously, and copy link_map.ld from C_API/buildsupport with override LDSCRIPT=./link_map.ld in the makefile. For data, keep hot structs on the stack or in a dtcm_mempool at the low end of the stack, with canaries at both ends checked once per update. For code in ITCM, start with one trivial function, keep -fPIC off, use short_call and longcall as described, and flush the icache after copying. To tame the performance lottery, add . = ALIGN(32) before each source file and section, consider -falign-loops=32, and sprinkle a few ALIGN(1024) directives with varying offsets, accepting 4+ kilobytes of wasted space.
How solid is it
This is one developer's field notes from a forum post, not a benchmarked study. The author gives few measurements: the 20 kb to 2 kb core shrink with a 'much faster' result, 'like 50% faster' as an example of build variation, and 'a big perf boost' for PlayGB from the stack-copy trick. No frame rates or speedup figures for the emulator are given. The author flags their own uncertainty: the Rev B cache size is written as 'apparently 16 kb for Rev B?', the 1024-address branch table is an estimate that 'could be completely wrong', and the benefit of ITCM on Rev B is unclear. The cache research comes from @StiNKz and is only linked to Discord. The available text is truncated in the final prefetching section.
Risks and caveats
The ITCM approach will likely crash on the first try if not done exactly right, and porting a large function is especially risky because it may call other functions by short call. Storing data at the low end of the stack depends on an empirically found offset (0x2180 was safe for the author, and the offset must stay below 10 kb); an overflow can corrupt it, hence the canaries. The author notes a better method would exist if the platform adds a feature for finding the low stack region. The Rev A and Rev B hardware may behave differently, with the cache size and ITCM benefit uncertain on Rev B. Branch-prediction explanations are unconfirmed, ALIGN(1024) wastes 4+ kilobytes of space each time several are used, and prefetching showed no benefit for the author.
“I found the performance lottery has almost disappeared since implementing these tricks.”
— Author of the Playdate developer forum post