![[WIP] Running FreeRTOS on a Dual-Issue RV32IM Core](/images/post/riscv/freertos.png)
[WIP] Running FreeRTOS on a Dual-Issue RV32IM Core
[WIP] This post is the record of getting FreeRTOS to run on the dual-issue RV32IM core from the previous posts. It is the same core that ran DOOM, and it is the same design the optimisation post measured at 0.7271 CPI on CoreMark. What it did not have was anything that makes a multitasking kernel possible: no working trap path, no timer interrupt, no machine timer registers and no usable CSR file. Most of the work was adding those pieces and then finding out which of them I had got wrong.
The demo that came out of this starts the scheduler, switches between two tasks on a timer tick, moves items through a queue, and ends the simulation through the exit register. There is also a second, harder program that runs several tasks at once, creates and deletes tasks in a loop, and checks its own results. Both run in simulation only. The main model is Verilator, and Icarus Verilog runs the same two programs as a cross-check. Nothing has been run on a board yet, so every cycle count here is a simulator count. Every number in the post came from a run that I repeated from a clean build before writing it down.
Before getting into the bugs, here is what the port asks of the hardware. FreeRTOS on RV32I runs in machine mode only. It needs CSR reads and writes, a machine timer exposed as two 64-bit memory-mapped registers, a trap vector, the interrupt enable bits in mstatus and mie, and an ecall to yield. The port reads the timer and writes the compare value with ordinary loads and stores, not CSR instructions. The port also refuses to compile unless both of its address settings are defined, which turned out to be useful: a missing setting failed at compile time, not on the first tick. This port has no memory protection, and the idle task spins rather than waiting for an interrupt, so the wait-for-interrupt instruction is not used.
The core I started from had the full pipeline, the hybrid predictor with its local, global and chooser tables, the BTB, and the return address stack. It also had a CSR file that aliased several registers onto one storage word, which did not matter for the bare-metal tests but matters the moment a kernel does a critical section. The kernel disables interrupts with csrc mstatus, 8 and turns them back on with csrs mstatus, 8. With the aliased file, that clear could land in another register, so interrupts would come back on inside a critical section. The demo’s source comment says it was written partly to catch this.
The additions were three pieces. The CSR file was rewritten with twelve decoded entries and two write ports, one per issue slot. When both slots in a pair touch the same CSR, the younger instruction’s value wins, and slot 1 reads the value slot 0 has just written. The trap overlay follows the standard rules: on a trap, MPP is set to machine mode, MPIE takes the old MIE, and MIE is cleared; mret restores MIE from MPIE. The mip register’s MTIP bit is driven directly from the timer comparison, and mtvec resets to 0x100.
The second piece is the trap unit in core_top.v. A synchronous trap, an ecall or an mret, is taken out of the ID stage, so that mepc holds the address of the trapping instruction. The check is written out in full:
assign id_sync_trap0 = trap_mode && if_id_valid0 && !hz_stall_id && !hz_flush_id &&
!ex_redirect0 && !ex_take_branch1 &&
!id_ex_jal1 && !id_ex_jalr1 &&
(id_ecall0 || id_is_mret0);
assign id_sync_mret = id_sync_trap0 && id_is_mret0;
assign id_sync_ecall = id_sync_trap0 && id_ecall0 && !id_is_mret0;
The timer interrupt is handled differently. It is pending when the timer is at or past the compare value, the timer enable bit is set and global interrupts are on. It is only taken when the fetch stage is in a state where taking it is safe:
assign irq_pending = irq_mtip && mie_w[7] && irq_global_en;
assign irq_take = irq_pending && irq_can_take && !id_sync_trap0;
assign trap_take = sync_take || irq_take;
assign trap_redirect_pc = id_sync_mret ? mepc_w : mtvec_w;
The trap redirect sits at the top of the next-PC selection, above every branch, so a trap always beats a prediction that was already in flight.
The third piece is the memory map. The core already had a data memory with two ports. I added a block at 0x1000_0000 that the port uses for the timer, the console and the exit. The table is in the header of mem_top.v, and the entries the port uses are these:
| Address | Register | Notes |
|---|---|---|
| 0x1000_0000 / 4 | mtimecmp low / high | read-write, resets to all ones |
| 0x1000_0008 / C | mtime low / high | read-only, counts every clock |
| 0x1000_0010 | putchar | one byte per store |
| 0x1000_0014 | exit | writing ends the simulation with that value |
| 0x1000_0018 | trap count | read-only, counts traps taken |
Each MMIO store is decoded per slot, as this excerpt shows:
wire putchar0 = wr0 && (mem_addr0 == `MMIO_PUTCHAR);
wire exit0 = wr0 && (mem_addr0 == `MMIO_EXIT);
wire cmplo0 = wr0 && (mem_addr0 == `MMIO_MTIMECMP_LO);
wire cmphi0 = wr0 && (mem_addr0 == `MMIO_MTIMECMP_HI);
wire [31:0] mtimecmp_lo_next = cmplo1 ? mem_wdata1 :
cmplo0 ? mem_wdata0 : mtimecmp[31:0];
The reason for the per-slot form is explained in one of the bugs below.
The startup code is short. It sets the stack pointer to the top of the 64 KB data window, writes mtvec, arms the trap controller, clears bss, and calls main. The arming step is the one that matters:
_start:
li sp, 0x8000FF00
la t0, freertos_risc_v_trap_handler
csrw mtvec, t0
csrwi 0x7C0, 1 # trap controller on: ecall traps instead of halting
la t0, __bss_start
la t1, __bss_end
1: bgeu t0, t1, 2f
sw x0, 0(t0)
addi t0, t0, 4
j 1b
2: call main
Bit 0 of the register at 0x7C0 chooses between the old behaviour, where an ecall stops the core, and the new one, where it traps. The older test programs all rely on the old behaviour, so the bit is off by default and the RTOS startup turns it on.
The memory layout is a 64 KB instruction window at address zero and a 64 KB data window at 0x8000_0000, with the stack at the top of the data window. The kernel config sets the CPU clock to 1 MHz and the tick to 1000 Hz, so one tick is 1000 cycles. It has eight priorities, a 128-word minimum stack, a 16 KB heap managed by heap_4, the timer task enabled, and task deletion enabled.
The first failures were in the plumbing. The first run halted on the first ecall instead of trapping, which meant the csrwi 0x7C0, 1 had done nothing. The immediate form of a CSR instruction carries its five-bit value in the rs1 field, and my write-data mux was taking the register value from that position. Nothing reported an error. The feature just looked like it did not exist. The fix was a select in the port map that uses the rs1 field as the write data for the immediate forms:
.wdata0(id_ex_funct3_0[2] ? {27'b0, id_ex_rs1_0} : ex_rs1_fwd0),
Bit 2 of funct3 is the immediate bit, so the select keys off that.
The address map failed in a way that looked like a hang. I had set both configuration values to the base of the timer block, which looked tidy. The port then read the compare register as if it were the current time, computed the next tick from that value, and never fired. The console printed one line and stopped, which looks like a deadlock, not a timer problem. The port needs two addresses: the low word of mtime, which is 0x10000008, and the low word of mtimecmp, which is 0x10000000. Once both were right, the first tick arrived.
A later failure only showed when the output looked wrong. Two putchar stores that landed in the same issue pair printed one character, and the two tasks’ output came out interleaved. Both slots shared one memory-mapped write path, so when both slots stored to a device in the same cycle, the second store was dropped. Per-slot decoding fixed it, and the rule for the rare case where both slots write the same register is that the younger slot wins. The register file already used that rule for its two write ports, so I did not have to invent anything.
The trap storm took the longest. After the timer worked, the console filled with ecall traps about every 110 cycles, all with mepc inside the idle task. The idle loop in this build is short:
600: lui a5,0x80001
604: li a4,1
608: lw a3,-996(a5) # pxReadyTasksLists[0].uxNumberOfItems
60c: bltu a4,a3,614 # if ready count > 1, go yield
610: j 610 # else spin here forever
614: ecall # portYIELD
618: j 608
With one ready task, the branch at 0x60c is not taken and the core should sit at the jump at 0x610. It reached the ecall constantly instead. I dumped the fetch stream for one iteration, cycle by cycle, with the PC of the pair being fetched and what sat in IF/ID:
43379 0x610 IF/ID = (608 lw, 60c bltu) squash_s1 = 1 inter-slot load RAW, replay at 0x60c
43380 0x60c IF/ID = bubble
43381 0x614 IF/ID = (60c bltu, 610 j) no redirect bltu resolves not taken, fetch ran on
43382 0x61c IF/ID = (614 ecall, 618 j) id_ex_jal1 = 1 the 610 j resolves in EX here
43383 0x5f00 trap taken on the ecall
Read carefully, the trace shows the cause. The fetch stage predicts only slot 0. When a pair has a jump in slot 1, the sequential fetch of the next pair is wrong-path filler until the EX-stage redirect arrives a cycle later. In cycle 43382 that filler contained the ecall, and the trap unit took it out of IF/ID on the same cycle as an older jump was already redirecting fetch. The interrupt path already required that no older control transfer was in flight. The synchronous trap path did not check for one. The fix was to give the synchronous path the same guard set, hz_flush_id, ex_redirect0, ex_take_branch1, id_ex_jal1 and id_ex_jalr1, which is the list in the trap unit excerpt above. A wrong-path ecall is now flushed and refetched on the correct path, and the flood stopped.
That fix changed the symptom. The console went quiet, which I first read as success. The next trace showed the core fetching the same pair and flushing it every cycle:
43390 0x610 IF/ID = (610 j, 614 ecall) flush_id = 1 squash_s1 = 1 irq_can_take = 0
43391 0x610 IF/ID = bubble
The second cause was in the hazard unit. A slot 1 ecall or mret is replayed, so that the trap unit sees it as slot 0 of the next pair. That replay keeps mepc precise when the ecall is the second instruction of a pair. The replay condition did not check whether slot 0 was itself redirecting control flow. When slot 0 was a jump that had already redirected fetch, the replay sent the pair back to fetch, the same pair was fetched again, flushed again, and so on. Interrupts could never be taken, because irq_can_take requires that IF/ID is not being flushed. The idle task was stuck with interrupts off, which is worse than a crash because nothing reports it. The corrected condition for the inter-slot trap case is:
wire inter_slot_trap = s1_id_trap_op && !inter_slot_ctrl && !s0_id_trap_op;
assign flush_id = do_flush || ((inter_slot_load_raw || inter_slot_trap ||
(inter_slot_ctrl && !s0_id_branch_pred_taken) ||
inter_slot_mem) && !eff_load_use_stall);
The !s0_id_trap_op term is the new part. If slot 0 is itself a trap, slot 0 must trap first, and squashing it would let the younger trap run in its place.
I took two rules from those two causes. Wrong-path fetch must not be able to have any architectural effect, and any replay that holds the fetch PC still has to be checked against the conditions for taking an interrupt. Without either rule, the core keeps running and never does anything useful, which is harder to see than a crash.
The demo is the first program that runs under the real kernel. Its source is 222 lines. A producer at priority 1 sends 20 items into a four-deep queue, pausing 50 ms between sends. A consumer at priority 2 blocks on the queue, prints each item with the tick it arrived on, and keeps a running sum. When the producer has sent its 20th item it suspends itself, and when the consumer receives item 20 it reports its stack high water mark and ends the run. The console output, with the simulator’s own trap log filtered out, reads like this:
[consumer] item 1 at tick 1, sum 1
[consumer] item 2 at tick 52, sum 3
[producer] sent 2 at tick 53
[consumer] item 3 at tick 103, sum 6
[producer] sent 3 at tick 104
...
[consumer] item 19 at tick 919, sum 190
[producer] sent 19 at tick 920
[consumer] item 20 at tick 970, sum 210
[consumer] done, stack high water mark 185 words left
FreeRTOS ran on the core.
The producer and the consumer print from different tasks, so the order of their lines is the order the scheduler ran them, and the simulator’s own trap log is interleaved with the console output. I filtered that out for the listing above.
The run takes 1,017,314 cycles and records 1036 traps. The testbench’s trap log only keeps the first 40 entries, so I did not split the total by cause. The total comes from the simulator’s own trap counter, not from the program.
The incident that cost the most time was not about FreeRTOS. I was editing the reset condition in core_top.v, and a scripted edit removed seven always blocks from the file. The project had no version control at that point, and I had no backup of that file. The demo then exited after about 4,200 cycles with no traps, which pointed at the core rather than the edit.
Recovery started from the earlier version of the core, which had the original form of the same file. The parent has the original combined reset and flush condition, so its six affected blocks still had the form I needed. A small script took each block’s assignments from the parent, split the combined condition into two branches, and wrote the result back. The script had three bugs of its own. It deduplicated lines on raw text. It emitted two branches where the split needs three. And it misread <= 0 clears as loads, which put some clears in the wrong branch.
I checked the output with a validator that compared every block against the parent. For each of the six blocks the validator compared the set of assigned signals, looked for duplicate assignments within one branch, and looked for assignments that were in the parent but missing from the repair. The block sizes matched: 32, 27, 12, 12, 8 and 8 assignments. The validator found five signals that exist only in the new version, all predictor state, and those were expected. It also found a real problem: the clear for id_ex_branch0 was in the reset and flush branches but missing from the load branch, where slot 0 needs it. I fixed that by hand after checking the parent’s text, then ran the validator again until every check was clean. The only structural difference left is the one I wanted: reset and flush clear in separate branches.
The repaired file went through the instruction regression, which passes 31 of 31, and through the parent’s own test programs. Their cycle counts match the numbers recorded before the incident: hello at 20 cycles, fibonacci at 239, binary_search at 426. I only made a backup copy of the repaired file after those runs were green, which is later than I should have done it.
The next problem was mine in a different way. After the repair, the first csr test gave 27 cycles and one trap, and my first reading was that the repair had broken the core again. The cause was the build script. It checked whether the simulator binary existed and skipped the rebuild, so the binary still reflected the broken file. A clean rebuild gave 468 cycles and two traps, the expected result. I now delete the build directory after every RTL change, and I do not trust a build step that only checks whether a file exists.
The stress test is the other half of the work, and it is the o work, and it is the one I was least careful with. It is one program with several tasks running at once, and it checks its own results. Two producers at priority 2 each push 30 items into a four-deep queue. An item is a 32-bit word with the producer id in the top half and a sequence number from 1 to 30 in the bottom half. Two consumers at priority 3 each take 30 items out. A counting semaphore with eight credits caps how many items can be in flight. A producer takes a credit before it sends, and a consumer gives one back after it receives:
xSemaphoreTake( xCredits, portMAX_DELAY );
xQueueSend( xQueue, &item, portMAX_DELAY );
The shared counters sit behind a mutex. The producers and consumers get 256-word stacks.
A churn task at priority 1 creates a child with a 128-word stack, waits on a semaphore the child gives once it is running, and the child deletes itself. That happens 24 times. A software timer with a 40 tick period, set to auto-reload, runs its callback on the timer task and bumps a heartbeat counter. Every stack and control block comes from heap_4.
The consumers keep a seen table with one entry per item, so a duplicate or a missing item is counted. When both consumers finish, the churn task checks the item totals, a checksum, the child counts and the heartbeat count, then writes PASS or FAIL to the exit register.
The timer was the first thing to go wrong, and the mistake was mine. I created it and never called xTimerStart, so the heartbeat stayed at zero. The final check caught it, and the fix was one call.
The heap took longer, and it is the more useful of the two problems. When a task deletes itself, its memory goes onto a list, and the idle task frees it later. Idle only runs when nothing else is ready. My churn loop went straight to the next child, so the memory was not freed in time and the heap ran out after 11 children. I added a one tick sleep after each child to give idle a turn. That was not enough. One run still ran out, this time after 17 children, and recovered no memory at all. The free heap dropped by 624 bytes every round, which is one 512-byte stack plus the control block. Nothing was being reclaimed.
What bothered me most is that when I printed the free heap from inside the loop, the run passed. The print changed the timing enough for idle to get the processor and reap the child. So the fixed delay was hiding the problem on some runs. I did not want a test that passes by luck, and a fixed delay is exactly that.
The final version has no fixed delay. Before each child is created, churn records the free heap. After the child deletes itself, churn waits up to 50 ticks for the heap to return to that level, and it records how long the wait was:
for( waited = 0; waited < 50U; waited++ )
{
if( ( unsigned int ) xPortGetFreeHeapSize() >= before )
{
break;
}
vTaskDelay( pdMS_TO_TICKS( 1 ) );
}
if( ( unsigned int ) xPortGetFreeHeapSize() < before )
{
ulReclaimFailures++;
}
If the heap never recovers, the round is counted as a failure and the final check fails the run. A leak, or an idle task that never runs, shows up as a FAIL with a count, not as a pass. The worst round in the passing run took three ticks.
The passing run ends at 270,403 cycles, with 450 traps and [stress] PASS. The console lines are:
[stress] FreeRTOS stress demo
[consumer] validated 60 items, errors 0, missing 0
[churn] created 24, deleted 24, shared counter 24
[timer] heartbeat fires 5
[heap] free bytes 14672, min ever 9504
[stack] consumer high water 181 words
[heap] reclaim waits max 3 ticks, failures 0
[stress] PASS
The free heap bottoms out at 9504 bytes and is back at 14672 bytes at the end. That is more than the start, because by then the other tasks have deleted themselves and their memory has been reclaimed. Icarus Verilog gives the same cycle count and trap count for the stress test and the demo.
The demo count changed once during this work. It was 1,017,264 before I turned task deletion on for the stress test. Task deletion was still off in the config when I started the stress test, so the churn task could not call vTaskDelete. Turning it on adds code to the demo, and the demo’s cycle count moved by 50. I checked that the demo’s output did not change, then recorded the new count.
Here is the full list of results on the repaired core, each from a clean build:
| Program | Cycles | Traps | Result |
|---|---|---|---|
| Instruction regression | n/a | n/a | 31 PASS, 0 FAIL |
| hello, fibonacci, binary_search (parent ELFs) | 20, 239, 426 | n/a | match recorded values |
| csr_trap_test | 468 | 2 | MMIO exit |
| CoreMark | 2,233,145 | 0 | legacy ecall halt |
| FreeRTOS demo | 1,017,314 | 1036 | MMIO exit, Verilator and Icarus agree |
| FreeRTOS stress | 270,403 | 450 | PASS, Verilator and Icarus agree |
| Directed trap-pairing tests | 450 programs | n/a | 0 failures |
The demo’s stack high water mark is 185 words. The stress test’s consumer peaks at 181 words of its 256. The CoreMark number is the one from before the incident, and it did not change after the repair.
The demo and the stress test reach a trap and an instruction next to each other only by accident,
and a kernel booting does not show that the trap logic is right for every pairing. So I wrote a
generator for those pairings, scripts/rtos/trap_pairs.py. It produces bare-metal programs that put
an ecall next to each instruction class the core can pair with it: ALU, multiply, load, store, taken
and not-taken branches, jal, jalr, a remainder, and a CSR read. Each pairing is placed in both fetch
slots, with zero, one or two nops in front so the trap lands on each slot position. Two ecalls in a
row, with no store between them, also run. There are 388 of these programs.
Every trap program has a reference twin. The twin is the same body with each ecall replaced by a nop and the timer left off. Trap entry and mret are supposed to be invisible to the program, so the twin and the trap program must leave the same value in every result word. The handler also records the cause and mepc of each trap. The count must equal the number of ecalls plus the interrupts taken, and the ecall mepc values must match the trap sites in program order. The twin is a comparison against the same core, not an independent model, so it catches a trap that changes the program’s results but cannot catch a bug that behaves the same way in both runs.
The second family is a timer sweep. One 20-block body contains two ecall sites, and the machine timer is armed at mtime plus K. I ran K over every cycle of the body, 62 values, and the interrupt fired at all 62 points. Each run has to match the reference twin, and the interrupt must be taken exactly once.
The first version of the generator missed the case that mattered most. Every block ended with a
result store, so two ecalls were never adjacent, and the situation where both fetch slots hold a
trap never came up. I checked this by removing the !s0_id_trap_op term from inter_slot_trap in a
copy of the RTL. The directed set still passed. After I added adjacent trap pairs at each alignment,
the same copy failed 150 of the 450 programs. The real RTL passes all 450. A second copy, with the
guard that stops a timer interrupt being taken alongside a synchronous trap removed, also passed. I
think that copy is equivalent under these checks, because the timer stays pending and is taken again
after mret, but I have not proved it.
The limits are plain. The sweep body is only 55 cycles long, the instruction mix is small, and the tests run only in Verilator. A reference model in the strict sense, an independent simulator that replays the same instruction stream, would be a stronger check, and I have not built one.
The flow is short. The FreeRTOS image is built from source with a script, converted to two hex files that the testbench loads, and the Verilator model is built with rm -rf obj_dir first. The model runs with a cycle limit passed on the command line. Icarus Verilog takes the same testbench file list with the memory sizes passed as parameters. The toolchain was Verilator 5.032, Icarus Verilog 12.0, riscv64-unknown-elf-gcc 14.2.0 and Python 3.13. The Icarus run of the demo takes around three minutes on the machine I used, which is much slower than Verilator, so the Verilator run is the one I would use day to day.
The limits are the ones I would want to know about. The port runs in machine mode with no memory protection, so nothing isolates one task from another. The idle task spins instead of waiting for an interrupt, which wastes cycles but is correct. The results are simulator results, and the board has not been tested, so timing on real silicon is unknown.
Sources I used, in roughly the order I needed them:
- The FreeRTOS RISC-V port documentation on freertos.org, for the MTIME and MTIMECMP configuration values and the chip-specific header the port expects.
- The neorv32-freertos project on GitHub, for how a small RV32I core with a machine timer and a UART is brought up under FreeRTOS, and which hooks the port calls.
- A SiFive forum thread on running FreeRTOS across several zones, which confirmed that portYIELD is an ecall the trap handler has to decode.
- A StackOverflow answer confirming that mtime and mtimecmp are memory-mapped and accessed with loads and stores, not CSR instructions.
- A HackMD note on bringing up FreeRTOS on VexRiscv, which I used as a checklist for mtvec, MPP, MIE, MEIE and MTIE, and handler alignment.
- The FreeRTOS-Kernel sources, including the RISC-V port files and heap_4.c, taken from a local copy of the upstream repository.
I also read the Zephyr RISC-V porting documentation while deciding on the approach, but did not follow it
