Getting back to our discussion of the raw performance that high-end cores need to demonstrate. What is one of the most common operations performed by any code? Array access. This is why x86 has addressing modes of the form [ebx + esi * 4] and ARM has [R0, R1, LSL #2]. Without it, you are forced, like an idiot, to shift a register left by two, then add it to another register, and only then use that to access memory. Three instructions for a single array access. The usual excuse given for this inexcusable lack of foresight is that “instruction fusion will fuse all those three instructions into one in fast cores”. Yeah… if anyone ever pulls that off, they will win many prizes. No core I am aware of fuses more than two consecutive instructions. None. The reason is quite obvious – the combinatorial explosion of the number of possible combinations to consider, track, and handle. So now that we’ve established the bullshit excuse is bullshit, what is there to be done? Well, a few YEARS after the spec was written, an extension was proposed to help this issue – Zba. It provides three instructions of the form SHxADD for x being 1, 2, or 3. This combines a shift with an add, basically shortening our array access from three instructions into two. This is still worse than having register + shifted register addressing mode, but at least now “the core can fuse them” becomes less bullshit and more believable, assuming someone produces such a core. One problem: SHxADD is always 4 bytes long, and the memory access instruction itself will be 2 or 4 bytes, so the array access becomes 6 or 8 bytes of code, to ARM’s 4. This is where all those people who were just shouting about the wonderfulness of the C extension for density and how great it is for high-perf cores quietly shut up and look at the floor. Yeah… For extra credit, the Zba extension was only ratified in 2021, over two years after the base spec. It took them TWO YEARS to realize that arrays exist!
Curiously, this would be easy to fix. Currently 3/4 of the encoding space is allocated to compressed instructions (all instructions whose lower 2 bits are not 0b11). Reusing some of that encoding space for better addressing modes is a no-brainer and would produce denser code with better array addressing ability. Unfortunately, it would make too much sense for anyone to actually do. Of all possible champions of sanity, Qualcomm … proposed doing this, and even prototyped it. It went nowhere… And you know you really fucked up bad when you manage to make Qualcomm sound like the voice of reason. But, back to our SHxADDs. There is no guarantee that you’d get to use them anyways, since Zba is an extension and is thus optional. Are you getting tired of hearing “optional” yet? Let’s talk about that next.
What does RISC-V have in common with USB-C and RCS? These things are all ostensibly standards, sure, but the interesting part is that claiming to be in compliance with one of these standards means NOTHING while being technically true. Is my USB-C cable wired only for USB 2.0 valid? Sure, USB 3 twisted pairs are optional. Is my non-e-marked cable valid? Sure, e-markers are optional. Can my USB 3.0 USB-C cable choose to not support 20Gbps? Sure, 20 Gbps support is optional. Can it support 20Gbps but not support 100W charging? Sure, that is optional too. Can my phone’s fully-compliant RCS implementation not support upgrading a text message to a video call? Sure! MIVC is optional. Can it fail to send pictures while I am on a call? Sure, that is optional! Can encryption not be supported? You bet, that is optional too! So what does it even mean to comply with the spec then, if everything is optional? It means the spec writers spent too much time engaging in mental masturbation and too little time in contact with the real world, basically. There are two ways this happens: academics who are not aware that outside their offices, there is such a thing as the real world, and design-by-committee situations, where the real world simply never gets a seat at the table – having failed to file a motion to be seated there in time for the chairman to bring it to a vote.
Just scope out this line: “The RISC-V B (Bit-Manipulation) extension is a standard collection of instruction set enhancements designed to improve performance and code density through efficient bit-level operations. It is split into distinct sub-extensions: Zba, Zbb, Zbc, and Zbs.” Only a design-by-committee process would ever unironically produce this sequence of words.
When writing a spec, every single thing you make optional, you split the possible implementations into two incompatible groups. Do this enough times and you end up with your spec being meaningless. And boy, did the designers of RISC-V screw the pooch here. Everything is optional! Multiplication – optional. Division – optional. Support for a user mode – optional. Supervisor mode – optional. CSRs – optional. Compressed instructions – optional. “Extra compressed instructions” – still optional. I bet that if they thought they could get away with it, they’d make addition optional!
The basic instruction set of RISC-V, at first publication time, included CSRs, which, as I had mentioned, are required for a standards-compliant method of handling interrupts as well as for support of differing privilege levels. That is not unreasonable; it is sane and not broken. Which is, of course, why they fixed it … ASAP! CSRs got pulled out into an extension called Zicsr, and now the base ISA lacks ability to handle interrupts or provide privilege separation. But it is actually, and hilariously, much more idiotic than that. Let’s say you have a lot of things that are optional. What is the first thing code would want to know? “Is feature X implemented on my hardware?” of course. CPUID is how you answer this question on x86. How do you do it on RISC-V? Well, I have good news and bad news. There is a CSR called misa which can answer some of those questions (not all of course, that would be too sane). Did you spot the problem yet? It is a CSR and CSR support is optional (Zicsr extension is not mandatory). If that was not enough of a crotch punch, misa is not required to be accurate if implemented – it is allowed to read as all zeroes – it being meaningful is itself entirely optional. Yup… the only way you have to detect optional features is optional. But wait, there is more!
The “M” in front of “misa” indicates that this is a machine-mode CSR. Machine mode is the highest privilege mode in RISC-V (and the only non-optional one, if you’re keeping track). This register is not readable from lower-privileged modes, even if you are lucky enough to (a) be running on hardware that implements Zicsr, (b) be running on hardware where misa is not hardwired to be all zeroes, and (c) running on hardware that implements other privilege modes. This means that tailoring your code to the capabilities of the hardware is not possible for normal user code. If your hardware implements the OPTIONAL supervisor mode, it also cannot detect the core features. The party line is “ask the machine mode supervisor”. The problem, obviously, is that you have no idea what that supervisor is or how to “ask” it. There is a common one in use called OpenSBI, but there is, of course, no way to detect if that is what your machine mode runs.
Another fun bit of optionality here is the system timer. RISC-V spec has a timer; it is optional, of course. It is not accessed using CSRs, because of course not! That would be too consistent. The official party lineexcuse is that this was done to conserve the encoding space in the CSRs. I guess this is because they expected to run out of … 4096 of them‽‽ A bit ambitious if you ask me – no current architecture comes even close, not even x86. But let’s move on. If the timer registers are not CSRs then where are they? They are memory mapped! Where? Well, since a timer is a core peripheral that any OS would need, and since the CPU core spec specified it, it, of course, is at a well defined address that you can rely on. Just kidding! Nothing in this spec is that sane! The address is “implementation defined” and can be anywhere at all. Good luck, have fun, don’t crash!
I shall tell you of just one more fun situation here, of the many I could: exception and interrupt vectoring. When an exception or an interrupt occurs (assuming the optional Zicsr is implemented), where does the CPU jump? Depending on a lot of optional and optionally-supported config regs, delegation regs, and all sorts of other overcomplicated nonsense, eventually the core will pick to use machine or supervisor vector register (mtvec or stvec). That CSR points to the handler, except its bottom two bits that determine its “mode”. What is a mode? There are two modes documented. The direct mode is when the lower two bits are 0b00, in which case all exception and interrupts just jump to the address in the higher bits. The vectored mode (lower bits 0b01) is meant to simplify and speed up interrupt handling. All exceptions jump to the address in the higher bits of the reg, while all interrupts jump to that address plus 4 times the interrupt number. So what is my problem with this seemingly sane design? That both of the modes are optional!!!! No part of the spec mandates even the simple direct mode! It is possible, at runtime, to detect if a given mode is implemented by writing the lower bits, reading them back, and seeing if they stuck. But it would be entirely valid to implement a core with only vectored mode supported. Or only direct mode supported, or both, or neither, if instead your core vendor invented their own separate mode. This makes writing any sort of a generic kernel very difficult – you literally have no idea what to expect. Why direct mode was not made mandatory I cannot fathom, but I can surely tell you that the person who decided that wore oversized shoes, had a big red nose, and wore a lot of white face makeup.
At this point in time you might jump to the defence of this indefensible idiocy by shouting one of two things: “other architectures also have optional features” and/or “hiding misa is needed to support virtualization, haven’t you read Popek & Goldberg?”. Let’s demolish these feeble excuses one at a time. For “other architectures” we’ll consider things in common use in the last few decades: x86, ARMv7, and Aarch64. x86, as previously mentioned, has CPUID which will happily tell you which features the current core has. It will do this quite easily in user mode, as one would expect. Arm has ID_AA64PFR0_EL1 and ID_AA64ISAR0_EL1 available to the kernel at least (though not to userspace). But there is a much more important point to be noted here, which explains why ARM’s design is not fatal. In both x86 and in ARM, optional features are of two clear classes: (1) high performance compute that is usually programmed using intrinsics or hand-rolled assembly for tight loops in special circumstances (video encoding, fluid simulations) or by libc (memcpy, memset, strlen), and (2) NOP-compatible optional things that can be safely run on hardware that does not support them since it will execute as a NOP and that is safe. For x86 that would be endbr64 and for aarch64 that would be almost all of the PAC instruction set. Note that at no point are things needed in completely normal compiled code optional. Multiplication, division, addressing modes, are always available. This means that a normal C compiler targeting these architectures does not face the impossible choice of: “compile for the lowest possible denominator to allow code to run on all arch versions” or “assume things like multiply and sane addressing modes exist and prepare to crash on a core that chose not to implement them”. Yes, of course, this is where people will say “just target your exact core, why don’t you?” Yes, I never said that the idiocy of this ISA cannot be overcome with enough contortion. I said that this sort of poor design was acceptable in the 1970s when we did not know better and is inexcusable in the 2000s, as now we do.
Now, on to the virtualization excuse. First of all, Popek & Goldberg talk about an architecture being virtualizable specifically in the context of it lacking special virtualization support. Indeed by trapping every instruction that acts differently in user and supervisor mode, one can virtualize any architecture. But the alternative is just building-in virtualization support. x86 is not virtualizable as per Popek & Goldberg, at least due to POPF instruction. And yet I have VMs running on my x86 box just fine, as x86 added support for virtualization. It was nontrivially difficult to bolt it on post-facto, but it was done. If doing it at architecture design time, it is trivial. Which is to say that ANY mention of Popek & Goldberg to justify decisions made at architecture design time is bullshit. Arch design time is precisely the time to do it right. Popek & Goldberg even mention that trapping everything is theoretically interesting for the proof of virtualization but not practical. NOT PRACTICAL. So what did the designers of RISC-V do? They justify misa being not exposed to supervisor and user mode with “but virtualization… what if the hypervisor wants to hide capabilities from a VM?”. Bull … let it arrive … shit! x86 and ARM both manage that just fine. And RISC-V could have too, simply by allowing the hypervisor to lie about misa’s contents while letting everyone read it still. It really looks like the authors had heard of Popek & Goldberg, but failed to read past the abstract.
Another thing they excuse by a vague hand wave in the direction of Popek & Goldberg is the inability of the executing code to detect what CPU mode it is in. This is, again, utter nonsense. x86 exposes this indirectly via POPF (for example) and ARM does not even make you trick it, exposing it completely openly in the CurrentEL MSR. Detecting the current mode on RISC-V is an adventure. It is sometimes possible, but not in all cases. Why might you need this? For example, if you are writing a kernel and want it to support all RISC-V cores. I spent a bit of time trying to make this work for my kernel for [rePalm](/?r=05.Projects&proj=27. rePalm), so I can walk you through the decision tree and show where each branch comes to life and whacks you in the gonads. First of all, if the core has no Zicsr, you are guaranteed to be running in machine mode, but, there is no way to know that there is no Zicsr, other than probing by doing a CSR read, but there are two problems: first, without Zicsr, there is no proper generic way to catch the resulting exception when an invalid instruction trap is generated. Second, what CSR to read? Don’t forget that likely all of them are optional. One might be tempted to go for misa, but do not forget that you are probing what mode you are in. If you are in supervisor mode, that probe would also fail, even if Zicsr was implemented. Ok, you can try reading sstatus. That one is readable to supervisor mode AND machine mode. You’re safe, right? You wish! Supervisor mode is optional, and if your core does not implement it, there is no sstatus register, so … you trap. Ok. Let’s simplify the problem. Let’s assume Zicsr exists. Can you then at least tell apart S mode from M mode? Nope! The following seems like a tempting solution: set stvec to point to your handler that simply adjusts sepc forward by 4 and returns (skipping the faulting instruction), then execute a read of misa. If you were in machine mode, it reads fine. If you were in supervisor mode, it traps, and any sane machine monitor would hand you an illegal instruction trap. Then, your handler would skip the instruction, and you’d note this and conclude you were in supervisor mode. You win, right? Almost… Once again: supervisor mode is optional. Let’s imagine you were in machine mode on a core without supervisor mode support. You’d trap as soon as you tried to set stvec, since it does not exist. You might be tempted to say: why not just catch that trap too? Because to do that, you need to set mtvec, and you cannot be sure you can do that since you might have been in supervisor mode all along. Thus the intersection of everything being optional and the authors’ complete misunderstanding of Popek & Goldberg lands the poor you in a pile of shite.
Despite having an extension for seemingly everything, including operations on the common kitchen sink, somehow a number of obviously-useful instructions are missing. I already covered the lack of register + register addressing modes, so we’ll not bother returning to that. There are a few other obvious low-hanging fruit that were seemingly ignored. And before you argue that RV32I was designed to be simple, all of the things I am about to suggest are trivial in the extreme, mostly reducing to simple wires on an ASIC.
First and foremost: test a bit and branch based on it. This one instruction replaces two (SLLI + BGEZ/BLTZ), but also it does not require a temporary register. To check how common this would be, if it existed, I grabbed a random aarch64 binary (the latest raspian kernel for raspberry pi, “vmlinuz-6.1.0-49-arm64”, sha256: B3B686DE 82CC7B84 EFEB8F6B 309A4E6C 53E7461F 281D3E56 F42E4AA3 B6207075), disassembled it, and counted the number of instances of TBZ/TBNZ. There were 35,393 instances. By comparison, there are 70,109 instances of RET, which means that on average every other function uses or could use an instruction to branch on the value of a bit. Implementing it is trivial in hardware, and indeed branching on a bit is extremely common in dissecting protocols or using bitfields. For big out-of-order cores, renaming is simpler when one fewer register gets clobbered, and also there is no need to try to fuse two instructions when one exists. For smaller MCU cores, where no fusion exists, this is a simple code size and speed win. Why this obvious thing was not done, I do not know.
My next major gripe - bitfield operations - bit field extract and bitfield insert. These are extremely useful for working on things like network packets and hardware registers. Bit field extract can be simulated using two instructions - SLLI + SRLI/SRAI based on the desired signedness. Bitfield insert takes a lot more work to simulate: create the inverse mask, AND with destination reg, shift source reg into place, OR into destination. Depending on the bits, it is 3-6 instructions easily. BFC (bit field clear) is a simpler special case that is also quite useful. It is doable in 2-3 instructions. In hardware though, it is just wires - no complex logic, no nothing! There is no need to be clever like aarch64 is, although the designers of RISC-V could learn a lesson or ten from aarch64 indeed, including clever bitfield handling. I did the same counting exercise using the same kernel image. There are 6,284 bitfield insert instructions and 8,881 bitfield extract instructions – two out of every 9 functions on average use these bitfield ops. To add insult to injury, there is a bit-ops extension for RISC-V – Zbs. By looking at it, you can tell the authors were academics. It is clean, simple, easy to explain, elegant, and completely useless. Who the hell ever needs to extract just one bit? Seriously, what a missed opportunity.
RISC-V is the first architecture I’ve ever encountered which scatters immediates randomly throughout the instruction with no immediately-clear reason for it. This makes emulating it a huge pain since it takes so very long to recombobulate the immediate values, compared to architectures like MIPS, or ARM. The former simply uses bits 0..15 for immediates, the latter has fancier encodings, but at least there are just a few. RISC-V designers’ justifications for this were “the same bits of the immediate come from the same bits of the instruction” – a justification so idiotic that it physically hurts to attempt to pretend to believe it. It is the sort of thing that a software person who’s never written verilog would think helps make things easier. You see, no matter how you spin it, you will need a mux for immediates, since they are of differing lengths and with differing number of trailing zeroes, depending on the instruction. And that mux, well … it does not care even a little which instruction bits are wired into its inputs, it is all just wires. But even if the reasoning for “why” is idiotic, let us inspect the claim itself, to see if they did accomplish what they claimed to have wanted. Do the same bits of immediates always come from the same place? Let’s take a look. For I-type instructions, bit 1 of the immediate comes from instruction bit 21, bit 11 of the immediate comes from instruction bit 31. For S-type instructions, the same bits of immediates come from instruction bits 8 and 31 respectively. For B-type instructions, they come from bits 8 and 7 respectively. And for J-type instructions, they come from bits 21 and 20 respectively. As you see, they indeed always come from the same place, as promised, if we merely ignore the meanings of the words “same” and “place”.
But this just barely touches the surface of the insanity of the encodings! For J-type instructions, the immediate value is scattered in the following order: 20 10 9 8 7 6 5 4 3 2 1 11 19 18 17 16 15 14 13 12. What possible justification could you imagine for this insanity, other than that the designers confused the chatter of a bingo parlor for the proper bit order for immediates. And yet, this is nothing compared to the mess that they made of the compressed instruction set…
There are no fewer than 9 (nine!) instruction formats here, and that does not include Zcb, which adds 8 more! But even that is not all! Depending on the instruction, the same format (eg: CI) encodes immediates differently in the same bit positions. Accounting for all of that, it is almost one instruction encoding format per instruction! It did not need to be like this! Thumb and microMIPS both give examples how to not fuck up this badly, and yet, despite easy availability of examples of how to do it right, RISC-V designers bid us hold their collective LSD-laced beers and went at it in the most pessimal way imaginable. Let us first, of course, look at the immediates in C, since “they always come from the same place in the instruction word”, you know ;)
When talking about immediate encodings henceforth, I shall use “x” to indicate when the immediate is broken into pieces and there is something else there in the instruction. L.LWSP (which loads a word from the stack) encodes the immediate in this order: 5 x x x x x 4 3 2 7 6, C.SWSP which is its sibling for storing to stack, instead, encodes the immediate as 5 4 3 2 7 6. C.LW (which loads a word from memory addressed by a register) encodes its immediate as: 5 4 3 x x x 2 6, naturally. Its sibling, C.SW uses the same encoding, indicating that the design team missed an opportunity to scramble some bits here. If you wanted to load a byte from memory, you’d use C.LB, whose immediate encoding is, of course, nothing like the above. It uses bit order: 0 1. If you wanted to load a halfword, you’d use C.LHU, whose bit order is just: 1. Because of course it is! I am not even going to touch on CM.PUSH and CM.POP because their encoding is so complex that the spec spends a whole chapter explaining how to decode them. Truly, a sign of a simple and intuitive encoding, if you ask me.
While claiming to make use of all possible encoding space in the 16-bit instruction space, some fruit remain so low-hanging as to require OSHA warnings! A simple example: logical shifts include 6 bits of immediate. The argument is that this is needed for 64-bit instructions, but there are already many encodings in the compressed instruction set that are RV64-only. That is to say that RV64C and RV32C are already incompatible. Given that, why the hell are all shift instructions in RV32C carrying an extra zero bit? You cannot shift a 32-bit register by more than 32 bits. The spec says that bit must be zero, and yet no encoding uses the space opened up by that bit being one. Self delusion is telling yourself you are making good use of encoding space while also carrying around the ability to shift 32-bit registers by 61 bits, my friends. RV32E is even more egregious since many of its compressed instructions carry an extra bit to encode registers that do not exist there. As RV32E is ABI-incompatible with RV32I and they would never share code, giving RV32E more useful encodings by using those extra bits that would have encoded registers x16..x31 would have been an excellent idea. That is probably why it was not done.
Moving on… C.J’s bit order could likely pass the NIST Statistical Test Suite for random number generators: 11 4 9 8 10 6 7 3 2 1 5. What even‽ C.BEQZ/C.BNEZ are also jumps, though conditional, so of course their encodings have almost nothing in common with the previous one, they use: 8 4 3 x x x 7 6 2 1 5. When encoding immediates for C.LI, the bit order is 5 x x x x x 4 3 2 1 0. That almost looks sane, but do not despair, more fun is coming. Let us say you wish to adjust the stack pointer. C.ADDI16SP is here for you, with its immediate encoded as: 9 x x x x 4 6 8 7 5. And if you wanted to get an address of a stack variable, C.ADDI4SPN is there, with its immediate encoded as: 5 4 9 8 7 6 2 3. All very logical, sane, and clear, as promised.
Imagine you are a CPU (or an emulator) trying to decide how to decode an instruction. If you are a sane CPU, most likely it goes like this: look at 1-3 bits to determine instruction format, from there look at 2-5 bits to figure out the instruction, and you are done, you are ready to execute. If you are RISC-V, things are a bit … more complex. First you look at the bottom 2 bits to figure out if the instr is 2 or 4 bytes, for 2-byte instructions, you look at the top three bits to determine what instruction this is. So far, so good. But then… you notice that this instruction has an immediate. You need to reassemble the jigsaw puzzle that that is. And then you recall that some register-register instructions are encoded in the immediate format, with a magic immediate value indicating that they are register-register ops. For example C.NOT is encoded this way, the magic value being 0b111101. C.ZEXT.W (if your hardware implements the proper mishmash of extensions to have it at all) uses the magic immediate 0b111100. Somehow microMIPS and Thumb managed to do without this insanity. How? Ancient secrets that apparently were irrecoverably lost before the RISC-V authors were born.