CPUSim64 Cycle Timing Model

Instructions on a modern RISC processor each take multiple cycles to pass through the CPU's execution stages (this is the instruction's latency). However, because these stages are pipelined—allowing several instructions to be processed simultaneously at different stages—the processor can often complete roughly one instruction per cycle in terms of throughput.

Some instructions break this ideal. Operations such as division are typically not fully pipelined and take many cycles, while events like interrupts incur additional overhead from flushing the pipeline and saving processor state. This table documents the effective number of cycles attributed to each instruction.

InstructionFormCycles
Simple ALU
NOP—1
CLEAR—1
MOVEreg-reg1
COMPL—1
AND, OR, XOR—1
TEST, CMP—1
LSHIFT, RSHIFT, ARSHIFT—1
LROTATE, RROTATE—1
PACK, PACK64, UNPACK, UNPACK64—1
ENDIAN—1
READONLY—1
Arithmetic
NEGATEinteger1
NEGATEFP3
ADD, SUBTRACTinteger1
ADD, SUBTRACTFP3
MULTIPLYinteger3
MULTIPLYFP3
DIVIDEinteger12
DIVIDE, RECIPFP10
Memory
LOAD—2
STORE—2
PUSH—2
POP—2
SAVEN registers1 + N
RESTOREN registers1 + N
CAS—3
Control Flow
JUMPunconditional1
JUMPconditional, taken2
JUMPconditional, not taken1
CALLunconditional3
CALLconditional, taken4
CALLconditional, not taken1
RETURN—3
STOP—1
I/O & System
IN—1
OUT—1
INTERRUPT—11
INTERRUPTconditional, not taken1
DEBUG—1

Notes