FOLLOW EACH STEP
CPU & von Neumann architecture
A program is a set of instructions. But where are they kept, who reads them, and how does a result reach a screen? Follow one input all the way through the machine.
Start with a task: read a number, add 5, display the answer.
Suppose you enter 7. To obtain 12, the computer needs more than the number 7: it also needs instructions telling it to read, add, and output. The von Neumann stored-program idea puts instructions and data in the same main memory. The CPU—the central processing unit—reads and executes those instructions.
Memory contains numbered locations. A location’s number is its address; the information stored there is its content. Asking for address 12 is like choosing a numbered drawer. It does not mean that the drawer must contain the number 12.
Information is represented by bits: each bit is a 0 or a 1. A bus is a set of connections carrying signals. An address bus identifies where to access; a data bus carries bits; control signals specify actions such as READ or WRITE. Instructions also travel as bits on the data bus. These colours distinguish signal roles, not three types of instruction.
How does the CPU know what to do next?
It needs small places to keep information while it works. These are registers: fast storage locations inside the CPU. The program counter (PC) remembers the next instruction’s address. The instruction register (IR) keeps the instruction being executed. The control unit (CU) reads that instruction and directs the other parts. The arithmetic and logic unit (ALU) does the calculation.
First fetch an instruction from memory. Next decode it: identify the operation and any operand, the value or location it uses. Then execute it. Each of these phases can require several smaller transfers, shown below as micro-operations. Press Step to complete one transfer; Play follows the whole program.
Functional cutaway · signal motion is slowed for inspection
Choose the next instruction
PC holds the address of the next instruction. Copy that address to MAR, which holds the location for a memory access.
Select a component: why does it need to exist?
PC · Program counter
Remembers the address of the next instruction. Otherwise the CPU would not know where to continue. In this word-addressed machine every instruction occupies one word, so sequential execution adds 1 to PC.
The Blender model shows a functional cutaway, with extra space between parts so the signals can be followed. It is not the transistor layout of a particular commercial chip. This machine uses one accumulator, 12-bit words, a 4-bit address space, and explicit input/output instructions. Modern CPUs may arrange these functions very differently.
The memory stores bits. Where did “ADD” come from?
A bit has two possible values, 0 and 1. A word is a fixed-size group of bits; here it has 12. Our instruction uses its first four bits as the operation code, or opcode. The final eight bits contain an operand field. For ADD and STORE, the low four operand bits identify one of the 16 memory locations; the remaining four are zero.
| Readable instruction | Opcode | operand | Meaning |
|---|---|---|
IN | 0001 | 00000000 | Input → ACC |
ADD [12] | 0010 | 00001100 | ACC + memory[12] → ACC |
STORE [13] | 0011 | 00001101 | ACC → memory[13] |
OUT | 0100 | 00000000 | ACC → output |
HALT | 0101 | 00000000 | Stop fetching instructions |
For example, 0010 selects ADD; 00001100 represents 12. So 0010 00001100 means “add the value at address 12”. IR keeps these bits, and the control unit decodes them. ADD is the readable name of that code, not a word that the hardware understands as English.
Why do instruction and data registers stay separate?
While executing ADD [12], MDR must receive the number 5 from memory. If MDR were our only copy of the instruction, that read would replace the instruction with 5. Fetching copied the instruction into IR first, so we can read data without losing the operation we are carrying out.
The clock coordinates when stored values change. A program instruction is not necessarily one clock cycle: fetching, decoding, and execution can need several cycles. The animation slows transfers so you can inspect them; its seconds are not a CPU speed measurement.
Must the next instruction wait for everything to finish?
Now that we can follow one instruction, imagine dividing its work between five stages. Once the first instruction leaves the fetch stage, that stage can begin fetching another instruction while the first is decoded elsewhere. This overlapping arrangement is a pipeline. It is an implementation technique, not a requirement of the von Neumann idea.
Read the instruction from memory.
Identify the operation and obtain its inputs.
Perform arithmetic or calculate an address.
Read or write data if the instruction needs it.
Store a result in its destination register.
The next demonstration uses a separate five-stage machine with registers R1, R2, … rather than a single accumulator. Ri names storage location i in the register bank. Each row is an instruction, each column one clock cycle. The gold banks between stages are pipeline registers: they hold intermediate values, destination numbers, and control bits between clock edges.
| Instruction / cycle → | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| I1: R1 ← R2 + R3 | IF | ID | EX | MEM | WB | |||
| I2: R4 ← R6 + R5 | IF | ID | EX | MEM | WB | |||
| I3: R7 ← R8 + R9 | IF | ID | EX | MEM | WB | |||
| I4: R10 ← R11 + R12 | IF | ID | EX | MEM | WB |
Different instructions can use different stages together.
In cycle 3, I1 is executing, I2 is decoding, and I3 is being fetched. Registers between stages retain each instruction’s intermediate information, so the next cycle can move each one forward without mixing them up.
Assume five stages of one cycle each, independent instructions, and no stalls. Completing four instructions separately takes 4 × 5 = 20 cycles. With overlap, the first takes 5 cycles, then one finishes in each of the next 3 cycles: 8 in total. For k instructions:
The first instruction still takes five cycles. What improves is throughput: how many instructions finish over time. Switch to the dependency example to see why real pipelines cannot always finish one instruction per cycle. A result that is not ready forces a wait unless another mechanism, such as forwarding, can safely supply it earlier.
Can fetch and data access use the same memory together?
Not if it has only one access port. A load or store in MEM can conflict with IF, requiring a stall or a different memory arrangement. The instructions in our table only add registers, so their MEM stages pass values along without accessing memory. The example therefore does not assume two simultaneous accesses to our single-port memory.
Check: does a blue address packet contain the answer?
No. It identifies a location. The data signal carries its contents, and a control signal says whether to read or write. To store an answer, the machine needs all three roles.