A computer program looks very abstract when we write it in C, Python, or another high-level language. But underneath the abstraction, the CPU eventually has to execute a sequence of bits. Consider this simple operation:
movl $42, -4(%rbp)
In plain English, this means: Take the number 42 and store it at the memory address obtained by taking the value in the RBP register and subtracting 4. Suppose RBP contains 0x2000. Then the target address is:
0x2000 - 4 = 0x1FFC
The assembler converts the assembly instruction into machine code:
c7 45 fc 2a 00 00 00
These seven bytes are what the CPU actually fetches from memory. This gives us a useful ladder:
High-level code ↓Assembly ↓Machine code ↓Binary bits ↓Electrical signals ↓Transistors and logic gates
The interesting part is that the CPU does not treat those seven bytes as an opaque blob. The x86-64 architecture defines exactly how the bytes should be interpreted.
The machine code has structure
For this particular instruction:
c7 45 fc 2a 00 00 00
the bytes have different roles:
c7 45 fc 2a 00 00 00│ │ │ ││ │ │ └── immediate value: 42│ │ └────────── displacement: -4│ └────────────────── ModR/M byte└─────────────────────────── opcode
C7 is the opcode. An opcode, short for operation code, tells the CPU what kind of operation this instruction represents.
The 45 byte is the ModR/M byte. It tells the CPU how to interpret the memory operand—in this case, using RBP as the base register with an 8-bit displacement. FC represents the displacement -4. Because x86 uses two’s-complement signed integers for this displacement, 0xFC represents -4. Finally: 2a 00 00 00 is the 32-bit immediate value. x86 stores multi-byte values in little-endian order, so these bytes represent:
0x0000002A = 42
So the CPU can reconstruct the meaning:
operation: movesource: constant 42destination: memoryaddress: RBP - 4
The assembly language is therefore not the thing the CPU directly understands. Assembly is a human-readable representation of the machine-code instructions defined by the CPU architecture.
What does an opcode actually do?
It is tempting to imagine that the CPU has a little table saying:
C7 → WRITE 42 TO MEMORY
A modern CPU contains enormous numbers of transistors arranged into functional circuits: arithmetic units, registers, multiplexers, shifters, branch logic, address-generation logic, load/store units, and many other components. The opcode helps determine which of these circuits should be activated and how data should flow through them. For example, different instructions require different operations.
- An
ADDinstruction needs arithmetic circuitry. - A shift instruction needs a shifter.
- A load instruction needs the load path.
- A store instruction needs the store path.
- A branch instruction needs branch and control-flow circuitry.
The opcode is therefore part of the information that ultimately generates the internal control signals that configure these circuits.
The CPU does not “know” what C7 means in the human sense. There is no little dictionary inside the processor.
Instead, the bit pattern enters a large network of electronic logic. That logic produces the control signals required to make the rest of the CPU perform the operation specified by the architecture.
From bytes to bits
Now take the first byte:
C7
In binary:
11000111
Physically, these are not little printed 1s and 0s.
They are electrical states.
A binary 1 can correspond to a relatively high voltage, while a binary 0 corresponds to a relatively low voltage. CMOS transistors are used to create circuits that reliably recognize and transform these voltage levels.
So conceptually, we can think of the instruction arriving at the decoder like this:
1 1 0 0 0 1 1 1│ │ │ │ │ │ │ │└─┴─┴─┴─┴─┴─┴─┴── electrical signals
The instruction decoder is itself built from transistors and logic gates.
When the appropriate combination of bits arrives, the decoder generates internal control signals.
For a store instruction, those signals eventually configure the CPU’s data path so that:
source data ↓store path ↓memory address ↓cache / memory system
Other circuits that are not needed for this instruction are not selected for the operation.
This is the important transition:
Machine code is not interpreted by software inside the CPU. The instruction bits control hardware logic that causes other hardware to operate.
Following the instruction through the CPU
We can simplify the process into several stages.
1. Fetch
The CPU has a program counter called RIP on x86-64. It identifies where the next instruction is located.
The instruction-fetch hardware uses that address to retrieve the instruction bytes.
For our example:
c7 45 fc 2a 00 00 00
The bytes are brought into the CPU and placed into internal structures used by the instruction-decoding machinery.
Conceptually:
RIP ↓Instruction Fetch ↓c7 45 fc 2a 00 00 00
Modern x86 CPUs are much more complicated than this simplified picture because they fetch and decode multiple instructions, use caches, queues, prediction, and other mechanisms. But this model is useful for understanding the basic idea.
2. Decode
The instruction bytes enter the decoder.
The decoder recognizes the opcode and the other instruction fields.
Conceptually:
c7 45 fc 2a 00 00 00 ↓ Decoder ↓ ┌───────────────────────┐ │ Store operation │ │ Base register = RBP │ │ Offset = -4 │ │ Immediate = 42 │ └───────────────────────┘
The decoder does not execute the instruction itself. It generates the information and control signals needed by the execution hardware.
The register file can then provide the current value of RBP.
Suppose:
RBP = 0x2000
The immediate data path provides:
42
and the address-generation hardware computes:
0x2000 + (-4) =0x1FFC
3. Execute
The address-generation unit calculates the destination address.
RBP = 0x2000displacement = -40x2000 + (-4) = 0x1FFC
At the same time, the immediate value is available as the data that needs to be stored:
42
The CPU now has the two essential pieces:
Address: 0x1FFCData: 42
4. Store
The CPU’s load/store machinery sends the store toward the memory hierarchy.
But here we need to make an important distinction.
The CPU normally does not directly take the address 0x1FFC and activate a DRAM word line itself.
The store normally goes first through the CPU’s cache hierarchy.
Conceptually:
CPU │ │ address = 0x1FFC │ data = 42 ↓Store / Cache │ ├── cache hit → cache is updated │ └── eventually → memory system → DRAM
The exact behavior is much more complicated in a modern multicore processor because of cache lines, store buffers, write-back policies, coherence, and memory ordering.
But the fundamental idea remains:
The CPU has transformed the original instruction into an electrical operation: put this data at this address.
What does the memory system actually receive?
At the CPU/memory-system level, we can simplify the store as:
Address:0001 1111 1111 1100 = 0x1FFCData:00000000 00000000 00000000 00101010 = 42Operation:WRITE
There is no longer any concept of:
movl $42, -4(%rbp)
The memory system does not know about assembly language.
It receives addresses, data, and control information represented by electrical signals.
This is the same pattern we have seen throughout computer hardware:
Software meaning ↓Instruction ↓Machine-code bits ↓Control signals ↓Data movement ↓Electrical voltage levels ↓Transistor switching
And underneath the cache?
If the data eventually has to be written to DRAM, the story continues.
The memory controller translates the memory request into the signaling required by the DRAM device.
Inside the DRAM chip is a huge array of memory cells.
The earlier example of rows, columns, word lines, and bit lines belongs inside the memory device, not directly to the x86 instruction decoder.
Conceptually:
CPU ↓Cache ↓Memory Controller ↓DRAM interface ↓DRAM row/column selection ↓Memory cells
A DRAM cell is fundamentally built from a transistor and a capacitor. The capacitor’s stored charge represents the bit, while the transistor provides controlled access to the cell.
The memory controller and DRAM circuitry work together to select the appropriate physical location and transfer the data.
This is different from SRAM cache, where the basic memory cell is built from a transistor-based latch and does not use a storage capacitor in the same way.
At the very bottom, none of these components understand C, assembly, or even the number 42 as a human concept.
There are only transistors switching between electrical states.
That is the remarkable continuity of a computer:
movl $42, -4(%rbp) ↓c7 45 fc 2a 00 00 00 ↓binary 0s and 1s ↓control signals ↓registers + arithmetic + address generation ↓cache / memory controller ↓DRAM circuitry ↓transistors switching
The layers look completely different to us, but each layer is simply a more concrete representation of the layer above it.
- A programmer sees an instruction.
- The assembler sees an encoding.
- The CPU decoder sees bit patterns.
- The hardware sees voltage levels.
And the transistor only responds to those voltage levels by changing whether current can flow.