VibeKoding / Ensiklopedia Β· Fondasi KuatEnsiklopedia Β· Fondasi Kuat / Computer Organization PrinciplesComputer Organization Principles
VK

Computer Organization PrinciplesComputer Organization Principles

πŸ“š Ensiklopedia Β· Fondasi KuatEnsiklopedia Β· Fondasi Kuat 🌏 Dual Bahasa (ID / EN) ⚑ VibeKoding Native

Ensiklopedia VibeKoding: Computer Organization Principles.Ensiklopedia VibeKoding: Computer Organization Principles.

πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

After building a CPU from transistors, how do computers form a complete system? In the previous chapter, we started from transistors and built adders, registers, arithmetic units, and finally assembled the CPU core. But a CPU alone is not enough β€” it needs to work in coordination with memory and I/O devices, requires buses to connect components, and needs an instruction set to drive operations. This chapter shifts our perspective from inside the CPU to the entire computer system, providing an in-depth understanding of the Von Neumann architecture, instruction sets, storage hierarchy, buses, and I/O principles.After building a CPU from transistors, how do computers form a complete system? In the previous chapter, we started from transistors and built adders, registers, arithmetic units, and finally assembled the CPU core. But a CPU alone is not enough β€” it needs to work in coordination with memory and I/O devices, requires buses to connect components, and needs an instruction set to drive operations. This chapter shifts our perspective from inside the CPU to the entire computer system, providing an in-depth understanding of the Von Neumann architecture, instruction sets, storage hierarchy, buses, and I/O principles.

What will you learn from this article?What will you learn from this article?

After completing this chapter, you will gain:After completing this chapter, you will gain:

ChapterContentCore Concepts
Chapter 1Von Neumann ArchitectureStored program, five major components, data path
Chapter 2Instruction Set ArchitectureInstruction format, addressing modes, CISC vs RISC
Chapter 3CPU Control UnitControl unit, micro-operations, instruction cycle
Chapter 4Storage HierarchyCache, main memory, virtual memory, paging
Chapter 5Bus and I/OBus arbitration, DMA, interrupt mechanism

------

0. Big Picture: Computer Hardware System0. Big Picture: Computer Hardware System

In the previous chapter "From Transistors to CPU," we understood how the CPU works internally β€” from fetch, decode, execute to write-back. But the CPU itself is just an execution unit. To make a computer truly "usable," a series of peripheral components must work together.In the previous chapter "From Transistors to CPU," we understood how the CPU works internally β€” from fetch, decode, execute to write-back. But the CPU itself is just an execution unit. To make a computer truly "usable," a series of peripheral components must work together.

πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

- Layer 1: CPU Core Responsible for instruction execution, including the control unit (issuing control signals) and the arithmetic unit (performing arithmetic and logic operations) - Layer 2: Register File High-speed storage units inside the CPU, including general-purpose registers and special-purpose registers (PC, IR, MAR, MDR, etc.) - Layer 3: Main Memory Memory for storing programs and data, accessed by the CPU through address and data buses - Layer 4: I/O Devices Input and output devices connected to the system bus through I/O controllers - Layer 5: System Bus Data channels connecting CPU, memory, and I/O, including address bus, data bus, and control bus- Layer 1: CPU Core Responsible for instruction execution, including the control unit (issuing control signals) and the arithmetic unit (performing arithmetic and logic operations) - Layer 2: Register File High-speed storage units inside the CPU, including general-purpose registers and special-purpose registers (PC, IR, MAR, MDR, etc.) - Layer 3: Main Memory Memory for storing programs and data, accessed by the CPU through address and data buses - Layer 4: I/O Devices Input and output devices connected to the system bus through I/O controllers - Layer 5: System Bus Data channels connecting CPU, memory, and I/O, including address bus, data bus, and control bus

------

1. Von Neumann Architecture: The Foundation of Modern Computers1. Von Neumann Architecture: The Foundation of Modern Computers

1.1 The Stored-Program Principle1.1 The Stored-Program Principle

In 1945, mathematician John von Neumann proposed the groundbreaking stored-program architecture concept. This idea laid the foundation for modern computers.In 1945, mathematician John von Neumann proposed the groundbreaking stored-program architecture concept. This idea laid the foundation for modern computers.

πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

Stored Program: The program itself, as a special kind of data, is stored in memory just like ordinary data. The CPU can read and execute program instructions stored in memory the same way it reads and writes data.Stored Program: The program itself, as a special kind of data, is stored in memory just like ordinary data. The CPU can read and execute program instructions stored in memory the same way it reads and writes data.

This means:This means:

1.2 Five Major Components1.2 Five Major Components

The Von Neumann architecture divides a computer into five core components:The Von Neumann architecture divides a computer into five core components:

ComponentEnglishFunctionMain Composition
Arithmetic UnitALU (Arithmetic Logic Unit)Performs arithmetic and logic operationsAdders, shifters, comparators
Control UnitCU (Control Unit)Directs and coordinates all componentsInstruction register, decoder, timing generator
MemoryMemoryStores programs and dataMemory Address Register (MAR), Memory Data Register (MDR)
Input DevicesInputInformation inputKeyboard, mouse, scanner
Output DevicesOutputInformation outputMonitor, printer

1.3 Data Path1.3 Data Path

The Data Path is the route through which data flows between functional components. Inside the CPU, the data path connects:The Data Path is the route through which data flows between functional components. Inside the CPU, the data path connects:

The width of the data path (how many bits can be transferred at once) directly affects computer performance.The width of the data path (how many bits can be transferred at once) directly affects computer performance.

1.4 The Von Neumann Bottleneck1.4 The Von Neumann Bottleneck

The Von Neumann architecture has a famous performance bottleneck:The Von Neumann architecture has a famous performance bottleneck:

> The data transfer speed between CPU and memory is far lower than the CPU's processing speed.> The data transfer speed between CPU and memory is far lower than the CPU's processing speed.

This causes the CPU to frequently be in an idle "waiting for data" state. Many optimization techniques in modern computers revolve around this problem:This causes the CPU to frequently be in an idle "waiting for data" state. Many optimization techniques in modern computers revolve around this problem:

Optimization TechniquePrinciple
CachePlace small, high-speed storage near the CPU
Instruction PipelineAllow multiple instructions to be in different stages simultaneously
SuperscalarIssue multiple instructions in the same clock cycle
Multi-core ParallelismMultiple CPU cores share computing tasks

------

2. Instruction Set Architecture: The Interface Between CPU and Software2. Instruction Set Architecture: The Interface Between CPU and Software

In the previous section, we learned the core idea of the Von Neumann architecture: programs are stored in memory just like data. But this raises a key question β€” what does a "program" stored in memory actually look like? How does the CPU understand it?In the previous section, we learned the core idea of the Von Neumann architecture: programs are stored in memory just like data. But this raises a key question β€” what does a "program" stored in memory actually look like? How does the CPU understand it?

The answer is the Instruction Set Architecture (ISA). If the CPU is a service, then the instruction set is its API documentation β€” it defines all the commands the CPU can understand, the format of each command, and the data range each command can operate on. Every line of code you write is ultimately translated by the compiler into a sequence of calls to this "API."The answer is the Instruction Set Architecture (ISA). If the CPU is a service, then the instruction set is its API documentation β€” it defines all the commands the CPU can understand, the format of each command, and the data range each command can operate on. Every line of code you write is ultimately translated by the compiler into a sequence of calls to this "API."

2.1 From Code to Instructions: Instruction Encoding and Execution2.1 From Code to Instructions: Instruction Encoding and Execution

First, let's establish a holistic understanding: the code you write in an editor and what the CPU actually executes are separated by several layers of translation.First, let's establish a holistic understanding: the code you write in an editor and what the CPU actually executes are separated by several layers of translation.

This translation chain is key to understanding instruction sets:This translation chain is key to understanding instruction sets:

LayerContentWho Can Understand It
High-level languageint a = 10 + 5;Humans
Assembly languageMOV R1, #10 / ADD R3, R1, R2Humans (with training)
Machine code0001 0001 0000 1010CPU
πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

- When you see a compilation error, you know it occurred at the "high-level language β†’ assembly" step - When you see a runtime crash, you know the problem is at the CPU instruction execution stage - When understanding performance optimization, you know what optimizations the compiler makes during "translation" - When choosing a CPU architecture (x86 vs ARM), you know the difference is in the "instruction set API"- When you see a compilation error, you know it occurred at the "high-level language β†’ assembly" step - When you see a runtime crash, you know the problem is at the CPU instruction execution stage - When understanding performance optimization, you know what optimizations the compiler makes during "translation" - When choosing a CPU architecture (x86 vs ARM), you know the difference is in the "instruction set API"

2.2 Instruction Formats2.2 Instruction Formats

Now that we know code gets translated into instructions, the next question is: what is the internal structure of an instruction?Now that we know code gets translated into instructions, the next question is: what is the internal structure of an instruction?

Each machine instruction is essentially a string of binary digits, but it has a strict internal format. The two most core parts are:Each machine instruction is essentially a string of binary digits, but it has a strict internal format. The two most core parts are:

Just as a sentence has a "verb + object" structure, an instruction has an "operation + target" structure:Just as a sentence has a "verb + object" structure, an instruction has an "operation + target" structure:

CODE
Instruction: ADD R3, R1, R2 ─── ────────── Opcode Operands (do addition) (R3 = R1 + R2)

Based on the number of operands, instruction formats range from simple to complex in four types:Based on the number of operands, instruction formats range from simple to complex in four types:

FormatStructureExampleUse Case
Zero-addressOpcode onlyRET (return)Stack machines; operands implicit at stack top
One-addressOpcode + 1 addressINC R1 (increment R1 by 1)Single-operand operations
Two-addressOpcode + 2 addressesMOV R1, R2Most common; data transfer and operations
Three-addressOpcode + 3 addressesADD R3, R1, R2Preserves source operands
πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

This is a trade-off between space and flexibility. Zero-address instructions are the shortest (saving memory) but require extra stack operations; three-address instructions are the most flexible (preserving source data) but occupy more bits. Different CPU architectures choose different combinations of instruction formats.This is a trade-off between space and flexibility. Zero-address instructions are the shortest (saving memory) but require extra stack operations; three-address instructions are the most flexible (preserving source data) but occupy more bits. Different CPU architectures choose different combinations of instruction formats.

2.3 Addressing Modes2.3 Addressing Modes

An instruction tells the CPU to "do addition," but where are the two numbers for the addition? They might be written directly in the instruction, in a register, or at some memory address. Addressing modes are the rules that tell the CPU "where to find the operands."An instruction tells the CPU to "do addition," but where are the two numbers for the addition? They might be written directly in the instruction, in a register, or at some memory address. Addressing modes are the rules that tell the CPU "where to find the operands."

Using a real-life analogy of "finding someone":Using a real-life analogy of "finding someone":

Addressing ModeAnalogyInstruction ExampleDescription
Immediate addressingThe person is standing right in front of youMOV R1, #100Data is written directly in the instruction; fastest
Register addressingCall an internal extension to reach a colleagueMOV R1, R2Data is in a CPU register; very fast
Direct addressingKnow the address, go directly to the doorMOV R1, [0x1000]Memory address is written in the instruction
Indirect addressingAsk the front desk "which room is Zhang San in?"MOV R1, [R2]The register contains an address; requires one extra lookup
Indexed addressing"Building 3 + Floor 5" to calculate the roomMOV R1, [R2+10]Base address + offset; used for array access

πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

Different scenarios require different "find data" strategies: - Constant assignment (x = 100) β†’ Immediate addressing; data is right in the instruction - Variable operations (a + b) β†’ Register addressing; data already loaded into registers - Array access (arr[i]) β†’ Indexed addressing; base address + index offset - Pointer operations (*ptr) β†’ Indirect addressing; register holds the address When you write arr[i], you don't think about addressing modes, but the compiler automatically selects the most appropriate one.Different scenarios require different "find data" strategies: - Constant assignment (x = 100) β†’ Immediate addressing; data is right in the instruction - Variable operations (a + b) β†’ Register addressing; data already loaded into registers - Array access (arr[i]) β†’ Indexed addressing; base address + index offset - Pointer operations (*ptr) β†’ Indirect addressing; register holds the address When you write arr[i], you don't think about addressing modes, but the compiler automatically selects the most appropriate one.

2.4 The CPU's Capability List β€” Instruction Categories2.4 The CPU's Capability List β€” Instruction Categories

Now that we know instruction formats and addressing modes, the last question is: what exactly can the CPU do?Now that we know instruction formats and addressing modes, the last question is: what exactly can the CPU do?

All instructions can be grouped into six categories that cover everything a computer can do:All instructions can be grouped into six categories that cover everything a computer can do:

TypeWhat It DoesRepresentative InstructionsMaps to Your Code
Data TransferMove data aroundMOV, LOAD, STORElet x = y, function parameter passing
ArithmeticAddition, subtraction, multiplication, divisionADD, SUB, MUL, DIVa + b, count++
LogicBitwise operationsAND, OR, NOT, XORflags & 0xFF, permission checks
ShiftShift left/rightSHL, SHRx << 2 (equivalent to multiplying by 4)
Control TransferJumps and callsJMP, CALL, RETif, for, function calls
Input/OutputCommunicate with peripheralsIN, OUTRead keyboard, write to screen
πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

All the code you write β€” no matter how complex the business logic or how flashy the UI animations β€” is ultimately broken down into combinations of these six basic operation types. The CPU's "intelligence" lies not in how complex its individual operations are, but in its ability to execute these simple operations billions of times per second.All the code you write β€” no matter how complex the business logic or how flashy the UI animations β€” is ultimately broken down into combinations of these six basic operation types. The CPU's "intelligence" lies not in how complex its individual operations are, but in its ability to execute these simple operations billions of times per second.

2.5 Two Design Philosophies: CISC vs RISC2.5 Two Design Philosophies: CISC vs RISC

There is a fundamental divide in instruction set design: should each instruction be as powerful as possible, or as simple as possible?There is a fundamental divide in instruction set design: should each instruction be as powerful as possible, or as simple as possible?

This divide created two camps that directly affect every device you use today:This divide created two camps that directly affect every device you use today:

An analogy to understand this:An analogy to understand this:

πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

- x86 (CISC) has dominated the PC and server market for 40 years, accumulating a massive software ecosystem. Switching architectures means all software must be recompiled - ARM (RISC) dominates mobile devices thanks to its low-power advantage. Phone batteries are small; every milliwatt counts - Apple Silicon proved that RISC can also deliver high performance β€” the M-series chips surpassed x86 competitors in both performance and power efficiency - RISC-V is an open-source RISC architecture rapidly rising in IoT, education, and AI chip sectors- x86 (CISC) has dominated the PC and server market for 40 years, accumulating a massive software ecosystem. Switching architectures means all software must be recompiled - ARM (RISC) dominates mobile devices thanks to its low-power advantage. Phone batteries are small; every milliwatt counts - Apple Silicon proved that RISC can also deliver high performance β€” the M-series chips surpassed x86 competitors in both performance and power efficiency - RISC-V is an open-source RISC architecture rapidly rising in IoT, education, and AI chip sectors

------

> Summary: The instruction set is the bridge between software and hardware. Your code is translated by the compiler into instructions, which tell the CPU what to do and to whom through opcodes and operands. Addressing modes determine where data comes from. Different instruction set designs (CISC/RISC) determine the CPU's performance characteristics and applicable scenarios.> Summary: The instruction set is the bridge between software and hardware. Your code is translated by the compiler into instructions, which tell the CPU what to do and to whom through opcodes and operands. Addressing modes determine where data comes from. Different instruction set designs (CISC/RISC) determine the CPU's performance characteristics and applicable scenarios.

>>

> Now we know the "static structure" of instructions β€” what they look like and what types exist. The next question is: how does the CPU execute these instructions step by step internally? That's the control unit's job.> Now we know the "static structure" of instructions β€” what they look like and what types exist. The next question is: how does the CPU execute these instructions step by step internally? That's the control unit's job.

------

3. Control Unit: The CPU's Control Logic3. Control Unit: The CPU's Control Logic

3.1 Components of the Control Unit3.1 Components of the Control Unit

The control unit is the "brain" of the CPU, responsible for coordinating all components to work according to instruction requirements:The control unit is the "brain" of the CPU, responsible for coordinating all components to work according to instruction requirements:

ComponentFunction
Program Counter (PC)Stores the address of the next instruction
Instruction Register (IR)Stores the currently executing instruction
Instruction DecoderParses the instruction's opcode and operands
Timing GeneratorGenerates clock beat signals to control component timing
Micro-operation Sequence GeneratorGenerates the series of control signals needed to execute instructions

3.2 Instruction Cycle3.2 Instruction Cycle

The CPU executes an instruction through a complete instruction cycle, typically including:The CPU executes an instruction through a complete instruction cycle, typically including:

  1. Fetch Cycle: Read instruction from memory into IRFetch Cycle: Read instruction from memory into IR
  2. Decode Cycle: Parse the instruction's meaningDecode Cycle: Parse the instruction's meaning
  3. Execute Cycle: Perform the operationExecute Cycle: Perform the operation
  4. Memory Access Cycle: Access memory if neededMemory Access Cycle: Access memory if needed
  5. Write-Back Cycle: Write the result back to a register or memoryWrite-Back Cycle: Write the result back to a register or memory
  6. 3.3 Micro-Operations3.3 Micro-Operations

    Micro-operations are the most basic operations driven by control signals. For example, the "fetch" phase can be decomposed into the following micro-operations:Micro-operations are the most basic operations driven by control signals. For example, the "fetch" phase can be decomposed into the following micro-operations:

    BeatMicro-OperationControl Signals
    T1PC β†’ MARPCout, MARin
    T2MEM β†’ MDRMEMout, MDRin
    T3MDR β†’ IRMDRout, IRin
    T4PC + 1 β†’ PCPC+1, PCin

    3.4 Hardwired vs. Microprogrammed Control3.4 Hardwired vs. Microprogrammed Control

    FeatureHardwired ControlMicroprogrammed Control
    ImplementationCombinational logic circuitsMicroinstruction sequences (firmware)
    SpeedFastSlightly slower
    Design difficultyComplexSimpler
    FlexibilityPoor (changes require circuit redesign)Good (just modify the microprogram)
    Typical applicationsRISC processorsEarly CISC processors

    ------

    4. Storage Hierarchy and Cache Principles4. Storage Hierarchy and Cache Principles

    4.1 Storage Hierarchy Structure4.1 Storage Hierarchy Structure

    A computer's storage devices form a pyramid structure:A computer's storage devices form a pyramid structure:

    LevelStorage TypeAccess TimeTypical CapacityLocation
    RegistersSRAM<1nsA few KBInside CPU
    L1 CacheSRAM~1ns32-64KBNear CPU core
    L2 CacheSRAM~3-10ns256KB-1MBOn CPU chip
    L3 CacheSRAM~10-20ns2-16MBOn CPU chip / shared
    Main Memory (RAM)DRAM~50-100ns8-64GBOn motherboard
    SSDFlash~10-100ΞΌs256GB-2TBOn motherboard
    HDDMagnetic disk~5-10ms1-10TBInside case
    πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

    If CPU accessing L1 cache is like grabbing a piece of paper from your desk: - Accessing main memory β†’ Taking the elevator to a convenience store downstairs to buy paper - Accessing SSD β†’ Driving to another city to buy paper - Accessing HDD β†’ Flying to another country to buy paper The speed difference can be millions of times!If CPU accessing L1 cache is like grabbing a piece of paper from your desk: - Accessing main memory β†’ Taking the elevator to a convenience store downstairs to buy paper - Accessing SSD β†’ Driving to another city to buy paper - Accessing HDD β†’ Flying to another country to buy paper The speed difference can be millions of times!

    4.2 Cache Principles4.2 Cache Principles

    Cache is fast storage located between the CPU and main memory. Its core idea is based on two locality principles:Cache is fast storage located between the CPU and main memory. Its core idea is based on two locality principles:

    πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

    - Temporal locality: If data was just accessed, it's likely to be accessed again soon - Spatial locality: If data was accessed, nearby data is likely to be accessed too- Temporal locality: If data was just accessed, it's likely to be accessed again soon - Spatial locality: If data was accessed, nearby data is likely to be accessed too

    How Cache WorksHow Cache Works

    1. Hit: The data the CPU needs is in the cache; read directlyHit: The data the CPU needs is in the cache; read directly
    2. Miss: The data is not in the cache; must be loaded from main memoryMiss: The data is not in the cache; must be loaded from main memory
    3. CODE
      Hit rate = Number of hits / Total number of accesses Average access time = Hit rate Γ— Cache time + (1 - Hit rate) Γ— Memory time
      

      4.3 Cache Mapping Methods4.3 Cache Mapping Methods

      MethodPrincipleAdvantageDisadvantage
      Direct-mappedEach memory block can only go to one fixed locationSimple and fastHigh conflict rate
      Set-associativeEach memory block can go to N locations (N-way)Balances speed and hit rateMore complex implementation
      Fully associativeAny locationLowest conflict rateHardest to implement (requires comparing all tags)

      4.4 Virtual Memory4.4 Virtual Memory

      Virtual memory is an important abstraction provided by the operating system:Virtual memory is an important abstraction provided by the operating system:

      • Each process thinks it has a complete virtual address spaceEach process thinks it has a complete virtual address space
      • The operating system translates virtual addresses to physical addressesThe operating system translates virtual addresses to physical addresses
      • Infrequently used pages can be swapped out to disk (swap space)Infrequently used pages can be swapped out to disk (swap space)
      πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

      Think of virtual memory as a hotel managing rooms: - You (the process) think the entire building is yours - In reality, the hotel (OS) only assigns you the rooms you currently need - Unused rooms get "swapped out" to storage (disk) - Needed rooms can be "swapped in" at any timeThink of virtual memory as a hotel managing rooms: - You (the process) think the entire building is yours - In reality, the hotel (OS) only assigns you the rooms you currently need - Unused rooms get "swapped out" to storage (disk) - Needed rooms can be "swapped in" at any time

      ------

      5. Bus and I/O Systems5. Bus and I/O Systems

      5.1 System Bus5.1 System Bus

      A Bus is the data channel connecting computer components:A Bus is the data channel connecting computer components:

      Bus TypeFunctionDirectionTypical Width
      Address BusTransfers memory addressesUnidirectional (CPU→Memory)32-bit/64-bit
      Data BusTransfers dataBidirectional32-bit/64-bit
      Control BusTransfers control signalsBidirectionalMultiple signal lines

      5.2 Bus Arbitration5.2 Bus Arbitration

      When multiple devices simultaneously request bus access, an arbitration mechanism determines who goes first:When multiple devices simultaneously request bus access, an arbitration mechanism determines who goes first:

      Arbitration MethodDescription
      Centralized arbitrationA central arbiter makes the decision
      Distributed arbitrationDevices negotiate among themselves

      5.3 I/O Device Access Methods5.3 I/O Device Access Methods

      MethodPrincipleAdvantageDisadvantage
      Programmed I/O (Polling)CPU polls I/O statusSimpleLow CPU utilization
      Interrupt-driven I/OI/O device actively notifies CPU when doneCPU can work in parallelInterrupt handling has overhead
      DMAI/O device accesses memory directlyCPU not involved at allRequires a DMA controller

      5.4 DMA Principles5.4 DMA Principles

      DMA (Direct Memory Access) allows I/O devices to exchange data directly with memory:DMA (Direct Memory Access) allows I/O devices to exchange data directly with memory:

      • Without DMA: The CPU participates in the entire data transfer process and can't do anything elseWithout DMA: The CPU participates in the entire data transfer process and can't do anything else
      • With DMA: The CPU tells the DMA controller "transfer from where to where, how much," then goes to do other tasks. The DMA notifies the CPU when completeWith DMA: The CPU tells the DMA controller "transfer from where to where, how much," then goes to do other tasks. The DMA notifies the CPU when complete
      πŸ’‘ Tips PraktisπŸ’‘ Pro Tip

      This is like ordering food delivery: - Without DMA: You go to the supermarket yourself, buy groceries, go home, wash vegetables, and cook (involved in the entire process) - With DMA: You place an order by phone, and the delivery person brings it straight to your kitchen (someone else handles it; you just "receive the goods" at the end)This is like ordering food delivery: - Without DMA: You go to the supermarket yourself, buy groceries, go home, wash vegetables, and cook (involved in the entire process) - With DMA: You place an order by phone, and the delivery person brings it straight to your kitchen (someone else handles it; you just "receive the goods" at the end)

      5.5 Interrupt Mechanism5.5 Interrupt Mechanism

      Interrupts are a very important mechanism in computer systems:Interrupts are a very important mechanism in computer systems:

      1. After an I/O device completes an operation, it sends an interrupt request to the CPUAfter an I/O device completes an operation, it sends an interrupt request to the CPU
      2. The CPU, currently executing an instruction, responds to the interrupt after completing the current instructionThe CPU, currently executing an instruction, responds to the interrupt after completing the current instruction
      3. The CPU saves its current state and jumps to the interrupt handlerThe CPU saves its current state and jumps to the interrupt handler
      4. After handling is complete, it restores the state and continues executionAfter handling is complete, it restores the state and continues execution
      5. ------

        6. CPU Performance Optimization: Pipeline Technology6. CPU Performance Optimization: Pipeline Technology

        6.1 Instruction Pipeline6.1 Instruction Pipeline

        Instruction pipelining is a parallel technique that maximizes CPU efficiency:Instruction pipelining is a parallel technique that maximizes CPU efficiency:

        How Pipelining WorksHow Pipelining Works

        CODE
        Sequential execution (5 instructions, 15 cycles): Instr 1: IF→ID→EX→MEM→WB Instr 2: IF→ID→EX→MEM→WB Instr 3: IF→ID→EX→MEM→WB ... Pipeline execution (5 instructions, 9 cycles): Instr 1: IF→ID→EX→MEM→WB Instr 2: IF→ID→EX→MEM→WB Instr 3: IF→ID→EX→MEM→WB ...
        

        Ideally, CPI (cycles per instruction) for N instructions β‰ˆ 1Ideally, CPI (cycles per instruction) for N instructions β‰ˆ 1

        6.2 Pipeline Hazards6.2 Pipeline Hazards

        While pipelining improves performance, it also introduces hazard problems:While pipelining improves performance, it also introduces hazard problems:

        TypeCauseSolution
        Structural hazardHardware resource conflictAdd hardware / stagger execution
        Data hazardLater instruction needs the result of an earlier oneData forwarding / bubbles / scheduling
        Control hazardBranch instructions change execution flowDelay slots / branch prediction

        ------

        7. Summary: The Computer Execution Model7. Summary: The Computer Execution Model

        Let's connect the entire process using professional terminology:Let's connect the entire process using professional terminology:

        > After a program starts, the operating system loads the executable file from disk into memory. The CPU's fetch unit (IF) reads instructions from memory into the instruction register (IR) via the address bus. The control unit decodes the instruction (ID), and after identifying the operation type, generates the corresponding control signals. The arithmetic unit (EX) performs arithmetic and logic operations. If memory access is needed, it accesses memory (MEM) via the data bus, and finally the result is written back (WB) to a register or memory. The entire process is driven by the clock, with micro-operation sequences generated by the control unit coordinating all components to work in an orderly manner.> After a program starts, the operating system loads the executable file from disk into memory. The CPU's fetch unit (IF) reads instructions from memory into the instruction register (IR) via the address bus. The control unit decodes the instruction (ID), and after identifying the operation type, generates the corresponding control signals. The arithmetic unit (EX) performs arithmetic and logic operations. If memory access is needed, it accesses memory (MEM) via the data bus, and finally the result is written back (WB) to a register or memory. The entire process is driven by the clock, with micro-operation sequences generated by the control unit coordinating all components to work in an orderly manner.

        ------

        Further ReadingFurther Reading

        TopicRecommended Resources
        Computer ArchitectureComputer Organization and Design: The Hardware/Software Interface - Patterson & Hennessy
        CPU MicroarchitectureComputer Systems: A Programmer's Perspective - Bryant & O'Hallaron
        Instruction Set ArchitectureARMv8 Architecture Manual, Intel x64 Manual
        Cache PrinciplesCache Coherence Protocol (MESI), Cache Write Policies
        Operating SystemsNext chapter: "Operating Systems"

        ------

        Next StepsNext Steps

        Now that you've mastered the professional knowledge of computer organization, you can continue learning:Now that you've mastered the professional knowledge of computer organization, you can continue learning:

        • [Operating Systems](./operating-systems.md): Understand how programs run on an operating system, and how processes, threads, and memory management are implemented[Operating Systems](./operating-systems.md): Understand how programs run on an operating system, and how processes, threads, and memory management are implemented
        • [Data Encoding, Storage, and Transmission](./data-encoding-storage.md): Deepen your understanding of how data is represented in computers[Data Encoding, Storage, and Transmission](./data-encoding-storage.md): Deepen your understanding of how data is represented in computers