8-bit Multicycle Processor

  • VHDL
  • GHDL
  • Quartus

Case Study

8-bit Multicycle Processor – RTL Verification & DE1 Port Preparation. A VHDL multicycle processor with a fixed internal program, extended from an existing university RTL design, verified with GHDL (including exhaustive ALU suites) and prepared for a Terasic DE1 port.

Context & Challenge

The starting point was an existing 8-bit multicycle processor RTL design from university coursework: a control FSM (FETCH, DECODE, LOAD_IMM, WAIT_UAL, CAPTURE, WRITEBACK, HALT) driving a fixed 16x8 instruction ROM, a 12-operation ALU with a start/done handshake, and result/flag registers. The challenge was to turn that design into something defensibly correct and board-ready: build a real verification environment around it, find and fix any boundary bugs, and prepare it for FPGA integration, without rewriting the processor’s own architecture.

Scope & Responsibilities

This repository extends and adapts an existing RTL design from university coursework; the complete processor was not built from scratch. What is specifically attributable to this repository’s work: the verification environment (all eight GHDL testbenches and the independent reference model in tb/tb_pkg.vhd), a program-counter boundary fix, the Terasic DE1 port preparation (wrapper, Quartus project, constraints and pin assignments), the continuous-integration workflow, and the documentation. The processor’s control FSM, ALU and datapath architecture are the original coursework design.

The program-counter fix addresses a specific boundary bug: PC is declared with range 0..15, but in the imported RTL both increments used a plain PC + 1, so fetching ROM[15] took PC out of range and stopped simulation before HALT was decoded. Both increments now use explicit modulo-16 arithmetic, and the regression exercises exactly this boundary.

Architecture / System Design

The datapath separates a fixed instruction ROM, an instruction register, two operand registers (A, B) loaded only from immediate ROM bytes, a registered ALU (ual_8) reached through a start/done handshake, and a result register with flags Z/N/C/V. ual_8 composes five combinational or sequential sub-units: an 8-bit adder/subtractor, logic unit, shifter, an unsigned 8x8 shift-add multiplier, and a 16-bit integer square-root core. R is written only by ALU operations and is never fed back into the datapath: there is no data memory, register file, branch, or stack, and instructions are limited to LOAD_A, LOAD_B, HALT and single ALU operations selected by the low nibble of the instruction byte.

Implementation

The fixed ROM program loads A = 0x19 and B = 0x07, then exercises every ALU operation except NOT and the unused opcodes, ending in HALT. Instruction timing is an architectural property of the control FSM and datapath: a LOAD sequence spans 3 clock edges, a single-step ALU operation 8 edges, MUL 16 edges, and SQRT 17 edges, derived directly from the FSM state graph and sub-unit latencies rather than measured runtime benchmarks. The DE1 wrapper instantiates the processor RTL unchanged and adds only board glue: a demonstration clock divider (giving an 8 Hz core clock so the ~110-edge result trace is observable), a two-register reset synchronizer, and seven-segment decoding for the result display.

Verification & Evidence

Every testbench is a self-checking GHDL testbench compared against an independent reference model that uses plain integer arithmetic rather than the RTL’s own bit-slice structure. Six of the eight testbenches are exhaustive: 1573888 vectors vectors are applied across the adder/subtractor, logic unit, shifter, multiplier, square-root core and the full ALU, and this scope applies only to the ALU and its sub-units, not to verification of the complete processor. The remaining two testbenches check the processor itself edge by edge against an instruction-level model, including the ROM[15]/HALT boundary and an asynchronous reset mid-operation. Across all eight testbenches together, the full repository regression result is 8 passed / 0 failed.

DE1 integration is prepared and statically reviewed: true.

Historical DE1 execution is a separate matter from the status of the current revision. The author attests that an earlier version of this design ran on a physical Terasic DE1 during the university coursework; this is an author attestation only, because no Quartus output, programming file, programmer log, report or photo is retained as independent evidence. That earlier version predates the modulo-16 program-counter fix, so it says nothing about the current fixed revision or the current wrapper, neither of which is independently evidenced on hardware.

Results

Across the test suites, the six exhaustive ALU/sub-unit testbenches (6 passed / 0 failed) and the full eight-testbench processor regression (8 passed / 0 failed) pass against the reference model, with the scope of each figure kept distinct: the ALU/sub-unit vector count is a separate, narrower figure from the eight-testbench full-regression result. No Fmax, place-and-route result, utilization figure or other implementation metric is claimed, and no FPGA-board execution of the current revision is claimed; the historical coursework execution is author-attested only, as described above.

Limitations

For the current revision and wrapper, no Quartus compilation, fit, assembly, DE1 programming or TimeQuest result is independently established, and no timing closure is claimed: current DE1 readiness is established only at the GHDL synthesis-check and static-review level, and the historical DE1 execution is author-attested only. The instruction set has no branches, data memory, register file or bus arbitration, and the ROM program is fixed at synthesis time. The multiplier silently truncates its high byte without setting C or V.

Key Takeaways

Taking an existing, non-trivial RTL design and making it trustworthy required building a verification environment from nothing: an independent reference model, exhaustive coverage where the state space allows it, and edge-by-edge processor checks where it does not. It also meant being precise about what “extends” a design actually means – fixing a real boundary bug without touching the architecture it sits inside, and being explicit in the regression itself about what has, and has not, been established for the FPGA target.

Artifacts / References

Public repository: github.com/mahdidou711/vhdl-8bit-multicycle-processor.