Sequential Search & Multiplier

  • VHDL
  • GHDL
  • Quartus

Case Study

Sequential Search & Multiplier. Two small, self-contained sequential VHDL designs – searchx, a sequential search engine over a fixed table, and mult_shift_add, an unsigned 8x8 iterative shift-add multiplier – each verified with a deterministic-latency GHDL testbench and packaged as a minimal Terasic DE1 demo.

Context & Challenge

Both designs originated from university FPGA coursework and were later cleaned up, given deterministic-latency testbenches, and packaged as a small, reproducible portfolio project sharing one architectural theme: simple FSM-controlled sequential processing (IDLE / working state / DONE), registered (“captured”) transaction inputs so later port changes cannot corrupt an in-flight transaction, and no combinational shortcuts.

Scope & Responsibilities

Provenance differs between the two cores. searchx – its RTL, verification, FPGA integration, CI and documentation – is Mehdi Bouama’s work. The mult_shift_add RTL was originally written by a fellow student on the assignment and is included in this repository with his permission; Mehdi Bouama’s contribution to it is its exhaustive verification, FPGA integration, CI and documentation, not its original design.

Architecture / System Design

searchx searches a fixed, hard-coded table of 16 distinct 8-bit constants: an FSM (IDLE -> SEARCH -> DONE) captures the search value on the acceptance edge, then compares exactly one table entry per clock cycle while in SEARCH, exposing a matched index or a not-found flag. mult_shift_add is a classic unsigned 8x8 shift-add multiplier: each of 8 CALC cycles conditionally adds the captured multiplicand to the running partial product when the current multiplier LSB is 1, then shifts the combined register and the captured multiplier right by one bit.

Implementation

Both cores share the capture-on-acceptance discipline: searchx’s search value and mult_shift_add’s operands are registered on the cycle a transaction is accepted, so later changes on those input ports cannot affect an in-flight transaction. Two independent, minimal Quartus DE1 projects package each core through a thin board-glue wrapper with no shared top-level and no unused peripherals: the multiplier’s wrapper exposes SW(7:0)/SW(8)/SW(9) for operand entry and launch and HEX0..HEX3 for the 16-bit product; the search wrapper exposes SW(7:0)/SW9 for the search value and debut, and LEDR(3:0)/LEDG(0)/LEDG(1) for the index, completion and not-found result.

Verification & Evidence

searchx’s testbench verifies all 16 table positions with their exact expected latency, several absent-value cases, held and pulsed debut behavior, asynchronous reset mid-search, and a capture-timing test that changes the search input immediately after acceptance to confirm the core still completes on the captured value: true. mult_shift_add’s testbench covers directed products, the exact 8-iteration latency, operand-capture behavior and reset mid-calculation, plus an exhaustive sweep of all 65,536 unsigned 8x8 operand combinations, which completes with zero mismatches against the expected product.

DE1 integration for both cores is prepared and statically reviewed: true. This repository’s continuous integration has already run: true.

Historical university coursework hardware experience is documented separately from the current public wrappers. For the multiplier, the mult_shift_add core ran on a physical Terasic DE1: the historical Quartus project completed Analysis & Synthesis and the Fitter, TimeQuest was run, programming files exist, and a retained photo is compatible with 13 x 11 = 143 = 0x008F. That historical core is byte-identical to the pinned core, but the historical top-level wrapper is not the current public wrapper, which is not hardware-validated. For the search core, a historical version was tested on a physical DE1 according to the coursework report and a historical Quartus implementation exists; that version predates the current X_reg transaction-capture fix, so the current search version is not independently hardware-validated.

  • Historical coursework DE1 execution of mult_shift_add: Measured, FPGA board, Historical coursework mult_shift_add core on a physical DE1 (original coursework wrapper; retained photo compatible with 13 x 11 = 143 = 0x008F); excluding Hardware execution of the current public wrappers, Exhaustive hardware verification, Timing closure and Fmax

    true

    • Provenance recorded — source not public
    • Provenance recorded — source not public
    • Provenance recorded — source not public
  • Historical coursework Quartus compilation flow of the multiplier: Source-validated, Static source review, Historical coursework Quartus compilation flow of the multiplier (Analysis & Synthesis, Fitter, TimeQuest stage executed); excluding Hardware execution of the current public wrappers, Timing closure and Fmax

    • Provenance recorded — source not public
  • Historical coursework Quartus programming files of the multiplier: Prepared, FPGA board, Historical coursework Quartus programming files of the multiplier (.sof/.pof/.cdf); excluding Hardware execution of the current public wrappers, Timing closure and Fmax

    • Provenance recorded — source not public
    • Provenance recorded — source not public
    • Provenance recorded — source not public
  • Historical coursework DE1 test of searchx (before the X_reg fix): Measured, FPGA board, Historical coursework version of searchx on a physical DE1 (before the current X_reg transaction-capture fix); excluding Hardware execution of the current search version (with the X_reg capture fix), Hardware execution of the current public wrappers, Timing closure and Fmax

    true

    • Provenance recorded — source not public
  • Historical coursework Quartus compilation flow of the search core: Source-validated, Static source review, Historical coursework Quartus compilation flow of the search core (before the current X_reg transaction-capture fix; Analysis & Synthesis, Fitter, TimeQuest stage executed); excluding Hardware execution of the current search version (with the X_reg capture fix), Timing closure and Fmax

    • Provenance recorded — source not public
  • Historical coursework Quartus programming files of the search core: Prepared, FPGA board, Historical coursework Quartus programming files of the search core (before the current X_reg transaction-capture fix; .sof/.pof); excluding Hardware execution of the current search version (with the X_reg capture fix), Timing closure and Fmax

    • Provenance recorded — source not public
    • Provenance recorded — source not public

Results

searchx completes with verified cycle-exact latency across every table position and absent-value case, and mult_shift_add completes its exhaustive sweep of all 65,536 unsigned 8x8 operand combinations with zero mismatches against the expected product.

No Fmax, timing-closure result or other physical-implementation metric is claimed, and no hardware validation is claimed for the current public wrappers or for the current search version.

Limitations

Both cores are deliberately small: a fixed-width 8x8 unsigned multiplier with no signed support or generic width, a fixed 16-entry search table rather than a parameterizable memory, no bus protocol around either core, no pipelining, and no debounce or synchronizer logic on the DE1 wrappers beyond what is noted above. For the current public wrappers, Quartus Analysis & Synthesis, Fitter, Assembler and TimeQuest results are not established and no timing closure is claimed; the historical coursework hardware experience described above does not extend to them. No license is currently provided for the repository’s FPGA integration layer pending resolution of its vendor-material provenance.

Key Takeaways

Packaging someone else’s RTL responsibly meant being explicit about exactly which parts of the repository are attributable to which author, and applying the same exhaustive-verification discipline to inherited RTL as to newly written RTL rather than trusting it by default. Building both cores around the same capture-on-acceptance FSM pattern also made the deterministic-latency testbenches reusable in shape, even though the two cores solve unrelated problems.

Artifacts / References

Public repository: github.com/mahdidou711/vhdl-search-multiplier.