Joseph Junior Mensah
BF CPU
← BACK TO PROJECTS
SiliconVerilogCPU2026

▦BF CPU

A CPU taped out on real silicon (Tiny Tapeout, GF180MCU) that runs the Brainmess esolang — a multicycle core with no memory of its own, fetching and executing every instruction over an SPI link to off-chip RAM.

Timeline

2026

Team

Joseph Wan & Joseph Mensah

Role

RTL design & verification

Skills

Silicon, Verilog, CPU

Built with

Verilog·cocotb·SPI·OpenLane·GF180MCU·TinyTapeout

LONG  STORY  SHORT

I taped out a real chip that runs Brainmess — a CPU with no memory of its own, reaching every byte of program and data over SPI.

Brainmess is a minimal, Turing-complete language with exactly eight instructions: > < move a data pointer, + - change the cell under it, . , do output and input, and [ ] form loops. It's small enough to build in hardware and complete enough that doing so means building a real CPU — a fetch/decode/execute machine with a program counter, a data pointer, and control flow.

So that's what this is: a multicycle Brainmess CPU, written in Verilog, run through the open-source ASIC flow, and manufactured as a 1×1 Tiny Tapeout tile on the GF180MCU process. It is a physical chip that runs programs.

A close crop of the hardened layout: standard-cell rows, local interconnect, and routed metal compressed into one Tiny Tapeout tile.
A close crop of the hardened layout: standard-cell rows, local interconnect, and routed metal compressed into one Tiny Tapeout tile.

What the chip looks like up close

The full GDS render is beautiful, but the useful detail is easier to see when you zoom in. The repeated green horizontal bands are standard-cell rows; the brighter yellow, orange, red, and blue shapes are routed layers and vias tying the cells together. This is the point where the project stops being a Verilog file and becomes geometry the foundry can manufacture.

That density is also why the architecture has to stay simple. There is no room to casually add large memories or wide datapaths, so the design leans into a small control core, a narrow byte datapath, and an external-memory protocol.

How it works

The core is a ten-state FSM — fetch, decode, the read/modify/write states for + -, the I/O states for . ,, and forward/backward scan states for loops — around a program counter and data pointer (both 16-bit) and 8-bit instruction and data registers. Loops are the interesting part: a small 6-deep hardware stack remembers the address of each open [, so a matching ] jumps back in one step; nesting deeper than the stack falls back to scanning the program for the matching bracket, so correctness never depends on the stack being big enough.

On-chip core

BF

  • 10-state multicycle FSM
  • PC + data pointer (16b)
  • 6-deep bracket stack
⇄mem read / write

On-chip bridge

SPI master

  • 32-bit command frames
  • 0x03 read / 0x02 write
  • one transaction per access
⇄SPI · mode 0

Off-chip

RP2040 SPI RAM

  • 64 KB, 23LC512 emulation
  • instructions 0x0000–0x7FFF
  • data tape 0x8000–0xFFFF
No on-chip RAM — the core reaches every byte of program and tape over SPI.

Fast path

[[[[[[

A six-entry hardware stack remembers recent loop-open PCs.

overflow→

Correctness path

+[+[-].]

If nesting is deeper, the FSM scans forward or backward to find the matching bracket.

Common loops jump in one step; deeper programs still run correctly by scanning for the matching bracket.

No memory of its own

The defining constraint: the chip has no on-chip RAM. Both the program and the data tape live off-chip in an RP2040 running Michael Bell's spi-ram-emu, which impersonates a 23LC512 — SPI mode 0, 0x03 to read, 0x02 to write, 16-bit addresses. A single 16-bit space is split in half: 0x0000–0x7FFF holds instructions, 0x8000–0xFFFF is the data tape.

Everything the core does therefore travels as an SPI transaction. An on-chip SPI master serialises each request into a 32-bit command/address/data frame and shifts it out MSB-first. Fetching an instruction is one read; . and , are one transfer each; and + or - is a read-modify-write — two transactions to touch a single cell.

0x0000-0x7FFF

Instruction memory

BF opcodes are fetched from the low half of the off-chip RAM.

0x8000-0xFFFF

Data tape

The data pointer walks byte cells in the high half of memory.

programsingle 16-bit address spacetape
The RP2040-backed RAM behaves like one 64 KB memory; the CPU treats the lower half as code and the upper half as the BF tape.
One 16-bit address space, split down the middle: instructions low, the data tape high.
One 16-bit address space, split down the middle: instructions low, the data tape high.
A routing-heavy crop from the GDS render — the physical consequence of turning each memory access into wires, cells, and metal.
A routing-heavy crop from the GDS render — the physical consequence of turning each memory access into wires, cells, and metal.

One instruction's journey

Because memory is remote, a single instruction is a little pipeline of bus traffic rather than a one-cycle affair — the FSM walks it through fetch, decode, and, for the arithmetic ops, a full read-modify-write against the tape before advancing.

  1. 1

    Fetch

    instr over SPI

  2. 2

    Decode

    one of 8 ops

  3. 3

    Read cell

    SPI read at DP

  4. 4

    Modify

    + / − / I/O

  5. 5

    Write cell

    SPI write back

  6. 6

    Advance

    PC++ or jump

Every instruction is at least one SPI transaction; + and − are a read-modify-write, so two.

CMD

0x03 / 0x02

ADDR[15:8]

high byte

ADDR[7:0]

low byte

DATA

8 bits

CS low32 SCK edges
Every memory access leaves the tile as a command, address, and data byte shifted through the SPI pins.

The pins

A Tiny Tapeout tile gives you eight dedicated inputs, eight outputs, and eight bidirectional pins — and this design uses all of them. The input and output bytes map straight onto Brainmess's , and .; the bidirectional pins carry the SPI link to memory plus a small handshake so the host can feed input, latch output, and start the machine.

ui[7:0]

The input byte the , instruction reads.

8 dedicated input pins

uo[7:0]

The output byte the . instruction writes.

8 dedicated output pins

uio[0:3]

SPI master out to the off-chip RAM.

CS · MOSI · MISO · SCK

uio[4:7]

Host I/O handshake and run control.

start · out_valid · in_valid · in_ack

One Tiny Tapeout tile: eight in, eight out, eight bidirectional.
A lower-left layout crop, where the regular edge structures make the tile boundary and I/O neighborhood easier to read.
A lower-left layout crop, where the regular edge structures make the tile boundary and I/O neighborhood easier to read.

Verifying a chip you can't probe

Silicon is unforgiving: once a design is taped out, there is no patch. So the design is tested at three levels before it ever becomes layout — the SPI master on its own, the BF core against a fake memory, and the whole chip against a mock SPI RAM, each with cocotb testbenches.

The SPI master

Driven against a mock SPI RAM, checking every frame — command, address, data, and timing.

Makefile.spi · cocotb

The BF core, isolated

The FSM run against a fake memory, verifying each instruction and nested-loop behaviour.

Makefile.bf · cocotb

The whole chip

The full design against a mock SPI RAM, end to end — real programs, real output.

Makefile.top · cocotb

Three levels of cocotb tests — a chip can't be probed after tapeout, so it's proven before.

From RTL to silicon

The Verilog runs through the open-source hardening flow and comes out as a placed-and-routed layout: standard cells packed into rows, wired across metal layers, dropped into one Tiny Tapeout tile. The image is the actual gds_render of this project — and it's explorable in 3D through the Tiny Tapeout GDS viewer linked above.

It's one thing to simulate a CPU; it's another to hold the constraint that there is no memory here and design a machine that works anyway by talking to the world one SPI frame at a time.

The finished die — standard cells placed and routed into a single GF180MCU tile.
The finished die — standard cells placed and routed into a single GF180MCU tile.

What comes next

The physical chip is expected later this fall, packaged and ready for bring-up. That is the part I am most excited for now: taking the design past simulation and layout, wiring it into the real Tiny Tapeout board environment, and proving that the packaged silicon behaves like the RTL said it would.

The next step is full post-silicon validation — loading real BF programs, exercising the SPI RAM path, checking the input/output handshake, and comparing observed behavior against the cocotb tests that shaped the design. I am looking forward to that process because it is where the project becomes more than "the build passed": it becomes a chance to learn how real chips are powered up, measured, debugged, and trusted.

Why I built it this way

  • A real CPU, not a toy. Brainmess's eight instructions still demand a program counter, a data pointer, control flow and I/O — the whole fetch/decode/execute loop, in hardware.
  • Own the memory constraint. No on-chip RAM forced a clean split: a core that computes, and an SPI master that fetches — every access an explicit transaction.
  • Loops in hardware, with a fallback. A 6-deep bracket stack makes common nesting O(1), while a scan path keeps arbitrarily deep programs correct.
  • Prove it before it's permanent. Three levels of cocotb tests — bus, core, and full chip — because you cannot debug a chip after it ships.

NEXT  UP  …