The FPGA Chronicles: Exploring The Tang Nano 20K

FPGAs used to be mysterious, expensive devices, but these days you can buy surprisingly capable boards for very little money. Some years ago, I did an FPGA Bootcamp over on Hackaday.io. Much of that material still applies, but the hardware is dated. So I decided it was time to update it, using the inexpensive Tang Nano 20K and its GOWIN GW2AR-18 FPGA as the main platform, with perhaps a few excursions into other FPGAs.

History and Motivation

Once upon a time, if you wanted to have a custom IC, you went with a wheelbarrow full of money to a semiconductor company. However, some smart person at a semiconductor fab eventually realized they could make a chip with a lot of uncommitted blocks on it and then, for a custom chip, only design the wiring that connected them together. This still required a wheelbarrow full of money, but it was a smaller wheelbarrow.

Then one day, someone realized they could do the same thing but make the electrical connections between the blocks configurable. Maybe have fuses you can blow, or use EEPROM or RAM cells to remember which blocks are connected to which. It is complicated, sure, but then you can make many of these chips and sell them to people who could, in theory, make their own custom chips without your help.

When do you need an FPGA? A classic classroom exercise for an FPGA, for example, is a traffic light because it shows off how to do state machines, which are important for some kinds of FPGA designs. But other than as a learning example, why would you do this? Even a simple 8-bit CPU can handle a traffic light.

Suppose instead that you have hundreds of digital sensors on a rocket, and any one of them must raise an alarm within a few microseconds. A processor has to sample inputs in groups, service interrupts, or rely on extra hardware. An FPGA can simply implement the equivalent of one enormous OR gate. It watches every input continuously, and unrelated logic elsewhere in the FPGA does not steal execution time from it. Can you do it with a microcontroller? Probably, but not easily. For some classes of problems, an FPGA is the better answer.

Of course, you can also build a CPU on your FPGA and some FPGAs have CPUs in the same package. This is often a sweet spot because then things that are easy to do in software, you do in software. Things that are easier to do in hardware, you do in the FPGA.

Continue reading “The FPGA Chronicles: Exploring The Tang Nano 20K” →

FPGA For All — CERN Releases “colibri” VHDL Library

Since you’re reading Hackaday, we’re pretty sure CERN needs no introduction, so we’ll get right to it– they’re giving back again, this time with a VHDL library called “colibri” containing over 100 components, functions, and procedures to help jumpstart your next FPGA project.

Like a lot of what CERN gives away under its CERN Open Hardware Licence, this library and the functions in it were developed in house to make CERN run better– specifically to streamline the development of gateway devices. As you might imagine, with the prodigious amount of data CERN’s various experiments spit out, FPGAs have become a key part of many of them. The best part is that because the fine folks at CERN don’t want to get locked in, everything here is vendor-independent and has been tested on multiple platforms. Speaking of tested, you get self-checking testbenches in there to make sure everything’s working, and there’s even formal verification, at least for some things. We’ve seen formal verification in software compilers, but its not common in the FPGA world. The whole thing is on GitLab if you want to take a look.

While CERN’s library might not have much to help you make a ternary processor or bus controller with everyone’s favourite programmable silicon, much like software libraries you can save some time at least not implementing say, SPI or i2c– both of which are in colibri, along with a whole lot more.

Thanks to [Alberto Perro] for the tip!

Homebrew 68K Machine Has A PCI Bus

The Peripheral Component Interconnect (PCI) bus was first introduced all the way back in 1992. It quickly became the standard way to interface add-on cards on the PC platform, supplanting earlier buses like ISA and various other oddball standards. You wouldn’t expect to see a PCI bus on a Motorola-based machine, but [maniek86]’s homebrew rig offers just that. 

That’s a lot of soldering.

This computer is a beautiful piece of homebrew engineering, constructed out of protoboard and loose wires rather than any fancy PCB. At the heart of the build lies a Motorola 68000 running at 10 MHz. It’s got 1 MB of SRAM, 4 KB of ROM, and a MC68681P acting as a UART, timer source, and I/O controller. Where things get special, though, is in the inclusion of a Xilinx Spartan II FPGA (XC2S100), which acts as a PCI bridge. It provides the machine with two 32-bit 5-volt PCI slots which are interrupt capable, albeit with no bus mastering. A XC95144XL CPLD also sits present to act as glue logic to help lace everything together.

[maniek86] does a great job of explaining exactly why the PCI bus was hard to implement, and how it was pulled off in the end. The guide also covers how the system was able to interface various cards, from a PCI serial expansion to a Cirrus VGA adapter. It’s all good stuff.

We’ve featured other work from [maniek86] before, too, like this brilliant 486-based single-board computer. Video after the break.

Continue reading “Homebrew 68K Machine Has A PCI Bus” →

Performance Improvements For Open-Source 80386

The Intel 80386 is a rather fascinating slice of computer history. It marked the first 32 bit X86 processor, and was a staple of early desktop computing. Like all chips, it has a number of quirks, one of which being the fact that all commands are executed in microcode. By this nature, it was a rather excellent prospect to be re-implemented in an FPGA core called the z386. However, it was lacking a feature native to the original 386, early start memory access. So to bring some performance to the z386 project, [nand2mario] went forth to fully implement this feature for FPGA 80386s.  

Instead of taking a cycle to find and allocate the memory required for executing the next instruction, the 386 would start this in the previous cycle. This is achieved in hardware by nature of having a separate memory management unit. In the FPGA, the key difficulty proved to be in getting the computation fast enough to execute within a single cycle. This change netted an approximate 9% performance benefit. However, for [nand2mario] this was too small a performance uplift. 

Some rewrites of the store cue allowed for cutting a cycle out of the process further improving the performance. However, more performance required slight deviations from the design of the original 386. Because code-branches are performance critical, the z386 project now computes the branch memory jump several cycles earlier than the 386, reducing the cycle time for the jumps from 9.25 to a mere 6. Some final changes to the microcode decode frontend rounded out the optimizations covered in this latest blog post.

The net result is an approximate 39% increase in performance in the all important DOOM benchmark. The z386 still not a complete project, the performance is still lacking compared to the 386, and it remains unable to boot Windows. X86 is complicated, which will take time, so make sure to stay tuned for more coverage! While you wait, make sure to check out our original writeup of the z386 project. 

Pauli Rautakorpi, CC BY 3.0.

 

 

Breaking Enigma With An FPGA, Just Like At Bletchley Park

The pioneering work done by Alan Turing and others at Bletchley Park in England was perhaps as important in the history of technology as it was the history of the war. Given the last 80-odd years of technological development, their revolutionary work should be within the realms of a student project — which it was, specifically in ECE 5760 at Cornell University. The work was done by [Erica Jiang], [Kelvin Resch], and [Isabella Frank].

Nowadays if someone told you there was a code to be broken, you wouldn’t be reaching for electromechanical devices, but you just might think of trying an FPGA. After all, the programmable gate arrays allow for much faster execution of fixed logic than software running on a traditional CPU. That won’t help much with modern RSA schemes, and for Enigma, it’s massively overkill, but doing it that way was a great learning opportunity for the students.

Their project emulates the whole Bletchley Park cryptography apparatus, not just the Bombe Machine, and if you’re interested in learning about this piece of history you could absolutely do worse than to examine their documentation. If you’re into video, you can check out the final presentation and demo video below. Meanwhile if you’re wondering what the opposition was up to, we have good explainer of the enigma machine here.

Continue reading “Breaking Enigma With An FPGA, Just Like At Bletchley Park” →

Z386: An Open-Source 80386 Built Around Original Microcode

There are many ways you can implement an Intel i386 CPU on an FPGA, with the use of original microcode probably being one of the most interesting approaches. This is what [nand2mario]’s z386 project does, with a recent blog post summarizing the development so far.

This effort is similar to the previously developed z8086 project, which as one may guess does something similar, except for the Intel 8086 CPU. By executing the original microcode you’re basically guaranteeing close compatibility with the original hardware, though of course the sheer scale of this microcode between an 8086 and 80386 is quite different.

There’s a much larger instruction set with a correspondingly much more complicated internal state to keep track of, including all those newfangled features like memory management, paging and register debugging, as well extensions to protected mode that began with the i286.

Currently z386 runs on a number of FPGAs, including the Altera Cyclone V and Gowin GW5A, with performance equivalent to a ~70 MHz i386 albeit with slightly worse cycle efficiency, some of which could be due to the limited 16 kB cache compared to the 32+ kB cache in the fastest i386 CPUs. Either way, it’s more than enough to run all kinds of software, including games like DOOM.

Important to note is that the goal here isn’t to be more performant than cores such as for example ao486, but more as an archaeological reconstruction of the original hardware and its interaction with said microcode.


Top image: line-up of Intel 286, 386 and 486 CPUs. (Credit: Sgroey, Wikimedia)

Build The CPU, Then Build The Calculator

It’s possible that among Hackaday readers are the largest community of people who have designed their own CPU in the world. We have featured many here, but it’s possible that not so many of them have gone on to power an everyday project. Step forward [Baltazar Studios] then, with a scientific calculator sporting a self-designed CPU on an FPGA.

The calculator itself is nice enough, with a smart 3D printed case, an OLED display which almost evokes a VFD, and very well made buttons. But it’s the CPU which is of most interest, because while it follows a conventional Harvard architecture with a 12-bit instruction set, it works with 4-bit nibbles. This choice follows one used by HP in their calculator designs, seemingly because it can be optimised for the binary coded decimal which the calculator uses.

With calculators being yet another app on our spartphones or comnputers, there seems to be less use of calculators outside of education in 2026. But if you are a calculator user there’s nothing like a calculator you made yourself, and with a CPU of your own design it has few equals. We like this project almost as much as we like the Flapulator!