Intel has built a chip that crunches encrypted data thousands of times faster than its own servers can manage.
Fully homomorphic encryption, or FHE, lets you compute on encrypted data without decrypting it, but it runs like treacle on standard CPUs and GPUs.
People want cloud AI without spilling their secrets, and they want medical risk calculations without handing their genome to whoever runs the server.
Chipzilla thinks Heracles is the way out, after showing the prototype at the IEEE International Solid-State Circuits Conference in San Francisco.
Intel security circuits research lead Sanu Mathew said, “Heracles is the first hardware that works at scale.”
Most FHE research chips sit at about 10 square millimetres or less, but Heracles is roughly 20 times larger and built on 3-nanometre FinFET technology.
It sits in a liquid-cooled package alongside two 24GB high-bandwidth memory chips, a setup usually reserved for GPUs used for AI training.
The demo was a private query to a secure server, dressed up as a voter checking their ballot had been registered correctly.
The state holds an encrypted database, the voter encrypts their ID and vote, and the server checks for a match without decrypting anything in the middle.
On a Chipzilla Xeon server CPU the process took 15 milliseconds, while Heracles did it in 14 microseconds.
At the human scale, that looks like nothing, but at 100 million ballot checks, it amounts to more than 17 days of CPU work, compared to 23 minutes on Heracles.
University of California, Irvine researcher Ro Cammarota said: “We have proven and delivered everything that we promised.”
FHE is a mathematical transformation that uses a quantum-computer-proof algorithm, then swaps in corollaries so the encrypted maths lands where the plain-text maths would.
Intel circuits research lab research scientist Anupam Golder told engineers: “Usually, the size of cypher text is the same as the size of plain text, but for FHE it’s orders of magnitude larger,” he said.
That ciphertext bloat is bad enough, then the compute turns ugly because FHE leans on massive integers that demand precision.
A CPU can do it, but in FHE integer addition and multiplication can take about 10,000 more clock cycles than the plain version.
GPUs do parallelism for breakfast, yet they have been tuned for less precise numbers, while FHE wants exactness and odd jobs like “twiddling” and “automorphism”.
Bootstrapping adds more pain because it is a compute-heavy noise-cancelling process that general-purpose processors handle badly.
Heracles started five years ago under a DARPA programme to accelerate FHE with purpose-built hardware.
Cammarota called it “a whole system-level effort that went all the way from theory and algorithms down to the circuit design,” which is a long walk for a short benchmark.
One early bet was to split enormous numbers into smaller chunks, then push them through 32-bit arithmetic blocks for parallelism without losing the precision FHE needs.
At the centre are 64 compute cores, called tile-pairs, arranged in an eight-by-eight grid, serving as SIMD engines for polynomial maths, twiddling, and the rest.
A 2D mesh network ties the tiles together with wide 512-byte buses, because nothing ruins encrypted compute like starving the cores.
The data footprint is huge, so Chipzilla linked 48GB of high-bandwidth memory to the processor with 819GB-per-second connections.
On chip, data stages in 64MB of cache, then moves through the array at 9.6TB per second by hopping from tile-pair to tile-pair.
To stop movement and maths tripping each other up, Heracles runs three synchronised instruction streams, one for off-chip movement, one for on-chip movement and one for arithmetic.
Heracles runs at 1.2GHz and takes 39 microseconds for a key FHE transformation, which Chipzilla says is 2,355 times quicker than a Xeon at 3.5GHz.
Across seven operations, Chipzilla claims Heracles is 1,074 to 5,547 times as fast, depending on the amount of shuffling required.







