Hacker Dudes// alternate deck. same feed.
FEED ~1/h SIGNAL
100%
SPOOLED 0/30 CYCLE 61319.852.53.27
28
AI is now capable of developing its own inference hardwaregithub.com
166pts/191 comments/4h/▼5
by fsbonetto / hn> / « feed
▲▼ fsbonetto16:23 hn>

After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.

▲▼ vatsachak16:40 hn>

I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.

▲▼ andai20:13 hn>

Vibe connoisseuring

▲▼ skybrian16:50 hn>

This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?

▲▼ fsbonetto17:00 hn>

Its a datacenter decommissioned board, really popular among hobbyists.

For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.

▲▼ pcarolan16:53 hn>

Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.

▲▼ hehimself16:55 hn>

They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.

▲▼ traverseda16:56 hn>

I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.

Also can't keep them closed source if you do that.

▲▼ skeskinen16:56 hn>

Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.

Also, it's hard to get fab capacity for any project. Let alone something so experimental.

▲▼ jcims16:59 hn>

Addressing these issues seems to a major driver behind the design of terrafab.

─ 1 more ─
▲▼ zitterbewegung16:58 hn>
▲▼ ohazi16:58 hn>
▲▼ slowin17:25 hn>

I think this company was recently acquired by AMD, so hopefully they'll start getting some this into production. I know OpenAI was working on model-on-a-chip too.

▲▼ yorwba17:28 hn>

8 months ago, Taalas claimed https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcomi... that "Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM. It is expected in our labs this spring and will be integrated into our inference service shortly thereafter. Following this, a frontier LLM will be fabricated using our second-generation silicon platform (HC2). HC2 offers considerably higher density and even faster execution. Deployment is planned for winter."

Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.

▲▼ birdatlaw17:03 hn>

From what I've read, not only are some labs doing it (other commenters already mentioned).

But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.

I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.

I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.

I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.

▲▼ zdragnar17:04 hn>

Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.

It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.

I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

▲▼ fhdkweig17:15 hn>

I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?

─ 14 more ─
▲▼ HoldOnAMinute17:45 hn>

At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.

All aboard! We're racing to the bottom now.

─ 7 more ─
▲▼ jb199117:49 hn>

I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…

─ 1 more ─
▲▼ _puk18:08 hn>

They have trillions..

Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.

─ 1 more ─
▲▼ thesz18:48 hn>

  > Model SOTA moves faster than chips can be designed or produced.
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.

Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252

Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.

▲▼ fsbonetto17:05 hn>

The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.

So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

▲▼ fnordpiglet17:16 hn>

Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.

I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.

Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.

But such is the cycle

─ 2 more ─
▲▼ schleck817:05 hn>

Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.

▲▼ pmarreck17:10 hn>

Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?

▲▼ fsbonetto17:12 hn>

GPUs are faster, but you can't make your own arch on GPUs. FPGAs offer you that possibility. Said that... There are a few beasty FPGAs used in crypto mining coming my way... I expect that OpenTPU will be able to run frontier models with those.

▲▼ dmitrygr17:12 hn>

In addition to some of the other replies you got, here is one more:

Much of a model are weights, and high-density ROMs are very very very hard.

▲▼ __MatrixMan__17:15 hn>

Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?

▲▼ sanderjd17:55 hn>

If I controlled a budget like this, I think I would put some portion of it toward paying to crystallize one of today's models in silicon, yes. Not 100%, but I do think this makes sense to invest in at this point. I would not have said so a year ago.

▲▼ jolt4217:22 hn>

Even dumber question: What is new or novel about this openTPU?

▲▼ fsbonetto17:28 hn>

First opensource arch that can do modern LLMs, while maximizing the potential of its hardware; First opensource TPU build by a recursive improvement loop...

It's upcoming second generation could run the inference of the models that are being used to improve it...

▲▼ samuelknight17:27 hn>

Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.

▲▼ AIblemblio17:29 hn>

We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough.

Your optimized hardware chip might be obsolete before its back from the fab.

SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.

Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.

▲▼ buriram17:37 hn>

Yes, and startups do exactly that. Check out Etched https://www.etched.com/ where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.

▲▼ jjcm17:49 hn>

Most responses here are along the lines of "model capabilites move too fast to build hardware for".

I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.

▲▼ sanderjd17:52 hn>

Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.

─ 2 more ─
▲▼ casta18:35 hn>
▲▼ root_axis18:36 hn>

Not sure that's actually practical at the scale of SOTA models.

▲▼ dualvariable18:43 hn>

This question would be better answered if people were careful about distinguishing between "models" and "transformer architecture".

If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.

▲▼ stronglikedan20:11 hn>

Cuz once they're in chips they can be put into robots, and once they're in robots it won't be so easy to reach the off switch, and once we can't easily reach the off switch, we're doomed.

▲▼ rfgplk16:58 hn>

Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.

▲▼ jetemple17:08 hn>

Which tools have made that leap? Faster design iteration makes sense, but what points to exponential hardware gains rather than shorter development cycles?

▲▼ Simboo19:14 hn>

What tools? Ummmm Synopsis type tools? I like FPGA’s and ASIC chips but you know, the tooling is ecosystem is a big black box of wtf. Verilator is great though. Any tools I should be looking up and looking out for? Thanks!

▲▼ platevoltage19:35 hn>

I'm not exactly in the chip making industry, but I don't think the big hurdle to progression is design capabilities.

▲▼ xg1517:05 hn>

"Recursive self-improvement will kill us all!"

Also: Here is our recursive self-improvement hard at work...

▲▼ lelanthran17:14 hn>
>"Recursive self-improvement will kill us all!"
>Also: Here is our recursive self-improvement hard at work...

Soon we will see

token-providers: "The torment nexus is a cautionary tale"

Also token-providers: "Finally, we have created the torment nexus that we first told you about!"

▲▼ nialse17:17 hn>

All will end up on same plateau eventually. RSI is just a phase on the way there.

▲▼ mrob17:29 hn>

The problem is that plateau is likely far beyond human capabilities. I don't care if ASI progress stalls after it's already killed all biological life as a useless waste of resources.

─ 15 more ─
▲▼ dumberquestions17:18 hn>

Technology has always contributed to improving next iterations of itself, it's only a concern when it's fully autonomous.

▲▼ buellerbueller17:47 hn>

Oh, like a virus?

▲▼ athrowaway3z17:07 hn>

I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December.

The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.

But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.

▲▼ felixgallo17:18 hn>

I suspect an AI could design a purpose-built FPGA-like replacement that would be, for its purpose, significantly more effective than the current general-purpose FPGAs.

▲▼ fsbonetto17:19 hn>

It could have a small improvement on power consumption, but the current design can already achieve 90% of the maximum theoretical speed of this hardware without giving up programability/flexibility

▲▼ chris_money20217:36 hn>

There doesn't exist a single FPGA that can fit an entire AI ASIC. You would need dozens stitched together, then comes the issue of clock speeds, FPGAs typically run far below reference. There also memory issues with FPGAs.

Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.

▲▼ bitwize17:11 hn>

Colossus is building Colossus II.

▲▼ rcarmo17:45 hn>

Feelis like working at Magrathea...

▲▼ ASalazarMX18:24 hn>

Colossus/Guardian engineers: We created AGI twice, simultaneously, on the first try, and we weren't even aiming for it!

It is still a very good read, but the machines are much better written than the people.

▲▼ srameshc17:17 hn>

This post brings me to question "What does it mean to be a software developer in future" ?

▲▼ amelius17:34 hn>

Basically, an unemployed plumber.

▲▼ ASalazarMX18:28 hn>

My bet is that promptgrammers will be so common they'll become standard full-stack engineers, and the few experienced programmers that still know software engineering from the ground up will become expensive gurus sought for critical tasks.

The elite gurus will get paid handsomely, while promptgrammers will be paid less since they've become a less-skilled commodity, and the company has to pay for the expensive tokens they'll avidly consume.

I've seen someone jump from Wordpress to deploying internet-facing APIs because 'they have PHP experience', and the holes in their knowledge were filled blindly by an LLM. I have also argued with a seasoned developer about how their code didn't need linting because LLMs 'already follow best practices'.

The future doesn't look bright when LLMs allow future generations to feign required knowledge.

▲▼ altcognito19:22 hn>

They may be running a linter and don't even realize it. Many LLMs do this by default, aka, as you say, follow best practices.

▲▼ BenzeneDream20:10 hn>

Of course, that will last for a time, until it won't.

I don't see a reason why expert humans will remain more expert than AIs.

▲▼ andai20:13 hn>

What of the workhorse in the face of mechanical workmen? Time for a mechanical pension!

▲▼ AnimalMuppet17:20 hn>

Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)

▲▼ chris_money20218:01 hn>

This is the smallest unit of a typical AI ASIC, for example Google's TPU would have several dozen more compute units inside of it per chip.

In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.

So its missing ALOT

▲▼ AnimalMuppet20:08 hn>

Thanks. But that wasn't my question. For this part, how is the performance? State of the art? Better? Or worse?

▲▼ gfalcao17:44 hn>

The birth of SkyNet

▲▼ rcarmo17:44 hn>

Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...

▲▼ QuantumNomad_17:47 hn>

Humans allegedly already took care of that

https://youtube.com/shorts/TC2jGXr0fig

▲▼ figassis17:56 hn>

Getting roundhouse kicked to extinction woudl not be an unfun way to go. We might even be proud of having passed the torch. These would not be boring inheritors to earth.

─ 1 more ─
▲▼ altmanaltman18:04 hn>

Seems too complex when you can just create a basic metal casing that can kill people. Why would they care if its anatomically accurate or not, it doesn't need us to relate to the characters like the movies do

▲▼ mbgerring17:59 hn>
>AI is now capable of developing its own inference hardware

No, it isn't.

A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.

▲▼ holmesworcester18:06 hn>

Can't LLMs also prompt LLMs?

Are we confident that no existing LLM is capable of similarly effective prompts to those this author used? (I agree it's a stretch, but would not reject it out of hand.)

Even if not yet, will the existence of this repo soon change that, because LLMs will soon ingest it?

▲▼ dpoloncsak18:18 hn>

An LLM can prompt an LLM when first prompted by a human.

I think OP is trying to convey the idea that LLMs do not take initiative to do anything, and these are not 'beings' capable of doing things. These are tools being used by humans.

─ 17 more ─
▲▼ besterman2318:06 hn>

For the benefit of a layman, can you explain why this is so much different than a human doing it?

Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?

▲▼ sailingparrot18:15 hn>

Because the hard part is getting the simulation to be very accurate such that if you design something that works in your simulator, it will actually work in real life. And I would be extremely surprised if any LLM can actually do that correctly.

Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.

─ 4 more ─
▲▼ superxpro1219:03 hn>

It cant "invent new things". It can only reuse what it already knows about.

Thats why the math breakthrough a few weeks ago was so hotly debated. Because OpenAI is desperate to demonstrate that AI isnt just a fancy regurgitation machine, but it can actually develop novel thought. Because that would be the stock price jumps to end all stock jumps.

But then it turned out it was really just listening in on a math professors supposed-to-be-private conversations with another instance of openai, and it used his novel work as the trigger to prove the breakthrough first.

The reason people conclude AI 'thinks' is because tt can reference obscure or poorly documented things quickly (which is its primary advantage along with processing natural language prompts into tasks), which is why a lot of people with emotions confuse that action with inventing things.

─ 15 more ─
▲▼ mrandish19:21 hn>
>For the benefit of a layman, can you explain why this is so much different than a human doing it?

More broadly than the existing answers (which are correct), For a layman, I'd also add that LLMs are essentially 'brains in a vat'. They can't confirm ground truth about physical reality. They only know what's in their training data and prompt, which is incomplete and can be incorrect. Even with real-time external sensors they are limited to the sensor's margin of error, range and trusting it's working correctly.

When properly trained, fine-tuned and prompted, LLMs can be very effective in well-defined, non-physical domains like logic, writing, math and code but making things function in the real-world quickly spirals into combinatorial complexity.

─ 3 more ─
▲▼ jmoggr18:11 hn>

perhaps, but prompting is quickly becoming another area where humans are no longer clearly superior.

Getting LLMs to prompt other LLMs in a loop is not hard, it doesn't produce great results most of the time, but that is changing.

▲▼ calmworm19:10 hn>

What is “getting LLMs to prompt other LLMs” if not a prompt?

─ 7 more ─
▲▼ knicholes18:11 hn>

AI (GPT-Sol-6:xhigh) isn't even capable of using Altium to re-layout a board without introducing a bunch of problems by a user who isn't an electrical engineer or familiar with board layouts.

▲▼ scarmig18:17 hn>

AI is capable of developing its own inference hardware, once a prompt is given. The fact that a human happens to kick it off here is not especially relevant to the fact that an AI is performing the task independently. There are plenty of ways a text prompt can be generated: a harness, another LLM, or just removing stop tokens so that once begun the AI will continue until its hardware fails.

It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.

▲▼ mbgerring18:20 hn>
>It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.

It's meant to assign agency and accountability where it actually lies instead of mystifying it with anthropomorphic language.

Failing to do so has real and harmful consequences, such as enabling OpenAI to escape accountability for clearly criminal behavior.

─ 10 more ─
▲▼ throw10101019:12 hn>

How many "prompts" from our parents, teachers, bosses, etc. it took for all of us to come here and discuss this? Or for that human to think of that prompt for that LLM?

These are clearly rhetorical questions, but think of the metaphysical implication of your contestation. Ex nihilo nihil fit.

What if that initial prompt never asked for this hardware to be developed, and it was just one piece of the puzzle to answer to that prompt? That it took a chain of thousands of agents to prompt each others to come up with that?

▲▼ cyanydeez19:26 hn>

well, let me know when they have proper context management, etc.

oh. they do. I built that. It's pretty fancy.

But there's still human direction behind most projects in the cybersphere.

─ 1 more ─
▲▼ deepsun18:06 hn>

Bulldozers, excavators and rollers are now capable of building roads.

▲▼ random__duck18:10 hn>

Opened the RTL, looked at the floating point math, learned that apparently you don't need correct floating point operations for LLMs, closed the page.

▲▼ fsbonetto18:32 hn>

Indeed there are some bugs/"non standard behavior" regarding very small or very big floating point values. All of those where proven harmless for LLM inference. Thanks for pointing out the bug. If you caught something outside of that, please point that out so I can fix it ;)

▲▼ mitxela18:35 hn>

Well, you don't. That's why 1.58-bit (ternary) quants are often used. But if that's what they're going for, no need to dress it up in a floating-point facade.

▲▼ jonahss18:39 hn>

I've been vibecoding an open source hardware AV1 decoder: https://github.com/Jonahss/openav1