13 comments

  • rcarmo 3 minutes ago
    Well, as long as it doesn't start developing anatomically accurate metal skeletons with red glowing eyes...
  • pcarolan 54 minutes ago
    Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
    • zdragnar 43 minutes ago
      Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.

      It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.

      I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

      • HoldOnAMinute 2 minutes ago
        At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.

        All aboard! We're racing to the bottom now.

      • fhdkweig 32 minutes ago
        I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?
        • jerf 19 minutes ago
          FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.
          • rjh29 9 minutes ago
            I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry?

            Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)

            • exmadscientist 1 minute ago
              You don't really get an FPGA for capability. CPUs are much more capable, and they're general-purpose so they can do absolutely anything with about the same efficiency and just a little more code.

              You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuffed in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.

              (Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)

        • zdragnar 24 minutes ago
          It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?

          Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.

          If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?

          There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.

          My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

        • monocasa 25 minutes ago
          They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.
        • fsbonetto 31 minutes ago
          They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC
        • LoganDark 26 minutes ago
          1. No

          2. They don't have enough capacity either

          The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).

          You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

    • skeskinen 51 minutes ago
      Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.

      Also, it's hard to get fab capacity for any project. Let alone something so experimental.

      • jcims 48 minutes ago
        Addressing these issues seems to a major driver behind the design of terrafab.
    • buriram 10 minutes ago
      Yes, and startups do exactly that. Check out Etched https://www.etched.com/ where they made a Transformer specific GPU (basically a form of ASIC) where they bet that transformers would be the dominant GPU architecture for running AI / LLM workload.
    • ohazi 49 minutes ago
      • yorwba 19 minutes ago
        8 months ago, Taalas claimed https://taalas.com/the-path-to-ubiquitous-ai/#:~:text=Upcomi... that "Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM. It is expected in our labs this spring and will be integrated into our inference service shortly thereafter. Following this, a frontier LLM will be fabricated using our second-generation silicon platform (HC2). HC2 offers considerably higher density and even faster execution. Deployment is planned for winter."

        Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.

      • slowin 22 minutes ago
        I think this company was recently acquired by AMD, so hopefully they'll start getting some this into production. I know OpenAI was working on model-on-a-chip too.
    • samuelknight 20 minutes ago
      Models fully deprecate in a few months. Why would you burn an algorithm that fully depreciates in value faster than a bag of potato chips. The 'inefficient' general purpose hardware is constantly renewed with every released model. Even 6 year old Ampere GPUs are still usable.
    • __MatrixMan__ 32 minutes ago
      Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?
    • birdatlaw 44 minutes ago
      From what I've read, not only are some labs doing it (other commenters already mentioned).

      But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.

      I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.

      I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.

      I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.

    • AIblemblio 18 minutes ago
      We are still in the middle of the AI race. Commodity hardware is easy to use, can do everything and is fast enough.

      Your optimized hardware chip might be obsolete before its back from the fab.

      SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.

      Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.

    • zitterbewegung 49 minutes ago
    • fsbonetto 42 minutes ago
      The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.

      So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

      • fnordpiglet 31 minutes ago
        Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

        The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.

        I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.

        Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.

        But such is the cycle

        • cestith 20 minutes ago
          You're starting to hint at compute-in-memory as a general replacement for CPU/DIMM layouts. That could be useful for far more than LLMs, world models, or any sort of AI. It takes a bit of a different software development stack than a standard architecture though.
    • schleck8 42 minutes ago
      Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.
    • pmarreck 37 minutes ago
      Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?
      • fsbonetto 35 minutes ago
        GPUs are faster, but you can't make your own arch on GPUs. FPGAs offer you that possibility. Said that... There are a few beasty FPGAs used in crypto mining coming my way... I expect that OpenTPU will be able to run frontier models with those.
    • traverseda 51 minutes ago
      I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.

      Also can't keep them closed source if you do that.

    • jolt42 25 minutes ago
      Even dumber question: What is new or novel about this openTPU?
      • fsbonetto 19 minutes ago
        First opensource arch that can do modern LLMs, while maximizing the potential of its hardware; First opensource TPU build by a recursive improvement loop...

        It's upcoming second generation could run the inference of the models that are being used to improve it...

    • hehimself 52 minutes ago
      They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.
    • dmitrygr 35 minutes ago
      In addition to some of the other replies you got, here is one more:

      Much of a model are weights, and high-density ROMs are very very very hard.

    • meowers1 15 minutes ago
      [flagged]
  • athrowaway3z 40 minutes ago
    I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December.

    The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.

    But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.

    • chris_money202 11 minutes ago
      There doesn't exist a single FPGA that can fit an entire AI ASIC. You would need dozens stitched together, then comes the issue of clock speeds, FPGAs typically run far below reference. There also memory issues with FPGAs.

      Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.

    • felixgallo 29 minutes ago
      I suspect an AI could design a purpose-built FPGA-like replacement that would be, for its purpose, significantly more effective than the current general-purpose FPGAs.
      • fsbonetto 28 minutes ago
        It could have a small improvement on power consumption, but the current design can already achieve 90% of the maximum theoretical speed of this hardware without giving up programability/flexibility
  • fsbonetto 1 hour ago
    After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.
  • xg15 42 minutes ago
    "Recursive self-improvement will kill us all!"

    Also: Here is our recursive self-improvement hard at work...

    • lelanthran 33 minutes ago
      > "Recursive self-improvement will kill us all!"

      > Also: Here is our recursive self-improvement hard at work...

      Soon we will see

      token-providers: "The torment nexus is a cautionary tale"

      Also token-providers: "Finally, we have created the torment nexus that we first told you about!"

    • dumberquestions 29 minutes ago
      Technology has always contributed to improving next iterations of itself, it's only a concern when it's fully autonomous.
    • nialse 30 minutes ago
      All will end up on same plateau eventually. RSI is just a phase on the way there.
      • mrob 18 minutes ago
        The problem is that plateau is likely far beyond human capabilities. I don't care if ASI progress stalls after it's already killed all biological life as a useless waste of resources.
        • Jtsummers 12 minutes ago
          > I don't care if ASI progress stalls after it's already killed all biological life

          What's your basis for thinking ASI will kill all biological life, and how do you think it's going to happen?

  • vatsachak 1 hour ago
    I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
  • gfalcao 3 minutes ago
    [delayed]
  • skybrian 57 minutes ago
    This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?
    • fsbonetto 47 minutes ago
      Its a datacenter decommissioned board, really popular among hobbyists.

      For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.

  • bitwize 36 minutes ago
    Colossus is building Colossus II.
    • rcarmo 2 minutes ago
      Feelis like working at Magrathea...
  • AnimalMuppet 28 minutes ago
    Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
    • sehw 18 minutes ago
      [dead]
  • srameshc 30 minutes ago
    This post brings me to question "What does it mean to be a software developer in future" ?
    • amelius 13 minutes ago
      Basically, an unemployed plumber.
  • fabiofachini92 1 hour ago
    [flagged]
    • pjmlp 1 hour ago
      I have seen this somewhere....
      • intrasight 56 minutes ago
        It was widely noted (at least 15 years ago, maybe more) that every generation of CPU is somewhat dependent upon the computational capabilities of the previous generation being used in its design.
      • cestith 16 minutes ago
        Until we get to Deep Thought creating the Earth, I think we're okay.
      • fsbonetto 1 hour ago
        Besides an specific movie ? There is this CPU auto improving loop as well: https://github.com/FeSens/auto-arch-tournament Opus 5.5 was the first to beat the human baseline
        • pjmlp 56 minutes ago
          Of course the joke was about a specific movie.
  • rfgplk 49 minutes ago
    Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.
    • jetemple 39 minutes ago
      Which tools have made that leap? Faster design iteration makes sense, but what points to exponential hardware gains rather than shorter development cycles?