GLM Built Its Own Inference Infrastructure

(z.ai)

122 points | by whiteros_e 4 hours ago

14 comments

  • throwa356262 1 minute ago

        "We implemented a series of aggressive memory optimizations, including..."
    
    
    This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.
  • zicohacks 18 minutes ago
    US chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chips
  • dada216 2 hours ago
    We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
    • freakynit 1 hour ago
      Most of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.
    • axiosgunnar 2 hours ago
      [dead]
  • chung8123 11 minutes ago
    I might be missing something but when I went to their site they are more expensive than Claude. Why would I pick GLM over Claude? Is it they just offer more tokens in their plans?
    • tokai 5 minutes ago
      For one you would have to use Claude if you pick it. But seriously there is no way for you to determine if one is a better offer than the other, when the usage/tokens/credits are vague, detached, and won't tell you much without trying both.
  • Havoc 1 hour ago
    Interesting that the tone of announcements between US and Chinese providers is converging.

    GLM has in the past been more technical rather than speculation about future development on RSI etc.

    Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.

    • HarHarVeryFunny 18 minutes ago
      Ziphu (who make GLM) use Huawei Ascend processors made by SMIC. Huawei use a combination of domestic memory from CXMT and leftover (pre-sanctions) memory from Samsung.

      Just like the rest of the world, including the US (Intel, Micron), SMIC are currently using ASML lithography equipment (DUV, not EUV), but Shanghai Aishengna are now moving into early production with their own DUV machines, with SMIC and CXMT as early customers.

      There is also a state sponsored Chinese EUV development underway.

    • dude250711 1 hour ago
      Any details on the latest approach to distillation would also be very interesting.
  • Argonautlabs 1 hour ago
    Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory.

    One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.

    Method and numbers: https://github.com/argonautlabsai/argodrive (built on antirez/ds4).

  • _aavaa_ 18 minutes ago
    [dead]
  • OhNoNotAgain_99 2 hours ago
    [dead]
  • tefkah 2 hours ago
    [flagged]
    • binsquare 1 hour ago
      Given the rate of improvement, why is this deranged?
      • sixeyes 1 hour ago
        because the rate of improvement is fairly stalled?
        • xyzsparetimexyz 1 hour ago
          Do you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?
          • koe123 1 hour ago
            You prove your own point no? You are asking for a benchmark to prove AGAINST ASI. Surely the burden of proof for such a scientific fiction concept should be the other way around.
            • phoghed 1 hour ago
              No, they asked for a reliable measurement to prove that model development has stalled. Nobody is talking about ASI except you.
              • jdiff 1 hour ago
                Nah, they're fine making that small leap. RSI is the new marketing term for the sci fi singularity.
                • phoghed 1 hour ago
                  >> because the rate of improvement is fairly stalled?

                  > Do you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?

                  They clearly aren’t talking about RSI here, but that model development has stalled in general.

        • pizza234 1 hour ago
          You're tragically misinformed; it isn't. Several metrics are actually growing exponentially. But if you want emprical information, you can just have al look at the nature of the late AI incidents.

          Ironically, many benchmarks being maxxed out, and quite quickly, so new ones have to be created.

          • ttoinou 52 minutes ago
            The AI "incidents" are pure marketing ploys to get free word of mouth. Like what you're doing.
          • glimshe 38 minutes ago
            If you knew what "exponentially" means, you probably wouldn't be saying that.
          • imp0cat 56 minutes ago

                Several metrics are actually growing exponentially.
            
            Power consumption and water consuption are the obvious ones. What are the others?
          • georgefloydd 1 hour ago
            [dead]
          • simonw_simonw_ 1 hour ago
            [dead]
        • owebmaster 39 minutes ago
          Kind of but there's a lot of improvements available to the competitors catching up
        • scarmig 1 hour ago
          Yeah, it's been over a week since a Millennium Problem was solved. AI has hit a wall.
          • etamponi 1 hour ago
            It was not solved. ~OpenAI~ Buckmaster and Alpöge found one (or a few) singularities in the forced version of the Navier-Stokes equations. Then magically 2 weeks later OpenAI found them too. Again, I am not saying this is not a great feat. I am just saying that everyone should be a bit more careful when making statements about RSI.
            • red75prime 46 minutes ago
              Buckmaster and Alpöge has found a forced finite-time singularity for the 3D incompressible Euler equations (and two other types) building on the work by Diego Córdoba and Luis Martínez-Zoroa with the assistance of Anthropic and OpenAI models. Then OpenAI found a forced finite-time singularity for the Navier-Stokes equations.

              TL;DR Buckmaster and Alpöge haven't solved Navier-Stokes blow up.

              How information can get so distorted when it's trivial to fact check?

      • ramon156 1 hour ago
        if you can be replaced by an algorithm, how useful were you really?
        • azan_ 1 hour ago
          Very useful. That’s a weird question.
          • fc417fc802 58 minutes ago
            I aspire to be at least as useful as bogosort.
        • simonw_simonw_ 1 hour ago
          [dead]
      • meyer3423 1 hour ago
        [dead]
      • simonw_simonw_ 1 hour ago
        [dead]
    • pizza234 1 hour ago
      > This is known as Recursive Self-Improvement, or RSI.

      Some call this "The singularity" (e.g. Hinton).

      This is actually a core danger postulated by the, let's call it, "worrying" scenario - see AI 2027 (to be clear, I think its timeline is not realistic).

      > Statements dreamed up by the utterly deranged.

      Evidently, and tragically, it will take catastrophes to show that deranged are the ones deriding the worried crowd.

  • almaight 1 hour ago
    [flagged]
  • embedding-shape 2 hours ago
    I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

    They must have hit really hard scaling limits if the prices were hiked so much so quickly.

    • _aavaa_ 2 minutes ago
      Their plans are still worth it if you use their models. You can see how many tokens you can except to get based on plan here: https://docs.z.ai/devpack/overview#estimated-token-allowance

      The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).

      Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.

      They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.

      They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).

    • Daviey 2 hours ago
      I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
      • world2vec 1 hour ago
        1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total.

        Can I ask where are you using all those tokens?

        • _0ffh 1 hour ago
          Well, there's essentially two major ways to use these models: Pair programming or fully autonomous fire-and-forget code generation. The second strategy needs essentially zero input, so the number of tokens you can blow is practically only limited by API speed.
          • rubslopes 2 minutes ago
            There's also a third way that can spend the most tokens: if the AI is used as part of the product, and not just a tool to build the product.
        • wartywhoa23 1 hour ago
          Something like this I guess: https://youtu.be/U-Rqv9dOB1U
          • p2detar 38 minutes ago
            This is such a good video. Instant sub. Next to tech bros, we should also put AI-cringe bros.
        • Daviey 13 minutes ago
          I have 3-5 agent harnesses with large context windows working on different applications concurrently.
        • buckle8017 1 hour ago
          That's easy to do with many agents independently told to find bugs in a large codebase.
        • tokai 1 hour ago
          300M for two weeks is surprisingly low. What are you doing that need so few tokens?
          • world2vec 1 hour ago
            It's not my main model (that would be Fable 5.1 Extra) but it's been doing agent-driven search and optimisation of a cross-trading ranking model (it's for work).
            • disiplus 1 hour ago
              I would suggest you to hook fable or 5.6 to check it regularly and its work because it gets lost easily on stuff it was not trained on. I'm doing some custom inference engine optimization and it's a workhorse but it can easily lose its way and if you don't recheck it you will get wrong answers in the end.
      • disiplus 1 hour ago
        I also have a legacy pro plan and the only limitation is if you are trying to work in the morning from Europe because you are in the 3x usage overlapping China time but after 12 or so you basically can run it at least for me at least 3 parallel sessions all the time.
    • Havoc 1 hour ago
      >I was gonna ask how people found their coding plans

      Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.

      >They must have hit really hard scaling limits if the prices were hiked so much so quickly.

      Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.

    • asp_hornet 2 hours ago
      The way I look at it, their coding plan doesn’t retain data or use it for training making it one of the cheaper plans for me.

      https://docs.z.ai/legal-agreement/privacy-policy

      • andy_ppp 2 hours ago
        You believe any of these companies care about the law? They care about winning and building the self improving AI as quickly as possible.
        • asp_hornet 1 hour ago
          I too am sceptical but I’ll take my chances. At least it’s helping the open weights.
        • criley2 1 hour ago
          I believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't.

          Also why Meta gets a +1, just charge less money on the training path.

          • orf 1 hour ago
            I’m not sure that follows. You’re assuming that all those claims have the same weight, without considering the size, jurisdiction, reputation or even the general vibe of the company making that claim.

            If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.

            These are not equal.

            • asp_hornet 53 minutes ago
              > I’m not sure that follows

              To be fair, none of us are sure of anything and I think that’s the part that’s most irritating

              • orf 36 minutes ago
                It’s more a polite way of saying “that’s crap”
    • probst 49 minutes ago
      Way to restrictive in terms of tokens provided. I am on their largest plan, and quickly run into their limits. And that is using it selectively in addition to codex.
    • broodbucket 2 hours ago
      Yeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground
      • lompad 2 hours ago
        It just gives a taste of what we are all going to have to pay soon, once the model providers actually have to make money. And the era of "let's charge a dollar for every 10 dollars running the infra actually costs" is rapidly coming to an end.

        And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.

        • chobbledotcom 1 hour ago
          This isn't true, you can pay for GLM 5.3 from a provider like Neuralwatt or Friendli who have no incentive to subsidize or loss-lead their inference APIs
          • breakingcups 41 minutes ago
            They didn't pay for training
          • jdiff 1 hour ago
            This introduces other incentives to cut corners and over-quantize.
      • pyrophane 2 hours ago
        What provider are you using currently?
    • bbor 2 hours ago
      It's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered.

      For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.

  • jonstewart 57 minutes ago
    Necessity is the mother of invention. The shortsighted protections put on chips, etc., by the US has forced Chinese AI industry to adapt or die. Guess what their response to this fitness function has been? Kudos to Z.ai on their inventions and excellent write-up, which reads like humans wrote it.
    • HarHarVeryFunny 7 minutes ago
      Wouldn't it be refreshing if OpenAI and Anthropic were this open, and spelled out how they were using their own models during development and rollout?!

      All I can recall reading from OpenAI about what they have actually done in the name of "RSI" is using one of their models to help automate the training process.

  • bbor 2 hours ago
    Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
    • pjc50 27 minutes ago
      Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/

      No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.

      • _aavaa_ 1 minute ago
        I'm not defending their actions, but we should be clear about where the law currently stands: Anthropic was found to infringe because of the torrenting, not because of the training.
    • woadwarrior01 1 hour ago
      That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.

      https://x.com/EricSimons/status/2099252922098061714

    • phoghed 1 hour ago
      We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.
    • jensb1 2 hours ago
      What is "illegal" about it?
      • bingud 2 hours ago
        breaking Anthropic TOS and misleading users
        • drbscl 1 hour ago
          Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.
      • bbor 1 hour ago
        Are you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA.

        In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.

        In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.

        I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(

        TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.

        [1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.

        • tuesdaynight 6 minutes ago
          You didn't explain why it's illegal or why distillation is bad.
        • pjc50 29 minutes ago
          > alignment crisis

          Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".

          If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.

        • dgellow 59 minutes ago
          What does any of this has to do with the legality of distilling Claude?
        • podocarp 30 minutes ago
          Source for 1? Are we sure those aren't hallucinations?
        • Bluestein 1 hour ago
          Nulla poena sine lege?
        • jLaForest 55 minutes ago
          Yes, wont somebody please think of the shareholders whose IP had been stolen...
        • jensb1 1 hour ago
          [dead]
    • butterNaN 56 minutes ago
      Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.
    • Laurel1234 1 hour ago
      [dead]
  • rob74 1 hour ago
    This article left me with one immediate question: "WTF is GLM?".

    Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...

    • jbonatakis 1 hour ago
      z.ai is a fairly well known AI lab out of China and their GLM models are probably the most popular outside of Anthropic or OpenAI’s. I don’t think it’s presumptuous for them to not introduce themselves in a post on their own blog, I think you’re just a bit out of the loop here.
    • HarHarVeryFunny 14 minutes ago
      Ziphu, aka Z.ai, is the company that makes GLM (a very competitive Chinese LLM).

      Why would you be reading their corporate blog posts if you don't even know who they are?!

    • bogdan 1 hour ago
      I don't get the outrage. Do you post this kind of stuff on every topic on hackernews that you are not knowledgeable about?
      • rob74 49 minutes ago
        Maybe my post sounded harsher than I intended, and yeah, it's probably on me that I'm not familiar with GLM. Actually the other major Chinese LLM Kimi does ring a bell, maybe it's because three-letter acronyms are a dime a dozen and annoy me because I'm confronted with them regularly at work too (people at my company seem to love acronyms), but that's obviously on me too...
        • tokai 34 minutes ago
          It didn't read as harsh. Only unaware and you broadcasted that you don't have the decency to do basic searches.
    • peri-cl 1 hour ago
      It's only the top open-weights LLM in the world,

      https://artificialanalysis.ai/#intelligence-category-tabs

    • fxwin 1 hour ago
      It's presumptuous for them to assume that a reader of their blog is familiar with their product?

      Also I feel like the obvious way to read the very first sentence is that GLM is a language model

      > As we develop GLM, the model sometimes exhibits capabilities that surprise us

    • Mashimo 1 hour ago
      A ai model family similar to Codex, Gemini or Claude.

      Where GLM-5.3-Flash is the newest "small / fast" model.

    • drbscl 1 hour ago
      >As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us.

      Come on now

      Also, why would they introduce themselves on their own blog?