• ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
    link
    fedilink
    arrow-up
    13
    arrow-down
    1
    ·
    7 days ago

    I expect we’ll start seeing stuff like Taalas where they print the model to the chip and other specialized chips like Xuantie C950 going forward. Neither of these requires DRAM, and Taalas is particularly clever since they just print the model right to an ASIC chip. So, the whole renting out LLMs business model isn’t going to last long I suspect.

    • MalReynolds@slrpnk.net
      link
      fedilink
      English
      arrow-up
      7
      ·
      6 days ago

      So a repeat of the crypto crash for graphics cards when ASICs ate their lunch. Mind you, that’s only for inference (although a super fast QWEN 3.8 would meet a lot of peoples needs).

      The argument for datacentres is for training the models, but then they’ll need to prove that they haven’t hit a diminishing returns wall, which will be hard if, as seems likely, they have. Also the Chinese have been doing it in a cave, with a box of scraps (figuratively), and gotten at least 90+% as good results.

      Seems like the recent advances have been in the frameworks, which don’t need no stinking (literally if fossil fueled) datacentres.

      • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
        link
        fedilink
        arrow-up
        7
        ·
        6 days ago

        Right, I’d argue that China proves you don’t need massive data centers for training. And yeah, I think something like Qwen 3.8 is more than enough for tasks most people do. There are a lot of tricks you can do as well with the harness, where there’s a lot of attention is shifting now. And it’s a lot cheaper and faster to develop better harnesses than train new models. I expect we’ll start seeing a shift towards neurosymbolic systems before long where the LLM acts as a stochastic component within a symbolic logic engine.