yeah, an SDK for all those times you want to hand-implement sparse attention and MoE kernels on an architecture where the only known code is from the vendor. a truly "educational experience" to be sure.
no one knows what Jaguar Shores will be like. it will probably use some Gaudi IP, but my...
As far as I know, no up-to-date software framework supports Ponte Vecchio or Gaudi 2 (Intel has essentially abandoned both after financial pressure last year).
Adding one more to my zoo:
16x V100 32GB SXM2 (two nodes), 2026: Volta has aged well. This eight-year-old monstrosity runs GLM-5.1 out to the full 202K token context (and should run GLM-5.2 out to 250K+ as well, but I haven't downloaded the weights yet). With many custom kernels (I had to...
Be careful here guys. OP only has posts in FS/FT, got really defensive in the other thread where they were selling $70K of RAM, local or wire only, and seems to post frequent drops of high-value hardware with timestamps that don't match their username.
@OP - if you're selling $20K of RAM, might...
I think if your only goal was to serve GLM5.1 or Kimi K2.6, its a viable choice - Intel does maintain a working vLLM implementation which is only a couple months behind upstream and covers most of the interesting models. 768GB for $18K really is a class of its own and I think if it creeps down...
In general I think multi-GPU scaling (or the lack thereof) is worth discussing here. Modern models are effectively "deep and narrow" during inference because they are all really sparse. If you are generating at 100 tokens per second, you have 10 ms per forward pass. If your model has 48 layers...
I think at this type of spend you really need to define your goals. If your goal is to not pay for Claude for non-work projects, just use GLM 5.1 or Kimi 2.6 on OpenRouter. The cloud providers get more out of their hardware than you do, because they can batch requests to increase GPU utilization...
I wouldn't go under 128GB of VRAM if your aim is to buy something to run LLMs on. 96 is really tight - it's barely enough for Qwen-122B at full context which means once more models with 1M context or something with 150B params rolls out you won't be able to take advantage of it. Here are a few...
The MoE will certainly be a worse model, and Gemma-26B and Qwen-35B are comparable, so if you're idea of trust is "made by an American tech company" then whatever floats your boat (though I would argue that Google is hardly a beacon of trust, and DeepMind's copyright violations during training...
(1) I really don't see why people need to have AI generate them a huge story out of something that could be summarized as "Tried to run Gemma-4-31B on a 20GB GPU, got 8 tokens/second with -ngl 54"
(2) If you want to run a dense model, Qwen3.6-27B is a bit smaller and a bit better for code
(3)...
Why wouldn't you sell them on eBay if that's the case? It's safer for you and the buyer if you sell them 4 or 8 at a time. At $6000 for a set of 4, I can't imagine any day job that pays high enough to make it not worth the 20 minutes to drop 4 of them in the mail.
Stuff that I've tried, in no particular order. The dates matter because model sizes and context lengths have changed a lot in the past three years:
A single 3060 (circa 2023): loved it. Ran 7B models in q4 at 16K context just fine, super handy for sundry text manipulation work (summarization...
Normally I'd agree, but DDR5 RDIMMs are currently something like $30/GB (there's a huge shortage of RDIMMs right now), so the RAM in that workstation is worth about $25K.
I do agree this is top, top dollar though, really taking advantage of current memory prices. I'd be surprised if someone bites.
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.