Slowdown in the homelab segment

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

BlueFox

Legendary Member Spam Hunter Extraordinaire
Oct 26, 2015
2,549
2,007
113
I don't want to invest in old tech though.

I rather see a high core/freq optimized CPU that comes with CPU/iGPU/NPU and power optimizations like it was a mobile cpu.
It can be packaged as a desktop cpu in size so more bits could be added for iGPU processing, etc.
Hopefully it would have quad memory channels and at least 16 pcie lanes.
Some of the AMD chips are kind of getting there.
Old? Arrow Lake is current and was only released last year. Already met all your requirements and now you're moving the goalposts. It has the same GPU and NPU as the mobile side too.

Aside from that, quad channel RAM already puts it in HEDT territory, which is the opposite of low TDP and quite unrealistic. Not sure what AMD CPUs you're referring to as being any better in that regard either.
 

marcoi

Well-Known Member
Apr 6, 2013
1,701
416
83
Gotha Florida
Old? Arrow Lake is current and was only released last year. Already met all your requirements and now you're moving the goalposts. It has the same GPU and NPU as the mobile side too.

Aside from that, quad channel RAM already puts it in HEDT territory, which is the opposite of low TDP and quite unrealistic. Not sure what AMD CPUs you're referring to as being any better in that regard either.
Originally you mentioned i7/i9, not core ultra. The I7/i9 14000 raptor lake came out like 23/24 and had the issue with instability and degradation. Not sure if the T-series was impacted. Also the e-core/p-core causes issues for hypervisors.

The ultra series still make use of the p/e cores which again can cause issues for hypervisors.

In my post i was saying i would like to see something like that, not that a product like that already exists.
It's more of a wish list for new tech.

For AMD chips like AMD Ryzen™ AI Max+ 395 would be a nice starting place for home lab.
 
  • Like
Reactions: abq and Greg_E

Greg_E

Active Member
Oct 10, 2024
516
165
43
As far as I know, all i5 or higher 13th and 14th gen Intel had problems where they slowly destroy themselves. I might be wrong on i5 13th gen.

Guess what I got for a classroom of video editing computers? i7 14th gen at a huge discount just before the news was released that all of them had problems, also before HP patched the bios to help.

Back to the lab talk, I think I need some pcie 3.0x8 sfp28 at 25g cards for my Harvester lab :( I was watching some cheap ones, but they were sold out last night. It's like an addict chasing their next high o_O
 

marcoi

Well-Known Member
Apr 6, 2013
1,701
416
83
Gotha Florida
I got most of my networking up to 10g or 2.5g. Some of those higher networking devices use a lot of power. I know i spent money getting the more efficient intel dual rj45 10gb based cards vs older gen cards.
 

Greg_E

Active Member
Oct 10, 2024
516
165
43
Like a dummy I put in an offer for three 25g cards. They are faster and at bus speed. What I have in there now is PCIe 2.0x8 and while they can hit 10g speeds, I'm wondering what else they might be slowing down. I was starting to go toward quad x710 cards, but then those got stupid expensive again. I have two of these, figured they would be great for a more advanced firewall when I got to that point. 10g across the local networks would be nice.

Have some sfp28 dac and multimode modules in my watchlist, probably go with the dac because the result is cheaper. The Supermicro cards are xxv710, so 1, 10, 25.

As I was writing this, seller accepted, so guess I need to buy some cables.

All save one of my mini-lab have an a+e i226 card in them, mostly just used for gigabit because I have no real 2.5g switches. But an extra Intel port can't hurt, came in handy when I had ESXi running and the built in Realtek was worthless.
 
  • Like
Reactions: nexox

T_Minus

Build. Break. Fix. Repeat
Feb 15, 2015
7,893
2,225
113
I'm running 13500T in one of my miniPCs hopefully it lasts :)

I did notice it was pretty warm (The entire case) at one point of testing, hopefully doesn't burn itself out...

Now I'm wondering if I need to redo the HS paste
 

Greg_E

Active Member
Oct 10, 2024
516
165
43
Don't overclock it! Make sure your main board doesn't overclock with the defaults, and make sure the bios is up to date.
 

WANg

Well-Known Member
Jun 10, 2018
1,519
1,159
113
48
New York, NY
As far as I know, all i5 or higher 13th and 14th gen Intel had problems where they slowly destroy themselves. I might be wrong on i5 13th gen.

Guess what I got for a classroom of video editing computers? i7 14th gen at a huge discount just before the news was released that all of them had problems, also before HP patched the bios to help.

Back to the lab talk, I think I need some pcie 3.0x8 sfp28 at 25g cards for my Harvester lab :( I was watching some cheap ones, but they were sold out last night. It's like an addict chasing their next high o_O
Eh, no - the Raptor Lake ring-bus over-goosing only impacts the H series laptop chips and the heavier desktop chips. For the U and P series chips (which are used in some TMM machines) they are not impacted.

HOWEVER I would still not use them as they are not that competitive in terms of efficiency or thermals, especially when you match it up against Zen3 based Rembrandt or Phoenix/Hawk Point APU based machines. Raptor/13th Gen and 14th Gen are not that much faster than the 12th Gen Alder Lakes.

My only beef with buying those AMD APUs is that they almost always use DDR5 SODIMMs, which went up, like, what, 6x in pricing within the past year and 48GB consumer SODIMMs are practically unobtainium?

As for QSFP28 boards, did we run out of old-but-reliable Mellanox ConnectX3-VPI (MCX354A) 40Gbit cards? They were around 30-40 bucks each with volume discounts up until maybe Q2 2025, but I haven’t really shopped for anything cutting edge for a while.
 
Last edited:

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,062
113
My son (not yet two) decided my keyboard was fun as I was typing this earlier.

Overall, I think the traditional homelab space has been driven by five main objectives:
  1. Storage
  2. Learning Networking
  3. Virtual machines/ Containers
  4. Hobby/ fun
  5. Home Automation
Storage is expensive. There is no real RJ45 path past 10Gbase-T, so if you want modern networking, you are talking higher-end gear. VMs and Container orchestration are largely solved, and will be managed by AI agents. Home automation is still very relevant.

Now you have local AI/ local AI agents, which are taking mindshare, but have a much higher cost of entry and are also a lot of the hobby/ fun space.

I think there are a couple of big things happening:
  1. Hardware started getting a lot more expensive and power-hungry. Early Xeon E5 V4 era 145W was a lot for a CPU. Now it is 400-500W. Heck, we have 2026 mini PCs that consume more power than the servers we reviewed in 2016. Also, pre-Cascade Lake CPUs, many folks will just not touch them due to Spectre/Meltdown.
  2. Power pricing has gone up, so the price of running bigger hardware has increased
  3. DRAM pricing has gone up. RAM is probably the easiest thing to recycle in systems.
  4. SSD/ HDD pricing is increasing. Storage has been the driver of home labs for 15+ years (before people used the term, even)
  5. AI is destroying some of the value equation. Why do I need to learn to set up Cisco switches or VMware when LLMs can do it for me? This is a bigger question because I spend a lot of time on vendor calls listening to how they are helping with setups and monitoring, and then thinking that a lot of this is now just applying a decent-sized model and an AI agent to go do. Of course, it never goes away, especially not overnight, but I think folks are figuring out that the next career move is going to value applying AI versus setting up a switch
  6. I also think that many people are better off with a $20/mo Anthropic account than with a low-cost Proxmox VE/ Ubuntu VM host for learning the most relevant skill.
  7. Next-generations of hardware will be more difficult to use
    1. Higher power required. How many people want 1kW+ CPU-only systems around for fun? It is not zero, but it is a different number than when a 1P ATX system was like 250W max
    2. Custom motherboards have become the norm for volume shipments.
    3. Liquid-cooled servers then mean you need a CDU
    4. ORv3 racks very few folks will have at home
    5. Since 2020, many of the Dell/ Lenovo CPUs have been vendor-locked with AMD PSB. AMD's market share is going over 50% this year, so PSB will hit an increasing portion of CPUs
    6. New AI models and techniques often make old hardware obsolete much faster than in previous cycles
    7. NVIDIA Grace shipments have been growing, but with soldered memory and such, they are not going to be as easy to re-purpose either
Not sure if that helps, but just a bit of my perspective.
 

WANg

Well-Known Member
Jun 10, 2018
1,519
1,159
113
48
New York, NY
Originally you mentioned i7/i9, not core ultra. The I7/i9 14000 raptor lake came out like 23/24 and had the issue with instability and degradation. Not sure if the T-series was impacted. Also the e-core/p-core causes issues for hypervisors.

The ultra series still make use of the p/e cores which again can cause issues for hypervisors.

In my post i was saying i would like to see something like that, not that a product like that already exists.
It's more of a wish list for new tech.

For AMD chips like AMD Ryzen™ AI Max+ 395 would be a nice starting place for home lab.
Well, the E/P core nonsense is really a KVM/linux kernel scheduling issue, and you could do core pinning to address them, like for example pinning your HA or storage server instance on E-Cores and the bursty stuff (like Jellyfin) on P-cores, but that's a freaking waste of power and more complications. The Core Ultra 1s and some Ultra 2s have an additional core type (LPE) that already earned itself an unflattering nickname (literally pointless engineering) which is not even used in the Windows scheduler. I personally favor the AMD Zen4/4c strategy on hypervisors myself, simply because it's not a massive gulf of performance difference between 2 (or 3) fundementally different core types.

That being said, it's really hard to be super-enthusiastic about the current homelab environment nowadays. Windows 11/12 isn't exactly pushing PC or corporate users to upgrade devices and this entire AI rigamarole is not exactly driving up end user device demand since as if you don't do LLMs locally, whatever you have is already plenty good for your existing needs, and older TMM secondary market pricing didn’t go down by that much. As for Strix Halo, unless your usecase is large-ish LLMs and you like the idea of paying extra for more soldered RAM, there’s really no need to buy one. From what I remember the 8060S in the Strix Halo in llamacpp for ROCm is roughly the same ballpark perf-wise as an Apple M4 Pro or M1 Max (roughly similar to each other), so you might be better off raiding the office retired asset pool for one.

Strix Point (Zen5) isn’t that much of a performance jump versus Phoenix/HawkPoint, and unless you got a Ryzen9 AI HX390 or Ryzen7 AI 370, the average case integrated GPU (860M) is either at par or weaker than the one from the previous gen (780M). Oh, and I think i saw numbers for 780M that’s roughly the same as an actively cooled M4 (aka the 500 dollar base M4 Mini). Do keep in mind that you can do somewhat useful things with small LLMs…like, say, facial detection on Google Coral Edge TPU or the iGPU on a Core Ultra 1, so unless you have something very specific in mind…why would you spend 2k on a Strix Halo box?
 
Last edited:
  • Like
Reactions: marcoi and Patrick

CyklonDX

Well-Known Member
Nov 8, 2022
1,819
660
113
Just few cents
NVIDIA Grace shipments have been growing, but with soldered memory and such, they are not going to be as easy to re-purpose either
Almost no one wants to buy those systems, funny reason i've heard "they didn't come with intel cpu, and its arm"

New AI models and techniques often make old hardware obsolete much faster than in previous cycles
Are they? You can still slap 5090 on broadwell, and if model fits whole on the gpu - it will perform about the same. If we add llama.cpp into the mix we can split load between different gpu's without expensive nvlink, sxm boxes
(the split has full freedom, but obviously the matrix will run at speed of the slowest card.); There's no real need to run 100G models, they are often as smart as your 8-20G models. The problem is that the ai didn't make a lot of nv cards obsolete quickly... rising prices on old quadro/geforce 16-24-48G cards... this includes even some amd cards.

DRAM pricing has gone up. RAM is probably the easiest thing to recycle in systems.
Doesn't this happen like every generation, the manufacturers find a way to raise them; Sure this time it rose a lot more - as scalping target is datacenter; but its nothing new.

Hardware started getting a lot more expensive and power-hungry. Early Xeon E5 V4 era 145W was a lot for a CPU. Now it is 400-500W. Heck, we have 2026 mini PCs that consume more power than the servers we reviewed in 2016. Also, pre-Cascade Lake CPUs, many folks will just not touch them due to Spectre/Meltdown.
in form of lot more expensive, the question is EoL and big companies flooding 2nd hand markets with lot of hardware. This too will happen again, all those datacenters will cycle hardware every 3-4 years; as it will be more profitable for them to run latest (just from power perspective). Once that happens we should expect enormous flood gates of cheap server grade hardware - if they really bought that much, and its not just scalping feast.
 
  • Like
Reactions: nexox

marcoi

Well-Known Member
Apr 6, 2013
1,701
416
83
Gotha Florida
in form of lot more expensive, the question is EoL and big companies flooding 2nd hand markets with lot of hardware. This too will happen again, all those datacenters will cycle hardware every 3-4 years; as it will be more profitable for them to run latest (just from power perspective). Once that happens we should expect enormous flood gates of cheap server grade hardware - if they really bought that much, and its not just scalping feast.
Even if that is the case, who will go out and buy a CPU that is 500watts to run in their own home lab?

If i had the capital and connections I would just start a company that builds systems for just for home lab and home automation use cases.
 

Greg_E

Active Member
Oct 10, 2024
516
165
43
The problem with the retired datacenter stuff from AI is the size, I know you've seen some of the STH video tours, most of that stuff is not lab friendly. Compute nodes from hypervisors would often be closer to the market.
 

WANg

Well-Known Member
Jun 10, 2018
1,519
1,159
113
48
New York, NY
My son (not yet two) decided my keyboard was fun as I was typing this earlier.

Overall, I think the traditional homelab space has been driven by five main objectives:
  1. Storage
  2. Learning Networking
  3. Virtual machines/ Containers
  4. Hobby/ fun
  5. Home Automation
Storage is expensive. There is no real RJ45 path past 10Gbase-T, so if you want modern networking, you are talking higher-end gear. VMs and Container orchestration are largely solved, and will be managed by AI agents. Home automation is still very relevant.

Now you have local AI/ local AI agents, which are taking mindshare, but have a much higher cost of entry and are also a lot of the hobby/ fun space.

I think there are a couple of big things happening:
  1. Hardware started getting a lot more expensive and power-hungry. Early Xeon E5 V4 era 145W was a lot for a CPU. Now it is 400-500W. Heck, we have 2026 mini PCs that consume more power than the servers we reviewed in 2016. Also, pre-Cascade Lake CPUs, many folks will just not touch them due to Spectre/Meltdown.
  2. Power pricing has gone up, so the price of running bigger hardware has increased
  3. DRAM pricing has gone up. RAM is probably the easiest thing to recycle in systems.
  4. SSD/ HDD pricing is increasing. Storage has been the driver of home labs for 15+ years (before people used the term, even)
  5. AI is destroying some of the value equation. Why do I need to learn to set up Cisco switches or VMware when LLMs can do it for me? This is a bigger question because I spend a lot of time on vendor calls listening to how they are helping with setups and monitoring, and then thinking that a lot of this is now just applying a decent-sized model and an AI agent to go do. Of course, it never goes away, especially not overnight, but I think folks are figuring out that the next career move is going to value applying AI versus setting up a switch
  6. I also think that many people are better off with a $20/mo Anthropic account than with a low-cost Proxmox VE/ Ubuntu VM host for learning the most relevant skill.
  7. Next-generations of hardware will be more difficult to use
    1. Higher power required. How many people want 1kW+ CPU-only systems around for fun? It is not zero, but it is a different number than when a 1P ATX system was like 250W max
    2. Custom motherboards have become the norm for volume shipments.
    3. Liquid-cooled servers then mean you need a CDU
    4. ORv3 racks very few folks will have at home
    5. Since 2020, many of the Dell/ Lenovo CPUs have been vendor-locked with AMD PSB. AMD's market share is going over 50% this year, so PSB will hit an increasing portion of CPUs
    6. New AI models and techniques often make old hardware obsolete much faster than in previous cycles
    7. NVIDIA Grace shipments have been growing, but with soldered memory and such, they are not going to be as easy to re-purpose either
Not sure if that helps, but just a bit of my perspective.
Well, I would weigh the 5 objectives a little more differently - I think most people run home labs for fun, and learning being secondary - when $dayjob has all sorts of proverbial tech junk sitting around and with power being "somewhat" cheap stateside (depends on $loc), any enterprising techie-of-culture would take stuff home provided that it's not expensive to own or run, and $sigoth (assert $sigoth != NULL) tolerates it ($patience converging to zero as tech junk accumulates either on take-home, or fleabay/FB marketplace/Craigslist/whatever). The issue of course is that as with everything within any culture, someone always takes it too far and ruins the fun for all.

The limiting factor for me has always been limited time to dedicate on getting the hardware going, then the limited space to have anything at home, that 30 Amp circuit for the main living space in the homestead and the real question of what I need to be doing with said hardware to make it worth my while. I might have old laptops, mini-boxes thin clients perfect for retro gaming, or a PowerPC G4 MacMini for running MorphOS (for AmigaOS fun)...but I have much less time to play with them than ever before.

Storage...is a weird subject, since it's essentially about data retention on mediums that has questionable longevity - one of the headaches that I have is that between me and the missus, our devices generate tens of GB of data each year, and one of the biggest lies of modern IT is "is it backed up into the onsite storage, placed securely into the cloud and easily recoverable?" Oh sure. My laptops are on TimeMachine, timemachine is on TrueNAS, TrueNAS is on S3...and the same goes for the smartphones whenever we plug them into a PC for backing up...but how frequently this needs to be done and how sad we'll be if we lose photos from an overseas trip or from a late aunt or whatever...is seriously up for debate. Of course, this is one of those marginally "actually useful" user cases with "AI" - large scale data classification and inference building. Feed the photos into Apple Photos on an M-series Mini, let their neural engine crunch through the EXIF data and then tell you "hey, back in mid-2016 you and the missus was in Montreal and spent 5 days - you also got photobombed by the owner's grandson while eating smoked meat sandwiches in Schwartz's deli on The Main...and you guys had some dupes of some photos...may I dedupe them?"

Home automation is basically tossing up a Zigbee or ZWave hub and running HA somewhere, getting everything to talk to each other, and then presenting an API to something that lets you see or control whatever IoT assets that you might havd. Technically you can run HA on a Raspberry Pi 3 if all you need is environmental control or condition checking, but the magic isn't simply to make it work - it's making it work reliably for as much as possible (UPSes, monthly battery checks, decent networking, etc). When the missus picks up her phone and open up HA, she's reading the room temperature/humidity levels, looking at the webcam and checking on the grow lamps in the sun room and the watering devices on the flower pots. It better work at 3a even if she's in a hotel in Osaka on a business trip and you are away from home on the other side of the North American continent..or you are gonna catch hell if her plants die. Nothing is more of a pain in the ass than playing IT for family - at least most employers have reasonable expectations of SLA, while family members thinks that the best time for you to "fix" their 10 year old Android tablet is NOW and your time spent has no value.

A major issue (at least for me) regarding agentic AI is data quality. Anyone who uses Microsoft copilot to find questions on Office 365 or Azure/Entra on MSDN or techhub will tell you that even Microsoft can't keep their own agents from spitting out bad information at you, and you can cringe at a typical corporate sharepoint/wikis/repo deployment crammed full of dead projects, changed initiatives, misguided security classifications and etc, and think about what the AI will come up with inhaling that stuff - garbage in, garbage out. A classic example that I give to interns of what agentic AI can't help you reason out would be something like asking how to resolve Microsoft Purview breaking when going from classical to new mode and having to use Azure cloudshell (cloudhell) to add a service provider (SP - which doesn't exist) on the primary tenant and then wire it with the correct graph API perms. Microsoft's own copilot agent will read a few of their own support links and assume that the response provided was correct...when actually it wasn't (the actual response had more to do with confirming the bill-ables on the primary tenant account, which wasn't a technical issue per--se, so the SP can be added to the list of enterprise apps).

The problem really isn't agentic as much as Microsoft's support forum guys not knowing how to edit stale responses and set expectations properly - some companies provide good guidance, most doesn't, and if the correct answer is paywalled behind "make a support case and we'll let someone try to figure it out", well, that's not great but at least it doesn't send someone down a blind valley. Unfortunately knowing how to use Anthropic or Gemini or ChatGPT is only half the battle, knowing that the answers given might be counter-intuitive, or nuanced, or requires an entire set of pre-conditions...is the other half. So yes, get that Anthropic subscription but never stop learning and working to sharpen up the resolution on that BS detector. Being able to learn and develop intuition based on imperfect data is an important skill that agentic AI can't really teach you. AI magnifies what it knows and hides what it doesn’t know, so you’ll still need to be able to check blind spots.

I would argue that there are much diminished opportunity nowadays to buy modern-ish hardware that's worth the power draw, thermals and/or real estate, or is actually fun and kinda useful for your needs (well, with caveats). I mean, the last truly fun and great bang-for-buck enterprise grade machine for your average data center nerd scalable to your home workload is the HP Microserver G7/G8/G10 (up to revision 2) before HPe turned it into this overpriced and kinda weak box that pleased no one (the G10 should probably have been like the EC200a with the quad bay disk shelf, but whether their support contract is worth it...is seriously up for debate). Some QNap boxes comes close, as does the Terramaster F6-424, the UGreen DXP6800, but...eh, buying those boxes is like buying a Minisforum or GMKTec or Geekom or whatever Chinese hardware maker is out there...you are not entirely sure what corners they cut to rush something to market and whether the money you save picking them is worth it. That's why I love TMMs and their red haired stepsisters the HP/Igel/Wyse thin clients or Seneca smart displays - they are sold by the thousands if not millions, designed to survive neglect and abuse, they are not that expensive to swap out and replace when they die, they can reuse standard parts for the most part, and getting a support contract for one with next business day parts or techie visit isn't too pricy. Sometimes if your home environment is down for 5 weeks because it needs parts slowboated from the PRC…well…that would suck, and yes, I know that those TMMs are probably made on the same Quanta/Wistron/Inventec factory floor as the Chinese mini-boxes. One has a CSM or techie that you can yell at along with a support depot with parts that they can overnight, and the other tries to pretend to not understand English when the proverbial hits the fan. Funnily enough, I consider the M4 Mac Mini with the swappable storage module to be an excellent machine for homelab use - low initial cost, very low power draw, good performance, and somewhat reasonable storage upgrade path...even if the RAM is soldered in. You could conceivably do some decent stuff with it.

I think my beef with the current crop of hardware is all the weird attempts to lock you into a set ecosystem - can you reuse the nVidia DGX as a general purpose (and very overpriced) Linux box? Probably. Then what? At least with a used Mac studio you can still flock it off to a graphics designer or video nerd as long as they are okay with MacOS. A Qualcomm WinARM machine though? Only functions 100% with WinARM. Want Linux? Firmware will always firewall you off the parts Qualcomm doesn't want to you have (the same also applies for Apple and their M-series machines). The same also applies for all of those "good enough" UEFI implementations on x64 hardware which requires more and more hoops to be jumped in order for you to do things with it. Want OpenBSD on a machine that doesn't do legacy boot? Tough nuggets. Bad Secureboot implementation on your hardware won't allow you to import a new pubcert into the boot chain so you can't load dkms drivers onto a UEFI Linux install (needed to get, say, Hailo8 or Google Coral Edge TPU drivers going)? Hah hah, sucks to be you.

Oh, the modern headaches of a less refined age.
 
Last edited:
  • Like
Reactions: uldise and marcoi

Greg_E

Active Member
Oct 10, 2024
516
165
43
Oh, and only 10g because I didn't hop on the mellanox bandwagen. I over spent to get my Extreme Networks switch because I use the same at work, and they happen to have four of the 25/10 ports on them, plus stacking if I end up with more of them. It did come with dual 900 watt supplies and is a 90watt per port version.