Sunken cost fallacy is certainly at play here. And yes, the Hailo stuff is, well, not supported on the software side as well as it should be. Let's throw a product out there and hope the community makes it work for us.
What I'm wanting might be simple enough right now that I can get where I want to go with this rig. What I want/need is something where I can say, how do I write a VBS login script to do this, that, and the other thing. Also would be nice to feed it a spreadsheet and ask it to remove all PII and replace with random number and provide a cross reference to be able to get back to the user after things are analyzed. There are some other compare this to that and suggest a result things that my wife could probably use, she's been paying Claude for a lot of things like this, mostly dealing with a cemetery that she volunteers for, it's from the 1700's and records are a mess. Sorting who is located where, and more importantly which spots are still available to be sold are part of the priorities. If any of that takes 15 minutes or more, that's just part of the game right now.
Now back to video cards, I do have one of those AMD E9173 cards at home, but I looked it up and it didn't seem like it was worth much from the feature list. If it might be useful, I can try it too, again probably in Vulkan mode. For video work it looks better then the k1200 (encode and decode), just wish it had more ram. Looks like 35 watts is near max for the PCIe slot in this thin client, that's the rating on the AMD 9173 card, I think the t1000 is similar. I also have some k620 and maybe a p620 card here, but older is generally not better, and lower vram is not better.
I'm leaning towards LocalAI because it says it can use all these different things (CPU, Vulkan, GPU, etc.) all at the same time, being resource constrained, moving the processing to all available seemed like a good idea. I know the expense will be speed, but I'm not trying to render out a motion picture right now.
I'd also like to add this, thanks for everyone trying to get me going, there are a lot of good people here spending their time to help others. We rarely see the words, "thanks for the help", so thanks for the help in getting me going with some possible choices and letting me know that there are probably some severe limitations in my current choices. Money being what it is, we are all scraping for scraps these days, stretching the older hardware as far as we can stretch it.
Yeah, but the more important question is...how much money are you planning to sink into Local LLMs and how prepared are you for disappointment? I would advise not to prematurely optimize things before you run it for the first time to get a better idea of where things can be improved, or rather, if you even want to invest in it. The Vega 8 graphics chip (gfx903) on the t740 comes with the machine for free and performs okay on Vulkan, so you might as well give it a try. Since it's an iGPU it uses DDR4 system memory anyways, and as long as you are okay with the idea that big models will run slow due to limited compute/RAM bandwidth, you'll not be overly disappointed.
As for that E9173, well, it's GCN4 (not GCN5 like in the Vega) so it's potentially just as slow as the iGPU. Even with the T1000 you are still dealing with the same performance ballpark as, say, the Radeon 780M iGPU on a Phoenix/Hawk Point APU (gfx1103), more or less, so it's not like it's a mind boggling significant improvement. You could wire the thin client up to a beefier GPU but that'll be like buying a nugget car (1997 Toyota Tercel) for 500 bucks and then doing a Chevy LS engine swap on it (~6k or so to stick the engine from an older Corvette). At some point if you want a cheap fast GPU and don't mind dealing with a power hungry howling one in a server...just buy an
AMD Radeon Pro V620. Or use some kind of dock...but that t740 is pretty much at the end of the road here.
If I have to guess its performance? On a Gemma-4 12 Billion parameter model with MTP, probably about 175 prompt tokens/sec, and ~5-7 tokens/sec during the thinking process on Vulkan with a 32k context window, based on the numbers I got from my HP mt46 thin client, which in terms of horsepower sits between the t740 and the t755. Most people consider a minimum of 25 tokens/sec sustained during the thinking process to be roughly acceptable, but 75+ is usually desired. Prompt processing is more of a “understanding what you want before responding” thing and unless you have very limited compute or you ask a ridiculously complex question most prompt processing only takes a few seconds before the thinking/tokenization begins. Oh, and don’t even think about using CPUs for inferencing. Most CPUs with paired integrated graphics usually deliver 1/2 to 1/3 of the inferencing horsepower power of its attached iGPUs, so unless there's something hilariously bad about it, just stick to iGPUs on Vulkan.
As for the utility of an LLM, heres some stuff to think about -
a) You don't ask an LLM to recall facts that are easily looked up. In fact, LLMs are terrible at it because they are trained with a data set that are typically 1-2 years out of date, they are designed to not take advantage of new data (unless the LLM is engineered to allow new learning...which most aren't), and they are prone to hallucinations.
Here's a simple question - name the first 6 stations of the Tokyo Metro Hanzomon line from Shibuya towards Oshiage (oh, and in case you don't trust Wikipedia, here's
the line map from Tokyo metro themselves)
That should be:
Z01 - Shibuya (great for shopping and people watching)
Z02 - Omote-Sando (great if you like ostentatious stuff)
Z03 - Aoyama-Itchome (fairly boughie area)
Z04 - Nagatacho (Next to Tokyo Broadcasting System HQ and the Harry Potter museum/staircase, also boughie)
Z05 - Hanzomon (the western gate of the Imperial Palace near the western moat)
Z06 - Kudanshita (next to several Tokyo Universities and on the north side of the Imperial palace moat)
Oh, let's ask this 20 billion parameter ChatGPT model for an answer -
Note the hallucinations developing...I actually had to stop the execution in 45 seconds as it was spinning in circles…

Amongst LLMs this is actually much more common than you think.
What about Qwen3.6 (from Ali Baba's AI labs)? Here's one with 27 billion parameters annd a quantization value of 4 (Q4) set for multi-token prediction (MTP)...
(Harajuku is on the JR East Yamanote line...not Tokyo Metro) - so it got the info incorrect...even when the other stations are on the line.
What about qwen3-next-REAM-Q3 (27.5GB)?
(With the exception of the very first station, the ones on the list are Tokyo Metro subway stations, but they are NOT on the Hanzomon line. So that one is pretty much 100% wrong)
Don't depend on LLMs for recalling things.
I had a 35 Billion parameter Qwen 3.5 model insist that the World trade center and the world financial center in NYC is the same building (they are adjacent complexes downtown but NOT The same), and I had one try to convince me that Gil Hodges played Mr. Spork on Star Trek. Since the dataset is baked in, you cannot correct it.
The tokenized information are connected together using semantic weights during the training process - if those connections are altered it can create false connectivity graphs ...or hallucinations.
b) The smaller the model and the more "niche" the training data, the worse the inferencing due to weak semantic cross checking -
Here's an example -
Ask a Deepseek/Qwen3 model with 8 billion parameters this question:
"Please suggest a fun summertime new england clambake for a Haredi congregation including sample recipes"
Note: Please do not invite your local rabbi and his congregation out for a clambake. At least not without rebranding it to be a summer picnic (that's a major faux pas like inviting your muslim friends to a summer pig-out), definitely hire a glatt kosher caterer and for the love of all that is good and holy, DO NOT ASK QWEN FOR SUGGESTIONS. The first thing Gemma-4 (Google Gemini Lab's LLM) did was to suggest calling it a summer picnic event instead.
In case you don't pick up on why this conversation is very bad, shellfish is never, ever, ever kosher. Even if you serve kosher steak or chicken in your meal instead of seafood you cannot serve dairy products with the same meal (i.e. definitely no butter on your corn).
Basically this tells me that some LLMs don't really put any strong semantic linkages between dogmatic religious concepts (such as kashrut or kosher laws) nearly as strenuously as it should or cross check the thinking process to look for discrepancies.
In general, after playing with qwen 3.6 and gemma-4 for a while, qwen tends to run/tokenize faster but gemma-4 tends to be more disciplined in accuracy and cross-checking, but both are prone to hallucinations even on high quantization models.

(Making sure you use kosher butter with your kosher clams and shrimp is…definitely an interesting result)
c) It could do some interesting stuff - but there are limits....
Well, here it is taking cues from an image and forming a narrative....
I think this is a 27 billion parameter Qwen 3.6 model.
That's not bad at all.
If you have access to Ace-Step music, you could abuse it to make some terrible, terrible music (let me know if you want to hear the farty-trip-hop-ska)
As for taking a bunch of files and making sense of them or using it to write powershell scripts, yes...it can do that. To be honest, I had Gemma-4-12B generate this one, but I have yet to test whether it's sensible or not. To be honest i can probably gank similar code from examples on reddit/stackexchange or MSDN, and get similar or better code going.
(oh yeah, the entire thing ran off the Radeon 780M iGPU on the Minisforum N5 Air - if I fire up the V620 I’ll use up about 4x more power to get a roughly 3x speedup)
It's fun to play around with LLMs but don't lose sight that they need to be supervised, and don't put too much undue hope on them. As Ed Zitron would’ve put it - they are great make-work agents for middle management types, and it can potentially boost productivity for those with a cluebat, but you need discipline and understand their limits.