update 9-1-26:
I think i settled into setup i want for my n5pro.
RUN SPLIT CO-HOST ENGINE (Isolated Dual-Model Concurrent Array)
Focus: High-concurrency cluster layout assigning isolated targets per card.
B70 Card 0: Qwen 3.8 Dense (27B) ^| 1 Lane x 64k
Serious Dev work via hermes CLI.
B70...
That's pretty cool app.
its funny, one of my ideas for testing app building with my local llm was a home based chat app that runs locally.
i never spent a ton of time on it but it did basic functions of chat messages stored in a db. I didnt code any of it. just helped the llm along the way...
i think the same as the card im testing with is 2080 ti and it doesn't support rebar. So hoping the new intel card will be fine. It comes Monday.
Luckily i been getting psu off woot/amazon when they go on sale so i had an extra psu. The deg2 i think was like 200-250 range on sale from amazon...
I just ordered a second intel b70 card to hopefully run on my deg2 dock with the other b70 card already setup in a deg1 dock. This is running on minisforum n5pro.
btw newegg has asrock b70 for 1299 and offers a free aio cooler worth 80 bucks.
The intel version was up to 1600 and the amd 9700 ai...
well some more tweaking on the b70 on my n5pro :)
the extra stuff in the set commands seem to up the performance a lot.
Im still working on the final tweaks so i can run two agents with max performance.
rest of the n5pro (npu/igpu) near max performance so unlikely ill get more out of tweaking...
yup :)
i ordered this card to test it.
https://www.amazon.com/dp/B0G2G9J4KT?th=1
while the video card did show up in windows it did not have rbar enabled. I would love to get two external docs going and put another intel b70 in use on the one system.
the resizable bar works on the oculink port on n5 pro. I also got a deg2 dock with both oculink and tb5 ports to test with.
now i recall the issue was the n5pro doesnt have tb4-5 on it so i got a pcie oculink card, that would not allow resizable bar to work on the deg2 video card.
I never...
i would have like it to also support oculink vs all those m.2 slots. would been nice to run external card. i know usb4 docks can be an alternative. but when i tested using one on my n5pro, the cards would show up but it would not support resizable bar.
how often are you guys updating your AI stack software?
I normally try for driver/llama.cpp/lemonaid ai once a week.
Usually tie it into windows patching that requires reboots.
For hermes agent, i usually do it once a day in the AM when i start work.
That includes honcho updates as well...
it does make sense to put something in the middle to check if data goes out and to compress long tokens to save on costs. not sure anyone developed that as a business use case.
for me, i havent used any paid ai/llm as of yet. all my stuff been local based only. I do use google ai mode to do...
i had hermes create a compression test to benchmark igpu and help fine tune it.
It checks the llama metrics website for pre and post to compare details of the run.
Here's the pre/post metrics table for both runs:
METRICS COMPARISON: Run 1 vs Run 2 (75k context, gpt-oss-20b on...
i have the b70 setup for hermes agents work. so that is the main LLM model for hermes.
AMD is setup only for compressing context in hermes when it gets close the context limit of 98k. it works really well for compressing.
I then use lemonaid server and npu for honcho memory and title...
This is benchmark from llama.cpp with intel sync on the b70 card
:: 2. OVERRIDE WITH YOUR PERFORMANCE TUNING CONFIGURATIONS
set GGML_SYCL_PINNED_MEMORY=1
set ONEAPI_DEVICE_SELECTOR=level_zero:gpu
set ZEX_NUMBER_OF_DEVICE_ALLOCATIONS_PER_POOL=1
::benchmark cmd
llama-bench.exe -m...
if you have time to play with these servers, then tinker with them.
As for AI the vid cards wont be much help. you might get away with a super tiny llm.
which may be okay if you just want to play with learning software like llama.cpp etc.
If you can get a modern card with 12gb ram, the pci...
@WANg i like it a lot. even the lego planes lol.
im kind of glad i didnt end up selling the n5pro before the end of PCs came about.
It running like a champ.
I did move the ubuntu VM off the n5pro to my Xen stack so i can bump the memory of the VM up and also give the iGPU 32GB of the 96gb ram...
seems like a used AMD Radeon Pro V620 32 gb can be had for 400-500 range if it fits your budget/components. I was considering getting one to play with but my current setup seems to be dialed in enough where im happy enough with it.
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.