Intel Optane PMEM 200 256GB DDR4 3200MHz $199

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

miraculix

Active Member
Mar 6, 2015
167
59
28
Seller accepted $181. I wouldn't call this a great deal per se, but with all the RAM price craziness it may be a good option for anyone looking to maximize overall usable memory on compatible systems (for example SM X12SPx boards with Ice Lake Xeon SP3s).

Pardon Our Interruption...

Just be sure you understand how these are used, operating modes, RDIMM vs PMEM mix, motherboard compatibility etc. before jumping in!
 
Last edited:

luckylinux

Well-Known Member
Mar 18, 2012
1,675
548
113
Seller accepted $181. I wouldn't call this a great deal per se, but with all the RAM price craziness it may be a good option for anyone looking to maximize overall usable memory on compatible systems (I use PMEM 200s on SM X12SPM boards with Ice Lake Xeon SP3s).

Pardon Our Interruption...

Just be sure you understand how these are used, operating modes, RDIMM vs PMEM mix, motherboard compatibility etc. before jumping in!
I already saw many of these Intel Xeon Scalable >= Gen 2 (?) "DIMMs" that where quite cheap but to be honest, from my somewhat poor Understanding at the Time, it seemed that they are more kind of a Flash NVMe SSD than a real RAM DIMM and as such they need to be paired with suitable RDIMM/LRDIMM.

Obviously I might be quite far off, but basically assume this is only usable as a Fast Cache Drive, unless proven otherwise.
 
  • Like
Reactions: abq

miraculix

Active Member
Mar 6, 2015
167
59
28
I already saw many of these Intel Xeon Scalable >= Gen 2 (?) "DIMMs" that where quite cheap but to be honest, from my somewhat poor Understanding at the Time, it seemed that they are more kind of a Flash NVMe SSD than a real RAM DIMM and as such they need to be paired with suitable RDIMM/LRDIMM.

Obviously I might be quite far off, but basically assume this is only usable as a Fast Cache Drive, unless proven otherwise.
It may depend on specific motherboard support but these can operate in either App Direct Mode (i.e. persistent/non-volatile storage similar to what you describe) or Memory Mode (volatile, the regular RDIMMs act as cache in from of these, results in a much greater pool of usable RAM).

This guide and this guide were handy for my X12SPM systems.
 
Last edited:
  • Like
Reactions: abq and luckylinux

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
I already saw many of these Intel Xeon Scalable >= Gen 2 (?) "DIMMs" that where quite cheap but to be honest, from my somewhat poor Understanding at the Time, it seemed that they are more kind of a Flash NVMe SSD than a real RAM DIMM and as such they need to be paired with suitable RDIMM/LRDIMM.

Obviously I might be quite far off, but basically assume this is only usable as a Fast Cache Drive, unless proven otherwise.
It depends on the mode you set for the PMEM. In memory mode, total ram seen by system is the total optane size. Ram then acts like cache. In app mode, Optane and Ram are separated but this also requires your application to support optane ram to even use it as persistent memory. Then you can also run in mixed mode of the two. The caveat of memory mode is that you will have the same operating ram speed up until the system memory goes beyond the size of the RAM cache. Memory size beyond your total RAM size then optane becomes your swap. The real problem is OS support. Windows will require specific Intel RST drivers and I think not all linux distro support it out of the box (I know Debian does).
 

iraqigeek

Active Member
Sep 17, 2018
137
115
43
Wasn't there a General Requirement that, in order to even Boot/Post, there was a minimum Requirement for RDIMM/LRDIMM anyway ?

Or can you run 100% in these PMEM without additional RDIMM/LRDIMM ?
IIRC, You need one RAM stick for each Optane stick, on the same channel (again, IIRC).
 

luckylinux

Well-Known Member
Mar 18, 2012
1,675
548
113
It may depend on specific motherboard support but these can operate in either App Direct Mode (i.e. persistent/non-volatile storage similar to what you describe) or Memory Mode (volatile, the regular RDIMMs act as cache in from of these, results in a much greater pool of usable RAM).

This guide and this guide were handy for my X12SPM systems.
Wasn't there a General Requirement that, in order to even Boot/Post, there was a minimum Requirement for RDIMM/LRDIMM anyway ?

Or can you run 100% in these PMEM without additional RDIMM/LRDIMM ?
 

iraqigeek

Active Member
Sep 17, 2018
137
115
43
It depends on the mode you set for the PMEM. In memory mode, total ram seen by system is the total optane size. Ram then acts like cache. In app mode, Optane and Ram are separated but this also requires your application to support optane ram to even use it as persistent memory. Then you can also run in mixed mode of the two. The caveat of memory mode is that you will have the same operating ram speed up until the system memory goes beyond the size of the RAM cache. Memory size beyond your total RAM size then optane becomes your swap. The real problem is OS support. Windows will require specific Intel RST drivers and I think not all linux distro support it out of the box (I know Debian does).
IIRC, there's also a mode in which you can run them as regular storage. It's seen by the OS as a regular block device, but you can't install/boot the OS from it, and you can partition the Optane pool you have however you want between those three modes.
 
  • Like
Reactions: nexox

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
IIRC, there's also a mode in which you can run them as regular storage. It's seen by the OS as a regular block device, but you can't install/boot the OS from it, and you can partition the Optane pool you have however you want between those three modes.
That's right and this is a really weird mode because I could only ever get a maximum of 512GB in this mode instead of the full 1TB i have.
 

iraqigeek

Active Member
Sep 17, 2018
137
115
43
That's right and this is a really weird mode because I could only ever get a maximum of 512GB in this mode instead of the full 1TB i have.
Single socket or dual? I had no issue allocating the full 1TB, but that's in a dual socket system with four 256GB sticks.

Intel did some weird things with optane that depended on which CPU you have and how much RAM that supports.
 

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
Single socket or dual? I had no issue allocating the full 1TB, but that's in a dual socket system with four 256GB sticks.

Intel did some weird things with optane that depended on which CPU you have and how much RAM that supports.
Single socket with 2.512gb sticks optane and 6 64gb ddr4 ram.
 

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
what CPU ? some can only 1TB
maybe you need to clear AD config first ?
Intel Xeon Platinum 8260Y (ES) QRC1 on Supermicro MotherBoard X11SPi-TF (bios updated to the latest). It could be because it's an ES butI've cleared the AD config a few times trying to figure out why only 512GB is allowed. It recognizes the full terabyte but when I allocate and save it changes to 512GB.
 

iraqigeek

Active Member
Sep 17, 2018
137
115
43
Intel Xeon Platinum 8260Y (ES) QRC1 on Supermicro MotherBoard X11SPi-TF (bios updated to the latest). It could be because it's an ES butI've cleared the AD config a few times trying to figure out why only 512GB is allowed. It recognizes the full terabyte but when I allocate and save it changes to 512GB.
I also have ES (QQ89). How much RAM do you have? The memory limit Rollo mentioned is for the combined RAM + Optane.
 

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
Which is why I am perplexed that only 512GB the limit. Oh well.
I took a deeper dive into this issue. There's Memory mode, Hybrid, and Appdirect for the 100 series. The hard limit of 512GB was actually my misunderstanding of the Supermicro menu. I can allocate more than 512GB toward Appdirect in hybrid mode but I would run into the ghost ram issue where the system treats the remaining optane as system RAM and the DDR4 RAM as cache for the Optane modules. Which means that if I allocate 90% to Appdirect in hybrid mode then remaining system memory is only 10% of the 1TB where by disregarding the 384GB of DDR4.

When you use Appdirect only then there's a hard limit of 768GB limit per cpu (the exception are M and L suffix cpu models which can support 2T-4T of total Ram) of Appdirect + DDR. Since I have 384GB of DDR4 then I can only allocate 192GB per stick to pure Appdirect and the remain chunk of Optane goes unused. To get high ratio of Appdirect to memory, I would need to reduce the DDR4 size.

It's all very confusing and it took me a good 5 or 6 hours to get there having wasting 3 hours trying to get ipmctl to compile in Debian 13 (there's a hard dependency with another library to include in the compile for Supermicro motherboards).

Also, hybrid mode only works with the 100 series and for a good reason, memory latency sucks when running in this mode (this is from research reading and not direct testing, take that knowledge with a grain of salt). Either run full memory mode or Appdirect.
 
Last edited:

zachj

Active Member
Apr 17, 2019
319
169
43
For the very vast majority of home lab use cases the negative impact of optane on memory latency is a nonissue.

the two primary use cases are great:
1. A higher performance SSD compared to NVMe
2. More system memory to run things like VMs and sql databases

it was never intended to host an LLM on optane persistent memory.

as long as your expectations are realistic there is nothing at all wrong with optane. It does exactly what it says on the box.
 
  • Like
Reactions: abq

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
For the very vast majority of home lab use cases the negative impact of optane on memory latency is a nonissue.

the two primary use cases are great:
1. A higher performance SSD compared to NVMe
2. More system memory to run things like VMs and sql databases

it was never intended to host an LLM on optane persistent memory.

as long as your expectations are realistic there is nothing at all wrong with optane. It does exactly what it says on the box.
It's perfect for LLM use case. I'm using it right now. It's intended for any software that supports appdirect mode which inference engines like llamacpp supports. Super fast model loading and better latency when storing cache that doesn't stutter in the case of NVME offloading.

The intention of optane was for low latency for applications that demands it. And LLM happens to be one of those use cases today.

Also fyi for anyone that read my last post above. Those quirks only apply to cascade lake xeon skus. Ice Lake and Sapphire Rapids does not have these arbitrary limitations.
 
  • Like
Reactions: abq and int0x2e

iraqigeek

Active Member
Sep 17, 2018
137
115
43
It's perfect for LLM use case. I'm using it right now. It's intended for any software that supports appdirect mode which inference engines like llamacpp supports. Super fast model loading and better latency when storing cache that doesn't stutter in the case of NVME offloading.
Do you mind sharing some details about how you're using it? I have four 256GB sticks sitting doing nothing and a dual Cascade Lake ES system (with six Mi50s) for LLM inference.
 

foureight84

Well-Known Member
Jun 26, 2018
477
411
63
Do you mind sharing some details about how you're using it? I have four 256GB sticks sitting doing nothing and a dual Cascade Lake ES system (with six Mi50s) for LLM inference.
I have the Optane modules set for AppDirect. This can be done in either the Bios or using impctl (depending in your distro, you might need to compile from source -- which I would recommend getting an AI agent to help since there's some complexity that isn't mentioned on the repo. It was difficult with my supermicro motherboard due to an external dependency--a package called edk2).

If you're using ipmctl then it's something like sudo ipmctl create -goal PersistentMemoryType=AppDirect .

You want to set your optane goal in the bios to use 100% of optane memory for appdirect (there's an appdirect max on no-suffix, M and L suffix CPUs. this cap was dropped in later gens).

Then use ndctl (you should be able to install this through your package manager). sudo ndctl create-namespace --mode=fsdax This turns the appdirect capacity to a DAX block device. Then format it to a DAX supported file extension, fs4, xfs and a few other. xfs will probably offer the maximum performance but I am just currently using fs4. (the block should show up as /dev/pmem0 if you run lsblk)

After this you can mount it and add it to your fstab for automount. Run ls -l /dev/disk/by-id and use the device id for fstab mounting.

The next step is to put your gguf on the DAX drive and use it with llamacpp. In my instance, I am using ik_llamacpp. I am seeing the models load within just a few seconds. You'll also want to put your llamacpp cache on there as well. This is super useful when you use llama-swap and now you can switch models for different usage scenarios from your llm agent and only have to wait a few seconds.

If you're running llamacpp or ik_llamacpp in docker, make sure to also passthrough the /dev/pmem0 device as well (I don't think it's necessary but I do it anyway).

You can also offload your docker data to the DAX drive as well.
sudo systemctl stop docker
Create a folder ex. /mnt/pmem0/docker-data
edit /etc/docker/daemon.json and add:

JavaScript:
{
  "data-root": "/mnt/pmem0/docker-data"
}
sudo rsync -aP /var/lib/docker/ /mnt/pmem0/docker-data
sudo systemctl start docker

If you're running a database on docker, instead of using docker volume mount, use bind mounting where the path for the database storage is bound to a folder on the DAX drive (e.g. /mnt/pmem0/postgres/data and docker run ... -v /mnt/pmem0/postgres/data:/data). Volume mounting uses docker's overlayfs which negates the benefits. Lastly, don't forget database specific configuration flags specifically for running with pmem.
 
Last edited: