Hi guys.
This is going to be a set of questions... various posts, as I unpack it myself.
First... Was wanting to do it / cluster as a 5 node, then came across a couple of comments, mentioning Tensor Parallel saying, 1 2 4 or 8...
What would be the impact if I could say only get hands onto 6 nodes... not 8.
Was thinking of adding a 2nd generic Intel/K8S cluster on the side, connected to the GPU cluster via 10GbE, allowing all GPU processing to be purely on the Spark's and all non / generic workload on the x86-64.
For models, and scripts etc, was looking at having a x86-64 as a jump box, doubling as a NAS, aka NFS mount... -> 4TB HDD... could go 2 x 4TB or single 1 x 8TB NVMe stick... Also allows me to run all monitoring etc from here, instead of on the GPU Cluster itself.
Fire away.
G
This is going to be a set of questions... various posts, as I unpack it myself.
First... Was wanting to do it / cluster as a 5 node, then came across a couple of comments, mentioning Tensor Parallel saying, 1 2 4 or 8...
What would be the impact if I could say only get hands onto 6 nodes... not 8.
Was thinking of adding a 2nd generic Intel/K8S cluster on the side, connected to the GPU cluster via 10GbE, allowing all GPU processing to be purely on the Spark's and all non / generic workload on the x86-64.
For models, and scripts etc, was looking at having a x86-64 as a jump box, doubling as a NAS, aka NFS mount... -> 4TB HDD... could go 2 x 4TB or single 1 x 8TB NVMe stick... Also allows me to run all monitoring etc from here, instead of on the GPU Cluster itself.
Fire away.
G