Building fastest possible stacking workstation
-
VideoMacro
- Posts: 1
- Joined: 05.08.2020 20:48
Building fastest possible stacking workstation
I am currently building a new workstation and want to know what hardware I should purchase for fastest possible stacking. Would helicon focus take advantage of a dual Titan rtx setup with a 64 core threadripper? Would I need to run multiple instances of the software to utilize all system resources? Is there a bottleneck I should be mindful of? Thanks for the info!
Re: Building fastest possible stacking workstation
We don't support dual graphics cards and that won't change in the foreseeable future, so one Titan RTX is all we can use. And you don't need a super powerful CPU if you're going to use OpenCL, something with a better balance between multi-threaded and single-threaded performance should improve the total rendering time over the 64-core monster that is destined to run at slower clock speed. Something like Ryzen 3950X or even 3900X will probably perform as good as it gets. If you intend to use the raw-in-DNG-out workflow, get a fast NVME storage drive for Helicon Focus cache folder, it will let you fully take advantage of our "Parallel image loading" option as well as the cache folder itself.
Re: Building fastest possible stacking workstation
My Ryzen 7 3700X works really well with Helicon, so a 3950X would be even faster.
I also use a PCIe 4 NMVE drive and a low end Quadro card. The combination means the stacking is very fast. I use a RAM cache drive for the Helicon cache folder (using SoftPerfect RAM Disk). That is even faster than the GIGABYTE GP-ASM2NE6100TTTD (SSD) which has reads speeds of 5,500 and write of 4,200 MBs. The RAM Drive is over 10,500 MBs read and 7,500 MBs writes using 3200MHz RAM.
If I overclock the CPU, my benchmarks go a lot higher. My 8 core CPU competes against 12 and 16 core CPUs in the Helicon benchmarks, if overclocked to 4GHz.
I have a good match between my CPU, MB, RAM, SSD and RAM drive with no obvious bottle necks. I stack using the CPU as my workstation graphics card is a few years old and OpenCL performance is weak. CPU performance when stacking is so good I did not feel the need to upgrade the graphics card.
I also use a PCIe 4 NMVE drive and a low end Quadro card. The combination means the stacking is very fast. I use a RAM cache drive for the Helicon cache folder (using SoftPerfect RAM Disk). That is even faster than the GIGABYTE GP-ASM2NE6100TTTD (SSD) which has reads speeds of 5,500 and write of 4,200 MBs. The RAM Drive is over 10,500 MBs read and 7,500 MBs writes using 3200MHz RAM.
If I overclock the CPU, my benchmarks go a lot higher. My 8 core CPU competes against 12 and 16 core CPUs in the Helicon benchmarks, if overclocked to 4GHz.
I have a good match between my CPU, MB, RAM, SSD and RAM drive with no obvious bottle necks. I stack using the CPU as my workstation graphics card is a few years old and OpenCL performance is weak. CPU performance when stacking is so good I did not feel the need to upgrade the graphics card.
-
clausgiloi
- Posts: 34
- Joined: 12.09.2024 16:27
Re: Building fastest possible stacking workstation
Came across this 5-year old thread and now I have the same question.
Has anything changed? I would be extremely interested in your current thoughts on configuring a PC for maximum performance. Are the below recommendations still completely current? What would be the current processor and graphics card that leaves nothing on the table? Thanks!
Has anything changed? I would be extremely interested in your current thoughts on configuring a PC for maximum performance. Are the below recommendations still completely current? What would be the current processor and graphics card that leaves nothing on the table? Thanks!
Re: Building fastest possible stacking workstation
That's easy to answer: top performance comes from the top hardware. AMD Threadripper Pro 9995WX, Nvidia RTX 5090.clausgiloi wrote: 15.01.2026 20:38 What would be the current processor and graphics card that leaves nothing on the table?
-
clausgiloi
- Posts: 34
- Joined: 12.09.2024 16:27
Re: Building fastest possible stacking workstation
So has the situation changed since you wrote this earlier in the thread?
"And you don't need a super powerful CPU if you're going to use OpenCL, something with a better balance between multi-threaded and single-threaded performance should improve the total rendering time over the 64-core monster that is destined to run at slower clock speed."
Thanks for your input. I want to run Helicon very fast, but don't want to spend blindly either.
Claus
"And you don't need a super powerful CPU if you're going to use OpenCL, something with a better balance between multi-threaded and single-threaded performance should improve the total rendering time over the 64-core monster that is destined to run at slower clock speed."
Thanks for your input. I want to run Helicon very fast, but don't want to spend blindly either.
Claus
Re: Building fastest possible stacking workstation
No, it hasn't changed, but you asked for nothing to be left on the table, not for best value.
-
clausgiloi
- Posts: 34
- Joined: 12.09.2024 16:27
Re: Building fastest possible stacking workstation
I process several hundred GB of image data in Helicon Focus pretty much every day, and I'm really looking to buy a faster - much faster - desktop setup, but I'm having a hard time finding the benchmark data to make a decision.
All I have seen is high end Windows workstations scoring around 2000 (which would be well short of the performance I'm looking for), and a single data point of the Apple M3 pro scoring way way higher.
At this point, I'd even consider an M3 or M4 studio even though I have never owned or used an Apple product, if it were clear that it is a better architectural match and will run HF faster.
I think I outlined my question badly as if money were no object. Of course it's always an object. I could really use advice more along the lines of: What is the best setup for stacking for $5k? $10k? and how much faster will the $10k setup be?
Do you have any benchmark numbers from higher end systems you could share with me?
Thank you Catherine,
Claus
All I have seen is high end Windows workstations scoring around 2000 (which would be well short of the performance I'm looking for), and a single data point of the Apple M3 pro scoring way way higher.
At this point, I'd even consider an M3 or M4 studio even though I have never owned or used an Apple product, if it were clear that it is a better architectural match and will run HF faster.
I think I outlined my question badly as if money were no object. Of course it's always an object. I could really use advice more along the lines of: What is the best setup for stacking for $5k? $10k? and how much faster will the $10k setup be?
Do you have any benchmark numbers from higher end systems you could share with me?
Thank you Catherine,
Claus
Re: Building fastest possible stacking workstation
That result is an error. The Apple systems are fast, but they're nothing special. They are more focused on performance per watt than absolute top of the line performance.
What's your workflow? Are you processing raw images? Do you open them directly, or export from Lightroom/Capture One/somewhere else?
-
clausgiloi
- Posts: 34
- Joined: 12.09.2024 16:27
Re: Building fastest possible stacking workstation
Helicon is first in my workflow - I stack right off the card (CF-A 900MB/s). Unless Helicon does repeated uncached reads of the same input file, I don't think I have anything to gain by copying to faster media first.
I use Adobe Raw-to-DNG converter to convert the 14-bit compressed RAW files (~24MB). Stacks are typically 200 images, about 4.8GB each.
I stack everything in batch mode, but have to limit myself to about 30-40 stacks per batch, otherwise the >1TB in temporary files may fill my disk. This is something I am hoping to improve with a bigger machine. I'd like to be able to batch load about 100 stacks, which would be about 480GB - a full CF-A card in the size I typically use.
I cull these outputs by deleting most, re-render and re-touch until I have only the keepers. Only then do I save to disk. Then I usually sharpen in Topaz and import into Lightroom cloud storage, where they remain as DNG+ACR for future editing flexibility.
A 200 image stack currently takes about 90 seconds off the card. Re-rendering the same stack from cached data takes about 10s.
6 seconds is the lower bound possible just based on the CF-A card read rate. Real world lower bound is surely higher, depends partly on how efficient HF/Windows is at reading off the card. I'm hoping 10-20s per stack is achievable.
I am truly grateful for any advice and info you have on this, thank you!
Claus
I use Adobe Raw-to-DNG converter to convert the 14-bit compressed RAW files (~24MB). Stacks are typically 200 images, about 4.8GB each.
I stack everything in batch mode, but have to limit myself to about 30-40 stacks per batch, otherwise the >1TB in temporary files may fill my disk. This is something I am hoping to improve with a bigger machine. I'd like to be able to batch load about 100 stacks, which would be about 480GB - a full CF-A card in the size I typically use.
I cull these outputs by deleting most, re-render and re-touch until I have only the keepers. Only then do I save to disk. Then I usually sharpen in Topaz and import into Lightroom cloud storage, where they remain as DNG+ACR for future editing flexibility.
A 200 image stack currently takes about 90 seconds off the card. Re-rendering the same stack from cached data takes about 10s.
6 seconds is the lower bound possible just based on the CF-A card read rate. Real world lower bound is surely higher, depends partly on how efficient HF/Windows is at reading off the card. I'm hoping 10-20s per stack is achievable.
I am truly grateful for any advice and info you have on this, thank you!
Claus
Re: Building fastest possible stacking workstation
Yes, there's no point.clausgiloi wrote: 28.01.2026 05:52 I don't think I have anything to gain by copying to faster media first.
There was in issue with managing the free space reserve in the recent versions of Helicon Focus, we have fixed it yesterday, try the new version. It should be working fine now: https://www.heliconsoft.com/downloads/HeliconFocus.execlausgiloi wrote: 28.01.2026 05:52 have to limit myself to about 30-40 stacks per batch, otherwise the >1TB in temporary files may fill my disk.
And if it doesn't - let us a know (the best way is to send a bug report after an issue occurs).
The second render is done completely from caches and does not read anything from the memory card. It also works with already raw-decoded files. The first render can never be as fast because it needs to read the files and decode raw files. 10 seconds for the 2nd run is already very fast and there's not much improvement possible, but if this is on your laptop then I expect a modern desktop GPU should shave a few more seconds off.
The primary improvement to be made is a beefy CPU to speed up decoding the raw files. A fast SSD is required so that it won't be a bottleneck, but your current one might be good enough already. But for new builds with high-end CPUs (8 cores or more) we recommend a PCI-E 5.0 SSD.
By the way, try changing the "Parallel image loading" setting to Enable if you have it on Auto, that might make it faster.