HighPoint Rocket 1608A add-in card review: More drives, more power

For when you really need 56 GB/s.

Highpoint Rocket 1608A charts and images
(Image credit: © Tom's Hardware)

Tom's Hardware Verdict

The HighPoint Rocket 1608A AIC is an excellent storage solution if you have the right platform and drives to spare. It’s impressive, but certainly not for everyone.

Pros

  • +

    High peak and sustained performance

  • +

    Up to 8 drives with no mobo bifurcation req’d

  • +

    Cooling and some nice additional features

  • +

    PCIe 5.0 in both directions

Cons

  • -

    A little pricey

  • -

    Not as full-featured as the 7608A

  • -

    Requires a dedicated x16 slot, preferably PCIe 5.0

  • -

    Performance issues on an Intel platforms

Why you can trust Tom's Hardware Our expert reviewers spend hours testing and comparing products and services so you can choose the best for you. Find out more about how we test.

If one high-end SSD isn’t fast enough for you, then how about eight? The HighPoint Rocket 1608A AIC (add-in card) allows you to assemble up to eight PCIe 5.0 SSDs with sixteen lanes of upstream bandwidth. If the Crucial T705’s 14 GB/s just isn’t quite doing getting the job done, you can have four or more of them work together for faster transfers and insanely high IOPS. If you can deliver the right workload, that is.

Our last review of an AIC like this from HighPoint was almost seven years ago, but we’ve had our eye on the Rocket 1608A for some months. With more PCIe 5.0 SSDs and platforms coming to the market as time goes on, there’s a natural enthusiast desire to push for more bandwidth, and this AIC delivers. You don't need to use PCIe 5.0 SSDs either — today, we’re using eight PCIe 4.0 Samsung 990 Pros. You could even use a non-PCIe 5.0 slot for that matter, as you can still benefit from steady state performance improvements, higher IOPS, and some management features of the hardware. But for maximum burst performance, you'll want to have a PCIe 5.0 x16 slot available.

The Rocket 1608A is an all-in-one solution as it provides cooling, connectivity, and an on-board PCIe switch so you don’t have to rely on motherboard bifurcation. The card and switch feature everything from indicator LEDs to deeper features, like synthetic mode. It’s quite possible to get 56 GB/s or more with the right hardware, and although the price seems steep it’s not unreasonable if you consider the advantage of not needing expensive 8TB drives to reach your capacity goals.

This solution is not for everyone, though, as the Rocket 1608A’s full potential is best met in an HEDT or enterprise environment where high levels of performance are possible and at times necessary. This ideally takes advantage of a full x16 PCIe 5.0 slot and a fast CPU that can keep up. Still, it’s worth a look as an interesting product that shows how far solid state storage has come even on the consumer end of things. If it happens to fit your budget, then it’s worth fully exploring the product and what it can do.

HighPoint Rocket Specifications

Latest Videos FromTom's Hardware
TOPICS
Shane Downing
Freelance Reviewer

Shane Downing is a Freelance Reviewer for Tom’s Hardware US, covering consumer storage hardware.

With contributions from
  • Amdlova
    60 seconds of speed after that slow as f.
    For normal user it's only burn money...
    For enterprise user? Nops
    For people with red cameras epic win!
    Reply
  • thestryker
    Would love to know what the issue with the Intel platforms is as I can't think of any logical reason for it.

    Broadcom is the reason why this card is so expensive as they massively inflated costs on PCIe switches after buying PLX Technology. There haven't been any reasonable priced switches at PCIe 3.0+ since.

    I have a pair of PCIe 3.0 x4 dual M.2 cards that I imported from China because the ones available in the US were the same generic cards but twice as much money and they were still ~$50 per. Dual slot cards with PCIe switches from western brands are typically $150+ (most of these are x8 with the exception of QNAP who has some x4).
    Reply
  • Amdlova
    thestryker said:
    Would love to know what the issue with the Intel platforms is as I can't think of any logical reason for it.

    Broadcom is the reason why this card is so expensive as they massively inflated costs on PCIe switches after buying PLX Technology. There haven't been any reasonable priced switches at PCIe 3.0+ since.

    I have a pair of PCIe 3.0 x4 dual M.2 cards that I imported from China because the ones available in the US were the same generic cards but twice as much money and they were still ~$50 per. Dual slot cards with PCIe switches from western brands are typically $150+ (most of these are x8 with the exception of QNAP who has some x4).
    Can you share the model?
    Reply
  • thestryker
    Amdlova said:
    Can you share the model?
    This is the version sold in NA which has come down in price ~$20 since I got mine (uses PCIe 3.0 switch):
    https://www.newegg.com/p/17Z-00SW-00037?Item=9SIAMKHK4H6020
    You should be able to find versions of on AliExpress for $50-60 and there are also x8 cards in the same price range, just make sure they have the right switch.

    The PCIe 2.0 one I got at the time this was the best pricing (you can find both versions for less elsewhere now): https://www.aliexpress.us/item/3256803722315447.htmlI didn't need PCIe 3.0 x4 worth of bandwidth because the drives I'm using are x2 which allowed me to save some more money. ASM2812 is the PCIe 3.0 switch and the ASM1812 is the PCIe 2.0 switch.
    Reply
  • Amdlova
    @thestryker Thanks
    Reply
  • razor512
    Amdlova said:
    60 seconds of speed after that slow as f.
    For normal user it's only burn money...
    For enterprise user? Nops
    For people with red cameras epic win!
    Sadly the issue is many SSD makers stopped offering MLC NAND. Consider the massive drop in write speeds after the pSLC cache runs out. Consider even back in the Samsung 970 Pro days.
    As long as the SOC was kept cool, the 1TB 970 Pro (a drive from 2018) could maintain its 2500-2600MB/s (2300MB/s for the 512GB drive) write speeds from 0% to 100% fill.

    So far we have not really seen a consumer m.2 NVMe SSD have steady state write speeds hitting even 1.8GB/s steady state until they started releasing PCIe 5.0 SSDs, and even then, you are only getting steady state write speeds exceeding that of old MLC drives at 2TB capacity; effectively requiring twice the capacity to achieve those speeds.

    The Samsung 990 Pro 2TB drops to 1.4GB/s writes when the pSLC cache runs out.

    If anything, I would have liked to see SSD makers continue to produce some MLC drives, since prices have come down, compared to many years ago, a modern 2TB SSD is still cheaper than 1TB drives MLC drives from those days. They could literally make a 1TB drive using faster MLC and charge the price of the 2TB TLC drives, and offer significantly higher steady state write performance for write intensive workloads.
    Reply
  • JarredWaltonGPU
    Amdlova said:
    60 seconds of speed after that slow as f.
    For normal user it's only burn money...
    For enterprise user? Nops
    For people with red cameras epic win!
    This is largely contingent on the SSDs being used. Samsung 990 Pro 2TB (which is what HighPoint provided) write at up to 6.2 GB/s or so for 25 seconds in Iometer Write Saturation testing, until the pSLC is full. Then they drop to ~1.4 GB/s, but with no "folding state" of lower performance. So best-case eight drives would be able to do 49.6 GB/s burst, and then 11.2 GB/s sustained.

    The R1608A doesn't quite hit those speeds, but having software RAID0 and a bit of other overhead is acceptable. I'm a bit bummed that we weren't provided eight Crucial T705 drives, or eight Sabrent Rocket 5 drives, because I suspect either one would have sustained ~30 GB/s in our write saturation test.
    razor512 said:
    Sadly the issue is many SSD makers stopped offering MLC NAND. Consider the massive drop in write speeds after the pSLC cache runs out. Consider even back in the Samsung 970 Pro days. As long as the SOC was kept cool, the 1TB 970 Pro (a drive from 2018) could maintain its 2500-2600MB/s (2300MB/s for the 512GB drive) write speeds from 0% to 100% fill.

    So far we have not really seen a consumer m.2 NVMe SSD have steady state write speeds hitting even 1.8GB/s steady state until they started releasing PCIe 5.0 SSDs, and even then, you are only getting steady state write speeds exceeding that of old MLC drives at 2TB capacity; effectively requiring twice the capacity to achieve those speeds.

    The Samsung 990 Pro 2TB drops to 1.4GB/s writes when the pSLC cache runs out.

    If anything, I would have liked to see SSD makers continue to produce some MLC drives, since prices have come down, compared to many years ago, a modern 2TB SSD is still cheaper than 1TB drives MLC drives from those days. They could literally make a 1TB drive using faster MLC and charge the price of the 2TB TLC drives, and offer significantly higher steady state write performance for write intensive workloads.
    You need to look at our many other SSD reviews, where there are tons of drives that sustain way more than 1.4 GB/s. The Samsung 990 Pro simply doesn't compete well with newer drives. Samsung used to be the king of SSDs, and now it's generally just okay. 980/980 Pro were a bit of a fumble, and 990 Pro/Evo didn't really recover.

    There are Maxio and Phison-based drives that clearly outperform the 990 Pro in a lot of metrics. Granted, most of the drives that sustain 3 GB/s or more are Phison E26 using the same basic hardware, and most of the others are Phison E18 drives. Here's a chart showing one E26 drive (Rocket 5), plus other E18 drives that broke 3 GB/s sustained, with one 4TB Maxio MAP1602 that did 2.6 GB/s:

    359
    Reply
  • abufrejoval
    It's a real shame that now that we have PCIe switches again, which are capable of 48 PCIe v5 speeds and can actually be bought, the mainboards which would allow easy 2x8 bifurcation are suddenly gone... they were still the norm on AM4 boards.

    Apart from the price of the AIC, I'd just be perfectly happy to sacrifice 8 lanes of PCIe to storage, in fact that's how I've operated many of my workstations for ages using 8 lanes for smart RAID adapters and 8 for the dGPU.

    One thing I still keep wondering about and for which I haven't been able to find an answer: do these PCIe switches effectively switch packets or just lanes? And having a look at the maximum size of the packet buffers along the path may also explain the bandwidth limits between AMD and Intel.

    Here is what I mean:

    If you put a full complement of PCIe v4 NVMe drives on the AIC, you'd only need 8 lanes of PCIe v5 to manage the bandwidth. But it would require fully buffering the packets that are being switched and negotiating PCIe bandwidths upstream and downstream independently.

    And from what I've been reading in the official PCIe specs, full buffering of the relatively small packets is actually the default operational mode in PCIe, so packets could arrive at one speed on its input and leave at another on its output.

    Yet what I'm afraid is happening is that lanes and PCIe versions/speeds seem to be negotiated both statically and based on the lowest common denominator. So if you have a v3 NVMe drive, it will only ever have its data delivered at the PCIe v5 slot at v3 speeds, even if within the switch four bundles of four v3 lanes could have been aggregated via packet switching to one bundle of four v5 lanes and thus deliver the data from four v3 drives in the same time slot on a v5 upstream bus.

    It the crucial difference between a lane switch and a packet switch and my impression is that a fundamentally packet switch capable hardware is reduced to lane switch performance by conservative bandwidth negotiations, which are made end-point-to-end-point instead of point-to-point.

    But I could have gotten it all wrong...
    Reply
  • JarredWaltonGPU
    thestryker said:
    Would love to know what the issue with the Intel platforms is as I can't think of any logical reason for it.
    I think it's just something to do with not properly providing the full x16 PCIe 5.0 bandwidth to non-GPU devices? Performance was bascially half of what I got from the AMD systems. And as noted, using non-HEDT hardware in both instances. Threadripper Pro or a newer Xeon (with PCIe 5.0 support) would probably do better.
    abufrejoval said:
    Here is what I mean:

    If you put a full complement of PCIe v4 NVMe drives on the AIC, you'd only need 8 lanes of PCIe v5 to manage the bandwidth. But it would require fully buffering the packets that are being switched and negotiating PCIe bandwidths upstream and downstream independently.
    Yeah, that's wrong. There are eight M.2 sockets, each with a 4-lane connection. So you need 32 lanes of PCIe 4.0 bandwidth for full performance... or 16 lanes of PCIe 5.0 offer the same total bandwidth. If you use eight PCIe 5.0 or 4.0 drives, you should be able to hit max burst throughput of ~56 GB/s (assuming 10% overhead for RAID and Broadcom and such). If you only have an x8 PCIe 5.0 link to the AIC, maximum throughput drops to 32 GB/s, and with overhead it would be more like ~28 GB/s.

    If you had eight PCIe 3.0 devices, then you could do an x8 5.0 connection and have sufficient bandwidth for the drives. :)
    Reply
  • abufrejoval
    JarredWaltonGPU said:
    Yeah, that's wrong. There are eight M.2 sockets, each with a 4-lane connection. So you need 32 lanes of PCIe 4.0 bandwidth for full performance... or 16 lanes of PCIe 5.0 offer the same total bandwidth. If you use eight PCIe 5.0 or 4.0 drives, you should be able to hit max burst throughput of ~56 GB/s (assuming 10% overhead for RAID and Broadcom and such). If you only have an x8 PCIe 5.0 link to the AIC, maximum throughput drops to 32 GB/s, and with overhead it would be more like ~28 GB/s.

    If you had eight PCIe 3.0 devices, then you could do an x8 5.0 connection and have sufficient bandwidth for the drives. :)
    Well wrong about forgetting that it's actually 8 slots instead of the usual 4 I get on other devices: happy to live with that mistake!

    But right, in terms of the aggregation potential would be very nice!

    So perhaps the performance difference can be explained by Intel and AMD using different bandwidth negotiation strategies? Intel doing end-to-end lowest common denominator and AMD something better?

    HWinfo can usually tell you what exactly is being negotiated and how big the buffers are at each step.

    But I guess it doesn't help identifying things when bandwidths are actually constantly being renegotiated for power management...
    Reply