By comparison, the NVIDIA GTX 780 Ti GPU has single precision throughput of 5046 GFLOPS, maximum TDP of 250W, and unit cost of $699. This yields 7.22 FLOPS per dollar and 20 FLOPS per Watt.
This makes the GPU board 8 times cheaper and at least 1.1 times more power efficient than the Parallella board.
Note that the average cost of electrical power in the US is about 10.5 cents per kilowatt-hour (kWh), which means a single GPU board running 24 hours at peak utilization would cost about 63 cents per day - so it would take over 3 years of 24/7 peak operation for the nominal power cost to equal the upfront unit cost. So while power efficiency is very important, that doesn't mean the unit cost can be ignored.
In addition to FLOPS/$ and FLOPS/W, one must also consider the processing power density, which essentially measures how many "FLOPS per cubic foot" are yielded by each device. This is important because these machines take up physical space and physical space (i.e. real estate) is costly. A single GPU board has 56 times the throughput of a single Parallella board, and I would say that a GPU board is only slightly larger than the Parallella board (mainly because GPU boards include massive cooling components which Parallella lacks). A machine that requires 56 times more real estate to achieve the same performance is clearly not competitive.
So far we've only looked at floating-point throughput. One must also consider memory capacity and bandwidth of each device. The GTX 780 Ti has 3GB GDDR5 memory, and I understand that Parallella has 1GB memory. I don't know how the memory bandwidth compares between the two.
It's clear that the Parallella board has a long way to go before it could be competitive for supercomputing applications. And as Parallella tries to catch up, the industry will keep moving - GPUs (and other accelerators, like Xeon Phi) we continue to improve along all these dimensions and I imagine that NVIDIA/Intel/AMD have vastly larger R&D budgets than Parallella/Adapteva. So it is difficult to see how this could ever be a viable supercomputing platform and not just an interesting hobbyist board.
If thats true and if it turns out their programming model is better then GPGPU programming then they could disrupt the parallel processing space. However they need to actually ship something, and the more time they take doing that the better Intel gets with MIC and AMD get's with HSA and at that point they would lose any advantage their architecture might have.
But you make a good point, if I were mining bitcoins, I would take the NVIDIA hands down.
There are far better things to do with a Parallela than mining, and far better things that work better for mining. I'd make a miner out of a Parallela system because I am curious about the tech and want practice at writing software for it not because I want a competitive miner.
However, it would be more challenging to put the GPU board on a quadcopter drone than it would be the Parallela.
Also your analysis only considers the "use in cluster" case, there may be cases where the Parallella provides the power and size budget needed for a single case. For example, using this in endpoint devices in control systems, sensor arrays, etc, may make a lot of sense.
Welcome to hardware development. Turnaround times for the software guys are stereotypically measured in minutes, for hardware its measured in days. Just kinda how it is.
I see no point in calling out, but pretty much any time someone who's never shipped hardware tries to ship their first hardware, it always takes two to ten times longer. If these guys are still together doing version 4 or something, those estimates will probably be close to reality. Its just a hardware design pattern to always be almost an order of magnitude more optimistic. They ALL do it, the RF guys, analog guys, digital guys. RF guys are by far the worst because of EMI/EMC and licensing reqs, if it makes you feel any better (LOL).
If they could have gotten the 16-core parallella out the door in 1Q 2013 then we could have hit the 2014 target of 1K-core parallella at 1.4 TFlops which is much more useful
It's just nice to help a group of people with a vision take a shot at moving the needle.
I stumbled across a mail from andreas@adapteva in my spam box the night before last. Just an apology for the delays, looking like a Feb delivery now. You've to send a mail to sales@ with your preference for t-shirt or case.
I just checked the Top500 list from November 1999 (just an arbitrary year to pick) and the lowest on the list hits 38.5 GFlop/s, so if the Parallella gets 90 GFlop/s, it's ell within the range of supercomputers of 14 years ago. In fact, it looks like it would still make the cut in the November 2000 list, but is hitting the edge in the June 2001 and solidly off the November 2001 list.
All, of course, assuming that the 90 GFlop/s quoted is comparable to that measured on the Top500 machines; but even if you toss in a factor of 2 or 4, that still only pushes you a few more years.
Of course, if you stretch that X back a few more years, almost any modern processor is a supercomputer on a chip. I remember when the Cray Y-MP was a supercomputer to dream about, and nowadays its performance would be considered mid-tier for a smartphone.