The real problem is getting 1 exaflop (or around it) within a reasonable power budget. The DOE's power budget for all of their supercomputing resources is 20 Megawatts, so at a full system level we would need to be at 50 GFLOPs per watt, while the best system right now is at 5.
That's single precision, which the DOE doesn't care about. The latest NVIDIA GPU's do 8 to 16 times better at single precision (32 bit) floating point compared to double precision (64 bit).
The best next generation DP GFLOPs/watt from one of the big players will most likely be the 2016 Xeon Phi, at ~10-12GFLOPs/watt... You are also forgetting that GPUs also have a ~100W+ CPU sitting next to it, which brings down total efficiency significantly.
Shameless self promotion: My startup (http://rexcomputing.com) is aiming for 64 double precision GFLOPs/watt, and 128 GFLOPs/watts single precision for its first chip next year.
Your chip looks cool. I guess it may be tricky to adapt software to run on the thing? Or else you could try to sell Obama 4 million of them for his new computer.
A huge amount of overall system power is spent in data transport. Plus, double everything for cooling. That brings the total system efficiency way down from what the actual computational components spend.
> "I'd say they're targeting around 60 megawatts, I can't imagine they'll get below that," [Mark Parsons] commented. "That's at least £60m a year just on your electricity bill."