TODO Move native libs out of here :) Documentation as to how to set up CUDA and native libs Find examples that actually benefit from the GPU performance-wise (this vector mult actually runs slower on GPU than on the CPU)