Thank you for publishing this @williamofai - we work in healthcare/lifesci/energy/critical path industries and this class of computational library for inference/ML work has come up multiple times in those spaces but oddly outside the finserv space the others tend to be extremely under-funded for foundational work like this even if they have cash pouring in for sexy chatbots they can't use (its a madhouse out there).
I've asked @brixen, one of the consultants on our team and a man "somewhat familiar with architectural semantics" to take a look at implementing this atop the tenstorrent interfaces to their RISC-V kit as the static allocation mechanics should permit granular targeting of host, device, and even SRAM in those things to make use of larger models feasible in smaller timeframes (sometimes clinically precise decisions have to come quick while being complex and conventional architectures like ARM/x86 aren't well suited for that presently). We have a BH QuietBox with 4xp150 cards in it to qualify the effort and can provide access through our private cloud edge for R&D if this thread is of interest.
Thanks again, hopefully we can help kick this effort up a notch toward production utilization in more use-cases.
Thank you for publishing this @williamofai - we work in healthcare/lifesci/energy/critical path industries and this class of computational library for inference/ML work has come up multiple times in those spaces but oddly outside the finserv space the others tend to be extremely under-funded for foundational work like this even if they have cash pouring in for sexy chatbots they can't use (its a madhouse out there).
I've asked @brixen, one of the consultants on our team and a man "somewhat familiar with architectural semantics" to take a look at implementing this atop the tenstorrent interfaces to their RISC-V kit as the static allocation mechanics should permit granular targeting of host, device, and even SRAM in those things to make use of larger models feasible in smaller timeframes (sometimes clinically precise decisions have to come quick while being complex and conventional architectures like ARM/x86 aren't well suited for that presently). We have a BH QuietBox with 4xp150 cards in it to qualify the effort and can provide access through our private cloud edge for R&D if this thread is of interest.
Thanks again, hopefully we can help kick this effort up a notch toward production utilization in more use-cases.