Update main readme to the point we can start the inference benchmark with ./nki-llama inference benchmark - #2
Update main readme to the point we can start the inference benchmark with ./nki-llama inference benchmark#2crizCraig wants to merge 4 commits into
./nki-llama inference benchmark#2Conversation
…with `./nki-llama inference benchmark`
./nki-llama inference benchmark
|
@arm-diaz Hello from AGI House! We are preparing for a trainium competition here and I was running into the following on a I've reduced the resources in my .env to the following in case it was a memory error Would very much appreciate any help. Best, |
|
I asked the above as there are no issues on this repo. However, this PR could also be merged. |
|
Hi Craig, this project only supports higher TP degrees at the moment. It is not intended to work with TP=2 given the current setup. |
|
@EmilyWebber Hi Emily, One logistical wrinkle we’re seeing: a single trn1.32xlarge requires the user’s account to have a 128-vCPU quota. Submitting a Service Quotas request through the standard support channel often takes several days to approve. That delay is manageable for long-running online entrants, but it could be a real hurdle for new participants and especially for our one-day in-person hackathons, where teams need to spin up a node immediately. Would love your thoughts on how we might streamline this (pre-approved event accounts, a smaller default instance, or an expedited quota path). Let me know what you think. cc @crizCraig |
No description provided.