@ysymyth @john-b-yang
Running the official repo with a variety of models including gpt-4, gpt-4o, gpt-4o-mini, gpt-3.5-turbo, etc. results in terrible performance. These days ranging from 0-10% success rate using ReAct.
Could you be so kind to re-run official results and post an updated version of results so that we can use these as "official" benchmark. (For example reviewers are often complaining that new results are much worse than original results). However, we tried reproducing result using your implementation, other implementations and our own implementation always yielding very low scores.
Thank you very much.
Concrete Ask:
If running the whole thing is too much, which we understand. Could you for example maybe just run the results on the first 30-50 examples and report the scores there for a few models (e.g. gpt-3.5-turbo, gpt-4o and gpt-4o-mini).
Thank you very much.
@ysymyth @john-b-yang
Running the official repo with a variety of models including gpt-4, gpt-4o, gpt-4o-mini, gpt-3.5-turbo, etc. results in terrible performance. These days ranging from 0-10% success rate using ReAct.
Could you be so kind to re-run official results and post an updated version of results so that we can use these as "official" benchmark. (For example reviewers are often complaining that new results are much worse than original results). However, we tried reproducing result using your implementation, other implementations and our own implementation always yielding very low scores.
Thank you very much.
Concrete Ask:
If running the whole thing is too much, which we understand. Could you for example maybe just run the results on the first 30-50 examples and report the scores there for a few models (e.g. gpt-3.5-turbo, gpt-4o and gpt-4o-mini).
Thank you very much.