Skip to content

Reproducing Results on Webshop #34

Description

@ai-nikolai

@ysymyth @john-b-yang

Running the official repo with a variety of models including gpt-4, gpt-4o, gpt-4o-mini, gpt-3.5-turbo, etc. results in terrible performance. These days ranging from 0-10% success rate using ReAct.

Could you be so kind to re-run official results and post an updated version of results so that we can use these as "official" benchmark. (For example reviewers are often complaining that new results are much worse than original results). However, we tried reproducing result using your implementation, other implementations and our own implementation always yielding very low scores.

Thank you very much.

Concrete Ask:
If running the whole thing is too much, which we understand. Could you for example maybe just run the results on the first 30-50 examples and report the scores there for a few models (e.g. gpt-3.5-turbo, gpt-4o and gpt-4o-mini).

Thank you very much.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions