Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Summary

Note

🧪 IN-PROGRESS EXPERIMENT

Warning

This write-up is an active work in progress - as the project itself is still under development. Some benchmarks and stats are changing as I continue to optimise the engine and benchmark.

Link to Snibble-Bench

Snibble-bench was pivoted from Snibble the game, when out of curiousity I tested OpenAI Codex 5.5 playing the game - and was surprised to find that it was able to play the game successfully (a surprise), but performed poorly against the games bots.

About

Can a word game become an LLM benchmark - testing spatial lexical and adversarial reasoning? Yes it can

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors