Our Explorer tool is a simple interface that lets you test models against each other on your own prompts, to see what performs best. We previously supported Best-of-N and LLM-as-judge strategies. Today we are launching a 4 stage Chain of Thought option. Try it out and send us feedback on any other functionality you need to effectively decide on the right models to use for your use case.
Discussion about this post
No posts



Interesting framing around "Our Explorer tool is a simple interface that lets you test models against each other on your own prompts, to see what pe". I wonder how this holds up when you scale past a single-agent setup though. The coordination overhead can change the calculus quite a bit.