Bench

Auto Benchmark

Enter a public OpenAI-compatible endpoint and model name. One click runs all 12 fixed tasks, scores each with the built-in judge, and appends the result to the leaderboard.

Target model

Base URL — `/chat/completions` is appended automatically. Include any prefix your provider needs (e.g. `/v1`).

Idle

Queued — not started.