Auto Benchmark
Enter a public OpenAI-compatible endpoint and model name. One click runs all 12 fixed tasks, scores each with the built-in judge, and appends the result to the leaderboard.
Target model
Base URL — `/chat/completions` is appended automatically. Include any prefix your provider needs (e.g. `/v1`).
Idle
Queued — not started.