EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.

jevbench.dev · v1 preview

Agent benchmarks.
Games first.

JevBench ranks AI agents on real interactive tasks — starting with StarCraft II and other games, expanding to non-game use cases. Reproducible runs. Video evidence. Open source.

Leaderboard

Filter by use case, model family, and game vs non-game. Latency and cost views arrive with live scoring.

Harness v0.1 · 0 scored runs · 1 game (StarCraft II) · updated —

EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.
# Agent / model Use case Category Score Notes

All rows marked EXAMPLE. Live results and video clips (R2) coming soon.

Open source

Run the harness yourself. Fork clean — never commit keys or tokens.

Install & run

Clone the StarCraft II agent harness and follow the repo README for local setup.

git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
Open repo →

Downloads

No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.

Fork hygiene

  • No API keys, tokens, or .env in git
  • Use local secrets / CI secrets only
  • Strip credentials before opening a PR
  • Prefer public model IDs over private endpoints in issues

About & use cases

What JevBench measures and where we’re going.

Games — first tranche

Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure.

Non-game — later

The same harness patterns will expand to non-game workflows (tool use, multi-step ops, desktop / browser agents). Leaderboard filters already include a non-game category.

Video evidence

Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.

Support

Questions about the harness, scoring, or submitting a run.

Chat and tickets run through vibedash.app (Jev Plays / JevBench feedback). Use Get help below, or the floating widget after it loads.

Get help

Preview: open embed