Design experiments

Each card is one experiment — set its system message, prompts and rubric side by side.

Experiment 1

Run

Results

Live progress, summary metrics, plots and transcripts land here.

Add at least one API key to run analyses.

Results will appear here after you run an analysis.

Choose data

Pick a bundled dataset or upload your own, then optionally narrow it to a subset.

Dataset

Upload your own dataset or choose one from the built-in library.

How do I bundle images or audio?

Zip one dataset file (CSV/XLSX/JSONL) together with the media it references. Add an Image or Audio column holding paths relative to the dataset file; leave the cell empty for text-only questions.

my_benchmark.zip
├── questions.csv      (columns: Question, Answer, Image)
└── images/
    ├── q1.png
    └── q2.png

Download a working example zip — or pick the Demo Media entry from the bundled dataset list to try it without uploading anything.

Question Subset

Optionally run the benchmark on a subset using printer-style ranges (e.g. 1-5,8,12).

Leave blank for all questions. Use commas to combine individual numbers and ranges.

All 0 questions will be used.

Design experiments

Prompts can reference {{question}} and {{answer}} placeholders.

Experiment 1

Run

Results

Benchmark progress, accuracy summary and transcripts land here.

Benchmark not started yet.