Drop yesterday's eval run and today's.

Backslide joins the two files, scores only the things that can honestly be called worse, and hands you the rows that broke — with the reason attached. It reads promptfoo, OpenAI Batch and Evals exports, raw completion dumps, JSONL and CSV.

Run A — before

Drop a file

or

Run B — after

Drop a file

or

What this cannot tell you