Regression testing for LLM applications — with the baseline in your repository, not on someone's server.
https://digline.dev ↗// readme
digline
Regression testing for LLM applications — with the baseline in your repository, not on someone’s server.
Your prompt worked on Tuesday. On Thursday it works a little less — not enough to break, enough for a user to notice in two weeks. No ordinary test catches it, because there is no correct output to compare against, only a better or a worse one.
digline gives you an approved reference — the baseline — and on every change
tells you whether you are below it: which case, which check, by how much. The
baseline is a JSON file in your repository, so it goes through code review and
it rolls back with git. No server, no account, no network call you have not
configured yourself.
$ digline compare --suite suite.py --run latest
2 checks got worse compared with the reference. Every case could be judged. No case is suspended. The suite is unchanged from the reference.
how-do-i-return · llm_rubric · Score fell from 1.000000 to 0.700000.
how-do-i-return · contains · Went from passing to failing (1.000000 → 0.000000).
Why digline
Most evaluation tools tell you whether an output is below a threshold. digline also tells you whether it is *worse…