Grade your chatbot's answers against its own sources, and measure retrieval before you ship a change.
answergrade measures recall@k and MRR for your search and fails a change that makes retrieval worse. A cautious AI judge grades each answer against the sources your bot used, and the report says what to fix.
Buy answergrade for $99One payment of $99, plus sales tax where it applies, and every future update. Delivered as a private GitHub repository you're invited to.
What it does
- Retrieval numbers with a regression gate
- Reports recall@1, recall@k and MRR for any search function, whether it returns text, ids or records. Save a baseline, and a change that drops a metric beyond your threshold exits non-zero.
- Vector coverage you can trust
- Counts only chunks that have a usable vector (present, numeric, non-zero, the right length), so semantic search over an empty or partial cache is caught.
- A cautious AI judge
- Flags only clear contradictions of your sources, invented details and claimed actions. A garbled, cut-off, hedged or self-contradicting reply is marked could not grade, never flagged. The judge is any function from prompt to text; there is no vendor SDK.
- A report that says what to fix
- Asks your bot over HTTP or as a Python function, grades each answer against the sources it returned or cited, and writes a text, JSON or HTML report with a fix card for each problem.
- Question sets from your logs
- Builds a probe set from chat logs, with greetings, thanks, quoted email history and duplicates removed, or generates questions spread across all your sources.
- Citation and gap counters
- Tracks which sources your bot cites and which have gone quiet, and records unanswered questions with emails, phone numbers, amounts and long digit runs scrubbed first.
- Check the judge before you trust it
- judge-check runs ten answers with known grades through the model you chose, and switches handle endpoints that reject max_tokens, temperature or JSON mode.
After you buy
pip install ./answergrade
answergrade eval --chunks answergrade/examples/chunks.jsonl --cases answergrade/examples/cases.jsonl --save baseline.json
answergrade demo & # a stand-in help bot to try it on
answergrade probe --questions answergrade/examples/questions.jsonl --bot http://127.0.0.1:8766/chat --judge stand-in --out report.html
What you get
answergrade 1.0.0: the Python library and the answergrade command, with a stand-in help bot to try it on, with 456 automated tests. Python 3.10 or newer.
A commercial license. Use it and change it in your own projects and your clients' projects. Don't share or resell the source. Read the license.
Every update. New versions land in the same repository; git pull to get them.
How buying works
- Enter your GitHub username. You'll see the account before you pay, so you can check it's yours.
- Pay $99, plus any sales tax shown at checkout, on Stripe's checkout page.
- Accept the invitation GitHub emails you. Signed in as that account, you'll also find it at
github.com/dominares-tools/answergrade/invitations. Invitations expire after 7 days; you can get a new one any time. - Clone and install:
git clone https://github.com/dominares-tools/answergrade.git pip install ./answergrade
Questions
- Do I need GitHub?
- Yes. Access is a read-only invitation to a private repository, sent to the GitHub account you choose.
- Can I get a refund?
- Yes, within 14 days, no questions asked. Access ends when the refund goes through. Refund policy.
- Who takes the payment?
- Stripe. Your purchase is sold through Link, Stripe's checkout service, which also works out sales tax and sends your receipt.
- Something wrong?
- Email josephlinares02@gmail.com. Every buyer of a tool shares its repository, so support happens by email rather than in GitHub issues.