All Tools

AI Answer Quality Checker

A well-written answer can still be wrong.

Check for vague advice, missing steps, and citation clues. Paste an answer to see what deserves a closer look.

An actual run · a synthetic answer

The calculation is wrong. Why did it score 82?

This answer has numbered steps, asks the reader to check the data, and states a limitation. It looks thorough. But the first step gets relative growth wrong.

“conversion rose from 20% to 30%, a relative increase of 10%, according to [1].”

Subtract the starting rate and you get 10 percentage points. Divide that increase by the original 20%, and relative growth is 50%.

(30% − 20%) ÷ 20% = 50%

The tool gave the full English answer 82/100. It asked for sources but did not identify the arithmetic error. [1] has no corresponding reference, yet counted as citation formatting; steps, verification words, and caveats added points.

That is where this checker helps—and where it misses things. It can find text signals; the calculations and citations still need checking.

Read the full answer and recorded output
1. First record the metric: conversion rose from 20% to 30%, a relative increase of 10%, according to [1].
2. Then check the starting and ending values in the file, for example by keeping the original counts and observation window for the same group.
3. Finally state the limitation: these are synthetic numbers, and other groups may differ; verify the denominator and sampling method, and preserve the calculation and source in the report.

Answer type: research or factual answer · 2026-09-08 · 82/100

  • Some claims need sources — Add sources or soften the wording.
  • Add sources, evidence, method boundaries, and uncertainty. — Add sources, evidence, method boundaries, and uncertainty.

This run uses the English original. The Chinese wording scores 73 because the two texts trigger different rules.

Check the numbers and sources before rewriting. Try a percentage exercise →

Check method and run records

Checks run locally in the browser without an AI call. Rules start at 92, subtract 18, 9, or 4 per high, medium, or low flag, and add points for structure, verification words, and caveats, bounded to 0–100. The score is not an accuracy rate.

The tool does not compare the answer with its original question. Brief correct answers may be told to add length; “guarantee” in a negation can trigger a rule. Claim scanning keeps only the first 160 segments longer than 12 characters.

Both demos use synthetic inputs, with outputs recorded from the tool functions. The review notes are AI-assisted and await maintainer review.

Research basis: FL-EVAL-001 currently has synthetic calibration material and no human-review results. It has not validated this tool's scores. Evaluation method and status

Download full inputs and outputs for both demos (JSON)