Fineuralab
Evaluating AI Answers as an Experiment
A small, repeatable evaluation design for comparing AI answers without rewarding fluency alone.
Research question
How can two AI answers be compared when correctness, evidence, actionability, and risk all matter?
Preference often follows tone and polish. An experimental comparison needs a fixed task, hidden answer identity, explicit dimensions, and recorded disagreements so that confidence does not become the main metric.
Fix the task contract
Write the user goal, supplied context, constraints, allowed assumptions, expected output, and unacceptable failures. Use the identical contract for every answer being compared.
Score independent dimensions
Use separate scores for requirement coverage, factual support, reasoning trace, actionability, calibration, privacy or safety risk, and unnecessary verbosity.
- Do not combine all dimensions into one intuition first.
- Require evidence for factual claims.
- Penalize invented constraints and omitted caveats.
Blind the surface when possible
Remove model names and normalize formatting before review. Keep original answers archived. Blinding does not remove all bias, but it reduces brand and presentation effects.
Record disagreements and abstentions
A reviewer should be able to mark insufficient evidence, not applicable, or unresolved. Disagreement identifies rubric ambiguity or a task that needs domain expertise.
- Keep reviewer notes next to each score.
- Resolve criteria before averaging scores.
- Escalate high-stakes claims to a qualified reviewer.
Choose an action, not a winner
The outcome may be accept, revise, verify, merge, or reject. The most fluent answer is not always the best artifact, and neither answer may be safe to use without further evidence.
Calibration appendix
Twenty bilingual synthetic fixtures test whether the rubric is clear enough to use. This supports the method; it is not a named-model leaderboard and contains no human-review result yet.
View the calibration appendixThis page records a current Fineuralab method, protocol, or maintainer judgment. Unless its status explicitly says reproduced or experiment complete, it should not be read as a reproduction claim for a specific paper.
Reviewed and updated: August 17, 2026