Fineuralab

AI Answer Review Calibration Appendix

A research appendix for FL-EVAL-001 with 20 bilingual synthetic fixtures, a local reviewer instrument, and explicit limits on what the pilot can establish.

What these fixtures calibrate

This appendix uses 20 synthetic answers to test whether the review dimensions are clear enough and whether two reviewers reach similar handling decisions. The full instrument remains at the end of the page and opens only when someone chooses to participate in calibration.

Material
20 bilingual fixtures
Reviewers
2 independent modes
Scoring
5 separate dimensions
Data
browser-local only
Current status: no human-review result yet

The built-in answers are synthetic calibration material deliberately written with different failure levels. They are not outputs from ChatGPT, Claude, Gemini, or any other named model, and the page does not expose a preset answer key.

How this becomes a real pilot

  1. Reviewer 1 completes all 20 tasks independently and exports JSON.
  2. Reviewer 2 repeats the review without seeing the first file.
  3. Import both records and inspect disagreement reasons before calculating any aggregate.
Optional research instrument Open the local two-reviewer workbench Use this only to calibrate the rubric or replace the synthetic answers with real blinded material.

FL-EVAL-001 · Calibration instrument ready; no human-review result yet

AI answer two-reviewer workbench

Reviewer 10/20
Reviewer 20/20
Comparable0
Action agreement

Reviewed and updated: August 17, 2026