Source docs/quality.md · 1de96aa

Quality judge (optional)

The quality judge is a model of another provider — Codex (a ChatGPT account), or any OpenAI-compatible or Anthropic-compatible API already set up in Settings › Providers — that reads each passage Libris translates, just before the passage is finished, as the last editor before publication. A passage it finds errors in is corrected on the judge’s model and read again. Nothing runs on the Libris machine besides the usual calls: a small VPS is enough.

Without a judge, Libris translates exactly as before.

Why a model, and not a metric

A translation metric was tried first, MetricX-24 (Google), on twenty paragraphs of an English → French light novel as Libris had translated them and as a person corrected them (2026-09-26). It found the gross errors — « J’ai abandonné » for “I abandoned her” (10.2 on 25, against 3.1 corrected), an omitted clause — but preferred the corrected paragraph only 8 times in 20 on the calques and false friends that make a text read as a translation, and it needed PyTorch, several gigabytes of memory and a large processor. A large model asked what an editor looks for sees both kinds of error, through an API an installation already has.

What the judge does

  1. It reads the finished passage (after the translation, the polish, the review and the revision) with the passage’s context: glossary, character sheets, style sheet, neighbouring passages, and the list of mistakes of its language pair. It answers like a review: each point names the paragraph, a category, the faulty words and a correction.
    • Error: what a professional editor would change before publishing — a change of meaning, a dropped or wrong pronoun, object or negation, an omission or an addition, a wrong name or locked term, a grammar or spelling mistake, a tense that breaks the narration, a calque or false friend a native reader notices, a register that breaks the character, text left untranslated.
    • Warning: a correct phrasing a good editor would still make more idiomatic.
  2. A passage with errors is corrected on the judge’s model, with the judge’s points as review points, through the polish’s guards (no paragraph shortened, no new error of the automatic checks), then read again. The correction is kept only if the judge now finds fewer errors; the version is labelled quality_rework in the passage’s history and the decision logged (Autopilot › Decisions, kind quality_rework).
  3. The errors the judge still finds in the text the passage keeps join its review points (marked by: judge); its warnings only feed its correction. The review’s own points stay, unless the text they are about was rewritten since (its revision, the judge’s correction). The judge’s correction removes a review point only when it changes that point’s paragraph. A passage with no point left is finished OK; otherwise it is To check, its quality score is lower, and the autopilot’s arbitration and the final review take its points up like any review point.

At most QUALITY_REWORK_SHARE of the passages a job set out to translate (25 % by default, at least one) are corrected, each once, even across a resumption. A book in fast quality is judged but never corrected. The judge reads the job’s instruction like every review, and is never swapped for the stronger model: it is the provider chosen to judge. The judge’s typography and pitfall rules are those of every reviewer (prompts/quality_judge.txt, then the rules of the language pair).

A judge that does not answer — outage, refused key, invalid answer — never fails a passage: the passage is finished without its reading and the reason is logged (kind quality_judge, action skipped). No fallback provider stands in for the judge.

Choose the judge

  • For the installation: Settings › Autopilot › Quality judge, or QUALITY_JUDGE_PROVIDER (a provider name or id) in .env. Only the installation’s own providers qualify (not a member’s, not a retired one, not a name two providers share): the judge reads every book of the installation, with the key of its provider. A value saved in the interface wins over the environment.
  • For a book: Book › Settings › Quality judge: empty, the installation’s; No judge for this book (a confidential book translated on a local model); or a provider the book’s owner may use. The choice is checked again at each passage: a provider that became unusable gives the installation’s judge.

Choose a model stronger than the one that translates, or at least another one: a model does not see its own mistakes (#180). Codex through a ChatGPT account costs no token; with an API, each passage costs one judging call (the passage twice — source and translation — plus its context, about the size of a review) and, for the passages corrected, a revision and a second reading. The provider’s budget and concurrency apply as for any call; the estimate shown before a translation does not count the judge’s calls yet.

VariableDefaultWhat it does
QUALITY_JUDGE_PROVIDERemptyName or id of the provider that judges. Empty: no judge, unless Settings › Autopilot or the book names one.
QUALITY_REWORK_SHARE0.25 (0–1)Largest share of a job’s passages corrected after the judge. 0: judged, never corrected.

Benchmark

python -m app.bench has the judge read translations, to compare two models, two prompt versions or two versions of Libris on the same text before a release. The calls are logged like any other; nothing is written to the books.

# A file: one JSON object per line, with "source", one field per system and an optional "reference".
# --book names the volume whose language pair and licence the calls belong to.
docker compose exec api python -m app.bench file /data/bench/pairs.jsonl --book <project id> [--judge <provider>]
# A sample of a volume (40 passages spread over the book), and the same passages of another volume.
docker compose exec api python -m app.bench book <project id> --compare <project id> --sample 40

It prints one line per system (or per volume): items, errors, warnings, with errors and chrF. errors and warnings are the judge’s points, with errors the items it would correct, and chrF the character n-gram F-score against the references (as sacreBLEU computes it). --json gives every point.

Semantic fidelity check

The check that compares each new version of a passage with its source (first translation, revision, polish, correction, arbitration) is read by the configured judge, not by the model that translated: two readings by the same model can share one misreading. The provider that read is recorded with each proof (provider_id).

ConfigurationWho reads the fidelity check
The book or the installation names a judgeThat judge
No judge anywhereThe book’s own provider
No judge for this bookThe book’s own provider: nothing else is contacted
The judge does not answer (outage, refused key, invalid answer)The book’s own provider, for that check

Privacy

The judge’s provider receives the passage, its translation and its context, like the book’s own provider. With Codex or an external API, that text leaves the machine for that provider, under its terms.

This applies to the fidelity check too: a book translated on a local model is sent to the installation’s judge as soon as one is configured, even without the autopilot. To keep such a book on the machine, choose No judge for this book in its settings, or name a local provider as its judge.