Source docs/development.md · 1de96aa

Development and releases

This page is for contributors and maintainers. It explains how to set up a development environment, run every kind of test, how the CI/CD pipeline works, how a release is cut and deployed, and how translation quality is evaluated. To run Libris rather than work on it, read the Docker guide.

Repository layout

DirectoryContents
backend/FastAPI application (app/), persistent worker (app/jobs/worker.py), Alembic migrations, pytest suite
frontend/React + TypeScript interface (src/), Playwright specs (e2e/)
prompts/Versioned model instructions, copied into the image
codex_bridge/Optional isolated Codex transport and its own image
scripts/Installation, deployment, test and release utilities
deploy/Production deployment, backup, restore and CI-runner maintenance procedures
docs/Documentation; docs/screenshots/ and docs/openapi/ are generated (see Screenshots and The OpenAPI description)
examples/Example client of the automation API (Python, standard library only)

How the pieces fit together is described in architecture.md; the interface conventions are in design-system.md.

Set up

Use Python 3.13 and Node.js 22, the versions of the images and of CI.

python3 -m venv .venv
.venv/bin/pip install -r backend/requirements.lock
.venv/bin/pip install -e './backend[test]'
npm --prefix frontend ci

backend/requirements.lock pins every Python dependency with its hashes. After changing a pin, regenerate the hashes with python3 scripts/hash_lock.py backend/requirements.lock (network needed); CI runs scripts/hash_lock.py --check and refuses an unhashed pin. Add a dependency only when it is needed, say why in the merge request, and update the matching lockfile.

Run the application locally

The simplest loop runs the real stack with Docker Compose and the interface with Vite’s development server.

  1. Create a local configuration once: python3 scripts/setup.py writes a .env with generated secrets and never overwrites an existing one.

  2. Add the Vite origin to ALLOWED_ORIGINS, for example ALLOWED_ORIGINS=http://127.0.0.1:8088,http://127.0.0.1:5173. COOKIE_SECURE=true works on localhost and 127.0.0.1; set it to false only if you open the development server through a network address.

  3. Build and start the stack from your working tree, under a local image name so that the published image is never overwritten:

    LIBRIS_IMAGE=libris:dev docker compose up -d --build
    
  4. Start the interface with hot reload. Vite serves it on http://127.0.0.1:5173 and forwards /api and /health to the API on port 8088:

    npm --prefix frontend run dev
    

Sign in with BOOTSTRAP_USERNAME and BOOTSTRAP_PASSWORD from .env. Every setting is described in configuration.md.

Tests

Backend

cd backend
../.venv/bin/ruff check app tests migrations
../.venv/bin/pytest -q

Ruff also covers the Alembic migrations (migrations): they run in production at every upgrade, so the backend CI job lints them with the same command.

The suite needs no service: it uses a temporary SQLite database and data directory, and mocked providers (tests/mock_server.py, respx). It also checks the migrations: tests/test_migrations.py upgrades, downgrades and upgrades the schema again.

pytest-xdist (in the test extra) spreads the tests over several processes; -n auto starts one per CPU:

../.venv/bin/pytest -q -n auto

Each process has its own temporary data directory and database. The schema is built once per process and every test starts on empty tables (tests/conftest.py); a test that changes the schema through the engine (CREATE, ALTER, DROP…) has the next one start on a freshly built schema. Fixtures that write users straight into the database take their password hash from tests/passwords.py, computed once per session. GitLab runs the SQLite and PostgreSQL suites with -n 6 on its 16-core runner, whose cores are also shared by the other jobs in the pipeline. Local runs can use -n auto; there is no second CI whose result could compensate for a failed GitLab pipeline.

Production runs PostgreSQL, so changes to queries, locking, concurrency or migrations must also pass on PostgreSQL 17. Point the suite at a throwaway database with LIBRIS_TEST_DATABASE_URL:

docker run -d --name libris-test-db -p 127.0.0.1:5432:5432 --tmpfs /var/lib/postgresql/data \
  -e POSTGRES_DB=libris_test -e POSTGRES_USER=libris -e POSTGRES_PASSWORD=libris_test_password postgres:17-bookworm
cd backend
export DATABASE_URL=postgresql+psycopg://libris:libris_test_password@127.0.0.1:5432/libris_test
export SECRET_KEY=local-test-secret-key-with-more-than-32-characters
export BOOTSTRAP_PASSWORD=local-test-password-123456
../.venv/bin/alembic upgrade head && ../.venv/bin/alembic downgrade base \
  && ../.venv/bin/alembic upgrade head && ../.venv/bin/alembic check
LIBRIS_TEST_DATABASE_URL="$DATABASE_URL" ../.venv/bin/pytest -q -n auto
docker rm -f libris-test-db

A serial run uses the database of LIBRIS_TEST_DATABASE_URL. Under -n, each worker creates its own database next to it (libris_test_gw0, libris_test_gw1…, dropped and created again at start, dropped at exit), so the user of the URL must be allowed to create databases, and two parallel runs must not share the same URL.

Every migration must stay reversible down to an empty schema, and alembic check must find no difference between the models and the migrations.

Tests marked epubcheck validate exports with the real EPUBCheck. They are skipped unless EPUBCHECK_JAR points to an EPUBCheck JAR and java is on the PATH; LIBRIS_REQUIRE_EPUBCHECK=1 turns a missing validator into a failure. CI runs them inside the built image, which ships EPUBCheck 5.3.0.

Tests and the clock

A test that reads the real time (time.time(), datetime.now()) and adds hours, days or months can pass all day and fail in the evening, at the end of a month or on 31 December (#167). Compute such instants from a fixed date (see tests/test_clock_boundaries.py and DAY in tests/test_work_window.py), or pass now explicitly to the function under test; never read “today” twice in one test.

scripts/clock_check.sh replays the suite with the clock moved by libfaketime (Debian package faketime; the clock keeps running from the given instant, the monotonic clock is untouched). One instant, in UTC, with an optional time zone, or every boundary near today at once:

cd backend
TZ=Pacific/Kiritimati PYTHON=../.venv/bin/python ../scripts/clock_check.sh "2026-12-31 23:30:00" -q -n auto
PYTHON=../.venv/bin/python ../scripts/clock_check.sh --boundaries -q -n auto

--boundaries runs 23:30 UTC today, 00:30 UTC tomorrow, 23:30 UTC on the last day of the month and on 31 December, the European changes of time (00:30 UTC on the last Sunday of March and of October, in Europe/Paris), and two time zones on either side of the date line (Pacific/Kiritimati, UTC+14, and Pacific/Pago_Pago, UTC−11). libfaketime does not move the times the kernel gives a file it writes, so the few tests that compare them with the clock carry the file_times marker and are skipped under the script (FAKETIME set); they run in the ordinary suite. The libfaketime of Debian 13 (trixie) is needed: the one of bookworm makes time.sleep() fail with EINVAL. Without the package, point LIBFAKETIME at libfaketimeMT.so.1, for example extracted with apt-get download libfaketime && dpkg -x libfaketime_*.deb /tmp/faketime.

The OpenAPI description

docs/openapi/libris-v1.json describes the automation API (/api/v1) and is generated from the code by backend/app/api/v1_openapi.py. After adding or changing a /api/v1 route, regenerate it and commit it:

.venv/bin/python scripts/export_openapi.py          # rewrite the file
.venv/bin/python scripts/export_openapi.py --check  # only compare (prints a diff)

tests/test_openapi_v1.py fails while the committed file differs from the code, checks that every /api/v1 operation is described with its scope and error responses, and compares the documented answers with real ones. A new route needs nothing else to appear: its tag comes from its path, its summary from the first line of its docstring, its scope from its require(...) dependency, and it gets the shared error responses. Describe what FastAPI cannot see (a body read by hand, the fields of a dict answer, specific error codes) in v1_openapi.py: OPERATIONS and SCHEMAS. tests/test_example_client.py runs examples/libris_client.py against the API: when a route, an upload option or a result format changes, update the client and its test in the same merge request.

Frontend

npm --prefix frontend run build

The build runs the strict TypeScript check before Vite. Visible text goes through the translation helpers described in design-system.md; a new string needs its English entry.

Browser tests (Playwright)

Specs are selected by a tag in their title:

TagNeedsWhere it runs
nonenothing: the spec mocks the API with page.routeCI frontend job, locally against vite preview
@journeya disposable stack with the synthetic model of docker-compose.test.ymlCI e2e job
@integrationa disposable stack holding the data created by scripts/smoke.pyby hand only

Without LIBRIS_E2E_URL, playwright.config.ts leaves out @integration and @journey, so a plain run is always safe on a workstation:

npm --prefix frontend run build
npm --prefix frontend exec playwright install chromium
npm --prefix frontend exec vite preview -- --host 127.0.0.1 --port 4173   # terminal 1
npm --prefix frontend run test:e2e                                         # terminal 2

The mocked specs open http://127.0.0.1:4173 unless SHOWCASE_URL says otherwise.

To run @journey or @integration specs, point them at a disposable installation on a loopback address. They never read the repository .env:

VariableMeaning
LIBRIS_E2E_URLBase URL of the disposable installation, on 127.0.0.1, localhost or ::1
LIBRIS_E2E_USERNAME, LIBRIS_E2E_PASSWORDIts administrator account
LIBRIS_E2E_CONFIRM_DISPOSABLEMust be 1: confirms that the target holds no real data
LIBRIS_E2E_MOCK_LLM_URL@journey only: the synthetic model as the backend sees it (default http://mock-llm:8091)
LIBRIS_E2E_PROJECT_ID or LIBRIS_E2E_STATEWorkspace specs: the smoke project, or the state file written by scripts/smoke.py (default /tmp/libris/epub-smoke.json)

A disposable stack needs a licence of its own

Libris refuses the work that costs words without a valid licence, and the key certificates are checked against is a constant of backend/app/licence/certificate.py — not a setting, so no disposable stack can simply be told to trust another server. A build may replace it: docker build --build-arg LICENCE_KEY=… writes another public key into the image, and that image trusts a licence server holding the private half. A released image is built without the argument and carries the release key.

The CI e2e job does exactly this: it draws an Ed25519 pair with openssl, builds the image on the public half, starts mock-licence (backend/tests/mock_licence.py, in docker-compose.test.yml) with the private one, and activates the installation through the API before the specs run. LICENCE_SERVER_URL points the stack at that container. To do the same by hand:

openssl genpkey -algorithm ed25519 -out /tmp/licence.pem
raw() { openssl pkey -in /tmp/licence.pem "$@" -outform DER | tail -c 32 | basenc --base64url | tr -d '=\n'; }
docker build --build-arg LICENCE_KEY="$(raw -pubout)" --tag libris-e2e:local .
# MOCK_LICENCE_SIGNING_KEY="$(raw)" and LICENCE_SERVER_URL=http://mock-licence:8093 in the env file,
# then activate with MOCK_LICENCE_KEY (default LIB-TEST-TEST-TEST-TEST) in Settings > Licence.

The backend suite needs none of this: tests/conftest.py gives every test a licensed installation, signed with a key drawn once per process. A test that wants an installation without one empties app.licence.store.

Full-stack smoke test

scripts/smoke.py drives a real Compose stack with the synthetic model: it imports a generated EPUB, analyses and translates it over HTTP, pauses and resumes, kills the worker with SIGKILL and checks recovery, then exports an EPUB and validates it with EPUBCheck. It reads the local account without printing secrets and refuses to run without --confirm-disposable. The commands below use the repository .env (its LIBRIS_IMAGE, PORT and account, also read by smoke.py, or another file with --env-file): run them from a development checkout whose .env names a local image (LIBRIS_IMAGE=libris:dev, as in Run the application locally) and a free port, never next to a real installation.

docker compose -p libris-smoke -f docker-compose.yml -f docker-compose.test.yml --profile test up -d --build
.venv/bin/python scripts/smoke.py --compose-project libris-smoke --confirm-disposable \
  --compose-file docker-compose.yml --compose-file docker-compose.test.yml
# ... run @integration specs against it if needed, then remove the test project:
.venv/bin/python scripts/smoke.py --compose-project libris-smoke --confirm-disposable --cleanup \
  --compose-file docker-compose.yml --compose-file docker-compose.test.yml
docker compose -p libris-smoke -f docker-compose.yml -f docker-compose.test.yml --profile test down --volumes

Related scripts, all for disposable installations only:

ScriptWhat it checks
scripts/check_installation.pyA fresh Compose installation on local port 4188 with a generated configuration and a random project name: login, health, migrations; then removes that stack and its volumes
scripts/resilience_smoke.pyProvider outage, manual pause, SIGTERM and retry on a synthetic project
scripts/make_browser_fixtures.pyWrites three small EPUBs under /tmp/libris/ for browser tests
scripts/cleanup_browser_fixtures.pyRemoves the known browser-test projects left by an interrupted run
scripts/check_epubcheck.pyRuns the bundled EPUBCheck on valid and invalid synthetic EPUBs (inside the image)

Never point any of these at a production library.

The resilience smoke has the same explicit boundary as the full smoke and additionally proves that a running mock-llm belongs to the requested Compose project before it creates a volume or stops its worker:

.venv/bin/python scripts/resilience_smoke.py --compose-project libris-smoke --confirm-disposable \
  --env-file .env --compose-file docker-compose.yml --compose-file docker-compose.test.yml

Screenshots

The images in docs/screenshots/ come from frontend/e2e/showcase.spec.ts, which renders the real interface against mocked API answers and fictional books. It uses no credentials and no model, checks the colour contrast of both themes, and writes to /tmp/libris/showcase unless told otherwise. To refresh the documentation images:

cd frontend
npx vite build
CI=1 SHOWCASE_SCREENSHOT_DIR=../docs/screenshots npx playwright test e2e/showcase.spec.ts

e2e/documentation.spec.ts captures the same views from a disposable API-backed installation instead; it only runs with LIBRIS_DOCS_CAPTURE=1 and LIBRIS_DOCS_CAPTURE_CONFIRM=disposable. Look at every image before committing it.

CI/CD pipeline

GitLab is the only repository and runs the release pipeline. Libris is proprietary and its source is not published.

What runs when

PipelineJobs
Merge request, branchbackend (Ruff + pytest on SQLite), backend-postgres (migration round trip + pytest on PostgreSQL 17), frontend (build, npm audit, mocked Playwright specs), e2e (@journey specs against the Compose stack), audit (version consistency, hashed lockfiles, pip-audit, Gitleaks)
Default branchthe same, then container-build, container-runtime, container-epubcheck, container-scan and, once everything passed, verified-image
Release tag vX.Y.Zrelease-policy, release-images, container-runtime, container-scan, publish-gitlab, release-gitlab, and the manual esoleau-bundle. Nothing is deployed anywhere
Schedule with LIBRIS_CLOCK_CHECKclock-check only: the backend suite at every clock boundary (see below)

A merge request that changes nothing under backend/, frontend/, codex_bridge/, prompts/, scripts/, deploy/, examples/, docs/openapi/, the Dockerfile, .dockerignore, the Compose files or .gitlab-ci.yml only runs audit, which still checks the version pins in the documentation.

Nightly clock check

clock-check runs scripts/clock_check.sh --boundaries in python:3.13-trixie with the faketime package (the libfaketime of bookworm, the base of $PYTHON_IMAGE, makes Python’s time.sleep() fail with EINVAL), and nothing else: it never runs on a merge request or a push, and a scheduled pipeline that sets LIBRIS_CLOCK_CHECK starts none of the other jobs. A maintainer creates the schedule once, in Build → Pipeline schedules → New schedule:

  • description: Clock check (#167);
  • interval pattern: 30 23 * * *, cron time zone UTC (the real time is then itself on a boundary);
  • target branch: main;
  • variable: LIBRIS_CLOCK_CHECK = 1;
  • activated.

Run in the schedule list starts it at once. A failure names the instants and time zones that failed; replay one locally with the single-instant form of the script, then fix the test (or the code) with a fixed date.

Images and verification

Only the protected default branch builds and pushes images, addressed by commit: sha-<commit> for the application and codex-sha-<commit> for the Codex bridge. Every later job uses those exact digests:

  • container-runtime starts the image and checks it;
  • container-epubcheck runs the epubcheck tests inside the image;
  • container-scan runs Trivy once per image: HIGH and CRITICAL findings with a fix fail the job, and the report is kept as a CycloneDX SBOM artifact;
  • verified-image adds verified-sha-<commit> when the whole default-branch pipeline passed.

A tag pipeline does not test again. release-policy checks that the tag is on the default branch, that it matches the version (scripts/check_version.py) and that CHANGELOG.md has its section. release-images waits up to 20 minutes for the verified-sha-<commit> marker, since a tag is often pushed while the branch pipeline is still running, then promotes those digests. If the branch pipeline failed, make it pass and retry release-images.

A tag vX.Y.Z publishes:

DestinationTagsPublished by
GitLab Container RegistryX.Y.Z, X.Y, latest (and codex- variants)publish-gitlab

Nowhere else: Libris is proprietary, its images are not published in any public registry and there is no public source repository. Customers reach that same GitLab registry at registry.libris-translate.com (registry.libris-translate.com/libris/libris:X.Y.Z, :latest, :codex-X.Y.Z), which authenticates every pull: each licence receives its own read-only registry credentials from the licence server, in the mail that carries its key, and loses them when the licence stops being valid. Nothing else distributes Libris: no archive, no download page, no deployment by the pipeline.

release-gitlab creates the GitLab release; its notes are exactly the CHANGELOG.md section of that version (python3 scripts/release_notes.py vX.Y.Z prints them). A retried job updates the existing release.

CI runner hygiene

All jobs run on a shell runner. Test, build and scan tools start with docker run; only small release steps (release-policy, release-gitlab) use the host’s git and python3. Every container is named libris-ci-<role>-$CI_JOB_ID, carries the label libris-ci-job=$CI_JOB_ID and runs under --init; every command that can hang is wrapped in timeout, and every job has its own timeout:. Cancelling a job only kills the Docker client, so each job’s after_script (which also runs on cancellation) removes the containers with its label, and e2e takes its Compose project libris-e2e-$CI_JOB_ID down with its volumes. As a last resort, audit removes any CI container or libris-e2e-* stack older than two hours. To look by hand: docker ps --all --filter label=libris-ci-job.

Volumes can still be left behind by a killed job. deploy/libris-runner-prune removes, on the runner host, unused anonymous volumes and the volumes and networks of libris-e2e-* projects older than LIBRIS_PRUNE_MIN_AGE_HOURS (default 6); it never touches named volumes of other projects, images or the build cache. Install it as root on the runner host:

install -m 0755 deploy/libris-runner-prune /usr/local/sbin/libris-runner-prune
install -m 0644 deploy/libris-runner-prune.service deploy/libris-runner-prune.timer /etc/systemd/system/
libris-runner-prune --dry-run          # see what it would remove
systemctl daemon-reload && systemctl enable --now libris-runner-prune.timer

The timer runs hourly; journalctl -u libris-runner-prune shows what was removed.

Required GitLab settings

  • Protect main and tags matching v*.
  • Enable the Container Registry and protect the sha-*, codex-sha-*, verified-sha-* and exact-version tags from being overwritten.
  • Keep no push mirror: Libris is not published anywhere public.
  • The predefined CI_REGISTRY* variables authenticate the GitLab registry, which is the only destination.

Never commit a password, token, .env, registry credential or user book.

Prepare dependency upgrades on GitLab. pip-audit, npm audit, the hashed Python locks and the Trivy image scan are the release gates; advisories observed elsewhere are inputs to an issue, never evidence that GitLab has already tested a fix.

Releasing

Libris follows Semantic Versioning. While in 0.x, any operationally breaking change is called out in CHANGELOG.md and therefore in the release notes. A tag publishes, it deploys nothing: the tag pipeline promotes the verified images to X.Y.Z, X.Y and latest and creates the GitLab release, and every installation that follows latest takes the new version the next time its administrator runs the installer. Agree on the release with the maintainer before tagging.

Prepare

  1. Set the new version everywhere scripts/check_version.py looks: backend/pyproject.toml, backend/app/__init__.py, frontend/package.json and its lockfile, codex_bridge/package.json and its lockfile, codex_bridge/rpc.py, the LIBRIS_VERSION argument of both Dockerfiles, scripts/install-docker.sh, and the LIBRIS_TAG=X.Y.Z pins of README.md, docs/docker.md and docs/docker.fr.md. Leave LIBRIS_IMAGE and LIBRIS_CODEX_IMAGE in .env.example on latest and codex-latest: the script requires it, and the installer replaces them with digests.
  2. In CHANGELOG.md, rename ## [Unreleased] to ## [X.Y.Z] - YYYY-MM-DD, and do the same in CHANGELOG.fr.md (## [Non publié], a complete translation): backend/tests/test_documentation.py fails when a release of the English changelog has no French section. Check the notes with python3 scripts/release_notes.py vX.Y.Z.
  3. When backend/app/licence/ changed since the last tag (git diff vPREVIOUS..HEAD -- backend/app/licence/), check that the licence server in production accepts what the new version sends and answers what it expects: its version is on https://sub.libris-translate.com/health. A Libris that needs a newer licence server is released after that server is deployed, never before.
  4. Run python3 scripts/check_version.py, then the backend, PostgreSQL, frontend and installation checks above.
  5. Merge to main and wait until the default-branch pipeline is green (the verified-sha-<commit> marker).

Tag

VERSION=X.Y.Z
python3 scripts/check_version.py "v${VERSION}"
git tag -s "v${VERSION}" -m "Libris ${VERSION}"
git push origin "v${VERSION}"

Use an annotated unsigned tag (git tag -a) only when signing is not configured. Never move or reuse a release tag.

Verify

A green pipeline is not proof of a release. Check the two destinations on their own: the GitLab registry and the GitLab release. Pull a published digest from registry.libris-translate.com with customer credentials, inspect its OCI version and revision labels, and query /health from a disposable stack installed with LIBRIS_TAG=X.Y.Z before announcing it.

Publish and announce

The release reaches customers through the website, libris-translate.com, which is published from its own repository once the images are out:

  1. The website regenerates its technical documentation from the tag (these docs/, the changelogs, the README and the OpenAPI description, as they are at vX.Y.Z): a correction made here after a tag is public only with the next release.
  2. It raises its version, which it publishes at https://libris-translate.com/version.json and in the pinned installer commands.
  3. The licence server reads that file every hour and announces the version to every installation in the latest_version field of its answers; administrators then see the update banner (see docker.md). Merging the website is therefore what announces a release.

The installer customers run, https://libris-translate.com/install.sh, is deploy/install.sh. Whatever copy starts, it hands over to the installer carried by the image it installs (/app/deploy/install.sh), so a change to the installer ships with the image; the website’s copy must still be refreshed from the tag, since it is the one a new customer starts with.

Deployment script for a host of your own

The pipeline deploys nothing. deploy/libris-production-deploy and deploy/librisctl remain in the repository for an operator who drives one host from a pipeline of their own; they are described in operations. A customer installation is updated with the installer (docker.md); a checkout, with scripts/deploy.sh (operations).

When a deployed release misbehaves and it added no migration, libris-production-deploy --rollback (or librisctl rollback --confirm) redeploys the retained previous-* images through the same guarded procedure (dump, stop, migrate as a no-op, start, /health, worker restart). It refuses, before touching anything, when the release added a migration: reverting an image does not revert a schema. In that case:

  1. Stop the application services: docker compose --project-name <project> --env-file <.env> --file <compose file> --profile codex stop api worker codex.
  2. Identify the dump taken before the release being reverted, not merely the newest dump. Restore it with librisctl restore <dump> --confirm --no-start, or the empty-database procedure in backup.md. The extra flag is essential: restarting the current image would immediately reapply its migrations. pg_restore --clean over the newer schema is not sufficient; tables added since the dump can block restoration through their foreign keys.
  3. Run libris-production-deploy --rollback again: the schema now matches.

A deployment never changes the books in /data or the .env; restore them from the scheduled backup only if they were damaged.

Restoration discards database changes made after the chosen dump, including newly issued review links. The CLI keeps a separate pre-restore-*.dump of the replaced state, with mode 0600 and no automatic pruning. Do not delete that rescue copy until the recovery is verified.

Before handing a build to anyone

  1. Review the whole Git history, not only the working tree, for secrets, private hostnames, books and test artifacts. Rotate any secret that was ever committed.
  2. Confirm the rights to the logo and every included asset.
  3. Never transfer a deployment .env, volumes or backups.
  4. Do not announce images or releases before they exist.
  5. Libris is proprietary (see LICENSE): the source is not published, and an image is handed over with the registry credentials that go with a licence. Nothing in the repository may say otherwise — scripts/check_version.py fails the build if a file declares another licence.

Evaluating translation quality

Passing tests, EPUBCheck and automatic scores show that the output is well formed, not that the translation is faithful. Before claiming a quality level for a language pair or a kind of book, run a documented evaluation.

Corpus

Use a text you are allowed to use, several dozen chapters long, with a reference reviewed by a bilingual reader. It must contain distant callbacks and variations in the source; repeating the same paragraph does not test narrative memory. Include at least:

  1. An invented name with typographic variants: no unjustified change of translation.
  2. A character identified late: ambiguous pronouns stay ambiguous before the reveal.
  3. An object given in chapter 1, reinterpreted in chapter 15, given again in chapter 30.
  4. Formal and informal address, with a motivated and an unmotivated change of register.
  5. An unreliable narrator: the characters’ beliefs stay distinct from the facts.
  6. A recurring idiom or joke: the effect stays consistent without artificial repetition.
  7. A human correction in chapter 20 that must influence chapter 35 and remain restorable.
  8. A sentence with emphasis, links and a note reference: meaning and markup preserved.

Controlled comparison

Duplicate the same project before translation and keep the model, parameters and prompts identical. Compare the internal, openviking and hybrid context engines. Keep the prompts actually sent, the retrieval choices and the prompt versions.

Score separately, blind if possible: fidelity (omissions, additions), stability of names, voices, pronouns and relations, handling of ambiguities and reveals, naturalness and literary effect, respect of human decisions, EPUB structure and formatting. Measure human corrections per 1,000 words, terminology violations, wrongly resolved references, useful and useless deep-retrieval calls, latency and token use. A model grading its own translation does not replace this review.

What the test suite already guards

Regression tests reproduce known failure cases on synthetic books: book text containing prompt delimiters, series terminology (a term locked in volume 1 kept in volume 3 despite a different unlocked term in volume 2, nothing leaking from later volumes or another owner), long CJK paragraphs split at sentence ends with ruby kept, small context windows, right-to-left output and translation-memory reuse. See tests/test_prompt_hardening.py, test_series_conventions.py, test_segmentation.py, test_book_structure.py, test_small_windows.py, test_rtl.py and test_translation_memory.py. On a real book, stats.translation_memory_reused in the project data gives the reuse rate.

To measure the prompt cost of a configuration without a real model, use scripts/measure_prompt_cost.py, described in operations.md.

Comparing the analysis modes

scripts/evaluate_analysis_modes.py analyses the same synthetic serial in the strict mode and in the parallel modes (parallel: every passage reconciled; parallel-flagged: only the ambiguous ones; parallel-unreconciled: none) and scores the memory each leaves against its ground truth: who each passage involves (pronoun referents included), late aliases resolved, identities kept together in the registry, relations, glossary proposals, and identity links shown to a passage before the text reveals them. scripts/benchmark_analysis.py measures the wall time, calls and tokens of the analysis of a long serial for several thread counts. Both use backend/tests/analysis_world.py (the serial, its ground truth and a simulated analyst that only knows what its prompt holds), which tests/test_parallel_analysis.py also uses: the parallel mode must score at least as well as the strict one, show no later fact to any passage, give the same memory for any number of threads and resume at every stage without asking the model twice. The results are in architecture. They measure what each mode delivers to each call, not a real model: before changing the default mode, compare both on a real book with the protocol above (duplicate the project, same model and prompts, then strict against parallel).

python scripts/evaluate_analysis_modes.py --chapters 60 --volumes 2 --seeds 1,2,3
python scripts/benchmark_analysis.py --chapters 400 --latency 0.2 --threads 1,4,8,16

evaluate_analysis_modes.py also takes --modes (a comma-separated subset of strict, parallel, parallel-flagged, parallel-unreconciled), --threads (8) and --json; benchmark_analysis.py takes --per-chapter (passages per chapter, min,max), --capacity (the provider’s concurrency, 16), --skip-strict and --json. Neither needs a licensed installation or the network: each works on a throwaway SQLite database seeded with the synthetic licence of the tests. With several seeds, the evaluation reports the mean of the per-seed ratios and the sum of the counts (aggregation in its JSON report, whose modes are under modes).