Development and releases
This page is for contributors and maintainers. It explains how to set up a development environment, run every kind of test, how the CI/CD pipeline works, how a release is cut and deployed, and how translation quality is evaluated. To run Libris rather than work on it, read the Docker guide.
Repository layout
| Directory | Contents |
|---|---|
backend/ | FastAPI application (app/), persistent worker (app/jobs/worker.py), Alembic migrations, pytest suite |
frontend/ | React + TypeScript interface (src/), Playwright specs (e2e/) |
prompts/ | Versioned model instructions, copied into the image |
codex_bridge/ | Optional isolated Codex transport and its own image |
scripts/ | Installation, deployment, test and release utilities |
deploy/ | Production deployment, backup, restore and CI-runner maintenance procedures |
docs/ | Documentation; docs/screenshots/ and docs/openapi/ are generated (see Screenshots and The OpenAPI description) |
examples/ | Example client of the automation API (Python, standard library only) |
How the pieces fit together is described in architecture.md; the interface conventions are in
design-system.md.
Set up
Use Python 3.13 and Node.js 22, the versions of the images and of CI.
python3 -m venv .venv
.venv/bin/pip install -r backend/requirements.lock
.venv/bin/pip install -e './backend[test]'
npm --prefix frontend ci
backend/requirements.lock pins every Python dependency with its hashes. After changing a pin, regenerate the
hashes with python3 scripts/hash_lock.py backend/requirements.lock (network needed); CI runs
scripts/hash_lock.py --check and refuses an unhashed pin. Add a dependency only when it is needed, say why in
the merge request, and update the matching lockfile.
Run the application locally
The simplest loop runs the real stack with Docker Compose and the interface with Vite’s development server.
-
Create a local configuration once:
python3 scripts/setup.pywrites a.envwith generated secrets and never overwrites an existing one. -
Add the Vite origin to
ALLOWED_ORIGINS, for exampleALLOWED_ORIGINS=http://127.0.0.1:8088,http://127.0.0.1:5173.COOKIE_SECURE=trueworks onlocalhostand127.0.0.1; set it tofalseonly if you open the development server through a network address. -
Build and start the stack from your working tree, under a local image name so that the published image is never overwritten:
LIBRIS_IMAGE=libris:dev docker compose up -d --build -
Start the interface with hot reload. Vite serves it on
http://127.0.0.1:5173and forwards/apiand/healthto the API on port 8088:npm --prefix frontend run dev
Sign in with BOOTSTRAP_USERNAME and BOOTSTRAP_PASSWORD from .env. Every setting is described in
configuration.md.
Tests
Backend
cd backend
../.venv/bin/ruff check app tests migrations
../.venv/bin/pytest -q
Ruff also covers the Alembic migrations (migrations): they run in production at every upgrade, so the
backend CI job lints them with the same command.
The suite needs no service: it uses a temporary SQLite database and data directory, and mocked providers
(tests/mock_server.py, respx). It also checks the migrations: tests/test_migrations.py upgrades, downgrades
and upgrades the schema again.
pytest-xdist (in the test extra) spreads the tests over several processes; -n auto starts one per CPU:
../.venv/bin/pytest -q -n auto
Each process has its own temporary data directory and database. The schema is built once per process and
every test starts on empty tables (tests/conftest.py); a test that changes the schema through the engine
(CREATE, ALTER, DROP…) has the next one start on a freshly built schema. Fixtures that write users
straight into the database take their password hash from tests/passwords.py, computed once per session.
GitLab runs the SQLite and PostgreSQL suites with -n 6 on its 16-core runner, whose cores are also shared by
the other jobs in the pipeline. Local runs can use -n auto; there is no second CI whose result could compensate
for a failed GitLab pipeline.
Production runs PostgreSQL, so changes to queries, locking, concurrency or migrations must also pass on
PostgreSQL 17. Point the suite at a throwaway database with LIBRIS_TEST_DATABASE_URL:
docker run -d --name libris-test-db -p 127.0.0.1:5432:5432 --tmpfs /var/lib/postgresql/data \
-e POSTGRES_DB=libris_test -e POSTGRES_USER=libris -e POSTGRES_PASSWORD=libris_test_password postgres:17-bookworm
cd backend
export DATABASE_URL=postgresql+psycopg://libris:libris_test_password@127.0.0.1:5432/libris_test
export SECRET_KEY=local-test-secret-key-with-more-than-32-characters
export BOOTSTRAP_PASSWORD=local-test-password-123456
../.venv/bin/alembic upgrade head && ../.venv/bin/alembic downgrade base \
&& ../.venv/bin/alembic upgrade head && ../.venv/bin/alembic check
LIBRIS_TEST_DATABASE_URL="$DATABASE_URL" ../.venv/bin/pytest -q -n auto
docker rm -f libris-test-db
A serial run uses the database of LIBRIS_TEST_DATABASE_URL. Under -n, each worker creates its own database
next to it (libris_test_gw0, libris_test_gw1…, dropped and created again at start, dropped at exit), so the
user of the URL must be allowed to create databases, and two parallel runs must not share the same URL.
Every migration must stay reversible down to an empty schema, and alembic check must find no difference
between the models and the migrations.
Tests marked epubcheck validate exports with the real EPUBCheck. They are skipped unless EPUBCHECK_JAR points
to an EPUBCheck JAR and java is on the PATH; LIBRIS_REQUIRE_EPUBCHECK=1 turns a missing validator into a
failure. CI runs them inside the built image, which ships EPUBCheck 5.3.0.
Tests and the clock
A test that reads the real time (time.time(), datetime.now()) and adds hours, days or months can pass all
day and fail in the evening, at the end of a month or on 31 December (#167). Compute such instants from a fixed
date (see tests/test_clock_boundaries.py and DAY in tests/test_work_window.py), or pass now explicitly to
the function under test; never read “today” twice in one test.
scripts/clock_check.sh replays the suite with the clock moved by libfaketime (Debian package faketime; the
clock keeps running from the given instant, the monotonic clock is untouched). One instant, in UTC, with an
optional time zone, or every boundary near today at once:
cd backend
TZ=Pacific/Kiritimati PYTHON=../.venv/bin/python ../scripts/clock_check.sh "2026-12-31 23:30:00" -q -n auto
PYTHON=../.venv/bin/python ../scripts/clock_check.sh --boundaries -q -n auto
--boundaries runs 23:30 UTC today, 00:30 UTC tomorrow, 23:30 UTC on the last day of the month and on
31 December, the European changes of time (00:30 UTC on the last Sunday of March and of October, in
Europe/Paris), and two time zones on either side of the date line (Pacific/Kiritimati, UTC+14, and
Pacific/Pago_Pago, UTC−11). libfaketime does not move the times the kernel gives a file it writes, so the
few tests that compare them with the clock carry the file_times marker and are skipped under the script
(FAKETIME set); they run in the ordinary suite. The libfaketime of Debian 13 (trixie) is needed: the one of
bookworm makes time.sleep() fail with EINVAL. Without the package, point LIBFAKETIME at
libfaketimeMT.so.1, for example extracted with
apt-get download libfaketime && dpkg -x libfaketime_*.deb /tmp/faketime.
The OpenAPI description
docs/openapi/libris-v1.json describes the automation API (/api/v1) and is generated from the code by
backend/app/api/v1_openapi.py. After adding or changing a /api/v1 route, regenerate it and commit it:
.venv/bin/python scripts/export_openapi.py # rewrite the file
.venv/bin/python scripts/export_openapi.py --check # only compare (prints a diff)
tests/test_openapi_v1.py fails while the committed file differs from the code, checks that every /api/v1
operation is described with its scope and error responses, and compares the documented answers with real ones.
A new route needs nothing else to appear: its tag comes from its path, its summary from the first line of its
docstring, its scope from its require(...) dependency, and it gets the shared error responses. Describe what
FastAPI cannot see (a body read by hand, the fields of a dict answer, specific error codes) in
v1_openapi.py: OPERATIONS and SCHEMAS. tests/test_example_client.py runs examples/libris_client.py
against the API: when a route, an upload option or a result format changes, update the client and its test in
the same merge request.
Frontend
npm --prefix frontend run build
The build runs the strict TypeScript check before Vite. Visible text goes through the translation helpers
described in design-system.md; a new string needs its English entry.
Browser tests (Playwright)
Specs are selected by a tag in their title:
| Tag | Needs | Where it runs |
|---|---|---|
| none | nothing: the spec mocks the API with page.route | CI frontend job, locally against vite preview |
@journey | a disposable stack with the synthetic model of docker-compose.test.yml | CI e2e job |
@integration | a disposable stack holding the data created by scripts/smoke.py | by hand only |
Without LIBRIS_E2E_URL, playwright.config.ts leaves out @integration and @journey, so a plain run is
always safe on a workstation:
npm --prefix frontend run build
npm --prefix frontend exec playwright install chromium
npm --prefix frontend exec vite preview -- --host 127.0.0.1 --port 4173 # terminal 1
npm --prefix frontend run test:e2e # terminal 2
The mocked specs open http://127.0.0.1:4173 unless SHOWCASE_URL says otherwise.
To run @journey or @integration specs, point them at a disposable installation on a loopback address.
They never read the repository .env:
| Variable | Meaning |
|---|---|
LIBRIS_E2E_URL | Base URL of the disposable installation, on 127.0.0.1, localhost or ::1 |
LIBRIS_E2E_USERNAME, LIBRIS_E2E_PASSWORD | Its administrator account |
LIBRIS_E2E_CONFIRM_DISPOSABLE | Must be 1: confirms that the target holds no real data |
LIBRIS_E2E_MOCK_LLM_URL | @journey only: the synthetic model as the backend sees it (default http://mock-llm:8091) |
LIBRIS_E2E_PROJECT_ID or LIBRIS_E2E_STATE | Workspace specs: the smoke project, or the state file written by scripts/smoke.py (default /tmp/libris/epub-smoke.json) |
A disposable stack needs a licence of its own
Libris refuses the work that costs words without a valid licence, and the key certificates are checked
against is a constant of backend/app/licence/certificate.py — not a setting, so no disposable stack can
simply be told to trust another server. A build may replace it: docker build --build-arg LICENCE_KEY=…
writes another public key into the image, and that image trusts a licence server holding the private half.
A released image is built without the argument and carries the release key.
The CI e2e job does exactly this: it draws an Ed25519 pair with openssl, builds the image on the public
half, starts mock-licence (backend/tests/mock_licence.py, in docker-compose.test.yml) with the private
one, and activates the installation through the API before the specs run. LICENCE_SERVER_URL points the
stack at that container. To do the same by hand:
openssl genpkey -algorithm ed25519 -out /tmp/licence.pem
raw() { openssl pkey -in /tmp/licence.pem "$@" -outform DER | tail -c 32 | basenc --base64url | tr -d '=\n'; }
docker build --build-arg LICENCE_KEY="$(raw -pubout)" --tag libris-e2e:local .
# MOCK_LICENCE_SIGNING_KEY="$(raw)" and LICENCE_SERVER_URL=http://mock-licence:8093 in the env file,
# then activate with MOCK_LICENCE_KEY (default LIB-TEST-TEST-TEST-TEST) in Settings > Licence.
The backend suite needs none of this: tests/conftest.py gives every test a licensed installation, signed
with a key drawn once per process. A test that wants an installation without one empties app.licence.store.
Full-stack smoke test
scripts/smoke.py drives a real Compose stack with the synthetic model: it imports a generated EPUB, analyses
and translates it over HTTP, pauses and resumes, kills the worker with SIGKILL and checks recovery, then exports
an EPUB and validates it with EPUBCheck. It reads the local account without printing secrets and refuses to run
without --confirm-disposable. The commands below use the repository .env (its LIBRIS_IMAGE, PORT and
account, also read by smoke.py, or another file with --env-file): run them from a development checkout whose
.env names a local image (LIBRIS_IMAGE=libris:dev, as in Run the application locally)
and a free port, never next to a real installation.
docker compose -p libris-smoke -f docker-compose.yml -f docker-compose.test.yml --profile test up -d --build
.venv/bin/python scripts/smoke.py --compose-project libris-smoke --confirm-disposable \
--compose-file docker-compose.yml --compose-file docker-compose.test.yml
# ... run @integration specs against it if needed, then remove the test project:
.venv/bin/python scripts/smoke.py --compose-project libris-smoke --confirm-disposable --cleanup \
--compose-file docker-compose.yml --compose-file docker-compose.test.yml
docker compose -p libris-smoke -f docker-compose.yml -f docker-compose.test.yml --profile test down --volumes
Related scripts, all for disposable installations only:
| Script | What it checks |
|---|---|
scripts/check_installation.py | A fresh Compose installation on local port 4188 with a generated configuration and a random project name: login, health, migrations; then removes that stack and its volumes |
scripts/resilience_smoke.py | Provider outage, manual pause, SIGTERM and retry on a synthetic project |
scripts/make_browser_fixtures.py | Writes three small EPUBs under /tmp/libris/ for browser tests |
scripts/cleanup_browser_fixtures.py | Removes the known browser-test projects left by an interrupted run |
scripts/check_epubcheck.py | Runs the bundled EPUBCheck on valid and invalid synthetic EPUBs (inside the image) |
Never point any of these at a production library.
The resilience smoke has the same explicit boundary as the full smoke and additionally proves that a running
mock-llm belongs to the requested Compose project before it creates a volume or stops its worker:
.venv/bin/python scripts/resilience_smoke.py --compose-project libris-smoke --confirm-disposable \
--env-file .env --compose-file docker-compose.yml --compose-file docker-compose.test.yml
Screenshots
The images in docs/screenshots/ come from frontend/e2e/showcase.spec.ts, which renders the real interface
against mocked API answers and fictional books. It uses no credentials and no model, checks the colour contrast
of both themes, and writes to /tmp/libris/showcase unless told otherwise. To refresh the documentation images:
cd frontend
npx vite build
CI=1 SHOWCASE_SCREENSHOT_DIR=../docs/screenshots npx playwright test e2e/showcase.spec.ts
e2e/documentation.spec.ts captures the same views from a disposable API-backed installation instead; it only
runs with LIBRIS_DOCS_CAPTURE=1 and LIBRIS_DOCS_CAPTURE_CONFIRM=disposable. Look at every image before
committing it.
CI/CD pipeline
GitLab is the only repository and runs the release pipeline. Libris is proprietary and its source is not published.
What runs when
| Pipeline | Jobs |
|---|---|
| Merge request, branch | backend (Ruff + pytest on SQLite), backend-postgres (migration round trip + pytest on PostgreSQL 17), frontend (build, npm audit, mocked Playwright specs), e2e (@journey specs against the Compose stack), audit (version consistency, hashed lockfiles, pip-audit, Gitleaks) |
| Default branch | the same, then container-build, container-runtime, container-epubcheck, container-scan and, once everything passed, verified-image |
Release tag vX.Y.Z | release-policy, release-images, container-runtime, container-scan, publish-gitlab, release-gitlab, and the manual esoleau-bundle. Nothing is deployed anywhere |
Schedule with LIBRIS_CLOCK_CHECK | clock-check only: the backend suite at every clock boundary (see below) |
A merge request that changes nothing under backend/, frontend/, codex_bridge/, prompts/, scripts/,
deploy/, examples/, docs/openapi/, the Dockerfile, .dockerignore, the Compose files or .gitlab-ci.yml only runs audit, which still checks the
version pins in the documentation.
Nightly clock check
clock-check runs scripts/clock_check.sh --boundaries in python:3.13-trixie with the faketime package
(the libfaketime of bookworm, the base of $PYTHON_IMAGE, makes Python’s time.sleep() fail with EINVAL), and
nothing else: it never runs on a merge request or a push, and a scheduled pipeline that sets
LIBRIS_CLOCK_CHECK starts none of the other jobs. A maintainer creates the schedule once, in
Build → Pipeline schedules → New schedule:
- description:
Clock check (#167); - interval pattern:
30 23 * * *, cron time zoneUTC(the real time is then itself on a boundary); - target branch:
main; - variable:
LIBRIS_CLOCK_CHECK=1; - activated.
Run in the schedule list starts it at once. A failure names the instants and time zones that failed; replay one locally with the single-instant form of the script, then fix the test (or the code) with a fixed date.
Images and verification
Only the protected default branch builds and pushes images, addressed by commit: sha-<commit> for the
application and codex-sha-<commit> for the Codex bridge. Every later job uses those exact digests:
container-runtimestarts the image and checks it;container-epubcheckruns theepubchecktests inside the image;container-scanruns Trivy once per image: HIGH and CRITICAL findings with a fix fail the job, and the report is kept as a CycloneDX SBOM artifact;verified-imageaddsverified-sha-<commit>when the whole default-branch pipeline passed.
A tag pipeline does not test again. release-policy checks that the tag is on the default branch, that it
matches the version (scripts/check_version.py) and that CHANGELOG.md has its section. release-images waits
up to 20 minutes for the verified-sha-<commit> marker, since a tag is often pushed while the branch pipeline
is still running, then promotes those digests. If the branch pipeline failed, make it pass and retry
release-images.
A tag vX.Y.Z publishes:
| Destination | Tags | Published by |
|---|---|---|
| GitLab Container Registry | X.Y.Z, X.Y, latest (and codex- variants) | publish-gitlab |
Nowhere else: Libris is proprietary, its images are not published in any public registry and there is no
public source repository. Customers reach that same GitLab registry at registry.libris-translate.com
(registry.libris-translate.com/libris/libris:X.Y.Z, :latest, :codex-X.Y.Z), which authenticates every pull:
each licence receives its own read-only registry credentials from the licence server, in the mail that carries
its key, and loses them when the licence stops being valid. Nothing else distributes Libris: no archive, no
download page, no deployment by the pipeline.
release-gitlab creates the GitLab release; its notes are exactly the CHANGELOG.md section of that version
(python3 scripts/release_notes.py vX.Y.Z prints them). A retried job updates the existing release.
CI runner hygiene
All jobs run on a shell runner. Test, build and scan tools start with docker run; only small release steps
(release-policy, release-gitlab) use the host’s git and python3. Every container is named
libris-ci-<role>-$CI_JOB_ID, carries the label libris-ci-job=$CI_JOB_ID and runs under --init; every
command that can hang is wrapped in timeout, and every job has its own timeout:. Cancelling a job only kills
the Docker client, so each job’s after_script (which also runs on cancellation) removes the containers with its
label, and e2e takes its Compose project libris-e2e-$CI_JOB_ID down with its volumes. As a last resort,
audit removes any CI container or libris-e2e-* stack older than two hours. To look by hand:
docker ps --all --filter label=libris-ci-job.
Volumes can still be left behind by a killed job. deploy/libris-runner-prune removes, on the runner host, unused
anonymous volumes and the volumes and networks of libris-e2e-* projects older than
LIBRIS_PRUNE_MIN_AGE_HOURS (default 6); it never touches named volumes of other projects, images or the build
cache. Install it as root on the runner host:
install -m 0755 deploy/libris-runner-prune /usr/local/sbin/libris-runner-prune
install -m 0644 deploy/libris-runner-prune.service deploy/libris-runner-prune.timer /etc/systemd/system/
libris-runner-prune --dry-run # see what it would remove
systemctl daemon-reload && systemctl enable --now libris-runner-prune.timer
The timer runs hourly; journalctl -u libris-runner-prune shows what was removed.
Required GitLab settings
- Protect
mainand tags matchingv*. - Enable the Container Registry and protect the
sha-*,codex-sha-*,verified-sha-*and exact-version tags from being overwritten. - Keep no push mirror: Libris is not published anywhere public.
- The predefined
CI_REGISTRY*variables authenticate the GitLab registry, which is the only destination.
Never commit a password, token, .env, registry credential or user book.
Prepare dependency upgrades on GitLab. pip-audit, npm audit, the hashed Python locks and the Trivy image
scan are the release gates; advisories observed elsewhere are inputs to an issue, never evidence that GitLab
has already tested a fix.
Releasing
Libris follows Semantic Versioning. While in 0.x, any operationally breaking change is called out in
CHANGELOG.md and therefore in the release notes. A tag publishes, it deploys nothing: the tag pipeline
promotes the verified images to X.Y.Z, X.Y and latest and creates the GitLab release, and every
installation that follows latest takes the new version the next time its administrator runs the installer.
Agree on the release with the maintainer before tagging.
Prepare
- Set the new version everywhere
scripts/check_version.pylooks:backend/pyproject.toml,backend/app/__init__.py,frontend/package.jsonand its lockfile,codex_bridge/package.jsonand its lockfile,codex_bridge/rpc.py, theLIBRIS_VERSIONargument of both Dockerfiles,scripts/install-docker.sh, and theLIBRIS_TAG=X.Y.Zpins ofREADME.md,docs/docker.mdanddocs/docker.fr.md. LeaveLIBRIS_IMAGEandLIBRIS_CODEX_IMAGEin.env.exampleonlatestandcodex-latest: the script requires it, and the installer replaces them with digests. - In
CHANGELOG.md, rename## [Unreleased]to## [X.Y.Z] - YYYY-MM-DD, and do the same inCHANGELOG.fr.md(## [Non publié], a complete translation):backend/tests/test_documentation.pyfails when a release of the English changelog has no French section. Check the notes withpython3 scripts/release_notes.py vX.Y.Z. - When
backend/app/licence/changed since the last tag (git diff vPREVIOUS..HEAD -- backend/app/licence/), check that the licence server in production accepts what the new version sends and answers what it expects: its version is onhttps://sub.libris-translate.com/health. A Libris that needs a newer licence server is released after that server is deployed, never before. - Run
python3 scripts/check_version.py, then the backend, PostgreSQL, frontend and installation checks above. - Merge to
mainand wait until the default-branch pipeline is green (theverified-sha-<commit>marker).
Tag
VERSION=X.Y.Z
python3 scripts/check_version.py "v${VERSION}"
git tag -s "v${VERSION}" -m "Libris ${VERSION}"
git push origin "v${VERSION}"
Use an annotated unsigned tag (git tag -a) only when signing is not configured. Never move or reuse a release
tag.
Verify
A green pipeline is not proof of a release. Check the two destinations on their own: the GitLab registry and the
GitLab release. Pull a published digest from registry.libris-translate.com with customer credentials, inspect
its OCI version and revision labels, and query /health from a disposable stack installed with
LIBRIS_TAG=X.Y.Z before announcing it.
Publish and announce
The release reaches customers through the website, libris-translate.com, which is published from its own repository once the images are out:
- The website regenerates its technical documentation from the tag (these
docs/, the changelogs, the README and the OpenAPI description, as they are atvX.Y.Z): a correction made here after a tag is public only with the next release. - It raises its version, which it publishes at
https://libris-translate.com/version.jsonand in the pinned installer commands. - The licence server reads that file every hour and announces the version to every installation in the
latest_versionfield of its answers; administrators then see the update banner (see docker.md). Merging the website is therefore what announces a release.
The installer customers run, https://libris-translate.com/install.sh, is deploy/install.sh. Whatever copy
starts, it hands over to the installer carried by the image it installs (/app/deploy/install.sh), so a change
to the installer ships with the image; the website’s copy must still be refreshed from the tag, since it is the
one a new customer starts with.
Deployment script for a host of your own
The pipeline deploys nothing. deploy/libris-production-deploy and deploy/librisctl remain in the
repository for an operator who drives one host from a pipeline of their own; they are described in
operations. A customer installation is updated with the
installer (docker.md); a checkout, with scripts/deploy.sh
(operations).
When a deployed release misbehaves and it added no migration, libris-production-deploy --rollback (or
librisctl rollback --confirm) redeploys the retained previous-* images through the same guarded procedure
(dump, stop, migrate as a no-op, start, /health, worker restart). It refuses, before touching anything, when
the release added a migration: reverting an image does not revert a schema. In that case:
- Stop the application services:
docker compose --project-name <project> --env-file <.env> --file <compose file> --profile codex stop api worker codex. - Identify the dump taken before the release being reverted, not merely the newest dump.
Restore it with
librisctl restore <dump> --confirm --no-start, or the empty-database procedure in backup.md. The extra flag is essential: restarting the current image would immediately reapply its migrations.pg_restore --cleanover the newer schema is not sufficient; tables added since the dump can block restoration through their foreign keys. - Run
libris-production-deploy --rollbackagain: the schema now matches.
A deployment never changes the books in /data or the .env; restore them from the scheduled backup only if they
were damaged.
Restoration discards database changes made after the chosen dump, including newly issued review links.
The CLI keeps a separate pre-restore-*.dump of the replaced state, with mode 0600 and no automatic
pruning. Do not delete that rescue copy until the recovery is verified.
Before handing a build to anyone
- Review the whole Git history, not only the working tree, for secrets, private hostnames, books and test artifacts. Rotate any secret that was ever committed.
- Confirm the rights to the logo and every included asset.
- Never transfer a deployment
.env, volumes or backups. - Do not announce images or releases before they exist.
- Libris is proprietary (see
LICENSE): the source is not published, and an image is handed over with the registry credentials that go with a licence. Nothing in the repository may say otherwise —scripts/check_version.pyfails the build if a file declares another licence.
Evaluating translation quality
Passing tests, EPUBCheck and automatic scores show that the output is well formed, not that the translation is faithful. Before claiming a quality level for a language pair or a kind of book, run a documented evaluation.
Corpus
Use a text you are allowed to use, several dozen chapters long, with a reference reviewed by a bilingual reader. It must contain distant callbacks and variations in the source; repeating the same paragraph does not test narrative memory. Include at least:
- An invented name with typographic variants: no unjustified change of translation.
- A character identified late: ambiguous pronouns stay ambiguous before the reveal.
- An object given in chapter 1, reinterpreted in chapter 15, given again in chapter 30.
- Formal and informal address, with a motivated and an unmotivated change of register.
- An unreliable narrator: the characters’ beliefs stay distinct from the facts.
- A recurring idiom or joke: the effect stays consistent without artificial repetition.
- A human correction in chapter 20 that must influence chapter 35 and remain restorable.
- A sentence with emphasis, links and a note reference: meaning and markup preserved.
Controlled comparison
Duplicate the same project before translation and keep the model, parameters and prompts identical. Compare the
internal, openviking and hybrid context engines. Keep the prompts actually sent, the retrieval choices and
the prompt versions.
Score separately, blind if possible: fidelity (omissions, additions), stability of names, voices, pronouns and relations, handling of ambiguities and reveals, naturalness and literary effect, respect of human decisions, EPUB structure and formatting. Measure human corrections per 1,000 words, terminology violations, wrongly resolved references, useful and useless deep-retrieval calls, latency and token use. A model grading its own translation does not replace this review.
What the test suite already guards
Regression tests reproduce known failure cases on synthetic books: book text containing prompt delimiters,
series terminology (a term locked in volume 1 kept in volume 3 despite a different unlocked term in volume 2,
nothing leaking from later volumes or another owner), long CJK paragraphs split at sentence ends with ruby kept,
small context windows, right-to-left output and translation-memory reuse. See tests/test_prompt_hardening.py,
test_series_conventions.py, test_segmentation.py, test_book_structure.py, test_small_windows.py,
test_rtl.py and test_translation_memory.py. On a real book, stats.translation_memory_reused in the project
data gives the reuse rate.
To measure the prompt cost of a configuration without a real model, use scripts/measure_prompt_cost.py,
described in operations.md.
Comparing the analysis modes
scripts/evaluate_analysis_modes.py analyses the same synthetic serial in the strict mode and in the
parallel modes (parallel: every passage reconciled; parallel-flagged: only the ambiguous ones;
parallel-unreconciled: none) and scores the memory each leaves
against its ground truth: who each passage involves (pronoun referents included), late aliases resolved,
identities kept together in the registry, relations, glossary proposals, and identity links shown to a passage
before the text reveals them. scripts/benchmark_analysis.py measures the wall time, calls and tokens of the
analysis of a long serial for several thread counts. Both use backend/tests/analysis_world.py (the serial, its
ground truth and a simulated analyst that only knows what its prompt holds), which
tests/test_parallel_analysis.py also uses: the parallel mode must score at least as well as the strict one, show
no later fact to any passage, give the same memory for any number of threads and resume at every stage without
asking the model twice. The results are in architecture. They measure what each
mode delivers to each call, not a real model: before changing the default mode, compare both on a real book with
the protocol above (duplicate the project, same model and prompts, then strict against parallel).
python scripts/evaluate_analysis_modes.py --chapters 60 --volumes 2 --seeds 1,2,3
python scripts/benchmark_analysis.py --chapters 400 --latency 0.2 --threads 1,4,8,16
evaluate_analysis_modes.py also takes --modes (a comma-separated subset of strict, parallel,
parallel-flagged, parallel-unreconciled), --threads (8) and --json; benchmark_analysis.py takes
--per-chapter (passages per chapter, min,max), --capacity (the provider’s concurrency, 16),
--skip-strict and --json. Neither needs a licensed installation or the network: each works on a
throwaway SQLite database seeded with the synthetic licence of the tests. With several seeds, the evaluation
reports the mean of the per-seed ratios and the sum of the counts (aggregation in its JSON report, whose
modes are under modes).