Source docs/operations.md · 1de96aa

Operations

This page is for administrators who keep Libris running: updating it without losing work, sizing its capacity, monitoring it, keeping its database small, controlling model costs and solving common problems. Installation is covered in the Docker guide, every setting in the configuration reference, and backups in backup and restore.

All commands run from the Libris directory unless stated otherwise.

Update Libris

The simple customer path is the one in the Docker guide: back up, then run the registry bootstrap again. It needs neither Git nor a source archive. It runs the migrations, then restarts the web application and the worker; running books resume from their last checkpoint, so no finished work is lost, but the passages being translated at that moment are sent again. The migrations run just before the worker is replaced, while the previous one may still be working: pause the running books first, as the Docker guide says, or use scripts/deploy.sh.

For a development checkout, scripts/deploy.sh gives finer control over when the worker restarts. Unlike the registry bootstrap, it does not change the image named in .env: set LIBRIS_IMAGE there to the new release first.

sed -i 's|^LIBRIS_IMAGE=.*|LIBRIS_IMAGE=registry.libris-translate.com/libris/libris:<new version>|' .env
LIBRIS_DEPLOY_SOURCE=pull ./scripts/deploy.sh --worker-when-idle
ModeWhat it does
--api-only (default)Gets the image, runs the migrations and restarts the web application only; the worker keeps running the previous version. Only for updates without a database migration: when the new version brings one and the worker is running, the script changes nothing and exits with status 5 (with the worker stopped, it migrates and leaves the worker stopped).
--worker-when-idleGets the image, waits up to 10 minutes for the job queue to be empty, stops the web application so that no new job starts, checks again, stops the worker, runs the migrations and starts both. If jobs stay active it changes nothing and exits with status 3 (4 if a job started at the last moment).
--force-workerGets the image, stops the web application and the worker at once, runs the migrations and starts both. Running jobs resume from their checkpoints.
VariableMeaning
LIBRIS_DEPLOY_SOURCEbuild (default) builds the image from the source tree and tags it with the LIBRIS_IMAGE of .env: give it a local name (for example LIBRIS_IMAGE=libris:local), not an official registry tag; pull downloads LIBRIS_IMAGE; loaded uses an image already present locally (LIBRIS_IMAGE required).
LIBRIS_COMPOSE_ENV_FILEFile passed to Compose with --env-file, for the values Compose itself reads (LIBRIS_IMAGE, PORT…). The containers still read LIBRIS_ENV_FILE (.env by default): set both when the configuration lives elsewhere.

A wrong mode or LIBRIS_DEPLOY_SOURCE exits with status 2.

If the new version fails to start, the script puts the previous image back (tagged libris:rollback) and restarts it. It cannot undo a migration: for that, see roll back a failed update.

With deploy.sh, the migrations never run while a worker of the previous version is running. Use --api-only for a quick fix of the web side, then finish with --worker-when-idle when books are done, so that the worker does not keep running an older version for long.

Production deployment script

deploy/libris-production-deploy deploys a tagged release, already loaded as images, onto a dedicated host. The project’s own pipeline no longer deploys anything (a tag only publishes images and a release, see development): the script stays for an operator who drives a host from a pipeline of their own, and is of no use to an installation made with the installer. Install it as /usr/local/sbin/libris-production-deploy.

Host-specific settings can be kept in /etc/libris-production.conf (root-owned, mode 0600), using deploy/libris-production.conf.example. This trusted shell file is sourced before the script resolves its settings, including for rollback and pruning; its assignments override the environment. It is not the application’s .env and must never be writable by an application account.

libris-production-deploy COMMIT VERSION APP_IMAGE_ID CODEX_IMAGE_ID   # deploy images already loaded
libris-production-deploy --rollback                                  # back to the images it replaced
libris-production-deploy --check-compose                             # is the installed Compose file the deployed version's?
libris-production-deploy --prune-images [--dry-run]                  # remove old Libris images

A deployment checks the images’ version and revision labels, refuses to go back to an older version (use --rollback), dumps the database into $LIBRIS_PRODUCTION_BASE/backups/ (last five kept), stops the application, installs the Compose file shipped inside the new image, runs the migrations, starts the API and the Codex bridge, checks /health, and only then starts the worker. If anything fails before the schema changed, it restarts the previous images; after a schema change it leaves the services stopped and names the dump to restore.

VariableDefaultMeaning
LIBRIS_PRODUCTION_CONFIG/etc/libris-production.confOptional trusted host configuration file. A missing file leaves environment settings in use.
LIBRIS_PRODUCTION_BASE/opt/libris-productionWorking directory: Compose file, recorded version, pre-deployment dumps.
LIBRIS_PRODUCTION_COMPOSE_FILE$LIBRIS_PRODUCTION_BASE/docker-compose.ymlInstalled Compose file.
LIBRIS_PRODUCTION_COMPOSE_OVERRIDEemptyOptional extra Compose file for local additions.
LIBRIS_PRODUCTION_SECRET_ENV$LIBRIS_PRODUCTION_BASE/.envThe installation’s .env. Without this variable the scripts take $LIBRIS_PRODUCTION_BASE/.env, then, if that file does not exist, the legacy /opt/epub-translator/.env, and say so once per run until librisctl migrate-env --confirm moves it.
LIBRIS_PRODUCTION_PROJECTepub-translatorCompose project name. Without this variable the scripts read $LIBRIS_PRODUCTION_BASE/project when that file exists (librisctl migrate-project writes it), and fall back to epub-translator.
LIBRIS_PRODUCTION_HEALTH_URLhttp://127.0.0.1:8088/healthHealth URL checked after deployment. Set it explicitly when the API is bound elsewhere.
LIBRIS_PRODUCTION_REGISTRY_REPOSITORYemptyExact registry repository whose old pulls --prune-images removes. Empty disables registry-pull cleanup; unused local Libris image tags can still be removed.

Change docker-compose.yml in the repository, never on the host: host-specific values belong in the .env (BIND_ADDRESS, PORT…) or in the override file, and the replaced file is kept as docker-compose.yml.before-<version>. When disk space is short, delete old files inside backups/, never the directory: without its Compose file the next deployment stops with Production configuration is not provisioned (exit 65) before touching anything. Reinstall /usr/local/sbin/libris-production-deploy whenever deploy/libris-production-deploy changes, and validate a deployment and a rollback on a disposable stack before trusting a new host.

Each deployment also refreshes /usr/local/sbin/librisctl from the image it deploys (the image carries deploy/librisctl in /app/deploy/), so the operator command line is always the one of the running version.

Operator command line

librisctl is the everyday command line of such a host: one command per task, with the Compose project, the configuration file and the Codex profile resolved exactly as the deployment script resolves them. It is installed as /usr/local/sbin/librisctl by each deployment, or by hand with install -m 0755 deploy/librisctl /usr/local/sbin/librisctl. Run it with no argument for the list, or librisctl <command> --help for one command. Its output is plain text without colour or pager, so it reads the same over SSH, in a script and in a log.

CommandWhat it does
statusDeployed version and commit, image ids, every container with its state and health, /health, the size of the data volumes and their free space, the most recent dump. Changes nothing.
psThe installation’s containers, stopped ones included.
logs [SERVICE...] [-f] [--since D] [--tail N]Container logs without colour codes.
journal [-f] [-n N] [FILTER...], journal export [FILTER...] [--output FILE] [--with-containers]The journal: every service in one chronology, the last events, followed, or a period exported.
restart [SERVICE...]Restarts api, worker and codex by default. Running books resume from their checkpoints. A container keeps the environment it was created with, so a restart does not apply a changed .env: when a service runs with a value the configuration no longer says, the command names that service, the settings concerned and up. A container that was never given a changed setting is left alone, so a database started long before an ALLOWED_ORIGINS change is not reported. Values are never printed.
start [SERVICE...], stop [SERVICE...]Start or stop services. stop keeps every container and volume; stopping the database (by name or with --all) needs --confirm.
upApplies the installed Compose file with the deployed images: recreates what changed, never builds, never pulls. This is what applies a changed .env — a setting such as ALLOWED_ORIGINS only takes effect once the containers are recreated.
health, versionIs the API answering with the deployed version (exit 1 if not); what is deployed.
jobsRead-only view of the queue from the database: jobs by state, live jobs with their priority and lease, expired leases. Ids only, never a book title.
shell SERVICE [COMMAND...], psql [ARGUMENT...]A shell, or psql as the installation’s own user, inside a container. Interactive only when the terminal is.
backupDumps the database now into $LIBRIS_PRODUCTION_BASE/backups/manual-<timestamp>.dump (mode 600) and prints the path. The full scheduled backup stays libris-backup.
restore DUMP --confirm [--no-start]Checks the archive, stops application services, saves a private rescue dump, recreates the application database and restores it. Normally restarts everything, migrations included. Use --no-start before rollback so the newer image cannot immediately reapply its schema. Rescue dumps pre-restore-*.dump are not pruned automatically.
rollback --confirmRuns libris-production-deploy --rollback.
prune-images [--dry-run] [--confirm]Runs libris-production-deploy --prune-images. No volume is ever touched.
check-composeRuns libris-production-deploy --check-compose; exits 1 and prints the difference on drift.
doctorOne pass over the whole host: configuration in place and private, containers running and healthy, /health, Compose drift, container hardening (read-only root, no-new-privileges, dropped capabilities, pids limit), free disk, a recent dump, the journal directory in place and private, the configuration the application expects (every mount of its Compose file, a journal that is written: a degraded one fails, an outdated one warns), and a worker whose leases are not expired. Exits 1 if any check fails.
migrate-env [--confirm]Moves a configuration file left at the legacy path under the production base, keeping its permissions.
migrate-project [NAME] [--confirm]Renames the Compose project (see below).

Exit codes: 0 success, 1 a check failed, 64 wrong usage, 65 production is not provisioned on this host or a precondition is not met, 77 refused (a destructive command without --confirm). librisctl down and friends are refused outright: the data volumes are the only copy of the books and of the database.

The common tasks:

librisctl status                     # what is running, and is it healthy?
librisctl doctor                     # everything that should be true, checked in one pass
librisctl restart worker             # the worker stopped taking jobs; books resume from their checkpoints
librisctl logs -f worker             # follow it
librisctl logs --since 30m api worker
librisctl journal -f                 # every service in one chronology, as it happens
librisctl journal export --since 2h --with-containers --output /root/incident.log
librisctl backup                     # a database dump before touching anything
librisctl rollback --confirm         # back to the version the last deployment replaced

Nothing prints the configuration or a key: status shows where the .env is, never what it holds.

Rename the Compose project

An installation made with the installer is renamed by the installer itself (Docker). On a production host driven by libris-production-deploy, the historical project name is epub-translator, and its volumes hold the books and the database, so it cannot simply be renamed. librisctl migrate-project [NAME] --confirm (libris by default) does it in one go: it refuses while a job is active, stops the installation, lets Compose create the new project’s volumes, copies each volume into its twin, compares their sizes, starts the installation under the new name and records it in $LIBRIS_PRODUCTION_BASE/project, which every script then follows. The old volumes are left exactly as they were; stateless service containers are recreated under the stable libris-* names. Removing $LIBRIS_PRODUCTION_BASE/project and running librisctl up returns to the old volumes, and docker volume rm epub-translator_database epub-translator_books epub-translator_codex-state removes them once the new name has proved itself. Add LIBRIS_PROJECT=<name> to the .env as well if you also run plain docker compose commands: the Compose file takes its project name from that variable.

The Compose options for a manual command on such a host, when librisctl is not enough:

docker compose --project-name epub-translator --env-file /opt/libris-production/.env \
  --file /opt/libris-production/docker-compose.yml --profile codex <command>

Capacity and concurrency

Two limits decide how much work runs at once:

  • Concurrent books, set on each provider in Settings › LLM providers: how many books may use that provider at the same time, all operations included (analysis, translation, review). Providers are independent: one set to 3 and another set to 1 allow four active books.
  • WORKER_BOOK_PARALLELISM: how many passages of one book are in flight at once, for the analysis (in the default parallel mode, ANALYSIS_MODE=parallel) as for the translation and the reviews. The default, 0, uses the provider’s capacity, shared between the books using it. 1 processes one passage at a time. A volume (Passages worked on at once in its settings), a launch or an API request (threads) can only ask for fewer. After a provider answers 429 or is overloaded, the job resumes at half its width and widens again by one passage per minute.

The strict analysis (ANALYSIS_MODE=strict) reads one passage after the other whatever these limits; the parallel mode makes about 1.8 times more analysis calls but ends several times sooner on a long book (see architecture).

When books wait, they do not start oldest first: the fair queue takes the highest priority, then the account and the API token with the fewest jobs running, in turn between accounts, so one person’s backlog does not hold everyone else back. On a shared installation, Settings › Queue (or QUEUE_*, see configuration) can also cap each account’s running and waiting jobs. The Queue page shows each waiting job’s place and what holds it.

A book stays queued while its provider has no free slot. To change a book’s provider mid-way, pause it, choose the new provider in its configuration and resume: the rest of the book uses the new one.

One worker is enough for most installations. It renews a 60-second lease on each job; if the worker stops abruptly, another start picks the job up after the lease expires.

Worker processes

A worker process runs on one CPU core, whatever the number of books it carries. On a machine with spare cores, WORKER_PROCESSES=N (2 to 16; default 1, the single process described above) makes the worker container run the books in N job processes:

  • The worker’s main process keeps everything that must run once: licence renewal, watched sources, retention, mail, webhooks, library publication, the memory catalog and OpenViking. It starts the N job processes, starts one again when it dies, and on docker compose stop passes the signal on and waits for them (25 seconds, inside the service’s 30-second stop_grace_period).
  • Each job process takes books from the same queue. Leases, the fair queue, a provider’s Concurrent books and “one volume of a series at a time” are kept in the database and hold across processes. A process leaves the next book to the others once it holds more than its share (books under way divided by N, rounded up), so the books spread over the processes.
  • If a job process dies, its books are picked up by another one when their 60-second lease expires; the log says process=died … role=jobs then process=started, and status=reclaimed for each book taken over.

Limits to know before raising it:

  • One book alone stays on one core: its passages are threads of one process. The setting helps when several books run at once, not a single long one.
  • One worker container only. Do not scale the worker service or start a second one: the loops kept in the main process (licence, watched sources, retention) are not protected against running twice.
  • Every process has its own database pool: raise max_connections with the number of processes (see the pool settings in .env.example).
  • The worker service’s pids_limit: 512 covers up to 8 job processes (a process uses at most about 45 threads); raise it in a Compose override beyond that.
  • A demo instance (DEMO_MODE) starts no job process.

The worker also sends memory updates to OpenViking and, when the cleanup is on, removes the OpenViking documents of deleted volumes and series. A removal holds a 5-minute lease too; a waiting one is visible, with its last error, in Settings › Memory · OpenViking › Cleanup log.

Database connections

Every process holds a connection pool of its own, DB_POOL_SIZE + DB_POOL_MAX_OVERFLOW connections (80 by default). PostgreSQL must accept all of them at once:

max_connections >= (API process + worker processes) x (DB_POOL_SIZE + DB_POOL_MAX_OVERFLOW) + 20

With WORKER_PROCESSES=1 there is one worker process. Above 1 there are N children and the parent that supervises them, so N + 1. The 20 are left for librisctl, a psql and the database’s own maintenance. The supplied Compose file reads max_connections from POSTGRES_MAX_CONNECTIONS in .env (200 by default, enough for the defaults with one worker: (1 + 1) x 80 + 20 = 180).

WORKER_PROCESSESWorker processesBudget with the default pool of 80POSTGRES_MAX_CONNECTIONS
1 (default)1(1 + 1) x 80 + 20 = 180200 (default)
23(1 + 3) x 80 + 20 = 340340
45(1 + 5) x 80 + 20 = 500500

Both processes check it at start-up. When max_connections is below the budget they log db_pool=over_budget max_connections=… needed=… hint=set POSTGRES_MAX_CONNECTIONS=… and start anyway: the failure would then show as pool_exhausted or refused connections under load. Fix it by raising POSTGRES_MAX_CONNECTIONS (and restarting the database service), or by lowering DB_POOL_SIZE and DB_POOL_MAX_OVERFLOW — the pool only needs the sum of the providers’ Concurrent books plus 16.

Monitoring

Health

GET /health answers {"status":"ok","version":"<version>","configuration":"ok"} when the API and its database connection work. It does not test model providers; test those in Settings.

configuration compares the containers with what the image’s own Compose file (/app/deploy/docker-compose.yml) expects: every mount it gives the application (/data, /etc/machine-id, /logs) and a journal that is written. ok; outdated when a mount is missing (the Compose file predates the version); degraded when a feature is off because of it — a journal that cannot be written. The API and the worker say it at start-up (configuration=… check=… in their output). The list of checks, each with the command that fixes it, is in the answer for a signed-in administrator, or for everyone with PUBLIC_HEALTH_DETAILS=true; administrators also see it in Settings › Installation health, and librisctl doctor prints it.

With the same details, openviking is the state the worker last saw (reachable, unreachable, unknown or not_configured), since when, and the send queue (pending, error, next_attempt); never its address. OpenViking is optional: an outage there never changes status nor configuration. See OpenViking.

curl --fail http://127.0.0.1:8088/health

Prometheus metrics

GET /metrics exposes the state of the installation in Prometheus text format. It is disabled until you set METRICS_TOKEN:

sed -i "s|^METRICS_TOKEN=.*|METRICS_TOKEN=$(openssl rand -hex 32)|" .env   # the line exists in a generated .env
grep '^METRICS_TOKEN=' .env                                                  # the value to give Prometheus
docker compose up -d --no-build --wait

Each scrape must send the token as Authorization: Bearer <token>; a browser session is not enough. Without the variable the endpoint answers 404, with a wrong token 401.

scrape_configs:
  - job_name: libris
    scrape_interval: 60s
    metrics_path: /metrics
    scheme: https
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/libris-metrics-token
    static_configs:
      - targets: ["books.example.com"]
MetricTypeContent
libris_jobs{operation,status}gaugeJobs by operation and state, finished ones included
libris_jobs_oldest_queued_age_secondsgaugeHow long the oldest job ready to run has been waiting (a resumed job counts from its resumption)
libris_jobs_expired_leasesgaugeRunning jobs whose 60-second lease expired: their worker stopped or hangs
libris_llm_requests_total{operation,status}counterFinished model requests by operation and outcome (success, error, refused, interrupted, abandoned); cached answers count as success
libris_llm_input_tokens_total{operation}, libris_llm_output_tokens_total{operation}counterTokens reported by the providers
libris_llm_wasted_input_tokens_total{operation}counterInput tokens of requests that failed, were refused or interrupted
libris_llm_cache_hits_total{operation}counterRequests answered from the response cache
libris_llm_cache_hit_ratiogaugeShare of all finished requests answered from the cache
libris_llm_requests_in_flight{provider}gaugeRequests in progress, by provider name
libris_segments{status}gaugePassages of all books, by state
libris_memory_outbox_pendinggaugeOpenViking updates not delivered yet

Counters are computed from the database: they count everything since installation, are the same in every process and survive restarts. Deleting a book deletes its requests, which Prometheus sees as a counter reset. Use rate() or increase() for a time window. No label contains a book title, text, id or provider address. The output is cached for 10 seconds, so scraping every 30 to 60 seconds is enough.

Useful alerts:

ConditionMeaning
libris_jobs_expired_leases > 0 for 5 minutesThe worker is stopped or stuck.
libris_jobs_oldest_queued_age_seconds > 900No worker takes jobs, or a provider is saturated.
sum(rate(libris_llm_wasted_input_tokens_total[1h])) / sum(rate(libris_llm_input_tokens_total[1h])) > 0.2More than one input token in five is spent on requests that produced nothing.

Behind a reverse proxy, expose /metrics only to your Prometheus network if you can.

Statistics in the interface

Statistics shows token usage by model, and each book shows what it has spent. Costs use the price recorded with each request, so a later price change on a provider does not rewrite past costs.

Logs

docker compose logs --since=30m api worker
docker compose logs -f worker

Each container keeps at most three log files of 10 MB. Review logs before sharing them: remove account names, addresses and anything that looks like a key.

Journal

To reconstruct an incident, read the journal instead: one chronology of the API, the worker, the migrations and the deployments, in files that survive a container’s recreation.

Where. Each process writes its own file in the journal directory: api.log, worker.log, migrate.log, and deploy.log for the production deployment script. A second process of the same service (several API workers, a docker compose run) writes api-<pid>-<random>.log instead, so no two processes ever append to, or rotate, the same file. On a host deployed by libris-production-deploy the directory is /opt/libris-production/logs ($LIBRIS_PRODUCTION_BASE/logs, or LIBRIS_PRODUCTION_LOG_DIR); elsewhere it is the logs volume of the Compose project (docker volume inspect libris_logs), or the host directory named by LIBRIS_LOG_DIR in the .env — create it first, owned by uid 10001 and mode 0700, or Docker creates it for root and the journal stays off.

What a line says. One line per event:

2026-09-24T08:15:03.123+00:00 ERROR    worker   pid=7 req=- job=6f0c1e2a-… epub.worker | job=6f0c1e2a-… status=failed …
    Traceback (most recent call last):
      File "/app/backend/app/jobs/worker.py", line 290, in _execute
    …

The time with its UTC offset (LOG_TIMEZONE), the level, the service, the process, the request (req=, the X-Request-ID every API response carries, taken from the client when it sends a well-formed one), the job (job=; deploy-<time>-<pid> for the steps of one deployment), the logger and the message. Starts, readiness and stops of each process, each API request (method, path, status, duration; /health and static files excepted), each job’s start (resumed=yes when it picks up a checkpoint), completion, pause and failure, the migrations run, and the steps of a deployment are all there, next to what the code already reported (status=database_unavailable, db_pool=…, licence=…).

A traceback follows its event on indented lines. It keeps every frame and the chain of causes, and the message of the exceptions whose message is a technical cause (a refused connection, a missing file, a TLS or HTTP error); for the others — which may quote a passage, a prompt or an account — and for database drivers, which repeat rejected values, only the type is kept (the SQLSTATE code for the latter).

What is never written. Passwords, cookies, Authorization headers, API tokens (their public lbr_… prefix is kept), licence keys, registry and other access tokens, private keys, credentials in URLs, and every value of the configuration whose name says it is a secret (SECRET_KEY, POSTGRES_PASSWORD, *_TOKEN…) are replaced by ***. Line breaks and control characters in a message are escaped (\n), so a file name or a header cannot forge an event, and a message is cut after 4,000 characters, so no chapter is ever copied whole. The same protections apply to the container output. An export still names accounts, books and addresses: review it before sharing it.

Reading it. On a production host:

librisctl journal                                  # the last 200 events of every service
librisctl journal -f --service api --service worker
librisctl journal --level WARNING --since 2h
librisctl journal --request 3f9c0a1b2c3d4e5f      # everything one request did
librisctl journal --job 6f0c1e2a-…                  # everything one job did
librisctl journal export --since 2026-09-24T07:00 --until 2026-09-24T09:00 --with-containers \
    --output /root/incident.log                     # a period, with the database and Codex output merged in

librisctl journal reads the files with the deployed image, in a container without network that mounts them read-only: it works while the installation is stopped. Without the production scripts, run the same tool in a container of the installation (python -m app.journal --help lists the options):

docker compose exec api python -m app.journal -n 100
docker compose exec api python -m app.journal follow
docker compose exec api python -m app.journal export --since 1d > incident.log

Size and permissions. Each file is rotated at LOG_MAX_MB (10 MB) and LOG_BACKUPS (5) rotated files are kept per process; files older than LOG_RETENTION_DAYS (14) go, but never the one a running process writes. The directory is mode 0700 and the files 0600, owned by the application’s user (uid 10001); root reads them on the host. See configuration.

When it cannot be written. A service never stops for its journal. If the directory cannot be opened at start-up (journal: disabled for api, … is not writable on the container output), or a write fails later (a full disk: journal: cannot write … events still go to this output, at most once every five minutes, with the count of the events missing from the file), the events keep going to the container output and the file is tried again at the next event. A deployment that cannot write deploy.log says so and goes on.

Data retention

Every model call leaves a row with its prompt and answer, which the request inspector shows. Without limits, this history grows quickly. The worker cleans it up once at start-up and then every hour, in small batches, following the RETENTION_* settings (configuration):

  • prompts, raw answers and context traces of requests older than 30 days are emptied (the inspector no longer shows them), while tokens, cost, duration, status and the cached answer are kept;
  • progress events older than 7 days are deleted, always keeping the last 500 of each book;
  • OpenViking updates already delivered are deleted after 7 days;
  • only the 20 most recent automatic Book Bible revisions of each book are kept (human revisions are all kept);
  • the per-passage resume state of jobs that ended more than 30 days ago is deleted. If such a job is resumed later, its finished passages are still not translated again (unless you force a new translation);
  • result files of automation requests are deleted after 30 days; asking for the result again rebuilds it;
  • files of expired imports are deleted from DATA_DIR/staging.

Before deleting anything, the same pass adds requests older than two hours to daily usage totals, which the statistics and /metrics read. With RETENTION_REQUEST_ROWS_DAYS set (for example 180), whole request rows older than that are then deleted; statistics stay correct, but the response cache and the inspector lose them.

See what a pass would remove, or run one now:

docker compose exec api python -m app.maintenance.retention --dry-run
docker compose exec api python -m app.maintenance.retention

PostgreSQL reuses the freed space, but gives it back to the system only after a full vacuum. It locks the table, so stop the worker first and make sure the disk has as much free space as the table’s useful size:

docker compose stop worker
docker compose exec database sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "VACUUM (FULL, ANALYZE) llm_requests;"'
docker compose start worker

Maintenance commands

CommandWhat it does
docker compose exec api python -m app.maintenance.retention [--dry-run]Runs the retention pass described above.
docker compose exec api python -m app.maintenance.usage [--dry-run]Adds finished requests to the daily usage totals now. The worker does it every hour.
docker compose exec api python -m app.maintenance.compact_request_logs [--dry-run] [--batch N]Rewrites old request rows in the compact form new rows use. Safe to interrupt and run again.
docker compose exec api python -m app.maintenance.compare_providers --project <book id> --providers <id>,<id> [--sample 5] [--output report.json]Compares providers on the same passages. See below.
python scripts/measure_prompt_cost.py [options]Measures the tokens a configuration sends per passage, without any real model. See below.
python scripts/benchmark_analysis.py [options]Wall time, calls and tokens of the analysis of a long synthetic serial, strict against parallel, for several thread counts (development).
python scripts/evaluate_analysis_modes.py [options]Memory quality of the analysis modes against a synthetic ground truth (development).
docker compose exec api alembic checkConfirms that the database schema matches the application.

Cost control

Each model call (translation, review, revision, polishing, final review) carries the same fixed context (instructions, glossary, Book Bible, character notes, neighbouring passages), whatever the length of the passage. Three things reduce what you pay:

  • Longer passages. PASSAGE_MAX_CHARS (3500 characters by default, 500 to 20000) sets the passage size of books imported afterwards. A volume can have its own (passage_max_chars in its configuration, applied to chapters added later), and so can an import. Longer passages share the fixed context between more text. Beyond 8000 to 10000 characters, the expected answer approaches many providers’ maximum output and a truncated answer costs more than it saves. Books already imported keep their cut.
  • Fused review. REVIEW_MODE=fused, or a volume’s review mode, reviews and corrects a passage in one call at high and maximum quality, instead of a review call followed by a revision call.
  • Prompt caching. Automatic: prompt sections go from the most stable to the most variable, so that providers with prefix caching can reuse the common start of successive calls.

Other levers: a lower quality level on books that do not need it, FINAL_REVIEW_ENABLED=false, and the per-book estimate shown before each launch.

To cap what a book or an integration may spend, give it a budget: a book’s own cap, the installation default (BUDGET_DEFAULT_BOOK or Settings › Budgets) or an API token’s monthly or total cap. Near the cap, a job moves to a cheaper fallback provider, or pauses until the cap is raised; see the user guide. Budgets only see calls made with a price: give every paid provider its input and output prices.

Measure a configuration

scripts/measure_prompt_cost.py runs the translation pipeline on a synthetic book against a simulated provider (no network, nothing billed) and reports calls and tokens per passage, the share a prefix cache could reuse, and for each operation where successive prompts start to differ. Run it from the repository with the backend’s development environment (development):

python scripts/measure_prompt_cost.py --quality high --passage-chars 3500 --review-issues 0.5
python scripts/measure_prompt_cost.py --review-mode fused --json

The figures compare settings of the same code; they do not predict the bill of a real book.

Compare providers

app.maintenance.compare_providers has several providers translate the same sample of a book’s passages (spread over its narrative chapters), with the prompt and context the pipeline would build, and without the response cache. Nothing is written to the book.

docker compose exec api python -m app.maintenance.compare_providers \
  --project <book id> --providers <provider id>,<provider id> --sample 5 --output comparison.json

The table lists, per provider, passages translated and failed, seconds per passage, tokens, cost and the findings of the automatic checks (locked glossary, markers, untranslated text…). The JSON report adds failure reasons, length ratios and the translations side by side. A relative --output path is written under DATA_DIR/tmp (/data/tmp in the container, since the application directory is read-only), and the command prints the final path; copy the file out with docker compose cp api:/data/tmp/comparison.json .. These calls are real and billed; they appear in the statistics under the operation provider_comparison.

Troubleshooting

SymptomCause and fix
A book stays queuedThe worker is not running (docker compose ps worker), or its provider has no free Concurrent books slot, or its account (or API token) already runs its quota of jobs. The Queue page gives the reason for each job.
Launching answers “Queue full” (HTTP 429 queue_full)The account or the API token already has its quota of waiting jobs (Settings › Queue, QUEUE_MAX_QUEUED_PER_ACCOUNT, or the token’s own limit). Wait until one starts, or raise the quota.
A book is waitingRead stop_reason and next_attempt first. A work window or daily cap defers it to an opening; analysis may wait for an earlier volume. For a provider outage, recovery starts after PROVIDER_RECOVERY_BASE_SECONDS (60 seconds by default), doubles up to an hour and respects Retry-After up to 24 hours. The autopilot eventually tries its configured fallback. This is not the same as a manual pause.
A series does not launch new workCheck its suspension banner. Resume the series clears persistent suspension; resuming selected jobs alone does not. Already sent provider calls may finish after suspension.
A job paused with stop_reason licenceThe licence refuses work that costs words: the quota of the cycle is spent (past LICENCE_QUOTA_OVERRUN_WORDS), or the licence is not activated, refused, expired, or held by another machine. Settings › Licence says which. Nothing resumes on its own: once the licence allows work again (a new cycle, a renewal, a larger quota), resume the book. See configuration.
A book is refused at the import with licence_quota_insufficient (or allowance_insufficient)Since 0.17.0 a book is counted when it is added: its words exceed what the quota cycle (or the account’s monthly allowance) has left, margin included. Nothing was imported or counted. Wait for the date of the next quota shown in Settings › Licence, or ask for a larger quota. A book added before 0.17.0 can still be refused before the queue for the same reason.
HTTP 402 automation_not_licensed or sharing_not_licensedThe plan does not include the automation API (tokens, /api/v1, /mcp) or book sharing (members, review links): only Studio and Pro do. Nothing is deleted: members already invited keep their access, and existing tokens and review links are suspended, then work again once the licence grants the right.
Activation answers “Every machine of this licence is taken, and the oldest one moved too recently to give up its place.”Every machine of the plan is taken and the oldest one took its place less than seven days ago (move_too_soon). Use Release this machine on the installation you leave, then activate again.
Activation or renewal refused with another reasonSettings › Licence gives the licence server’s reason in the interface’s language: unknown key, revoked or suspended licence, expired subscription, a signature the server does not recognise, a report received twice (the installation runs in two places), a clock that is off (set NTP), or too many attempts. A reason this version does not know yet is shown with its code.
An email did not arriveInspect Settings › Mail and its pending/sent/failed queue, then the worker. A queued test is not proof of SMTP acceptance; SMTP acceptance is not proof that the recipient’s inbox accepted it. Correct configuration before retrying eligible failures; request a fresh recovery email rather than replaying an old password-reset link.
A book is blockedThe provider rejected the credentials. Fix the key or sign in again, then resume. Under the autopilot, the next fallback provider takes over, or the job fails if none is left.
A book failed with “providers exhausted”Every provider in the autopilot chain was unavailable. Add a fallback provider in Settings › Autopilot, then resume.
Passages are refused or kept in the source languageSee autopilot for the recovery ladder and how to retranslate them with another provider.
libris_jobs_expired_leases is above 0The worker stopped abruptly or hangs: docker compose logs worker, then docker compose restart worker. Jobs resume from their checkpoints. On a host driven by the production scripts, librisctl doctor reports it and librisctl restart worker fixes it.
The live progress stops updatingToo many tabs open (limit EVENT_STREAMS_PER_USER), or a proxy buffering server-sent events.
An EPUB export is refusedRead the error message. Missing or refused passages must be resolved first, or use the partial export that keeps the source text. “Too many EPUB validations in progress” means EPUBCHECK_CONCURRENCY is reached: try again in a moment.
The worker logs status=pool_exhausted, or status=database_unavailable while the database is healthyThe connection pool ran out, not the database. Look for db_pool=undersized in the start-up log: it gives the pool’s capacity and the sum of the providers’ Concurrent books. Raise DB_POOL_SIZE (and max_connections on the database service if you go past 80 per process), or lower a provider’s capacity. Jobs are not lost: they go back to waiting and resume.
The database grows fastCheck the retention settings and run the retention pass with --dry-run.
Provider keys “must be entered again”SECRET_KEY changed. Restore the original .env, or enter each key again.

For installation, network and sign-in problems, see the Docker guide.