Operations
This page is for administrators who keep Libris running: updating it without losing work, sizing its capacity, monitoring it, keeping its database small, controlling model costs and solving common problems. Installation is covered in the Docker guide, every setting in the configuration reference, and backups in backup and restore.
All commands run from the Libris directory unless stated otherwise.
Update Libris
The simple customer path is the one in the Docker guide: back up, then run the registry
bootstrap again. It needs neither Git nor a source archive. It runs the migrations, then restarts the web application and the worker; running books resume
from their last checkpoint, so no finished work is lost, but the passages being translated at that moment are
sent again. The migrations run just before the worker is replaced, while the previous one may still be working:
pause the running books first, as the Docker guide says, or use scripts/deploy.sh.
For a development checkout, scripts/deploy.sh gives finer control over when the worker restarts. Unlike the registry bootstrap, it does not change
the image named in .env: set LIBRIS_IMAGE there to the new release first.
sed -i 's|^LIBRIS_IMAGE=.*|LIBRIS_IMAGE=registry.libris-translate.com/libris/libris:<new version>|' .env
LIBRIS_DEPLOY_SOURCE=pull ./scripts/deploy.sh --worker-when-idle
| Mode | What it does |
|---|---|
--api-only (default) | Gets the image, runs the migrations and restarts the web application only; the worker keeps running the previous version. Only for updates without a database migration: when the new version brings one and the worker is running, the script changes nothing and exits with status 5 (with the worker stopped, it migrates and leaves the worker stopped). |
--worker-when-idle | Gets the image, waits up to 10 minutes for the job queue to be empty, stops the web application so that no new job starts, checks again, stops the worker, runs the migrations and starts both. If jobs stay active it changes nothing and exits with status 3 (4 if a job started at the last moment). |
--force-worker | Gets the image, stops the web application and the worker at once, runs the migrations and starts both. Running jobs resume from their checkpoints. |
| Variable | Meaning |
|---|---|
LIBRIS_DEPLOY_SOURCE | build (default) builds the image from the source tree and tags it with the LIBRIS_IMAGE of .env: give it a local name (for example LIBRIS_IMAGE=libris:local), not an official registry tag; pull downloads LIBRIS_IMAGE; loaded uses an image already present locally (LIBRIS_IMAGE required). |
LIBRIS_COMPOSE_ENV_FILE | File passed to Compose with --env-file, for the values Compose itself reads (LIBRIS_IMAGE, PORT…). The containers still read LIBRIS_ENV_FILE (.env by default): set both when the configuration lives elsewhere. |
A wrong mode or LIBRIS_DEPLOY_SOURCE exits with status 2.
If the new version fails to start, the script puts the previous image back (tagged libris:rollback)
and restarts it. It cannot undo a migration: for that, see
roll back a failed update.
With deploy.sh, the migrations never run while a worker of the previous version is running. Use --api-only for a quick fix of
the web side, then finish with --worker-when-idle when books are done, so that the worker does not keep running
an older version for long.
Production deployment script
deploy/libris-production-deploy deploys a tagged release, already loaded as images, onto a dedicated host. The
project’s own pipeline no longer deploys anything (a tag only publishes images and a release, see
development): the script stays for an operator who drives a host from a pipeline of
their own, and is of no use to an installation made with the installer. Install it as
/usr/local/sbin/libris-production-deploy.
Host-specific settings can be kept in /etc/libris-production.conf (root-owned, mode 0600),
using deploy/libris-production.conf.example. This trusted shell file is sourced before the script
resolves its settings, including for rollback and pruning; its assignments override the environment.
It is not the application’s .env and must never be writable by an application account.
libris-production-deploy COMMIT VERSION APP_IMAGE_ID CODEX_IMAGE_ID # deploy images already loaded
libris-production-deploy --rollback # back to the images it replaced
libris-production-deploy --check-compose # is the installed Compose file the deployed version's?
libris-production-deploy --prune-images [--dry-run] # remove old Libris images
A deployment checks the images’ version and revision labels, refuses to go back to an older version (use
--rollback), dumps the database into $LIBRIS_PRODUCTION_BASE/backups/ (last five kept), stops the application,
installs the Compose file shipped inside the new image, runs the migrations, starts the API and the Codex bridge,
checks /health, and only then starts the worker. If anything fails before the schema changed, it restarts the
previous images; after a schema change it leaves the services stopped and names the dump to restore.
| Variable | Default | Meaning |
|---|---|---|
LIBRIS_PRODUCTION_CONFIG | /etc/libris-production.conf | Optional trusted host configuration file. A missing file leaves environment settings in use. |
LIBRIS_PRODUCTION_BASE | /opt/libris-production | Working directory: Compose file, recorded version, pre-deployment dumps. |
LIBRIS_PRODUCTION_COMPOSE_FILE | $LIBRIS_PRODUCTION_BASE/docker-compose.yml | Installed Compose file. |
LIBRIS_PRODUCTION_COMPOSE_OVERRIDE | empty | Optional extra Compose file for local additions. |
LIBRIS_PRODUCTION_SECRET_ENV | $LIBRIS_PRODUCTION_BASE/.env | The installation’s .env. Without this variable the scripts take $LIBRIS_PRODUCTION_BASE/.env, then, if that file does not exist, the legacy /opt/epub-translator/.env, and say so once per run until librisctl migrate-env --confirm moves it. |
LIBRIS_PRODUCTION_PROJECT | epub-translator | Compose project name. Without this variable the scripts read $LIBRIS_PRODUCTION_BASE/project when that file exists (librisctl migrate-project writes it), and fall back to epub-translator. |
LIBRIS_PRODUCTION_HEALTH_URL | http://127.0.0.1:8088/health | Health URL checked after deployment. Set it explicitly when the API is bound elsewhere. |
LIBRIS_PRODUCTION_REGISTRY_REPOSITORY | empty | Exact registry repository whose old pulls --prune-images removes. Empty disables registry-pull cleanup; unused local Libris image tags can still be removed. |
Change docker-compose.yml in the repository, never on the host: host-specific values belong in the .env
(BIND_ADDRESS, PORT…) or in the override file, and the replaced file is kept as
docker-compose.yml.before-<version>. When disk space is short, delete old files inside backups/, never the
directory: without its Compose file the next deployment stops with Production configuration is not provisioned (exit 65) before touching anything. Reinstall /usr/local/sbin/libris-production-deploy whenever
deploy/libris-production-deploy changes, and validate a deployment and a rollback on a disposable stack
before trusting a new host.
Each deployment also refreshes /usr/local/sbin/librisctl from the image it deploys (the image carries
deploy/librisctl in /app/deploy/), so the operator command line is always the one of the running version.
Operator command line
librisctl is the everyday command line of such a host: one command per task, with the Compose project, the
configuration file and the Codex profile resolved exactly as the deployment script resolves them. It is
installed as /usr/local/sbin/librisctl by each deployment, or by hand with
install -m 0755 deploy/librisctl /usr/local/sbin/librisctl. Run it with no argument for the list, or
librisctl <command> --help for one command. Its output is plain text without colour or pager, so it reads the
same over SSH, in a script and in a log.
| Command | What it does |
|---|---|
status | Deployed version and commit, image ids, every container with its state and health, /health, the size of the data volumes and their free space, the most recent dump. Changes nothing. |
ps | The installation’s containers, stopped ones included. |
logs [SERVICE...] [-f] [--since D] [--tail N] | Container logs without colour codes. |
journal [-f] [-n N] [FILTER...], journal export [FILTER...] [--output FILE] [--with-containers] | The journal: every service in one chronology, the last events, followed, or a period exported. |
restart [SERVICE...] | Restarts api, worker and codex by default. Running books resume from their checkpoints. A container keeps the environment it was created with, so a restart does not apply a changed .env: when a service runs with a value the configuration no longer says, the command names that service, the settings concerned and up. A container that was never given a changed setting is left alone, so a database started long before an ALLOWED_ORIGINS change is not reported. Values are never printed. |
start [SERVICE...], stop [SERVICE...] | Start or stop services. stop keeps every container and volume; stopping the database (by name or with --all) needs --confirm. |
up | Applies the installed Compose file with the deployed images: recreates what changed, never builds, never pulls. This is what applies a changed .env — a setting such as ALLOWED_ORIGINS only takes effect once the containers are recreated. |
health, version | Is the API answering with the deployed version (exit 1 if not); what is deployed. |
jobs | Read-only view of the queue from the database: jobs by state, live jobs with their priority and lease, expired leases. Ids only, never a book title. |
shell SERVICE [COMMAND...], psql [ARGUMENT...] | A shell, or psql as the installation’s own user, inside a container. Interactive only when the terminal is. |
backup | Dumps the database now into $LIBRIS_PRODUCTION_BASE/backups/manual-<timestamp>.dump (mode 600) and prints the path. The full scheduled backup stays libris-backup. |
restore DUMP --confirm [--no-start] | Checks the archive, stops application services, saves a private rescue dump, recreates the application database and restores it. Normally restarts everything, migrations included. Use --no-start before rollback so the newer image cannot immediately reapply its schema. Rescue dumps pre-restore-*.dump are not pruned automatically. |
rollback --confirm | Runs libris-production-deploy --rollback. |
prune-images [--dry-run] [--confirm] | Runs libris-production-deploy --prune-images. No volume is ever touched. |
check-compose | Runs libris-production-deploy --check-compose; exits 1 and prints the difference on drift. |
doctor | One pass over the whole host: configuration in place and private, containers running and healthy, /health, Compose drift, container hardening (read-only root, no-new-privileges, dropped capabilities, pids limit), free disk, a recent dump, the journal directory in place and private, the configuration the application expects (every mount of its Compose file, a journal that is written: a degraded one fails, an outdated one warns), and a worker whose leases are not expired. Exits 1 if any check fails. |
migrate-env [--confirm] | Moves a configuration file left at the legacy path under the production base, keeping its permissions. |
migrate-project [NAME] [--confirm] | Renames the Compose project (see below). |
Exit codes: 0 success, 1 a check failed, 64 wrong usage, 65 production is not provisioned on this host or
a precondition is not met, 77 refused (a destructive command without --confirm). librisctl down and friends
are refused outright: the data volumes are the only copy of the books and of the database.
The common tasks:
librisctl status # what is running, and is it healthy?
librisctl doctor # everything that should be true, checked in one pass
librisctl restart worker # the worker stopped taking jobs; books resume from their checkpoints
librisctl logs -f worker # follow it
librisctl logs --since 30m api worker
librisctl journal -f # every service in one chronology, as it happens
librisctl journal export --since 2h --with-containers --output /root/incident.log
librisctl backup # a database dump before touching anything
librisctl rollback --confirm # back to the version the last deployment replaced
Nothing prints the configuration or a key: status shows where the .env is, never what it holds.
Rename the Compose project
An installation made with the installer is renamed by the installer itself (Docker). On a production
host driven by libris-production-deploy, the historical project name is epub-translator, and its volumes hold the
books and the database, so it cannot simply be renamed. librisctl migrate-project [NAME] --confirm (libris by default) does it in one go: it
refuses while a job is active, stops the installation, lets Compose create the new project’s volumes, copies each
volume into its twin, compares their sizes, starts the installation under the new name and records it in
$LIBRIS_PRODUCTION_BASE/project, which every script then follows. The old volumes are left exactly as they
were; stateless service containers are recreated under the stable libris-* names. Removing
$LIBRIS_PRODUCTION_BASE/project and running librisctl up returns to the old volumes, and
docker volume rm epub-translator_database epub-translator_books epub-translator_codex-state removes them once
the new name has proved itself. Add LIBRIS_PROJECT=<name> to the .env as well if you also run plain
docker compose commands: the Compose file takes its project name from that variable.
The Compose options for a manual command on such a host, when librisctl is not enough:
docker compose --project-name epub-translator --env-file /opt/libris-production/.env \
--file /opt/libris-production/docker-compose.yml --profile codex <command>
Capacity and concurrency
Two limits decide how much work runs at once:
- Concurrent books, set on each provider in Settings › LLM providers: how many books may use that provider at the same time, all operations included (analysis, translation, review). Providers are independent: one set to 3 and another set to 1 allow four active books.
WORKER_BOOK_PARALLELISM: how many passages of one book are in flight at once, for the analysis (in the default parallel mode,ANALYSIS_MODE=parallel) as for the translation and the reviews. The default,0, uses the provider’s capacity, shared between the books using it.1processes one passage at a time. A volume (Passages worked on at once in its settings), a launch or an API request (threads) can only ask for fewer. After a provider answers429or is overloaded, the job resumes at half its width and widens again by one passage per minute.
The strict analysis (ANALYSIS_MODE=strict) reads one passage after the other whatever these limits; the
parallel mode makes about 1.8 times more analysis calls but ends several times sooner on a long book (see
architecture).
When books wait, they do not start oldest first: the fair queue takes the highest
priority, then the account and the API token with the fewest jobs running, in turn between accounts, so one
person’s backlog does not hold everyone else back. On a shared installation, Settings › Queue (or QUEUE_*, see
configuration) can also cap each account’s running and waiting jobs. The Queue
page shows each waiting job’s place and what holds it.
A book stays queued while its provider has no free slot. To change a book’s provider mid-way, pause it, choose the new provider in its configuration and resume: the rest of the book uses the new one.
One worker is enough for most installations. It renews a 60-second lease on each job; if the worker stops abruptly, another start picks the job up after the lease expires.
Worker processes
A worker process runs on one CPU core, whatever the number of books it carries. On a machine with spare cores,
WORKER_PROCESSES=N (2 to 16; default 1, the single process described above) makes the worker container run
the books in N job processes:
- The worker’s main process keeps everything that must run once: licence renewal, watched sources, retention,
mail, webhooks, library publication, the memory catalog and OpenViking. It starts the N job processes, starts
one again when it dies, and on
docker compose stoppasses the signal on and waits for them (25 seconds, inside the service’s 30-secondstop_grace_period). - Each job process takes books from the same queue. Leases, the fair queue, a provider’s Concurrent books and “one volume of a series at a time” are kept in the database and hold across processes. A process leaves the next book to the others once it holds more than its share (books under way divided by N, rounded up), so the books spread over the processes.
- If a job process dies, its books are picked up by another one when their 60-second lease expires; the log says
process=died … role=jobsthenprocess=started, andstatus=reclaimedfor each book taken over.
Limits to know before raising it:
- One book alone stays on one core: its passages are threads of one process. The setting helps when several books run at once, not a single long one.
- One worker container only. Do not scale the
workerservice or start a second one: the loops kept in the main process (licence, watched sources, retention) are not protected against running twice. - Every process has its own database pool: raise
max_connectionswith the number of processes (see the pool settings in.env.example). - The worker service’s
pids_limit: 512covers up to 8 job processes (a process uses at most about 45 threads); raise it in a Compose override beyond that. - A demo instance (
DEMO_MODE) starts no job process.
The worker also sends memory updates to OpenViking and, when the cleanup is on, removes the OpenViking documents of deleted volumes and series. A removal holds a 5-minute lease too; a waiting one is visible, with its last error, in Settings › Memory · OpenViking › Cleanup log.
Database connections
Every process holds a connection pool of its own, DB_POOL_SIZE + DB_POOL_MAX_OVERFLOW connections (80 by
default). PostgreSQL must accept all of them at once:
max_connections >= (API process + worker processes) x (DB_POOL_SIZE + DB_POOL_MAX_OVERFLOW) + 20
With WORKER_PROCESSES=1 there is one worker process. Above 1 there are N children and the parent that
supervises them, so N + 1. The 20 are left for librisctl, a psql and the database’s own maintenance.
The supplied Compose file reads max_connections from POSTGRES_MAX_CONNECTIONS in .env (200 by default,
enough for the defaults with one worker: (1 + 1) x 80 + 20 = 180).
WORKER_PROCESSES | Worker processes | Budget with the default pool of 80 | POSTGRES_MAX_CONNECTIONS |
|---|---|---|---|
| 1 (default) | 1 | (1 + 1) x 80 + 20 = 180 | 200 (default) |
| 2 | 3 | (1 + 3) x 80 + 20 = 340 | 340 |
| 4 | 5 | (1 + 5) x 80 + 20 = 500 | 500 |
Both processes check it at start-up. When max_connections is below the budget they log
db_pool=over_budget max_connections=… needed=… hint=set POSTGRES_MAX_CONNECTIONS=… and start anyway: the
failure would then show as pool_exhausted or refused connections under load. Fix it by raising
POSTGRES_MAX_CONNECTIONS (and restarting the database service), or by lowering DB_POOL_SIZE and
DB_POOL_MAX_OVERFLOW — the pool only needs the sum of the providers’ Concurrent books plus 16.
Monitoring
Health
GET /health answers {"status":"ok","version":"<version>","configuration":"ok"} when the API and its database
connection work. It does not test model providers; test those in Settings.
configuration compares the containers with what the image’s own Compose file (/app/deploy/docker-compose.yml)
expects: every mount it gives the application (/data, /etc/machine-id, /logs) and a journal that is written.
ok; outdated when a mount is missing (the Compose file predates the version); degraded when a feature is off
because of it — a journal that cannot be written. The API and the worker say it at start-up
(configuration=… check=… in their output). The list of checks, each with the command that fixes it, is in the
answer for a signed-in administrator, or for everyone with PUBLIC_HEALTH_DETAILS=true; administrators also see it
in Settings › Installation health, and librisctl doctor prints it.
With the same details, openviking is the state the worker last saw (reachable, unreachable, unknown or
not_configured), since when, and the send queue (pending, error, next_attempt); never its address.
OpenViking is optional: an outage there never changes status nor configuration. See
OpenViking.
curl --fail http://127.0.0.1:8088/health
Prometheus metrics
GET /metrics exposes the state of the installation in Prometheus text format. It is disabled until you set
METRICS_TOKEN:
sed -i "s|^METRICS_TOKEN=.*|METRICS_TOKEN=$(openssl rand -hex 32)|" .env # the line exists in a generated .env
grep '^METRICS_TOKEN=' .env # the value to give Prometheus
docker compose up -d --no-build --wait
Each scrape must send the token as Authorization: Bearer <token>; a browser session is not enough. Without the
variable the endpoint answers 404, with a wrong token 401.
scrape_configs:
- job_name: libris
scrape_interval: 60s
metrics_path: /metrics
scheme: https
authorization:
type: Bearer
credentials_file: /etc/prometheus/libris-metrics-token
static_configs:
- targets: ["books.example.com"]
| Metric | Type | Content |
|---|---|---|
libris_jobs{operation,status} | gauge | Jobs by operation and state, finished ones included |
libris_jobs_oldest_queued_age_seconds | gauge | How long the oldest job ready to run has been waiting (a resumed job counts from its resumption) |
libris_jobs_expired_leases | gauge | Running jobs whose 60-second lease expired: their worker stopped or hangs |
libris_llm_requests_total{operation,status} | counter | Finished model requests by operation and outcome (success, error, refused, interrupted, abandoned); cached answers count as success |
libris_llm_input_tokens_total{operation}, libris_llm_output_tokens_total{operation} | counter | Tokens reported by the providers |
libris_llm_wasted_input_tokens_total{operation} | counter | Input tokens of requests that failed, were refused or interrupted |
libris_llm_cache_hits_total{operation} | counter | Requests answered from the response cache |
libris_llm_cache_hit_ratio | gauge | Share of all finished requests answered from the cache |
libris_llm_requests_in_flight{provider} | gauge | Requests in progress, by provider name |
libris_segments{status} | gauge | Passages of all books, by state |
libris_memory_outbox_pending | gauge | OpenViking updates not delivered yet |
Counters are computed from the database: they count everything since installation, are the same in every process
and survive restarts. Deleting a book deletes its requests, which Prometheus sees as a counter reset. Use
rate() or increase() for a time window. No label contains a book title, text, id or provider address. The
output is cached for 10 seconds, so scraping every 30 to 60 seconds is enough.
Useful alerts:
| Condition | Meaning |
|---|---|
libris_jobs_expired_leases > 0 for 5 minutes | The worker is stopped or stuck. |
libris_jobs_oldest_queued_age_seconds > 900 | No worker takes jobs, or a provider is saturated. |
sum(rate(libris_llm_wasted_input_tokens_total[1h])) / sum(rate(libris_llm_input_tokens_total[1h])) > 0.2 | More than one input token in five is spent on requests that produced nothing. |
Behind a reverse proxy, expose /metrics only to your Prometheus network if you can.
Statistics in the interface
Statistics shows token usage by model, and each book shows what it has spent. Costs use the price recorded with each request, so a later price change on a provider does not rewrite past costs.
Logs
docker compose logs --since=30m api worker
docker compose logs -f worker
Each container keeps at most three log files of 10 MB. Review logs before sharing them: remove account names, addresses and anything that looks like a key.
Journal
To reconstruct an incident, read the journal instead: one chronology of the API, the worker, the migrations and the deployments, in files that survive a container’s recreation.
Where. Each process writes its own file in the journal directory: api.log, worker.log, migrate.log,
and deploy.log for the production deployment script. A second process of the same service (several API
workers, a docker compose run) writes api-<pid>-<random>.log instead, so no two processes ever append to, or
rotate, the same file. On a host deployed by libris-production-deploy the directory is
/opt/libris-production/logs ($LIBRIS_PRODUCTION_BASE/logs, or LIBRIS_PRODUCTION_LOG_DIR); elsewhere it is
the logs volume of the Compose project (docker volume inspect libris_logs), or the host directory
named by LIBRIS_LOG_DIR in the .env — create it first, owned by uid 10001 and mode 0700, or Docker creates
it for root and the journal stays off.
What a line says. One line per event:
2026-09-24T08:15:03.123+00:00 ERROR worker pid=7 req=- job=6f0c1e2a-… epub.worker | job=6f0c1e2a-… status=failed …
Traceback (most recent call last):
File "/app/backend/app/jobs/worker.py", line 290, in _execute
…
The time with its UTC offset (LOG_TIMEZONE), the level, the service, the process, the request (req=, the
X-Request-ID every API response carries, taken from the client when it sends a well-formed one), the job
(job=; deploy-<time>-<pid> for the steps of one deployment), the logger and the message. Starts, readiness
and stops of each process, each API request (method, path, status, duration; /health and static files
excepted), each job’s start (resumed=yes when it picks up a checkpoint), completion, pause and failure, the
migrations run, and the steps of a deployment are all there, next to what the code already reported
(status=database_unavailable, db_pool=…, licence=…).
A traceback follows its event on indented lines. It keeps every frame and the chain of causes, and the message of the exceptions whose message is a technical cause (a refused connection, a missing file, a TLS or HTTP error); for the others — which may quote a passage, a prompt or an account — and for database drivers, which repeat rejected values, only the type is kept (the SQLSTATE code for the latter).
What is never written. Passwords, cookies, Authorization headers, API tokens (their public lbr_… prefix
is kept), licence keys, registry and other access tokens, private keys, credentials in URLs, and every value of
the configuration whose name says it is a secret (SECRET_KEY, POSTGRES_PASSWORD, *_TOKEN…) are replaced
by ***. Line breaks and control characters in a message are escaped (\n), so a file name or a header
cannot forge an event, and a message is cut after 4,000 characters, so no chapter is ever copied whole. The
same protections apply to the container output. An export still names accounts, books and addresses: review
it before sharing it.
Reading it. On a production host:
librisctl journal # the last 200 events of every service
librisctl journal -f --service api --service worker
librisctl journal --level WARNING --since 2h
librisctl journal --request 3f9c0a1b2c3d4e5f # everything one request did
librisctl journal --job 6f0c1e2a-… # everything one job did
librisctl journal export --since 2026-09-24T07:00 --until 2026-09-24T09:00 --with-containers \
--output /root/incident.log # a period, with the database and Codex output merged in
librisctl journal reads the files with the deployed image, in a container without network that mounts them
read-only: it works while the installation is stopped. Without the production scripts, run the same tool in
a container of the installation (python -m app.journal --help lists the options):
docker compose exec api python -m app.journal -n 100
docker compose exec api python -m app.journal follow
docker compose exec api python -m app.journal export --since 1d > incident.log
Size and permissions. Each file is rotated at LOG_MAX_MB (10 MB) and LOG_BACKUPS (5) rotated files are
kept per process; files older than LOG_RETENTION_DAYS (14) go, but never the one a running process writes.
The directory is mode 0700 and the files 0600, owned by the application’s user (uid 10001); root reads them on
the host. See configuration.
When it cannot be written. A service never stops for its journal. If the directory cannot be opened at
start-up (journal: disabled for api, … is not writable on the container output), or a write fails later (a
full disk: journal: cannot write … events still go to this output, at most once every five minutes, with the
count of the events missing from the file), the events keep going to the container output and the file is
tried again at the next event. A deployment that cannot write deploy.log says so and goes on.
Data retention
Every model call leaves a row with its prompt and answer, which the request inspector shows. Without limits, this
history grows quickly. The worker cleans it up once at start-up and then every hour, in small batches, following
the RETENTION_* settings (configuration):
- prompts, raw answers and context traces of requests older than 30 days are emptied (the inspector no longer shows them), while tokens, cost, duration, status and the cached answer are kept;
- progress events older than 7 days are deleted, always keeping the last 500 of each book;
- OpenViking updates already delivered are deleted after 7 days;
- only the 20 most recent automatic Book Bible revisions of each book are kept (human revisions are all kept);
- the per-passage resume state of jobs that ended more than 30 days ago is deleted. If such a job is resumed later, its finished passages are still not translated again (unless you force a new translation);
- result files of automation requests are deleted after 30 days; asking for the result again rebuilds it;
- files of expired imports are deleted from
DATA_DIR/staging.
Before deleting anything, the same pass adds requests older than two hours to daily usage totals, which the
statistics and /metrics read. With RETENTION_REQUEST_ROWS_DAYS set (for example 180), whole request rows
older than that are then deleted; statistics stay correct, but the response cache and the inspector lose them.
See what a pass would remove, or run one now:
docker compose exec api python -m app.maintenance.retention --dry-run
docker compose exec api python -m app.maintenance.retention
PostgreSQL reuses the freed space, but gives it back to the system only after a full vacuum. It locks the table, so stop the worker first and make sure the disk has as much free space as the table’s useful size:
docker compose stop worker
docker compose exec database sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB" -c "VACUUM (FULL, ANALYZE) llm_requests;"'
docker compose start worker
Maintenance commands
| Command | What it does |
|---|---|
docker compose exec api python -m app.maintenance.retention [--dry-run] | Runs the retention pass described above. |
docker compose exec api python -m app.maintenance.usage [--dry-run] | Adds finished requests to the daily usage totals now. The worker does it every hour. |
docker compose exec api python -m app.maintenance.compact_request_logs [--dry-run] [--batch N] | Rewrites old request rows in the compact form new rows use. Safe to interrupt and run again. |
docker compose exec api python -m app.maintenance.compare_providers --project <book id> --providers <id>,<id> [--sample 5] [--output report.json] | Compares providers on the same passages. See below. |
python scripts/measure_prompt_cost.py [options] | Measures the tokens a configuration sends per passage, without any real model. See below. |
python scripts/benchmark_analysis.py [options] | Wall time, calls and tokens of the analysis of a long synthetic serial, strict against parallel, for several thread counts (development). |
python scripts/evaluate_analysis_modes.py [options] | Memory quality of the analysis modes against a synthetic ground truth (development). |
docker compose exec api alembic check | Confirms that the database schema matches the application. |
Cost control
Each model call (translation, review, revision, polishing, final review) carries the same fixed context (instructions, glossary, Book Bible, character notes, neighbouring passages), whatever the length of the passage. Three things reduce what you pay:
- Longer passages.
PASSAGE_MAX_CHARS(3500 characters by default, 500 to 20000) sets the passage size of books imported afterwards. A volume can have its own (passage_max_charsin its configuration, applied to chapters added later), and so can an import. Longer passages share the fixed context between more text. Beyond 8000 to 10000 characters, the expected answer approaches many providers’ maximum output and a truncated answer costs more than it saves. Books already imported keep their cut. - Fused review.
REVIEW_MODE=fused, or a volume’s review mode, reviews and corrects a passage in one call at high and maximum quality, instead of a review call followed by a revision call. - Prompt caching. Automatic: prompt sections go from the most stable to the most variable, so that providers with prefix caching can reuse the common start of successive calls.
Other levers: a lower quality level on books that do not need it, FINAL_REVIEW_ENABLED=false, and the per-book
estimate shown before each launch.
To cap what a book or an integration may spend, give it a budget: a book’s own cap, the installation default
(BUDGET_DEFAULT_BOOK or Settings › Budgets) or an API token’s monthly or total cap. Near the cap, a job moves
to a cheaper fallback provider, or pauses until the cap is raised; see the
user guide. Budgets only see calls made with a price: give every paid
provider its input and output prices.
Measure a configuration
scripts/measure_prompt_cost.py runs the translation pipeline on a synthetic book against a simulated provider (no
network, nothing billed) and reports calls and tokens per passage, the share a prefix cache could reuse, and for
each operation where successive prompts start to differ. Run it from the repository with the backend’s
development environment (development):
python scripts/measure_prompt_cost.py --quality high --passage-chars 3500 --review-issues 0.5
python scripts/measure_prompt_cost.py --review-mode fused --json
The figures compare settings of the same code; they do not predict the bill of a real book.
Compare providers
app.maintenance.compare_providers has several providers translate the same sample of a book’s passages (spread
over its narrative chapters), with the prompt and context the pipeline would build, and without the response
cache. Nothing is written to the book.
docker compose exec api python -m app.maintenance.compare_providers \
--project <book id> --providers <provider id>,<provider id> --sample 5 --output comparison.json
The table lists, per provider, passages translated and failed, seconds per passage, tokens, cost and the findings
of the automatic checks (locked glossary, markers, untranslated text…). The JSON report adds failure reasons,
length ratios and the translations side by side. A relative --output path is written under DATA_DIR/tmp
(/data/tmp in the container, since the application directory is read-only), and the command prints the final path;
copy the file out with docker compose cp api:/data/tmp/comparison.json .. These calls are real and billed; they appear in the
statistics under the operation provider_comparison.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| A book stays queued | The worker is not running (docker compose ps worker), or its provider has no free Concurrent books slot, or its account (or API token) already runs its quota of jobs. The Queue page gives the reason for each job. |
Launching answers “Queue full” (HTTP 429 queue_full) | The account or the API token already has its quota of waiting jobs (Settings › Queue, QUEUE_MAX_QUEUED_PER_ACCOUNT, or the token’s own limit). Wait until one starts, or raise the quota. |
| A book is waiting | Read stop_reason and next_attempt first. A work window or daily cap defers it to an opening; analysis may wait for an earlier volume. For a provider outage, recovery starts after PROVIDER_RECOVERY_BASE_SECONDS (60 seconds by default), doubles up to an hour and respects Retry-After up to 24 hours. The autopilot eventually tries its configured fallback. This is not the same as a manual pause. |
| A series does not launch new work | Check its suspension banner. Resume the series clears persistent suspension; resuming selected jobs alone does not. Already sent provider calls may finish after suspension. |
A job paused with stop_reason licence | The licence refuses work that costs words: the quota of the cycle is spent (past LICENCE_QUOTA_OVERRUN_WORDS), or the licence is not activated, refused, expired, or held by another machine. Settings › Licence says which. Nothing resumes on its own: once the licence allows work again (a new cycle, a renewal, a larger quota), resume the book. See configuration. |
A book is refused at the import with licence_quota_insufficient (or allowance_insufficient) | Since 0.17.0 a book is counted when it is added: its words exceed what the quota cycle (or the account’s monthly allowance) has left, margin included. Nothing was imported or counted. Wait for the date of the next quota shown in Settings › Licence, or ask for a larger quota. A book added before 0.17.0 can still be refused before the queue for the same reason. |
HTTP 402 automation_not_licensed or sharing_not_licensed | The plan does not include the automation API (tokens, /api/v1, /mcp) or book sharing (members, review links): only Studio and Pro do. Nothing is deleted: members already invited keep their access, and existing tokens and review links are suspended, then work again once the licence grants the right. |
| Activation answers “Every machine of this licence is taken, and the oldest one moved too recently to give up its place.” | Every machine of the plan is taken and the oldest one took its place less than seven days ago (move_too_soon). Use Release this machine on the installation you leave, then activate again. |
| Activation or renewal refused with another reason | Settings › Licence gives the licence server’s reason in the interface’s language: unknown key, revoked or suspended licence, expired subscription, a signature the server does not recognise, a report received twice (the installation runs in two places), a clock that is off (set NTP), or too many attempts. A reason this version does not know yet is shown with its code. |
| An email did not arrive | Inspect Settings › Mail and its pending/sent/failed queue, then the worker. A queued test is not proof of SMTP acceptance; SMTP acceptance is not proof that the recipient’s inbox accepted it. Correct configuration before retrying eligible failures; request a fresh recovery email rather than replaying an old password-reset link. |
| A book is blocked | The provider rejected the credentials. Fix the key or sign in again, then resume. Under the autopilot, the next fallback provider takes over, or the job fails if none is left. |
| A book failed with “providers exhausted” | Every provider in the autopilot chain was unavailable. Add a fallback provider in Settings › Autopilot, then resume. |
| Passages are refused or kept in the source language | See autopilot for the recovery ladder and how to retranslate them with another provider. |
libris_jobs_expired_leases is above 0 | The worker stopped abruptly or hangs: docker compose logs worker, then docker compose restart worker. Jobs resume from their checkpoints. On a host driven by the production scripts, librisctl doctor reports it and librisctl restart worker fixes it. |
| The live progress stops updating | Too many tabs open (limit EVENT_STREAMS_PER_USER), or a proxy buffering server-sent events. |
| An EPUB export is refused | Read the error message. Missing or refused passages must be resolved first, or use the partial export that keeps the source text. “Too many EPUB validations in progress” means EPUBCHECK_CONCURRENCY is reached: try again in a moment. |
The worker logs status=pool_exhausted, or status=database_unavailable while the database is healthy | The connection pool ran out, not the database. Look for db_pool=undersized in the start-up log: it gives the pool’s capacity and the sum of the providers’ Concurrent books. Raise DB_POOL_SIZE (and max_connections on the database service if you go past 80 per process), or lower a provider’s capacity. Jobs are not lost: they go back to waiting and resume. |
| The database grows fast | Check the retention settings and run the retention pass with --dry-run. |
| Provider keys “must be entered again” | SECRET_KEY changed. Restore the original .env, or enter each key again. |
For installation, network and sign-in problems, see the Docker guide.