OpenViking memory
This page is for administrators who want to connect Libris to an OpenViking server as an external, semantic memory for their books. It explains what OpenViking adds, how to set it up, what Libris writes there, what a passage is allowed to read back, and how to maintain it.
OpenViking is optional. Libris keeps every memory in its own database (PostgreSQL in production),
and the default internal backend reads it directly. OpenViking is an extra semantic index of the same
memories:
- everything written to it can be rebuilt from the database at any time;
- a failure of OpenViking never removes or undoes a result stored in Libris. With the
hybridoropenvikingbackend, translation simply continues with the database’s memory, and the context inspector says so. The same holds when OpenViking cannot be searched: its stored key became unreadable (theSECRET_KEYchanged; the inspector’s retrieval error asks to save the key again) or Semantic search is switched off.
Set it up
-
Run an OpenViking server that Libris can reach on your private network. Keep it private and maintained separately.
-
In Settings › Memory · OpenViking (administrators), fill in:
Field Meaning OpenViking URL Base URL of the server. Dedicated viking://rootA dedicated sub-directory under viking://resources/(defaultviking://resources/epub-translator). The root itself is refused.API key Sent as X-API-Key. Stored encrypted; never shown again.Authentication API key (recommended), or trusted mode, which also sends the account and user headers you provide ( X-OpenViking-Account,X-OpenViking-User).Context budget, retrieval budget The ceiling of the optional context of every call, whatever the memory backend (default 32000 estimated tokens: a call takes three quarters of what its provider’s window leaves, up to this ceiling; the former default 12000 reads as 32000), and how much of it may come from OpenViking (default 6000). See the budget. Minimum score Relevance threshold of search hits (default 0.15). Timeout Seconds per call (default 20). Semantic search (find), deep search (search) Which OpenViking search modes Libris may use. Test connection checks that the server answers, that the key is accepted and that the root can be searched. The first three settings can also come from
OPENVIKING_URL,OPENVIKING_API_KEYandOPENVIKING_ROOT_URI; values saved in the interface win (see configuration). -
Choose the memory backend per volume (the book’s Settings) or as a series default:
Backend Behaviour internalThe database only. Nothing is written to OpenViking. openvikingMemories are mirrored to OpenViking and retrieved from it. hybridBoth: retrieval uses OpenViking when it answers and the database otherwise.
The automation API accepts the same values in pipeline.context_backend (API guide).
What Libris writes
Libris writes with idempotent replace operations on stable URIs, without waiting for indexing. Every
identifier in a URI is a database id chosen by Libris, never a title or a path taken from a book:
<root>/<owner_id>/series/<series_id>/volumes/<project_id>/events/<memory_id>.json event, volume of a series
<root>/<owner_id>/series/<series_id>/volumes/<project_id>/book.md … catalog of that volume
<root>/<owner_id>/standalone/<project_id>/events/<memory_id>.json event, standalone volume
<root>/<owner_id>/standalone/<project_id>/book.md … catalog of that volume
Events
An event is one memory of the book: the analysis of a passage, its narrative state, or a validated human decision. Its document is computed from the database at the moment it is written, so OpenViking always holds the current event at its current place.
| Field | Meaning |
|---|---|
schema_version | 2 |
owner_id, series_id, project_id, volume_number | Where the memory belongs. |
chapter_id, chapter_position, chapter_number | Its chapter, its order in the volume, and the author’s number. |
segment_id, position | The passage and its narrative position in the volume. |
type | analysis, narrative or human_decision. |
identities | Characters named by the memory (canonical names and who knows them). |
validated | A person validated it. |
created_at | When the memory was created. |
content | The memory itself. |
The send queue only carries the work of writing a document; failed writes are retried with a growing delay (about half an hour at most).
Catalog documents
For each volume that has an analysis or a Book Bible, Libris also publishes a small named catalog,
refreshed every MEMORY_CATALOG_INTERVAL_SECONDS (60) when something changed:
| Document | Content |
|---|---|
book.md | Title, author, languages, analysis progress and links to the other documents. |
book-bible.json | The Book Bible (editorial synthesis), without the characters. |
characters.json and characters-NNNN.json | Identities, aliases, role, description, gender, pronouns, speech style, and whether a person confirmed them. |
relationships.json and relationships-NNNN.json | Relations between characters, with evidence, provenance and validation. |
These are compact projections; the database keeps the complete and exact data. Catalog documents are never used as narrative evidence (see below).
What a passage may read
A passage of volume N at position P searches with target_uri set to its series’ volumes
directory, or to its own events directory for a standalone volume. That only narrows the search on
the server. Libris then keeps a hit only if:
- its URI is exactly the URI of an event the database admits for this passage: a memory of the same volume at an earlier position (a person’s validated analysis of the passage itself included), or a memory of an earlier volume of the same series, owner and language pair (volume number lower than N);
- it is not a superseded human analysis, nor a human decision on a passage edited since;
- the document read back is identical to the event computed from the database.
Everything else is rejected and listed in the context inspector with its reason
(outside_narrative_allowlist, low_relevance, differs_from_canonical_event,
invalid_structured_memory, retrieval_budget): later passages, later volumes, directory summaries,
catalogs, another owner’s documents, stale or tampered documents. The Book Bible and character sheets
are editorial knowledge built from a whole reading, which is why they are never admitted as evidence
of what a passage may know. Continuous webnovel flows and unnumbered volumes have no “earlier volume”:
they rely on their own chapters.
State of OpenViking and of the send queue
In Settings › Memory · OpenViking, the OpenViking status card shows:
- whether OpenViking is reachable: the worker asks its
/healthat most once a minute; when it is unreachable, since when and the cause (ConnectError,HTTP 503…); - the events waiting to be sent and the events in error (volumes that use
openvikingorhybridonly), the last error, the most frequent errors, and the next planned attempt; - Retry the failed events: every failed event is sent at the worker’s next turn instead of after its backoff delay.
When OpenViking goes from reachable to unreachable, or back, the worker writes it once to the journal
(openviking=unreachable previous=reachable reason=ConnectError, openviking=reachable previous=unreachable requeued=140). When it is reachable again, the failed events are queued again at once, without waiting
for their delay.
The same state is the openviking element of GET /health, for a signed-in administrator or for
everyone with PUBLIC_HEALTH_DETAILS=true (see operations). Through the API
(administrators): GET /api/settings/memory/status and POST /api/settings/memory/outbox/retry.
Maintenance
On a series page, the Memory tab (Series OpenViking memory) shows, per volume, the events waiting, written and failed, with the last errors, and offers:
- Resynchronize: retry now everything waiting, without waiting for the backoff delay;
- Rebuild: rewrite every event and catalog of the series from the database, including memories whose queue rows the retention already removed;
- Reindex: ask OpenViking to recompute its index of the series directory.
On a volume, the Book Bible tab has an OpenViking memory panel with the queue state, the root, links that open the documents OpenViking actually holds (read through Libris, without exposing the key), and these actions:
- Synchronize the book and graph: queue the catalog again (works even during an analysis);
- Check in OpenViking: read the catalog documents back and compare them with what Libris wrote, and look for them in the index (this confirms the documents found, not the indexing of every event);
- Reindex in OpenViking and Rebuild from the database, as for a series.
The same actions are available through the interface’s API: GET /api/series/{id}/memory,
POST /api/series/{id}/memory/{resync|rebuild|reindex}, GET /api/projects/{id}/memory/status,
POST /api/projects/{id}/memory/{synchronize|rebuild|reindex|check}.
When a volume moves into or out of a series, or is renumbered, its events and catalog are written
again at the new place automatically; until then, its remote hits are rejected and translation relies
on the database. The same automatic rewrite moves volumes written with the older
<root>/<owner_id>/<project_id>/… layout: no manual step is needed after an upgrade.
Cleanup of deleted volumes and series
By default Libris never deletes anything in OpenViking. Deleting a volume or a series, or moving a volume, leaves the old documents in place; the database check simply no longer admits them.
An administrator can switch on the cleanup in Settings › Memory · OpenViking, card OpenViking
cleanup: tick Remove OpenViking documents on deletion, then Save the cleanup (or set
OPENVIKING_CLEANUP_ON_DELETE=true; a value saved in the interface wins until Go back to the environment
value). Nothing is removed while the OpenViking URL is empty. When it is on:
- deleting a volume queues the removal of its directories: its current place
(
…/series/<series_id>/volumes/<project_id>or…/standalone/<project_id>), the older<root>/<owner_id>/<project_id>place, and any other place the send queue shows it was written to (a volume that changed series); - deleting a series (possible once it has no volume left) queues the removal of
<root>/<owner_id>/series/<series_id>as a whole.
The removal is queued in the same transaction as the deletion and done by the worker, never by the
request: the deletion is immediate even when OpenViking is down. The worker waits about 30 seconds (so a
write already on its way lands first), then, for each directory, lists its documents and removes it
recursively. A failure is retried with a growing delay (one hour at most); a restarted worker resumes
where the previous one stopped. A cleanup queued while the switch was on still runs if it is turned
off later. If the OpenViking URL is emptied, waiting cleanups wait; if the root changes, they stay
waiting with the error The OpenViking root changed since the deletion and remove nothing.
The card now counts these holds separately from retryable failures, across the entire queue (not
only the visible log page). Each pending entry exposes its original root and a blocked_reason:
root_changed, not_configured, retry_error, or an empty string.
An administrator can Retire this cleanup after confirmation. No remote request is sent, the directories and earlier results are untouched, and the row remains in the log as Cleanup retired. This transition is refused once a worker has claimed the cleanup, even if its lease expired: recovery must finish first. Retrying cannot reactivate a retired row. Libris deliberately does not retarget an old deletion into a new root: restore the original configuration, or retire it and run a fresh dry scan of the new root before selecting anything to remove.
What a cleanup may remove is bounded twice:
- only a whole item directory directly under the configured root: a series
(
<owner_id>/series/<series_id>), a volume (<owner_id>/series/<series_id>/volumes/<project_id>,<owner_id>/standalone/<project_id>, or<owner_id>/<project_id>), with database identifiers only. The root, an owner’s directory and anything else are refused; - at the moment of the removal, the database must confirm that nothing lives there any more. A directory that is in use again is kept and logged as such.
Orphans
Documents of items deleted before the cleanup was switched on remain. In the same card, Look for
orphans (dry run) lists the directories under the root that the database no longer owns, with the
reason (volume_deleted, volume_moved, series_deleted, legacy_layout for the older layout), and
removes nothing. Remove these N directories then queues their removal, after a confirmation; each one is
checked against the database again, both when it is queued and when it is removed. Directories whose
name is not a Libris identifier are ignored. The dry run and the removal work whether the automatic
cleanup is on or off.
Scan for orphans daily, without deleting anything is a separate, off-by-default setting. The worker stores the date, root, count and first 500 results in the database; the card displays the last date and count, or a failure. A restart keeps the next attempt, and multiple workers share the same claim. The scheduled scan runs at most once every 24 hours, takes at most 30 seconds, and never queues a deletion. Disabling the option during a scan wins over its completion. An interrupted scan waits until its next daily attempt; a manual dry run is always available sooner.
Prune the orphan vectors of the Libris root daily is another, off-by-default setting of the same
daily turn. OpenViking can keep vector records whose document no longer exists (after the incident of
2026-09-24, 1.9 million records for about 6,000 indexed documents). With OpenViking 0.4.21 or later, the
worker calls POST /api/v1/content/reindex with {"uri": <root>, "mode": "prune_orphans", "wait": true, "dry_run": …} on the configured root only. Dry run: count without removing anything is ticked by
default: untick it only after reading a dry run. The result (the counters OpenViking returns:
scanned_records, would_delete_records in a dry run, deleted_records otherwise, failed_records, duration_ms,
and the number of warnings) is shown in the card and written to the journal
(openviking_prune status=done dry_run=True … scanned_records=… would_delete_records=…). The call waits at
most 5 minutes; a failure is tried again the next day. A server that does not know the mode (HTTP 400 or
422) is noted once as too old and not asked again until the setting is saved again (after an upgrade).
Through the API: PUT /api/settings/memory/orphans/schedule with {"prune": true} and
{"prune_dry_run": false}; only the fields sent change.
Both manual and scheduled inventories refuse a scope exceeding 500 directory-list requests, 1,000 entries in a single directory, or 10,000 candidate item directories. An incomplete inventory is never reported as “no orphans”. Reduce the root scope when this limit is reached. After a scheduled report, use a fresh manual dry run and its explicit deletion confirmation; an old report is not a deletion order.
Log
The Cleanup log of the card lists the last cleanups: what was deleted (volume, series or orphans, with its title), the state (waiting, running, done, retired), the attempts and the last error, and per directory whether it was removed (with the number and the first 100 names of its documents), already absent, or kept (Directories in detail). Retry now skips the waiting delay. Everything removed can be written again from the database (Rebuild) as long as the volume still exists.
The same actions are available through the interface’s API (administrators):
GET, PUT, DELETE /api/settings/memory/cleanup, GET /api/settings/memory/cleanups,
POST /api/settings/memory/cleanups/{id}/retry, POST /api/settings/memory/orphans/scan (dry run) and
POST /api/settings/memory/orphans/clean (body {"uris": [...]}, from the dry run).
POST /api/settings/memory/cleanups/{id}/retire retires a pending cleanup without deleting anything.
GET and PUT /api/settings/memory/orphans/schedule read or change the daily dry scan ({"enabled": true}).
The original /cleanups list shape is retained; the global blocked counts are in /cleanup.
Retention
RETENTION_OUTBOX_SENT_DAYS (7) deletes queue rows that were already written. This affects neither
retrieval, which is checked against the database, nor a rebuild, which recreates the rows it needs.
See data retention.
The cleanup log is small (one row per deleted volume or series, or per orphan cleanup) and is kept.